Sales data sparsity compensation and adaptive visualization method and system

By constructing a sales data knowledge graph and combining it with bidirectional time series processing, graph attention, and Monte Carlo tree search, the sparsity problem of sales data is solved, the accuracy and visualization of data analysis are improved, and it can adaptively compensate and display multi-dimensional relationships.

CN120689091APending Publication Date: 2025-09-23GUANGZHOU SAIER CULTURE TECHNOLOGY CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510736178.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Existing technologies are unable to effectively solve the sparsity problem of sales data, which affects the accuracy and reliability of data analysis results. At the same time, visualization methods are unable to intuitively display multidimensional relationships and dynamic change trends.

Method used

By constructing a sales data knowledge graph and combining it with a bidirectional time series processing mechanism, a graph attention mechanism, and a Monte Carlo tree search, we can achieve precise compensation for the sparsity of sales data and intuitive visualization.

Benefits of technology

It improves the accuracy of sales data analysis and decision-making support capabilities, can effectively capture spatiotemporal correlation characteristics, adaptively focus on important node relationships, discover hidden data correlation patterns, and intuitively display multidimensional relationships and dynamic change trends.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120689091A_ABST
    Figure CN120689091A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data analysis and visualization, in particular to a sales data sparsity compensation and self-adaptive visualization method and a sales data sparsity compensation and self-adaptive visualization system, a knowledge graph containing a sales node time sequence relation is constructed through spatio-temporal clustering and attribute division, a bidirectional time sequence processing mechanism is adopted, historical evolution and future trend are considered at the same time, and the sales data sparsity compensation and self-adaptive visualization method is established. Historical and future relation strength among nodes is obtained, sales association probability among the nodes is predicted, comprehensive mastering of node relations is realized, prediction accuracy is improved, a missing value regression model of a multi-dimensional relation is constructed on the basis of a graph attention mechanism, attributes and states needing to be filled are identified, and the probability of sales association between the nodes is predicted. Potential relation reasoning is carried out through Monte Carlo tree search, an adaptive node attribute filling result is obtained, a visual layout strategy is designed according to node weights and attribute values, a sales flow change trend is displayed through a thermodynamic model, space-time correlation characteristics of sales data are effectively captured, and a solid foundation is provided for sparsity compensation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data analysis and visualization technology, and specifically to a sales data sparsity compensation and adaptive visualization method and system, which is suitable for scenarios such as enterprise sales data analysis, market trend forecasting, and sales decision support. Background Art

[0002] As enterprises advance their digital transformation, sales data analysis plays an increasingly important role in business decision-making. However, actual sales data often suffers from sparsity, manifested primarily in incomplete data, uneven distribution, and time series breakpoints. This severely impacts the accuracy and reliability of data analysis results.

[0003] Existing methods for addressing sales data sparsity primarily include simple interpolation, mean filling, and regression prediction. While simple interpolation is easy to use, it fails to capture complex relationships between data; mean filling ignores the temporal and spatial correlations of data; and while regression prediction takes some data correlations into account, it struggles to model complex, multidimensional relationships. Furthermore, existing visualization methods often struggle to intuitively display the multidimensional relationships and dynamic trends of sales data, limiting the depth and breadth of data analysis.

[0004] Therefore, there is an urgent need for a technical solution that can effectively solve the sparsity problem of sales data and provide intuitive visual expression to improve the accuracy and intuitiveness of sales data analysis. Summary of the Invention

[0005] The purpose of the present invention is to provide a sales data sparsity compensation and adaptive visualization method and system. By constructing a sales data knowledge graph and combining advanced technologies such as bidirectional time series processing mechanism, attention mechanism and Monte Carlo tree search, accurate compensation and intuitive visualization of sales data sparsity can be achieved, thereby improving the accuracy of sales data analysis and decision support capabilities.

[0006] The present invention proposes a sales data sparsity compensation and adaptive visualization method, including:

[0007] Constructing a sales data knowledge graph through spatiotemporal clustering and attribute partitioning, wherein the knowledge graph includes the temporal relationship of sales nodes;

[0008] Based on the bidirectional time series processing mechanism, the historical and future relationship strengths between nodes are obtained to predict the probability of sales association between nodes;

[0009] Based on the graph attention mechanism, a missing value regression model for multi-dimensional relationships is constructed to identify the attributes and states that need to be filled;

[0010] Monte Carlo tree search is used to infer potential relationships and obtain adaptive node attribute filling results;

[0011] Design a visual layout strategy based on node weights and attribute values, and demonstrate sales flow trends through thermodynamic models.

[0012] Preferably, the sales data knowledge graph is constructed by spatiotemporal clustering and attribute partitioning, and the knowledge graph contains the temporal relationship of sales nodes, specifically including:

[0013] Acquire sales behavior data, including channel data, salesperson data, customer data, and their time series attributes;

[0014] Preprocessing the sales behavior data, including preliminary filling of missing values ​​and data standardization;

[0015] Based on the time series attributes as the distance variable of similarity measurement, the sales behavior data is clustered by spatiotemporal proximity through KNN analysis;

[0016] Calculate geographical distance, time difference and attribute difference between sales nodes;

[0017] The transaction relationship between channels and customers is calculated using the consistency of category attributes, and sales records are classified and aggregated by category and sales volume to obtain sales node data under the dimensions of channel, time, and type.

[0018] Preferably, the method of obtaining the historical and future relationship strengths between nodes and predicting the probability of sales association relationships between nodes based on a bidirectional time series processing mechanism specifically includes:

[0019] Divide channel data into dynamic attributes and static attributes, the dynamic attributes include sales volume, region and brand, and the static attributes include model and ownership;

[0020] Based on the channel-customer dimension, node data is imported into the graph database to build a time series relationship graph of sales nodes;

[0021] A bidirectional processing mechanism is used to encode the temporal relationship status of sales nodes. The historical evolution trend of the nodes is captured through forward processing, and the key factors leading to the current status are analyzed through backward processing.

[0022] Based on the node relationship change rate, the correlation between the future node status and the future node sales behavior is predicted, and the probability of the existence of a sales correlation relationship between the predicted nodes is obtained.

[0023] Preferably, the bidirectional processing mechanism is used to encode the time sequence relationship state of the sales nodes, specifically including:

[0024] Expand the relationships of node data using channels and customers as primary keys, and transform data containing time series relationships into a node-oriented time series relationship graph;

[0025] Extract channel sales, customer sales, channel product share, and customer share by date;

[0026] Construct a bidirectional network structure and encode the temporal relationship state of sales nodes through the forward network and backward network respectively;

[0027] The temporal relationship parameters of the temporal relationship graph are grouped by channel and customer to obtain a relationship matrix containing time steps;

[0028] Design a business logic interface, import the relationship matrix and historical data into the data set according to the business logic interface, and complete the input of business logic.

[0029] Preferably, the missing value regression model of multi-dimensional relationships is constructed based on the graph attention mechanism to identify the attributes and states that need to be filled, specifically including:

[0030] Generate an initial feature vector for each node in the graph, including node type, attributes, historical behavior and other information;

[0031] Calculate the attention weight coefficient according to different types of node relationships to achieve differentiated processing of the importance of nodes;

[0032] Based on the degree of association between nodes, the attention weight is dynamically adjusted to enable the system to focus on more important node relationships;

[0033] Aggregate neighborhood node information, update the current node feature representation, and achieve effective information transmission in the graph structure;

[0034] By stacking multiple attention layers, high-order features are extracted layer by layer to capture complex relationship patterns between nodes.

[0035] Preferably, the method of performing potential relationship reasoning through Monte Carlo tree search to obtain an adaptive node attribute filling result specifically includes:

[0036] Select the node data with the current channel as the primary key as the analysis object, and use the node state code to represent the variable value of the graph relationship matrix;

[0037] Analyze the correlation between variable values ​​and the current channel temporal state, and select attribute dimensions under the channel and customer dimensions as missing state variables;

[0038] Create a Monte Carlo tree structure containing several state sets, each state set contains multiple nodes, and the nodes contain two attributes: predicted value and confidence;

[0039] Based on the node confidence ranking results, the state set with the highest prediction value is selected for in-depth exploration;

[0040] Determine whether the maximum predicted value in the state set exceeds a preset threshold and calculate the multi-dimensional correlation strength between the state set and the target node;

[0041] The optimal exploration path is selected based on the association strength, and iteration is continued until the optimal path that meets the conditions is found, and the adaptive node attribute filling result is obtained.

[0042] Preferably, the state set with the highest prediction value is selected for in-depth exploration based on the node confidence ranking result, specifically including:

[0043] Implement differentiated filling strategies for different attribute types, using regression models for numerical attributes and classification models for categorical attributes;

[0044] Perform historical consistency verification, relationship consistency verification, and business rule verification, forming a three-level verification mechanism;

[0045] Quantitatively evaluate the reliability of the filling results and generate a confidence score;

[0046] Based on the difference between actual sales data and filling results, the model parameters are continuously adjusted to form a closed-loop optimization mechanism;

[0047] Different temporal relationship state encoding methods are used for storage for different nodes, supporting a bidirectional storage mechanism of forward filling and backward filling.

[0048] Preferably, the visual layout strategy is designed based on node weights and attribute values, and the sales flow change trend is displayed through a thermodynamic model, specifically including:

[0049] Design node distribution algorithms to calculate graph data layout and structure;

[0050] Set node size, node coordinates, and node center position in the layout algorithm based on node weights and attribute values;

[0051] Generate a visual multi-dimensional relationship map of sales data to support user interactive exploration;

[0052] Design a thermodynamic diffusion model and evolve it based on the dynamic trend of the network;

[0053] Through the evolution results of the thermodynamic diffusion model, a thermodynamic diffusion trend diagram is obtained, which shows the change direction and intensity of sales attributes of different nodes.

[0054] Preferably, the design of the thermodynamic diffusion model and the evolution of the diffusion model in combination with the dynamic change trend of the network specifically include:

[0055] Set thermodynamic gradients for different time periods and calculate thermodynamic layout states;

[0056] Build a time series chart display model and adjust the matching relationship between the status of different time periods and the node layout coordinates;

[0057] Realize the display of time series changes of each node in the sales process;

[0058] Adjacent nodes are aggregated together iteratively until the vectors under all relationship types converge;

[0059] The similarity between nodes is calculated, and the layout is automatically optimized according to the similarity between the nodes, so as to realize the dynamic update of the thermodynamic diffusion trend diagram.

[0060] Sales data sparsity compensation and adaptive visualization system, including:

[0061] A basic data module is used to construct a sales data knowledge graph through spatiotemporal clustering and attribute partitioning, wherein the knowledge graph includes the temporal relationship between sales nodes;

[0062] The graph learning module is used to obtain the historical and future relationship strengths between nodes based on a bidirectional time series processing mechanism and predict the probability of sales association relationships between nodes;

[0063] The filling prediction module is used to build a missing value regression model for multi-dimensional relationships based on the graph attention mechanism, and perform potential relationship reasoning through Monte Carlo tree search to obtain adaptive node attribute filling results;

[0064] Visual modeling module, used to design visual layout strategies based on node weights and attribute values, and to demonstrate sales flow trends through thermodynamic models;

[0065] Among them, the basic data module includes a data preprocessing component and a data extraction component; the graph learning module includes a graph relationship learner, a temporal association relationship extractor and an attribute association relationship extractor; the filling prediction module includes a filling model learning component, a filling attribute learning component and a Monte Carlo tree search component; the visual modeling module includes a temporal change visualization component and a stream data visualization component.

[0066] The beneficial effects of the present invention include:

[0067] 1. By constructing a sales data knowledge graph containing temporal relationships, we can effectively capture the spatiotemporal correlation characteristics of sales data, providing a solid foundation for sparsity compensation.

[0068] 2. Adopting a bidirectional time series processing mechanism, taking into account both historical evolution and future trends, it can fully grasp the node relationships and improve prediction accuracy;

[0069] 3. Multi-dimensional relationship modeling based on the graph attention mechanism can adaptively focus on important node relationships and improve the ability to extract key information;

[0070] 4. Use Monte Carlo tree search to reason about potential relationships, effectively discovering hidden data association patterns and achieving more accurate sparsity compensation;

[0071] 5. Visual design based on node weights and thermodynamic models intuitively displays the multi-dimensional relationships and dynamic trends of sales data, enhancing data interpretation capabilities;

[0072] 6. Build a complete closed-loop optimization system to continuously improve model performance through multi-level verification and feedback mechanisms to ensure the reliability of compensation results. BRIEF DESCRIPTION OF THE DRAWINGS

[0073] Figure 1 This is a flowchart of the sales data sparsity compensation and adaptive visualization method of the present invention;

[0074] Figure 2 A schematic diagram of the structure of the module for building the sales data knowledge graph;

[0075] Figure 3 This is a working principle diagram of the bidirectional timing processing mechanism;

[0076] Figure 4 This is the network structure diagram of the graph attention mechanism;

[0077] Figure 5 Schematic diagram of the decision process of Monte Carlo tree search;

[0078] Figure 6 A flow chart for the process of filling in adaptive node attributes;

[0079] Figure 7 Schematic diagram for visualizing the working of the layout strategy and thermodynamic model;

[0080] Figure 8 This is a module composition framework diagram of the system of the present invention. DETAILED DESCRIPTION

[0081] Please refer to the attached Figure 1-8 The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood by those skilled in the art that these embodiments are only used to illustrate the present invention and are not intended to limit the present invention.

[0082] Reference Figure 1The sales data sparsity compensation and adaptive visualization method provided by the present invention includes the following steps: constructing a sales data knowledge graph through spatiotemporal clustering and attribute partitioning; obtaining the historical and future relationship strengths between nodes based on a bidirectional time series processing mechanism; constructing a missing value regression model for multidimensional relationships based on a graph attention mechanism; performing potential relationship reasoning through Monte Carlo tree search; and designing a visualization layout strategy based on node weights and attribute values.

[0083] Reference Figure 2 The process of constructing a sales data knowledge graph through spatiotemporal clustering and attribute partitioning is as follows:

[0084] First, obtain sales behavior data, including channel data, salesperson data, customer data, and their time series attributes. Preferably, the sales behavior data spans three months to one year to ensure temporal continuity and integrity. In practical applications, for example, a large electrical appliance chain contains sales records for 326 stores, 5,600 sales personnel, and nearly one million customers nationwide, spanning six months.

[0085] Next, the sales data is preprocessed, including initial filling of missing values ​​and data normalization. Small missing values ​​(missing rates below 5%) are filled using the mean or median. Numerical features are normalized using the min-max method, converting them to the [0, 1] range. Categorical features are converted to numerical representations using one-hot encoding. For example, missing daily sales figures in a salesperson's performance data can be initially filled using the salesperson's average daily sales figures over the past 30 days.

[0086] Then, based on the time series attribute as the distance variable of the similarity measure, the sales behavior data is clustered in time and space by KNN analysis. In this embodiment, the K value is preferably set to between 5 and 10 and dynamically adjusted according to the data density. Specifically, when the data is relatively dense (such as the sales data of stores in first-tier cities), the K value can be set to a larger value (such as 8-10) to reduce the impact of noise; when the data is relatively sparse (such as the sales data of stores in third- and fourth-tier cities), the K value can be set to a smaller value (such as 5-6) to avoid the introduction of interference from irrelevant data.

[0087] During the clustering process, the geographical distance, time difference, and attribute difference between sales nodes are calculated. The geographical distance is calculated using the Euclidean distance; the time difference is calculated using the weighted time decay function, whose mathematical expression is:

[0088]

[0089] Where W(t) is the weight value at time point t; t represents a historical time point in days; t0 represents the current time point in days; λ is the time decay coefficient, dimensionless, with a preferred value range of 0.01-0.1, which is adjusted according to the timeliness requirements of the business scenario. For example, for fast-moving consumer goods sales data, λ can be set to 0.05-0.1 to emphasize the importance of recent data; for durable consumer goods sales data, λ can be set to 0.01-0.03 to balance the impact of long-term and short-term data. In the mobile phone sales scenario, if λ is set to 0.03, the sales data from 30 days ago will weigh approximately 40% of the current data, and the sales data from 60 days ago will weigh approximately 16% of the current data, which is consistent with the seasonal characteristics of mobile phone sales.

[0090] Attribute differences are calculated using weighted cosine similarity, and its mathematical expression is:

[0091]

[0092] Among them, Sim(i,j) is the similarity between node i and node j, dimensionless, and the value range is [-1,1]; a i ,k and a j,k Represents the k-th attribute value of node i and node j respectively; w k The dimensionless weight of the kth attribute, where all weights sum to 1; n represents the total number of attributes. Attribute weights are set based on business importance. For example, sales weight is set to 0.4, region weight is set to 0.3, brand weight is set to 0.2, and other attribute weights are set to 0.1. In a real-world application, an electronics retailer clustered consumer purchasing behavior and found that customer groups with a similarity greater than 0.8 had similar response rates to promotions, which could serve as a foundation for subsequent precision marketing.

[0093] Finally, the transaction relationship between channels and customers is calculated using the consistency of category attributes. Sales records are classified and aggregated by category and sales volume to obtain sales node data under the dimensions of channel, time, and type. The aggregated sales node data contains information such as node ID, node type, node attribute set, and node relationship set, forming the basic data structure of the sales data knowledge graph. For example, the sales records of high-end mobile phones in Beijing are aggregated by channel (online flagship stores, offline specialty stores, large electrical appliance stores) and customer type (individual consumers, corporate customers) to obtain sales node data for different channels and customer types, providing a basis for subsequent analysis.

[0094] Reference Figure 3 The present invention is based on a bidirectional time series processing mechanism. The process of obtaining the historical and future relationship strengths between nodes is as follows:

[0095] First, channel data is divided into dynamic and static attributes. Dynamic attributes include sales volume, region, and brand, while static attributes include model number and ownership. Dynamic attributes change over time and require time series processing; static attributes are relatively stable and can serve as basic node characteristics. For example, in the channel data of a home appliance retailer, dynamic attributes include monthly store sales and the proportion of specific brands, while static attributes include store area and the business district it belongs to.

[0096] Next, based on the channel-customer dimension, the node data is imported into a graph database to construct a time-series relationship graph of sales nodes. In this embodiment, the Neo4j graph database is preferably used for storage, as it has high query efficiency and supports the expression of complex relationship patterns. For example, the sales records of a home appliance chain's 10 core brands, 326 stores, and nearly one million customers are imported into a Neo4j database, forming a time-series relationship graph of sales nodes containing approximately 2 million nodes and 30 million relationships.

[0097] Then, a bidirectional processing mechanism is employed to encode the temporal relationship states of sales nodes. Specifically, forward processing captures the historical evolution trends of nodes, while backward processing analyzes the key factors leading to the current state. The core of this bidirectional processing mechanism lies in simultaneously considering both forward temporal evolution and backward causal relationships, thereby fully grasping the dynamic characteristics of node relationships. In sales data analysis, forward processing can identify sales growth trends, while backward processing can uncover key drivers influencing sales. The combination of the two provides a more comprehensive analysis of sales dynamics.

[0098] In the specific implementation, node data is expanded using channels and customers as primary keys, converting data containing time-series relationships into a node-oriented time-series relationship graph. Preferably, channel sales, customer sales, channel product share, and customer share are extracted by date to form a time-series feature sequence for the node. For example, for an e-commerce platform, time-series features such as daily sales of different product types, brand share, and user repurchase rate are extracted to construct a time-series relationship graph covering the past six months.

[0099] A bidirectional network structure is constructed, and the temporal relationship state of the sales nodes is encoded through the forward network and the backward network respectively. The state update formula of the forward network is:

[0100]

[0101] in, Represents the hidden state of the forward network at time point t, which is a d-dimensional vector, where d is the hidden layer dimension, preferably set to 64-128; represents the hidden state of the forward network at time point t-1, which is also a d-dimensional vector; x t Represents the input feature at time point t, which is an m-dimensional vector, where m is the feature dimension; Wf represents the weight matrix of the forward network, with a dimension of d×(d+m); b f represents the bias vector of the forward network, with dimension d; f represents the activation function, preferably the ReLU function; [·,·] represents the vector concatenation operation.

[0102] The state update formula of the backward network is:

[0103]

[0104] in, Represents the hidden state of the backward network at time point t, which is a d-dimensional vector; represents the hidden state of the backward network at time point t+1, which is also a d-dimensional vector; x t Represents the input feature at time point t, which is an m-dimensional vector; W b Represents the weight matrix of the backward network, with a dimension of d×(d+m); b b represents the bias vector of the backward network, with a dimension of d; f represents the activation function, and the ReLU function is also preferably used.

[0105] The forward and backward state fusion formula is:

[0106]

[0107] Among them, h t Represents the node state after fusion, which is a d-dimensional vector; W h represents the fusion weight matrix, with a dimension of d×2d; b h Represents the fusion bias vector with dimension d. For sales data, state fusion takes into account historical trends and future predictions, and can more comprehensively express the state characteristics of nodes at different time points.

[0108] The temporal relationship parameters of the temporal relationship graph are grouped by channel and customer to obtain the relationship matrix R containing the time steps. t,c ∈R n×m×d , where n represents the number of channels, m represents the number of customers, and d represents the dimension of the relationship characteristics. In practical applications, for example, a retail company analyzes the sales relationship between 20 major cities, 50 core channels and 100,000 customers. The dimension of the relationship matrix is ​​about 20×50×10 5 ×64, using sparse matrix storage to improve computational efficiency.

[0109] Design a business logic interface and import the relationship matrix and historical data into the dataset based on the business logic interface to complete the business logic input. The business logic interface defines the data organization and interface specifications to ensure the consistency and integrity of data flow between different modules. In the retail business scenario, the business logic interface may include functional modules such as sales forecasting, inventory optimization, and customer segmentation, each of which corresponds to different data requirements and processing logic.

[0110] Through this bidirectional time series processing mechanism, the present invention can comprehensively capture the temporal evolution characteristics of sales nodes, providing rich temporal relationship features for subsequent missing value regression models. For example, in seasonal product sales analysis, bidirectional processing can simultaneously consider historical seasonal patterns and recent sales trends, improving the accuracy of sales forecasts.

[0111] Reference Figure 4 The present invention uses the graph attention mechanism as the basis to construct a missing value regression model for multi-dimensional relationships. The specific process is as follows:

[0112] First, an initial feature vector is generated for each node in the graph, containing information such as node type, attributes, and historical behavior. In this embodiment, the dimension of the node feature vector is preferably set between 64 and 128 to balance expressiveness and computational complexity. For example, for a store node in a retail scenario, its feature vector may contain information such as store type, area, location, and historical sales performance.

[0113] The initial eigenvector is generated as follows:

[0114] X i =[E type ,E attr ,E hist ],

[0115] Among them, X i represents the initial feature vector of node i, with dimension d; E type Embedding vector representing the node type (dimension is 16-32), such as store type, customer type, etc.; E attr Embedding vector representing node attributes (dimension is 32-64), such as store area, customer spending power, etc.; E hist Embedding vectors (16-32 dimensions) representing a node's historical behavior, such as sales history and purchase frequency. [·,·,·] represents a vector concatenation operation. In an actual application on an e-commerce platform, a user's demographic characteristics, historical purchase behavior, and browsing history are combined into a 128-dimensional feature vector for subsequent purchase intention prediction.

[0116] Next, the attention weight coefficient is calculated according to different types of node relationships to achieve differentiated processing of the importance of nodes. The core of the graph attention mechanism is to calculate the attention coefficient between nodes, and its calculation formula is:

[0117] e ij =a(W·X i ,W·X j ),

[0118] Among them, e ij It represents the attention coefficient of node i to node j, which is a scalar; W represents the linear transformation matrix with dimension d ′ ×d,d ′ is the feature dimension after transformation, preferably set to 0.5-1 times of d; a represents the attention function, which is usually implemented using a single-layer neural network; X i and X j They represent the feature vectors of node i and node j respectively, and their dimensions are both d.

[0119] The mathematical expression of the attention function is:

[0120] a(W·X i ,W·X j )=LeakyReLU(A T ·[W·X i ||W·X j ]),

[0121] Among them, A represents the parameter vector of the attention network, with a dimension of 2d ′ ; || represents the vector concatenation operation; LeakyReLU represents the leaky ReLU activation function, and its mathematical expression is:

[0122]

[0123] Where x is the input value, which can be a scalar or vector; α is the dimensionless leakage coefficient, preferably set to 0.2. Compared with the traditional ReLU function, the LeakyReLU function still has a smaller gradient when the input is negative, which helps alleviate the vanishing gradient problem.

[0124] Then, based on the degree of association between nodes, the attention weight is dynamically adjusted so that the system can focus on more important node relationships. The normalized formula for attention weight is:

[0125]

[0126] Among them, α ij represents the normalized attention weight, dimensionless, ranging from [0,1]; N irepresents the set of neighboring nodes of node i; exp represents the natural exponential function. Normalization ensures that the sum of all attention weights is 1, facilitating subsequent weighted aggregation operations. In sales network analysis, attention mechanisms can automatically identify important sales relationships. For example, a large retailer discovered through attention mechanisms that the relationship weight between high-end consumers and brand flagship stores is significantly higher than that between other channels. This discovery guided subsequent precision marketing strategies.

[0127] Aggregate neighboring node information, update the current node feature representation, and achieve effective information transmission in the graph structure. The node feature update formula is:

[0128]

[0129] Among them, X′ i represents the updated node feature vector, with a dimension of d′; σ represents the activation function, preferably the ReLU function; N i represents the set of neighbor nodes of node i; α ij represents the attention weight of node i to node j; W represents the linear transformation matrix with dimension d′×d; X j Represents the feature vector of node j, with dimension d.

[0130] Finally, by stacking multiple attention layers, high-level features are extracted layer by layer, capturing complex patterns of node relationships. In this embodiment, it is preferred to stack 2-3 attention layers, adding skip connections (residual connections) between each attention layer to alleviate the gradient vanishing problem that deep networks may face. The final node representation is a concatenation or weighted average of the outputs of each layer.

[0131] The overall architecture of the multi-layer attention network can be expressed as:

[0132] X (l+1) =σ(A (l) ·X (l) W (l) )+X (l) ,

[0133] Among them, X (l) Represents the node feature matrix of the lth layer, with a dimension of n×d (l) ,n is the number of nodes, d (l) is the feature dimension of the lth layer; A (l) W represents the attention coefficient matrix of the lth layer, with a dimension of n×m; (l) Represents the weight matrix of layer l, with dimension d (l) ×d (l+1) ; σ represents the activation function, preferably the ReLU function; the plus sign represents the residual connection, which directly adds the input features to the output features.

[0134] In practical applications, a large supermarket chain used a three-layer graph attention network to analyze sales relationships across 367 stores and tens of thousands of products nationwide. The network successfully identified patterns in over 90% of missing sales data, providing a reliable basis for subsequent data filling. Especially for seasonal products and new product launches, the graph attention mechanism can accurately predict missing sales data based on the sales patterns of similar products and similar stores.

[0135] Through this graph attention mechanism, the present invention can effectively model the multidimensional and complex relationships in sales data, providing powerful feature representation capabilities for missing value regression. Compared to traditional missing value filling methods, the graph attention mechanism considers the complex interactions between nodes and can more accurately capture the inherent patterns in sales data.

[0136] Reference Figure 5 The process of performing potential relationship reasoning through Monte Carlo tree search in the present invention is as follows:

[0137] First, we select node data with the current channel as the primary key as the analysis object, and use node state encoding to represent the variable values ​​of the graph relationship matrix. Node state encoding uses the node representation obtained by the bidirectional time series processing mechanism and graph attention mechanism described above. For example, for the 3C digital category of an e-commerce platform, we can select each sales channel of that category (official flagship stores, specialty stores, authorized dealers, etc.) as the analysis object to study the sales correlation between different channels.

[0138] Analyze the correlation between variable values ​​and the current channel time series status, and select attribute dimensions under the channel and customer dimensions as missing state variables. Correlation analysis is calculated using the Pearson Correlations Coefficient. Attribute dimensions with an absolute value of a correlation coefficient greater than 0.3 are prioritized as missing state variables. The formula for calculating the Pearson Correlations Coefficient is:

[0139]

[0140] Among them, r is the Pearson correlation coefficient, dimensionless, and the value range is [-1,1]; X i and Y i Represents the i-th observation value of the two variables respectively; and where represents the mean of the two variables, and n represents the number of observations. In practice, one retail company found that the correlation coefficient between store area and sales of high-end appliances was 0.68, while the correlation coefficient with sales of low-end appliances was only 0.12. This finding guided store layout and merchandise display strategies.

[0141] Create a Monte Carlo tree structure consisting of several state sets. Each state set contains multiple nodes, each of which has two attributes: a predicted value and a confidence level. The predicted value represents the prediction of a missing value, and the confidence level indicates the degree of confidence in the accuracy of the predicted value. The initial Monte Carlo tree structure consists of a root node and several child nodes, preferably 5-10, to balance exploration breadth and computational complexity. In a sales forecasting scenario, the root node can represent the current sales status, while the child nodes represent possible outcomes under different sales strategies or market conditions.

[0142] Based on the node confidence ranking results, the state set with the highest prediction value is selected for in-depth exploration. The confidence calculation formula is:

[0143]

[0144] Where C(s) represents the confidence of state s, dimensionless; Q(s) represents the cumulative reward value of state s; N(s) represents the number of times state s is visited; parent(s) represents the parent node of state s; ln represents the natural logarithm function; c represents the exploration coefficient, dimensionless, and is preferably set to 1.414 (i.e. ). The first term of this formula stands for exploitation, which is choosing the best action based on known information; the second represents exploration, i.e., trying actions that have not been fully explored. The coefficient c controls the balance between exploration and exploitation, with larger c values ​​encouraging more exploration.

[0145] Determine whether the maximum predicted value in the state set exceeds the preset threshold and calculate the multi-dimensional correlation strength between the state set and the target node. The preset threshold is set according to business needs. For example, for sales forecast, the threshold can be set to 80% of the historical average sales; for category forecast, the threshold can be set to a confidence level of 0.7. The formula for calculating multi-dimensional correlation strength is:

[0146] Strength(s,t)=w t Sim time (s,t)+w s Sim space (s,t)+w a Sim attr (s,t),

[0147] Among them, Strength(s,t) represents the multidimensional correlation strength between state s and target node t, dimensionless, and the value range is usually [0,1]; Sim time 、Sim space and Sim attrRespectively represent the similarity of time dimension, space dimension and attribute dimension, dimensionless, and the value range is [0,1]; w t 、w s and w a The weights of the three dimensions are dimensionless and sum to 1. The preferred settings are 0.3, 0.3, and 0.4. In retail scenarios, the time dimension considers the seasonality and cyclicality of sales, the spatial dimension considers geographic location and store layout, and the attribute dimension considers product characteristics and customer preferences.

[0148] The optimal exploration path is selected based on the association strength, and iterations are continued until the optimal path that meets the conditions is found, resulting in the adaptive node attribute filling result. The iterative process consists of four stages: selection, expansion, simulation, and backpropagation.

[0149] In the selection phase, the most valuable nodes are selected for exploration based on the above confidence formula. In the expansion phase, child nodes are added to the selected nodes to expand the search space. In the simulation phase, starting from the expanded nodes, a random strategy is used to simulate until the termination condition is reached (such as reaching the maximum depth or finding the target node). In the backtracking phase, the value and visit count of the node are updated according to the simulation results. The update formula is:

[0150] Q(s)=Q(s)+V(result),

[0151] N(s)=N(s)+1,

[0152] Where Q(s) represents the cumulative reward value of state s; V(result) represents the value assessment of the simulation result, calculated based on the association strength and prediction accuracy; N(s) represents the number of times state s is visited. The value assessment function V(result) is usually designed as a weighted sum of the association strength and prediction accuracy, that is:

[0153] V(result)=β·Strength(s,t)+(1-β)·Accuracy(s),

[0154] Where β is the weight coefficient, dimensionless, ranging from [0, 1], and preferably set to 0.6; Accuracy(s) represents the prediction accuracy of state s, dimensionless, ranging from [0, 1].

[0155] To fill missing values ​​in sales data, a differentiation strategy is employed: for numerical attributes (such as sales volume), a regression model is used to predict the specific value; for categorical attributes (such as brand and model), a classification model is used to predict the category label. The differentiation strategy is selected automatically based on the attribute type, without manual intervention. In practical application, a supermarket chain used this method to successfully predict the product mix and sales levels of newly opened stores, with a prediction accuracy of over 85%, providing strong support for new store operations.

[0156] Through the Monte Carlo tree search mechanism described above, the present invention can effectively discover potential relationships in sales data, providing a richer information foundation for sparsity compensation. Compared to traditional linear regression or simple interpolation methods, Monte Carlo tree search can explore more complex nonlinear relationships and is particularly suitable for handling the sparsity and irregularity in sales data.

[0157] Reference Figure 7 The present invention designs a visualization layout strategy based on node weights and attribute values, and uses a thermodynamic model to display the sales flow change trend. The specific process is as follows:

[0158] First, we designed a node distribution algorithm to calculate the graph data layout and structure. The node distribution algorithm uses the principle of force-directed layout to achieve a balanced graph layout by simulating a physical mechanical system. The mechanical model consists of two parts: attractive and repulsive forces, and its mathematical expression is:

[0159] F attract (d) = d 2 / k,

[0160] F repulse (d) = k 2 / d,

[0161] Among them, F attract (d) represents the attraction between two nodes with a distance d, and the unit is the same as d 2 / k consistent; F repulse (d) represents the repulsive force between two nodes with a distance d, and its unit is the same as k 2 / d is consistent; d represents the distance between nodes, which can be expressed in pixels; k represents the ideal side length, which is the same as d and is preferably set to 10% of the diagonal length of the graph layout area. The attractive force increases with distance, while the repulsive force decreases with distance. A stable layout is achieved when the two reach equilibrium.

[0162] The node size, node coordinates, and node center position are set in the layout algorithm based on the node weight and attribute values. The node size is proportional to the node weight, and the calculation formula is:

[0163]

[0164] Among them, Size(i) represents the size of node i, and the unit can be pixels; S min and S max Represent the minimum and maximum sizes of nodes (preferably set to 10 and 50 pixels); W(i) represents the weight of node i, dimensionless; W min and W max The minimum and maximum values ​​of the node weight are dimensionless. In sales network visualization, node size can reflect the sales volume and facilitate intuitive identification of important nodes.

[0165] The node coordinates are iteratively calculated using the force-directed algorithm. The iterative formula is:

[0166] Pos(i) t+1 =Pos(i) t +Δt·F(i) t ,

[0167] Among them, Pos(i) t represents the position coordinates of node i at the tth iteration, which is a two-dimensional or three-dimensional vector; Δt represents the time step, dimensionless, and is preferably set to 0.1; F(i) t represents the net force acting on node i, with the same units as Pos(i), and is the vector sum of all attractive and repulsive forces. Iterations typically continue until the system energy drops below a preset threshold, or until a maximum number of iterations (e.g., 100) is reached.

[0168] Generates a visual multidimensional relationship map of sales data, supporting interactive exploration. Interactive features include node selection, relationship filtering, time range adjustment, and attribute display switching, improving user experience and analysis efficiency. For example, in a retail sales network visualization, users can select a specific time period (such as a promotional period or holiday) to view changes in sales relationships or filter sales trends for specific product categories.

[0169] Design a thermodynamic diffusion model and evolve it based on the dynamic trend of the network. The thermodynamic diffusion model is based on the heat conduction equation, and its mathematical expression is:

[0170]

[0171] in, It represents the rate of change of temperature T at position x with time t, with the unit of temperature / time; α represents the thermal diffusivity, with the unit of area / time, which describes the speed of heat propagation; represents the Laplace operator, which is a spatial second-order derivative operator; T(x,t) represents the temperature at position x at time t, with the unit being temperature. In terms of graph structure, the discretized heat diffusion process can be expressed as:

[0172]

[0173] in, represents the temperature of node i at time step t, which can be expressed as a dimensionless value; A ij represents the adjacency relationship between nodes i and j, typically 0 or 1, or a weighted value (such as sales relationship strength); N(i) represents the set of neighbors of node i. In sales data visualization, temperature can represent sales activity, and thermal diffusion models can simulate the spread of sales activity, such as the diffusion of promotional effects from a core store to surrounding stores.

[0174] Set thermodynamic gradients for different time periods and calculate the thermodynamic layout state. Thermodynamic gradients are based on the rate of change in sales data. Areas with higher rates of change are assigned higher temperature gradients to highlight important trends. For example, during holiday sales, the temperature gradient can be increased to highlight sales changes; during regular sales periods, the temperature gradient can be smaller to reflect a stable sales pattern.

[0175] A time-series chart display model was constructed, adjusting the matching relationship between the states of different time periods and the node layout coordinates to display the time-series changes of each node during the sales process. This time-series display uses animation to smoothly transition and illustrate the changes in node states over time. In practical applications, a large retail group used time-series visualization to intuitively demonstrate the communication effects of different holiday promotion strategies. They found that electronic product promotions spread approximately 30% faster in first-tier cities than in third- and fourth-tier cities, guiding the development of differentiated promotion strategies.

[0176] Adjacent nodes are aggregated together in an iterative manner until the vectors under all relationship types converge. The aggregation process adopts a weighted average method, and the calculation formula is:

[0177]

[0178] in, represents the vector representation of node i at time step t, with dimension d; β represents the aggregation coefficient, dimensionless, ranging from [0, 1], and preferably set to 0.15; A ij represents the adjacency relationship between nodes i and j; N(i) represents the set of neighbors of node i. The aggregation operation makes the representation of related nodes closer, which facilitates the identification of community structure and relationship patterns.

[0179] Node similarity is calculated and the layout is automatically optimized based on this similarity, enabling dynamic updates of the thermodynamic diffusion trend graph. Node similarity is calculated using cosine similarity, and the optimized layout is achieved by minimizing the energy function. In sales network visualization, nodes with high similarity (such as stores with similar sales models) are automatically clustered, forming intuitive visual groupings that facilitate identification of patterns and trends.

[0180] Through this adaptive visualization layout strategy, the present invention can intuitively display the multidimensional relationships and dynamic trends of sales data, enhancing data interpretation and decision support. Compared to traditional static charts, the present visualization solution is more dynamic and interactive, capable of more comprehensively displaying the complex relationships and changes in sales data.

[0181] Reference Figure 8 The present invention provides a sales data sparsity compensation and adaptive visualization system comprising: a basic data module 10, a graph learning module 20, a filling prediction module 30 and a visualization modeling module 40.

[0182] The basic data module 10 is used to construct a sales data knowledge graph through spatiotemporal clustering and attribute partitioning. The knowledge graph includes the temporal relationship of sales nodes. The basic data module 10 includes a data preprocessing component 11 and a data extraction component 12.

[0183] Data preprocessing component 11 is responsible for cleaning, standardizing, and initially filling raw sales data. This includes operations such as outlier detection, missing value processing, and data format conversion. The goal of data preprocessing is to improve data quality and lay the foundation for subsequent analysis. In practice, data preprocessing component 11 can process approximately 5 million sales records daily, with an outlier detection accuracy rate exceeding 95%.

[0184] The data extraction component 12 imports sales behavior data into the data pool in batches, organized by customer and channel, and categorizes the data based on the type and attributes of each node. The data extraction component 12 supports incremental data updates, automatically identifying newly added data and incorporating it into the analysis process. For example, a retail company's data extraction component processes sales data from 326 stores nationwide daily and categorizes it by product category, customer group, and sales channel, providing structured data support for subsequent analysis.

[0185] The input of the basic data module 10 is raw sales behavior data, and the output is a preprocessed sales data knowledge graph, which provides basic data support for subsequent graph learning. The knowledge graph contains nodes (such as products, customers, channels) and relationships (such as purchase, recommendation, and supply), which can fully express the multidimensional relationships of sales data.

[0186] The graph learning module 20 is used to obtain the historical and future relationship strengths between nodes based on a bidirectional time series processing mechanism and predict the probability of sales association relationships between nodes. The graph learning module 20 includes a graph relationship learner 21, a time series relationship extractor 22, and an attribute relationship extractor 23.

[0187] Graph Relationship Learner 21 uses deep learning techniques to extract relationship matrices from temporal relationship graphs across different dimensions, capturing complex interaction patterns between nodes. Graph Relationship Learner 21 employs a bidirectional network structure that simultaneously considers both forward evolution and backward causal relationships. In an application on an e-commerce platform, the Graph Relationship Learner successfully predicted over 80% of users' next purchases by analyzing user-item-category interaction histories.

[0188] The temporal correlation extractor 22 extracts the current channel's time series data and time series relationship data from the relationship matrix and calculates the channel's temporal state encoding value. The temporal correlation extractor 22 can identify key nodes and turning points in the time series, providing important reference for sparsity compensation. For example, when analyzing holiday sales patterns, the temporal correlation extractor can automatically identify the warm-up period before the sales peak and the decay period after the peak, helping companies optimize promotional timing.

[0189] The attribute association extractor 23 extracts the attribute parameter values ​​and temporal relationship data of the current channel from the relationship matrix and calculates the attribute temporal state encoding value of the channel. The attribute association extractor 23 can discover potential associations between attributes and enrich the feature representation of nodes. In retail analysis, the attribute association extractor can discover association patterns between product attributes (such as price, brand, and packaging) and sales channels, guiding product positioning and channel selection.

[0190] The input of the graph learning module 20 is the sales data knowledge graph, and the output is the time series relationship graph, time series state coding value and attribute time series state coding value, which provides feature support for subsequent filling prediction.

[0191] The Fill Prediction Module 30 is used to construct a missing value regression model for multidimensional relationships based on the graph attention mechanism. It then uses Monte Carlo tree search to infer potential relationships and obtain adaptive node attribute filling results. The Fill Prediction Module 30 includes a Fill Model Learning Component 31, a Fill Attribute Learning Component 32, and a Monte Carlo Tree Search Component 33.

[0192] The Filling Model Learning Component 31 uses a forward algorithm to simulate the future states and trends of nodes in historical data, building a predictive model based on historical patterns. This component captures the cyclical, trending, and sudden characteristics of data, improving forecast accuracy. For a retail company forecasting seasonal product sales, the Filling Model Learning Component accurately predicted over 90% of peak sales periods by analyzing historical sales patterns, providing strong support for inventory management.

[0193] The Fill Attribute Learning Component 32 uses a backward algorithm to simulate attribute trends over time and uncover the underlying patterns of attribute change. This component employs a multi-task learning framework, enabling it to simultaneously handle prediction tasks for multiple related attributes, enhancing the model's generalization capabilities. For example, when analyzing customer purchasing behavior, the Fill Attribute Learning Component can simultaneously predict a customer's purchase frequency, average order value, and category preferences, creating a comprehensive customer profile.

[0194] The Monte Carlo Tree Search component 33 infers potential relationships and optimizes population results. Through extensive simulation and evaluation, the Monte Carlo Tree Search component 33 finds the optimal population strategy, balancing exploration and exploitation. In new product sales forecasting, the Monte Carlo Tree Search component can explore sales potential under different market conditions based on historical sales patterns of similar products, providing more comprehensive forecasts.

[0195] The input of the filling prediction module 30 is the time series relationship graph, the time series state encoding value and the attribute time series state encoding value, and the output is the adaptive node attribute filling result, which provides data support for subsequent visual modeling.

[0196] The visualization modeling module 40 is used to design a visualization layout strategy based on node weights and attribute values, and to display sales flow change trends through a thermodynamic model. The visualization modeling module 40 includes a time series change visualization component 41 and a flow data visualization component 42.

[0197] The Time Series Change Visualization Component 41 calculates layout coordinates based on the output of the time series relationship analysis engine, simulating the time series changes of nodes and enabling interactive time series analysis of the graph. It supports features such as time window sliding, playback speed adjustment, and key moment marking, enhancing time series analysis capabilities. In a retail company's sales analysis system, the Time Series Change Visualization Component intuitively displays sales trends for stores in different regions, helping managers identify rapidly growing and declining stores and adjust operational strategies in a timely manner.

[0198] The stream data visualization component 42 uses a visualization engine to adjust node state weights and calculate layout coordinates, achieving a layout display that aligns with temporal trends. It employs a thermodynamic diffusion model to intuitively display sales flow trends and diffusion paths. For example, when analyzing the effectiveness of a promotional campaign, the stream data visualization component can display the dynamic diffusion of promotional information from a core area to surrounding areas, allowing assessment of the promotion's scope and duration.

[0199] The input of the visual modeling module 40 is the adaptive node attribute filling result, and the output is a visual layout and a thermodynamic diffusion trend diagram, providing the user with an intuitive data analysis interface.

[0200] The modules of the system of the present invention form a closely coordinated workflow, realizing the full-process automated processing of sales data from collection, processing, analysis to visualization, which greatly improves the efficiency and accuracy of sales data analysis. In actual application of a large retail group, the system shortened the sales data analysis time from the original 2-3 days to a few hours, and the prediction accuracy was improved by 15 percentage points, which brought significant economic benefits to the enterprise. In summary, the sales data sparsity compensation and adaptive visualization method and system provided by the present invention, by constructing a sales data knowledge graph, combined with advanced technologies such as bidirectional time series processing mechanism, attention mechanism and Monte Carlo tree search, realizes accurate compensation and intuitive visualization of sales data sparsity, and provides strong support for enterprise sales decision-making.

Claims

1. Sales data sparsity compensation and adaptive visualization method, characterized by: include: Constructing a sales data knowledge graph through spatiotemporal clustering and attribute partitioning, wherein the knowledge graph includes the temporal relationship of sales nodes; Based on the bidirectional time series processing mechanism, the historical and future relationship strengths between nodes are obtained to predict the probability of sales association between nodes; Based on the graph attention mechanism, a missing value regression model for multi-dimensional relationships is constructed to identify the attributes and states that need to be filled; Monte Carlo tree search is used to infer potential relationships and obtain adaptive node attribute filling results; Design a visual layout strategy based on node weights and attribute values, and demonstrate sales flow trends through thermodynamic models.

2. The method according to claim 1, characterized in that The sales data knowledge graph is constructed through spatiotemporal clustering and attribute partitioning. The knowledge graph contains the temporal relationship of sales nodes, specifically including: Acquire sales behavior data, including channel data, salesperson data, customer data, and their time series attributes; Preprocessing the sales behavior data, including preliminary filling of missing values ​​and data standardization; Based on the time series attributes as the distance variable of similarity measurement, the sales behavior data is clustered by spatiotemporal proximity through KNN analysis; Calculate geographical distance, time difference and attribute difference between sales nodes; The transaction relationship between channels and customers is calculated using the consistency of category attributes, and sales records are classified and aggregated by category and sales volume to obtain sales node data under the dimensions of channel, time, and type.

3. The method according to claim 1, characterized in that The bidirectional time series processing mechanism is used to obtain the historical and future relationship strengths between nodes and predict the probability of sales association relationships between nodes, specifically including: Divide channel data into dynamic attributes and static attributes, the dynamic attributes include sales volume, region and brand, and the static attributes include model and ownership; Based on the channel-customer dimension, node data is imported into the graph database to build a time series relationship graph of sales nodes; A bidirectional processing mechanism is used to encode the temporal relationship status of sales nodes. The historical evolution trend of the nodes is captured through forward processing, and the key factors leading to the current status are analyzed through backward processing. Based on the node relationship change rate, the correlation between the future node status and the future node sales behavior is predicted, and the probability of the existence of a sales correlation relationship between the predicted nodes is obtained.

4. The method according to claim 3, characterized in that The bidirectional processing mechanism is used to encode the time sequence relationship state of the sales node, specifically including: Expand the relationships of node data using channels and customers as primary keys, and transform data containing time series relationships into a node-oriented time series relationship graph; Extract channel sales, customer sales, channel product share, and customer share by date; Construct a bidirectional network structure and encode the temporal relationship state of sales nodes through the forward network and backward network respectively; The temporal relationship parameters of the temporal relationship graph are grouped by channel and customer to obtain a relationship matrix containing time steps; Design a business logic interface, import the relationship matrix and historical data into the data set according to the business logic interface, and complete the input of business logic.

5. The method according to claim 1, wherein Based on the graph attention mechanism, a missing value regression model for multi-dimensional relationships is constructed to identify the attributes and states that need to be filled, specifically including: Generate an initial feature vector for each node in the graph, including node type, attributes, historical behavior and other information; Calculate the attention weight coefficient according to different types of node relationships to achieve differentiated processing of the importance of nodes; Based on the degree of association between nodes, the attention weight is dynamically adjusted to enable the system to focus on more important node relationships; Aggregate neighborhood node information, update the current node feature representation, and achieve effective information transmission in the graph structure; By stacking multiple attention layers, high-order features are extracted layer by layer to capture complex relationship patterns between nodes.

6. The method according to claim 1, characterized in that The Monte Carlo tree search is used to perform potential relationship reasoning to obtain the adaptive node attribute filling result, which specifically includes: Select the node data with the current channel as the primary key as the analysis object, and use the node state code to represent the variable value of the graph relationship matrix; Analyze the correlation between variable values ​​and the current channel temporal state, and select attribute dimensions under the channel and customer dimensions as missing state variables; Create a Monte Carlo tree structure containing several state sets, each state set contains multiple nodes, and the nodes contain two attributes: predicted value and confidence; Based on the node confidence ranking results, the state set with the highest prediction value is selected for in-depth exploration; Determine whether the maximum predicted value in the state set exceeds a preset threshold and calculate the multi-dimensional correlation strength between the state set and the target node; The optimal exploration path is selected based on the association strength, and iteration is continued until the optimal path that meets the conditions is found, and the adaptive node attribute filling result is obtained.

7. The method according to claim 6, characterized in that Based on the node confidence ranking results, the state set with the highest prediction value is selected for in-depth exploration, specifically including: Implement differentiated filling strategies for different attribute types, using regression models for numerical attributes and classification models for categorical attributes; Perform historical consistency verification, relationship consistency verification, and business rule verification, forming a three-level verification mechanism; Quantitatively evaluate the reliability of the filling results and generate a confidence score; Based on the difference between actual sales data and filling results, the model parameters are continuously adjusted to form a closed-loop optimization mechanism; Different temporal relationship state encoding methods are used for storage for different nodes, supporting a bidirectional storage mechanism of forward filling and backward filling.

8. The method according to claim 1, characterized in that The visualization layout strategy is designed based on node weights and attribute values, and the sales flow change trend is displayed through a thermodynamic model, specifically including: Design node distribution algorithms to calculate graph data layout and structure; Set node size, node coordinates, and node center position in the layout algorithm based on node weights and attribute values; Generate a visual multi-dimensional relationship map of sales data to support user interactive exploration; Design a thermodynamic diffusion model and evolve it based on the dynamic trend of the network; Through the evolution results of the thermodynamic diffusion model, a thermodynamic diffusion trend diagram is obtained, which shows the change direction and intensity of sales attributes of different nodes.

9. The method according to claim 8, characterized in that The design of the thermodynamic diffusion model and the evolution of the diffusion model in combination with the dynamic change trend of the network specifically include: Set thermodynamic gradients for different time periods and calculate thermodynamic layout states; Build a time series chart display model and adjust the matching relationship between the status of different time periods and the node layout coordinates; Realize the display of time series changes of each node in the sales process; Adjacent nodes are aggregated together iteratively until the vectors under all relationship types converge; The similarity between nodes is calculated, and the layout is automatically optimized according to the similarity between the nodes, so as to realize the dynamic update of the thermodynamic diffusion trend diagram.

10. Sales data sparsity compensation and adaptive visualization system, characterized by: include: A basic data module is used to construct a sales data knowledge graph through spatiotemporal clustering and attribute partitioning, wherein the knowledge graph includes the temporal relationship between sales nodes; The graph learning module is used to obtain the historical and future relationship strengths between nodes based on a bidirectional time series processing mechanism and predict the probability of sales association relationships between nodes; The filling prediction module is used to build a missing value regression model for multi-dimensional relationships based on the graph attention mechanism, and perform potential relationship reasoning through Monte Carlo tree search to obtain adaptive node attribute filling results; Visual modeling module, used to design visual layout strategies based on node weights and attribute values, and to demonstrate sales flow trends through thermodynamic models; Among them, the basic data module includes a data preprocessing component and a data extraction component; the graph learning module includes a graph relationship learner, a temporal association relationship extractor and an attribute association relationship extractor; the filling prediction module includes a filling model learning component, a filling attribute learning component and a Monte Carlo tree search component; the visual modeling module includes a temporal change visualization component and a stream data visualization component.

Citation Information

Cited By

  • User portrait generation method and system based on big data

    CN121092969A

  • A user portrait generation method and system based on big data

    CN121092969B

  • Intelligent logistics path optimization method based on knowledge graph

    CN121503836A