State transition processing method and system based on time sequence diagram vector storage
Patent Information
- Application Number
- CN202610699290.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-20
- Publication Date
- 2026-09-25
AI Technical Summary
随着客户规模扩大、行为维度增多、业务链路不断拉长,传统方案已难以满足企业对高精度、自动化、前瞻性客户状态管理的需求
Smart Images

Figure CN122817675A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, specifically to a state transition processing method and system based on time-series graph vector storage, applicable to customer lifecycle state management and prediction in enterprise customer relationship management (CRM) scenarios. Background Technology
[0002] In the process of enterprise informatization and digital transformation, Customer Relationship Management (CRM) systems are the core tools for managing the entire customer lifecycle. In modern enterprise customer management systems, the customer lifecycle typically includes multiple states such as leads, intent, follow-up, bidding, signing, maintenance, and renewal. The ability to manage these customer states directly impacts an enterprise's customer conversion efficiency, resource utilization efficiency, and customer satisfaction. The accuracy of predicting customer state transitions, the timeliness of process triggering, and the rationality of resource allocation directly determine an enterprise's customer acquisition efficiency, operating costs, and customer retention rate.
[0003] To achieve customer status management and assessment, the industry commonly employs technologies such as customer tagging systems, rule engines, or basic machine learning models for status determination and process advancement. Traditional solutions typically store static customer information in relational databases, record behavioral flows using time-series logs, and determine whether a customer meets the conversion criteria through rule configuration or simple models. When preset conditions are met, status changes and resource scheduling are triggered manually or semi-automatically. As customer scale expands, behavioral dimensions increase, and business chains lengthen, traditional solutions are no longer sufficient to meet enterprises' needs for high-precision, automated, and forward-looking customer status management.
[0004] Existing customer state transition prediction technologies have certain limitations: First, feature extraction is insufficient and cannot simultaneously characterize customer behavior features; second, they have weak ability to identify early signals such as potential conversion intentions and churn risks; third, they treat customer state, behavior, and resources as independent objects, ignoring the correlation and transmission between states, customers, and tasks, making it impossible to identify correlation triggers and root cause attribution, and easily leading to duplicate alarms, false triggers, and resource misallocation; fourth, it is difficult to achieve rapid matching and migration decisions for similar customers, similar behaviors, and similar states; fifth, prediction, process, and scheduling are isolated from each other, lacking closed-loop feedback and real-time iteration mechanisms, and cannot continuously optimize prediction accuracy and scheduling strategies based on dynamic changes in customers and processing effects.
[0005] The existing customer status management and prediction technologies mentioned above have the technical problem of being unable to comprehensively and accurately predict the customer status of a target enterprise, and no effective solution has yet been proposed. Summary of the Invention
[0006] This invention provides a state transition processing method and system based on time-series graph vector storage, which can at least solve some of the problems existing in the prior art, improve the accuracy, real-time performance and automation level of customer state prediction, optimize resource allocation efficiency, and reduce customer churn risk and manual intervention costs.
[0007] A first aspect of this invention provides a customer state transition intelligent prediction method based on time-series graph vector storage, comprising: Step 1: Create a time-series-graph-vector fusion storage architecture. Construct a fusion storage model that integrates static attributes, dynamic behavioral time-series data, state transition relationships, and high-dimensional vector representations. Treat customers, states, behaviors, resources, and time as core entity nodes, and establish complex relational graph relationships through connecting edges. Simultaneously, generate high-dimensional vector representations for all nodes and edges. Design a storage architecture with a triple sharding strategy of "horizontal sharding + vertical sharding + vector partitioning" and a four-level indexing system to achieve triple retrieval of customer data based on time-series relationships, graph structure, and vector features.
[0008] Step 2: For the aforementioned triple retrieval, a dual-branch attention prediction model is created to accurately mine the triple implicit features. A dual-branch attention prediction model is designed, fusing graph attention-vector enhancement layers and temporal convolution-vector enhancement layers. The graph-vector branch extracts the structural relationships and vector semantic features of customer-state-behavior-resource, while the temporal-vector branch captures the temporal dependencies of behavior and dynamic vector changes. A gating fusion mechanism and vector similarity weighting are used to achieve deep fusion of the two types of features. A multi-task prediction head simultaneously outputs the target state, transition probability, and transition time window.
[0009] Step 3: Establish a dual-dimensional confidence judgment mechanism based on "transition probability + vector similarity". Divide confidence into three categories of high, medium and low confidence based on dual thresholds, and trigger different state transition processes and resource scheduling strategies respectively. Resource allocation is based on resource-customer vector similarity matching, combined with greedy algorithm and load balancing strategy to achieve optimal allocation. Establish a feedback optimization mechanism to iteratively optimize the model, storage architecture and process rules based on actual execution results and vector feature changes.
[0010] Step 4: Define a multi-dimensional anomaly system that includes vector anomaly indicators; detect anomalies in real time based on the similarity matching between customer vectors and anomaly vector pattern libraries; locate the cause of anomalies through vector source tracing and generate personalized intervention suggestions that include vector feature analysis; monitor changes in vector features during the intervention process in real time to achieve proactive discovery, accurate analysis and rapid intervention of anomalies, and reduce the risk of customer churn.
[0011] In one optional implementation, the construction of the converged storage model in step 1 includes the following sub-steps: 1.1 Definition of Core Entity Nodes and Vector Representations: Five core nodes are identified: customer, status, behavior, resource, and time. Attribute dictionaries for each node are configured, and globally unique UUIDs are assigned. Based on node attributes and historical data, high-dimensional node vectors are generated through feature fusion and vector training to ensure that the vectors can accurately represent the core features of the nodes. The integrity of node attributes and vectors is verified, redundant attributes are removed, and necessary vector generation fields are added.
[0012] 1.2 Definition of Node-to-Node Association Edges and Vector Enhancement: Define five categories of association edges, configure edge attributes, generate edge vectors based on the node vectors at both ends of the edge and association features, and enhance the vector-level expression of association relationships; clarify the association rules of association edges, verify the matching rationality of association edges with nodes and vectors, and ensure that the association relationship is consistent with the vector semantics.
[0013] 1.3 Integrated Storage and Sharding Strategy Design: A triple sharding strategy of "customer ID hash horizontal sharding + time granularity vertical sharding and vector similarity partitioning" is adopted to achieve balanced data storage; a four-level index system of node ID, edge association, timestamp, and vector similarity is constructed to improve retrieval efficiency; and a CDC incremental synchronization + vector real-time update mechanism is designed.
[0014] 1.4 Implementation of Multi-Dimensional CRUD Interfaces: Implement interfaces for adding and deleting nodes, maintaining edge relationships, and updating time-series-vector attributes; implement multi-dimensional retrieval interfaces, support parallel data reading and writing and vector retrieval, and ensure the efficiency and stability of the interfaces.
[0015] In an optional implementation, a data acquisition and preprocessing step is further included between step 1 and step 2, specifically: A data acquisition and preprocessing module is constructed to collect various types of customer data throughout their entire lifecycle in real time. This module performs data cleaning, standardization, vector generation, and annotation, eliminating invalid data, standardizing data formats, and generating a high-quality dataset containing vector representations to provide input for model training and real-time prediction. Specifically, this includes: 1.2.1 Full-dimensional data collection: Connects to multiple data sources such as CRM core business database, customer interaction logs, and resource management system to collect four types of data in real time: static basic customer data, dynamic behavior time series data, historical status transition data, and resource allocation data; aggregates collected behavior data and updates vectors, and collects status data triggered by changes and recalculates edge vectors to ensure the timeliness of data and vectors.
[0016] 1.2.2 Data cleaning and vector preprocessing: Remove duplicate, invalid, and abnormal data; convert text data into text vectors after speech-to-text conversion or keyword extraction; use the 3σ criterion to handle numerical outliers to ensure that the data meets the requirements for vector training.
[0017] 1.2.3 Data Standardization and Vector Generation: Static attributes are encoded and mapped and combined with text vectors to generate static feature vectors; time series data is timestamped, aggregated, and feature-enhanced to generate time series feature vectors; state transition sequences are encoded to generate path vectors; the three types of vectors are fused to generate a high-dimensional customer global vector, which is then normalized using LayerNorm.
[0018] 1.2.4 Data annotation and vector label association: Four core labels are labeled using a combination of manual and rule-based annotation. The labels are converted into label vectors and stored in association with customer data and vector representations. The training / validation / test sets are divided into layers in a 7:2:1 ratio to ensure a balanced distribution of the dataset.
[0019] In an optional implementation, the construction of the dual-branch attention prediction model in step 2 includes the following sub-steps: 2.1 Graph-Vector Feature Branch Construction: Input the adjacency matrix, node attribute matrix and vector representation, configure 3 layers of GAT-VE (8 attention heads), fuse node attributes and vectors to calculate attention weights, generate high-dimensional graph-vector fusion feature vectors through triple pooling; verify the effectiveness of branch features to ensure that the vector contribution of key nodes is significant.
[0020] 2.2 Temporal-Vector Feature Branch Construction: Input temporal sequence and temporal vector, configure 4 layers of TCN-VE (expansion coefficient 1 / 2 / 4 / 8), capture temporal dependencies and fuse vector semantics, generate high-dimensional temporal-vector fused feature vector through temporal-vector attention layer; verify the branch temporal dependency capture capability to ensure feature enhancement of key temporal nodes.
[0021] 2.3 Feature Fusion and Prediction Layer Construction: Gated fusion and vector similarity weighted fusion of two types of features are adopted to build a multi-task prediction head (classification + dual regression), and a total loss function including vector similarity loss is configured; regularization is added to ensure the rationality and stability of the prediction results.
[0022] 2.4 Model Training and Optimization: Xavier initialization and AdamW optimizer were used, with batch size consistent with vector dimension, 100 training epochs, and early stopping mechanism and learning rate decay enabled; model performance was evaluated on the test set, and model lightweighting was achieved through pruning, quantization and vector compression; the model was saved and version management was performed.
[0023] In an optional implementation, the two-dimensional confidence judgment mechanism in step 3 specifically refers to: High confidence: When the transition probability is ≥0.9 and the vector similarity is ≥0.85, the automatic state transition process is triggered, a process instance is automatically created, the executor is matched based on the vector similarity and the pending tasks are pushed, and the state and vector are updated after the process is completed; Medium confidence level: 0.7 ≤ conversion probability < 0.9 or 0.7 ≤ vector similarity < 0.85, triggering manual review + automatic preparation process, generating preparation suggestions and vector analysis report, pre-allocating resources, and triggering the process after manual confirmation; Low confidence: Transition probability < 0.7 and vector similarity < 0.7, only the prediction result and vector similarity analysis are recorded, and the process is not triggered; The resource scheduling strategy includes: matching resources based on target status, customer value, and vector similarity; querying resource status and vector information; allocating resources using a greedy + load balancing + vector sorting strategy, supporting the preemption of high-value customer resources and replenishment within 24 hours; and real-time monitoring of resource load and vector similarity to trigger expansion warnings.
[0024] In one optional implementation, the feedback optimization mechanism in step 3 includes the following sub-steps: 3.1 Full-dimensional data feedback collection: Collect actual results of state transitions, process / resource scheduling results, and vector matching data, and synchronize them to the fusion storage model in real time to update vectors.
[0025] 3.2 Deviation and Vector Similarity Analysis: Calculate the predicted deviation value and analyze the reasons for the deviation (including insufficient vector accuracy); locate the customer groups and vector patterns in the deviation cluster through vector clustering.
[0026] 3.3 Model and Vector Iteration: When the bias or vector matching accuracy is not up to standard, iteration is triggered to supplement new data and optimize vectors, incrementally train the model, and ensure performance improvement.
[0027] 3.4 Optimized Unified Storage Architecture: Based on feedback, optimize sharding strategies, index structures, and vector attribute fields to improve storage and retrieval efficiency.
[0028] 3.5 Process and rule optimization: Optimize process templates, resource scheduling rules and anomaly detection rules, increase vector matching weights, and improve automation and accuracy.
[0029] In one optional implementation, the implementation of the multi-dimensional anomaly system in step 4 includes the following sub-steps: 4.1 Definition of Multi-Dimensional Anomaly Indicators and Vector Patterns: Define five types of anomaly indicators: status, behavior, process, resource, and vector. Construct an anomaly vector pattern library based on historical data.
[0030] 4.2 Real-time anomaly detection and vector matching: Read data and vectors in real time, calculate anomaly index values and vector anomaly degree, and determine an anomaly and classify when the similarity is ≥0.7.
[0031] 4.3 Anomaly Cause Analysis and Vector Tracing: Input the time series-graph-vector data of abnormal customers, combine the rule base and pattern base to analyze the causes, and locate key nodes through vector tracing.
[0032] 4.4 Personalized intervention suggestion generation: Based on the cause of the anomaly and vector analysis, targeted intervention suggestions are generated, along with anomaly vector characteristics and success cases.
[0033] 4.5 Intervention Implementation and Vector Monitoring: Push abnormal information and suggestions, monitor the intervention effect and vector changes in real time, end the intervention and record logs after the abnormality is recovered.
[0034] A second aspect of the present invention provides a customer state transition intelligent prediction system based on time-series graph vector storage, used to implement the method described in any of the above claims, the system comprising: The time-series-graph-vector fusion storage module builds and deploys time-series-graph-vector fusion storage, completing functions such as node / edge definition, vector representation generation, sharding strategy design, and interface implementation. It supports time-series-structured-vectorized storage, real-time synchronization, and multi-dimensional retrieval of customer data, providing data and vector support for the entire system.
[0035] Data Acquisition and Preprocessing Module: Collects various types of customer data in real time throughout the entire customer lifecycle, performs data cleaning, standardization, vector generation, and annotation, removes invalid data, unifies data format, and generates high-quality datasets containing vector representations to provide input for model training and real-time prediction.
[0036] Dual-branch attention prediction module: responsible for building a dual-branch attention prediction model, completing model training, optimization, lightweighting and deployment, receiving customer data and vector representations output by the fusion storage module, and generating real-time state transition prediction results and vector similarity verification reports.
[0037] Real-time prediction and transformation trigger module: responsible for accessing customer real-time data and updating vector representation, calling prediction models for real-time prediction and vector verification, triggering corresponding state transformation processes based on dual-dimensional confidence, monitoring process progress and updating customer status and vector representation.
[0038] The intelligent resource scheduling module is responsible for matching resource scheduling rules based on target status, customer value level, and vector representation, querying resource availability status and vector information, and performing resource allocation, load monitoring, and expansion warning operations to ensure optimal resource utilization.
[0039] Feedback Optimization and Iteration Module: Responsible for collecting actual execution results and vector matching data, analyzing deviations and changes in vector characteristics, triggering model iteration, storage architecture optimization, and process rule optimization to achieve continuous improvement in system performance.
[0040] Anomaly Detection and Intervention Module: Responsible for defining multi-dimensional anomaly indicators and vector anomaly patterns, detecting anomaly states in real time, analyzing the causes of anomalies, generating personalized intervention suggestions, pushing them to operations personnel, and monitoring the intervention effects.
[0041] Vector Management and Optimization Module: Independently responsible for the entire process management of vector generation, updating, compression, and retrieval, supporting vector algorithm iteration and similarity calculation method switching to ensure the accuracy of vector features and retrieval efficiency.
[0042] In one optional implementation, the time-series-graph-vector fusion storage module adopts a self-developed distributed time-series-graph-vector database engine, which supports dynamic expansion and synchronous updates of nodes, edges, and vectors; the dual-branch attention prediction module adopts a containerized deployment method (Docker+K8s), which supports rapid model deployment and version management; each module is connected through API interfaces to realize real-time interaction between data and vectors, adapting to the existing IT architecture of enterprises.
[0043] In one optional implementation, the system is deployed in a distributed cluster, comprising a distributed cluster of 10 servers. Three servers are used for the time-series-graph-vector fusion storage module, two for the data acquisition and preprocessing module, two for the dual-branch attention prediction module (supporting parallel model training and real-time prediction), one for the real-time prediction and transformation triggering module, one for the intelligent resource scheduling module, and one for the feedback optimization, anomaly detection, and vector management module. The servers are configured with 32 CPU cores, 128GB of memory, and 10TB of hard disk space to ensure smooth system operation. Attached Figure Description
[0044] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0045] Figure 1 This is an overall flowchart of the intelligent prediction method for customer state transition based on time-series-graph-vector fusion storage in this embodiment of the invention.
[0046] Figure 2 This is a schematic diagram of the node, edge, and vector association structure of the time-series-graph-vector fusion storage model in an embodiment of the present invention.
[0047] Figure 3 This is a schematic diagram of the structure of the dual-branch attention prediction model in an embodiment of the present invention.
[0048] Figure 4This is a schematic diagram of the triggering logic for the two-dimensional confidence judgment and state transition process in an embodiment of the present invention.
[0049] Figure 5 This is a schematic diagram of the complete process of anomaly detection and intervention in an embodiment of the present invention.
[0050] Figure 6 This is a logical diagram illustrating intelligent resource scheduling and vector matching in an embodiment of the present invention. Detailed Implementation
[0051] The present invention will be further described below with reference to the accompanying drawings and specific embodiments. For example... Figure 1-6 As shown, a first aspect of this invention provides a customer state transition intelligent prediction method based on time-series graph vector storage, comprising: Step 1: Create a time-series-graph-vector fusion storage architecture. Construct a fusion storage model that integrates static attributes, dynamic behavioral time-series data, state transition relationships, and high-dimensional vector representations. Treat customers, states, behaviors, resources, and time as core entity nodes, and establish complex relational graph relationships through connecting edges. Simultaneously, generate high-dimensional vector representations for all nodes and edges. Design a storage architecture with a triple sharding strategy of "horizontal sharding + vertical sharding + vector partitioning" and a four-level indexing system to achieve triple retrieval of customer data based on time-series relationships, graph structure, and vector features.
[0052] Step 2: For the aforementioned triple retrieval, a dual-branch attention prediction model is created to accurately mine the triple implicit features. A dual-branch attention prediction model is designed that fuses Graph Attention-Vector Enhancement Layer (GAT-VE) and Temporal Convolution-Vector Enhancement Layer (TCN-VE). The graph-vector branch extracts the structural relationships and vector semantic features of customer-state-behavior-resource, while the temporal-vector branch captures the temporal dependencies of behavior and dynamic changes in vectors. A gating fusion mechanism and vector similarity weighting are used to achieve deep fusion of the two types of features. The target state, transition probability, and transition time window are simultaneously output through a multi-task prediction head.
[0053] Step 3: Establish a dual-dimensional confidence judgment mechanism based on "transition probability + vector similarity". Divide confidence into three categories of high, medium and low confidence based on dual thresholds, and trigger different state transition processes and resource scheduling strategies respectively. Resource allocation is based on resource-customer vector similarity matching, combined with greedy algorithm and load balancing strategy to achieve optimal allocation. Establish a feedback optimization mechanism to iteratively optimize the model, storage architecture and process rules based on actual execution results and vector feature changes.
[0054] Step 4: Define a multi-dimensional anomaly system that includes vector anomaly indicators; detect anomalies in real time based on the similarity matching between customer vectors and anomaly vector pattern libraries; locate the cause of anomalies through vector source tracing and generate personalized intervention suggestions that include vector feature analysis; monitor changes in vector features during the intervention process in real time to achieve proactive discovery, accurate analysis and rapid intervention of anomalies, and reduce the risk of customer churn.
[0055] In one optional implementation, the construction of the converged storage model in step 1 includes the following sub-steps: 1.1 Definition of Core Entity Nodes and Vector Representations: Five core nodes are identified: customer, status, behavior, resource, and time. Attribute dictionaries for each node are configured, and globally unique UUIDs are assigned. Based on node attributes and historical data, high-dimensional node vectors are generated through feature fusion and vector training to ensure that the vectors can accurately represent the core features of the nodes. The integrity of node attributes and vectors is verified, redundant attributes are removed, and necessary vector generation fields are added.
[0056] 1.2 Definition of Node-to-Node Association Edges and Vector Enhancement: Define five categories of association edges, configure edge attributes (weight, effective / ineffective time, etc.), generate edge vectors based on the node vectors at both ends of the edge and association features, and enhance the vector-level expression of association relationships; clarify the association rules of association edges (many-to-many, one-to-many, etc.), verify the matching rationality of association edges with nodes and vectors, and ensure that the association relationship is consistent with the vector semantics.
[0057] 1.3 Integrated Storage and Sharding Strategy Design: A triple sharding strategy of "customer ID hash horizontal sharding + time granularity vertical sharding (hour / day / month) + vector similarity partitioning" is adopted to achieve balanced data storage; a four-level index system (B+ tree structure) of node ID, edge association, timestamp, and vector similarity is constructed to improve retrieval efficiency; and a CDC incremental synchronization + vector real-time update mechanism is designed.
[0058] 1.4 Implementation of Multi-Dimensional CRUD Interfaces: Implement interfaces for adding and deleting nodes, maintaining edge relationships, and updating time-series-vector attributes; implement multi-dimensional retrieval interfaces (query by customer ID, status, time, and vector similarity), support parallel data reading and writing and vector retrieval, and ensure the efficiency and stability of the interfaces.
[0059] In an optional implementation, a data acquisition and preprocessing step is further included between step 1 and step 2, specifically: A data acquisition and preprocessing module is constructed to collect various types of customer data throughout their entire lifecycle in real time. This module performs data cleaning, standardization, vector generation, and annotation, eliminating invalid data, standardizing data formats, and generating a high-quality dataset containing vector representations to provide input for model training and real-time prediction. Specifically, this includes: 1.2.1 Full-dimensional data collection: Connects to multiple data sources such as CRM core business database, customer interaction logs, and resource management system to collect four types of data in real time: static basic customer data, dynamic behavior time series data, historical status transition data, and resource allocation data; aggregates collected behavior data and updates vectors, and collects status data triggered by changes and recalculates edge vectors to ensure the timeliness of data and vectors.
[0060] 1.2.2 Data cleaning and vector preprocessing: Remove duplicate, invalid, and abnormal data; convert text data into text vectors after speech-to-text conversion or keyword extraction; use the 3σ criterion to handle numerical outliers to ensure that the data meets the requirements for vector training.
[0061] 1.2.3 Data Standardization and Vector Generation: Static attributes are encoded and mapped and combined with text vectors to generate static feature vectors; time series data is timestamped, aggregated, and feature-enhanced to generate time series feature vectors; state transition sequences are encoded to generate path vectors; the three types of vectors are fused to generate a high-dimensional customer global vector, which is then normalized using LayerNorm.
[0062] 1.2.4 Data annotation and vector label association: Four core labels are labeled using a combination of manual and rule-based annotation. The labels are converted into label vectors and stored in association with customer data and vector representations. The training / validation / test sets are divided into layers in a 7:2:1 ratio to ensure a balanced distribution of the dataset.
[0063] In an optional implementation, the construction of the dual-branch attention prediction model in step 2 includes the following sub-steps: 2.1 Graph-Vector Feature Branch Construction: Input the adjacency matrix, node attribute matrix and vector representation, configure 3 layers of GAT-VE (8 attention heads), fuse node attributes and vectors to calculate attention weights, generate high-dimensional graph-vector fusion feature vectors through triple pooling; verify the effectiveness of branch features to ensure that the vector contribution of key nodes is significant.
[0064] 2.2 Temporal-Vector Feature Branch Construction: Input temporal sequence and temporal vector, configure 4 layers of TCN-VE (expansion coefficient 1 / 2 / 4 / 8), capture temporal dependencies and fuse vector semantics, generate high-dimensional temporal-vector fused feature vector through temporal-vector attention layer; verify the branch temporal dependency capture capability to ensure feature enhancement of key temporal nodes.
[0065] 2.3 Feature Fusion and Prediction Layer Construction: Gated fusion and vector similarity weighted fusion of two types of features are adopted to build a multi-task prediction head (classification + dual regression), and a total loss function including vector similarity loss is configured; regularization is added to ensure the rationality and stability of the prediction results.
[0066] 2.4 Model Training and Optimization: Xavier initialization and AdamW optimizer were used, with batch size consistent with vector dimension, 100 training epochs, and early stopping mechanism and learning rate decay enabled; model performance was evaluated on the test set, and model lightweighting was achieved through pruning, quantization and vector compression; the model was saved and version management was performed.
[0067] In an optional implementation, the two-dimensional confidence judgment mechanism in step 3 specifically refers to: High confidence: When the transition probability is ≥0.9 and the vector similarity is ≥0.85, the automatic state transition process is triggered, a process instance is automatically created, the executor is matched based on the vector similarity and the pending tasks are pushed, and the state and vector are updated after the process is completed; Medium confidence level: 0.7 ≤ conversion probability < 0.9 or 0.7 ≤ vector similarity < 0.85, triggering manual review + automatic preparation process, generating preparation suggestions and vector analysis report, pre-allocating resources, and triggering the process after manual confirmation; Low confidence: Transition probability < 0.7 and vector similarity < 0.7, only the prediction result and vector similarity analysis are recorded, and the process is not triggered; The resource scheduling strategy includes: matching resources based on target status, customer value, and vector similarity; querying resource status and vector information; allocating resources using a greedy + load balancing + vector sorting strategy, supporting the preemption of high-value customer resources and replenishment within 24 hours; and real-time monitoring of resource load and vector similarity to trigger expansion warnings.
[0068] In one optional implementation, the feedback optimization mechanism in step 3 includes the following sub-steps: 3.1 Full-dimensional data feedback collection: Collect actual results of state transitions, process / resource scheduling results, and vector matching data, and synchronize them to the fusion storage model in real time to update vectors.
[0069] 3.2 Deviation and Vector Similarity Analysis: Calculate the predicted deviation value and analyze the reasons for the deviation (including insufficient vector accuracy); locate the customer groups and vector patterns in the deviation cluster through vector clustering.
[0070] 3.3 Model and Vector Iteration: When the bias or vector matching accuracy is not up to standard, iteration is triggered to supplement new data and optimize vectors, incrementally train the model, and ensure performance improvement.
[0071] 3.4 Optimized Unified Storage Architecture: Based on feedback, optimize sharding strategies, index structures, and vector attribute fields to improve storage and retrieval efficiency.
[0072] 3.5 Process and rule optimization: Optimize process templates, resource scheduling rules and anomaly detection rules, increase vector matching weights, and improve automation and accuracy.
[0073] In one optional implementation, the implementation of the multi-dimensional anomaly system in step 4 includes the following sub-steps: 4.1 Definition of Multi-Dimensional Anomaly Indicators and Vector Patterns: Define five types of anomaly indicators: status, behavior, process, resource, and vector. Construct an anomaly vector pattern library based on historical data.
[0074] 4.2 Real-time anomaly detection and vector matching: Read data and vectors in real time, calculate anomaly index values and vector anomaly degree, and determine an anomaly and classify when the similarity is ≥0.7.
[0075] 4.3 Anomaly Cause Analysis and Vector Tracing: Input the time series-graph-vector data of abnormal customers, combine the rule base and pattern base to analyze the causes, and locate key nodes through vector tracing.
[0076] 4.4 Personalized intervention suggestion generation: Based on the cause of the anomaly and vector analysis, targeted intervention suggestions are generated, along with anomaly vector characteristics and success cases.
[0077] 4.5 Intervention Implementation and Vector Monitoring: Push abnormal information and suggestions, monitor the intervention effect and vector changes in real time, end the intervention and record logs after the abnormality is recovered.
[0078] A second aspect of the present invention provides a customer state transition intelligent prediction system based on time-series graph vector storage, used to implement the method described in any of the above claims, the system comprising: The time-series-graph-vector fusion storage module builds and deploys time-series-graph-vector fusion storage, completing functions such as node / edge definition, vector representation generation, sharding strategy design, and interface implementation. It supports time-series-structured-vectorized storage, real-time synchronization, and multi-dimensional retrieval of customer data, providing data and vector support for the entire system.
[0079] The data acquisition and preprocessing module collects various types of customer data throughout their entire lifecycle in real time, performs data cleaning, standardization, vector generation, and annotation, removes invalid data, unifies data formats, and generates high-quality datasets containing vector representations to provide input for model training and real-time prediction.
[0080] The dual-branch attention prediction module is responsible for building a dual-branch attention prediction model, completing model training, optimization, lightweighting and deployment, receiving customer data and vector representations output by the fusion storage module, and generating real-time state transition prediction results and vector similarity verification reports.
[0081] The real-time prediction and transformation triggering module is responsible for accessing real-time customer data and updating vector representations, calling prediction models for real-time prediction and vector verification, triggering corresponding state transition processes based on dual-dimensional confidence levels, monitoring process progress, and updating customer status and vector representations.
[0082] The intelligent resource scheduling module is responsible for matching resource scheduling rules based on target status, customer value level, and vector representation, querying resource availability status and vector information, and performing resource allocation, load monitoring, and expansion warning operations to ensure optimal resource utilization.
[0083] The feedback optimization and iteration module is responsible for collecting actual execution results and vector matching data, analyzing deviations and changes in vector characteristics, triggering model iteration, storage architecture optimization, and process rule optimization to achieve continuous improvement in system performance.
[0084] The anomaly detection and intervention module is responsible for defining multi-dimensional anomaly indicators and vector anomaly patterns, detecting anomaly states in real time, analyzing the causes of anomalies, generating personalized intervention suggestions, pushing them to operations personnel, and monitoring the intervention effects.
[0085] The vector management and optimization module is independently responsible for the entire process management of vector generation, updating, compression, and retrieval. It supports vector algorithm iteration and similarity calculation method switching to ensure the accuracy of vector features and retrieval efficiency.
[0086] In one optional implementation, the time-series-graph-vector fusion storage module adopts a self-developed distributed time-series-graph-vector database engine, which supports dynamic expansion and synchronous updates of nodes, edges, and vectors; the dual-branch attention prediction module adopts a containerized deployment method (Docker+K8s), which supports rapid model deployment and version management; each module is connected through API interfaces to realize real-time interaction between data and vectors, adapting to the existing IT architecture of enterprises.
[0087] In one optional implementation, the system is deployed in a distributed cluster, comprising a distributed cluster of 10 servers. Three servers are used for the time-series-graph-vector fusion storage module, two for the data acquisition and preprocessing module, two for the dual-branch attention prediction module (supporting parallel model training and real-time prediction), one for the real-time prediction and transformation triggering module, one for the intelligent resource scheduling module, and one for the feedback optimization, anomaly detection, and vector management module. The servers are configured with 32 CPU cores, 128GB of memory, and 10TB of hard disk space to ensure smooth system operation.
Claims
1. A customer state transition intelligent prediction method based on time-series graph vector storage, characterized in that, Includes the following steps: Step 1: Create a time-series-graph-vector fusion storage architecture, construct a fusion storage model that integrates static attributes, dynamic behavioral time-series, state transition associations, and high-dimensional vector representations. Customers, states, behaviors, resources, and time are used as core entity nodes, and complex association graph relationships are established through association edges. Simultaneously, high-dimensional vector representations are generated for all nodes and edges. A storage architecture with a triple sharding strategy of "horizontal sharding + vertical sharding + vector partitioning" and a four-level indexing system is designed to achieve triple retrieval of customer data based on time-series associations, graph structure, and vector features. Step 2: For the aforementioned triple retrieval, a dual-branch attention prediction model is created to accurately mine the triple implicit features; a dual-branch attention prediction model is designed that fuses graph attention-vector enhancement layer (GAT-VE) and temporal convolution-vector enhancement layer (TCN-VE). The graph-vector branch extracts the structural association and vector semantic features of customer-state-behavior-resource, while the temporal-vector branch captures the temporal dependencies of behavior and the dynamic changes of vectors; a gating fusion mechanism and vector similarity weighting are used to achieve deep fusion of the two types of features, and the target state, transition probability, and transition time window are simultaneously output through a multi-task prediction head; Step 3: Establish a two-dimensional confidence judgment mechanism based on "transition probability + vector similarity", divide confidence into three categories of high, medium and low confidence based on dual thresholds, and trigger different state transition processes and resource scheduling strategies respectively; Resource allocation is based on resource-customer vector similarity matching, combined with a greedy algorithm and load balancing strategy to achieve optimal allocation; a feedback optimization mechanism is established to iteratively optimize the model, storage architecture and process rules based on actual execution results and vector feature changes; Step 4: Define a multi-dimensional anomaly system that includes vector anomaly indicators; detect anomalies in real time based on the similarity matching between customer vectors and anomaly vector pattern libraries; locate the cause of anomalies through vector source tracing and generate personalized intervention suggestions that include vector feature analysis; monitor changes in vector features during the intervention process in real time to achieve proactive discovery, accurate analysis and rapid intervention of anomalies.
2. The method according to claim 1, characterized in that, The construction of the converged storage model described in step 1 includes the following sub-steps: 1.1 Definition of Core Entity Nodes and Vector Representations: Five core nodes are identified: customer, status, behavior, resource, and time. Attribute dictionaries for each node are configured, and globally unique UUIDs are assigned. Based on node attributes and historical data, high-dimensional node vectors are generated through feature fusion and vector training. The integrity of node attributes and vectors is verified. 1.2 Definition of Inter-node Association Edges and Vector Enhancement: Define five categories of association edges, configure edge attributes, generate edge vectors based on the node vectors at both ends of the edge and association features, clarify the association rules of the association edges and verify the matching rationality; 1.3 Integrated Storage and Sharding Strategy Design: A triple sharding strategy of "customer ID hash horizontal sharding + time granularity vertical sharding + vector similarity partitioning" is adopted to build a four-level index system of node ID, edge association, timestamp, and vector similarity, and a CDC incremental synchronization + vector real-time update mechanism is designed. 1.4 Implementation of multi-dimensional CRUD interfaces: Implement interfaces for adding and deleting nodes, maintaining edge relationships, updating time-series-vector attributes, and multi-dimensional retrieval interfaces, supporting parallel data reading and writing and vector retrieval.
3. The method according to claim 1, characterized in that, Between Step 1 and Step 2, there is also a data collection and preprocessing step, which is as follows: connect to multiple data sources to collect various types of customer data throughout the entire life cycle in real time, perform data cleaning, standardization, vector generation, and labeling operations, remove invalid data, unify data format, and generate a high-quality dataset containing vector representations; use manual and rule-based labeling methods to label core tags, convert the tags into tag vectors and store them in association with customer data and vector representations, and divide the training / validation / test sets according to a preset ratio.
4. The method according to claim 1, characterized in that, The construction of the dual-branch attention prediction model in step 2 includes the following sub-steps: 2.1 Graph-vector feature branch construction: Input adjacency matrix, node attribute matrix and vector representation, configure multi-layer GAT-VE layer, fuse node attributes and vector to calculate attention weight, and generate high-dimensional graph-vector fusion feature vector through triple pooling; 2.2 Temporal-Vector Feature Branch Construction: Input temporal sequence and temporal vector, configure multi-layer TCN-VE layers, capture temporal dependencies and fuse vector semantics, and generate high-dimensional temporal-vector fused feature vectors through temporal-vector attention layers; 2.3 Feature Fusion and Prediction Layer Construction: Gated fusion and vector similarity weighted fusion of two types of features are adopted to build a multi-task prediction head, and a total loss function including vector similarity loss is configured and regularization processing is added; 2.4 Model Training and Optimization: The model is lightweight by using a preset initialization method and optimizer, setting the batch size and number of training rounds, enabling early stopping mechanism and learning rate decay, and implementing pruning, quantization and vector compression.
5. The method according to claim 1, characterized in that, The dual-dimensional confidence judgment mechanism described in step 3 is as follows: high confidence corresponds to a conversion probability ≥ 0.9 and a vector similarity ≥ 0.85, triggering an automatic state transition process; medium confidence corresponds to a conversion probability < 0.9 or a vector similarity < 0.85, triggering a manual review + automatic preparation process; low confidence corresponds to a conversion probability < 0.7 and a vector similarity < 0.7, only recording the prediction result and vector similarity analysis; the resource scheduling strategy includes resource-customer vector similarity calculation, resource status query, resource allocation based on greedy algorithm + load balancing + vector sorting, and resource preemption and expansion early warning mechanism.
6. The method according to claim 1, characterized in that, The multi-dimensional anomaly system described in step 4 includes five types of anomaly indicators: state, behavior, process, resources, and vector. An anomaly vector pattern library is constructed based on historical data. Anomaly detection calculates the anomaly indicator value and vector anomaly degree in real time. When the similarity is ≥0.7, it is judged as an anomaly and classified. Vector tracing is used to locate the key behavioral nodes and vector feature changes that cause the anomaly. Personalized intervention suggestions include anomaly vector feature analysis and historical successful cases.
7. The method according to claim 1, characterized in that, The feedback optimization mechanism described in step 3 includes the following sub-steps: 3.1 Full-dimensional data feedback collection: Collect actual results of state transitions, process / resource scheduling results, and vector matching data, synchronize them to the fused storage model, and update the vectors; 3.2 Deviation and Vector Similarity Analysis: Calculate the predicted deviation value, analyze the causes of the deviation, and locate the customer groups and vector patterns in the deviation cluster through vector clustering; 3.3 Model and Vector Iteration: When the bias or vector matching accuracy is not up to standard, supplement with new data and optimize vectors to incrementally train the model; 3.4 Optimization of Integrated Storage Architecture and Process Rules: Based on feedback, optimize sharding strategies, index structures, vector attribute fields, process templates, resource scheduling rules, and anomaly detection rules.
8. A customer state transition intelligent prediction system based on time-series graph vector storage, characterized in that, The system, used to implement the method according to any one of claims 1-7, comprises: a time-series-graph-vector fusion storage module, used to construct and deploy time-series-graph-vector fusion storage, complete functions such as node / edge definition, vector representation generation, sharding strategy design, and interface implementation, and support time-series-structured-vectorized storage, real-time synchronization, and multi-dimensional retrieval of customer data; a data acquisition and preprocessing module, used to collect various types of customer data throughout their entire lifecycle in real time, perform data cleaning, standardization, vector generation, and annotation operations, and generate high-quality datasets; a dual-branch attention prediction module, used to build a dual-branch attention prediction model, complete model training, optimization, lightweighting, and deployment, and generate real-time state transition prediction results and vector similarity verification reports; and a real-time prediction and transition triggering module, used to access real-time customer data and update vectors. The system comprises several modules: Representation, which invokes the prediction model for real-time prediction and vector validation, and triggers corresponding state transition processes based on dual-dimensional confidence levels; Intelligent Resource Scheduling, which matches resource scheduling rules with target status, customer value level, and vector representation, and executes resource allocation, load monitoring, and expansion warning operations; Feedback Optimization and Iteration, which collects actual execution results and vector matching data, analyzes deviations and vector feature changes, and triggers model iteration, storage architecture optimization, and process rule optimization; Anomaly Detection and Intervention, which defines multi-dimensional anomaly indicators and vector anomaly patterns, detects anomaly states in real time, analyzes the causes of anomalies, generates personalized intervention suggestions, and monitors the intervention effects; and Vector Management and Optimization, which independently manages the entire process of vector generation, updating, compression, and retrieval, and supports vector algorithm iteration and similarity calculation method switching.
9. The system according to claim 8, characterized in that, The time-series-graph-vector fusion storage module adopts a self-developed distributed time-series-graph-vector database engine, which supports dynamic expansion and synchronous updates of nodes, edges, and vectors; the dual-branch attention prediction module adopts a containerized deployment method (Docker+K8s), which supports rapid model deployment and version management; each module is connected through API interfaces to realize real-time interaction between data and vectors, adapting to the existing IT architecture of enterprises.
10. The system according to claim 8, characterized in that, The system adopts a distributed cluster deployment, including multiple servers, each corresponding to a functional module. The server configuration meets the performance requirements for data storage, model training, real-time prediction, and resource scheduling. After deployment, the system undergoes a system test for a preset period of time to ensure that indicators such as data retrieval efficiency, model prediction accuracy, real-time response speed, and anomaly detection accuracy meet the preset requirements.