A dynamic tag processing method, system, medium and product for customer leads

CN121479356BActive Publication Date: 2026-09-29BEIJING YOU TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511703981.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-19
Publication Date
2026-09-29
Estimated Expiration
2045-11-19

AI Technical Summary

Technical Problem

[0004]然而现有技术通常是对某个时间点的客户数据进行一次“快照式”分析,得到的标签本质上仍是静态的

Benefits of technology

1、通过在客户线索数据处理流程中引入客户实体解析、客户关系图谱构建及特征聚合增强等技术手段,使得客户状态的建模不仅基于个体属性特征,更融合了客户之间的关联关系,从而生成能够动态反映客户当前综合状态的目标特征向量,并在持续的时间周期内进行增量聚类与动态演化指标计算,进而动态生成客户标签,相较于现有技术中基于静态快照分析的标签方法,能够有效刻画客户群体的演化趋势,提高了动态标签处理的准确性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121479356B_ABST
    Figure CN121479356B_ABST
Patent Text Reader

Abstract

A dynamic label processing method, system, medium and product of customer clues, wherein the method comprises: acquiring a customer clue data stream; generating a corresponding customer main clue data for each customer entity; performing feature extraction on the target customer attribute field and the target field value in the customer main clue data to determine the individual intention feature of the customer entity; constructing a customer relationship graph; performing aggregation enhancement on the individual intention feature to obtain a target feature vector of the customer entity; performing incremental clustering on the target feature vector in each preset time period to obtain a plurality of customer clustering clusters, and calculating a dynamic evolution index of each customer clustering cluster in a plurality of continuous preset time periods; and generating a customer label for each customer entity based on the dynamic evolution index. The application can improve the accuracy of dynamic label processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of dynamic tagging technology, specifically to a method, system, medium, and product for dynamic tagging of customer leads. Background Technology

[0002] With the development of digital marketing and intelligent customer relationship management, enterprises need to process massive amounts of customer lead data and accurately identify and classify customers. Customer lead tagging is an important means of achieving refined customer operations. By assigning different tags to customers, enterprises can better understand customer needs, predict customer behavior, and thus develop targeted marketing strategies and service solutions.

[0003] Existing technologies typically employ clustering algorithms (such as K-Means, DBSCAN, etc.) to analyze customers' historical behavioral data, demographic data, etc., automatically dividing customers with similar characteristics into different groups and assigning a label to each group, such as "price-sensitive young group" or "high-loyalty business group".

[0004] However, existing technologies typically perform a "snapshot" analysis of customer data at a specific point in time, resulting in tags that are essentially static. They describe a customer's state over a period of time, tagging based on past events, but lack insight into the evolving trends of the customer group, thus reducing the accuracy of dynamic tagging. Summary of the Invention

[0005] This application provides a method, system, medium, and product for processing dynamic tags for customer leads, which improves the accuracy of dynamic tag processing.

[0006] The first aspect of this application provides a method for dynamic tagging of customer leads, the method comprising: Acquire a customer lead data stream, the customer lead data stream including multiple target customer lead data, the target customer lead data including identity information, customer attribute fields and field values ​​corresponding to the customer attribute fields; Based on the identity information, the customer lead data stream is parsed to generate corresponding customer lead data for each customer entity. The customer lead data includes the target customer attribute field of the customer entity and the target field value corresponding to the target customer attribute field. One customer entity corresponds to one customer lead data. Feature extraction is performed on the target customer attribute field and the target field value in the customer lead data to determine the individual intent features of the customer entity; Obtain the relationships between multiple customer entities, and construct a customer relationship graph based on the relationships and the customer lead data; Based on the customer relationship graph, the individual intent features are aggregated and enhanced to obtain the target feature vector of the customer entity, which is used to characterize the current comprehensive state of the customer entity; Incremental clustering is performed on the target feature vector within each preset time period to obtain multiple customer clusters. The dynamic evolution index of each customer cluster is calculated within multiple consecutive preset time periods. The dynamic evolution index is used to characterize the customer lead cycle status of the customer group represented by the customer cluster. Customer tags are generated for each customer entity based on the dynamic evolution indicators.

[0007] Optionally, based on the identity information, customer entity parsing is performed on the customer lead data stream to generate corresponding customer lead data for each customer entity, specifically including: Based on the identity information, the association analysis is performed on multiple target customer lead data in the customer lead data stream, and multiple target customer lead data belonging to the same customer entity are combined into a target customer lead dataset, with each target customer lead dataset corresponding to a customer entity; The data in the target customer lead dataset corresponding to the customer entity is deduplicated to obtain the customer lead data of the customer entity.

[0008] Optionally, the data in the target customer lead dataset corresponding to the customer entity is deduplicated to obtain the customer lead data of the customer entity, specifically including: If the same target customer attribute field in the target customer lead dataset has multiple conflicting field values, and the target customer lead data corresponding to the conflicting field values ​​meets the preset conflict deadlock condition, then the historical customer lead data of the customer entity within the preset historical time period is obtained, and a time decay weight is set for the historical customer lead data. The consistency score is obtained by comparing the multiple conflicting field values ​​with the historical field values ​​corresponding to the target customer attribute field, which have the time decay weight. The conflict field value with the highest consistency score is determined as the target field value of the target customer attribute field in the customer lead data, thus obtaining the customer lead data.

[0009] Optionally, based on the customer relationship graph, the individual intent features are aggregated and enhanced to obtain the target feature vector of the customer entity, specifically including: The individual intent features of adjacent customer entities that have a direct relationship with the target customer entity in the customer relationship graph are combined into neighborhood common features, where the target customer entity is any of the customer entities in the customer relationship graph; Calculate the neighborhood expression intensity corresponding to each first feature component in the common neighborhood features; For any second feature component in the individual intent features, find the neighborhood expression strength corresponding to the second feature component in the neighborhood common features; If the neighborhood expression intensity is less than a preset expression intensity threshold, or if there is no corresponding neighborhood expression intensity for the second feature component in the common features of the neighborhood, then the second feature component is determined as a unique feature component of the target customer entity. The target feature vector is determined based on the unique feature components.

[0010] Optionally, determining the target feature vector based on the unique feature components specifically includes: If the unique feature component exists in the individual intent feature, the neighborhood common feature and the individual intent feature are combined into a preliminary fusion feature, and the expression weight of the unique feature component in the preliminary fusion feature is increased based on a preset weight adjustment rule to obtain the target feature vector; If the unique feature component is not present in the individual intent feature, then the neighborhood common feature and the individual intent feature are combined to obtain the target feature vector.

[0011] Optionally, the dynamic evolution index of each customer cluster is calculated within multiple consecutive preset time periods, specifically including: Calculate the member overlap between the first customer cluster within the first preset time period and the second customer cluster within the second preset time period, wherein the first preset time period and the second preset time period are adjacent, and the first preset time period is before the second preset time period. If the overlap of the members is greater than the preset inheritance threshold, then the first customer cluster and the second customer cluster are determined as a pair of continuous customer clusters; For any of the continuous customer clusters, the first period cluster core state vector is determined based on the target feature vectors of all customer entities within the first preset time period. Based on the target feature vectors of all customer entities within the second preset time period, determine the core state vector of the second period cluster; Calculate the state drift between the first periodic cluster core state vector and the second periodic cluster core state vector, and determine the state drift as the dynamic evolution index.

[0012] Optionally, customer tags are generated for each customer entity based on the dynamic evolution indicators, specifically including: Obtain the second customer cluster to which the target customer entity belongs within the second preset time period; If there exists a first customer cluster that forms the continuous customer cluster with the second customer cluster, then the number of first members of the first customer cluster and the number of second members of the second customer cluster are counted respectively, and the scale change rate is calculated. The scale change rate is the ratio of the difference between the number of second members and the number of first members to the number of first members. If the dynamic evolution index is less than or equal to the preset drift threshold, and the absolute value of the scale change rate is less than the preset stability threshold, or there is no first customer cluster that forms the continuous customer cluster with the second customer cluster, then a steady-state attribution label is generated for the target customer entity as the customer label. Otherwise, a migration trend label is generated for the target customer entity as the customer label.

[0013] In a second aspect, embodiments of this application provide a dynamic tagging system for customer leads. The dynamic tagging system for customer leads includes one or more processors and a memory. The memory is coupled to the one or more processors and is used to store computer program code, which includes computer instructions. The one or more processors invoke the computer instructions to cause the dynamic tagging system for customer leads to perform the method described in the first aspect and any possible implementation thereof.

[0014] Thirdly, embodiments of this application provide a computer-readable storage medium including instructions that, when executed on a dynamic tagging system for customer leads, cause the dynamic tagging system for customer leads to perform the method described in the first aspect and any possible implementation thereof.

[0015] Fourthly, embodiments of this application provide a computer program product containing instructions that, when the computer program product is run on a dynamic tagging system for customer leads, cause the dynamic tagging system for customer leads to execute the method described in the first aspect and any possible implementation thereof.

[0016] In summary, one or more technical solutions provided in this application have at least the following technical effects or advantages: 1. By introducing techniques such as customer entity parsing, customer relationship graph construction, and feature aggregation enhancement into the customer lead data processing flow, the modeling of customer status is not only based on individual attribute features, but also integrates the relationships between customers. This generates a target feature vector that can dynamically reflect the current comprehensive status of customers, and performs incremental clustering and dynamic evolution index calculation over a continuous time period to dynamically generate customer tags. Compared with the existing tagging methods based on static snapshot analysis, this can effectively depict the evolution trend of customer groups and improve the accuracy of dynamic tagging processing.

[0017] 2. By performing association analysis based on identity information, multiple target customer leads belonging to the same customer entity are combined into a unified dataset. A data deduplication mechanism is then introduced to generate uniquely corresponding master customer leads, ensuring the consistency and completeness of customer attribute information. Especially when dealing with multiple conflicting field values ​​in target customer attributes, historical master customer data and its time decay weights are further incorporated to perform consistency scoring on conflicting field values. This effectively integrates the stability of historical data with the timeliness of current data, thereby determining the optimal field value based on the scoring results, improving the accuracy and robustness of the master customer lead data. This solves the data confusion problem caused by conflicting customer attribute fields, enhances the stability of customer entity identification and the credibility of master lead data, and provides a high-quality data foundation for subsequent customer modeling and tag generation.

[0018] 3. By calculating the overlap of members among customer clusters over multiple consecutive preset time periods, continuous customer clusters with evolutionary relationships are identified. Based on the target feature vectors of customer entities in the clusters at different times, a core state vector of the periodic clusters is constructed. Then, the state drift between clusters is calculated as a dynamic evolution indicator. This not only quantifies the state change trend of customer groups in the feature space but also effectively captures the dynamic evolutionary relationship between customer clusters. Furthermore, by analyzing the rate of change in the member size of continuous customer clusters, combined with the state drift and preset thresholds, it is determined whether the customer group is in a relatively stable state or a behavioral migration trend. This generates steady-state affiliation labels or migration trend labels for customer entities, realizing the transformation of customer tags from static description to dynamic evolution. This solves the problem that existing tagging systems cannot reflect the evolutionary characteristics of customer groups, making customer tags more timely, behaviorally perceptive, and predictive, providing enterprises with accurate dynamic customer insights and more forward-looking marketing decision support. Attached Figure Description

[0019] Figure 1 This is a flowchart illustrating a dynamic tagging method for customer leads in an embodiment of this application. Figure 2This is a schematic diagram of the process for generating customer lead data in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of the electronic device in the embodiments of this application.

[0020] Explanation of reference numerals in the attached drawings: 301, Central Processing Unit; 302, Read-Only Memory; 303, Random Access Memory; 304, Bus; 305, Input / Output Interface; 306, Input Section; 307, Output Section; 308, Storage Section; 309, Communication Section; 310, Driver; 311, Removable Media. Detailed Implementation

[0021] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.

[0022] In the description of the embodiments of this application, the words "for example" or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design that is described as "for example" or "for instance" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design options. Rather, the use of the words "for example" or "for instance" is intended to present the relevant concepts in a specific manner.

[0023] In the description of the embodiments of this application, the term "multiple" means two or more. For example, multiple systems means two or more systems, and multiple screen terminals means two or more screen terminals. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the indicated technical features. Thus, a feature defined with "first" or "second" may explicitly or implicitly include one or more of that feature. The terms "comprising," "including," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.

[0024] Figure 1 This is a flowchart illustrating a dynamic tagging method for customer leads according to an embodiment of this application.

[0025] Please see Figure 1 This application provides a method for processing dynamic tags for customer leads, the method comprising: S101. Obtain customer lead data stream, wherein the customer lead data stream includes multiple target customer lead data, wherein the target customer lead data includes identity information, customer attribute fields and field values ​​corresponding to the customer attribute fields; In implementing step S101, it is first necessary to clarify the meaning and role of "customer lead data stream." Customer lead data stream refers to the collection of potential customer-related data that has not yet completed a closed-loop conversion, collected in real-time or in batches from different sources (such as website visits, app interactions, ad clicks, third-party data platforms, CRM systems, etc.) within a digital marketing system. Due to the diversity of customer online behavior and information sources, customer lead data streams exhibit high heterogeneity and dynamism. Therefore, unified structured processing is required in the initial stage of data entry into the system to support subsequent customer identification and dynamic tag generation.

[0026] In practical implementation, the system first utilizes a multi-source data access module, which can be implemented based on Apache Kafka technology. For example, by using the Kafka Connect component, the system can extract data in real time from various data sources (such as database logs from enterprise CRM systems, user behavior event streams from WeChat official accounts / mini-programs, call records from call centers, etc.). This data is then unified into a JSON Schema format and published to different Kafka topics, forming a high-throughput, low-latency distributed customer lead data stream, providing a data foundation for subsequent real-time processing steps. Customer lead data from multiple channels is collected, and the unstructured or semi-structured raw data is converted into a unified format of target customer lead data using a data access adapter. Each target customer lead must contain three key fields: identity information, customer attribute fields, and the corresponding field values ​​for the customer attribute fields. The identity information is used to uniquely identify potential customers, such as phone numbers, email addresses, device IDs, cookie identifiers, and social media account IDs. This information enables unified aggregation of customer behavior across channels and time periods. Customer attribute fields are structured field names used to describe customer characteristics, such as "age", "occupation", "interests", "recently visited pages", etc., while field values ​​are the specific behaviors of customers under the above fields, such as "25 years old", "teacher", "interested in financial content", "visited insurance product pages".

[0027] After data collection is completed, the target customer lead data undergoes quality control through a data preprocessing module. This includes handling missing field values, cleaning illegal characters, validating and standardizing field types to ensure data format compliance and semantic consistency. This step also introduces a field mapping rule library to unify synonymous fields that may have naming differences across different channels. For example, "user_age" and "age" are uniformly mapped to "age," improving compatibility for subsequent processing.

[0028] By implementing step S101 as described above, a customer lead data stream with a clear structure and explicit semantics is obtained, establishing a reliable data foundation for subsequent customer entity parsing and feature extraction. The effect of this step is that it enables data standardization and structuring processing to be completed in the first stage of customer data collection, effectively reducing the anomaly rate and data conflict probability of subsequent model inputs, and improving the stability and accuracy of the entire customer tagging system.

[0029] S102. Based on the identity information, perform customer entity parsing on the customer lead data stream, and generate corresponding customer lead data for each customer entity. The customer lead data includes the target customer attribute field of the customer entity and the target field value corresponding to the target customer attribute field. One customer entity corresponds to one customer lead data. In step S102, to achieve unified collection and standardized processing of customer lead data, customer entity parsing is required based on identity information to analyze the customer lead data stream. The goal is to accurately attribute multiple target customer lead data from diverse sources and with heterogeneous content to their respective real customer individuals, and to generate a unique corresponding customer master lead data for each customer entity. This customer master lead data, as the basic data unit for subsequent feature extraction and relationship modeling, should have the characteristics of unified structure, deduplication of content, and accurate field values. Therefore, S102 not only involves the entity identification process but also needs to solve the problem of how to filter and integrate representative customer attribute information from multiple related data. Thus, it needs to be further refined into two key processing steps: first, by using association analysis to combine target customer lead data belonging to the same customer entity into a unified dataset; and then, based on this, by handling field value conflicts and deduplicating data in the dataset to obtain the final customer master lead data.

[0030] Figure 2 This is a flowchart illustrating the process of generating customer lead data in an embodiment of this application. The following is a summary of the process. Figure 2 A detailed explanation of step S102 is provided below: S201. Based on the identity information, perform correlation analysis on multiple target customer lead data in the customer lead data stream, and combine multiple target customer lead data belonging to the same customer entity into a target customer lead dataset, with each target customer lead dataset corresponding to a customer entity; In implementing step S201, to accurately generate customer lead data, it is first necessary to aggregate and unify multiple target customer lead data in the customer lead data stream at the customer entity level. Since customer lead data typically originates from multiple heterogeneous systems and channels, such as e-commerce platforms, app usage records, ad click logs, and public account interaction behavior, target customer lead data collected from different sources often contains multiple, duplicate, or even partially missing records, and may be represented with different field names or structural formats. Therefore, it is impossible to directly determine which records belong to the same real customer. To address this, "identity identification information" needs to be introduced as the core association basis, using this type of information to perform cross-source integration and entity identification of seemingly isolated lead data.

[0031] Identification information typically includes, but is not limited to, mobile phone numbers, email addresses, unique device identifiers (such as IMEI and MAC addresses), cookie IDs, account IDs, or unique user identifiers provided by third-party platforms (such as OpenID and UnionID). This information possesses uniqueness or high-confidence characteristics that can be used to identify individual customers. Therefore, during implementation, the system performs batch comparison and merging of all target customer lead data in the customer lead data stream using identification information matching rules. Specifically, an identity mapping table and matching rule library can be built in the data processing engine. Methods such as hash mapping, fuzzy matching, and field weight scoring can be used to jointly match multiple identification information fields to determine whether they belong to the same customer entity. For example, if two target customer lead data sets contain the same mobile phone number or highly similar combinations of device ID and account ID, then they can be determined to belong to the same customer entity.

[0032] In a preferred embodiment, the association analysis goes beyond simple exact matching. For identity information such as names, phone numbers, and email addresses, which may have slight differences, the system can use the Jaro-Winkler similarity algorithm combined with the TF-IDF algorithm for fuzzy matching and similarity scoring. For example, a similarity threshold (e.g., 0.92) can be set. When the combined similarity score calculated from the key identity information fields in two customer leads is higher than this threshold, they can be determined to belong to the same customer entity. This method can effectively identify duplicate customer records caused by input errors or inconsistent formats.

[0033] Furthermore, the customer entities and their relationships (such as recommendation relationships, family relationships, etc., up to 72 dimensions) collected after identity verification can be stored in graph databases such as Neo4j. Each customer entity is represented as a master data node in the graph database, ensuring the uniqueness of the customer master data and supporting subsequent complex association queries at the millisecond level.

[0034] After identity verification, all target customer lead data determined to belong to the same customer entity are merged into a single target customer lead dataset, which is then assigned a unique customer entity identifier for unified referencing in subsequent processing stages. This process not only improves data organization efficiency but also provides a context-rich data foundation for deduplication of conflicting data and subsequent feature extraction.

[0035] By implementing step S201 in the above manner, the transformation from "leads" to "entities" is achieved at the data granularity level. This enables the effective collection and organization of multi-source customer leads, significantly improving the integration of customer data, identification accuracy, and the accuracy and stability of subsequent feature modeling and tag generation. Simultaneously, this step plays a crucial bridging role in customer data lifecycle management, serving as a prerequisite and foundation for dynamic customer tagging.

[0036] To further illustrate the implementation process of step S201, the following example details how to perform correlation analysis on multiple target customer lead data in the customer lead data stream based on identity information, and combine them into a target customer lead dataset corresponding to a single customer entity: Suppose a company is operating an online insurance sales platform. Customers may access the platform's services through multiple channels, such as: Channel A: WeChat Official Account, where users access the product page by clicking on the promotional link via WeChat; Channel B: Official website registration system, users fill in registration information to apply for a trial calculation; Channel C: Telephone customer service system, where users call a hotline to inquire about product details; Channel D: Mobile App, where users can browse policy details and submit questionnaires.

[0037] In each channel, customer behavior is recorded separately as different target customer lead data, for example: Clue 1 (from WeChat Official Account): Identity information: OpenID=wx123abc; Customer attribute fields: Gender=Male, Interest category=Medical insurance; Clue 2 (from official website registration): Identity information: Mobile number=13800001111; Customer attribute fields: Age=32, City=Shanghai; Clue 3 (from telephone customer service): Identity information: Mobile number=13800001111; Customer attribute fields: Inquired product=Critical illness insurance, Call duration=12 minutes; Clue 4 (from APP): Identity information: DeviceID=device_xyz456, Mobile number=13800001111; Customer attribute fields: Viewed page=Policy details page, Questionnaire result=High risk preference.

[0038] During step S201, the system first extracts the identity information field from each target customer lead and inputs it into the entity recognition module for matching analysis. By constructing mapping rules including fields such as mobile phone number, OpenID, and DeviceID, the system finds that leads 2, 3, and 4 all have the same mobile phone number 13800001111, thus belonging to the same customer entity A. Although lead 4 also contains a DeviceID, this information can also be used to identify other potential leads. Regarding the OpenID wx123abc in lead 1, if the system already has a WeChat account linked to the mobile phone number, or if the user has previously established a mapping relationship through linking, lead 1 can also be classified as customer entity A.

[0039] After matching is complete, the system merges the four leads into a single target customer lead dataset, labeled as the dataset corresponding to customer entity A. This dataset will contain the customer's behavioral records, attribute information, and interaction context across different channels, forming a multi-dimensional, cross-time, and cross-platform customer profile data set.

[0040] Through this implementation method, customer leads that were originally scattered across different systems and had different structures are effectively merged into a unified dataset. This not only avoids customer identification errors caused by information fragmentation, but also provides structured input for subsequent operations such as data deduplication, feature extraction, and relationship graphing. This significantly improves the consistency and accuracy of customer data processing and ensures that dynamic tagging processing has entity consistency and data integrity.

[0041] S202. Deduplicate the data in the target customer lead dataset corresponding to the customer entity to obtain the customer lead data of the customer entity.

[0042] In step S202, to ensure the structural accuracy and semantic uniqueness of customer lead data, field-level data deduplication is required for multiple data records in the target customer lead dataset corresponding to the customer entity. Since lead data generated by the same customer entity at different times and through different channels may contain duplicate entries, updated information, or data conflicts, especially in the target customer attribute fields where multiple distinct field values ​​may appear, failure to address this will directly affect the reliability of the lead data. Therefore, step S202 not only includes routine field merging and redundancy removal operations but also further addresses the value inconsistency caused by field conflicts. Particularly in scenarios where conflicts meet preset deadlock conditions—that is, where the validity of field values ​​cannot be directly determined by time sequence or weight priority—historical customer lead data and a time decay mechanism need to be introduced to perform consistency scoring on conflicting field values, thereby selecting the most representative field value to write into the customer lead data. Specifically, this may include the following steps: If the same target customer attribute field in the target customer lead dataset has multiple conflicting field values, and the target customer lead data corresponding to the conflicting field values ​​meets the preset conflict deadlock condition, then the historical customer lead data of the customer entity within the preset historical time period is obtained, and a time decay weight is set for the historical customer lead data. The consistency score is obtained by comparing the multiple conflicting field values ​​with the historical field values ​​corresponding to the target customer attribute field, which have the time decay weight. The conflict field value with the highest consistency score is determined as the target field value of the target customer attribute field in the customer lead data, thus obtaining the customer lead data.

[0043] In the process of deduplicating data in the target customer lead dataset corresponding to a customer entity, it is first necessary to identify conflicting data fields, especially when multiple distinct field values ​​appear for the same target customer attribute field. Such conflicts typically arise from different information submitted by customers at different times or through different channels; for example, a user might fill in "teacher" as their occupation on a public WeChat account, but "freelancer" when registering on the official website. When the system cannot directly determine which field value is more accurate based on chronological order or source credibility, a pre-defined conflict deadlock condition is met. This pre-defined conflict deadlock condition is derived from extensive experience in handling historical customer lead data conflicts and common data conflict patterns in actual business scenarios. It is used to identify special cases where conventional strategies cannot effectively determine the validity of field values. During the generation of customer lead data, the pre-defined conflict deadlock condition is used to identify complex situations where field value conflicts cannot be resolved using conventional strategies (such as time priority, source credibility priority, etc.). The pre-defined conflict deadlock condition primarily aims to determine whether there is sufficient contextual information or strategic basis to directly select one as the unique field value when multiple distinct field values ​​appear under a target customer attribute field. When conflicting field values ​​have similar source confidence levels, no significant difference in data timestamps, and semantic differences exceeding a set threshold (e.g., based on natural language similarity or classification coding distance), a deadlock is considered to have occurred. Furthermore, if multiple conflicting field values ​​originate from highly reliable but different channels, and these channels are not explicitly prioritized in the system settings, a deadlock condition will also be triggered. Field conflicts meeting these conditions cannot be directly determined as valid values ​​using conventional logic; therefore, the system must rely on historical customer lead data and time decay mechanisms for consistency scoring to achieve deduplication decisions with stronger contextual semantic support. By setting such deadlock conditions, the semantic consistency and structural stability of lead data can be effectively improved, avoiding customer profile bias caused by merging erroneous field values. After identifying such deadlock conflicts, the system needs to introduce historical data with greater contextual reference value to assist in determining the validity of field values.

[0044] To resolve the aforementioned conflicts, it is necessary to obtain historical customer lead data for customer entities within a preset historical time period and introduce a time decay weighting mechanism into this historical data. Time decay weighting is a weighting strategy that reflects the gradual decrease in the influence of field values ​​over time. Its basic principle is that data closer to the current time is more reliable and better represents the customer's current true state. The system assigns progressively decreasing weight coefficients to field values ​​based on their historical occurrence time. For example, field values ​​from the last 30 days have a weight of 1.0, decaying to 0.8 from 31 to 60 days, and so on. This method allows for a more reasonable measurement of the influence of historical field values ​​in current field conflict decisions, thereby avoiding the misleading influence of old data on the current situation.

[0045] After obtaining historical field values ​​with time decay weights, the system performs consistency scoring on multiple conflicting field values ​​in the current target customer attribute fields, comparing them with historical field values ​​of the same field at different times. Consistency scoring compares the similarity between current and historical field values ​​in dimensions such as semantics, expression, and related behaviors, and then weights these similarities with time decay weights to obtain a final score for each conflicting field value. For example, if there are two conflicting field values, "financial industry" and "teacher," and "financial industry" appears more frequently in the past and was the most recent occurrence, then that field value is likely to have a higher consistency score. The scoring model can be implemented using weighted cosine similarity, edit distance, TF-IDF matching, etc., and the specific algorithm can be selected and optimized based on the field type (categorical field or text field).

[0046] Ultimately, the system identifies the conflicting field value with the highest consistency score as the target field value for that customer attribute in the customer lead data, thus completing the deduplication decision for conflicting field values. In this way, the system not only avoids information loss or incorrect merging but also fully utilizes the accumulated value of historical data, achieving intelligent resolution of field conflicts and effectively improving the accuracy and stability of customer lead data. For example, if customer A has five lead records in the past three months, with four entries listing "IT Engineer" and one "Product Manager," and a new lead contains conflicting field values ​​for "Operations Specialist" and "IT Engineer," after weighting based on historical records, "IT Engineer" has a higher consistency score and is therefore automatically selected as the customer's lead field value by the system.

[0047] This series of processing steps constructs a complete closed loop from conflict identification to historical data utilization and field value selection, enabling customer lead data to maintain data consistency and semantic accuracy in the face of complex, multi-source and dynamic data update scenarios, providing a high-quality and reliable data foundation for subsequent feature extraction, customer clustering and tag generation.

[0048] S103. Extract features from the target customer attribute field and the target field value in the customer lead data to determine the individual intent features of the customer entity; In implementing step S103, to accurately characterize customers' current behavioral preferences and potential business needs, it is necessary to extract features from the target customer attribute fields and their corresponding target field values ​​in the customer lead data, thereby determining the individual intent features of the customer entity. Individual intent features refer to a set of features that reflect a customer's sustained attention or inclination towards a certain type of product, service, or behavioral content within a specific time window, used to support the accuracy of subsequent customer state modeling, group clustering, and tag generation.

[0049] During implementation, the customer lead data is first subjected to field semantic recognition and structure mapping to extract target customer attribute fields that express intent, such as "recently viewed pages," "search keywords," "type of product inquired about," "behavior timestamp," and "click frequency." These fields often exist in the form of behavior logs, user input, or system interactions, reflecting the customer's true intent to actively or passively participate in platform interactions. The system uses a field classifier and contextual semantic rules to identify and filter fields with behavioral significance as feature candidates.

[0050] Subsequently, for each field value with behavioral semantics, a feature encoding mechanism is used for numerical processing. Feature encoding can be performed using different methods depending on the field type. For example, for categorical fields (such as "consultation product type" being "commercial insurance" or "health insurance"), one-hot encoding or embedding can be used; for numerical fields (such as "page dwell time" or "number of visits"), direct normalization can be performed; for textual fields (such as "search keywords"), semantic feature vectors can be extracted using TF-IDF, word vector models (such as Word2Vec, BERT), etc. The encoded field values ​​constitute the initial feature vector. In another embodiment, to capture contextual semantics more deeply, a pre-trained language model based on the Transformer architecture, such as BERT (Bidirectional Encoder Representations from Transformers), can be used to encode the customer's text and behavioral sequence data, generating high-dimensional feature vectors that accurately reflect the customer's individual intent.

[0051] To further identify customer intent trends within a specific time period, the system introduces a time window mechanism, limiting customer lead data to a sliding time window for feature statistics and aggregation. The time window can be a fixed length (e.g., 7 days, 30 days) or behavior-driven (e.g., the last 5 interactions). Within this timeframe, the system performs cumulative frequency statistics, behavior sequence modeling, or click path analysis on behavioral fields, extracting feature indicators such as "high-frequency interest categories," "repeated inquiry content," and "behavioral path depth," forming a set of customer intent features for the current period.

[0052] The individual intent features generated through the above methods can dynamically reflect the focus and behavioral trends of customer entities, showcasing the potential match between customers and business objectives across multiple dimensions. The effect of this step is that it not only improves the behavioral resolution of customer profiles but also provides semantically rich and structurally clear input for subsequent customer relationship graph construction and feature vector aggregation. For example, if a customer's lead data shows that they have repeatedly browsed health insurance product pages, clicked the "free consultation" button, and submitted a questionnaire within the past 7 days, the system, after extracting their individual intent features, will include tagged features such as "High Intent Focus: Health Insurance," "High Behavioral Depth," and "Recent Consultation Activity," thereby effectively revealing their current business preferences and conversion potential.

[0053] S104. Obtain the association relationships between multiple customer entities, and construct a customer relationship graph based on the association relationships and the customer lead data; In implementing step S104, to gain a more comprehensive understanding of the social structure, business communication paths, and potential influencing relationships behind customer behavior, it is necessary to obtain the relationships between multiple customer entities based on the customer lead data, and construct a customer relationship graph accordingly. A customer relationship graph is a structured graph data constructed with customer entities as nodes and certain or high-probability relationships between customers as edges. It is used to reveal the networked distribution and dynamic interaction relationships of customers in actual business or social environments. The construction of the relationship graph not only helps enrich the dimensions of customer feature expression but also provides structural enhancement support for subsequent feature aggregation, group modeling, and tag generation.

[0054] The relationships between customer entities can originate from various data clues, such as invitation relationships, shared contact methods, shared devices, transactions, belonging to the same company, social interactions, geographical proximity, and similar behaviors. In implementation, the system first extracts potentially meaningful combinations of fields from the customer lead data, such as "mapped referrer's mobile phone number to the referred user's mobile phone number," "multiple customers logging in using the same IP address," "account registration with the same corporate email suffix," and "multiple account behaviors with the same terminal device ID." By constructing a rule base and a decision model, these fields are matched and confidence scores are calculated to form a set of candidate edges between customer entity pairs.

[0055] Based on the candidate edge set, the system introduces a correlation metric to assign weights to the edges between each pair of customer entities. The correlation metric can be calculated based on factors such as interaction frequency, number of shared behaviors, overlap of behavior time, and path similarity, ultimately forming a weighted edge structure. Higher weights indicate stronger behavioral synergy or structural similarity between customers. The weight calculation model can employ graph embedding-based similarity scoring algorithms, such as Node2Vec and LINE, or probabilistic relationship models based on Bayesian inference, to improve edge weight accuracy and control graph sparsity.

[0056] After completing the entity node extraction and relation edge construction, the system organizes all customer entities and related edges into graph-structured data, which is then stored and managed using a graph database (such as Neo4j) or a graph computing engine (such as GraphX ​​or DGL) to construct a complete customer relationship graph. This graph is queryable, iterable, and scalable, supporting subsequent tasks such as graph mining, community discovery, and customer influence assessment based on structural features.

[0057] By constructing a customer relationship graph, the problem of customer lead data containing only individual data and lacking structural context can be effectively addressed, further improving the expressive dimensions of customer profiles and the accuracy of behavioral reasoning. For example, if customer A and customer B have no obvious similarities in the lead data, but the graph construction reveals that they have logged in multiple times using the same device within the past month and both appeared in a group-buying activity, the system can establish a high-weight edge in the graph, marked as "same device + same activity participation," and introduce this structural relationship into the subsequent model as a behavioral enhancement factor, improving the synergy and predictive accuracy of customer intent recognition and status assessment.

[0058] S105. Based on the customer relationship graph, the individual intent features are aggregated and enhanced to obtain the target feature vector of the customer entity, and the target feature vector is used to characterize the current comprehensive state of the customer entity; In one specific embodiment, the aggregation enhancement process based on the customer relationship graph can be implemented using a graph neural network (GNN). Here, individual intent features (e.g., vectors generated by the BERT model) serve as the initial features of customer nodes. The GNN updates the feature representation of each customer node by propagating and aggregating information from neighboring nodes on the customer relationship graph (e.g., stored in Neo4j). The resulting target feature vector integrates both its own deep semantics and the influence of its social or relational circles. This hybrid model combining BERT and GNN enables feature enhancement driven by both semantics and relationships, significantly improving the expressive power and accuracy of the target feature vector. To more comprehensively depict the current overall state of customer entities, structured contextual information needs to be introduced on top of the original individual intent features, and feature aggregation enhancement is achieved through the customer relationship graph. Since customers are often not isolated in actual business scenarios, their behavioral patterns and preference characteristics may be influenced by other customer entities with whom they have social, business, or behavioral relationships. Therefore, relying solely on individual data easily overlooks group synergy effects or anomaly identification capabilities. To this end, based on the direct relationships between customer entities provided in the customer relationship graph, the system combines and aggregates the individual intent features of customer entities adjacent to the target customer entity to form neighborhood common features, and analyzes the expression intensity of the neighborhood features to identify which feature components in the target customer entity have significant uniqueness. By judging the scarcity of individual feature expression in the neighborhood, unique feature components reflecting differentiated customer behavior can be further screened out, and a target feature vector can be constructed based on this, thereby achieving a more discriminative and semantically deep expression of the current comprehensive state of the customer entity. Specifically, this may include the following steps: The individual intent features of adjacent customer entities that have a direct relationship with the target customer entity in the customer relationship graph are combined into neighborhood common features, where the target customer entity is any of the customer entities in the customer relationship graph; Calculate the neighborhood expression intensity corresponding to each first feature component in the common neighborhood features; For any second feature component in the individual intent features, find the neighborhood expression strength corresponding to the second feature component in the neighborhood common features; If the neighborhood expression intensity is less than a preset expression intensity threshold, or if there is no corresponding neighborhood expression intensity for the second feature component in the common features of the neighborhood, then the second feature component is determined as a unique feature component of the target customer entity. The target feature vector is determined based on the unique feature components.

[0059] In the process of feature enhancement for target customer entities in a customer relationship graph, it is necessary to fully utilize the individual intent features of adjacent customer entities connected in the graph structure to uncover potential group behavior patterns and identify individual differences. To this end, the system uniformly extracts and combines the individual intent features of adjacent customer entities directly related to the target customer entity in the customer relationship graph, forming neighborhood common features. Neighborhood common features refer to the set of intent features exhibited by the neighboring nodes of the target customer entity within a specific time window, possessing a certain degree of behavioral stability and structural consistency, and reflecting the behavioral tendencies of a certain customer group. The system traverses the adjacent nodes of the target customer entity in the graph, extracts the individual intent feature vector corresponding to each neighbor, and performs vector-level merging and frequency statistics on all neighbor feature vectors to form a feature distribution within the neighborhood.

[0060] After constructing the common features of the neighborhood, it is necessary to further calculate the neighborhood expression strength of each first feature component. Neighborhood expression strength refers to the frequency or weight of a specific feature component appearing in the common features of the neighborhood. A higher value indicates that the feature is more common among neighboring customers, possibly just part of group behavior rather than a personalized expression of the target customer entity. In calculating the neighborhood expression strength, the system aggregates and statistically analyzes the individual intent features of neighboring customers based on the edge weight relationship between the target customer entity and its neighboring customer entities in the graph structure, thereby quantifying the representativeness of each feature component within the neighborhood. Specifically, the system first extracts the individual intent features of all neighboring customer entities directly related to the target customer entity, and statistically analyzes the frequency and distribution range of each feature component in the neighborhood. Second, it introduces an edge weight factor as an importance indicator of neighboring nodes, weighting and accumulating the edge weight values ​​of each neighboring customer entity exhibiting that feature component. Furthermore, if the individual intent feature includes feature weights (such as behavioral intensity, feature confidence, etc.), the neighborhood expression strength of that feature component will be multiplied by the expression weight of that neighboring node on that feature. Ultimately, the neighborhood expression strength can be expressed as: Neighborhood Expression Strength = ∑(Edge Weight × Feature Weight), where the summation covers all neighbor nodes possessing that feature component. Through this weighting mechanism, the system not only considers the prevalence of features within the neighborhood but also integrates the actual influence of neighbor nodes on the target customer entity. This allows for a more accurate identification of which features are common to the customer group and which are unique to the target customer, providing a reliable basis for subsequent feature enhancement and tag generation. For example, if the feature "frequent browsing of health insurance products" appears three times among five neighbors, and the edge weight between the neighbor and the target entity is high, then the neighborhood expression strength of this feature will be relatively high.

[0061] Subsequently, for any second feature component in the individual intent characteristics of the target customer entity, the system compares it against the aforementioned common neighborhood features to determine whether the second feature component also appears in the neighborhood and obtains its corresponding neighborhood expression strength. If the feature component cannot be found in the common neighborhood features, it means that neighboring customers do not exhibit similar behavior, and the neighborhood expression strength is set to zero. The purpose of this operation is to identify behavioral features that are individually different and uncommon among neighboring customers, in order to perform reinforcement modeling.

[0062] Furthermore, to select truly representative individual feature components, the system compares the neighborhood expression strength of each second feature component with a preset expression strength threshold. If the neighborhood expression strength of a feature component is less than the threshold, it indicates that the feature is weak in the neighborhood and has strong individual uniqueness, thus being identified as a unique feature component of the target customer entity. The identification of unique feature components is significant, as it highlights the customer's distinctive behavioral patterns within the group, helps construct more discriminative target feature vectors, and thereby improves the accuracy of customer status judgment and the targeting of label generation.

[0063] For example, in an insurance platform, if a target customer entity's individual intent characteristics include the feature "visited health insurance product comparison pages for 5 consecutive days," and only one of its five adjacent customers in the graph has exhibited similar behavior, then the neighborhood expression strength of this feature is low, below the system's set threshold of 0.3. Therefore, it is automatically marked as a unique feature component by the system. Ultimately, this feature is used to construct the target customer entity's target feature vector, reflecting its relatively independent and continuous attention to health insurance products, which helps the system subsequently identify it as a high-potential conversion customer.

[0064] Through the above processing, the system not only integrates structural information and individual behavior, but also realizes differentiated expression of customer characteristics, avoiding the problem of customer status ambiguity caused by the homogenization of group behavior, and providing a more robust data foundation for the accurate generation of dynamic tags.

[0065] After identifying the unique feature components of the target customer entity, in order to construct a target feature vector that truly reflects customer differences and dominant behavioral intentions, it is necessary to fuse individual intention features and neighborhood common features based on these unique feature components. Since customer behavior often includes both group characteristics consistent with neighboring customers and personalized characteristics unique to themselves, simply concatenating these two types of features may mask the importance of individual features, leading to a final target feature representation biased towards group averages and losing discriminative power. Therefore, during the construction of the target feature vector, the system determines whether the individual intention features contain unique feature components. If so, its weight in the fused features is strengthened to ensure the dominant role of this feature in the overall feature representation; conversely, if the individual intention features are completely covered by neighborhood common features, equal-weighted fusion can be performed directly. Specifically, this may include the following steps: If the unique feature component exists in the individual intent feature, the neighborhood common feature and the individual intent feature are combined into a preliminary fusion feature, and the expression weight of the unique feature component in the preliminary fusion feature is increased based on a preset weight adjustment rule to obtain the target feature vector; If the unique feature component is not present in the individual intent feature, then the neighborhood common feature and the individual intent feature are combined to obtain the target feature vector.

[0066] In determining the target feature vector based on unique feature components, the system needs to determine whether the individual intent features of the target customer entity contain previously identified unique feature components. The core purpose of this determination is to differentiate the importance of individual differences in feature fusion, in order to achieve a more discernible expression of customer state. When unique feature components are indeed present in the individual intent features, it indicates that the customer exhibits a unique tendency in behavior that distinguishes them from their neighboring customers, possessing high individual identification value. To prevent this difference from being masked by common features in the neighborhood, the system combines common features in the neighborhood with individual intent features when constructing the target feature vector, generating preliminary fused features, and then boosts the weight of the unique feature components in the fused features based on a preset weight adjustment rule. This weight adjustment rule can employ exponential weighting, normalization readjustment, or feature weight allocation strategies based on attention mechanisms, so that the unique feature components obtain higher representation intensity in the final target feature vector, thereby highlighting the customer's individualized behavioral characteristics. The preset weight adjustment rules include: multiplying the weight value of the unique feature component by a preset enhancement coefficient (e.g., between 1.2 and 2.0), or using an attention mechanism during feature fusion to assign a higher attention score to the unique feature component. For example, if "frequent browsing of health insurance comparison pages" is identified as a unique feature component with an original feature weight of 0.4, the system can increase its expression weight in the fused features to 0.7 to strengthen the dominance of this feature in building customer profiles, thereby improving the model's ability to identify customers with specific behaviors.

[0067] When the individual intent features of a target customer entity do not contain unique feature components, it means that its current behavior has high commonality among neighboring customer groups and does not show significant individual differences. In this case, the system does not need to perform differentiated weighting on the feature components, but directly fuses the individual intent features with the common features of the neighborhood with equal or weighted values ​​to form the final target feature vector. To maintain the stability of the fused features, the system adopts feature normalization and dimension alignment strategies during the combination process to ensure that features from different sources are comparable in the same vector space. The fusion method can take the form of vector concatenation, average pooling, or weighted addition, and the specific strategy can be configured according to the needs of different business models. For example, if both the individual intent features and the common features of the neighborhood contain the feature "recently interested in commercial insurance", the system maintains consistency in the expression weights of this feature during the fusion process to ensure that the final target feature vector can objectively reflect the customer's behavioral status within its group. By using this fusion method that does not increase the weight of individual features, the system can effectively reduce the risk of overfitting to non-differentiated features, enhance the model's ability to understand group behavior, and help with stability control when performing customer clustering and label assignment in the future.

[0068] The different designs of the two paths in the fusion strategy enable the system to have greater flexibility and precision control when dealing with different types of customers. It can highlight individual characteristics while retaining the commonalities of the group, providing a more complete and hierarchical feature expression basis for dynamic tag generation.

[0069] S106. Incremental clustering is performed on the target feature vector within each preset time period to obtain multiple customer clusters, and the dynamic evolution index of each customer cluster is calculated within multiple consecutive preset time periods. The dynamic evolution index is used to characterize the customer lead cycle status of the customer group represented by the customer cluster. In step S106, in order to dynamically track the phased changes in customer group behavior, the system performs incremental clustering on the target feature vector of customer entities within each preset time period to generate multiple customer clusters, and uses these clusters to identify customer groups with similar behavioral characteristics.

[0070] During the incremental clustering of the target feature vectors within each preset time period, the core objective of the system is to promptly capture the changing trends of customer behavior and dynamically maintain the structural division of customer groups to support the accurate generation of subsequent customer tags. Since the target feature vectors of customer entities are continuously updated with the data stream, their behavioral intentions and overall states have certain time-sensitivity and evolutionary characteristics. Therefore, static clustering methods cannot meet the business requirements for real-time performance and continuity. To address this, the system employs an incremental clustering algorithm to cluster the target feature vectors of all customer entities within each preset time period (e.g., hour). Incremental clustering is a clustering method that can quickly absorb new data based on existing clustering results. Its principle is to use the cluster centers from the previous moment as initial reference points and perform local updates based on the current input data, thereby avoiding the resource waste and cluster drift problems caused by training from scratch.

[0071] In practical implementation, the system can use Mini-Batch K-Means, Online DBSCAN, or a custom incremental clustering model based on density or graph structures to calculate the similarity between newly generated or updated target feature vectors in the current period and historical clusters. If the similarity is higher than a set threshold, the feature vector is assigned to the existing cluster; otherwise, a new cluster is generated or a cluster adjustment operation is triggered. During the clustering process, the system also periodically corrects the drift of cluster centers to adapt to the gradual changes in customer features over time. Furthermore, to ensure clustering quality, the system introduces metrics such as intra-cluster compactness and inter-cluster separation for real-time evaluation. When the cluster structure is significantly disturbed, the system can trigger a local re-clustering mechanism to ensure the representativeness and stability of the clusters.

[0072] For example, within an hourly cycle, the system has formed 10 customer clusters. Of the 2000 new customer entities entering the system in the current cycle, 1500 customers have target feature vectors with a similarity higher than 0.85 to existing cluster centers. The system assigns these to their respective clusters. For the remaining 500 customers, due to their more dispersed feature distribution, the system identifies three new behavioral patterns, forming new clusters for each and updating the cluster structure. Through this incremental clustering method, the system can continuously maintain the dynamic structure of the customer group, improve the real-time performance and accuracy of customer tag generation, and provide a structured foundation for subsequent calculations of customer group evolution trends.

[0073] Because customer status is time-sensitive and evolutionary, relying solely on static clustering results cannot reflect the potential changes in customer groups over time. Therefore, it is necessary to compare and analyze clustering results over multiple consecutive time periods to quantify the evolutionary trajectory of customer clusters. To this end, the system introduces a dynamic evolution index as a key parameter to measure the degree of change in customer group status. By analyzing the member inheritance relationships and state vector changes between customer clusters in adjacent time periods, it identifies whether structural changes such as migration, splitting, or merging have occurred in the customer group. Specifically, this may include the following steps: Calculate the member overlap between a first customer cluster within a first preset time period and a second customer cluster within a second preset time period, wherein the first preset time period and the second preset time period are adjacent and the first preset time period precedes the second preset time period; if the member overlap is greater than a preset inheritance threshold, then the first customer cluster and the second customer cluster are determined as a pair of continuous customer clusters; for any continuous customer cluster, determine the core state vector of the first period cluster based on the target feature vectors of all customer entities within the first preset time period; determine the core state vector of the second period cluster based on the target feature vectors of all customer entities within the second preset time period; calculate the state drift between the core state vectors of the first period cluster and the core state vectors of the second period cluster, and determine the state drift as the dynamic evolution index.

[0074] In modeling the dynamic evolution trend of customer groups, the system first needs to identify the correspondence between customer clusters within adjacent time periods to track the state change trajectory of the same customer group at different times. To this end, the system calculates the overlap of members between customer clusters within two consecutive preset time periods. Member overlap is an indicator used to measure the degree of overlap of customer entities in two clusters, typically calculated using the Jaccard coefficient, intersection-union ratio, or overlap rate. It is calculated as the ratio of the number of common customers in the two clusters to the number of customers in the union of the two cluster sets. By setting a predetermined inheritance threshold (e.g., 0.5 or 0.6), the system can determine whether two clusters from different time periods constitute a continuous evolutionary relationship. If the overlap exceeds this threshold, it indicates a high degree of consistency between the two clusters, and the system marks them as a pair of "continuous customer clusters." This determination provides a stable customer group mapping basis for subsequent state drift analysis.

[0075] After identifying a pair of continuous customer clusters, the system needs to calculate their core state vectors for two time periods to further track the state changes of this customer group. The core state vector is a centralized representation of the target feature vectors of all customer entities in the cluster in the feature space. It is typically calculated using weighted centroids or average vectors and represents the dominant behavioral characteristics and focus of the customer group in the current time period. The system normalizes the target feature vectors of all customer entities in the first preset time period and calculates the average or weighted average across all dimensions to obtain the core state vector for the first period cluster. Similarly, the same process is performed on customer entities in the second preset time period to obtain the core state vector for the second period cluster. The generation of the core state vector ensures the comparability and continuity of the customer group's feature-level expression, providing a quantitative basis for subsequent calculations of state changes.

[0076] After obtaining the core state vectors for two cycles, the system further calculates the state drift between them as a dynamic evolution indicator quantifying the magnitude of changes in customer group behavior. The state drift can be calculated using methods such as Euclidean distance, cosine similarity difference, Manhattan distance, or Mahalanobis distance to reflect the degree of shift in the overall behavioral characteristics of the same customer group over different periods. For example, the state drift can be calculated as the Euclidean distance or 1-cos(θ) between two core state vectors, where θ is the angle between the two vectors. A small drift indicates stable group behavior and no significant change in customer intent; a large drift indicates a significant shift in customer group behavior characteristics, possibly due to a shift in product interest, market activity stimulation, or changes in lifecycle stages. The system can automatically label customer groups as "stable," "volatile," or "transitional" based on a drift threshold, providing support for subsequent tag generation and operational strategy formulation.

[0077] For example, assuming the overlap of members in a customer cluster between October and November 2025 is 0.68, exceeding the inheritance threshold of 0.6, the system identifies it as a continuous customer cluster. Next, the system extracts the target feature vectors of all customers within these two periods and calculates their average, obtaining the core state vectors for each period. Further calculation shows the Euclidean distance between these two vectors is 0.83, exceeding the preset state drift threshold of 0.6. Based on this, the system determines that the behavioral characteristics of this customer group have changed significantly, possibly shifting from "insurance product comparison behavior" to "financial product ordering behavior," thus marking this group as "high-conversion stage" customers, providing a high-value reference for subsequent marketing strategies. Through this series of steps, the system achieves full-process quantitative modeling of customer groups from structural identification to behavioral evolution, improving the sensitivity and strategy adaptability of dynamic labels.

[0078] S107. Generate customer tags for each customer entity based on the dynamic evolution indicators.

[0079] After calculating the dynamic evolution indicators for customer clusters, the system further utilizes these indicators to identify behavioral trends and classify the states of individual customer entities, thereby achieving accurate generation of customer tags. Since the dynamic evolution indicators can reflect the drift of behavioral characteristics of customer groups across different time periods, they can serve as an important basis for judging whether a customer's state is stable or has migrated. However, relying solely on state drift may not fully reveal the overall picture of changes in the customer group; therefore, the system also introduces the cluster size change rate as an auxiliary basis for judging whether a customer's affiliation has undergone structural changes. By comprehensively analyzing the state drift and size change of the customer's cluster, the system determines whether the customer's current behavior remains within the stable pattern of their original group or shows a trend of migrating to other behavioral groups. Specifically, this may include the following steps: Obtain the second customer cluster to which the target customer entity belongs within the second preset time period; if there exists a first customer cluster that forms a continuous customer cluster with the second customer cluster, then count the number of first members in the first customer cluster and the number of second members in the second customer cluster, and calculate the scale change rate, which is the ratio of the difference between the number of second members and the number of first members to the number of first members; if the dynamic evolution index is less than or equal to the preset drift threshold, and the absolute value of the scale change rate is less than the preset stability threshold, or, there is no first customer cluster that forms a continuous customer cluster with the second customer cluster, then generate a steady-state affiliation label for the target customer entity as the customer label; otherwise, generate a migration trend label for the target customer entity as the customer label.

[0080] During the dynamic tag generation phase, to accurately identify the current behavioral status of customers and provide actionable tag information, the system first needs to determine the customer cluster to which each target customer entity belongs within the current time period. Since the previous steps have already generated multiple customer clusters using incremental clustering, and each customer entity has been assigned to a specific cluster during the clustering process, the system can directly locate its second customer cluster within the clustering results of the second preset time period based on the target customer entity's identification information. This step ensures accurate identification of customer status during subsequent tag generation, providing a cluster-level behavioral group foundation for judging evolutionary trends.

[0081] After identifying the second customer cluster to which the target customer entity belongs, the system further determines whether there exists a first customer cluster that forms a continuous customer cluster with it. If so, it indicates that the customer cluster has clear temporal continuity, which can be used to calculate the evolution trend of the customer group. Based on this, the system separately counts the number of first members in the first customer cluster and the number of second members in the second customer cluster, and calculates the customer group size change rate between the two time periods. The size change rate is obtained by calculating the ratio of the difference between the number of second members and the number of first members to the number of first members, reflecting the expansion or contraction trend of the customer group in terms of size. By introducing this indicator, the system not only considers the characteristic changes of customer status (represented by dynamic evolution indicators), but also comprehensively considers the structural changes of the customer group in the time dimension, which helps to identify customer migration behavior caused by marketing activities, product promotion, or changes in the market environment. For example, if a customer cluster grows from 500 to 750 people in two periods, the size change rate is 0.5, indicating that the group has attracted a large number of new customers and may be in a stage of explosive growth.

[0082] After acquiring dynamic evolution indicators and scale change rates, the system needs to dynamically assess customer status and generate customer tags. If the state drift (i.e., dynamic evolution indicator) between a pair of consecutive customer clusters is less than or equal to a preset drift threshold, it indicates that the customer group has relatively small changes in behavioral characteristics. Simultaneously, if the absolute value of its scale change rate is also less than a preset stability threshold, it indicates that the group structure is relatively stable. Therefore, it can be determined that the current customer entity is still in a stable state within its original behavioral scenario, and the system will generate a "steady-state attribution tag" for that customer entity. This tag indicates that the customer's behavior has not undergone significant migration and has high behavioral persistence and predictive value. However, if the dynamic evolution indicator exceeds the drift threshold, or the scale change rate deviates significantly from the stability threshold, it indicates that the customer group has undergone significant changes in behavior or structure. The system will generate a "migration trend tag" for that customer entity, suggesting that the customer may be in a critical stage of behavioral transition and requires targeted outreach or personalized recommendations in conjunction with marketing strategies. For example, if customer B belongs to the "insurance consultation customer group" in the previous period and is classified into the "financial management order customer group" in the next period, and the corresponding state drift is 0.75, which exceeds the system's set threshold of 0.6, the system will mark it as a customer with "migration trend" to support strategy optimization.

[0083] Through the aforementioned tag generation process, the system can accurately identify the stability and drift of customer status based on both dynamic evolution indicators and structural changes, achieving highly sensitive customer tag generation at the individual level. This not only improves the real-time nature and accuracy of customer insights but also provides a more forward-looking and operational basis for business strategy decisions.

[0084] Furthermore, to support rapid querying and analysis of massive amounts of customer tags and data, a preferred embodiment of the present invention employs a hybrid storage architecture. In this architecture, the generated customer tags and textual information from customer lead data are stored in an inverted index database (such as Elasticsearch) to leverage its powerful full-text search capabilities, enabling millisecond-level tag filtering and keyword searching. Simultaneously, structured data such as customer behavior details are stored in a columnar storage OLAP database (such as Apache Druid). Druid, through its columnar storage, bitmap indexing, and dimension pre-aggregation features, significantly accelerates complex analytical queries on large-scale datasets (e.g., statistically analyzing the behavioral trends of a certain type of customer group over the past quarter).

[0085] To further improve query performance, this hybrid storage architecture can also introduce query optimization mechanisms. For example, the system can analyze and predict historical query logs through Monte Carlo simulations to identify users' "popular query paths" and pre-calculate and generate materialized views (Top-N views) based on the prediction results. When users initiate queries that match these popular paths, the system can directly return results from the materialized views, thereby avoiding real-time scanning of massive amounts of raw data, improving query speed by several times (e.g., 15 times), and supporting high-concurrency queries (e.g., more than 200 concurrent queries).

[0086] Furthermore, to lower the barrier to entry for the system, embodiments of the present invention can also provide a natural language interaction layer. Users do not need to write complex query statements; they can directly input colloquial queries, such as finding VIP customers who have consulted about product A in the last three months but have not placed an order. The system uses an integrated natural language processing model (e.g., a model built based on the SpaCy library) to convert this natural language query into a machine-executable domain-specific language (DSL) query in real time. This DSL query is then sent to the aforementioned hybrid storage system (Elasticsearch and Druid) for parallel execution, ultimately returning the query results to the user in sub-seconds (e.g., within 100 milliseconds), and may even include visualized relationship graphs, greatly improving the efficiency of data querying and the user experience.

[0087] In a preferred embodiment, to efficiently store, query, and analyze the billions of dynamic customer tags and related massive amounts of customer data generated by the method of this invention, the system can further employ a high-performance hybrid storage and query architecture. This architecture solves the performance bottleneck problem faced by traditional single databases when dealing with complex query scenarios, significantly improving the data insight efficiency for business personnel. Specifically, the hybrid storage architecture includes: Inverted index database: such as Elasticsearch. The final customer tags generated by this invention (such as "steady-state attribution tags" and "migration trend tags"), along with unstructured or semi-structured text information (such as user comments and consultation records) from the customer lead data, are stored in this database. Leveraging its powerful full-text search and multi-dimensional filtering capabilities, it can achieve millisecond-level rapid filtering of any tag combination and keyword; for example, quickly locating all customers with the "migration trend tag" who have recently inquired about "high-end financial products."

[0088] Columnar storage OLAP databases, such as Apache Druid or ClickHouse, store structured behavioral details of customers, such as time-series data like page view logs, click events, and transaction records. Through their columnar storage, bitmap indexes, and pre-aggregation features, they can support ultra-low latency aggregation and analysis queries on large-scale datasets. For example, they can complete complex statistics such as "the average daily activity of a certain customer group over the past quarter" within seconds.

[0089] To further enhance query performance and support high-concurrency access, this architecture can also introduce query optimization mechanisms. For example, the system can predict and identify users' "popular query paths" by performing Monte Carlo simulation analysis on historical query logs, and pre-calculate and generate materialized views (e.g., Top-N views) based on the prediction results. When users subsequently initiate queries that match these popular paths, the system can directly return the pre-calculated results from the materialized views, thereby avoiding real-time scanning of massive amounts of raw data. This can improve the response speed of typical queries by several times (e.g., more than 15 times) and stably support high-concurrency queries (e.g., more than 200 concurrent users).

[0090] Furthermore, to significantly lower the barrier to entry for the system and empower business personnel without a technical background, this embodiment also provides a natural language interaction layer. Users do not need to write complex SQL or DSL queries; they can directly input conversational query commands, such as "find VIP customers who have consulted about product A in the last three months but have not placed an order." The system, through an integrated Natural Language Processing (NLP) model (e.g., built based on the spaCy library or a finely tuned Transformer model), converts this NLP query into a machine-executable Domain-Specific Language (DSL) query in real time. This DSL query is then distributed to the aforementioned hybrid storage system (Elasticsearch and Druid) for parallel execution, ultimately returning accurate query results to the user in sub-seconds (e.g., within 100 milliseconds), and may even return visualized relationship graphs.

[0091] The following uses an e-commerce scenario as an example to comprehensively illustrate the application of this application. Assume that a customer group A's core state vector in the first period is mainly reflected in browsing and comparing prices of "entry-level cameras," and the system labels them with a stable "photography enthusiast - potential purchase" tag. In the second period, because some core members of this group (identified through relationship graphs) purchase "high-end lenses," and through graph aggregation enhancement, the core state vector of the entire group significantly shifts towards "professional photography accessories" (dynamic evolution index > drift threshold), and the group size expands (size change rate > stability threshold). At this time, the system updates the tag of all members in the group (including those who have not yet purchased high-end lenses) to the transition trend tag "photography advancement - high-value conversion," triggering a targeted marketing strategy to push high-profit accessories such as professional tripods and filters to them. This demonstrates the significant advantages of this invention compared to traditional static tags in understanding customer lifecycle transitions and guiding forward-looking marketing.

[0092] Please see Figure 3 This is a schematic diagram of the dynamic tagging system for customer leads in this embodiment of the application.

[0093] It should be noted that, Figure 3 The structure of the dynamic tagging system for customer leads shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of the present invention.

[0094] like Figure 3 As shown, a dynamic tagging system for customer leads includes a central processing unit 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory 302 or a program loaded from a storage section 308 into a random access memory 303, such as performing the methods described in the above embodiments. The random access memory 303 also stores various programs and data required for system operation. The central processing unit 301, the read-only memory 302, and the random access memory 303 are interconnected via a bus 304. An input / output interface 305 is also connected to the bus 304.

[0095] The following components are connected to the input / output interface 305: an input section 306 including audio input devices, push-button switches, etc.; an output section 307 including an LCD display, audio output devices, indicator lights, etc.; a storage section 308 including a hard disk, etc.; and a communication section 309 including a network interface card such as a LAN (Local Area Network) card, modem, etc. The communication section 309 performs communication processing via a network such as the Internet. A drive 310 is also connected to the input / output interface 305 as needed. A removable medium 311, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 310 as needed so that computer programs read from it can be installed into the storage section 308 as needed.

[0096] It should be noted that specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory, flash memory, optical fiber, portable compact disk read-only memory, optical storage devices, magnetic storage devices, or any suitable combination thereof. In this invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0097] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. Each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those shown in the drawings.

[0098] Specifically, the dynamic tagging system for customer leads in this embodiment includes a processor and a memory. The memory stores a computer program, and when the computer program is executed by the processor, it implements the dynamic tagging system method for customer leads provided in the above embodiment.

[0099] In another aspect, the present invention also provides a computer-readable storage medium, which may be included in a dynamic tagging system for customer leads described in the above embodiments; or it may exist independently and not assembled into the dynamic tagging system for customer leads. The storage medium carries one or more computer programs that, when executed by a processor of the dynamic tagging system for customer leads, cause the dynamic tagging system for customer leads to implement the dynamic tagging method for customer leads provided in the above embodiments.

Claims

1. A method for dynamic tagging of customer leads, characterized in that, The method includes: Acquire a customer lead data stream, the customer lead data stream including multiple target customer lead data, the target customer lead data including identity information, customer attribute fields and field values ​​corresponding to the customer attribute fields; Based on the identity information, the customer lead data stream is parsed to generate corresponding customer lead data for each customer entity. The customer lead data includes the target customer attribute field of the customer entity and the target field value corresponding to the target customer attribute field. One customer entity corresponds to one customer lead data. Feature extraction is performed on the target customer attribute field and the target field value in the customer lead data to determine the individual intent features of the customer entity; Obtain the relationships between multiple customer entities, and construct a customer relationship graph based on the relationships and the customer lead data; Based on the customer relationship graph, the individual intent features are aggregated and enhanced to obtain the target feature vector of the customer entity, which is used to characterize the current comprehensive state of the customer entity; The overlap of members between customer clusters is calculated within multiple consecutive preset time periods to identify continuous customer clusters with evolutionary relationships. Based on the target feature vectors of customer entities in the customer clusters in different periods, a core state vector of the periodic cluster is constructed. The state drift between the customer clusters is calculated as a dynamic evolution index. The dynamic evolution index is used to characterize the periodic state of customer leads of the customer group represented by the customer clusters. Analyze the rate of change in the size of the members of the continuous customer cluster, and combine the state drift amount with a preset threshold to determine whether the customer group is in a relatively stable state or a behavioral migration trend, and generate customer tags for the customer entities. The customer tags are either steady-state affiliation tags or migration trend tags.

2. The method according to claim 1, characterized in that, The step of parsing the customer lead data stream based on the identity information and generating corresponding customer lead data for each customer entity specifically includes: Based on the identity information, the association analysis is performed on multiple target customer lead data in the customer lead data stream, and multiple target customer lead data belonging to the same customer entity are combined into a target customer lead dataset, with each target customer lead dataset corresponding to a customer entity; The data in the target customer lead dataset corresponding to the customer entity is deduplicated to obtain the customer lead data of the customer entity.

3. The method according to claim 2, characterized in that, The step of deduplicating the data in the target customer lead dataset corresponding to the customer entity to obtain the customer lead data of the customer entity specifically includes: If the same target customer attribute field in the target customer lead dataset has multiple conflicting field values, and the target customer lead data corresponding to the conflicting field values ​​meets the preset conflict deadlock condition, then the historical customer lead data of the customer entity within the preset historical time period is obtained, and a time decay weight is set for the historical customer lead data. The consistency score is obtained by comparing the multiple conflicting field values ​​with the historical field values ​​corresponding to the target customer attribute field, which have the time decay weight. The conflict field value with the highest consistency score is determined as the target field value of the target customer attribute field in the customer lead data, thus obtaining the customer lead data.

4. The method according to claim 1, characterized in that, The step of aggregating and enhancing the individual intent features based on the customer relationship graph to obtain the target feature vector of the customer entity specifically includes: The individual intent features of adjacent customer entities that have a direct relationship with the target customer entity in the customer relationship graph are combined into neighborhood common features, where the target customer entity is any of the customer entities in the customer relationship graph; Calculate the neighborhood expression intensity corresponding to each first feature component in the common neighborhood features; For any second feature component in the individual intent features, find the neighborhood expression strength corresponding to the second feature component in the neighborhood common features; If the neighborhood expression intensity is less than a preset expression intensity threshold, or if there is no corresponding neighborhood expression intensity for the second feature component in the common features of the neighborhood, then the second feature component is determined as a unique feature component of the target customer entity. The target feature vector is determined based on the unique feature components.

5. The method according to claim 4, characterized in that, The determination of the target feature vector based on the unique feature components specifically includes: If the unique feature component exists in the individual intent feature, the neighborhood common feature and the individual intent feature are combined into a preliminary fusion feature, and the expression weight of the unique feature component in the preliminary fusion feature is increased based on a preset weight adjustment rule to obtain the target feature vector; If the unique feature component is not present in the individual intent feature, then the neighborhood common feature and the individual intent feature are combined to obtain the target feature vector.

6. The method according to claim 1, characterized in that, The process of calculating the overlap of members among customer clusters within multiple consecutive preset time periods, identifying continuous customer clusters with evolutionary relationships, and constructing a core state vector for periodic clusters based on the target feature vectors of customer entities within the customer clusters in different periods, and calculating the state drift between customer clusters as a dynamic evolution indicator, specifically includes: Calculate the member overlap between the first customer cluster within the first preset time period and the second customer cluster within the second preset time period, wherein the first preset time period and the second preset time period are adjacent, and the first preset time period is before the second preset time period. If the overlap of the members is greater than the preset inheritance threshold, then the first customer cluster and the second customer cluster are determined as a pair of continuous customer clusters; For any of the continuous customer clusters, the first period cluster core state vector is determined based on the target feature vectors of all customer entities within the first preset time period. Based on the target feature vectors of all customer entities within the second preset time period, determine the core state vector of the second period cluster; Calculate the state drift between the first periodic cluster core state vector and the second periodic cluster core state vector, and determine the state drift as the dynamic evolution index.

7. The method according to claim 6, characterized in that, The analysis of the rate of change in the size of the continuous customer clusters, combined with the state drift and a preset threshold, determines whether the customer group is in a relatively stable state or exhibiting a behavioral migration trend. Customer tags are then generated for each customer entity. These customer tags are either steady-state affiliation tags or migration trend tags, specifically including: Obtain the second customer cluster to which the target customer entity belongs within the second preset time period; If there exists a first customer cluster that forms the continuous customer cluster with the second customer cluster, then the number of first members of the first customer cluster and the number of second members of the second customer cluster are counted respectively, and the scale change rate is calculated. The scale change rate is the ratio of the difference between the number of second members and the number of first members to the number of first members. If the dynamic evolution index is less than or equal to the preset drift threshold, and the absolute value of the scale change rate is less than the preset stability threshold, or there is no first customer cluster that forms the continuous customer cluster with the second customer cluster, then a steady-state attribution label is generated for the target customer entity as the customer label. Otherwise, a migration trend label is generated for the target customer entity as the customer label.

8. A dynamic tagging system for customer leads, characterized in that, The customer lead dynamic tagging processing system includes: one or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code including computer instructions, and the one or more processors invoke the computer instructions to cause the customer lead dynamic tagging processing system to perform the method as described in any one of claims 1-7.

9. A computer-readable storage medium comprising instructions, characterized in that, When the instruction is executed on the dynamic tagging system for customer leads, the dynamic tagging system for customer leads performs the method as described in any one of claims 1-7.

10. A computer program product, characterized in that, When the computer program product is run on the dynamic tagging system for customer leads, the dynamic tagging system for customer leads performs the method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Retrieval enhancement generation-based verbal skill generation method and device, equipment and storage medium

    CN120670572A

  • Method and system for searching for and navigating to user content and other user experience pages in a financial management system with a customer self-service system for the financial management system

    US20180108092A1