A method and system for building and expanding a social network of successful customers

By deeply preprocessing and semantic analysis of unstructured data in customer social networks, combining structured data to calculate association indicators, and building a multi-dimensional correlation matrix and social network diagram, the problem of insufficient utilization of unstructured data in the existing technology is solved, and a more accurate and in-depth construction of customer relationship networks is achieved.

CN119578422BActive Publication Date: 2025-06-17XIAMEN CUSTOM ELF TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510144788.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-06-17
Estimated Expiration
2045-02-10

AI Technical Summary

Technical Problem

The prior art lacks the full utilization of unstructured data when building customer social networks, especially in scenarios based on text or non-traditional interactive data, resulting in insufficient mining of complex relationships between customers.

Method used

By obtaining the original data set of the target customer, including structured and unstructured data, preprocessing, converting the unstructured data into standardized text, applying an embedded vector model to convert the text into high-dimensional vector representation, calculating transaction frequency correlation and interaction semantic similarity, building a multi-dimensional correlation matrix and social network graph, and calculating central indicators to identify high-value potential customers.

Benefits of technology

Effective utilization of structured and unstructured data enhances the semantic depth and accuracy of customer relationship networks, improves the accuracy and expression ability of social network graphs, and can more accurately identify high-value customer nodes, overcoming the problem of insufficient utilization of unstructured data in the existing technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119578422B_ABST
    Figure CN119578422B_ABST
Patent Text Reader

Abstract

The present invention provides a method and system for building and expanding a social network of a closed customer, and relates to the technical field of data processing. The method comprises: obtaining an original data set of a target customer; preprocessing unstructured data in the original data set; calculating the association index between the target customer and other customers respectively according to a set of structured data and unstructured data feature vectors; building a multidimensional association matrix between the target customer and other customers according to the association index; converting the element values ​​in the multidimensional association matrix into edge weights, and according to a preset edge weight threshold, eliminating the relationship with an edge weight lower than the threshold, building a graph structure including nodes and edges, and generating a customer social network graph; calculating the centrality index of each customer node according to the customer social network graph; and selecting the top N customer nodes as high-value potential customers according to the centrality index sorting. The present invention improves the autonomy and accuracy of social network management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a method and system for constructing and expanding a social network of transaction customers. Background Art

[0002] With the rapid development of big data and artificial intelligence technologies, enterprises are increasingly relying on social network analysis technology to explore potential customer relationships and provide decision-making basis for marketing strategies. At present, existing technologies collect and analyze customer data and use social network construction algorithms to build relationship networks between customers. These methods usually rely on information such as customer historical behavior, transaction records, and interaction data. Specifically, the system will first collect basic customer data, and generate a correlation matrix between customers through association analysis or machine learning models, and then display the direct and indirect relationships between customers in the form of a graph structure.

[0003] However, existing technologies lack full utilization of unstructured data, especially in scenarios where customer relationship networks are constructed based on text or non-traditional interaction data, such as comment content or non-standard communication data on social media. For example, when companies identify potential customer relationships through customers' social media interactions, such as comments, likes, and reposts, existing technologies usually simply rely on keyword matching or preset rules, ignoring the semantic information implicit in these unstructured data. This approach leads to insufficient mining of complex relationships between customers, and may not accurately identify high-value potential customer nodes, thereby limiting the scalability and practicality of social networks. Summary of the invention

[0004] The purpose of the present invention is to provide a method and system for building and expanding a social network of successful customers, aiming to solve the problems mentioned in the background technology.

[0005] In order to solve the above technical problems, the technical solution of the present invention is as follows:

[0006] In a first aspect, a method for building and expanding a social network of closed customers is provided, the method comprising:

[0007] Obtaining an original data set of target customers, wherein the original data set includes structured data and unstructured data, wherein the unstructured data includes comment content and information platform interaction data;

[0008] Preprocess the unstructured data in the original dataset, including:

[0009] Convert comment content and information platform interaction data into standardized text;

[0010] Apply the text segmentation method to segment the standardized text and generate word sequence data;

[0011] Apply the embedded vectorization model to convert the word sequence data into high-dimensional vector representations, obtaining a set of unstructured data feature vectors;

[0012] According to the structured data and the set of unstructured data feature vectors, calculate the association metrics between the target customer and other customers respectively. The association metrics include the transaction frequency correlation degree and the interactive semantic similarity; where:

[0013] Calculate the transaction frequency correlation degree based on the structured data;

[0014] Calculate the interactive semantic similarity based on the set of unstructured data feature vectors;

[0015] According to the association metrics, construct a multi-dimensional association matrix between the target customer and other customers, where the rows represent the target customer, the columns represent other customers, and the element values of the matrix are the weighted combination values of the association metrics;

[0016] Convert the element values in the multi-dimensional association matrix into edge weights, and according to the preset edge weight threshold, eliminate the relationships with edge weights lower than the threshold, construct a graph structure including nodes and edges, and generate a customer social network graph, where the nodes represent customers and the edges represent the association relationships between customers;

[0017] According to the customer social network graph, calculate the centrality metrics of each customer node;

[0018] Sort according to the centrality metrics and select the top N customer nodes as high-value potential customers.

[0019] Preferably, convert the comment content and information platform interaction data into standardized text, including:

[0020] Perform syntactic parsing on the comment content, extract the subject, predicate, and core semantic relationships in the syntactic structure, and generate text data;

[0021] According to the interaction types of the information platform interaction data, assign preset type weight tags to different interaction data, and embed the preset type weight tags into the text data to generate weighted text data;

[0022] Decompose the weighted text data into sentences, eliminate special symbols and meaningless phrases, and rearrange the sentence structure according to the semantic hierarchy to generate standardized text.

[0023] Preferably, apply the text tokenization method to perform tokenization on the standardized text to generate word sequence data, including:

[0024] Analyze the semantic context of the standardized text to identify long and short phrases in the text;

[0025] Attach semantic weight tags to long and short phrases according to the word frequency statistics value, semantic relevance score, and the dependence strength of long and short phrases in the context to generate participle semantic data. The calculation formula for the semantic weight is:

[0026] , where is the semantic weight of the th long and short phrase in the word sequence, is the number of contexts containing the th long and short phrase, is the term frequency-inverse document frequency value of the th long and short phrase in the th context, is the syntactic dependence degree of the th long and short phrase in the th context, , is the depth value of the th long and short phrase in the syntax tree, is the context influence factor, , and respectively represent the th dimension value of the embedding vector of the th long and short phrase in the th context and the th context, is the dimension number of the embedding vector, is the term frequency-inverse document frequency value of the th long and short phrase in the th context, is a constant;

[0027] Generate word frequency distribution data according to the participle semantic data, screen out low-frequency words for classification, and generate feature keyword data;

[0028] Construct preliminary word sequence data according to the feature keyword data, and optimize the arrangement according to the semantic weight sorting and context dependence relationship to generate word sequence data.

[0029] Preferably, apply an embedded vectorization model to convert the word sequence data into a high-dimensional vector representation to obtain a set of unstructured data feature vectors, including:

[0030] Perform vectorization processing on each word in the word sequence data to generate a word vector matrix;

[0031] Calculate the dependence degree between adjacent words according to the word vector matrix, dynamically adjust the dependence degree through a weighting function, and update the word vectors to generate a set of weighted word vectors;

[0032] Use the attention mechanism to aggregate the weighted word vector set to generate a sentence vector set;

[0033] The vector dimensions of the sentence vector set are unified and normalized to generate a set of unstructured data feature vectors.

[0034] Preferably, calculating the transaction frequency correlation based on structured data includes:

[0035] Extract the structured transaction record data of the target customers based on the structured data to generate transaction behavior data, which includes transaction time, transaction amount and transaction frequency;

[0036] Calculate the transaction interaction frequency between target customers and other customers based on transaction behavior data and generate a transaction frequency matrix;

[0037] According to the transaction frequency matrix, the transaction behaviors in different time periods are weighted to generate a weighted transaction frequency matrix;

[0038] The values ​​in the weighted transaction frequency matrix are normalized, and the correlation values ​​with low transaction frequency are eliminated by presetting the transaction frequency threshold to generate a transaction frequency correlation data set; the calculation formula for transaction frequency correlation is:

[0039] ,in, For target customers With other customers The transaction frequency correlation of For time period Internal target customers With other customers The number of transactions, For time period Internal target customers With other customers The total transaction amount, and is the weight coefficient, is the adjustment coefficient, The length of the time segment.

[0040] Preferably, the interactive semantic similarity is calculated based on the unstructured data feature vector set, including:

[0041] Obtain unstructured data feature vectors of target customers and unstructured data feature vectors of other customers, and generate a target vector set and a comparison vector set;

[0042] The initial similarity value between the target vector set and the comparison vector set is calculated by the cosine similarity formula to generate an initial semantic similarity matrix;

[0043] Normalize the values in the initial semantic similarity matrix to generate a normalized semantic similarity matrix;

[0044] Calculate the dynamic weights based on the similarity values in the normalized semantic similarity matrix and the semantic features of the target customer;

[0045] Apply the dynamic weights to each item in the normalized semantic similarity matrix to generate a dynamically adjusted semantic similarity matrix;

[0046] Perform grouped calculations based on the semantic association strength of the dynamically adjusted semantic similarity matrix to generate an interactive semantic similarity dataset; The formula for interactive semantic similarity is:

[0047] , where is the target customer and other customers 's interactive semantic similarity, is the total number of long and short phrases, is the global semantic weight of the th long and short phrase, is the target customer and other customers on the th long and short phrase's cosine similarity, , where and are respectively the th dimensional values of the embedding vectors of the target customer and other customers on the th long and short phrase, and are respectively the Euclidean norms of the embedding vectors of the target customer and other customers on the th long and short phrase, is the dynamic weight of the target customer on the th long and short phrase, , is the target customer on the th long and short phrase's term frequency - inverse document frequency value, is the target customer on the th long and short phrase's term frequency - inverse document frequency value.

[0048] Preferably, according to the customer social network graph, calculate the centrality index of each customer node, including:

[0049] Extract the connection information and edge weights of each node according to the customer social network graph, calculate the node degree centrality index of each node, and generate a node degree centrality index data set;

[0050] Conduct path analysis on the customer social network graph, calculate the betweenness centrality index of each node, and generate a betweenness centrality index data set of nodes;

[0051] According to the node degree centrality index data set and the betweenness centrality index data set of nodes, eliminate the nodes whose node degree centrality index and betweenness centrality index are both lower than the preset centrality threshold, and generate an optimized node data set;

[0052] According to the optimized node data set, calculate the closeness between each node and the target stage, and generate a node centrality index data set; The calculation formula of the node centrality index is:

[0053] , where is the node centrality index of node , is the degree centrality index of node , , where is the edge weight between node and node , is the set of nodes connected to node , is the betweenness centrality index of node , , where is the number of shortest paths from node to node , is the number of shortest paths passing through node when node reaches node , is the closeness centrality index of node , , where is the shortest path distance between node and node , , and are weight coefficients.

[0054] In the second aspect, a system for constructing and expanding the social network of transaction customers, the system includes:

[0055] A data acquisition module for acquiring the original data set of target customers, where the original data set includes structured data and unstructured data, and the unstructured data includes comment content and information platform interaction data;

[0056] A preprocessing module for preprocessing the unstructured data in the original data set, including:

[0057] Converting the comment content and information platform interaction data into standardized text;

[0058] Applying a text tokenization method to tokenize the standardized text to generate word sequence data;

[0059] Applying an embedded vectorization model to convert the word sequence data into a high-dimensional vector representation to obtain a set of unstructured data feature vectors;

[0060] An association index calculation module for calculating the association indexes between the target customer and other customers according to the structured data and the set of unstructured data feature vectors, including:

[0061] Calculating the transaction frequency correlation based on the structured data;

[0062] Calculating the interaction semantic similarity based on the set of unstructured data feature vectors;

[0063] An association matrix construction module for constructing a multi-dimensional association matrix of the target customer and other customers according to the association indexes, where the rows represent the target customer, the columns represent other customers, and the element values of the matrix are the weighted combination values of the association indexes;

[0064] A social network graph construction module for converting the element values in the multi-dimensional association matrix into edge weights, and removing the relationships with edge weights lower than the threshold according to a preset edge weight threshold, constructing a graph structure including nodes and edges, and generating a customer social network graph, where the nodes represent customers and the edges represent the association relationships between customers;

[0065] A centrality index calculation module for calculating the centrality indexes of each customer node according to the customer social network graph;

[0066] A customer expansion module for selecting the top N customer nodes with the highest centrality indexes as high-value potential customers according to the centrality index ranking.

[0067] The above solution of the present invention has at least the following beneficial effects:

[0068] First, this method not only utilizes traditional structured data, such as transaction records and customer historical behaviors, to construct association relationships, but also conducts in-depth semantic mining on unstructured data, such as comment content and information platform interaction data. It transforms the comment content and interaction data into standardized text, and extracts the subject, predicate, and core semantic relationships of the text by means of syntactic parsing and semantic analysis, solving the deficiency of simply relying on keyword matching in the existing technology. Through semantic hierarchical reorganization and weight embedding, the utilization of implicit semantic information in unstructured data is enhanced, ensuring that the generated customer relationship network has greater semantic depth and accuracy.

[0069] Secondly, in terms of data feature expression, this method transforms unstructured data into high-dimensional vector representations through steps such as word segmentation and embedded vectorization, and combines it with structured data to generate a multi-dimensional association matrix. This matrix not only includes the traditional transaction frequency correlation degree, but also incorporates the interactive semantic similarity calculated based on unstructured data. Through this multi-combination of association metrics, the relationship strength between customers can be comprehensively described and reflected in the form of edge weights in the customer social network graph. Compared with the existing technology, this method of constructing a multi-dimensional association matrix breaks through the limitation of a single data source, significantly improving the accuracy and expressive ability of the social network graph.

[0070] Thirdly, in the process of constructing the customer social network graph, this method eliminates low-value edges by presetting edge weight thresholds, thereby ensuring that the generated graph structure is more compact and efficient. At the same time, through the calculation of centrality metrics, this method can not only intuitively display the direct relationships between customers, but also deeply explore the potential relationships between customers. Compared with the customer network generated by the existing technology through simple association analysis only, the customer social network graph generated by this method is richer in relationship recognition and scalability.

[0071] Finally, this method can also quickly screen out the top N high-value potential customers through the centrality ranking of the customer social network graph. This function is not only applicable to the mining of potential customers in the precise marketing of enterprises, but also can provide strong support for other scenarios, such as customer relationship management and customized service recommendation. Especially with the support of diverse data sources, this method can significantly improve the accuracy of identifying high-value customer nodes, overcoming the identification deficiencies caused by keyword matching or simple rules in the existing technology.

[0072] In summary, by fully combining structured data and unstructured data and deeply mining semantic information, this method greatly improves the construction quality of the customer social network, breaking through the bottleneck of insufficient utilization of unstructured data in the existing technology. At the same time, through multi-dimensional association metrics and an optimized social network graph, this method can provide more comprehensive and accurate support for enterprises in customer relationship mining and marketing decision-making, significantly enhancing the practicality and scalability of the system. Description of the Drawings

[0073] Figure 1 It is a flowchart of a method for constructing and expanding a social network of transaction customers provided by an embodiment of the present invention. Detailed Embodiment

[0074] Hereinafter, exemplary embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be completely conveyed to those skilled in the art.

[0075] As Figure 1 shown, an embodiment of the present invention proposes a method for constructing and expanding a social network of transaction customers, and the method includes:

[0076] S100. Obtain an original data set of target customers, where the original data set includes structured data and unstructured data, and the unstructured data includes comment content and information platform interaction data;

[0077] S200. Preprocess the unstructured data in the original data set, including:

[0078] Convert the comment content and information platform interaction data into standardized text;

[0079] Apply a text tokenization method to tokenize the standardized text to generate word sequence data;

[0080] Apply an embedded vectorization model to convert the word sequence data into a high-dimensional vector representation to obtain a set of unstructured data feature vectors;

[0081] S300. Calculate the association metrics between the target customer and other customers according to the structured data and the set of unstructured data feature vectors, where the association metrics include the transaction frequency correlation degree and the interaction semantic similarity; where:

[0082] Calculate the transaction frequency correlation degree based on the structured data;

[0083] Calculate the interaction semantic similarity based on the set of unstructured data feature vectors;

[0084] S400. Construct a multi-dimensional association matrix between the target customer and other customers according to the association metrics, where the rows represent the target customer, the columns represent other customers, and the element values of the matrix are the weighted combination values of the association metrics;

[0085] S500. Convert the element values in the multi-dimensional association matrix into edge weights, and based on a preset edge weight threshold, eliminate the relationships with edge weights lower than the threshold, construct a graph structure including nodes and edges, and generate a customer social network graph, where nodes represent customers and edges represent the association relationships between customers;

[0086] S600. Calculate the centrality index of each customer node according to the customer social network graph;

[0087] S700. Select the top N customer nodes with the highest rankings as high-value potential customers according to the ranking of the centrality index.

[0088] In the embodiment of the present invention, through this method, the original data set of the target customers is comprehensively obtained and processed first. The original data set covers two types of structured data and unstructured data. The structured data mainly includes transaction records, which detail information such as the transaction time, transaction amount, and transaction frequency of customers; the unstructured data includes comment content and information platform interaction data, such as likes, comments, and forwards. The diversity of this data source lays the foundation for subsequent comprehensive association calculations.

[0089] In the data preprocessing stage, through text standardization of the unstructured data, the consistency of the data format is ensured, and through text tokenization and vectorization models, the unstructured text is successfully transformed into a feature set represented by high-dimensional vectors. At the same time, using the transaction frequency association model and the interactive semantic similarity model, various association relationships between the target customer and other customers are quantified into a unified association index. Based on these indexes, a multi-dimensional association matrix is constructed, where the rows of the matrix correspond to the target customer, the columns correspond to other customers, and the element values represent the association strength between customers.

[0090] By converting the multi-dimensional association matrix into a graph structure, including nodes and edges, the weights of the edges reflect the association strength between customers, and at the same time, low-weight relationship edges are eliminated through a preset edge weight threshold, and finally, the target customer social network graph is generated. Centrality calculations are performed on the social network graph to identify the centrality index of each customer node. According to the ranking of the centrality index, the top N customer nodes with the highest rankings are selected as high-value potential customers. Through this method, the implicit associations between customers can be accurately mined, a clearer and more efficient customer relationship network graph can be constructed, and strong data support and decision-making basis are provided for subsequent precision marketing.

[0091] More specifically, obtain the original data set of the target customer, the original data set includes structured data and unstructured data, and the unstructured data includes comment content and information platform interaction data, including:

[0092] Structured data refers to data with a clear organizational form, usually existing in the form of tables or relational databases. The structured data in this method mainly includes the transaction records and historical behavior information of customers. Transaction records are the key data sources, which contain field information such as transaction time, transaction amount, and transaction frequency. By processing the structured data, the transaction behavior characteristics between the target customer and other customers can be directly extracted. For example, the transaction frequencies and amounts between a certain customer and the target customer in different time periods can reflect the interaction intensity between the two, providing data support for the calculation of the subsequent transaction frequency correlation degree.

[0093] Unstructured data refers to free-form data that has not been strictly organized, including comment content and information platform interaction data.

[0094] Comment content: This part of the data usually comes from the comment records of the target customer and other customers on social media or online platforms, such as user evaluations of products or services, discussions with other users, etc. These comment contents exist in the form of natural language texts and have not been structured. Direct use may lead to information redundancy or misinterpretation. To effectively utilize the comment content, this method combines natural language processing technologies, such as grammar parsing and semantic analysis, to extract the semantic relationships contained in the text, such as the subject, predicate, and object, providing support for the subsequent calculation of semantic similarity.

[0095] Information platform interaction data: This type of data reflects the interaction behaviors of the target customer and other customers on social platforms, such as likes, comments, forwards, etc. Different interaction behaviors have different importance, so corresponding weight tags need to be set for each behavior. For example, the forward behavior may reflect a stronger customer correlation, so a higher weight can be set, while a simple like behavior may reflect a weaker correlation and can be set with a lower weight. By assigning weight tags to the interaction data and embedding them into the text data, the utilization value of unstructured data is further improved.

[0096] More specifically, according to the correlation indicators, a multi-dimensional correlation matrix of the target customer and other customers is constructed, where the rows represent the target customer, the columns represent other customers, and the element values of the matrix are the weighted combination values of the correlation indicators, including:

[0097] Definition of matrix rows and columns:

[0098] The rows of the multi-dimensional correlation matrix represent the target customer, and the columns represent other customers. The elements of the matrix represent the correlation intensity between the target customer and other customers.

[0099] Calculation of element values:

[0100] The element values of the matrix are the weighted combination values of the correlation indicators, and the calculation formula is as follows:

[0101] , where and are weight coefficients used to balance the impacts of two correlation indicators. is the target customer and other customers 's transaction interaction frequency. is the target customer and other customers 's interactive semantic similarity.

[0102] Normalize the matrix element values to ensure that all values of the matrix are within a unified range, such as [0, 1].

[0103] In a preferred embodiment of the present invention, convert the comment content and information platform interaction data into standardized text, including:

[0104] Perform syntactic parsing on the comment content, extract the subject, predicate, and core semantic relationships in the syntactic structure, and generate text data.

[0105] According to the interaction types of the information platform interaction data, assign preset type weight tags to different interaction data, and embed the preset type weight tags into the text data to generate weighted text data.

[0106] Decompose the weighted text data into sentences, remove special symbols and meaningless phrases, and rearrange the sentence structure according to the semantic level to generate standardized text.

[0107] In the embodiment of the present invention, through a series of natural language processing technologies, comprehensively process the original comment content and interaction data, and successfully realize the conversion of unstructured data into standardized text. First, perform syntactic parsing on the comment content to extract the subject, predicate, and core semantic relationships of each comment. This process effectively filters out redundant information in the comments and enhances the semantic expression ability of the comment content. Secondly, combine the interaction types of the information platform interaction data, such as likes, comments, forwards, etc., with the corresponding preset weight tags, and embed them into the comment text to generate weighted text data. This embedding method effectively integrates the user's interaction behavior and semantic information, further enhancing the depth and expression ability of the data.

[0108] Based on the weighted text data, through the steps of sentence decomposition and reconstruction, remove meaningless symbols and phrases, and rearrange the text sentences according to the semantic level. This process not only optimizes the semantic logic of the text, but also enhances the normativity and parsability of the text. The finally generated standardized text content has a highly consistent formatted structure, providing a high-quality data basis for subsequent semantic analysis and feature extraction.

[0109] More specifically, perform syntactic parsing on the comment content, extract the subject, predicate, and core semantic relationships in the syntactic structure to generate the data of this article, including:

[0110] This embodiment uses a syntactic parsing algorithm in natural language processing (NLP) technology to analyze each sentence of the comment content. Common syntactic parsing tools include Stanford Parser based on statistical models, dependency tree parsing models, etc. These tools can decompose natural language text into a tree-like structure containing syntactic relationships.

[0111] Subject extraction: Use the dependency tree structure generated by syntactic parsing to locate the subject part in the syntactic markers. For example, the "nsubj" tag in the sentence, and extract the corresponding word or phrase of the subject.

[0112] Predicate extraction: Find the corresponding predicate part from the dependency tree structure, such as the "ROOT" tag, and extract the predicate verb and its modifiers.

[0113] Core semantic relationship extraction: Extract the dependency relationship between the subject and the predicate, such as the object, adverbial, etc., to ensure that the generated data can completely express the core semantics of the comment content.

[0114] Recombine the extracted subject, predicate, and semantic relationships into structured text data according to a preset format. For example, for the comment "The user likes the product", the extracted subject is "user", the predicate is "likes", and the semantic relationship is "the product", and the generated data format is a triple of "subject-predicate-semantic relationship", such as: user-likes-the product.

[0115] More specifically, rearrange the sentence structure according to the semantic level, including:

[0116] Based on the results of syntactic parsing, analyze the primary and secondary levels among the subject, predicate, and semantic relationships. For example:

[0117] The part directly describing the core event in the sentence is the main semantic level;

[0118] The subsidiary information describing time, place, conditions is the secondary semantic level.

[0119] Sort the components of the sentence according to the semantic weight, place the main semantic components in the core position of the sentence, and arrange the secondary semantic components in sequence. For example, the original sentence "Last year, the company's users gave positive reviews on the new product" can be rearranged as "The company's users positively reviewed the new product (last year)".

[0120] The rearranged text content retains the core semantics of the original sentence and optimizes the logic and readability of the semantic level.

[0121] In a preferred embodiment of the present invention, a text segmentation method is applied to segment the standardized text to generate word sequence data, including:

[0122] Analyze the semantic context of the standardized text to identify long and short phrases in the text;

[0123] Attach semantic weight tags to the long and short phrases according to the word frequency statistical value, semantic relevance score, and the dependence strength of the long and short phrases in the context to generate segmented word semantic data; the calculation formula of the semantic weight is:

[0124] , where is the semantic weight of the th long and short phrase in the word sequence, is the number of contexts containing the th long and short phrase, is the term frequency-inverse document frequency value of the th long and short phrase in the th context, is the syntactic dependence degree of the th long and short phrase in the th context, , is the depth value of the th long and short phrase in the syntax tree, is the context influence factor, , and respectively represent the th long and short phrase in the th context and the th context, and the th dimensional value of the embedded vector, is the dimensionality of the embedded vector, is the term frequency-inverse document frequency value of the th long and short phrase in the th context, is a constant;

[0125] Generate word frequency distribution data according to the segmented word semantic data, screen out low-frequency words for classification, and generate feature keyword data;

[0126] Construct preliminary word sequence data according to the feature keyword data, and optimize the arrangement according to the semantic weight sorting and context dependence relationship to generate word sequence data.

[0127] In the embodiments of the present invention, semantic context analysis is carried out on standardized texts, and long and short phrases in the texts are accurately identified through intelligent word segmentation technology. On the basis of word segmentation, semantic weight tags are attached to each long and short phrase respectively according to the word frequency statistical value, semantic relevance, and the dependence strength of the long and short phrases in the context. The importance of each long and short phrase is comprehensively calculated through the semantic weight formula, further refining the features of each word segmentation unit in the text. The attached semantic weights not only reflect the relative importance of the word segmentation in the sentence, but also combine the structural depth in the syntax tree, making the generated word segmentation results have high accuracy.

[0128] The generation of word segmentation semantic data further constructs a word frequency distribution table. After the low-frequency vocabulary is removed and classified, a set of optimized feature keyword data is generated. The feature keywords are sorted by semantic weights and arranged into preliminary word sequence data. On this basis, it is optimized in combination with the context dependence relationship. The finally generated word sequence data completely retains the semantic information of the original text, while improving the usability and computational efficiency of the data.

[0129] More specifically, the semantic context of the standardized text is analyzed to identify long and short phrases in the text, including:

[0130] Adopt a phrase recognition algorithm combining statistics and rules:

[0131] Statistical-based method: Use word frequency statistics and mutual information (MI) calculation to identify common word combinations. For example, calculate the ratio of the probability that words A and B appear simultaneously to the probability that they appear separately to measure the importance of their combination.

[0132] Rule-based method: According to grammar rules, such as the noun phrase structure: attributive + noun, directly extract phrases that conform to the rules.

[0133] The classification of long and short phrases is as follows:

[0134] Long phrase: A fixed expression usually composed of multiple words, such as "product user experience", whose meaning is inseparable in the context.

[0135] Phrase: A simple semantic unit composed of one or two words, such as "user", "experience", which can express semantics alone.

[0136] Analyze the context dependence relationship of long and short phrases in the sentence. By analyzing their positions in the sentence and their relationships with other words, the accuracy of phrase extraction is further optimized.

[0137] More specifically, generate word frequency distribution data according to the word segmentation semantic data, screen out low-frequency vocabulary for classification, and generate feature keyword data, including:

[0138] Perform word frequency statistics on the extracted set of long and short phrases, calculate the frequency of each phrase appearing in the text, and generate a word frequency distribution table sorted by frequency.

[0139] According to the set word frequency threshold, such as a word frequency less than 5, filter out the set of low-frequency words.

[0140] Classify the phrases in the set of low-frequency words, for example, aggregate them according to semantic categories:

[0141] Classification 1: Phrases related to user behavior, such as "liking" and "commenting".

[0142] Classification 2: Phrases related to product attributes, such as "performance" and "appearance".

[0143] The filtered set of low-frequency words forms a new set of feature keyword data through classification, further optimizing the semantic feature expression.

[0144] More specifically, construct preliminary word sequence data based on the feature keyword data, and optimize the arrangement according to semantic weight sorting and context dependence relationship to generate word sequence data, including:

[0145] According to the classification results of the feature keyword data, generate preliminary word sequences in turn according to semantic categories. For example, the keywords in category 1 are arranged as "user, interaction, comment", and category 2 is arranged as "performance, appearance, design".

[0146] Assign weights to words using the semantic weight formula.

[0147] Sort the words according to the calculated semantic weights, and the words with higher weights are arranged first.

[0148] On the basis of sorting, optimize the arrangement of words in combination with the context dependence relationship. For example, if "user" and "experience" frequently appear together, then arrange them together to form the phrase "user experience".

[0149] In a preferred embodiment of the present invention, apply an embedded vectorization model to convert the word sequence data into a high-dimensional vector representation to obtain an unstructured data feature vector set, including:

[0150] Perform vectorization processing on each vocabulary in the word sequence data to generate a word vector matrix;

[0151] Calculate the dependence degree between adjacent words according to the word vector matrix, dynamically adjust the dependence degree through a weighting function and update the word vectors to generate a weighted word vector set;

[0152] Use the attention mechanism to aggregate the weighted word vector set to generate a sentence vector set;

[0153] Unify the vector dimensions of the sentence vector set and perform normalization processing to generate a set of unstructured data feature vectors.

[0154] In the embodiments of the present invention, through an embedding model, the word sequence data is converted into a high-dimensional vector representation. First, the context-sensitive embedding model is used to process the word sequence data word by word to generate a high-dimensional word vector matrix containing all vocabulary. On this basis, according to the dependence degree between vocabulary, the word vectors are optimized and adjusted through a dynamic weighting function to generate a set of weighted word vectors.

[0155] Furthermore, the attention mechanism is used to aggregate the weighted word vectors to successfully generate a sentence-level vector set. This method not only retains the overall semantic structure of the sentence but also combines the importance differences between the words inside the sentence. Finally, by unifying the dimensions and normalizing the sentence vectors, a complete set of unstructured data feature vectors is generated. This set not only efficiently reduces the dimension of the original text but also provides highly structured basic data for subsequent semantic analysis and relationship calculation.

[0156] More specifically, each vocabulary in the word sequence data is vectorized to generate a word vector matrix, including:

[0157] The context-sensitive embedding model is used to dynamically generate word vectors according to the context semantics of the vocabulary in the sentence. Common models include:

[0158] Based on bidirectional LSTM, it can capture the dynamic semantics of words in the context.

[0159] Using the Transformer architecture, it supports bidirectional context analysis and is suitable for complex sentence and polysemous word scenarios.

[0160] Taking the word sequence data as the model input, the model extracts the semantic features of the vocabulary layer by layer through multiple deep neural network layers, such as the TransformerEncoder layer.

[0161] For each vocabulary, the model dynamically generates a high-dimensional vector according to its context information.

[0162] Generating word vectors for all the vocabulary in the sequence and arranging these vectors in the order of the words to form a word vector matrix.

[0163] Each row represents the high-dimensional vector of a word, such as 768 dimensions, and each column represents the semantic features of the vocabulary in different dimensions.

[0164] More specifically, calculating the dependence degree between adjacent vocabulary according to the word vector matrix to generate a set of weighted word vectors, including:

[0165] Use a syntax parsing tool, such as Stanford Parser or an open-source tool based on dependency syntax, to analyze the dependency relationships between adjacent words in the word vector matrix.

[0166] The degree of dependency is determined by the semantic distance and syntactic relationship between words. For example, the degree of dependency in a subject-predicate relationship may be higher than that in an attributive modification relationship.

[0167] For each pair of adjacent words, assign a dynamic weight according to their degree of dependency. The range of weight values can be defined based on the importance of syntactic relationships.

[0168] The role of the dynamic weight is to adjust the feature representation of words in the vector. For example, for "like" and "product" in "The user likes the product", if the context of the sentence emphasizes the verb "like", a higher weight is assigned to it.

[0169] According to the assigned weights, adjust the original word vectors to generate a new set of weighted word vectors. Each weighted word vector not only retains the original semantic information but also combines the dependency characteristics of the context, showing higher semantic accuracy.

[0170] More specifically, use the attention mechanism to aggregate the set of weighted word vectors to generate a set of sentence vectors, including:

[0171] The goal of aggregation is to integrate the information in the set of weighted word vectors into a sentence-level vector while retaining the semantic features of important words in the sentence.

[0172] Assign attention weights to each vector in the set of weighted word vectors. The magnitude of the weight is determined by the importance of the word. For example, subject-predicate words are usually more important in the core semantics of a sentence; the importance of modifying words, such as adjectives and adverbs, may be lower.

[0173] The assignment of attention weights is usually calculated by a feed-forward neural network, and the Transformer architecture can be used for specific implementation.

[0174] According to the assigned attention weights, perform a weighted sum on the set of word vectors to integrate all word vectors into a sentence vector of a fixed length.

[0175] For each input sentence data, generate the corresponding sentence vector and store all sentence vectors as a set for subsequent processing.

[0176] More specifically, unify the vector dimensions of the set of sentence vectors and perform normalization processing to generate a set of unstructured data feature vectors, including:

[0177] The dimensions of different sentence vectors may vary due to the length of the input sentence or context characteristics. To ensure consistency in subsequent processing, all sentence vectors need to be adjusted to a unified dimension, for example, 512 dimensions.

[0178] The operation of unifying dimensions can be done through zero padding or dimensionality reduction, such as PCA dimensionality reduction.

[0179] To make the numerical range of the sentence vector uniform, each component in the vector is mapped to a fixed range, such as [0, 1]. Commonly used normalization methods include: maximum and minimum normalization or L2 normalization.

[0180] The normalized sentence vector set is stored as a feature vector set. Each vector in the feature vector set represents the semantic feature of an unstructured data, providing standardized input for subsequent calculations.

[0181] In a preferred embodiment of the present invention, calculating the transaction frequency correlation based on structured data includes:

[0182] Extract the structured transaction record data of the target customers based on the structured data to generate transaction behavior data, which includes transaction time, transaction amount and transaction frequency;

[0183] Calculate the transaction interaction frequency between target customers and other customers based on transaction behavior data and generate a transaction frequency matrix;

[0184] According to the transaction frequency matrix, the transaction behaviors in different time periods are weighted to generate a weighted transaction frequency matrix;

[0185] The values ​​in the weighted transaction frequency matrix are normalized, and the correlation values ​​with low transaction frequency are eliminated by presetting the transaction frequency threshold to generate a transaction frequency correlation data set; the calculation formula for transaction frequency correlation is:

[0186] ,in, For target customers With other customers The transaction frequency correlation of For time period Internal target customers With other customers The number of transactions, For time period Internal target customers With other customers The total transaction amount, and is the weight coefficient, is the adjustment coefficient, The length of the time segment.

[0187] In the embodiments of the present invention, for the transaction record data of target customers, through a series of formulas and data processing, the method successfully quantifies the transaction frequency correlation degree between target customers and other customers. First, transaction behavior information is extracted from structured data, including transaction time, transaction amount, and transaction frequency, and a transaction behavior data set is constructed based on this. Through the transaction frequency formula, the transaction interaction frequency between the target customer and other customers is accurately calculated and quantified as a transaction frequency matrix. The transaction frequency matrix is further combined with a time weight function to assign higher weights to recent transaction behaviors, emphasizing the time sensitivity of dynamic transaction behaviors.

[0188] In addition, by normalizing the data, customer relationship data with low transaction frequencies is eliminated, generating a high-quality transaction frequency correlation data set. The finally generated result can accurately reflect the strength of transaction behaviors between customers, providing reliable data support for the construction of subsequent correlation matrices and social network graphs.

[0189] In a preferred embodiment of the present invention, the interactive semantic similarity is calculated based on the unstructured data feature vector set, including:

[0190] Obtain the unstructured data feature vectors of the target customer and other customers, generating a target vector set and a comparison vector set;

[0191] Calculate the initial similarity value between the target vector set and the comparison vector set through the cosine similarity formula, generating an initial semantic similarity matrix;

[0192] Normalize the values in the initial semantic similarity matrix to generate a normalized semantic similarity matrix;

[0193] Calculate the dynamic weight based on the similarity values in the normalized semantic similarity matrix in combination with the semantic features of the target customer;

[0194] Apply the dynamic weight to each item in the normalized semantic similarity matrix to generate a dynamically adjusted semantic similarity matrix;

[0195] Perform grouped calculations according to the semantic association strength of the dynamically adjusted semantic similarity matrix to generate an interactive semantic similarity data set; The calculation formula for the interactive semantic similarity is:

[0196] , where is the target customer and other customers 's interactive semantic similarity, is the total number of long and short phrases, is the th global semantic weight of the long and short phrase, is the target customer Cosine similarity with other customers at the th long and short phrase, where and are the th dimensional values of the embedding vectors of the target customer and other customers at the th long and short phrase, and are the th Euclidean norms of the embedding vectors of the target customer and other customers at the th long and short phrase, is the dynamic weight of the target customer at the , is the th term frequency - inverse document frequency value of the target customer at the th long and short phrase, is the th term frequency - inverse document frequency value of the target customer

[0197] In the embodiment of the present invention, the method successfully calculates the interactive semantic similarity between the target customer and other customers through the unstructured data feature vector set. First, the feature vectors of the target customer and other customers are respectively constructed into a target vector set and a comparison vector set, and the initial semantic similarity value is calculated through the cosine similarity formula. This stage completes the basic semantic similarity quantification.

[0198] On the basis of the initial semantic similarity, further combined with the dynamic semantic features of the target customer, a set of dynamically adjusted weights is calculated and applied to the normalized semantic similarity matrix to generate a set of dynamically adjusted semantic similarity matrix data. Finally, by performing grouped calculations on the dynamically adjusted matrix, an interactive semantic similarity data set is generated. This data set fully reflects the strength of the semantic association between customers and provides in - depth semantic support for the construction and optimization of the customer social network.

[0199] In a preferred embodiment of the present invention, according to the customer social network graph, the centrality index of each customer node is calculated, including:

[0200] According to the customer social network graph, the connection information and edge weights of each node are extracted, and the node degree centrality index of each node is calculated to generate a node degree centrality index data set;

[0201] Perform path analysis on the customer social network graph, calculate the betweenness centrality index of each node, and generate a dataset of betweenness centrality indices of nodes;

[0202] According to the dataset of degree centrality indices of nodes and the dataset of betweenness centrality indices of nodes, remove the nodes whose degree centrality index and betweenness centrality index are both lower than the preset centrality threshold, and generate an optimized node dataset;

[0203] According to the optimized node dataset, calculate the closeness between each node and the target stage, and generate a dataset of node centrality indices; The calculation formula of the node centrality index is:

[0204] , where is the node centrality index of node , is the degree centrality index of node , , where is the edge weight between node and node , is the set of nodes connected to node , is the betweenness centrality index of node , , where is the number of shortest paths from node to node , is the number of shortest paths passing through node when node reaches node , is the closeness centrality index of node , , where is the shortest path distance between node and node , , and are weight coefficients.

[0205] In the embodiments of the present invention, through the construction of the customer social network graph, the method further calculates the centrality indices of each customer node, including the degree centrality, betweenness centrality, and closeness centrality of the node. The degree centrality of the node directly reflects the connection strength between the node and other nodes, while the betweenness centrality measures the ability of the node to act as a bridge in the customer network. The closeness centrality reflects the importance of the node in the network by calculating the path length between nodes.

[0206] By removing nodes with both degree centrality and betweenness centrality below the threshold, an optimized node dataset was generated. This optimization not only reduces the interference of redundant nodes on the network structure but also significantly enhances the significance of key nodes in the network. Finally, based on the centrality index ranking of customer nodes, high-value customer nodes with the closest connection to the target customer were successfully screened out, providing reliable support for precision marketing and efficient customer relationship management.

[0207] More specifically, the node degree centrality index includes:

[0208] For an unweighted network, the weight of an edge is set to 1, indicating the existence of a relationship; for a weighted network, the weight value of an edge can directly reflect the association strength.

[0209] For a target node, count the number of nodes it is directly connected to in an unweighted network, or the sum of the weights of all its connecting edges in a weighted network.

[0210] For example, in a weighted network, if a customer has multiple high-frequency trading relationships with other customers, its degree centrality value is relatively high.

[0211] Customers with high degree centrality usually indicate that they have direct connections with more customers in the network and are key interactors in the network.

[0212] Enterprises can focus on analyzing these customers, for example, by preferentially covering these customers through marketing activities to achieve a higher network penetration rate.

[0213] The degree centrality index is mainly applicable to measuring active nodes in the network, helping to discover core nodes that can directly influence more customers, and is suitable for identifying "active participants" in social networks.

[0214] More specifically, the betweenness centrality index includes:

[0215] Through graph algorithms such as Dijkstra's algorithm or Floyd-Warshall algorithm, calculate the shortest paths between all pairs of nodes in the network. For example, calculate whether the shortest path between nodes s and q passes through node v.

[0216] For each pair of nodes' shortest paths, count whether the target node appears on this path and record the number of shortest paths passing through this node.

[0217] Nodes with high betweenness centrality values are usually irreplaceable bridge nodes in the network, responsible for connecting different customer groups in the network.

[0218] Enterprises can analyze these nodes to determine the key nodes connecting different customer groups and utilize their special status in the network to enhance the overall network connectivity.

[0219] The betweenness centrality index is mainly applicable to measuring the "bridge" role of nodes and can identify key customers who connect different customer groups in a social network. By optimizing resource allocation for these nodes, the information dissemination efficiency and the stability of the network structure can be enhanced.

[0220] More specifically, the closeness centrality index includes:

[0221] Use the shortest path algorithm to calculate the path lengths from the target node to all other nodes. The path length can be expressed as the number of edges or the sum of edge weights.

[0222] For the target node, count the shortest path lengths to all other nodes and calculate their average value. The shorter the path from the node to other nodes, the higher the closeness centrality value.

[0223] Customers with high closeness centrality are usually in the central position of the network and can quickly reach all other customers in the network.

[0224] Enterprises can give priority to contacting these nodes because they play an important role in information dissemination and customer relationship coordination throughout the network.

[0225] The closeness centrality index is mainly used to measure the "global position" of nodes in the network and can identify core nodes in the network. Enterprises can regard customers with high closeness centrality as potential partners or marketing priorities to quickly expand their influence.

[0226] An embodiment of the present invention also provides a system for constructing and expanding a social network of transaction customers, and the system includes:

[0227] A data acquisition module, configured to acquire an original data set of target customers, where the original data set includes structured data and unstructured data, and the unstructured data includes comment content and information platform interaction data;

[0228] A preprocessing module, configured to preprocess the unstructured data in the original data set, including:

[0229] Convert the comment content and information platform interaction data into standardized text;

[0230] Apply a text tokenization method to tokenize the standardized text to generate word sequence data;

[0231] Apply an embedded vectorization model to convert the word sequence data into a high-dimensional vector representation to obtain a set of unstructured data feature vectors;

[0232] An association index calculation module, configured to calculate association indexes between the target customer and other customers respectively according to the structured data and the set of unstructured data feature vectors, including:

[0233] Calculate the transaction frequency correlation based on structured data;

[0234] Calculate the interactive semantic similarity based on the unstructured data feature vector set;

[0235] An association matrix construction module, configured to construct a multi-dimensional association matrix between a target customer and other customers according to the association index, where the rows represent the target customer, the columns represent other customers, and the element values of the matrix are the weighted combination values of the association index;

[0236] A social network graph construction module, configured to convert the element values in the multi-dimensional association matrix into edge weights, and remove the relationships with edge weights lower than the threshold according to the preset edge weight threshold, construct a graph structure including nodes and edges, and generate a customer social network graph, where the nodes represent customers and the edges represent the association relationships between customers;

[0237] A centrality index calculation module, configured to calculate the centrality index of each customer node according to the customer social network graph;

[0238] A customer expansion module, configured to select the top N customer nodes as high-value potential customers according to the ranking of the centrality index.

[0239] It should be noted that this system corresponds to the above method. All implementation manners in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.

[0240] An embodiment of the present invention further provides a computing device, including: a processor and a memory storing a computer program. When the computer program is run by the processor, it executes the method as described above. All implementation manners in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.

[0241] An embodiment of the present invention further provides a computer-readable storage medium storing instructions. When the instructions are run on a computer, the computer is made to execute the method as described above. All implementation manners in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.

[0242] The above is the preferred implementation manner of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle described in the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A method for building and expanding a social network of successful customers, characterized in that: The method comprises: Obtaining an original data set of target customers, wherein the original data set includes structured data and unstructured data, wherein the unstructured data includes comment content and information platform interaction data; Preprocess the unstructured data in the original dataset, including: Convert comment content and information platform interaction data into standardized text; Apply the text segmentation method to segment the standardized text and generate word sequence data; The embedded vectorization model is used to transform word sequence data into high-dimensional vector representation to obtain a set of unstructured data feature vectors; According to the feature vector set of structured data and unstructured data, the correlation indexes between the target customer and other customers are calculated respectively. The correlation indexes include transaction frequency correlation and interaction semantic similarity. Among them: Calculate transaction frequency correlation based on structured data; Calculate the interaction semantic similarity based on the unstructured data feature vector set; According to the correlation index, a multi-dimensional correlation matrix between the target customer and other customers is constructed, in which the rows represent the target customers, the columns represent other customers, and the element values ​​of the matrix are the weighted combination values ​​of the correlation index; Convert the element values ​​in the multidimensional association matrix into edge weights, and according to the preset edge weight threshold, remove the relationships with edge weights lower than the threshold, build a graph structure containing nodes and edges, and generate a customer social network graph, where nodes represent customers and edges represent the associations between customers; According to the customer social network graph, calculate the centrality index of each customer node; According to the centrality index sorting, the top N customer nodes are selected as high-value potential customers; Calculate the interactive semantic similarity based on the unstructured data feature vector set, including: Obtain unstructured data feature vectors of target customers and unstructured data feature vectors of other customers, and generate a target vector set and a comparison vector set; The initial similarity value between the target vector set and the comparison vector set is calculated by the cosine similarity formula to generate an initial semantic similarity matrix; Normalizing the values ​​in the initial semantic similarity matrix to generate a normalized semantic similarity matrix; According to the similarity values ​​in the normalized semantic similarity matrix, the dynamic weight is calculated in combination with the semantic features of the target customers; Applying the dynamic weight to each item in the normalized semantic similarity matrix to generate a dynamically adjusted semantic similarity matrix; The interactive semantic similarity data set is generated by grouping calculation based on the semantic association strength of the dynamically adjusted semantic similarity matrix. The calculation formula for interactive semantic similarity is: ,in, For target customers With other customers The interactive semantic similarity of is the total number of long phrases, For the The global semantic weight of a long phrase, For target customers With other customers In the The cosine similarity on long phrases, ,in and Target customers and other customers In the The embedding vector of the long phrase Dimension value, and Target customers and other customers In the The Euclidean norm of the embedding vector on the long phrase, For target customers In the Dynamic weights on long phrases, , For target customers In the The term frequency-inverse document frequency value on the long phrase, For target customers In the The term frequency-inverse document frequency value on a long phrase.

2. A method for building and expanding a customer social network according to claim 1, characterized in that: Convert comment content and information platform interaction data into standardized text, including: Perform grammatical analysis on the review content, extract the subject, predicate and core semantic relationship in the grammatical structure, and generate text data; According to the interaction type of the information platform interaction data, a preset type weight label is assigned to different interaction data, and the preset type weight label is embedded into the text data to generate weighted text data; The weighted text data is decomposed into sentences, special symbols and meaningless phrases are removed, and the sentence structure is rearranged according to the semantic level to generate standardized text.

3. A method for building and expanding a customer social network according to claim 2, characterized in that: Apply the text segmentation method to segment the standardized text and generate word sequence data, including: Analyze the semantic context of standardized text and identify long phrases in the text; According to the word frequency statistics, semantic relevance scores, and the dependency strength of long phrases in the context, semantic weight tags are added to long phrases to generate word segmentation semantic data; the calculation formula for semantic weight is: ,in, The first The semantic weight of a long phrase, To include The number of contexts for long phrases, For the The long phrase is in The term frequency in the context - the inverse document frequency value, For the The long phrase is in The grammatical dependency in the context, , It is The depth value of a long phrase in the syntax tree, is the contextual impact factor, , and Respectively represent The long phrase is in context and The embedding vector of the context Dimension value, is the number of dimensions of the embedding vector, For the The long phrase is in The term frequency in the context - the inverse document frequency value, is a constant; Generate word frequency distribution data based on word segmentation semantic data, filter out low-frequency words for classification, and generate feature keyword data; Preliminary word sequence data is constructed based on the characteristic keyword data, and the word sequence data is generated by optimizing the arrangement according to the semantic weight sorting and context dependency.

4. A method for building and expanding a customer social network according to claim 3, characterized in that: The embedded vectorization model is used to transform word sequence data into high-dimensional vector representation, and a set of unstructured data feature vectors is obtained, including: Vectorize each word in the word sequence data to generate a word vector matrix; The dependency between adjacent words is calculated based on the word vector matrix, and the dependency is dynamically adjusted and the word vector is updated through the weighting function to generate a weighted word vector set; Use the attention mechanism to aggregate the weighted word vector set to generate a sentence vector set; The vector dimensions of the sentence vector set are unified and normalized to generate a set of unstructured data feature vectors.

5. A method for building and expanding a customer social network according to claim 4, characterized in that: Calculate transaction frequency correlation based on structured data, including: Extract the structured transaction record data of the target customers based on the structured data to generate transaction behavior data, which includes transaction time, transaction amount and transaction frequency; Calculate the transaction interaction frequency between target customers and other customers based on transaction behavior data and generate a transaction frequency matrix; According to the transaction frequency matrix, the transaction behaviors in different time periods are weighted to generate a weighted transaction frequency matrix; The values ​​in the weighted transaction frequency matrix are normalized, and the correlation values ​​with low transaction frequency are eliminated by presetting the transaction frequency threshold to generate a transaction frequency correlation data set; the calculation formula for transaction frequency correlation is: ,in, For target customers With other customers The transaction frequency correlation of For time period Internal target customers With other customers The number of transactions, For time period Internal target customers With other customers The total transaction amount, and is the weight coefficient, is the adjustment coefficient, is the length of the time period.

6. A method for building and expanding a customer social network according to claim 1, characterized in that: According to the customer social network graph, the centrality index of each customer node is calculated, including: According to the customer social network graph, the connection information and edge weight of each node are extracted, the node degree centrality index of each node is calculated, and the node degree centrality index data set is generated; Perform path analysis on the customer social network graph, calculate the betweenness centrality index of each node, and generate a node betweenness centrality index dataset; According to the node degree centrality index data set and the node betweenness centrality index data set, nodes whose node degree centrality index and betweenness centrality index are both lower than the preset centrality threshold are eliminated to generate an optimized node data set; According to the optimized node data set, the closeness between each node and the target stage is calculated to generate a node centrality index data set; the calculation formula of the node centrality index is: ,in, For Node The node centrality index of For Node The degree centrality index of ,in Is a node With Node The edge weights of Is with the node A collection of connected nodes, For Node The betweenness centrality index of ,in It is a slave node To Node The number of shortest paths, Is a node To Node Passing Node The number of shortest paths, For Node The closeness centrality index of ,in Is a node With Node The shortest path distance between , and is the weight coefficient.

7. A system for building and expanding a social network of successful customers, characterized in that: Applied to the method according to any one of claims 1 to 6, the system comprises: A data acquisition module is used to acquire the original data set of the target customer, wherein the original data set includes structured data and unstructured data, and the unstructured data includes comment content and information platform interaction data; The preprocessing module is used to preprocess the unstructured data in the original data set, including: Convert comment content and information platform interaction data into standardized text; Apply the text segmentation method to segment the standardized text and generate word sequence data; The embedded vectorization model is used to transform word sequence data into high-dimensional vector representation to obtain a set of unstructured data feature vectors; The correlation index calculation module is used to calculate the correlation index between the target customer and other customers according to the feature vector set of structured data and unstructured data. The correlation index includes transaction frequency correlation and interaction semantic similarity; wherein: Calculate transaction frequency correlation based on structured data; Calculate the interaction semantic similarity based on the unstructured data feature vector set; The association matrix construction module is used to construct a multi-dimensional association matrix between target customers and other customers based on association indicators, where rows represent target customers, columns represent other customers, and element values ​​of the matrix are weighted combination values ​​of association indicators; A social network graph construction module is used to convert the element values ​​in the multidimensional association matrix into edge weights, and according to a preset edge weight threshold, remove the relationships with edge weights lower than the threshold, construct a graph structure containing nodes and edges, and generate a customer social network graph, in which nodes represent customers and edges represent the association relationship between customers; A centrality index calculation module is used to calculate the centrality index of each customer node based on the customer social network graph; The customer development module is used to select the top N customer nodes as high-value potential customers based on the sorting of centrality indicators.

8. A computing device, characterized in that include: one or more processors; A storage device for storing one or more programs, when the one or more programs are executed by the one or more processors, the one or more processors implement the method as claimed in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a program, which, when executed by a processor, implements the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • User value prediction method and system

    CN111695719A

  • Banking business target customer determination method and device

    CN117078374A