User value mining method and system based on database fusion

Through database fusion technology and multi-level network division model, the problems of large data migration workload and limited data sources in the existing technology are solved, and efficient data processing and user value model are achieved.

CN120067172APending Publication Date: 2025-05-30CHINA UNITED NETWORK COMM GRP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510149568.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In the prior art, data migration is large, data redundancy is high, data model uses limited data sources, and it has failed to make full use of multiple heterogeneous database resources.

Method used

Through kafka technology, metadata synchronization between heterogeneous databases and cross-database access between heterogeneous data is realized to achieve database fusion. Then, the user's business characteristics and customer service requirements characteristics are extracted, and the user attribute data is integrated to form, and the user's attribute data is constructed through a multi-level network division model to perform user clustering analysis.

Benefits of technology

It improves data real-timeness, reduces redundancy, improves the efficiency of model building, and improves the effectiveness of user value models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067172A_ABST
    Figure CN120067172A_ABST
Patent Text Reader

Abstract

The invention provides a user value mining method and system based on database fusion, electronic equipment and a storage medium, and aims at solving the problems that the workload of data migration is large, the data redundancy is high, and data sources applied by a data model are limited. Metadata synchronization between the heterogeneous databases and cross-database access between the heterogeneous data are carried out through a kafka technology so as to realize database fusion; based on the data fused in the database, extracting service features of the user, and extracting customer service demand features of the user through text vector processing; the service features and the customer service demand features are fused to form overall user attribute data; and constructing a multi-level network division model, hierarchically dividing the user into a specified number of groups and sub-groups under the groups according to the user attribute data, and completing clustering analysis. The data real-time performance can be improved, redundancy is reduced, the model building efficiency is improved, and the model effect is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of data processing, and in particular, to a method for mining user value based on database fusion, a system for mining user value based on database fusion, an electronic device, and a computer-readable storage medium. Background Art

[0002] With the development of big data technology, the need for data analysis and mining, and the application of artificial intelligence technology, the databases in enterprises have become diverse. The big data platform provides capabilities such as massive computing and storage of big data, offline or real-time data processing, and interactive analysis and query; the analytical database provides scenario applications such as data mining and analysis. The database fusion system can fuse the above heterogeneous databases, efficiently convert the heterogeneous databases, and efficiently load the data of the big data platform for enterprise production into the analytical database, forming an efficient conversion of the enterprise data stream. At the same time, it saves data storage resources, meets the enterprise production requirements, and is suitable for scenarios of data mining and analysis applications relying on the enterprise big data platform.

[0003] After years of operation, the enterprise's data warehouse stores data such as user-desensitized consumption habits, behavior preferences, and product tariff characteristics. Mining user value through these data is of great significance to the development of the enterprise. Due to historical development reasons and the needs of the business system functions themselves during the enterprise operation process, various user data are stored and precipitated in relatively independent and diverse databases. Existing mining methods often integrate databases through migration technology and build data models using machine learning methods. However, traditional data migration methods have a large data migration workload, a high data redundancy rate, and consume a lot of resources. Moreover, the data sources used by traditional data models are limited, and they fail to make full use of various heterogeneous database resources. Summary of the Invention

[0004] In order to at least solve the problems in the prior art, such as large data migration workload, high data redundancy rate, limited data sources used by data models, and failure to make full use of various heterogeneous database resources. The present disclosure provides a method for mining user value based on database fusion, a system for mining user value based on database fusion, an electronic device, and a computer-readable storage medium, which can improve data real-time performance, reduce redundancy, improve the efficiency of building models, and enhance the effect of user value models.

[0005] In a first aspect, the present disclosure provides a method for mining user value based on database fusion, the method comprising:

[0006] For preset multi-source heterogeneous data, perform metadata synchronization between heterogeneous databases and cross-database access between heterogeneous data through kafka technology to achieve database fusion;

[0007] Extract the business characteristics of users based on the multi-source heterogeneous data after database fusion, and extract the customer service demand characteristics of users through text vector processing;

[0008] Fuse the business characteristics and customer service demand characteristics to form the overall user attribute data;

[0009] Construct a multi-level network division model, divide users into groups with a specified number of levels and subgroups under the groups according to the user attribute data, and complete the clustering analysis.

[0010] Furthermore, the metadata synchronization between the heterogeneous databases includes:

[0011] Configure the listening address, authentication method between the heterogeneous databases, and the field correspondence relationship between the databases;

[0012] Initiate operations related to metadata changes on the big data platform;

[0013] The monitoring platform generates topics according to the Catalog and writes the listening information into the kafka topic;

[0014] The target database synchronization service listens to the kafka topic messages and calls the metadata synchronization method to complete the metadata synchronization;

[0015] Monitor the mapping relationship between the Catalog of the big database instance and the target database instance and the DB (database), synchronize the permissions on both sides, and synchronously call the corresponding authorization SQL (Structured Query Language) of JDBC (Java Database Connectivity) to implement permission changes.

[0016] Furthermore, the cross-database access between the heterogeneous data includes:

[0017] Initiate a data access request to the target database through the big database platform;

[0018] Read the original data of the target database;

[0019] Based on the metadata synchronization, convert the format of the original data of the target database through the heterogeneous database format compatibility technology;

[0020] Fuse the converted data with the data on the big database platform.

[0021] Furthermore, the business characteristics of the users include:

[0022] Consumption habits, behavior preferences, product tariff characteristics.

[0023] Further, the extraction of the customer service demand characteristics of the user through text vector processing includes:

[0024] Obtain SMS reminders and customer service call data, including SMS reminder types, user characteristics, reminder times, frequency characteristics, user feedback characteristics, feedback frequencies, and customer service call characteristic data;

[0025] Perform text vectorization processing, including:

[0026] Preprocessing, cleaning the text data and replacing special characters with special vocabulary;

[0027] Word segmentation, performing word segmentation processing on the text data;

[0028] Stop word processing, on the basis of referring to the conventional stop word library, supplementing the special stop words that conform to the enterprise characteristics and performing stop word processing;

[0029] Parameter tuning, combining business characteristics and data characteristics, and repeatedly tuning the parameters to achieve the preset word segmentation effect;

[0030] Add corpus supplement, add industry-specific materials and supplement the word library;

[0031] Encoding, encoding the word segmentation;

[0032] Extract the customer service demand characteristics of the user through text vectorization processing and supplement them to the user value model characteristic data.

[0033] Further, the construction of the multi-level network division model includes:

[0034] Successively calculate the distance relationship between a certain characteristic attribute value of a user and the characteristic attribute values of other users, determine the threshold, and determine the number of users adjacent to each user attribute value according to the threshold to obtain the "number of adjacent users";

[0035] Search one by one for the user with the largest "number of adjacent users" among the users adjacent to a certain user and larger than its own, retain its relevant relationship, and filter out other relevant relationships;

[0036] Traverse according to the above results until the "number of adjacent users" pointing to the user is not less than the "number of adjacent users" of the current user, and record the shortest distance during the traversal process;

[0037] According to the traversal results, determine the center point and the range of its group, divide the users into a specified number of groups at different levels, and sub-groups under the groups to complete the clustering analysis.

[0038] In a second aspect, the present disclosure provides a user value mining system based on database fusion, and the system includes:

[0039] A data fusion module, which is configured to perform metadata synchronization between heterogeneous databases and cross-database access between heterogeneous data on preset multi-source heterogeneous data through Kafka technology to achieve database fusion;

[0040] A feature extraction module, which is configured to extract the business features of users based on the multi-source heterogeneous data after database fusion, and extract the customer service demand features of users through text vector processing;

[0041] A feature fusion module, which is configured to fuse the business features and customer service demand features to form overall user attribute data;

[0042] A clustering analysis module, which is configured to construct a multi-level network partitioning model, divide users into groups with a specified number of levels and subgroups under the groups according to the user attribute data, and complete the clustering analysis.

[0043] Furthermore, the data fusion module is specifically configured as follows:

[0044] Configure the listening address, authentication method between heterogeneous databases, and the corresponding relationship of fields between databases;

[0045] Initiate operations related to metadata changes on the big data platform;

[0046] The monitoring platform generates a topic according to the Catalog and writes the listening information into the Kafka topic;

[0047] The target database synchronization service listens to the Kafka topic message and calls the metadata synchronization method to complete the metadata synchronization;

[0048] Listen to the mapping relationship between the Catalog and DB of the big database instance and the target database instance, synchronize the permissions on both sides, and synchronously call the corresponding authorization SQL of JDBC to implement permission changes;

[0049] Initiate a data access request to the target database through the big database platform;

[0050] Read the original data of the target database;

[0051] Based on metadata synchronization, convert the format of the original data of the target database through heterogeneous database format compatibility technology;

[0052] Fuse the converted data with the data on the big database platform.

[0053] In a third aspect, the present disclosure provides an electronic device, including a memory and a processor. A computer program is stored in the memory. When the processor runs the computer program stored in the memory, the processor executes the user value mining method based on database fusion as described in any one of the first aspect.

[0054] In a fourth aspect, the present disclosure provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, it implements the user value mining method based on database fusion as described in any one of the first aspect.

[0055] Advantageous effects:

[0056] The user value mining method based on database fusion, the user value mining system based on database fusion, the electronic device and the storage medium provided by the present disclosure; utilize database fusion technology to efficiently fuse multiple heterogeneous databases, improve data real-time performance, reduce redundancy, reduce resource occupancy, and improve the efficiency of building models. Introduce customer service demand feature data to optimize and segment the user value mining model; utilize multi-level network partitioning technology to enhance the effect of the user value model. Description of the drawings

[0057] Figure 1 It is a schematic flowchart of a user value mining method based on database fusion provided by Embodiment 1 of the present disclosure;

[0058] Figure 2 It is a schematic diagram of a database fusion structure provided by an embodiment of the present disclosure;

[0059] Figure 3 It is a schematic diagram of the overall structure of an optimized user value model provided by an embodiment of the present disclosure;

[0060] Figure 4 It is a schematic flowchart of a real-time feedback mechanism provided by an embodiment of the present disclosure;

[0061] Figure 5 It is an architecture diagram of a user value mining system based on database fusion provided by Embodiment 2 of the present disclosure;

[0062] Figure 6 It is an architecture diagram of an electronic device provided by Embodiment 3 of the present disclosure. Detailed implementation manners

[0063] To enable those skilled in the art to better understand the technical solutions of the present disclosure, the present disclosure will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments and drawings described herein are only for explaining the present invention, rather than limiting the present invention.

[0064] It should be noted that the terms "first", "second", etc. in the description, claims and above-mentioned drawings of the present disclosure are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence; moreover, without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other arbitrarily.

[0065] Among them, the terms used in the embodiments of the present disclosure are only for the purpose of describing specific embodiments, and are not intended to limit the present disclosure. The singular forms of "a", "the" and "said" used in the embodiments of the present disclosure and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise.

[0066] In the following description, suffixes such as "module", "component" or "unit" used to represent elements are only for the convenience of description of the present disclosure, and have no specific meaning in themselves. Therefore, "module", "component" or "unit" can be used interchangeably.

[0067] The following uses specific embodiments to elaborate in detail on the technical solutions of the present disclosure and how the technical solutions of the present disclosure solve the technical problems existing in the prior art. It can be understood that in the embodiments of the present application, the execution subject can execute some or all of the steps in the embodiments of the present application. These steps or operations are only examples, and the embodiments of the present application can also execute other operations or various deformations of the operations. In addition, each step can be executed in a different order presented in the embodiments of the present application, and it is possible not to execute all the operations in the embodiments of the present application. Moreover, these several specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0068] Figure 1 The following is a schematic flowchart of a method for mining user value based on database fusion provided for Embodiment 1 of the present disclosure, as Figure 1 shown, the method includes:

[0069] Step S101: Perform metadata synchronization between heterogeneous databases and cross-database access between heterogeneous data on preset multi-source heterogeneous data through Kafka technology to achieve database fusion;

[0070] Step S102: Extract the business characteristics of users based on the multi-source heterogeneous data after database fusion, and extract the customer service demand characteristics of users through text vector processing;

[0071] Step S103: Integrate the business characteristics and customer service demand characteristics to form overall user attribute data;

[0072] Step S104: Build a multi-level network division model, hierarchically divide users into a specified number of groups and subgroups under the groups according to user attribute data, and complete cluster analysis.

[0073] Due to the reasons of the enterprise's historical development and business characteristics, there are multiple database systems coexisting within the enterprise, such as Mysql, OceanBase, Hive, etc. And for user value mining, a large number of data sources are required, so the situation of crossing different types of databases is inevitable. The traditional data preparation method is to perform data migration, migrate the data of each system to the target host or target database for data mining analysis according to a unified interface file specification; then build a data model through the migrated data. The advantage of this method is that the standard is unified and the execution process is relatively simple. The disadvantage is that it takes a long time and causes a large data delay. Another disadvantage of data migration is that there is a large amount of data redundancy, occupying a lot of storage resources.

[0074] Aiming at the disadvantages of the traditional data model construction, the embodiments of the present disclosure use the "database fusion" technology instead of the "data migration" technology. And permission division is added to ensure the security of data at the data governance level. Applying the database fusion technology to the data preparation stage of model construction can reduce data redundancy and the occupied resources of data storage, and improve the real-time performance of data. According to the heterogeneous databases of each enterprise, combined with the kafka technology, implement the database fusion technology in the massive database, and optimize the adaptability in combination with the business scenario, so as to introduce the database fusion technology into the user value mining model project.

[0075] Kafka is a high-throughput distributed publish-subscribe messaging system. Usually, it is used to build a data pipeline between systems or utilities using kafka to transform or respond to real-time data, enabling the timely transmission of business data.

[0076] The database fusion structure is as Figure 2 shown. Through the kafka technology, metadata synchronization between heterogeneous databases and cross-database access between heterogeneous data are carried out. By uniformly collecting, processing, and distributing the data of different databases through Kafka, database fusion can be realized. It can effectively solve the problems of data synchronization, integration, and real-time processing between different databases.

[0077] After completing data fusion, extract the business characteristics of users, and extract the customer service demand characteristics of users through text vector processing;

[0078] Integrate business features and customer service demand features to form overall user attribute data. Through the constructed multi-level network partitioning model, divide users into groups with a specified number of levels and subgroups under the groups according to the user attribute data to complete clustering analysis. Through user clustering analysis using multi-level network partitioning technology, the relationships and characteristics between users can be deeply explored, providing more valuable insights and decision-making support for enterprises.

[0079] Traditional user value mining models are based on data such as users' desensitized consumption habits, behavior preferences, and product tariff features to construct user value mining models, forming product recommendations, benefit recommendations, payment method recommendations, etc. for users. On this basis, the embodiments of the present disclosure introduce customer service demand data such as SMS reminders and customer service calls to more deeply explore user needs, accurately capture user value, and thus optimize and segment the output results such as product recommendations, benefit recommendations, and payment method recommendations in combination with SMS reminder preferences, payment preferences, customer service demand preferences, etc., making them more in line with user needs.

[0080] The embodiments of the present disclosure perform data fusion on the multi-source heterogeneous data of an enterprise, and through the fused data, introduce the customer service demand characteristics of users to conduct user value mining, improve data real-time performance, reduce redundancy, reduce resource occupation, and improve the efficiency of model construction. Then introduce customer service demand characteristic data such as SMS reminders and customer service calls to optimize and segment the user value mining model; use multi-level network partitioning technology to improve the effect of the user value model. Further optimize and refine the user value analysis to make it more in line with enterprise development and user needs.

[0081] Further, the metadata synchronization between the heterogeneous databases includes:

[0082] Configure the listening addresses, authentication methods, and field correspondence relationships between the heterogeneous databases;

[0083] Initiate operations related to metadata changes on the big data platform;

[0084] The monitoring platform generates topics according to the Catalog and writes the listening information into the kafka topic;

[0085] The target database synchronization service listens to the kafka topic messages and calls the metadata synchronization method to complete metadata synchronization;

[0086] Monitor the mapping relationships between the Catalog and DB of the big database instance and the target database instance, synchronize the permissions on both sides, and synchronously call the corresponding privilege-granting SQL of JDBC to implement permission changes.

[0087] Metadata changes between heterogeneous databases refer to the operations of synchronizing metadata such as table structures, fields, indexes, and constraints between different types of databases (such as MySQL, PostgreSQL, Oracle, MongoDB, etc.). Since there may be differences in the metadata definitions and operation syntax of different databases, special attention needs to be paid to issues such as data type mapping, syntax conversion, and dependency relationships.

[0088] Among them, the metadata table further includes changes in data table attributes, changes in fields in the data table, changes in field attributes, etc.; the listening information is the information triggered by real-time changes.

[0089] The mapping relationship between Catalog and DB refers to the correspondence between data tables and data fields in heterogeneous databases; by listening to the mapping relationship between Catalog and DB, it can be ensured that the permissions of both databases are consistent, thus ensuring the security and consistency of data access.

[0090] Invoking the metadata synchronization method to complete metadata synchronization specifically includes:

[0091] 1. Use a CDC tool (such as Debezium) to capture metadata change events in the source database.

[0092] 2. Send the change events to Kafka.

[0093] 3. Listen for change events in Kafka in the target database and execute the corresponding DDL (Data Definition Language) statements according to the syntax of the target database.

[0094] By implementing metadata change synchronization between heterogeneous databases, the consistency and stability of the system are ensured.

[0095] Furthermore, the cross-database access between the heterogeneous data includes:

[0096] Initiate a data access request to the target database through the large database platform;

[0097] Read the original data in the target database;

[0098] Based on metadata synchronization, through the heterogeneous database format compatibility technology, convert the format of the original data in the target database;

[0099] Fuse the converted data with the data in the large database platform.

[0100] After completing the metadata synchronization between heterogeneous databases, cross-database access can be carried out through the big data platform. The original data of the target database is formatted and converted into the format of the big data platform, so as to integrate the converted data with the data of the big database platform.

[0101] Furthermore, the business characteristics of the user include:

[0102] Consumption habits, behavior preferences, and product tariff characteristics.

[0103] In user value mining, consumption habits, behavior preferences, and product tariffs are key elements that can help enterprises better understand user needs, optimize products and services, form user product recommendations, rights and interests recommendations, payment method recommendations, etc., and enhance user satisfaction and loyalty; by analyzing these data, enterprises can more accurately identify user needs, optimize products and services, enhance user satisfaction and loyalty, and ultimately achieve business growth.

[0104] Furthermore, the extraction of the user's customer service demand characteristics through text vector processing includes:

[0105] Obtain SMS reminders and customer service call data, including SMS reminder types, user characteristics, reminder times, frequency characteristics, user feedback characteristics, feedback frequencies, and customer service call characteristic data;

[0106] Perform text vectorization processing, including:

[0107] Preprocessing, cleaning the text data and replacing special characters with special vocabulary;

[0108] Word segmentation, performing word segmentation on the text data;

[0109] Stop word processing, supplementing special stop words that conform to the enterprise characteristics on the basis of referring to the conventional stop word library for stop word processing;

[0110] Parameter tuning, repeatedly tuning the parameters in combination with business characteristics and data characteristics to achieve the preset word segmentation effect;

[0111] Add corpus supplement, add industry-specific materials to supplement the word library;

[0112] Encoding, encoding the word segmentation;

[0113] Extract the user's customer service demand characteristics through text vectorization processing and supplement them to the user value model feature data.

[0114] Taking an operator enterprise as an example, for customer service demand characteristics, it is mainly SMS reminders and customer service calls;

[0115] SMS reminders refer to SMS messages sent by enterprises to users, such as consumption reminders, shutdown reminders, and contract expiration reminders. SMS reminders are a window for enterprises to provide warm services to users. SMS links connect the relationship between enterprises and users. Users can quickly handle business through the link of SMS reminders, saving users' time and improving the service quality of enterprises. The combination of SMS reminders and SMS business halls enables users to efficiently handle communication consumption matters, playing an irreplaceable role in maintaining user relationships.

[0116] Different user groups have different responses to SMS reminders, some welcome, some reject, some block, and some contact customer service for consultation. At this time, the introduction and combination of customer service call data can obtain an overall portrait of users' responses to SMS reminders and customer needs. Through in-depth mining and segmentation of customer service call data, we can perceive the various demands of users.

[0117] By introducing SMS reminders, customer service calls and other data into the model, the model includes SMS reminder types, user characteristics, reminder time, frequency characteristics, user feedback characteristics, feedback frequency, customer service call characteristics, etc. After integration and improvement, the output of product recommendations, rights and interests recommendations, payment method recommendations and other results are optimized and segmented in combination with SMS reminder preferences, payment preferences, customer service demand preferences, etc., which better meet user needs.

[0118] In the process of SMS reminder and customer service call data preprocessing and model building, multi-level network partitioning technology was introduced to process text and build models. The construction process is text vectorization processing (including preprocessing, word segmentation, stop word processing, parameter tuning, adding corpus supplements, encoding), and multi-level network partitioning model construction. The specific processing steps are:

[0119] 1. Text vectorization processing

[0120] 1.1 Preprocessing

[0121] Clean the text data. Text information mainly comes from SMS reminders and customer service call data. Combined with business characteristics, different from traditional methods, preprocessing is not simply to clean special characters, but to replace special characters with special words to retain key features to the maximum extent.

[0122] 1.2 Word segmentation

[0123] The text data is segmented. Due to the obvious business characteristics, a special vocabulary library for the industry is specially created during the segmentation process to make the segmentation more accurate and more in line with the business characteristics.

[0124] 1.3 Stop word processing

[0125] On the basis of citing the conventional stop word library, special stop words that meet the characteristics of the target enterprise are supplemented.

[0126] 1.4 Parameter Tuning

[0127] Combined with business characteristics and data characteristics, the parameters are repeatedly tuned to achieve the preset word segmentation effect; the preset word segmentation effect can be determined according to actual inspection to ensure that the word segmentation accuracy reaches the predetermined effect, such as more than 95%. The parameter tuning of text word segmentation includes mode selection, stop word selection, etc. The tuning method is to use the grid search method combined with the actual business and data characteristics to determine the optimal parameters.

[0128] 1.5 Adding Corpus Supplement

[0129] Add industry-specific materials to supplement the thesaurus.

[0130] Through the above process, the customer service demand characteristics of users are extracted and supplemented into the feature data of the user value model.

[0131] 1.6 Encoding

[0132] Encode the word segmentation.

[0133] Fuse the data obtained by vectorizing the text with other business data to form the overall user attribute data.

[0134] Furthermore, the construction of the multi-level network partitioning model includes:

[0135] Successively calculate the distance relationship between a certain feature attribute value of a user and other user feature attribute values, determine the threshold, and determine the number of users adjacent to each user attribute value according to the threshold to obtain the "number of adjacent users";

[0136] Search one by one for the user with the largest "number of adjacent users" among the users adjacent to a certain user and larger than its own, retain its relevant relationship, and filter out other relevant relationships;

[0137] Traverse according to the above results until the "number of adjacent users" pointing to the user is not less than the "number of adjacent users" of the current user, and record the shortest distance during the traversal process;

[0138] Based on the traversal results, determine the center point and the scope of its group, divide the users into a specified number of groups and subgroups under the group to complete the clustering analysis.

[0139] The multi-level network partitioning model is an algorithm framework for decomposing a complex network into multiple sub-networks or modules. Its core goal is to gradually simplify the network structure in a hierarchical manner and finally achieve efficient and accurate partitioning. Through the multi-level network partitioning technology for user clustering analysis, the relationships and characteristics between users can be deeply explored, providing valuable insights and decision-making support for enterprises.

[0140] User clustering is to divide users with similar characteristics or behaviors into the same group.

[0141] Network partitioning is to divide the nodes (users) in the network into several communities or sub-networks, so that the connections within the communities are dense and the connections between the communities are sparse.

[0142] Multi-level network partitioning is to partition the network at different levels to capture more complex user relationships.

[0143] The construction process of the multi-level network partitioning model specifically includes:

[0144] 2.1. Sequentially calculate the distance relationship between a certain attribute value of a user and the attribute values of other users, determine the threshold, and determine the number of users adjacent to each user attribute value (i.e., the "adjacent number") according to the threshold.

[0145] Calculate the distance formula and generate a distance matrix through the following formula:

[0146] For the vector X = [X 1 , X 2 , …, X n , its distance formula is ||X||∞ = max 1≤i≤n |X i |, which is the infinity norm (Infinity Norm) of the vector X = [X 1 , X 2 , …, X n , also known as the maximum norm. It represents the component with the largest absolute value in the vector X. The symbol on the left side of the equal sign represents the distance matrix, and the right side represents the calculation formula of the distance formula matrix.

[0147] Distance relationship vector and identification number:

[0148] The distance matrix is arranged in order to form a distance relationship vector, and identification numbers are formed from small to large. The values of the X vector are sequentially compared with each value in the distance relationship vector, and the threshold is determined based on the returned data relationship.

[0149] Determine the threshold and the "adjacent number":

[0150] Determine the number of users adjacent to each user attribute value (i.e., the "adjacent number") according to the threshold.

[0151] 2.2. Based on the adjacent relationship identification number network diagram generated in 2.1, search one by one for the user with the largest "adjacent number" among the users adjacent to a certain user, retain its relevant relationship, and filter out other relevant relationships. That is, retain the user relationship with the largest "adjacent number", and remove other user relationships.

[0152] Traverse each user.

[0153] For each user, traverse its adjacent users.

[0154] Check whether the "number of adjacent users" of an adjacent user is greater than that of the user.

[0155] Find the user that meets the condition and has the largest "number of adjacent users", and retain the relevant relationship between this user and the currently traversed user. Filter out the relationships between the current traversal and other adjacent users.

[0156] 2.3. Traverse according to the result of 2.2 (i.e., the above result) until the "number of adjacent users" of the pointed user is not less than that of the current user. Record the shortest distance during the traversal process.

[0157] The specific process includes:

[0158] Initialization:

[0159]

[0160] Q ← {(u 0 , 0)}, where Q is a queue that stores the users to be visited and their current distances.

[0161] Traversal process:

[0162] When Q is not empty, perform the following steps:

[0163] Take out an element (u, dist) from Q.

[0164] If C(u) ≥ C(u 0 ), then update S as S ← S ∪ {dist} (if multiple shortest distances need to be recorded, use a set or a list; if only one shortest distance is needed, use a variable and update the minimum value).

[0165] For each v ∈ N(u) and v has not been visited, add (v, dist + 1) to Q.

[0166] End condition:

[0167] When Q is empty, the traversal ends.

[0168] Result:

[0169] If S is not empty, the minimum value in S is the shortest distance from u 0 to the user whose number of adjacent users is not less than u 0 . If S is empty, there is no user that meets the condition.

[0170] Among them, S is an empty stack, Q is a queue, u is an element to be visited, and N represents the neighborhood.

[0171] 2.4. Based on the traversal result in 2.3, determine the center point and the scope of its group, divide the users into a specified number of groups at different levels, and sub - groups under the groups to complete the clustering analysis.

[0172] In the process of model construction in the embodiments of the present disclosure, comparisons are made with the fasttext and LSTM models. The processing efficiency is higher, the calculation speed is faster, and the accuracy is relatively high. The model effect can be further improved through the integration of multiple technologies.

[0173] The overall structure of the optimized model is as Figure 3 shown: After integrating various data in the database, data integration is carried out, through machine learning, natural language processing, and multi - level network division; a user value mining model is obtained. By using the database fusion technology to efficiently fuse heterogeneous databases, data such as SMS reminders and customer service calls are introduced to optimize and subdivide the user value mining model; machine learning models and multi - level network division technologies are used to improve the effect of the user value mining model.

[0174] Furthermore, the method further includes:

[0175] Introducing a real - time feedback mechanism to update the user value model.

[0176] Introducing a real - time feedback mechanism into the user value mining model. The real - time feedback data source comes from user behavior data, user customer service calls, etc. The real - time feedback information forms a linkage mechanism, which acts on links such as corpus supplementation, stop - word processing, model construction and subsequent links respectively. The feedback content includes data feedback and real - time result feedback of model training and prediction. The flowchart is as Figure 4 shown.

[0177] In the embodiments of the present disclosure, by efficiently fusing multiple heterogeneous databases, the real - time performance of data is improved, redundancy is reduced, resource occupation is reduced, and the efficiency of model construction is improved. Customer service demand feature data is introduced to optimize and subdivide the user value mining model; multi - level network division technology is used to improve the effect of the user value model.

[0178] Embodiment 2 of the present disclosure also provides a user value mining system based on database fusion, as Figure 5 shown, the system includes:

[0179] A data fusion module 11, which is configured to perform metadata synchronization between heterogeneous databases and cross - database access between heterogeneous data for preset multi - source heterogeneous data through kafka technology to achieve database fusion;

[0180] A feature extraction module 12, which is configured to extract the business features of users based on the multi - source heterogeneous data after database fusion, and extract the customer service demand features of users through text vector processing;

[0181] A feature fusion module 13, which is configured to fuse business features and customer service requirement features to form overall user attribute data;

[0182] A clustering analysis module 14, which is configured to construct a multi-level network partitioning model, and hierarchically partition users into a specified number of groups and subgroups under the groups according to the user attribute data to complete the clustering analysis.

[0183] Further, the data fusion module 11 is specifically configured as follows:

[0184] Configure the listening address, authentication method between heterogeneous databases, and the field correspondence relationship between databases;

[0185] Initiate operations related to metadata changes on the big data platform;

[0186] The monitoring platform generates a topic according to the Catalog and writes the monitoring information into the kafka topic;

[0187] The target database synchronization service listens to the kafka topic messages and calls the metadata synchronization method to complete the metadata synchronization;

[0188] Listen based on the mapping relationship between the Catalog and the DB of the big database instance and the target database instance, synchronize the permissions on both sides, and synchronously call the corresponding privilege-granting SQL of JDBC to implement permission changes;

[0189] Initiate a data access request to the target database through the big data platform;

[0190] Read the original data of the target database;

[0191] Based on metadata synchronization, through heterogeneous database format compatibility technology, convert the format of the original data of the target database;

[0192] Fuse the converted data with the data on the big data platform.

[0193] Further, the business features of the user include:

[0194] Consumption habits, behavior preferences, product tariff features.

[0195] Further, the feature extraction module 12 is specifically configured as follows:

[0196] Obtain SMS reminder and customer service call data, including SMS reminder type, user characteristics, reminder time, frequency characteristics, user feedback characteristics, feedback frequency, and customer service call characteristic data;

[0197] Perform text vectorization processing, including:

[0198] Preprocessing: Clean the text data and replace special characters with specific vocabulary;

[0199] Word segmentation: Perform word segmentation on the text data;

[0200] Stop word processing: Based on the citation of a conventional stop word library, supplement special stop words that conform to the characteristics of the enterprise and perform stop word processing;

[0201] Parameter tuning: Combine business characteristics and data characteristics to repeatedly tune the parameters to achieve the preset word segmentation effect;

[0202] Increase corpus supplementation: Add industry-specific materials to supplement the word library;

[0203] Encoding: Encode the word segmentation;

[0204] Extract the customer service demand characteristics of the user through text vectorization processing and supplement them to the feature data of the user value model.

[0205] Furthermore, the clustering analysis module 14 is specifically set as follows:

[0206] Successively calculate the distance relationship between a certain feature attribute value of a user and the feature attribute values of other users, determine the threshold, and determine the number of users adjacent to each user attribute value according to the threshold to obtain the "number of adjacent users";

[0207] Search one by one for the user among the adjacent users of a certain user whose "number of adjacent users" is larger and the largest than its own, retain its relevant relationship, and filter out other relevant relationships;

[0208] Traverse according to the above results until the "number of adjacent users" pointing to the user is not less than the "number of adjacent users" of the current user, and record the shortest distance during the traversal process;

[0209] Based on the traversal results, determine the center point and the range of its group, divide the users into a specified number of groups at different levels and subgroups under the groups to complete the clustering analysis.

[0210] The user value mining system based on database fusion in the embodiments of the present disclosure is used to implement the user value mining method based on database fusion in Embodiment 1 of the method. Therefore, the description is relatively simple. For specific details, reference can be made to the relevant descriptions in the previous method embodiments, and details will not be repeated here.

[0211] In addition, as Figure 6 shown, Embodiment 3 of the present disclosure further provides an electronic device, including a memory 100 and a processor 200. A computer program is stored in the memory 100. When the processor 200 runs the computer program stored in the memory 100, the processor 200 executes the above various possible methods.

[0212] Among them, the memory 100 is connected to the processor 200. The memory 100 can be a flash memory, a read-only memory, or other memories, and the processor 200 can be a central processing unit or a single-chip microcomputer.

[0213] In addition, an embodiment of the present disclosure further provides a computer-readable storage medium, on which a computer program is stored, and the computer program is executed by the processor to implement the above various possible methods.

[0214] The computer-readable storage medium includes volatile or non-volatile, removable or non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, computer program modules, or other data). The computer-readable storage medium includes, but is not limited to, RAM (Random Access Memory), ROM (Read-Only Memory), EEPROM (Electrically Erasable Programmable Read Only Memory), flash memory, or other memory technologies, CD-ROM (Compact Disc Read-Only Memory), digital versatile disc (DVD, Digital Video Disc), or other optical disc storage, magnetic cassette, tape, magnetic disk storage, or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer.

[0215] It can be understood that the above embodiments are merely exemplary embodiments adopted to illustrate the principle of the present disclosure, but the present disclosure is not limited thereto. For those of ordinary skill in the art, various modifications and improvements can be made without departing from the spirit and essence of the present disclosure, and these modifications and improvements are also regarded as the protection scope of the present disclosure.

Claims

1. A user value mining method based on database fusion, characterized in that: The method comprises: For preset multi-source heterogeneous data, Kafka technology is used to synchronize metadata between heterogeneous databases and access heterogeneous data across databases to achieve database integration. Based on the multi-source heterogeneous data after database fusion, the user's business characteristics are extracted, and the user's customer service demand characteristics are extracted through text vector processing; Integrate business characteristics and customer service demand characteristics to form overall user attribute data; Construct a multi-level network partitioning model, divide users into a specified number of groups and sub-groups under the groups according to user attribute data, and complete cluster analysis.

2. The method according to claim 1, characterized in that: The metadata synchronization between the heterogeneous databases includes: Configure the listening addresses and authentication methods between heterogeneous databases, as well as the field correspondence between databases; Initiate metadata change related operations on the big data platform; The monitoring platform generates a topic based on the catalog and writes the monitoring information into the Kafka topic. The target database synchronization service listens to the Kafka topic message and calls the metadata synchronization method to complete metadata synchronization; Monitor the mapping relationship between the Catalog and database DB of the large database instance and the target database instance, synchronize the permissions on both sides, and synchronously call the corresponding authorization SQL of the Java database connection JDBC to implement permission changes.

3. The method according to claim 2, characterized in that The cross-database access between heterogeneous data includes: Initiate a data access request to the target database through the big database platform; Read the original data of the target database; Based on metadata synchronization, the original data of the target database is converted into a new format through heterogeneous database format compatibility technology. Integrate the converted data with the big database platform data.

4. The method according to claim 1, characterized in that: The user's service characteristics include: Consumption habits, behavioral preferences, and product pricing characteristics.

5. The method according to claim 1, characterized in that The extracting of the user's customer service demand features through text vector processing includes: Obtain SMS reminder and customer service call data, including SMS reminder type, user characteristics, reminder time, frequency characteristics, user feedback characteristics, feedback frequency, and customer service call characteristic data; Perform text vectorization processing, including: Preprocessing: cleaning the text data and replacing special characters with specialized words; Word segmentation: perform word segmentation on text data; Stop word processing: on the basis of citing the conventional stop word database, special stop words that meet the characteristics of the enterprise are added to carry out stop word processing; Parameter tuning: Combine business characteristics and data characteristics to repeatedly tune parameters to achieve the preset word segmentation effect; Increase corpus supplements, add industry-specific materials, and supplement the vocabulary; Encoding: encode the word segmentation; Through text vectorization processing, the user's customer service demand characteristics are extracted and supplemented into the user value model feature data.

6. The method according to claim 1, characterized in that The multi-level network partitioning model construction includes: Calculate the distance relationship between a certain characteristic attribute value of a user and the characteristic attribute values ​​of other users one by one, determine the threshold, and determine the number of users adjacent to each user attribute value according to the threshold to obtain the "adjacent number"; Search one by one for the largest user whose "number of neighbors" is greater than that of a user, keep the relevant relationships with them, and filter out other relevant relationships; Traverse according to the above results until the "adjacent number" pointing to the user is not less than the "adjacent number" of the current user, and record the shortest distance during the traversal process; Based on the traversal results, the scope of the center point and its group is determined, and the users are hierarchically divided into a specified number of groups and sub-groups under the groups to complete the cluster analysis.

7. A user value mining system based on database fusion, characterized in that: The system comprises: The data fusion module is configured to synchronize metadata between heterogeneous databases and perform cross-database access between heterogeneous data through Kafka technology for preset multi-source heterogeneous data, so as to achieve database fusion; A feature extraction module is configured to extract the user's business features based on the multi-source heterogeneous data after database fusion, and extract the user's customer service demand features through text vector processing; A feature fusion module is configured to fuse business features and customer service demand features to form overall user attribute data; The cluster analysis module is configured to construct a multi-level network partitioning model, and divides users into a specified number of groups and sub-groups under the groups according to user attribute data to complete cluster analysis.

8. The system according to claim 7, characterized in that The data fusion module is specifically configured as follows: Configure the listening addresses and authentication methods between heterogeneous databases, as well as the field correspondence between databases; Initiate metadata change related operations on the big data platform; The monitoring platform generates a topic based on the Catalog and writes the monitoring information into the Kafka topic; The target database synchronization service listens to the Kafka topic message and calls the metadata synchronization method to complete metadata synchronization; Monitor the mapping relationship between the Catalog and DB of the large database instance and the target database instance, synchronize the permissions on both sides, and synchronously call the corresponding authorization SQL of JDBC to implement permission changes; Initiate a data access request to the target database through the big database platform; Read the original data of the target database; Based on metadata synchronization, the original data of the target database is converted into a new format through heterogeneous database format compatibility technology. Integrate the converted data with the big database platform data.

9. An electronic device, characterized in that: It comprises a memory and a processor, wherein the memory stores a computer program, and when the processor runs the computer program stored in the memory, the processor executes the user value mining method based on database fusion as described in any one of claims 1 to 6.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the user value mining method based on database fusion according to any one of claims 1 to 6 is implemented.