Intelligent Business Decision-making Method Based on Multi-source Heterogeneous Data
Through data quality evaluation and blood relationship diagram construction, combined with distributed computing system, the unified standardization and personalized service problems of multi-source heterogeneous data are solved, efficient data processing and real-time application are realized, and the company's business decision-making support capabilities are improved.
Patent Information
- Application Number
- CN202510466707.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-04-15
AI Technical Summary
In enterprise data management, how to achieve unified and standardized processing of multi-source heterogeneous data, avoid repeated calculations, and provide personalized data services to meet the needs of different business scenarios, especially in the process of data collection, processing and application, to ensure the accuracy and timeliness of data.
Using data quality evaluation model and cleaning process, a data relationship diagram is constructed, a distributed computing system is used to process data in parallel, business-related features are extracted, data assets for different business scenarios are generated, and real-time updates are provided through a RESTful-style data service interface.
It realizes the intelligence of the entire process from data collection to application, improves the efficiency of data assets, and improves marketing effectiveness, risk control capabilities and user experience.
Smart Images

Figure CN119988477B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information technology, and in particular to an intelligent business decision-making method based on multi-source heterogeneous data. Background Art
[0002] In enterprise data management, accurate tracking and enriching data before it goes live presents a complex technical challenge. First, different business systems generate data in varying formats and types. Standardizing and standardizing this data during the data collection phase is a primary concern when tracing data sources. Second, multiple copies of data are often generated during production and circulation. Key considerations during the data enrichment phase include building a data lineage diagram to document data conversion and calculation processes, thereby avoiding duplicate calculations and wasted resources. Third, large enterprises generate massive amounts of data daily. Timely processing and application of this data for business analysis requires data platforms to respond within seconds, necessitating the development of highly available, high-performance distributed computing engines. Finally, different business personnel have varying data priorities. Data governance platforms must consider how to design targeted, personalized data services, provide data assets tailored to specific business scenarios, and meet business demands such as precision marketing, risk control, and product recommendations. Summary of the invention
[0003] The present invention provides an intelligent business decision-making method based on multi-source heterogeneous data, which mainly includes:
[0004] Acquire multi-source heterogeneous data from different business scenarios, use a data quality assessment model to determine whether the data meets the preset quality requirements, and use the data that meets the quality requirements as available multi-source heterogeneous data; construct a data lineage relationship diagram based on the available multi-source heterogeneous data, where the data lineage relationship diagram records the copy generation, conversion, and calculation process of the available multi-source heterogeneous data during the production and circulation process; based on the data lineage relationship diagram, shard the information in the data lineage relationship diagram to obtain multiple data shards, and use a distributed computing system to process the multiple data shards in parallel; the processing of each data shard includes:
[0005] Extract behavioral data from different users, extract user behavioral features from different user behavioral data, use collaborative filtering algorithms to generate user profiles for different users, and based on the user profiles of different users, obtain marketing recommendation lists corresponding to different users. Integrate the data format of the marketing recommendation lists to generate data assets for precision marketing scenarios;
[0006] Extract historical transaction data, extract risk characteristics from the historical transaction data, input the risk characteristics into a preset risk assessment model to obtain a risk warning list. Among them, the risk assessment model is trained by the decision tree algorithm using historical risk characteristics, integrate the data format of the risk warning list to generate a data asset for the risk control scenario;
[0007] Extract product functional attribute data, extract potential product characteristics from the product functional attribute data, and calculate the similarity between user preferences and products based on the potential product characteristics to obtain a product recommendation list. Integrate the data format of the product recommendation list to generate a data asset for the product recommendation scenario; Integrate the marketing recommendation list, risk warning list, and product recommendation list generated from each shard data into a data service interface and provide it for the business system to call to realize the real-time application of data assets. The data service interface adopts the RESTful style and supports batch query and real-time query to ensure the timeliness and consistency of data; According to the marketing recommendation list, risk warning list, and product recommendation list, use the Redis cache mechanism to realize real-time data update to ensure that the latest data is obtained by calling the interface, and realize efficient query and real-time update of data.
[0008] The technical solution provided by the embodiments of the present invention may include the following beneficial effects:
[0009] The present invention discloses an intelligent business decision-making method based on multi-source heterogeneous data. This method ensures data quality through a data quality assessment model and a cleaning process, constructs a data lineage graph to avoid duplicate calculations. Utilize distributed computing technology to efficiently process data, extract business-related characteristics, and generate data assets for different business scenarios. For the precision marketing scenario, risk control scenario, and product recommendation scenario, use collaborative filtering, decision tree, and matrix factorization algorithms respectively to generate a marketing recommendation list, a risk assessment model, and a product recommendation list. Finally, integrate these results into a RESTful-style data service interface for the business system to call. The present invention realizes the full-process intelligence from data collection, processing to application, improves the utilization efficiency of data assets, provides accurate and timely business decision support for enterprises, and effectively improves the marketing effect, risk control ability, and user experience. Brief Description of the Drawings
[0010] Figure 1 It is a flowchart of the intelligent business decision-making method based on multi-source heterogeneous data of the present invention. Detailed Embodiment
[0011] Next, the technical solutions in the embodiments of the present invention will be clearly and detailedly described in conjunction with the accompanying drawings in the embodiments of the present invention. The described embodiments are only a part of the embodiments of the present invention.
[0012] Such asFigure 1 , the intelligent service decision-making method based on multi-source heterogeneous data in this embodiment may specifically include:
[0013] Step S101, obtain multi-source heterogeneous data in different business scenarios, and use a data quality assessment model to determine whether the data meets the preset quality requirements. The data that meets the quality requirements is used as available multi-source heterogeneous data.
[0014] Use a preset data acquisition module to obtain multi-source heterogeneous data from different business scenarios, input the data into a preset data quality assessment model, and determine whether the data meets the preset quality requirements. If it meets the quality requirements, the data that meets the quality requirements is used as available multi-source heterogeneous data. If it does not meet the quality requirements, trigger the data cleaning module to clean the data. After cleaning, the data is input into the quality assessment model for re-judgment. If it meets the quality requirements, the cleaned data is used as available multi-source heterogeneous data. If the data still does not meet the quality requirements after cleaning, use clustering analysis in machine learning algorithms to classify the data, and use regression analysis to optimize the data distribution for different categories. The optimized data is input into the quality assessment model for judgment again. If it meets the quality requirements, the optimized data is used as available multi-source heterogeneous data. If the optimized data still does not meet the quality requirements, mark the optimized data as abnormal data and store it in the abnormal database. At the same time, trigger the data source optimization module to optimize the data source, and input the data provided after the data source optimization into the data acquisition module for cyclic processing.
[0015] Exemplarily, the data acquisition module is the starting point of the entire data processing flow. It obtains multi-source heterogeneous data from different business scenarios. Taking an e-commerce platform as an example, it can collect various types of data such as user browsing records, purchase history, and evaluation information. These data have diverse sources and different formats, and need to be uniformly processed. The data quality assessment model is a key link to ensure data availability. It is usually a multi-dimensional quantitative assessment framework that combines a rule engine, statistical analysis, and machine learning techniques. Specifically, it may consist of models such as a statistical distribution model, a verification model based on a rule engine, a correlation verification model, a machine learning anomaly detection model, and a text quality analysis model. Through the organic combination of rules and machine learning, this model not only ensures the verification efficiency but also has the adaptability to complex scenarios. It can evaluate data from multiple dimensions such as integrity, accuracy, and consistency. For example, for user registration information, it can check whether required fields are empty, whether the mobile phone number format is correct, and whether the age matches the date of birth. By setting preset requirements, such as requiring more than 90% of the fields to meet the specifications, it is determined whether the data quality meets the standard. The data cleaning module is responsible for processing unqualified data. Common cleaning operations include deduplication, filling missing values, and correcting outliers. For example, for duplicate order records, deduplication can be performed based on the order number; for missing user ages, the average age of the user group can be used for filling; for significantly abnormal commodity prices, the historical average price can be used for replacement. Cluster analysis in machine learning algorithms can help better understand the data structure. Taking customer segmentation as an example, based on features such as user consumption behavior and browsing habits, users can be divided into different categories, such as high-frequency low-amount users and low-frequency high-amount users. This classification can help optimize data targeted. Regression analysis is used to optimize the data distribution. For example, for a sales prediction model, linear regression can be used to analyze historical sales data, find the key factors affecting sales, and adjust abnormal data points accordingly to make the data distribution more reasonable. The principal component analysis method is a commonly used dimensionality reduction method. When processing high-dimensional data, exemplarily, the optimized data may contain ten related indicators (such as daily active duration, number of search keywords, add-to-cart rate, number of favorite products, order cancellation rate, average order value, promotion sensitivity, etc.). Calculate the covariance matrix of the optimized data, obtain the principal component directions (eigenvectors) and the variance ratio (eigenvalues) they explain through eigenvalue decomposition. Select the eigenvectors corresponding to the top 3 largest eigenvalues to form 3 principal components. Principal component 1 (purchasing power): a linear combination of high-weight indicators such as the number of orders and consumption amount, reflecting the user's consumption ability; Principal component 2 (activity): a linear combination of high-weight indicators such as click-through rate and browsing duration, reflecting the user's activity level; Principal component 3 (return tendency): a linear combination of high-weight indicators such as return rate and number of customer service complaints, reflecting the user's return sensitivity. Project the original 10-dimensional data onto these 3 principal components, and cumulatively retain 85% of the original data variance, realizing the simplification from complex indicators to core features.Three principal component indicators are used to replace the original ten indicators, retaining 85% of the core information. This not only reduces the data complexity but also retains the key information, thereby improving the efficiency of subsequent processing. The handling of abnormal data is crucial for enhancing the overall data quality. Marking and storing the data that cannot be processed by conventional methods can provide a reference for subsequent data source optimization. For example, after discovering that a large amount of user age data is abnormal, it may be necessary to redesign the age input interface and add rationality checks. Data source optimization is an ongoing process of improvement. By analyzing abnormal data, problems in the data collection process, such as sensor failures and human input errors, can be identified. For these problems, corresponding measures can be taken, such as replacing equipment and optimizing operation processes, to improve data quality at the source. It can be understood that the available multi-source heterogeneous data includes various formats of user behavior, product basic attributes, product functional attributes, historical transactions, and other aspects of data. Although the available multi-source heterogeneous data has been preliminarily processed, it may not have been integrated and its structure is not suitable for directly extracting features, and subsequent steps are needed for further processing. This multi-level, iterative data processing method can effectively improve data quality and lay a solid foundation for subsequent data analysis and applications. By continuously optimizing the data processing process, enterprises can make better use of data assets and improve the efficiency and accuracy of decision-making.
[0016] Step S102: Construct a data lineage graph based on the available multi-source heterogeneous data, where the data lineage graph records the copy generation, transformation, and calculation processes of the available multi-source heterogeneous data during production and circulation.
[0017] Obtain data copy generation information, data transformation information, and data calculation information from the available multi-source heterogeneous data. Record the obtained data copy generation information in the data lineage graph. Record the data transformation information in the data lineage graph. Record the data calculation information in the data lineage graph. During the process of recording the data calculation information in the data lineage graph, detect whether there is a duplicate calculation part. If there is, mark the duplicate calculation part. Use a greedy algorithm to optimize the marked duplicate calculation part to generate an optimized calculation information record. Update the optimized calculation information record to the data lineage graph to replace the original duplicate calculation part. Through the data lineage graph construction module, continuously obtain the latest information on the data copy generation, transformation, and calculation processes, and dynamically update the data lineage graph at a frequency of once per minute.
[0018] For example, a data lineage diagram is a visual tool for tracking data flows and transformations. It documents the entire lifecycle of data from source to end use, including data generation, transformation, and computation. In practical applications, this type of diagram is crucial for understanding complex data processing workflows, optimizing data management, and ensuring data quality. Obtaining data replica generation information from multi-source heterogeneous data can be understood as extracting a snapshot from the original data for subsequent processing. For example, on an e-commerce platform, user browsing history may originate from web pages, mobile devices, and mini-programs. Data replica generation information records the generation time (e.g., 2025-04-01), location (e.g., Server A), and source (e.g., "browsing log table"). This recording method facilitates tracing the origin of data and ensures traceability. When recording data transformation information in a data lineage diagram, the conversion rules and methods must be clearly defined. Specifically, suppose the amount field in a user purchase record is converted from US dollars to RMB. The conversion rule might be "multiply by the exchange rate of 6.5," and the conversion method is field mapping. The specific conversion steps include reading the original field, applying the exchange rate, and outputting a new field. This detailed record helps understand how data changes from one form to another, providing a basis for subsequent audits. Recording data calculation information is more complex. In one embodiment, to calculate a user's average spending over the past 30 days, the calculation process might be "read purchase records, filter the time range, sum and divide by the number of orders." The calculation parameter is "time range = 30 days," and the result is "average spending = 150 yuan." Recording this information in the data lineage relationship graph clearly demonstrates the calculation logic. It should be noted that if duplicate calculations are detected during the calculation process, such as multiple calculations of the same user's total spending, these are marked, such as "The total calculation for user ID 001 was repeated 3 times." Preferably, when using a greedy algorithm to optimize duplicate calculations, the path with the least computational effort can be prioritized. For example, for repeated calculations of user spending totals, the result of the first calculation can be directly reused, rather than rereading and calculating the original data each time. This optimized calculation information is updated to the data lineage relationship graph, replacing the original redundant parts, thereby reducing resource waste. For example, dynamically updating the data lineage relationship graph once a minute ensures real-time data availability. In one possible implementation, suppose an e-commerce platform adds a batch of new order data. The data replica generation information records the addition date (2025-04-01), the conversion information records the change in order status from "pending payment" to "completed," and the calculation information updates the user's total purchase amount. Continuously acquiring this information and updating the lineage relationship graph can fully reflect the data flow process. As you can see, constructing a data lineage relationship graph not only records the data's "past" but also provides a foundation for future analysis. For example, by examining the lineage relationship graph, you can quickly determine whether an anomalous data entry is the result of a collection issue or a conversion error.This clear context helps improve the efficiency and transparency of data management. In one embodiment, if it is found that the total consumption amount of a certain user is abnormally high, tracing the lineage graph may show that it is due to the failure to optimize duplicate calculations in a timely manner. The optimized record then shows that the total amount has been adjusted from 1000 yuan to 500 yuan, reflecting the true situation. This method ensures the reliability of data through multi-level verification. Specifically, from multiple perspectives, the information generated by data replicas provides an anchor for time and source, the transformation information clarifies the rules of data changes, the calculation information reveals the origin of the results, and the optimization process reduces redundancy. This multi-dimensional recording and continuous update jointly support the integrity and practicality of the data lineage graph. For example, managers can quickly determine which data needs further verification and which calculations can be directly reused through the lineage graph, thereby improving the response speed of the overall process. Generally speaking, the core functions of the data lineage graph are as follows: 1. Solve the data silo problem: The available multi-source heterogeneous data in step S101 may be independent data sets scattered in different business systems (such as CRM user data, ERP transaction data, and behavior data in the log system). Data lineage establishes cross-system logical associations by recording the data flow path, enabling the originally isolated data to have business semantic connections. 2. Eliminate the risk of duplicate calculations: In the scenario of precision marketing, the construction of user portraits may require simultaneous invocation of historical order data and behavior buried-point data. If calculated directly from the original data, it may lead to duplicate calculations such as order amount aggregation and behavior path analysis for the same user in different scenarios. Data lineage can mark the completed intermediate calculation results to achieve the reuse of calculation results. 3. Ensure feature interpretability: Anti-fraud features in the risk control scenario (such as "the number of device changes of a user within 7 days") need to track the change records of device IDs across multiple systems. Data lineage can clearly show the data source path of this feature, facilitating compliance auditing and model interpretability verification. The data lineage graph is not only a technical tool but also a strategic asset for data management. It can improve the efficiency and quality of data processing, enhance the interpretability and traceability of data, and thus provide more reliable support for data-driven decision-making.
[0019] Step S103, based on the data lineage graph, slice the information of the data lineage graph to obtain multiple data slices, and use a distributed computing system to perform parallel processing on the multiple data slices.
[0020] Extract the latest information on the data copy generation, transformation, and calculation processes from the data lineage graph, use the key fields in the extracted latest information on the data copy generation, transformation, and calculation processes as the input of the hash algorithm to obtain multiple different shard indexes, and allocate the data corresponding to the same shard index together to form a data shard. Use Apache Spark to evenly distribute multiple data shards to different preset computing nodes for parallel computing, where each computing node independently processes the data shards allocated to it. Obtain the shard computing status of each computing node. If the shard computing status of a certain computing node is data skew or processing bottleneck, use Apache Spark to evenly distribute the data of this computing node to this computing node and multiple standby computing nodes for re-parallel computing.
[0021] Exemplarily, a data lineage graph is a key tool for tracking data flow and transformation, from which the latest information is extracted to lay the foundation for subsequent processing. Taking an e-commerce platform as an example, the graph may contain multiple dimensions such as user browsing records, product information, order data, etc. When extracting this information, it is necessary to pay attention to the generation time, source, and transformation process of the data to provide a clear context for subsequent analysis. The hash algorithm plays an important role in data sharding to ensure uniform distribution of data. It is crucial to select appropriate key fields as input, and input key fields such as user ID or product category in the latest information of the generated, transformed, and calculated process of the extracted data copy into the hash algorithm. The output of the hash algorithm is usually an integer value of a fixed length (such as a 32-bit or 64-bit integer), and this value is converted into a specific shard number through further calculation (such as modulo operation). The shard number is used as a shard index to guide data allocation. For example, for user behavior data, the user ID can be selected as the hash input to obtain a shard index with the user ID as the key field, and data of the same user is allocated to the shard with the same number for subsequent calculation. As a distributed computing system, Apache Spark evenly distributes multiple data shards to different computing nodes for parallel computing. This distribution not only balances the computing load but also improves data locality and reduces cross-node data transmission. Parallel computing is the core advantage of Apache Spark, which means that multiple computing nodes process the allocated data simultaneously, and through distributed task scheduling, the computing resources are fully utilized. Through parallel computing, the system can process data of multiple user groups simultaneously, significantly shortening the response time. It should be noted that Apache Spark generates the shard calculation status of each node during the parallel computing process. The shard calculation status refers to the intermediate or final records generated after the data shards are parallelly calculated by each computing node. If the shard calculation status of a certain computing node is data skew or processing bottleneck, in a possible implementation, a new field with a granularity smaller than the key field is selected and re-input into the hash algorithm to obtain a new shard index for the data in this computing node, and the data corresponding to the same new shard index is allocated together to form at least one new data shard. Among them, the new field with finer granularity refers to a new field formed by adding new data attributes (such as time, subcategory, random factor) on the basis of the original key field, and by increasing the field dimension, the logical division of the formed data shards is made more fragmented and dispersed, thereby reducing the data volume of a single shard. For example, assume that the above steps form 20 data shards (numbered S1 - S20), and 10 computing nodes (N1 - N10) are preset. When node N3 processes data shard S9, it may take much longer to process due to the large amount of data of a certain user, showing data skew. Exemplarily, the system detects that the CPU occupancy rate of N3 continuously exceeds 90%, while that of other nodes is only 50%, and confirms it as a processing bottleneck.Preferably, for computing nodes with data skew or processing bottlenecks, the system will select fields with finer granularity to re-shard. For example, the user ID combined with the browsing timestamp is used as the new input for the hashing algorithm to generate a new shard index, and data for the same user in different time periods is assigned to one new shard. Exemplarily, the original shard S9 may be split into S9-1 and S9-2, each containing records from different time periods. Through Apache Spark, multiple new data shards are evenly distributed among the original computing nodes and multiple standby computing nodes for parallel computing again, such as N9 and N10. The introduction of standby nodes ensures the elastic expansion of computing resources. For example, after re-sharding, the new shards S9-1 and S9-2 are evenly distributed among the original computing node N3 and the standby computing nodes N9 and N10 for parallel computing again. The load of N3 drops to 60%, and the processing time of N3, N9, and N10 is consistent with the average level of other computing nodes. If the shard computing status of a certain computing node is data skew or processing bottleneck, in another possible implementation, through Apache Spark, the data of this computing node is evenly distributed among this node and multiple standby computing nodes for parallel computing again. Exemplarily, through Apache Spark, the data of the original computing node is directly evenly distributed among the original computing node and multiple standby computing nodes for parallel computing again. For example, the data of the original shard S9 is evenly distributed among the original computing node N3 and the standby computing nodes N9 and N10 for parallel computing again. This refined adjustment can better balance the computing load. Through continuous optimization, the system can ultimately achieve second-level response. The whole process reflects the closed-loop optimization of data processing: obtaining information from the data lineage graph, generating shard indexes through the hashing algorithm, performing parallel computing using Apache Spark, and then adjusting the strategy according to the feedback of Apache Atlas to ensure that the data has second-level response capabilities during the transfer and computing process. It can be understood that the advantage of this method is that it can quickly respond to abnormal situations in data processing. For example, when the data volume of a certain product category surges, the system can quickly adjust by recalculating the shard index to avoid overloading a single node. This flexibility is particularly important for large-scale data processing on e-commerce platforms and can significantly improve the response speed and accuracy of the recommendation system. Generally speaking, the core values of the distributed computing system are as follows: First, it processes complex data topologies. When constructing a commodity similarity matrix in the product recommendation scenario, it involves the joint calculation of commodity attribute data (structured), user reviews (unstructured), and clickstream graph data (graph structure). The distributed system realizes the parallel processing of multi-modal data through sharding strategies (such as sharding by commodity ID). Second, it optimizes the computing process. Analyzing user behavior sequences (such as session partitioning) requires sliding window calculations over time.A distributed system can ensure that the behavioral data of the same user is concentrated in the same data shard through pre-sharding (such as sharding by hashing user IDs), avoiding performance losses caused by cross-node data transmission. Third, it supports dynamic expansion capabilities: in scenarios with traffic peaks such as "Double 11", based on the marked data dependencies in the data lineage graph, the computing resources of specific feature calculation links (such as the real-time user preference feature calculation cluster) can be dynamically expanded without the need for full-link capacity expansion. Exemplarily, taking the "associated network features" in risk control as an example: directly using the data in step S101 requires reconstructing the multi-layer associated network of users - merchants - geographical locations from the original transaction records each time, and the single calculation takes more than 2 hours. After processing through steps S102 - S103: based on the lineage relationship, the existing intermediate calculation results are identified, and through distributed graph computing for incremental updates, the feature generation time is shortened to 15 minutes. This architecture design enables the system to maintain a daily data processing volume of over 1 billion while controlling the P99 latency of feature production within 5 minutes and reducing the computing resource consumption by 40%. This is precisely the engineering practice value brought by steps S102 - S103, far exceeding the optimization effect that can be achieved by simple data cleaning.
[0022] Step S104: Extract the behavioral data of different users, extract the user behavioral features of different users from the behavioral data of different users, generate user portraits of different users using a collaborative filtering algorithm, and based on the user portraits of different users, obtain a marketing recommendation list corresponding to different users, and integrate the data format of the marketing recommendation list to generate data assets for the precise marketing scenario.
[0023] Extract the behavioral data of the different users; extract the behavioral features of different users from the behavioral data of different users, including the number of user clicks, browsing duration, and purchase frequency. Based on the behavioral features of different users, calculate the similarity between users using a user-based collaborative filtering algorithm. Generate user portraits of different users according to the user similarity. Combine the basic product attributes, and perform a dot product calculation on the preference vectors in the user portraits of different users and the basic product attribute vectors to generate a matching matrix reflecting the matching degree between users and basic product attributes. Each element in the matrix represents the preference intensity of a specific user for a certain basic product attribute, where the basic product attributes include category attributes, physical attributes, and price attributes. Generate an initial recommendation list corresponding to different users according to the matching matrix. Extract the category click proportion data in the historical behaviors of different users from the behavioral features of different users, and sort the initial recommendation lists of different users correspondingly according to the category click proportion data in the historical behaviors of different users to generate a marketing recommendation list corresponding to different users. Integrate each marketing recommendation list according to the preset data format to generate data assets for the precise marketing scenario.
[0024] Exemplarily, user behavior data is the basis of precision marketing. The time window statistical method is used to extract the feature of the number of user clicks in the recent 7 days, and the sliding time window is used to quantify the short-term interest intensity of users, reflecting the timeliness of behavior; the page type correlation analysis method is used to extract the browsing duration feature by distinguishing the browsing depth of different page types, and the industry dynamics frequency calculation method is used to extract the purchase frequency feature by counting the number of purchases based on a sliding window. User behavior features such as the number of user clicks, browsing duration, and purchase frequency can comprehensively depict the online activity patterns of users. For example, an e-commerce platform may find that user A has browsed mobile phone products 50 times in the past month, stayed for an average of 3 minutes each time, and finally purchased 2 mobile phones. These data reflect user A's high interest in mobile phone products. The collaborative filtering algorithm based on users realizes personalized recommendation by calculating the similarity between users, and can quantify the similarity degree of user behavior. Suppose user A and user B both frequently browse smartphones and have similar purchase patterns, the system will consider them to have a high similarity. This similarity calculation lays the foundation for subsequent user profiling and product recommendation. User profiling is a comprehensive description of user preferences. By analyzing the browsing and purchase history of user A, the following profile may be obtained: aged 25-35, a technology enthusiast, preferring high-end smartphones, and paying special attention to the photography function. Such a profile can guide the formulation of more precise marketing strategies. Constructing a matching matrix of user preferences and product basic attributes is the core link of the recommendation system. Taking smartphones as an example, the matrix may include dimensions such as price, brand, screen size, and camera pixel. User A's preferences may be manifested as attaching great importance to camera performance and being relatively insensitive to price. The dot product calculation is performed on the preference vector (attaching great importance to camera performance and being relatively insensitive to price) in user A's profile and the product basic attribute vector to generate a matching matrix reflecting the matching degree of user-product basic attributes. Each element in the matrix represents the preference intensity of a specific user for the basic attributes of a certain product. This matching relationship provides a basis for generating the initial recommendation list. The generation of the initial recommendation list takes into account the matching degree of user preferences and product basic attributes. For user A, the system may recommend 10 high-end smartphones, including the latest released flagship models and professional photography phones. This list reflects the system's preliminary judgment of user needs. Sorting the initial recommendation list is a key step in optimizing the user experience. The system may rank the mobile phones with high-pixel cameras at the top of the list according to the browsing and click ratio of mobile phones with strong camera functions in user A's historical behavior, generating a marketing recommendation list corresponding to user A. The marketing recommendation list needs to be integrated according to a preset data format to form a data asset for the precision marketing scenario. This may include information such as user ID, recommended product list, and recommended reasons for each product.For example, the marketing recommendation list for user A may contain personalized descriptions such as "Based on your interest in high-quality photography, we recommend the latest flagship mobile phone of XX brand. Its 100-megapixel main camera will bring you a professional-level shooting experience." Store the generated data assets in a dedicated database to support subsequent marketing activities. This database may adopt distributed storage technology to ensure high availability and fast access to data. By regularly updating and analyzing these data, enterprises can continuously optimize their recommendation algorithms and marketing strategies, improving user satisfaction and conversion rates.
[0025] Step S105: Extract historical transaction data, extract risk features from the historical transaction data, input the risk features into a preset risk assessment model to obtain a risk warning list. Among them, the risk assessment model is trained by the decision tree algorithm using historical risk features. Integrate the data format of the risk warning list to generate data assets for the risk control scenario.
[0026] Extract the historical transaction data. Clean the historical transaction data to remove missing values and outliers. Extract risk features including transaction amount, transaction frequency, and transaction time interval from the cleaned historical transaction data; among them, the transaction amount is directly obtained, the transaction frequency is obtained by counting the number of transactions of the same user within 30 days, and the transaction time interval is obtained by calculating the time difference between adjacent transactions. Input the risk features into the risk assessment model to calculate the risk score of each transaction; combine a preset risk score threshold to judge the transaction risk level. A score greater than the preset risk score threshold is a high risk, and a score less than or equal to the preset risk score threshold is a low risk; generate a risk warning list containing transaction numbers, risk scores, and risk levels; integrate the risk warning list according to the preset data format to generate data assets for the risk control scenario; among them, the risk assessment model is pre-trained by the decision tree algorithm using historical risk features, including: pre-establishing risk score labels and classifying transactions into two categories: high risk and low risk.
[0027] Calculate the contribution degree of historical risk features to the risk score using information gain, where the historical risk features include historical transaction amount, historical transaction frequency, and historical transaction time interval; select the feature with the largest contribution degree as the splitting node of the decision tree according to the information gain calculation result; generate a decision tree by recursively dividing the data set, set the minimum number of samples of the node as the recursive termination condition, and train to generate a risk assessment model.
[0028] Exemplarily, obtaining historical transaction data is the first step in risk assessment. This data may include information such as user ID, transaction amount, transaction time, etc. During the data cleaning process, the following situations may be encountered: A user has multiple abnormally large transactions in a day, which may be due to data entry errors or potential fraud. By setting reasonable thresholds, such as the total daily transaction amount not exceeding 10 times the user's monthly income, such outliers can be effectively identified and processed. Extracting risk features is the basis for constructing a risk assessment model. The transaction amount directly reflects the scale of the transaction and is crucial for risk assessment. The transaction frequency can reveal the user's behavior pattern. For example, if an ordinary wage earner's account suddenly makes frequent large transfers in a short period, this may indicate abnormal fund flows. The transaction time interval is also of great significance. If an account makes two large transfers consecutively at 3:00 am and 3:05 am, this unusual time pattern may imply automated fraud. Using historical risk features, a risk assessment model is pre-trained through a decision tree algorithm. Specifically, a risk score label is established in advance to provide a training target for the model. High-risk transactions may include large overseas transfers, frequent small transfers, etc., while low-risk transactions may be regular transactions such as salary payments and daily consumption. The information gain is used to calculate the contribution of the transaction amount, transaction frequency, and transaction time interval to the risk score. Information Gain is the core indicator for feature selection in the decision tree algorithm. Its essence is to evaluate the discrimination ability of a feature by measuring the reduction in the uncertainty of data classification caused by the feature. In the risk control scenario, the information gain is used to quantify the contribution of these three features, namely the transaction amount, transaction frequency, and transaction time interval, to the risk score. For example, it may be found that the information gain of the transaction amount is the highest, which means that the transaction amount is the most effective feature for distinguishing high-risk and low-risk transactions. The process of constructing a decision tree is a process of gradually refining the risk judgment criteria. The decision tree is generated by recursively partitioning the data set, and the minimum number of samples at a node is set as the recursive termination condition. Combining historical transaction data and the pre-established risk score label, a risk assessment model is trained. In one embodiment, the decision tree may first be divided into two groups according to the transaction amount, less than 1000 yuan and greater than 1000 yuan, and then refined according to the transaction frequency, and finally a risk assessment model is generated. When calculating the risk score, the model outputs a numerical value according to the feature value. For example, if the transaction amount of user D is 3000 yuan and the frequency is 3 times / 30 days, the score may be 0.7. Combining a preset risk score threshold such as 0.6, transactions with a score greater than 0.6 are determined to be high-risk. Such a scoring mechanism intuitively reflects the potential risk. The finally generated risk warning list is an important tool for risk management. The risk warning list needs to be integrated according to a preset data format to form a data asset for the risk control scenario. It can be determined that the data asset for the risk control scenario is a risk management tool for enterprises.It may contain the following information: transaction number T20240622001, risk score 85, and high risk level. Such data assets enable the e-commerce platform to quickly identify and process high-risk transactions. By continuously updating and optimizing this risk assessment model, the e-commerce platform can improve the processing efficiency of normal transactions while ensuring security, providing a better service experience for customers. Store the generated data assets in a dedicated database to support subsequent risk control. This database may adopt distributed storage technology to ensure high availability and fast access to data.
[0029] Step S106, extract product functional attribute data, extract product potential features from the product functional attribute data, calculate the similarity between user preferences and products based on the product potential features, obtain a product recommendation list, and integrate the data format of the product recommendation list to generate data assets for the product recommendation scenario.
[0030] Obtain the product functional attribute data, preprocess the product functional attribute data, use the interpolation method to process missing values, and use the box plot method to remove outliers. Among them, the product functional attribute data includes the functional characteristics of the product, the technical parameters of the product, and the usage scenarios of the product. From the preprocessed product functional attribute data, use factor analysis to extract product potential features, and use the principal component analysis method for dimensionality reduction to obtain the dimensionality-reduced product potential features. Among them, the product potential features refer to the features that can reveal the internal relationship between product functional attributes. According to the user behavior data, obtain the user's rating data for the product, combine it with the dimensionality-reduced product potential feature vector, and construct a user-product potential feature rating matrix. Each element in the matrix is the user's rating of the product potential feature. Use the singular value decomposition algorithm to decompose the user-product potential feature rating matrix into a user feature matrix and a product potential feature matrix, and obtain the user preference vector and the product potential feature weight vector. Among them, each row of the user feature matrix corresponds to the user preference vector, and each column of the product feature matrix corresponds to the product potential feature weight vector. According to the user preference vector and the product potential feature weight vector, use the cosine similarity to calculate the similarity between user preferences and products. Obtain the recommended products with a similarity greater than the preset similarity threshold. Based on the recommended products, generate a product recommendation list that includes user numbers, product numbers, and recommended products; integrate the product recommendation list according to the preset data format to generate data assets for the product recommendation scenario.
[0031] Exemplarily, the construction of the product recommendation list begins with data acquisition and preprocessing. Extracting product functional attribute data is a crucial step. These data may include features such as the functional characteristics of the product, the technical parameters of the product, and the usage scenarios of the product. In practical applications, there may be missing values and outliers in the product functional attribute data, which need to be cleaned and processed. The interpolation method is a commonly used method for handling missing values. For example, for the missing waterproof rating data of a certain mobile phone, the waterproof ratings of other models in the same brand and series can be used for filling. The box plot method can effectively identify outliers. For example, if the battery capacity of a certain tablet computer is much larger than the normal range, it may be an error in data entry and should be excluded. The preprocessed data lays the foundation for subsequent analysis. Factor analysis can extract the product latent features that can reveal the internal relationships between product functional attributes from numerous variables. For example, when analyzing a smartwatch, it may be found that the surface features such as battery capacity, battery life, and charging speed actually reflect a common product latent feature: battery performance. Principal component analysis further reduces the data dimension and generates the most representative reduced-dimension product latent features. This is particularly important when dealing with high-dimensional data. For example, smart home products may involve dozens or even hundreds of parameters. Through principal component analysis, a new set of fewer key features can be generated, and these new key features can replace the original hundreds of parameters. When obtaining the rating data based on user behavior data, it can be understood that user browsing, purchasing, or commenting behaviors can be converted into ratings. For example, if user A frequently purchases large-screen mobile phones, it can be inferred that user A has a relatively high rating for "display effect", which can be set to 4.5 points. Combining with the reduced-dimension product latent feature vectors, a user-product latent feature rating matrix is constructed. In the matrix, the rows represent users, and the columns represent features. For example, user A rates the "display effect" of a certain mobile phone as 4.5 and the "battery life" as 3.0. Using the singular value decomposition algorithm, the user-product latent feature rating matrix (R) is decomposed into the product of three matrices R = UΣV T , where U is an m×m orthogonal user feature matrix (left singular vector), representing user preferences; Σ is an m×n diagonal matrix, and the diagonal elements are singular values (sorted from largest to smallest), reflecting the importance of features; V T is an n×n orthogonal product latent feature matrix (transpose of the right singular vector), representing the product latent feature space. After decomposition, dimensionality reduction truncation is performed: retain the first k largest singular values (e.g., k = 50), truncate U to m×k, Σ to k×k, and V T to k×n, obtaining the user feature matrix (U = m×k) and the product latent feature matrix (V T= k × n). Finally, the output user feature matrix, with each row representing a user preference vector and each column representing a product attribute weight vector, is used for similarity calculation. For example, user A's preference vector might be [0.8, 0.3, 0.5], reflecting their preferences for display, battery life, and camera performance; while a particular phone's product potential attribute weight vector might be [0.9, 0.4, 0.6], representing its performance in these three areas. Using the cosine similarity method to calculate the similarity between the two, the result is 0.998, exceeding the similarity threshold of 0.7. The phone is then marked as a recommended product. When generating a product recommendation list based on the recommended products, the list might specifically include user ID U001, product ID P2025, and the recommended product "a certain brand of large-screen, long-battery-life mobile phone." Preferably, a pre-set data format is used to integrate the product recommendation list to generate a data asset tailored for product recommendation scenarios. This data asset supports personalized recommendations and enhances the user experience. It should be noted that combining factor analysis and singular value decomposition can uncover implicit relationships between attributes and recommend products that better meet user needs. For example, if a user prefers a large screen but doesn't prioritize photos, the system can accurately recommend a model with a prominent screen and an average camera. This approach can effectively improve matching efficiency in data-driven recommendation scenarios. Understandably, unlike marketing recommendation lists, which recommend products based on basic product attributes, product recommendation lists make more in-depth recommendations based on product features and the similarity between user preferences and products.
[0032] In step S107, the marketing recommendation list, risk warning list, and product recommendation list generated by each shard data are integrated into the data service interface and provided to the business system for calling to realize the real-time application of data assets. The data service interface adopts the RESTful style and supports batch query and real-time query to ensure the timeliness and consistency of the data.
[0033] The data service interface is designed in RESTful style, and the request method and parameters of the interface are defined. The interface supports GET requests, and the parameters include user ID and timestamp, which are used to distinguish between batch queries and real-time queries. The marketing recommendation list, risk warning list, and product recommendation list generated by each shard data are obtained from the pre-established database. The marketing recommendation list, risk warning list, and product recommendation list are integrated into the data service interface and returned to the caller in JSON format. When the business system calls the data service interface, it obtains the corresponding data based on the user ID and timestamp parameters. If the timestamp is the current time, a real-time query is triggered to obtain the latest data from each data source. If the timestamp is a historical time, a batch query is triggered to obtain historical data from the cache.
[0034] Exemplarily, the RESTful-style data service interface design provides a standardized method for inter-system communication. Taking the user profile service as an example, an interface " / user-profile / {userId}" can be defined to obtain the profile data of a specific user through a GET request. Here, "user-profile" represents "user information" or "user profile", and "{userId}" represents a path parameter used to dynamically specify the specific user to be operated on. The timestamp parameter "timestamp" is used to distinguish between real-time queries and batch queries. For example, "?timestamp=2025022012" means to obtain the historical data at that time point. The design of the interface parameters directly affects the efficiency and flexibility of data acquisition. Using the user ID as a path parameter makes the request more semantic, while using the timestamp as a query parameter provides the ability to access temporal data. This design allows business systems to flexibly perform real-time or batch data acquisition according to requirements. The pre-stored marketing recommendation list, risk warning list, and product recommendation list in the database constitute a complete user data ecosystem. The marketing recommendation list may include attributes such as age, occupation, and consumption habits; the risk warning list may identify the user's credit risk level; and the product recommendation list generates personalized suggestions based on user characteristics and historical behaviors. The JSON format, as the carrier of data transmission, has the advantages of being lightweight and easy to parse. A typical JSON response may be as follows: {"userId": "12345", "age": 30, "occupation": "engineer", "riskLevel": "low", "recommendations": ["Product A", "Product B"]}, where "userId" refers to the string or number used to uniquely identify a user; "age" refers to the age; "occupation" refers to the occupation; "riskLevel" refers to the risk level; and "recommendations" refers to the recommendation list. This format is both intuitive and convenient for front-end applications to process. The mechanism for distinguishing between real-time queries and batch queries reflects the system's performance optimization considerations. When the timestamp is the current time, the system will obtain the latest data from various original data sources to ensure data real-time. For example, when a user has just completed a large transaction, the risk warning system can immediately reflect this change. For batch queries of historical data, the system quickly retrieves from the cache, greatly improving the response speed. The Redis cache mechanism plays a key role in this design. It not only accelerates data access but also enables real-time data updates. For example, when the user profile changes, the system will immediately update the corresponding data in Redis. This ensures that even in high-concurrency scenarios, business systems can obtain the latest and most accurate user information. The advantage of this design is that it balances data real-time and system performance.By reasonably utilizing caching and real-time queries, the system can not only meet the demand for the latest data but also maintain high efficiency when dealing with a large number of historical data queries. This is particularly important for e-commerce systems that need to handle both real-time transactions and historical data analysis. Generally speaking, this solution combining RESTful API design with a caching mechanism not only provides a standardized data access interface but also realizes efficient, flexible, and reliable data services through intelligent query strategies and cache update mechanisms. This lays a foundation for building a large-scale system with rapid response and scalability.
[0035] Step S108: According to the marketing recommendation list, risk warning list, and product recommendation list, implement real-time data update through the Redis caching mechanism to ensure that the latest data is obtained when the interface is called, and realize efficient query and real-time update of data.
[0036] In the real-time query scenario, after obtaining the marketing recommendation list, store the data in the cache through the ZADD command of Redis SortedSet. After obtaining the risk warning list, store the result in the cache through the List command of Redis. After obtaining the product recommendation list, store the list in the cache through the RPUSH command of Redis. In the batch query scenario, read the marketing recommendation list through the ZRANGE command of Redis, and read the risk warning list and product recommendation list through the LRANGE command of Redis to ensure data consistency. Set the survival time of the data stored in the cache through the EXPIRE command of Redis. When the survival time expires, Redis will automatically delete the data stored in the cache.
[0037] Exemplarily, in real-time query and batch query scenarios, various data structures and commands in Redis support the efficient storage and reading of data. The following analyzes the storage and reading mechanisms for marketing recommendation lists, risk warning lists, and product recommendation lists, and illustrates specific implementation methods through multi-faceted examples to highlight the business logic and effects of each theme. For the storage and reading of marketing recommendation lists, Redis's Sorted Set stores data according to priority through the ZADD command, which is suitable for dynamically adjusting the recommendation order in real-time scenarios. Exemplarily, assume that an e-commerce system generates personalized marketing activities for users, and the recommended content includes coupons, recommended products, etc. Each time a recommendation is generated, the system calculates the score of each recommendation based on the user's recent behavior, such as browsing records or transaction frequency. For example, browsing the recommended product page adds 0.5 points, and completing a transaction adds 1 point. The Sorted Set stores the recommended items with the user ID as the key and the score as the sorting basis. In batch queries, the ZRANGE command can quickly extract items within a score range. For example, extract recommended items with scores between 0.5 and 2.0. This method supports quick screening of high-priority recommendations to ensure that the business system can promptly display the most relevant marketing content. It should be noted that the storage of risk warning lists utilizes Redis's List structure, and the latest warning results are appended to the head of the list through the LPUSH command, which is suitable for real-time recording of user risk status. In a possible implementation, a user triggers a high-risk warning due to frequent large transfers recently, and the LPUSH command pushes this record into the List with the user ID as the key. In batch queries, LRANGE extracts data from the List according to the index range. For example, when the system queries the user's last 10 risk warning records, LRANGE returns the results with the key "risk:user123:alert" and the index range from 0 to 9. The LRANGE command can also extract a specified range of recommended items, such as the last 5 recommendations. This method is suitable for the business system to display recommended content in chronological order, with clear logic and easy maintenance. It can be understood that to achieve efficient query and real-time update of data, the EXPIRE command in Redis needs to be used to set the data survival time. For example, the real-time query cache is set to "EXPIRE marketing:U10086 3600", indicating that the data expires after 1 hour, and the data in the cache is forcibly deleted to maintain real-time performance; the batch query cache can be set to "EXPIRE risk:U10086 86400" (deleted after 24 hours) to reduce the load caused by frequent refreshes. This mechanism balances real-time performance and system resource occupancy. In one embodiment, if the user "U10086" has just completed a transaction, the real-time query immediately updates the cache, while the batch query still uses the snapshot data of the previous day, with clear logic and high efficiency. For example, in a high-concurrency scenario, the cache expiration time can be adjusted according to user activity.Set a short expiration time (e.g., 1 hour) for active users to ensure frequent updates; set a long expiration time (e.g., 12 hours) for low-active users to reduce unnecessary refreshes. This differential strategy improves resource utilization while ensuring efficient query and real-time update of data.
[0038] Only some preferred embodiments of the present invention are listed above, but the present invention is not limited thereto, and many improvements and transformations can be made. As long as the improvements and transformations are made based on the basic principles of the present invention, they should be regarded as falling within the protection scope of the present invention.
Claims
1. An intelligent business decision-making method based on multi-source heterogeneous data, characterized in that, The method includes: Obtaining multi-source heterogeneous data in different business scenarios, using a data quality assessment model to determine whether the data meets the preset quality requirements, and taking the data that meets the quality requirements as available multi-source heterogeneous data; Based on the available multi-source heterogeneous data, constructing a data lineage graph, where the data lineage graph records the copy generation, transformation, and calculation processes of the available multi-source heterogeneous data during production and circulation; Based on the data lineage graph, performing sharding processing on the information of the data lineage graph to obtain multiple data shards, and using a distributed computing system to perform parallel processing on the multiple data shards; The processing of each data shard includes: Extracting the behavior data of different users, extracting the user behavior characteristics of different users from the behavior data of different users, generating user portraits of different users using a collaborative filtering algorithm, and based on the user portraits of different users, obtaining a marketing recommendation list for corresponding different users, and integrating the data formats of the marketing recommendation list to generate data assets for the precision marketing scenario; Extracting historical transaction data, extracting risk characteristics from the historical transaction data, inputting the risk characteristics into a preset risk assessment model to obtain a risk warning list, where the risk assessment model is trained by a decision tree algorithm using historical risk characteristics, and integrating the data formats of the risk warning list to generate data assets for the risk control scenario; Extracting product functional attribute data, extracting product potential characteristics from the product functional attribute data, and calculating the similarity between user preferences and products based on the product potential characteristics to obtain a product recommendation list, and integrating the data formats of the product recommendation list to generate data assets for the product recommendation scenario; Integrating the marketing recommendation list, risk warning list, and product recommendation list generated by each shard of data into a data service interface and providing it for the business system to call to realize the real-time application of data assets. The data service interface adopts the RESTful style, supports batch query and real-time query, and ensures the timeliness and consistency of data; According to the marketing recommendation list, risk warning list, and product recommendation list, realizing real-time data update through the Redis cache mechanism to ensure that the latest data is obtained by calling the interface, and realizing efficient query and real-time update of data.
2. The method according to claim 1, wherein The obtaining of multi-source heterogeneous data in different business scenarios, using a data quality assessment model to determine whether the data meets the preset quality requirements, and taking the data that meets the quality requirements as available multi-source heterogeneous data includes: Using a preset data collection module to obtain multi-source heterogeneous data from different business scenarios, inputting the data into a preset data quality assessment model to determine whether the data meets the preset quality requirements; If it meets the quality requirements, taking the data that meets the quality requirements as available multi-source heterogeneous data; If it does not meet the quality requirements, triggering the data cleaning module to clean the data, inputting the cleaned data into the quality assessment model for re-judgment, and if it meets the quality requirements, taking the cleaned data as available multi-source heterogeneous data; If the data still does not meet the quality requirements after data cleaning, clustering analysis in machine learning algorithms is used to classify the data. For different categories, regression analysis is adopted to optimize the data distribution. The optimized data is input into the quality assessment model again for judgment. If it meets the quality requirements, the optimized data is used as available multi-source heterogeneous data; If the optimized data still does not meet the quality requirements, the optimized data is marked as abnormal data and stored in the abnormal database. At the same time, the data source optimization module is triggered to optimize the data source. The data provided after the data source optimization is input into the data collection module for cyclic processing.
3. The method according to claim 1, wherein Based on the available multi-source heterogeneous data, a data lineage graph is constructed. Among them, the data lineage graph records the copy generation, transformation, and calculation processes of the available multi-source heterogeneous data during production and circulation, including: Obtain data copy generation information, data transformation information, and data calculation information from the available multi-source heterogeneous data; Record the obtained data copy generation information into the data lineage graph; Record the data transformation information into the data lineage graph; Record the data calculation information into the data lineage graph; During the process of recording the data calculation information into the data lineage graph, detect whether there is a duplicate calculation part. If so, mark the duplicate calculation part; Use the greedy algorithm to optimize the marked duplicate calculation part and generate an optimized calculation information record; Update the optimized calculation information record to the data lineage graph to replace the original duplicate calculation part; Through the data lineage graph construction module, continuously obtain the latest information on the data copy generation, transformation, and calculation processes, and dynamically update the data lineage graph at a frequency of once per minute.
4. The method according to claim 1, wherein Based on the data lineage graph, the information in the data lineage graph is sliced to obtain multiple data slices, and a distributed computing system is used to perform parallel processing on the multiple data slices, including: Extract the latest information on the data copy generation, transformation, and calculation processes from the data lineage graph. Use the key fields in the extracted latest information on the data copy generation, transformation, and calculation processes as the input of the hash algorithm to obtain multiple different slice indexes, and allocate the data corresponding to the same slice index together to form a data slice; Use Apache Spark to evenly distribute the multiple data slices to different preset computing nodes for parallel computing. Among them, each computing node independently processes the data slice allocated to it; Obtain the slice calculation status of each computing node. If the slice calculation status of a certain computing node is data skew or processing bottleneck, use Apache Spark to evenly distribute the data of this computing node to this computing node and multiple standby computing nodes for re-parallel computing.
5. The method according to claim 1, characterized in that Extract the behavioral data of different users, extract the user behavioral characteristics of different users from the behavioral data of different users, generate user portraits of different users using a collaborative filtering algorithm, and based on the user portraits of different users, obtain a marketing recommendation list corresponding to different users, and integrate the data format of the marketing recommendation list to generate data assets for the precision marketing scenario, including: Extract the behavioral data of the different users; Extract the behavioral characteristics of different users from the behavioral data of different users, including the number of user clicks, browsing duration, and purchase frequency; Based on the behavioral characteristics of different users, calculate the similarity between users using a user-based collaborative filtering algorithm; Generate user portraits of different users according to the user similarity; Combined with the basic product attributes, perform a dot product calculation on the preference vector in the user portraits of different users and the basic product attribute vector to generate a matching matrix reflecting the user-basic product attribute matching degree. Each element in the matrix represents the preference intensity of a specific user for a certain basic product attribute. Among them, the basic product attributes include category attributes, physical attributes, and price attributes; Generate an initial recommendation list corresponding to different users according to the matching matrix; Extract the category click ratio data in the historical behaviors of different users from the behavioral characteristics of different users, and correspondingly sort the initial recommendation lists of different users according to the category click ratio data in the historical behaviors of different users to generate a marketing recommendation list corresponding to different users; Integrate each marketing recommendation list according to the preset data format to generate data assets for the precision marketing scenario.
6. The method according to claim 1, wherein Extract the historical transaction data, extract risk characteristics from the historical transaction data, input the risk characteristics into a preset risk assessment model to obtain a risk warning list. Among them, the risk assessment model is trained by a decision tree algorithm using historical risk characteristics, and integrate the data format of the risk warning list to generate data assets for the risk control scenario, including: Extract the historical transaction data; Clean the historical transaction data to remove missing values and outliers; Extract risk characteristics including transaction amount, transaction frequency, and transaction time interval from the cleaned historical transaction data; among them, the transaction amount is directly obtained, the transaction frequency is obtained by counting the number of transactions of the same user within 30 days, and the transaction time interval is obtained by calculating the time difference between adjacent transactions; Input the risk characteristics into the risk assessment model to calculate the risk score of each transaction; Combine the preset risk score threshold to judge the transaction risk level. A score greater than the preset risk score threshold is a high risk, and a score less than or equal to the preset risk score threshold is a low risk; Generate a risk warning list including transaction number, risk score, and risk level; Integrate the risk warning list according to the preset data format to generate data assets for the risk control scenario; Among them, a risk assessment model is pre-trained by a decision tree algorithm using historical risk characteristics, including: Pre-establish risk score labels and classify transactions into two categories: high risk and low risk; The contribution degree of historical risk features to the risk score is calculated using information gain, where the historical risk features include historical transaction amount, historical transaction frequency, and historical transaction time interval; Based on the information gain calculation results, the feature with the largest contribution degree is selected as the splitting node of the decision tree; A decision tree is generated by recursively partitioning the data set, and the minimum number of samples of the node is set as the recursive termination condition to train and generate a risk assessment model.
7. The method according to claim 1, characterized in that The product functional attribute data is extracted, the product potential features are extracted from the product functional attribute data, and the similarity between the user preference and the product is calculated based on the product potential features to obtain a product recommendation list, and the data format of the product recommendation list is integrated to generate a data asset for the product recommendation scenario, including: The product functional attribute data is obtained, and the product functional attribute data is preprocessed. The interpolation method is used to process the missing values, and the box plot method is used to remove the outliers. Among them, the product functional attribute data includes the functional characteristics of the product, the technical parameters of the product, and the usage scenario of the product; From the preprocessed product functional attribute data, the factor analysis method is used to extract the product potential features, and the principal component analysis method is used for dimensionality reduction to obtain the product potential features after dimensionality reduction. Among them, the product potential features refer to the features that can reveal the internal relationship between the product functional attributes; According to the user behavior data, the rating data of the user for the product is obtained, and combined with the product potential feature vector after dimensionality reduction, a user-product potential feature rating matrix is constructed, and each element in the matrix is the rating of the user for the product potential feature; The singular value decomposition algorithm is used to decompose the user-product potential feature rating matrix into a user feature matrix and a product potential feature matrix to obtain a user preference vector and a product potential feature weight vector. Among them, each row of the user feature matrix corresponds to the user preference vector, and each column of the product feature matrix corresponds to the product potential feature weight vector; According to the user preference vector and the product potential feature weight vector, the cosine similarity is used to calculate the similarity between the user preference and the product; The recommended products with similarity greater than the preset similarity threshold are obtained; Based on the recommended products, a product recommendation list including user ID, product ID, and recommended products is generated; According to the preset data format, the product recommendation list is integrated to generate a data asset for the product recommendation scenario.
8. The method according to claim 1, characterized in that, The marketing recommendation list, risk warning list, and product recommendation list generated from each shard data are integrated into the data service interface and provided for the business system to call to realize the real-time application of the data asset. The data service interface adopts the RESTful style, supports batch query and real-time query, and ensures the timeliness and consistency of the data, including: The data service interface is designed in the RESTful style, and the request method and parameters of the interface are defined; The interface supports GET requests, and the parameters include user ID and timestamp, which are used to distinguish batch query and real-time query; The marketing recommendation list, risk warning list, and product recommendation list generated from each shard data are obtained from the pre-established database; The marketing recommendation list, risk warning list, and product recommendation list are integrated into the data service interface and returned to the caller in JSON format; When the business system calls the data service interface, it obtains corresponding data according to the user ID and timestamp parameters; If the timestamp is the current time, a real-time query is triggered to obtain the latest data from each data source; If the timestamp is a historical time, a batch query is triggered to obtain historical data from the cache.
9. The method according to claim 1, characterized in that According to the marketing recommendation list, risk warning list, and product recommendation list, the data is updated in real time through the Redis cache mechanism to ensure that the latest data is obtained when the interface is called, and efficient query and real-time update of the data are realized, including: In the real-time query scenario, after obtaining the marketing recommendation list, the data is stored in the cache through the ZADD command of Redis SortedSet; After obtaining the risk warning list, the result is stored in the cache through the List command of Redis; After obtaining the product recommendation list, the list is stored in the cache through the RPUSH command of Redis; In the batch query scenario, the marketing recommendation list is read through the ZRANGE command of Redis, and the risk warning list and product recommendation list are read through the LRANGE command of Redis to ensure data consistency; The survival time of the data stored in the cache is set through the EXPIRE command of Redis. When the survival time expires, Redis will automatically delete the data stored in the cache.
Citation Information
Patent Citations
Intelligent data asset storage data evaluation method
CN118503236A
Augmenting search results based on relevancy and utility
US11630829B1