Intelligent service decision-making method based on multi-source heterogeneous data

By obtaining multi-source heterogeneous data in enterprise data management, cleaning and building data relationship diagrams, and using distributed computing and multiple algorithms to generate data assets for different business scenarios, multiple challenges in data management are solved, and efficient data processing and accurate business decision support are achieved.

CN119988477AActive Publication Date: 2025-05-13GUANGDONG TOPWAY NETWORK

Patent Information

Application Number
CN202510466707.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-05-13
Estimated Expiration
2045-04-15

AI Technical Summary

Technical Problem

In enterprise data management, it is difficult to achieve unified and standardized processing of multi-source heterogeneous data, construction of data ties, high availability and high performance of data processing, and personalized data services for different business scenarios.

Method used

By obtaining multi-source heterogeneous data, cleaning and screening using the data quality evaluation model, building a data relationship diagram, using a distributed computing system for parallel processing, extracting user behavior characteristics, risk characteristics and product potential characteristics, using collaborative filtering, decision tree and matrix decomposition algorithms to generate marketing recommendation lists, risk warning lists and product recommendation lists, and integrating them into the RESTful-style data service interface.

Benefits of technology

It realizes the intelligence of the entire process from data collection, processing to application, improves the efficiency of data assets utilization, provides enterprises with accurate and timely business decision-making support, and improves marketing effectiveness, risk control capabilities and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988477A_ABST
    Figure CN119988477A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of information, and provides an intelligent service decision-making method based on multi-source heterogeneous data, which comprises the following steps: acquiring multi-source heterogeneous data of different service scenes, judging whether the data meets a preset quality requirement by adopting a data quality evaluation model, and outputting available multi-source heterogeneous data; according to available multi-source heterogeneous data, a data blood relationship graph is constructed, the data copy generation, conversion and calculation process is recorded, and repeated calculation is avoided; generating a marketing recommendation list, a risk early warning list and a product recommendation list according to the data blood relationship graph; according to data assets of the three scenes, real-time data updating is achieved through a Redis cache mechanism, it is ensured that latest data are obtained through a calling interface, and efficient query and real-time updating of the data are achieved. According to the technical scheme of the invention, the whole process intelligence from data acquisition, processing to application is realized, the utilization efficiency of data assets is improved, accurate and timely business decision support is provided for enterprises, and the marketing effect, the risk management and control capability and the user experience are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information technology, and in particular to an intelligent business decision-making method based on multi-source heterogeneous data. Background Art

[0002] In enterprise data management, accurate tracking and data enrichment before data goes online is a complex technical challenge. First, the data formats and types generated by different business systems are different. How to standardize and process them uniformly during the data collection stage is the primary issue in tracking the source of data. Secondly, multiple copies of data are often generated during the production and circulation process. How to build a data lineage relationship diagram, record the data conversion and calculation process, and avoid repeated calculations and waste of resources are key points to consider in the data enrichment stage. Thirdly, large enterprises generate massive amounts of data every day. Timely processing of data and applying it to business analysis requires data platforms to have second-level response capabilities, and there is an urgent need to build a highly available and high-performance distributed computing engine. Finally, different business personnel have different concerns about data. How to design personalized data services in a targeted manner, provide data assets for different business scenarios, and meet business demands such as precision marketing, risk control, and product recommendations are issues that must be considered by the data governance platform. Summary of the invention

[0003] The present invention provides an intelligent business decision-making method based on multi-source heterogeneous data, which mainly includes: Obtain multi-source heterogeneous data from different business scenarios, use a data quality assessment model to determine whether the data meets the preset quality requirements, and use the data that meets the quality requirements as available multi-source heterogeneous data; construct a data lineage relationship diagram based on the available multi-source heterogeneous data, where the data lineage relationship diagram records the copy generation, conversion and calculation process of the available multi-source heterogeneous data during production and circulation; on the basis of the data lineage relationship diagram, shard the information of the data lineage relationship diagram to obtain multiple data shards, and use a distributed computing system to process the multiple data shards in parallel; the processing of each data shard includes: Extract the behavioral data of different users, extract the user behavior characteristics of different users from the behavioral data of different users, use the collaborative filtering algorithm to generate user portraits of different users, and based on the user portraits of different users, obtain the marketing recommendation lists corresponding to different users, integrate the data format of the marketing recommendation lists, and generate data assets for precision marketing scenarios; Extract historical transaction data, extract risk features from the historical transaction data, input the risk features into the preset risk assessment model, and obtain a risk warning list. The risk assessment model is trained by the decision tree algorithm using historical risk features. The risk warning list is integrated into data format to generate data assets for risk control scenarios. Extract product function attribute data, extract product potential features from the product function attribute data, and calculate the similarity between user preferences and products based on the product potential features to obtain a product recommendation list, integrate the product recommendation list in data format, and generate data assets for product recommendation scenarios; integrate the marketing recommendation list, risk warning list, and product recommendation list generated by each shard data into the data service interface, and provide them to the business system for calling to realize the real-time application of data assets. The data service interface adopts the RESTful style, supports batch query and real-time query, and ensures the timeliness and consistency of data; according to the marketing recommendation list, risk warning list, and product recommendation list, the Redis cache mechanism is used to realize real-time data update, ensure that the calling interface obtains the latest data, and realize efficient query and real-time update of data.

[0004] The technical solution provided by the embodiment of the present invention may have the following beneficial effects: The present invention discloses an intelligent business decision-making method based on multi-source heterogeneous data. The method ensures data quality through a data quality assessment model and a cleaning process, and constructs a data lineage relationship diagram to avoid repeated calculations. Distributed computing technology is used to efficiently process data, extract business-related features, and generate data assets for different business scenarios. For precision marketing scenarios, risk control scenarios, and product recommendation scenarios, collaborative filtering, decision trees, and matrix decomposition algorithms are used to generate marketing recommendation lists, risk assessment models, and product recommendation lists. Finally, these results are integrated into a RESTful-style data service interface for business system calls. The present invention realizes the intelligence of the entire process from data collection, processing to application, improves the utilization efficiency of data assets, provides enterprises with accurate and timely business decision-making support, and effectively improves marketing effects, risk management capabilities, and user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0005] Figure 1 The present invention is a flowchart of an intelligent business decision-making method based on multi-source heterogeneous data. DETAILED DESCRIPTION

[0006] The following will describe the technical solutions in the embodiments of the present invention in detail in conjunction with the accompanying drawings in the embodiments of the present invention. The described embodiments are only part of the embodiments of the present invention.

[0007] like Figure 1 The intelligent business decision-making method based on multi-source heterogeneous data in this embodiment may specifically include: Step S101, obtain multi-source heterogeneous data of different business scenarios, use a data quality assessment model to determine whether the data meets the preset quality requirements, and use the data that meets the quality requirements as available multi-source heterogeneous data.

[0008] The preset data acquisition module is used to obtain multi-source heterogeneous data from different business scenarios, and the data is input into the preset data quality assessment model to determine whether the data meets the preset quality requirements. If the quality requirements are met, the data that meets the quality requirements are used as available multi-source heterogeneous data. If the quality requirements are not met, the data cleaning module is triggered to clean the data, and the cleaned data is input into the quality assessment model for further judgment. If the quality requirements are met, the cleaned data is used as available multi-source heterogeneous data. If the data still does not meet the quality requirements after cleaning, the clustering analysis in the machine learning algorithm is used to classify the data, and regression analysis is used to optimize the data distribution for different categories. The optimized data is input into the quality assessment model again for judgment. If the quality requirements are met, the optimized data is used as available multi-source heterogeneous data. If the optimized data still does not meet the quality requirements, the optimized data is marked as abnormal data and stored in the abnormal database, and the data source optimization module is triggered to optimize the data source, and the data provided after the data source optimization is input into the data acquisition module for cyclic processing.

[0009] Exemplarily, the data acquisition module is the starting point of the entire data processing process. It obtains multi-source heterogeneous data from different business scenarios. Taking the e-commerce platform as an example, various types of data such as user browsing records, purchase history, and evaluation information can be collected. These data come from various sources and in different formats, and need to be processed uniformly. The data quality assessment model is a key link to ensure data availability. It is usually a multi-dimensional quantitative assessment framework that combines rule engines, statistical analysis, and machine learning technologies. It may be composed of statistical distribution models, rule engine-based verification models, correlation verification models, machine learning anomaly detection models, and text quality analysis models. The model combines rules and machine learning to ensure verification efficiency and adaptability to complex scenarios. It can evaluate data from multiple dimensions such as completeness, accuracy, and consistency. For example, for user registration information, you can check whether the required fields are empty, whether the mobile phone number format is correct, and whether the age and date of birth match. By presetting requirements, such as requiring more than 90% of the fields to meet the specifications, it is determined whether the data quality meets the standards. The data cleaning module is responsible for processing unqualified data. Common cleaning operations include deduplication, filling missing values, and correcting outliers. For example, for duplicate order records, you can remove duplicates based on the order number; for missing user ages, you can fill in the average age of the user group; for obviously abnormal commodity prices, you can replace them with historical average prices. Cluster analysis in machine learning algorithms can help you better understand the data structure. Taking customer segmentation as an example, users can be divided into different categories based on their consumption behavior, browsing habits and other characteristics, such as high-frequency low-amount users, low-frequency high-amount users, etc. This classification can help to optimize data in a targeted manner. Regression analysis is used to optimize data distribution. For example, for a sales forecasting model, you can use linear regression to analyze historical sales data, find out the key factors affecting sales, and adjust abnormal data points accordingly to make the data distribution more reasonable. Principal component analysis is a commonly used dimensionality reduction method. When processing high-dimensional data, for example, the optimized data may contain ten related indicators (such as daily active time, number of search keywords, add-to-cart rate, number of favorite products, order cancellation rate, average order value, promotion sensitivity, etc.), calculate the covariance matrix of the optimized data, and obtain the principal component directions (eigenvectors) and the proportion of variance they explain (eigenvalues) through eigenvalue decomposition. Select the eigenvectors corresponding to the first three largest eigenvalues ​​to form three principal components: principal component 1 (purchasing power): a linear combination of high-weight indicators such as the number of orders and the amount of consumption, reflecting the user's consumption ability; principal component 2 (activity): a linear combination of high-weight indicators such as the number of clicks and the length of browsing time, reflecting the user's activity level; principal component 3 (return tendency): a linear combination of high-weight indicators such as the return rate and the number of customer service complaints, reflecting the user's sensitivity to returns. The original 10-dimensional data is projected onto these three principal components, and a total of 85% of the original data variance is retained, achieving simplification from complex indicators to core features.Replacing the original 10 indicators with 3 principal component indicators retains 85% of the core information, which not only reduces the complexity of the data, but also retains key information, thereby improving the efficiency of subsequent processing. The processing of abnormal data is crucial to improving the overall data quality. Marking and storing data that cannot be processed by conventional methods can provide a reference for subsequent data source optimization. For example, after discovering a large number of abnormal user age data, it may be necessary to redesign the age input interface and add rationality verification. Data source optimization is a process of continuous improvement. By analyzing abnormal data, problems in the data collection process, such as sensor failure and human input errors, can be found. For these problems, corresponding measures can be taken, such as replacing equipment and optimizing operating procedures, to improve data quality from the source. It is understandable that the available multi-source heterogeneous data contains various formats of user behavior, product basic attributes, product functional attributes, historical transactions and other aspects of data. Although the available multi-source heterogeneous data has been initially processed, it may not have been integrated and the structure is not suitable for direct feature extraction, and needs to be processed in subsequent steps. This multi-level, cyclic iterative data processing method can effectively improve data quality and lay a solid foundation for subsequent data analysis and application. By continuously optimizing data processing processes, companies can better utilize data assets and improve decision-making efficiency and accuracy.

[0010] Step S102, constructing a data lineage relationship diagram based on the available multi-source heterogeneous data, wherein the data lineage relationship diagram records the copy generation, conversion and calculation process of the available multi-source heterogeneous data during the production and circulation process.

[0011] Obtain data replica generation information, data conversion information, and data calculation information from available multi-source heterogeneous data. Record the acquired data replica generation information in the data lineage graph. Record the data conversion information in the data lineage graph. Record the data calculation information in the data lineage graph. In the process of recording the data calculation information in the data lineage graph, detect whether there is a repeated calculation part, and if so, mark the repeated calculation part. Use a greedy algorithm to optimize the marked repeated calculation part to generate an optimized calculation information record. Update the optimized calculation information record to the data lineage graph to replace the original repeated calculation part. Through the data lineage graph construction module, continuously obtain the latest information on the data replica generation, conversion, and calculation process, and dynamically update the data lineage graph at a frequency of once a minute.

[0012] Exemplarily, a data lineage diagram is a visual tool for tracking data flow and conversion processes. It records the entire life cycle of data from the source to the final use, including the generation, conversion and calculation process of data. In practical applications, this kind of diagram is essential for understanding complex data processing processes, optimizing data management and ensuring data quality. When obtaining data copy generation information from multi-source heterogeneous data, it can be understood as extracting a snapshot from the original data for subsequent processing. For example, in an e-commerce platform, user browsing records may come from web pages, mobile terminals and mini programs. The data copy generation information records the generation time of these data, such as 2025-04-01, the location such as server A, and the source such as "browsing log table". This recording method makes it easy to track where the data comes from and ensure traceability. When recording data conversion information to the data lineage diagram, the conversion rules and methods must be clearly defined. Specifically, assuming that the amount field in the user's purchase record is converted from US dollars to RMB, the conversion rule may be "multiply by the exchange rate of 6.5", and the conversion method is field mapping. The specific conversion steps include reading the original field, applying the exchange rate, and outputting the new field. This detailed record helps to understand how data changes from one form to another, and provides a basis for subsequent audits. The record of data calculation information is more complicated. In one embodiment, the average consumption amount of a user in the past 30 days is calculated. The calculation process may be "reading purchase records, filtering time ranges, summing and dividing by the number of orders", the calculation parameters are "time range = 30 days", and the calculation result is "average consumption = 150 yuan". Recording this information in the data lineage relationship diagram can clearly display the calculation logic. It should be noted that if repeated calculations are found during the calculation process, such as calculating the total consumption of the same user multiple times, it is marked, such as "the total calculation of user ID001 is repeated 3 times". Preferably, when a greedy algorithm is used to optimize the repeated calculation part, the path with the least amount of calculation can be preferred. For example, for the total consumption of users that is repeatedly calculated, the first calculation result can be reused directly instead of re-reading the original data and calculating it each time. After this optimized calculation information record is updated to the data lineage relationship diagram, the original redundant part is replaced, thereby reducing resource waste. Exemplarily, when the data lineage relationship diagram is dynamically updated, the frequency of once per minute ensures the real-time nature of the data. In one possible implementation, suppose that an e-commerce platform has added a batch of order data, the data copy generation information records the time of addition 2025-04-01, the conversion information records the change of order status from "pending payment" to "completed", and the calculation information updates the user's total consumption. Continuously obtaining this information and updating the lineage relationship diagram can fully reflect the flow process of data. It is understandable that the construction of the data lineage relationship diagram not only records the "past" of the data, but also provides a basis for future analysis. For example, by viewing the lineage relationship diagram, you can quickly locate whether the source of a certain abnormal data is a collection problem or a conversion error.This clear context helps improve the efficiency and transparency of data management. In one embodiment, if it is found that the total consumption of a user is abnormally high, the traceability lineage diagram may show that it is caused by repeated calculations that have not been optimized in time. The optimized record shows that the total amount has been adjusted from 1,000 yuan to 500 yuan, reflecting the actual situation. This method ensures the reliability of the data through multi-level verification. Specifically, from multiple aspects, the data copy generation information provides anchor points for time and source, the conversion information clarifies the rules of data changes, the calculation information reveals the origin of the results, and the optimization process reduces redundancy. This multi-dimensional record and continuous update jointly support the integrity and practicality of the data lineage diagram. For example, managers can use the lineage diagram to quickly determine which data needs further verification and which calculations can be directly reused, thereby improving the response speed of the overall process. In general, the core functions of the data lineage diagram are as follows: 1. Solve the data island problem: The available multi-source heterogeneous data in step S101 may be independent data sets scattered in different business systems (such as CRM user data, ERP transaction data, and behavioral data in the log system). The data lineage relationship records the data flow path and establishes a logical association across systems, so that the originally isolated data can be connected in terms of business semantics. 2. Eliminate the risk of repeated calculations: In the precision marketing scenario, the construction of user portraits may require the simultaneous call of historical order data and behavioral embedding data. If calculated directly from the original data, the same user may repeat calculations such as order amount aggregation and behavioral path analysis in different scenarios. The data lineage relationship can mark the completed intermediate calculation results to achieve the reuse of the calculation results. 3. Ensure feature interpretability: The anti-fraud features in the risk control scenario (such as "the number of device changes within 7 days for users") need to track the change records of device IDs between multiple systems. The data lineage relationship can clearly show the data source path of the feature, which is convenient for compliance audits and model interpretability verification. The data lineage relationship diagram is not only a technical tool, but also a strategic asset for data management. It can improve the efficiency and quality of data processing, enhance the interpretability and traceability of data, and thus provide more reliable support for data-driven decision-making.

[0013] Step S103, based on the data lineage relationship diagram, the information of the data lineage relationship diagram is sliced ​​to obtain multiple data slices, and a distributed computing system is used to process the multiple data slices in parallel.

[0014] Extract the latest information of data copy generation, conversion and calculation process from the data lineage relationship diagram, use the key fields in the latest information of data copy generation, conversion and calculation process as the input of the hash algorithm, obtain multiple different shard indexes, and distribute the data corresponding to the same shard index together to form a data shard. Use ApacheSpark to evenly distribute multiple data shards to different preset computing nodes for parallel computing, where each computing node independently processes its assigned data shard. Obtain the shard computing status of each computing node. If the shard computing status of a computing node is data skew or processing bottleneck, use ApacheSpark to evenly distribute the data of the computing node to the computing node and multiple backup computing nodes for parallel computing again.

[0015] For example, the data lineage diagram is a key tool for tracking data flow and conversion, and extracting the latest information from it lays the foundation for subsequent processing. Taking the e-commerce platform as an example, the diagram may contain multiple dimensions such as user browsing history, product information, and order data. When extracting this information, it is necessary to pay attention to the generation time, source, and conversion process of the data to provide a clear context for subsequent analysis. The hash algorithm plays an important role in data sharding to ensure that the data is evenly distributed. It is crucial to select appropriate key fields as input, and input the key fields such as user ID or product category in the latest information of the extracted data copy generation, conversion, and calculation process into the hash algorithm. The output of the hash algorithm is usually a fixed-length integer value (such as a 32-bit or 64-bit integer), which is converted into a specific shard number through further calculation (such as modulo operation), and the shard number is used as the shard index to guide data allocation. For example, for user behavior data, the user ID can be selected as the hash input to obtain a shard index with the user ID as the key field, and the data of the same user can be allocated to the shard with the same number for subsequent calculation. As a distributed computing system, ApacheSpark evenly distributes multiple data shards to different computing nodes for parallel computing. This distribution not only balances the computing load, but also improves data locality and reduces cross-node data transmission. Parallel computing is the core advantage of ApacheSpark, which means that multiple computing nodes simultaneously process the assigned data and fully utilize computing resources through distributed task scheduling. Through parallel computing, the system can simultaneously process data from multiple user groups, significantly shortening the response time. It should be noted that ApacheSpark will generate the shard computing status of each node during the parallel computing process. The shard computing status refers to the intermediate or final records generated after the data shard is parallelly computed by each computing node. If the shard calculation status of a computing node is data skew or processing bottleneck, in a possible implementation method, a new field with a granularity smaller than the key field is selected and re-entered into the hash algorithm to obtain a new shard index of the data in the computing node, and the data corresponding to the same new shard index is allocated together to form at least one new data shard, wherein the new field with finer granularity refers to a new field composed of new data attributes (such as time, sub-classification, random factors) added on the basis of the original key field, and the logical division of the formed data shard is made more fragmented and dispersed by adding field dimensions, thereby reducing the amount of data in a single shard. For example, assuming that the above steps form 20 data shards (numbered S1-S20), and 10 computing nodes (N1-N10) are preset, when node N3 processes data shard S9, the processing time may be much longer than other nodes due to the large amount of data of a certain user, which is manifested as data skew. Exemplarily, the system detects that the CPU occupancy rate of N3 is continuously higher than 90%, while that of other nodes is only 50%, confirming that it is a processing bottleneck.Preferably, for computing nodes with data skew or processing bottlenecks, the system will select fields with finer granularity for re-sharding, such as user ID combined with browsing timestamp, as new inputs to the hash algorithm, generate new sharding indexes, and assign data of different time periods of the same user to one new shard. Exemplarily, the original shard S9 may be split into S9-1 and S9-2, each containing records of different time periods. Through Apache Spark, multiple new data shards are evenly distributed to the original computing nodes and multiple standby computing nodes, such as N9 and N10, for re-parallel calculation. The introduction of standby nodes ensures the elastic expansion of computing resources. For example, after re-sharding, the new shards S9-1 and S9-2 are evenly distributed to the original computing node N3 and the standby computing nodes N9 and N10 for re-parallel calculation, and the load of N3 drops to 60%, while the processing time of N3, N9 and N10 is consistent with the average level of other computing nodes. If the shard computing status of a computing node is data tilt or processing bottleneck, in another possible implementation method, the data of the computing node is evenly distributed to the node and multiple standby computing nodes through Apache Spark for re-parallel computing. For example, the data of the original computing node is evenly distributed to the original computing node and multiple standby computing nodes through Apache Spark for re-parallel computing, such as evenly distributing the data of the original shard S9 to the original computing node N3 and the standby computing nodes N9 and N10 for re-parallel computing. This refined adjustment can better balance the computing load. Through continuous optimization, the system can eventually achieve a second-level response. The whole process reflects the closed-loop optimization of data processing: obtaining information from the data lineage relationship diagram, generating shard indexes through hash algorithms, using Apache Spark for parallel computing, and then adjusting the strategy according to the feedback of Apache Atlas to ensure that the data has a second-level response capability during the circulation and calculation process. It can be understood that the advantage of this method is that it can quickly respond to abnormal situations in data processing. For example, when the amount of data in a certain commodity category surges, the system can quickly adjust by recalculating the shard index to avoid overloading a single node. This flexibility is particularly important for large-scale data processing on e-commerce platforms, and can significantly improve the response speed and accuracy of the recommendation system. In general, the core value of distributed computing systems is as follows: First, it handles complex data topology: When the product recommendation scenario requires the construction of a product similarity matrix, it involves the joint calculation of product attribute data (structured), user comments (unstructured), and click stream graph data (graph structure). Distributed systems achieve parallel processing of multimodal data through sharding strategies (such as sharding by product ID). Second, it optimizes the computing process: user behavior sequence analysis (such as session partitioning) requires time window sliding calculations.Distributed systems can ensure that the behavior data of the same user is concentrated in the same data shard through pre-sharding (such as sharding by user ID hash), avoiding performance loss caused by cross-node data transmission. Third, it supports dynamic expansion capabilities: in traffic peak scenarios such as "Double 11", based on the data dependencies marked in the data lineage relationship graph, the computing resources of specific feature calculation links (such as real-time user preference feature calculation clusters) can be dynamically expanded without full-link expansion. For example, taking the "association network features" in risk control as an example: directly using step S101 data: it is necessary to reconstruct the multi-layer association network of user-merchant-geographic location from the original transaction record each time, and a single calculation takes more than 2 hours. After processing through steps S102-S103: based on the lineage relationship to identify the existing intermediate calculation results, through the distributed graph calculation incremental update, the feature generation time is shortened to 15 minutes. This architectural design enables the system to control the feature production P99 delay within 5 minutes while maintaining an average daily processing level of more than 1 billion data items, and reduce computing resource consumption by 40%. This is exactly the engineering practice value brought by steps S102-S103, which far exceeds the optimization effect that can be achieved by simple data cleaning.

[0016] Step S104, extracting the behavioral data of different users, extracting the user behavioral features of different users from the behavioral data of different users, using a collaborative filtering algorithm to generate user portraits of different users, and based on the user portraits of different users, obtaining marketing recommendation lists corresponding to different users, integrating the data format of the marketing recommendation lists, and generating data assets for precision marketing scenarios.

[0017] Extract the behavioral data of the different users; extract the behavioral characteristics of the different users from the behavioral data of the different users, including the number of user clicks, browsing time and purchase frequency. Based on the behavioral characteristics of different users, a user-based collaborative filtering algorithm is used to calculate the similarity between users. Generate user portraits of different users according to the user similarity. Combined with the basic attributes of the product, the preference vectors in the user portraits of different users are dot-producted with the basic attribute vectors of the product to generate a matching matrix reflecting the user-product basic attribute matching degree, and each element in the matrix represents the preference intensity of a specific user for a basic attribute of a certain product, wherein the basic attributes of the product include category attributes, physical attributes and price attributes. Generate an initial recommendation list corresponding to different users according to the matching matrix. Extract the category click ratio data in the historical behaviors of different users from the behavioral characteristics of different users, sort the initial recommendation lists of different users according to the category click ratio data in the historical behaviors of different users, and generate marketing recommendation lists corresponding to different users. Integrate the marketing recommendation lists according to the preset data format to generate data assets for precision marketing scenarios.

[0018] For example, user behavior data is the basis of precision marketing. The time window statistics method is used to extract the number of user clicks in the last 7 days, and the sliding time window is used to quantify the short-term interest intensity of users to reflect the timeliness of behavior; the page type association analysis method is used to extract the browsing time characteristics by distinguishing the browsing depth of different page types, and the industry dynamic frequency calculation method is used to extract the purchase frequency characteristics by counting the number of purchases based on the sliding window. User behavior characteristics such as user clicks, browsing time and purchase frequency can fully characterize the user's online activity pattern. For example, an e-commerce platform may find that user A has browsed mobile phone products 50 times in the past month, staying for an average of 3 minutes each time, and finally bought 2 mobile phones. These data reflect user A's high interest in mobile phone products. The user-based collaborative filtering algorithm realizes personalized recommendations by calculating the similarity between users, which can quantify the similarity of user behavior. Assuming that user A and user B both browse smartphones frequently and have similar purchasing patterns, the system will consider them to have a high degree of similarity. This similarity calculation lays the foundation for subsequent user portraits and product recommendations. User portraits are a comprehensive description of user preferences. By analyzing the browsing and purchase history of user A, the following profile may be obtained: 25-35 years old, technology enthusiast, preference for high-end smartphones, and special attention to photography functions. Such a profile can guide the formulation of more accurate marketing strategies. Constructing a matching matrix between user preferences and basic product attributes is the core link of the recommendation system. Taking smartphones as an example, the matrix may contain dimensions such as price, brand, screen size, and camera pixels. User A's preference may be highly valued for camera performance and relatively insensitive to price. The preference vector in the user A profile (highly valued for camera performance and relatively insensitive to price) is dot-producted with the product basic attribute vector to generate a matching matrix that reflects the user-product basic attribute matching degree. Each element in the matrix represents the preference intensity of a specific user for a certain product basic attribute. This matching relationship provides a basis for generating the initial recommendation list. The generation of the initial recommendation list takes into account the matching degree between user preferences and product basic attributes. For user A, the system may recommend 10 high-end smartphones, including the latest flagship models and professional photography phones. This list reflects the system's preliminary judgment of user needs. Sorting the initial recommendation list is a key step in optimizing user experience. The system may rank mobile phones with high-pixel cameras at the top of the list based on the percentage of browsing and clicking on mobile phones with strong camera functions in user A's historical behavior, and generate a marketing recommendation list for user A. The marketing recommendation list needs to be integrated according to the preset data format to form a data asset for precision marketing scenarios. This may include information such as user ID, recommended product list, and the reason for recommending each product.For example, the marketing recommendation list for user A may include a personalized description such as "Based on your interest in high-quality photography, we recommend the latest flagship phone of brand XX, whose 100-megapixel main camera will bring you a professional-level shooting experience." The generated data assets are stored in a dedicated database to support subsequent marketing activities. This database may use distributed storage technology to ensure high availability and fast access to data. By regularly updating and analyzing this data, companies can continuously optimize their recommendation algorithms and marketing strategies to improve user satisfaction and conversion rates.

[0019] Step S105, extract historical transaction data, extract risk features from the historical transaction data, input the risk features into a preset risk assessment model, and obtain a risk warning list, wherein the risk assessment model is trained using a decision tree algorithm using historical risk features, and the risk warning list is integrated into a data format to generate data assets for risk control scenarios.

[0020] Extract the historical transaction data. Clean the historical transaction data to remove missing values ​​and outliers. Extract risk features including transaction amount, transaction frequency and transaction time interval from the cleaned historical transaction data; wherein, the transaction amount is directly obtained, the transaction frequency is obtained by counting the number of transactions of the same user within 30 days, and the transaction time interval is obtained by calculating the time difference between adjacent transactions. Input the risk features into the risk assessment model and calculate the risk score for each transaction; judge the transaction risk level in combination with the preset risk score threshold, a score greater than the preset risk score threshold is high risk, and a score less than or equal to the preset risk score threshold is low risk; generate a risk warning list containing transaction number, risk score and risk level; integrate the risk warning list according to the preset data format to generate data assets for risk control scenarios; wherein, use historical risk features to pre-train a decision tree algorithm to generate a risk assessment model, including: pre-establishing risk score labels to divide transactions into high-risk and low-risk categories; Information gain is used to calculate the contribution of historical risk features to the risk score, where historical risk features include historical transaction amount, historical transaction frequency and historical transaction time interval; according to the information gain calculation result, the feature with the largest contribution is selected as the split node of the decision tree; the decision tree is generated by recursively partitioning the data set, and the minimum number of node samples is set as the recursive termination condition to train and generate a risk assessment model.

[0021] For example, obtaining historical transaction data is the first step in risk assessment. This data may include information such as user ID, transaction amount, and transaction time. During the data cleaning process, the following situation may be encountered: a user has multiple abnormally large transactions in one day, which may be a data entry error or potential fraud. By setting a reasonable threshold, such as the total transaction amount in a single day not exceeding 10 times the user's monthly income, such outliers can be effectively identified and processed. Extracting risk features is the basis for building a risk assessment model. The transaction amount directly reflects the scale of the transaction and is crucial for risk assessment. The transaction frequency can reveal the user's behavior pattern. For example, an ordinary salaried worker's account suddenly frequently makes large transfers in a short period of time, which may indicate abnormal capital flow. The transaction time interval is also important. If an account makes two large transfers in a row at 3 am and 3:05 am, this unusual time pattern may indicate automated fraud. The historical risk features are used to pre-train the decision tree algorithm to generate a risk assessment model. Specifically, the risk score label is pre-established to provide a training target for the model. High-risk transactions may include large overseas transfers, frequent small transfers, etc., while low-risk transactions may be routine transactions such as salary payments and daily consumption. Information gain is used to calculate the contribution of transaction amount, transaction frequency and transaction time interval to risk scoring. Information gain is the core indicator for feature selection in the decision tree algorithm. Its essence is to evaluate the distinguishing ability of the feature by measuring the degree to which the feature reduces the uncertainty of data classification. In the risk control scenario, information gain is used to quantify the contribution of the three features of transaction amount, transaction frequency and transaction time interval to risk scoring. For example, it may be found that the information gain of the transaction amount is the highest, which means that the transaction amount is the most effective feature for distinguishing high-risk and low-risk transactions. The process of building a decision tree is a process of gradually refining the risk judgment criteria. The decision tree is generated by recursively dividing the data set, setting the minimum number of node samples as the recursive termination condition, and combining historical transaction data and pre-established risk score labels to train and generate a risk assessment model. In one embodiment, the decision tree may first be divided into two groups of less than 1,000 yuan and greater than 1,000 yuan according to the transaction amount, and then refined according to the transaction frequency to finally generate a risk assessment model. When calculating the risk score, the model outputs a numerical value based on the feature value. For example, if the transaction amount of user D is 3,000 yuan and the frequency is 3 times / 30 days, the score may be 0.7. Combined with the preset risk score threshold such as 0.6, transactions greater than 0.6 are judged as high risk. Such a scoring mechanism intuitively reflects the potential risks. The risk warning list finally generated is an important tool for risk management. The risk warning list needs to be integrated according to the preset data format to form a data asset for risk control scenarios. It can be determined that data assets for risk control scenarios are risk management tools for enterprises.It may contain the following information: transaction number T20240622001, risk score 85, high risk level. Such data assets enable e-commerce platforms to quickly identify and process high-risk transactions. By continuously updating and optimizing this risk assessment model, e-commerce platforms can improve the processing efficiency of normal transactions while ensuring security and provide customers with a better service experience. The generated data assets are stored in a dedicated database to provide support for subsequent risk control. This database may use distributed storage technology to ensure high availability and fast access to data.

[0022] Step S106, extract product function attribute data, extract product potential features from the product function attribute data, and calculate the similarity between user preferences and products based on the product potential features to obtain a product recommendation list, integrate the data format of the product recommendation list, and generate data assets for product recommendation scenarios.

[0023] The product function attribute data is obtained, the product function attribute data is preprocessed, the missing values ​​are processed by interpolation method, and the outliers are removed by box plot method, wherein the product function attribute data includes the functional characteristics of the product, the technical parameters of the product and the usage scenarios of the product. From the preprocessed product function attribute data, the factor analysis method is used to extract the product potential features, and the principal component analysis method is used to reduce the dimension to obtain the reduced product potential features, wherein the product potential features refer to the features that can reveal the intrinsic connection between the product function attributes. According to the user behavior data, the user's product rating data is obtained, and the user-product potential feature rating matrix is ​​constructed in combination with the reduced product potential feature vector, wherein each element in the matrix is ​​the user's rating of the product potential feature. The singular value decomposition algorithm is used to decompose the user-product potential feature rating matrix into a user feature matrix and a product potential feature matrix, and the user preference vector and the product potential feature weight vector are obtained, wherein each row of the user feature matrix corresponds to the user preference vector, and each column of the product feature matrix corresponds to the product potential feature weight vector. According to the user preference vector and the product potential feature weight vector, the cosine similarity is used to calculate the similarity between the user preference and the product. Get recommended products whose similarity is greater than the preset similarity threshold. Based on the recommended products, generate a product recommendation list containing user ID, product ID and recommended products; integrate the product recommendation list according to the preset data format to generate data assets for product recommendation scenarios.

[0024] Exemplarily, the construction of a product recommendation list begins with data acquisition and preprocessing. Extracting product functional attribute data is a key step. These data may include features such as product functional characteristics, product technical parameters, and product usage scenarios. In actual applications, product functional attribute data may have missing values ​​and outliers, which need to be cleaned and processed. Interpolation is a common method for processing missing values. For example, for missing waterproof rating data of a certain mobile phone, the waterproof rating of other models of the same brand and series can be used to fill it. The box plot rule effectively identifies outliers. For example, the battery capacity of a tablet computer is far beyond the normal range, which may be a data entry error and should be eliminated. The preprocessed data lays the foundation for subsequent analysis. Factor analysis can extract product potential features that can reveal the intrinsic connection between product functional attributes from many variables. For example, when analyzing smart watches, it may be found that the surface features of battery capacity, battery life, and charging speed actually reflect a common product potential feature: battery performance. Principal component analysis further reduces the data dimension and generates the most representative product potential features after dimensionality reduction. This is particularly important when dealing with high-dimensional data. For example, smart home products may involve dozens or even hundreds of parameters. Principal component analysis can generate new key features with a smaller number, and use the generated new key features to replace the original hundreds of parameters. When obtaining rating data based on user behavior data, it is understandable that user browsing, purchasing or commenting behavior can be converted into ratings. For example, user A frequently purchases large-screen mobile phones, so it can be inferred that he has a high score for "display effect", which is set at 4.5 points. Combined with the product potential feature vector after dimensionality reduction, a user-product potential feature rating matrix is ​​constructed. In the matrix, rows represent users and columns represent features. For example, user A scored a mobile phone's "display effect" of 4.5 and a "battery life" score of 3.0. Using the singular value decomposition algorithm, the user-product potential feature rating matrix (R) is decomposed into the product of three matrices R=UΣV T , where U is an m×m orthogonal user feature matrix (left singular vector), representing user preferences; Σ is an m×n diagonal matrix, with diagonal elements being singular values ​​(sorted from large to small), reflecting feature importance; V T is an n×n orthogonal product potential feature matrix (right singular vector transpose), representing the product potential feature space. After decomposition, dimensionality reduction and truncation are performed: retain the first k largest singular values ​​(such as k=50), truncate U to m×k, Σ to k×k, and V T =k×n, and we get the user feature matrix (U=m×k) and the product potential feature matrix (V T=k×n). Finally, each row of the user feature matrix is ​​output to represent the user preference vector, and each column of the product feature matrix is ​​output to represent the product attribute weight vector, which is used for similarity calculation. For example, the preference vector of user A may be [0.8, 0.3, 0.5], reflecting his preference for display, battery life, and photography; the product potential feature weight vector of a certain mobile phone may be [0.9, 0.4, 0.6], indicating its performance in these three aspects. The cosine similarity method is used to calculate the similarity between the two, and the result is similarity = 0.998, which is higher than the similarity threshold of 0.7, then the mobile phone is marked as a recommended product. When generating a product recommendation list based on the recommended product, specifically, the list may include user number U001, product number P2025, and recommended product "a certain brand of large-screen and high-battery mobile phone". Preferably, the preset data format integrates the product recommendation list to generate a data asset for the product recommendation scenario. This data asset supports personalized recommendations and improves user experience. It should be noted that the combination of factor analysis and singular value decomposition can mine implicit relationships between attributes and recommend products that better meet user needs. For example, if a user prefers a large screen but does not care about taking photos, the system can accurately recommend models with prominent screens and ordinary cameras. This method can effectively improve matching efficiency in data-driven recommendation scenarios. It is understandable that unlike the marketing recommendation list that recommends products based on basic product attributes, the product recommendation list makes more in-depth recommendations based on product functional attributes and the similarity between user preferences and products.

[0025] Step S107, integrate the marketing recommendation list, risk warning list and product recommendation list generated by each shard data into the data service interface, and provide it to the business system for calling to realize the real-time application of data assets. The data service interface adopts RESTful style, supports batch query and real-time query, and ensures the timeliness and consistency of data.

[0026] The data service interface is designed in RESTful style, and the request method and parameters of the interface are defined. The interface supports GET requests, and the parameters include user ID and timestamp, which are used to distinguish batch queries from real-time queries. The marketing recommendation list, risk warning list, and product recommendation list generated by each shard data are obtained from the pre-established database. The marketing recommendation list, risk warning list, and product recommendation list are integrated into the data service interface and returned to the caller in JSON format. When the business system calls the data service interface, it obtains the corresponding data based on the user ID and timestamp parameters. If the timestamp is the current time, a real-time query is triggered to obtain the latest data from each data source. If the timestamp is a historical time, a batch query is triggered to obtain historical data from the cache.

[0027] For example, the RESTful-style data service interface design provides a standardized method for communication between systems. Taking the user portrait service as an example, an interface " / user-profile / {userId}" can be defined to obtain the portrait data of a specific user through a GET request, where user-profile means "user information" or "user profile", and {userId} means a path parameter, which is used to dynamically specify the specific user to be operated. The timestamp parameter "timestamp" is used to distinguish between real-time queries and batch queries. For example, "?timestamp=2025022012" means obtaining historical data at that time point. The design of interface parameters directly affects the efficiency and flexibility of data acquisition. User ID as a path parameter makes the request more semantic, while timestamp as a query parameter provides the ability to access temporal data. This design allows business systems to flexibly obtain real-time or batch data according to needs. The marketing recommendation list, risk warning list, and product recommendation list pre-stored in the database constitute a complete user data ecosystem. The marketing recommendation list may include attributes such as age, occupation, and consumption habits; the risk warning list may identify the user's credit risk level; and the product recommendation list generates personalized suggestions based on user characteristics and historical behavior. As a carrier of data transmission, JSON format has the advantages of being lightweight and easy to parse. A typical JSON response may be as follows: {"userId":"12345","age":30,"occupation":"engineer","riskLevel":"low","recommendations":["product A","product B"]}, where "userId" refers to a string or number used to uniquely identify a user; "age" refers to age; "occupation" refers to occupation; "riskLevel" refers to risk level; and "recommendations" refers to a recommendation list. This format is both intuitive and easy for front-end applications to process. The distinction between real-time query and batch query reflects the performance optimization considerations of the system. When the timestamp is the current time, the system obtains the latest data from each original data source to ensure the real-time nature of the data. For example, if a user has just completed a large transaction, the risk warning system can immediately reflect this change. For batch queries of historical data, the system quickly retrieves from the cache, greatly improving the response speed. The Redis cache mechanism plays a key role in this design. It not only speeds up data access, but also enables real-time data updates. For example, when a user profile changes, the system will immediately update the corresponding data in Redis. This ensures that even in high-concurrency scenarios, the business system can obtain the latest and most accurate user information. The advantage of this design is that it balances data real-time and system performance.By making proper use of cache and real-time query, the system can meet the demand for the latest data while maintaining high efficiency when processing large amounts of historical data queries. This is especially important for e-commerce systems that need to handle both real-time transactions and historical data analysis. In general, this RESTful API design combined with a cache mechanism not only provides a standardized data access interface, but also achieves efficient, flexible, and reliable data services through intelligent query strategies and cache update mechanisms. This lays the foundation for building large-scale, responsive, and scalable systems.

[0028] Step S108, based on the marketing recommendation list, risk warning list and product recommendation list, the data is updated in real time through the Redis cache mechanism to ensure that the calling interface obtains the latest data and realizes efficient query and real-time update of the data.

[0029] In the real-time query scenario, after obtaining the marketing recommendation list, use the Redis SortedSet ZADD command to store the data in the cache. After obtaining the risk warning list, use the Redis List command to store the result in the cache. After obtaining the product recommendation list, use the Redis RPUSH command to store the list in the cache. In the batch query scenario, use the Redis ZRANGE command to read the marketing recommendation list, and use the Redis LRANGE command to read the risk warning list and product recommendation list to ensure data consistency. Use the Redis EXPIRE command to set the expiration time of the data stored in the cache. When the expiration time expires, Redis will automatically delete the data stored in the cache.

[0030] For example, in real-time query and batch query scenarios, Redis's various data structures and commands provide support for efficient storage and reading of data. The following is an analysis of the storage and reading mechanisms of marketing recommendation lists, risk warning lists, and product recommendation lists, and uses examples from multiple aspects to illustrate the specific implementation methods, highlighting the business logic and effects of each topic. For the storage and reading of marketing recommendation lists, Redis's Sorted Set uses the ZADD command to store data by priority, which is suitable for dynamically adjusting the recommendation order in real-time scenarios. For example, assume that an e-commerce system generates personalized marketing activities for users, and the recommended content includes coupons, recommended products, etc. Each time a recommendation is generated, the system calculates the score of each recommendation based on the user's recent behavior, such as browsing history or transaction frequency, for example, browsing the recommended product page adds 0.5 points, and completing the transaction adds 1 point. Sorted Set stores the recommended items with the user ID as the key and the score as the sorting basis. In batch queries, the ZRANGE command can quickly extract by score range, for example, extracting recommended items with scores between 0.5 and 2.0. This method supports fast screening of high-priority recommendations, ensuring that the business system can display the most relevant marketing content in a timely manner. It should be noted that the storage of the risk warning list uses the List structure of Redis, and the latest warning result is appended to the head of the list through the LPUSH command, which is suitable for real-time recording of user risk status. In one possible implementation, a user triggers a high-risk warning due to frequent large-amount transfers recently. The LPUSH command pushes this record into the List with the user ID as the key. In batch queries, LRANGE extracts data from the List by index range. For example, the system queries the user's most recent 10 risk warning records, and LRANGE returns the results with the key "risk:user123:alert" and the index range 0 to 9. The LRANGE command can also extract recommended items in a specified range, such as extracting the most recent 5 recommendations. This method is suitable for business systems to display recommended content in chronological order, with clear logic and easy maintenance. It is understandable that in order to achieve efficient query and real-time update of data, it is also necessary to cooperate with the Redis EXPIRE command to set the data expiration time. For example, the real-time query cache is set to "EXPIREmarketing:U100863600", which means that the data will expire after 1 hour, and the data in the cache will be deleted forcibly to maintain real-time performance; the batch query cache can be set to "EXPIRErisk:U1008686400" (deleted after 24 hours) to reduce the load caused by frequent refreshes. This mechanism balances real-time performance and system resource usage. In one embodiment, if user "U10086" has just completed a transaction, the real-time query immediately updates the cache, while the batch query still uses the snapshot data of the previous day, with clear logic and high efficiency. For example, in a high-concurrency scenario, the cache expiration time can be adjusted according to user activity.For active users, a short expiration time (such as 1 hour) is set to ensure frequent updates; for low-activity users, a long expiration time (such as 12 hours) is set to reduce unnecessary refreshes. This differentiated strategy improves resource utilization while ensuring efficient query and real-time update of data.

[0031] The above only lists some preferred embodiments of the present invention, but the present invention is not limited thereto, and many improvements and changes can be made. As long as the improvements and changes are made on the basis of the basic principles of the present invention, they should be regarded as falling within the protection scope of the present invention.

Claims

1. An intelligent business decision-making method based on multi-source heterogeneous data, characterized in that: The method comprises: Obtain multi-source heterogeneous data for different business scenarios, use data quality assessment models to determine whether the data meets the preset quality requirements, and use data that meets the quality requirements as available multi-source heterogeneous data; Based on the available multi-source heterogeneous data, a data lineage relationship diagram is constructed, wherein the data lineage relationship diagram records the copy generation, conversion and calculation process of the available multi-source heterogeneous data during the production and circulation process; Based on the data lineage relationship graph, the information of the data lineage relationship graph is processed in pieces to obtain multiple data pieces, and a distributed computing system is used to process the multiple data pieces in parallel; The processing of each data shard includes: Extract the behavioral data of different users, extract the user behavior characteristics of different users from the behavioral data of different users, use the collaborative filtering algorithm to generate user portraits of different users, and based on the user portraits of different users, obtain the marketing recommendation lists corresponding to different users, integrate the data format of the marketing recommendation lists, and generate data assets for precision marketing scenarios; Extract historical transaction data, extract risk features from the historical transaction data, input the risk features into the preset risk assessment model, and obtain a risk warning list. The risk assessment model is trained by the decision tree algorithm using historical risk features. The risk warning list is integrated into data format to generate data assets for risk control scenarios. Extract product function attribute data, extract product potential features from the product function attribute data, and calculate the similarity between user preferences and products based on the product potential features to obtain a product recommendation list, integrate the product recommendation list into a data format, and generate data assets for product recommendation scenarios; The marketing recommendation list, risk warning list, and product recommendation list generated by each shard data are integrated into the data service interface and provided to the business system for calling to realize the real-time application of data assets. The data service interface adopts the RESTful style and supports batch query and real-time query to ensure the timeliness and consistency of data. According to the marketing recommendation list, risk warning list and product recommendation list, real-time data update is achieved through the Redis cache mechanism to ensure that the calling interface obtains the latest data and realizes efficient data query and real-time update.

2. The method according to claim 1, characterized in that The method of obtaining multi-source heterogeneous data in different business scenarios, using a data quality assessment model to determine whether the data meets the preset quality requirements, and using the data that meets the quality requirements as available multi-source heterogeneous data includes: Use the preset data acquisition module to obtain multi-source heterogeneous data from different business scenarios, input the data into the preset data quality assessment model, and determine whether the data meets the preset quality requirements; If the quality requirements are met, the data meeting the quality requirements will be used as available multi-source heterogeneous data; If the data does not meet the quality requirements, the data cleaning module is triggered to clean the data. The cleaned data is input into the quality assessment model for further judgment. If the data meets the quality requirements, the cleaned data is used as available multi-source heterogeneous data. If the data still does not meet the quality requirements after cleaning, cluster analysis in the machine learning algorithm is used to classify the data, and regression analysis is used to optimize the data distribution for different categories. The optimized data is input into the quality assessment model again for judgment. If it meets the quality requirements, the optimized data is used as available multi-source heterogeneous data; If the optimized data still does not meet the quality requirements, the optimized data will be marked as abnormal data and stored in the abnormal database. At the same time, the data source optimization module will be triggered to optimize the data source, and the data provided after the data source optimization will be input into the data acquisition module for cyclic processing.

3. The method according to claim 1, characterized in that The data lineage relationship diagram is constructed based on the available multi-source heterogeneous data, wherein the data lineage relationship diagram records the copy generation, conversion and calculation process of the available multi-source heterogeneous data in the production and circulation process, including: Obtain data copy generation information, data conversion information, and data calculation information from available multi-source heterogeneous data; Record the acquired data copy generation information into the data lineage relationship diagram; Record data conversion information into the data lineage diagram; Record data calculation information into the data lineage relationship diagram; In the process of recording data calculation information into the data lineage relationship diagram, detecting whether there is a repeated calculation part, and if so, marking the repeated calculation part; A greedy algorithm is used to optimize the marked repeated calculation parts to generate optimized calculation information records; Update the optimized calculation information record to the data lineage relationship diagram to replace the original repeated calculation part; By building a module through the data lineage relationship graph, we can continuously obtain the latest information on the data copy generation, conversion and calculation process, and dynamically update the data lineage relationship graph at a frequency of once a minute.

4. The method according to claim 1, characterized in that: Based on the data lineage relationship graph, the information of the data lineage relationship graph is processed in pieces to obtain multiple data pieces, and a distributed computing system is used to process the multiple data pieces in parallel, including: Extract the latest information of data copy generation, conversion and calculation process from the data lineage relationship graph, use the key fields in the latest information of data copy generation, conversion and calculation process as the input of the hash algorithm, obtain multiple different shard indexes, and allocate the data corresponding to the same shard index together to form a data shard; Apache Spark evenly distributes multiple data shards to different preset computing nodes for parallel computing, where each computing node independently processes its assigned data shards; Get the shard computing status of each computing node. If the shard computing status of a computing node is data skew or processing bottleneck, use Apache Spark to evenly distribute the data of the computing node to the computing node and multiple standby computing nodes for re-parallel computing.

5. The method according to claim 1, characterized in that The extracting of behavioral data of different users, extracting user behavioral features of different users from the behavioral data of different users, using collaborative filtering algorithms to generate user portraits of different users, and obtaining marketing recommendation lists corresponding to different users based on the user portraits of different users, integrating the marketing recommendation lists in data format, and generating data assets for precision marketing scenarios, including: Extracting behavioral data of the different users; Extract behavioral characteristics of different users from their behavioral data, including number of clicks, browsing time, and purchase frequency; Based on the behavioral characteristics of different users, a user-based collaborative filtering algorithm is used to calculate the similarity between users; Generate user profiles of different users based on user similarity; Combined with the basic attributes of the product, the preference vectors in the user portraits of different users are dot-producted with the basic attribute vectors of the product to generate a matching matrix that reflects the matching degree between the user and the basic attributes of the product. Each element in the matrix represents the strength of preference of a specific user for a basic attribute of a product. The basic attributes of the product include category attributes, physical attributes, and price attributes. Generate an initial recommendation list corresponding to different users based on the matching matrix; Extract the category click share data in different users' historical behaviors from the behavioral characteristics of different users, sort the initial recommendation lists of different users accordingly based on the category click share data in different users' historical behaviors, and generate marketing recommendation lists corresponding to different users; Integrate various marketing recommendation lists according to the preset data format to generate data assets for precision marketing scenarios.

6. The method according to claim 1, characterized in that The historical transaction data is extracted, risk features are extracted from the historical transaction data, and the risk features are input into a preset risk assessment model to obtain a risk warning list, wherein the risk assessment model is trained by a decision tree algorithm using historical risk features, and the risk warning list is integrated in data format to generate data assets for risk control scenarios, including: Extracting the historical transaction data; Clean historical transaction data to remove missing values ​​and outliers; Extract risk features including transaction amount, transaction frequency and transaction time interval from the cleaned historical transaction data; the transaction amount is directly obtained, the transaction frequency is obtained by counting the number of transactions of the same user within 30 days, and the transaction time interval is obtained by calculating the time difference between adjacent transactions; Input the risk characteristics into the risk assessment model to calculate the risk score for each transaction; The transaction risk level is determined based on the preset risk score threshold. A score greater than the preset risk score threshold indicates high risk, and a score less than or equal to the preset risk score threshold indicates low risk. Generate a risk warning list containing transaction number, risk score and risk level; Integrate the risk warning list according to the preset data format to generate data assets for risk control scenarios; Among them, the risk assessment model is generated by pre-training the decision tree algorithm using historical risk characteristics, including: Pre-establish risk scoring labels to categorize transactions into high-risk and low-risk categories; Information gain is used to calculate the contribution of historical risk features to risk scores, where historical risk features include historical transaction amounts, historical transaction frequencies, and historical transaction time intervals; According to the information gain calculation results, the feature with the largest contribution is selected as the split node of the decision tree; A decision tree is generated by recursively partitioning the data set, setting the minimum number of node samples as the recursive termination condition, and training to generate a risk assessment model.

7. The method according to claim 1, characterized in that The extracting of product function attribute data, extracting product potential features from the product function attribute data, and calculating the similarity between user preferences and products based on the product potential features to obtain a product recommendation list, integrating the product recommendation list into a data format, and generating data assets for product recommendation scenarios, including: Obtaining the product function attribute data, preprocessing the product function attribute data, using interpolation to process missing values, and using a box plot method to remove outliers, wherein the product function attribute data includes product functional characteristics, product technical parameters, and product usage scenarios; From the preprocessed product function attribute data, factor analysis is used to extract product potential features, and principal component analysis is used to reduce the dimension to obtain the reduced product potential features, where product potential features refer to features that can reveal the internal connection between product function attributes; Based on user behavior data, obtain user rating data on products, and build a user-product potential feature rating matrix based on the reduced dimension product potential feature vector. Each element in the matrix is ​​the user's rating of the product potential feature. The singular value decomposition algorithm is used to decompose the user-product potential feature rating matrix into a user feature matrix and a product potential feature matrix, and the user preference vector and the product potential feature weight vector are obtained. Each row of the user feature matrix corresponds to the user preference vector, and each column of the product feature matrix corresponds to the product potential feature weight vector. Based on the user preference vector and the product potential feature weight vector, the cosine similarity is used to calculate the similarity between user preference and product; Obtain recommended products whose similarity is greater than a preset similarity threshold; Based on the recommended products, a product recommendation list including the user ID, the product ID and the recommended products is generated; Integrate the product recommendation list according to the preset data format to generate data assets for product recommendation scenarios.

8. The method according to claim 1, characterized in that The marketing recommendation list, risk warning list and product recommendation list generated by each shard data are integrated into the data service interface and provided to the business system for calling to realize the real-time application of data assets. The data service interface adopts RESTful style, supports batch query and real-time query, and ensures the timeliness and consistency of data, including: Design data service interface in RESTful style and define the request method and parameters of the interface; The interface supports GET requests, and the parameters include user ID and timestamp, which are used to distinguish batch queries from real-time queries; Obtain marketing recommendation lists, risk warning lists, and product recommendation lists generated by each shard data from the pre-established database; Integrate the marketing recommendation list, risk warning list and product recommendation list into the data service interface and return them to the caller in JSON format; When the business system calls the data service interface, it obtains the corresponding data based on the user ID and timestamp parameters; If the timestamp is the current time, a real-time query is triggered to obtain the latest data from each data source; If the timestamp is a historical time, a batch query is triggered to obtain historical data from the cache.

9. The method according to claim 1, characterized in that: According to the marketing recommendation list, risk warning list and product recommendation list, real-time data update is achieved through the Redis cache mechanism to ensure that the calling interface obtains the latest data and realizes efficient query and real-time update of data, including: In the real-time query scenario, after obtaining the marketing recommendation list, the data is stored in the cache through the ZADD command of Redis SortedSet; After obtaining the risk warning list, use the Redis List command to store the results in the cache; After obtaining the product recommendation list, store the list in the cache using the Redis RPUSH command; In batch query scenarios, use the Redis ZRANGE command to read the marketing recommendation list, and use the Redis LRANGE command to read the risk warning list and product recommendation list to ensure data consistency; Use the Redis EXPIRE command to set the expiration time of the data stored in the cache. When the expiration time expires, Redis will automatically delete the data stored in the cache.

Citation Information

Patent Citations

  • Intelligent data asset storage data evaluation method

    CN118503236A

  • Augmenting search results based on relevancy and utility

    US11630829B1

Cited By

  • Illegal behavior analysis method, device and equipment based on big data model prediction

    CN120408461A

  • Intelligent decision support system based on ERP data

    CN120450661A

  • Large-scale unstructured data joint processing method and system

    CN120892608A

  • Enterprise management data processing method based on big data analysis

    CN121118074A