Business big data analysis system based on artificial intelligence

Through the business big data analysis system based on artificial intelligence, the problem that the sparseness of user historical behavior data affects the accuracy of recommendation results is solved. Through the cooperation of data integration and intelligent analysis module, matrix decomposition is used to complete and predict user behavior data, and a personalized recommendation list is generated, which improves the accuracy of recommendation results and user satisfaction.

CN120106985AInactive Publication Date: 2025-06-06SHANDONG UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510204200.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

When building user portraits and making personalized recommendations, the prior art is affected by the sparsity or incompleteness of user historical behavior data, resulting in inaccurate recommendation results or low user satisfaction.

Method used

Adopt a business big data analysis system based on artificial intelligence, including data integration module, intelligent analysis module, real-time analysis module, recommendation system module and decision support module. Through the data integration module, a variety of data sources are collected and preprocessed, the intelligent analysis module performs data analysis and model optimization, the real-time analysis module performs real-time data processing, and the recommendation system module uses matrix decomposition to complete and predict user historical behavior data to generate a personalized recommendation list.

Benefits of technology

By completing and predicting user historical behavior data, the problem of data sparsity is alleviated, the missing values ​​in the data are significantly reduced, the data set is more complete, the accuracy of recommended results is improved, and user satisfaction is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120106985A_ABST
    Figure CN120106985A_ABST
Patent Text Reader

Abstract

The invention discloses a business big data analysis system based on artificial intelligence. The system comprises a data integration module which supports access of multiple data sources; the intelligent analysis module comprises a target definition and data type identification sub-module and an algorithm application and optimization sub-module, and the target definition and data type identification sub-module is responsible for clearly analyzing targets and identifying data types and structures; the algorithm application and optimization sub-module selects and optimizes an algorithm according to an analysis target and a data type to realize an analysis function; the real-time analysis module is responsible for carrying out real-time processing and analysis on the data by utilizing Apache Kafka; and the recommendation system module is responsible for carrying out customer subdivision, formulating a personalized marketing strategy according to a subdivision result, complementing and predicting historical behavior data of the user by utilizing matrix decomposition, calculating which financial products the user is most interested in by utilizing the matrix decomposition, and generating a personalized recommendation list.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of business analysis technology, and in particular to a business big data analysis system based on artificial intelligence. Background Art

[0002] Big data technology can process and analyze massive amounts of data, mine the value in it, and provide support for business decisions. Artificial intelligence technology, especially machine learning and deep learning technology, provides powerful tools for big data analysis. By training models, artificial intelligence can learn patterns from data, make predictions, and make decisions. With the deepening of digital transformation, all industries are seeking ways to drive business growth through data. Business big data analysis systems based on artificial intelligence are important tools to meet this demand.

[0003] In the prior art, when building user portraits and making personalized recommendations, the sparsity or incompleteness of user historical behavior data may affect the recommendation results, resulting in inaccurate or low user satisfaction. Therefore, an artificial intelligence-based business big data analysis system is proposed. Summary of the invention

[0004] The purpose of the present invention is to solve the shortcomings of the prior art, that is, when building user portraits and making personalized recommendations, the user's historical behavior data may be affected by the sparsity or incompleteness, resulting in inaccurate recommendation results or low user satisfaction, and to propose a business big data analysis system based on artificial intelligence.

[0005] In order to achieve the above object, the present invention adopts the following technical solutions: Business big data analysis system based on artificial intelligence, including: Data integration module: supports access to multiple data sources and collects data from them, including databases, log files, social media, IoT devices, etc., to achieve automatic synchronization and incremental update of data. The data integration module pre-processes the collected data, and the pre-processed data is stored in the data repository; Intelligent analysis module: including target definition and data type identification submodule and algorithm application and optimization submodule. The target definition and data type identification submodule is responsible for clarifying the analysis target and identifying the data type and structure; the algorithm application and optimization submodule selects and optimizes the algorithm according to the analysis target and data type to realize specific analysis functions, including classification, regression, clustering, etc. The algorithm application and optimization submodule optimizes the model performance through K-fold cross validation and grid search. The intelligent analysis module obtains data from the data repository for analysis and mining. The target definition and data type identification submodule selects and optimizes the algorithm according to the analysis target and data type provided by the target definition and data type identification submodule. The intelligent analysis module interacts with the real-time analysis module to share the analysis results and algorithm model; Real-time analysis module: responsible for real-time data processing and analysis using Apache Kafka. The real-time analysis module obtains data from the data repository for analysis and mining. The real-time analysis results obtained by the real-time analysis module are passed to the intelligent analysis module and the decision support module. Recommendation system module: responsible for customer segmentation, formulating personalized marketing strategies based on the segmentation results, using matrix decomposition to complete and predict user historical behavior data, using matrix decomposition to calculate which financial products users are most interested in, and generating personalized recommendation lists; Decision support module: responsible for integrating the data and analysis results of each module, and using FP-Growth to formulate decision recommendations and business strategy plans based on the analysis results and business needs.

[0006] The above technical solution further includes: Furthermore, the data integration module includes a data source unit, a data extraction unit, a data cleaning unit, a data conversion unit, a data synchronization and monitoring unit, and a data repository. The data source unit provides original data, including databases, log files, social media, IoT devices, etc. The data extraction unit extracts data from the data source unit, including data parsing, format conversion, etc. The data extraction unit receives data from the data source unit and passes the processed data to the data cleaning unit. The data cleaning unit cleans the extracted data, including removing invalid data, processing abnormal values, completing data, etc., and passes the cleaned data to the data conversion unit. The data conversion unit converts the data into the format and structure required by the target storage system, and passes the converted data to the data repository. The data repository stores the processed data, including relational databases, non-relational databases, data warehouses, etc. The data synchronization and monitoring unit is responsible for automatic synchronization and incremental update of data, and monitors the data synchronization process to ensure the integrity and timeliness of the data. The data synchronization and monitoring unit interacts with the data source unit and the data repository to achieve data synchronization and update.

[0007] Furthermore, the target definition and data type identification submodule includes a user interface unit, a target definition unit, a data type identification unit and a data storage and retrieval unit. The user interface unit is responsible for receiving the analysis target and data input by the user. The user interface unit passes the analysis target and data input by the user to the target definition unit. The target definition unit clarifies the specific analysis target based on the user input, converts the specific target into executable instructions, and passes the clarified target to the data type identification unit and the algorithm application and optimization submodule. The data type identification unit identifies the type and structure of the data and generates a data type report. The data type identification unit receives the data transmitted by the user interface unit and passes the identification result to the algorithm application and optimization submodule. The data storage and retrieval unit stores the identified data type information and analysis target for subsequent rapid retrieval and reuse. The data storage and retrieval unit receives data from the target definition unit and the data type identification unit and provides data retrieval services.

[0008] Furthermore, the algorithm application and optimization submodule includes an algorithm library unit, an algorithm selection and configuration unit, a model training and optimization unit, a model evaluation and verification unit, a prediction and result output unit, and a data storage and retrieval unit. The algorithm library unit is responsible for storing and managing various machine learning and deep learning algorithms, including classification algorithms, regression algorithms, clustering algorithms, etc. The algorithm library unit provides algorithm options to the algorithm selection and configuration unit and receives algorithm information selected by the algorithm selection and configuration unit. The algorithm selection and configuration unit selects an algorithm from the algorithm library according to the analysis target and data type, and performs configuration, such as setting hyperparameters, etc. The selection and configuration unit receives the analysis target and data type information transmitted by the target definition and data type identification submodule, requests algorithm options from the algorithm library unit, and transmits the selected algorithm and its configuration information to the model training and optimization unit. The model training and optimization unit uses the training data set to train the selected algorithm, and optimizes the model performance through K-fold cross validation and grid search. The model training and optimization unit receives the algorithm selection information. The algorithm and its configuration information transmitted by the configuration unit, receive the training data set provided by the data storage and retrieval unit, and pass the trained model and its performance evaluation results to the model evaluation and verification unit. The model evaluation and verification unit uses the verification data set to evaluate the trained model to ensure that its performance meets the analysis requirements. If the performance does not meet the requirements, it needs to return to the model training and optimization unit for further optimization. The model evaluation and verification unit receives the trained model and performance evaluation results transmitted by the model training and optimization unit, receives the verification data set provided by the data storage and retrieval unit, and passes the evaluation results and the judgment on whether the requirements are met to the prediction and result output unit. The prediction and result output unit applies the trained model to the new data set, performs prediction or classification, and outputs the analysis results. The prediction and result output unit receives the model that meets the requirements transmitted by the model evaluation and verification unit, receives the new data set provided by the data storage and retrieval unit, and the data storage and retrieval unit stores the training data set, the verification data set, the new data set, the trained model and the evaluation results.

[0009] Furthermore, the real-time analysis module includes a Kafka producer unit, a Kafka topic unit, a Kafka consumer unit, a real-time analysis and decision unit, a result output unit and a monitoring unit. The Kafka producer unit receives real-time data from the data integration module, writes the data into the Kafka topic, and communicates with the Kafka cluster. The Kafka topic unit is responsible for storing and managing real-time data streams for the Kafka consumer unit to read and process, and provides data streams to the Kafka consumer unit. The Kafka consumer unit reads the real-time data stream from the Kafka topic, performs preprocessing and stream processing, extracts valuable information, and sends the processing results to the real-time analysis and decision unit or the result output unit. The real-time analysis and decision unit performs data analysis, model prediction and real-time decision-making based on the data provided by the Kafka consumer unit, and sends the decision results to the result output unit or the monitoring system. The result output unit receives the decision results sent by the real-time analysis and decision unit, and outputs the results to a specified location. The monitoring unit obtains monitoring data from the Kafka topic unit, the Kafka consumer unit, and the real-time analysis and decision unit, and performs analysis and alarm.

[0010] Furthermore, the recommendation system module includes a data collection unit, a user portrait construction unit, a user portrait library unit, a recommendation algorithm unit and a result display unit. The data collection unit is responsible for the user's historical behavior data and product feature data. The user portrait construction unit is responsible for obtaining user data from the data collection unit and constructing user portraits using data mining and machine learning techniques. The user portraits include key information such as the user's financial preferences, risk tolerance, and investment habits. During the construction process, the user portraits are enriched by combining external data sources (such as social media data, IoT device data, etc.). The user portrait construction unit stores the constructed user portraits in the user portrait library for use by the recommendation algorithm unit. The unit interacts with the data collection unit and the user portrait library unit. The user portrait library unit is responsible for storing user portraits. The recommendation algorithm unit is responsible for obtaining user portraits from the user portrait library unit. The recommendation algorithm unit uses matrix decomposition to calculate which financial products users are most interested in, and generates a personalized recommendation list, which is passed to the result display unit for display. The recommendation algorithm unit interacts with the user portrait library unit, the data management module and the result display unit. The result display unit is responsible for receiving the recommendation list generated by the recommendation algorithm unit and displaying it in a certain format and layout. The result display unit provides a user interaction interface.

[0011] Furthermore, the data type identification unit identifies the type and structure of the data, specifically in the following steps: Data type identification: Numerical data: Identify numeric fields in the data, such as stock prices, trading volumes, and yields, which are used to calculate statistics (such as mean, variance, standard deviation, etc.) and perform numerical analysis; Text data: Identify text fields in the data, such as stock codes, company names, news headlines, etc. The text fields are processed (such as word segmentation, stop word removal, stemming, etc.) to extract useful information; Date / time data: Identify date or time fields in the data, such as transaction date, timestamp, etc. The date or time fields are used for time series analysis or as part of features; Image / video data (for certain specific applications, such as stock chart analysis): Identify image or video information in the data, perform image preprocessing (such as scaling, cropping, grayscale, etc.) and feature extraction (such as edge detection, texture analysis, etc.) on the image or video information; Data structure identification: Tabular data: Identify whether the data is stored in a tabular format, where each field corresponds to a column and each record corresponds to a row. Tabular data is one of the most common data structures and is suitable for financial data analysis tasks. Time series data: Identify whether the data is arranged in chronological order, such as the change of stock prices over time. The time series data is used for time series analysis, such as trend prediction, cycle detection, etc. Graph data (used for certain specific applications, such as social network analysis): Identify whether the data is stored in the form of a graph, where nodes represent entities (such as companies, individuals, etc.) and edges represent relationships between entities (such as transaction relationships, cooperative relationships, etc.). The graph data is used for graph theory analysis and network analysis; Data structure and type verification Use Python and Pandas to write code to verify whether the data type and structure meet expectations. For numerical data, calculate statistics (such as mean, variance, etc.) and draw histograms, box plots and other charts for visual analysis; for text data, perform experimental analysis such as word frequency statistics and text classification; for time series data, draw time series graphs, autocorrelation graphs and other charts for visual analysis.

[0012] Furthermore, the algorithm selected by the selection and configuration unit and its configuration information specifically include the following steps: Receiving analysis target and data type information, which is transmitted by the target definition and data type identification submodule to the algorithm selection and configuration unit; Analysis goal: predict stock prices; Data types: historical stock price data (opening price, highest price, lowest price, closing price), trading volume data, macroeconomic indicator data (such as GDP growth rate, unemployment rate, etc.); The algorithm selection and configuration unit selects an algorithm from an algorithm library according to an analysis target (predicting stock prices) and a data type (time series data, numerical data); In the financial field, commonly used prediction algorithms include time series analysis algorithms (such as ARIMA, GARCH), machine learning algorithms (such as linear regression, decision tree, random forest, support vector machine, neural network, etc.) and deep learning algorithms (such as recurrent neural network RNN, long short-term memory network LSTM, etc.); When selecting an algorithm, the algorithm selection and configuration unit considers the accuracy, robustness, and computational efficiency of the algorithm. For time series data, time series analysis algorithms such as ARIMA and GARCH or deep learning algorithms such as RNN and LSTM are given priority because they can capture the dynamic characteristics of time series data. For numerical data, machine learning algorithms such as linear regression and decision trees are used, and a trade-off is made according to specific data characteristics. After the algorithm is selected, the algorithm selection and configuration unit is configured, including setting hyperparameters. For example, for a neural network algorithm, it may be necessary to set hyperparameters such as the number of layers of the network, the number of neurons in each layer, and the learning rate; After the configuration is completed, the algorithm selection and configuration unit transmits the selected algorithm and its configuration information to the model training and optimization unit, and the model training and optimization unit uses this information to train the model and optimize the model to improve the prediction accuracy.

[0013] Furthermore, the recommendation algorithm unit uses matrix decomposition to complete and predict the user's historical behavior data. At the same time, it uses matrix decomposition to calculate which financial products the user is most interested in and generates a personalized recommendation list. The specific steps are: User-product interaction matrix: Construct a user-product interaction matrix R, where represents the rating or interaction behavior (such as purchase, click, etc.) of user i on product j. This matrix is ​​usually sparse because most users have only interacted with a few products. Matrix decomposition: Decompose the user-product interaction matrix R into the product of two low-dimensional matrices U and V, namely , where U is the user feature matrix, V is the product feature matrix, and the matrix decomposition is performed by optimizing the loss function through gradient descent. The loss function measures the reconstruction matrix The difference between the original matrix R; Completion and prediction: Through matrix decomposition, missing values ​​in the user-product interaction matrix are completed and user-product interactions are predicted; Calculate the user's interest in financial products: After obtaining the user feature matrix U and the product feature matrix V, the user's interest in each financial product is obtained by calculating the dot product of the user feature vector and the product feature vector, expressed as , sort the financial products according to their interest scores and select the products with the highest scores as part of the recommendation list; Generate a personalized recommendation list: Generate a personalized recommendation list based on the calculated user interest score for financial products.

[0014] The present invention has the following beneficial effects: In the present invention, matrix decomposition is used to complete and predict user historical behavior data to alleviate the data sparsity problem, significantly reduce missing values ​​in the data, make the data set more complete, and the completed data can more accurately reflect the user's real behavior pattern, thereby avoiding deviations caused by data sparsity. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 This is a system block diagram of the artificial intelligence-based business big data analysis system proposed in the present invention. DETAILED DESCRIPTION

[0016] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0017] See also Figure 1 As shown, the present invention is a business big data analysis system based on artificial intelligence, comprising: Data integration module: supports access to multiple data sources and collects data from them, including databases, log files, social media, IoT devices, etc., to achieve automatic synchronization and incremental update of data. The data integration module pre-processes the collected data, and the pre-processed data is stored in the data repository; Intelligent analysis module: including target definition and data type identification submodule and algorithm application and optimization submodule. The target definition and data type identification submodule is responsible for clarifying the analysis target and identifying the data type and structure; the algorithm application and optimization submodule selects and optimizes the algorithm according to the analysis target and data type to realize specific analysis functions, including classification, regression, clustering, etc. The algorithm application and optimization submodule optimizes the model performance through K-fold cross validation and grid search. The intelligent analysis module obtains data from the data repository for analysis and mining. The target definition and data type identification submodule selects and optimizes the algorithm according to the analysis target and data type provided by the target definition and data type identification submodule. The intelligent analysis module interacts with the real-time analysis module to share the analysis results and algorithm model; Real-time analysis module: responsible for real-time data processing and analysis using Apache Kafka. The real-time analysis module obtains data from the data repository for analysis and mining. The real-time analysis results obtained by the real-time analysis module are passed to the intelligent analysis module and the decision support module. Recommendation system module: responsible for customer segmentation, formulating personalized marketing strategies based on the segmentation results, using matrix decomposition to complete and predict user historical behavior data, using matrix decomposition to calculate which financial products users are most interested in, and generating personalized recommendation lists; Decision support module: responsible for integrating the data and analysis results of each module, and using FP-Growth to make decision recommendations and business strategy planning based on the analysis results and business needs. Use the FP-Growth algorithm to process the pre-processed data and analysis results of each module to construct a frequent pattern tree (FP-tree). Through FP-tree, frequent item sets are mined, which represent the patterns or combinations that frequently appear in business data. Based on the mined frequent item sets, association rules are further generated. By setting appropriate support and confidence thresholds, meaningful association rules are screened out. Pattern recognition: Analyze the mined frequent item sets and association rules to identify potential patterns, trends or associations in the business. Evaluate the specific impact of these patterns on the business, such as increasing sales, increasing user stickiness, optimizing product portfolio, etc. Market strategy adjustment: Adjust market strategies based on the mined association rules, such as optimizing promotional activities, improving product recommendation algorithms, etc. Long-term goal setting: Set long-term business development goals based on the analysis results and business status.

[0018] In one embodiment, for the above-mentioned data integration module, the data integration module includes a data source unit, a data extraction unit, a data cleaning unit, a data conversion unit, a data synchronization and monitoring unit, and a data repository. The data source unit provides original data, including databases, log files, social media, IoT devices, etc. The data extraction unit extracts data from the data source unit, including data parsing, format conversion, etc. The data extraction unit receives data from the data source unit and passes the processed data to the data cleaning unit. The data cleaning unit cleans the extracted data, including removing invalid data, processing abnormal values, completing data, etc., and passes the cleaned data to the data conversion unit. The data conversion unit converts the data into the format and structure required by the target storage system, and passes the converted data to the data repository. The data repository stores the processed data, including relational databases, non-relational databases, data warehouses, etc. The data synchronization and monitoring unit is responsible for automatic synchronization and incremental update of data, and monitors the data synchronization process to ensure the integrity and timeliness of the data. The data synchronization and monitoring unit interacts with the data source unit and the data repository to achieve data synchronization and update.

[0019] The data source unit accesses the data source, uses technologies such as JDBC (Java Database Connectivity) or ODBC (Open Database Connectivity) to connect to relational databases such as MySQL, Oracle, etc., to obtain structured data, obtains unstructured or semi-structured data from third-party service providers (such as credit rating agencies, social media platforms) through RESTful API or SOAP API, etc., uses log parsing tools (such as Logstash, Fluentd) to read log files generated by applications, servers, etc., extract useful information therein, and collects real-time data from IoT devices (such as smart wearable devices, sensors) through MQTT (Message Queuing Telemetry Transport) protocol or other IoT communication protocols. In the financial industry, structured data such as customer transaction records and account balances are obtained from the bank system; unstructured data such as customer public information and comments are obtained from social media platforms; and real-time data such as customer consumption habits and health status are collected from IoT devices.

[0020] In one embodiment, for the above-mentioned target definition and data type identification submodule, the target definition and data type identification submodule includes a user interface unit, a target definition unit, a data type identification unit and a data storage and retrieval unit. The user interface unit is responsible for receiving the analysis target and data input by the user. The user interface unit passes the analysis target and data input by the user to the target definition unit. The target definition unit clarifies the specific analysis target based on the user input, and converts the specific target into executable instructions, and passes the clarified target to the data type identification unit and the algorithm application and optimization submodule. The data type identification unit identifies the type and structure of the data and generates a data type report. The data type identification unit receives the data transmitted by the user interface unit and passes the identification result to the algorithm application and optimization submodule. The data storage and retrieval unit stores the identified data type information and analysis target for subsequent rapid retrieval and reuse. The data storage and retrieval unit receives data from the target definition unit and the data type identification unit and provides data retrieval services.

[0021] The user inputs the analysis target (such as predicted sales, customer classification, etc.) and data source (such as database, log file, etc.) through the interface. The target definition unit parses the user input, clarifies the analysis target, and passes it to the data type identification unit. The data type identification unit uses preset rules or models to identify the data type (such as numerical type, categorical type, time series, etc.) and structure (such as table, graph structure, etc.), and generates a detailed data type report.

[0022] In one embodiment, for the above-mentioned algorithm application and optimization submodule, the algorithm application and optimization submodule includes an algorithm library unit, an algorithm selection and configuration unit, a model training and optimization unit, a model evaluation and verification unit, a prediction and result output unit, and a data storage and retrieval unit. The algorithm library unit is responsible for storing and managing various machine learning and deep learning algorithms, including classification algorithms, regression algorithms, clustering algorithms, etc. The algorithm library unit provides algorithm options to the algorithm selection and configuration unit and receives the algorithm information selected by the algorithm selection and configuration unit. The algorithm selection and configuration unit selects an algorithm from the algorithm library according to the analysis target and data type, and performs configuration, such as setting hyperparameters, etc. The selection and configuration unit receives the analysis target and data type information transmitted by the target definition and data type identification submodule, requests the algorithm option from the algorithm library unit, and transmits the selected algorithm and its configuration information to the model training and optimization unit. The model training and optimization unit uses the training data set to train the selected algorithm, and optimizes the model performance through K-fold cross validation and grid search. The model training and optimization unit performs the following steps: The optimization unit receives the algorithm and its configuration information transmitted by the algorithm selection and configuration unit, receives the training data set provided by the data storage and retrieval unit, and transmits the trained model and its performance evaluation results to the model evaluation and verification unit. The model evaluation and verification unit uses the verification data set to evaluate the trained model to ensure that its performance meets the analysis requirements. If the performance does not meet the requirements, it needs to be returned to the model training and optimization unit for further optimization. The model evaluation and verification unit receives the trained model and performance evaluation results transmitted by the model training and optimization unit, receives the verification data set provided by the data storage and retrieval unit, and transmits the evaluation results and the judgment on whether the requirements are met to the prediction and result output unit. The prediction and result output unit applies the trained model to the new data set, performs prediction or classification, and outputs the analysis results. The prediction and result output unit receives the model that meets the requirements transmitted by the model evaluation and verification unit, receives the new data set provided by the data storage and retrieval unit, and stores the training data set, the verification data set, the new data set, the trained model and the evaluation results.

[0023] Forecast sales: Data preparation: Collect historical sales data, seasonal factors, promotional information, etc. as feature variables, and use sales as the target variable.

[0024] Analysis goal definition: The clear analysis goal is to predict sales for the next month.

[0025] Data type identification: Identify data as time series data, including numerical and categorical features.

[0026] Algorithm selection and configuration: Select a time series analysis algorithm (such as ARIMA, LSTM, etc.) from the algorithm library as the forecasting model. Configure algorithm parameters, such as the order of the ARIMA model, the number of layers and neurons of the LSTM model, etc. ARIMA model parameters: order p=2, d=1, q=2; LSTM model parameters: number of layers=3, number of neurons=128, learning rate=0.001, batch size=32.

[0027] Model training and optimization: The selected algorithm is trained using historical sales data, and model parameters are optimized through K-fold cross validation and grid search.

[0028] Prediction and result output: Apply the trained model to predict sales data for the next month, and output the prediction results and confidence intervals.

[0029] In one embodiment, for the above-mentioned real-time analysis module, the real-time analysis module includes a Kafka producer unit, a Kafka topic unit, a Kafka consumer unit, a real-time analysis and decision unit, a result output unit and a monitoring unit. The Kafka producer unit receives real-time data from the data integration module, writes the data into the Kafka topic, and communicates with the Kafka cluster. The Kafka topic unit is responsible for storing and managing real-time data streams for the Kafka consumer unit to read and process, and provides data streams to the Kafka consumer unit. The Kafka consumer unit reads the real-time data stream from the Kafka topic, performs preprocessing and stream processing, extracts valuable information, and sends the processing results to the real-time analysis and decision unit or the result output unit. The real-time analysis and decision unit performs data analysis, model prediction and real-time decision-making based on the data provided by the Kafka consumer unit, and sends the decision results to the result output unit or the monitoring system. The result output unit receives the decision results sent by the real-time analysis and decision unit, and outputs the results to a specified location. The monitoring unit obtains monitoring data from the Kafka topic unit, the Kafka consumer unit, and the real-time analysis and decision unit, and performs analysis and alarm.

[0030] Real-time inventory alert: The Kafka producer unit receives real-time data from data sources such as the order system and inventory system of the e-commerce platform. The Kafka topic unit stores and manages these data streams. The Kafka consumer unit reads the real-time data stream, performs preprocessing (such as calculating order quantity, inventory quantity, etc.), and extracts inventory features (such as inventory quantity, order growth trend, etc.). The real-time analysis and decision unit applies early warning algorithms (such as threshold-based early warnings, trend prediction-based early warnings, etc.) to generate real-time early warning information based on inventory characteristics. The result output unit sends the early warning information to inventory management personnel or triggers the automatic replenishment process.

[0031] Real-time marketing effect analysis: The Kafka producer unit receives real-time data from data sources such as the e-commerce platform's advertising delivery system and user behavior logs. The Kafka topic unit stores and manages these data streams. The Kafka consumer unit reads the real-time data stream, performs preprocessing (such as calculating ad clicks, conversion rates, etc.), and extracts marketing features (such as ad exposures, clicks, conversions, etc.). The real-time analysis and decision-making unit applies marketing effect evaluation algorithms (such as A / B testing, multivariate analysis, etc.) and generates real-time marketing effect evaluation reports based on marketing features. The result output unit displays the evaluation report to marketers or triggers automatic optimization strategies.

[0032] In one embodiment, for the above-mentioned recommendation system module, the recommendation system module includes a data collection unit, a user portrait construction unit, a user portrait library unit, a recommendation algorithm unit and a result display unit. The data collection unit is responsible for the user's historical behavior data and product feature data. The user portrait construction unit is responsible for obtaining user data from the data collection unit and constructing a user portrait using data mining and machine learning techniques. The user portrait includes key information such as the user's financial preferences, risk tolerance, and investment habits. During the construction process, the user portrait is enriched by combining external data sources (such as social media data, IoT device data, etc.). The user portrait construction unit stores the constructed user portrait in the user The user portrait library is used by the recommendation algorithm unit. The user portrait construction unit interacts with the data collection unit and the user portrait library unit. The user portrait library unit is responsible for storing user portraits. The recommendation algorithm unit is responsible for obtaining user portraits from the user portrait library unit. The recommendation algorithm unit uses matrix decomposition to calculate which financial products the user is most interested in, and generates a personalized recommendation list, which is passed to the result display unit for display. The recommendation algorithm unit interacts with the user portrait library unit, the data management module and the result display unit. The result display unit is responsible for receiving the recommendation list generated by the recommendation algorithm unit.

[0033] In one embodiment, for the data type identification unit, the data type identification unit identifies the type and structure of data, specifically in the following steps: Data type identification: Numerical data: Identify numeric fields in the data, such as stock prices, trading volumes, and yields, which are used to calculate statistics (such as mean, variance, standard deviation, etc.) and perform numerical analysis; Text data: Identify text fields in the data, such as stock codes, company names, news headlines, etc. The text fields are processed (such as word segmentation, stop word removal, stemming, etc.) to extract useful information; Date / time data: Identify date or time fields in the data, such as transaction date, timestamp, etc. The date or time fields are used for time series analysis or as part of features; Image / video data (for certain specific applications, such as stock chart analysis): Identify image or video information in the data, perform image preprocessing (such as scaling, cropping, grayscale, etc.) and feature extraction (such as edge detection, texture analysis, etc.) on the image or video information; Data structure identification: Tabular data: Identify whether the data is stored in a tabular format, where each field corresponds to a column and each record corresponds to a row. Tabular data is one of the most common data structures and is suitable for financial data analysis tasks. Time series data: Identify whether the data is arranged in chronological order, such as the change of stock prices over time. The time series data is used for time series analysis, such as trend prediction, cycle detection, etc. Graph data (used for certain specific applications, such as social network analysis): Identify whether the data is stored in the form of a graph, where nodes represent entities (such as companies, individuals, etc.) and edges represent relationships between entities (such as transaction relationships, cooperative relationships, etc.). The graph data is used for graph theory analysis and network analysis; Data structure and type verification Use Python and Pandas to write code to verify whether the data type and structure meet expectations. For numerical data, calculate statistics (such as mean, variance, etc.) and draw histograms, box plots and other charts for visual analysis; for text data, perform experimental analysis such as word frequency statistics and text classification; for time series data, draw time series graphs, autocorrelation graphs and other charts for visual analysis; Suppose we have a CSV file containing stock price data. The file contains the following fields: stock symbol (text type), transaction date (date type), opening price (numeric type), high price (numeric type), low price (numeric type), closing price (numeric type), and trading volume (numeric type).

[0034] Data type identification: Stock code: text data used to uniquely identify each stock.

[0035] Transaction date: Date type data, indicating the date when the transaction occurred.

[0036] Opening price, highest price, lowest price, closing price: numerical data, indicating the trading price of the stock within a day.

[0037] Trading volume: Numerical data, indicating the number of stock transactions in a day.

[0038] Data structure identification: The data is stored in a table format, with each field corresponding to a column and each record corresponding to a row.

[0039] The data is sorted by transaction date to form time series data.

[0040] Data structure and type verification: Use the Pandas library to read the CSV file and verify that the data type and structure of each field are as expected.

[0041] Calculate stock price statistics (such as mean, variance, etc.) and draw time series graphs for visual analysis.

[0042] In one embodiment, for the above-mentioned selection and configuration unit, the algorithm selected by the selection and configuration unit and its configuration information are specifically as follows: Receiving analysis target and data type information, which is transmitted by the target definition and data type identification submodule to the algorithm selection and configuration unit; Analysis goal: predict stock prices; Data types: historical stock price data (opening price, highest price, lowest price, closing price), trading volume data, macroeconomic indicator data (such as GDP growth rate, unemployment rate, etc.); The algorithm selection and configuration unit selects an algorithm from an algorithm library according to an analysis target (predicting stock prices) and a data type (time series data, numerical data); In the financial field, commonly used prediction algorithms include time series analysis algorithms (such as ARIMA, GARCH), machine learning algorithms (such as linear regression, decision tree, random forest, support vector machine, neural network, etc.) and deep learning algorithms (such as recurrent neural network RNN, long short-term memory network LSTM, etc.); When selecting an algorithm, the algorithm selection and configuration unit considers the accuracy, robustness, and computational efficiency of the algorithm. For time series data, time series analysis algorithms such as ARIMA and GARCH or deep learning algorithms such as RNN and LSTM are given priority because they can capture the dynamic characteristics of time series data. For numerical data, machine learning algorithms such as linear regression and decision trees are used, and a trade-off is made according to specific data characteristics. After the algorithm is selected, the algorithm selection and configuration unit is configured, including setting hyperparameters. For example, for a neural network algorithm, it may be necessary to set hyperparameters such as the number of layers of the network, the number of neurons in each layer, and the learning rate; After the configuration is completed, the algorithm selection and configuration unit transmits the selected algorithm and its configuration information to the model training and optimization unit, and the model training and optimization unit uses this information to train the model and optimize the model to improve the prediction accuracy; Suppose we have chosen the LSTM algorithm to predict stock prices. Here is an example of configuring the LSTM algorithm:

[0043] Input layer: The size of the input layer should be equal to the number of features of the input data. In our case, the input data may include historical stock price data (4 features: open price, high price, low price, close price) and trading volume data (1 feature), so the size of the input layer can be set to 5.

[0044] Hidden layers: The number of hidden layers and the number of neurons in each layer are hyperparameters that need to be tuned based on experiments. For example, we can choose 2 hidden layers with 100 neurons in each layer.

[0045] Output layer: The size of the output layer should be equal to the number of prediction targets. In our case, the prediction targets are stock prices, so the size of the output layer can be set to 1.

[0046] Learning rate: The learning rate is an important hyperparameter that controls how fast the model updates weights during training. The choice of learning rate needs to be adjusted based on the specific problem and data. For example, we can choose 0.001 as the initial learning rate.

[0047] These configuration information will be passed to the model training and optimization unit for training and optimizing the LSTM model.

[0048] In one embodiment, for the above-mentioned recommendation algorithm unit, the recommendation algorithm unit uses matrix decomposition to complete and predict the user's historical behavior data, and at the same time, uses matrix decomposition to calculate which financial products the user is most interested in, and generates a personalized recommendation list. The specific steps are: User-product interaction matrix: Construct a user-product interaction matrix R, where represents the rating or interaction behavior (such as purchase, click, etc.) of user i on product j. This matrix is ​​usually sparse because most users have only interacted with a few products. Matrix decomposition: Decompose the user-product interaction matrix R into the product of two low-dimensional matrices U and V, namely , where U is the user feature matrix, V is the product feature matrix, and the matrix decomposition is performed by optimizing the loss function through gradient descent. The loss function measures the reconstruction matrix The difference between the original matrix R; Completion and prediction: Through matrix decomposition, missing values ​​in the user-product interaction matrix are completed and user-product interactions are predicted; Calculate the user's interest in financial products: After obtaining the user feature matrix U and the product feature matrix V, the user's interest in each financial product is obtained by calculating the dot product of the user feature vector and the product feature vector, expressed as , sort the financial products according to their interest scores and select the products with the highest scores as part of the recommendation list; Generate a personalized recommendation list: Generate a personalized recommendation list based on the calculated user interest score for financial products.

[0049] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. The business big data analysis system based on artificial intelligence is characterized by: include: Data integration module: supports access to multiple data sources and collects data from them. The data integration module pre-processes the collected data, and the pre-processed data is stored in a data repository; Intelligent analysis module: including target definition and data type identification submodule and algorithm application and optimization submodule, wherein the target definition and data type identification submodule is responsible for identifying data types; The algorithm application and optimization submodule selects algorithms for analysis according to the analysis objectives and data types, including classification, regression and clustering. The algorithm application and optimization submodule optimizes model performance through K-fold cross validation and grid search; Real-time analysis module: responsible for real-time data processing and analysis using Apache Kafka. The real-time analysis module obtains data from the data repository for analysis and mining. The real-time analysis results obtained by the real-time analysis module are passed to the intelligent analysis module and the decision support module. Recommendation system module: responsible for customer segmentation, formulating personalized marketing strategies based on the segmentation results, using matrix decomposition to complete and predict user historical behavior data, using matrix decomposition to calculate which financial products users are most interested in, and generating personalized recommendation lists; Decision support module: responsible for integrating the data and analysis results of each module, and using FP-Growth to formulate decision recommendations and business strategy plans based on the analysis results and business needs.

2. The business big data analysis system based on artificial intelligence according to claim 1 is characterized in that: The data integration module includes a data source unit, a data extraction unit, a data cleaning unit, a data conversion unit, a data synchronization and monitoring unit and a data repository. The data source unit provides original data, the data extraction unit extracts data from the data source unit, the data extraction unit receives data from the data source unit, and passes the processed data to the data cleaning unit. The data cleaning unit cleans the extracted data and passes the cleaned data to the data conversion unit. The data conversion unit converts the data into the format and structure required by the target storage system and passes the converted data to the data repository. The data repository stores the processed data. The data synchronization and monitoring unit is responsible for automatic synchronization and incremental update of data, and monitors the data synchronization process. The data synchronization and monitoring unit interacts with the data source unit and the data repository.

3. The business big data analysis system based on artificial intelligence according to claim 1 is characterized in that: The target definition and data type identification submodule includes a user interface unit, a target definition unit, a data type identification unit and a data storage and retrieval unit. The user interface unit is responsible for receiving the analysis target and data input by the user. The user interface unit passes the analysis target and data input by the user to the target definition unit. The target definition unit clarifies the specific analysis target based on the user input, converts the specific target into executable instructions, and passes the clarified target to the data type identification unit and the algorithm application and optimization submodule. The data type identification unit identifies the type and structure of the data and generates a data type report. The data type identification unit receives the data passed by the user interface unit and passes the identification result to the algorithm application and optimization submodule. The data storage and retrieval unit stores the identified data type information and analysis target. The data storage and retrieval unit receives data from the target definition unit and the data type identification unit.

4. The business big data analysis system based on artificial intelligence according to claim 1 is characterized in that: The algorithm application and optimization submodule includes an algorithm library unit, an algorithm selection and configuration unit, a model training and optimization unit, a model evaluation and verification unit, a prediction and result output unit, and a data storage and retrieval unit. The algorithm library unit is responsible for storing and managing various machine learning and deep learning algorithms. The algorithm library unit provides algorithm options to the algorithm selection and configuration unit and receives algorithm information selected by the algorithm selection and configuration unit. The algorithm selection and configuration unit selects an algorithm from the algorithm library according to the analysis target and data type and performs configuration. The selection and configuration unit receives the analysis target and data type information transmitted by the target definition and data type identification submodule, requests algorithm options from the algorithm library unit, and transmits the selected algorithm and its configuration information to the model training and optimization unit. The model training and optimization unit uses the training data set to train the selected algorithm and optimizes the model performance through K-fold cross validation and grid search. Receive the algorithm and its configuration information transmitted by the algorithm selection and configuration unit, receive the training data set provided by the data storage and retrieval unit, and transmit the trained model and its performance evaluation results to the model evaluation and verification unit. The model evaluation and verification unit uses the verification data set to evaluate the trained model. The model evaluation and verification unit receives the trained model and the performance evaluation results transmitted by the model training and optimization unit, receives the verification data set provided by the data storage and retrieval unit, and transmits the evaluation results and the judgment of whether the requirements are met to the prediction and result output unit. The prediction and result output unit applies the trained model to the new data set, performs prediction or classification, and outputs the analysis results. The prediction and result output unit receives the model that meets the requirements transmitted by the model evaluation and verification unit, receives the new data set provided by the data storage and retrieval unit, and stores the training data set, the verification data set, the new data set, the trained model and the evaluation results.

5. The business big data analysis system based on artificial intelligence according to claim 1 is characterized in that: The real-time analysis module includes a Kafka producer unit, a Kafka topic unit, a Kafka consumer unit, a real-time analysis and decision unit, a result output unit and a monitoring unit. The Kafka producer unit receives real-time data from the data integration module, writes the data into the Kafka topic, and communicates with the Kafka cluster. The Kafka topic unit is responsible for storing and managing real-time data streams for reading and processing by the Kafka consumer unit, and provides data streams to the Kafka consumer unit. The Kafka consumer unit reads the real-time data stream from the Kafka topic, performs preprocessing and stream processing, and sends the processing results to the real-time analysis and decision unit or the result output unit. The real-time analysis and decision unit performs data analysis, model prediction and real-time decision based on the data provided by the Kafka consumer unit, and sends the decision results to the result output unit or the monitoring system. The result output unit receives the decision results sent by the real-time analysis and decision unit, and outputs the results to a specified location. The monitoring unit obtains monitoring data from the Kafka topic unit, the Kafka consumer unit, and the real-time analysis and decision unit, and performs analysis and alarm.

6. The business big data analysis system based on artificial intelligence according to claim 1 is characterized in that: The recommendation system module includes a data collection unit, a user portrait construction unit, a user portrait library unit, a recommendation algorithm unit and a result display unit. The data collection unit is responsible for the user's historical behavior data and product feature data. The user portrait construction unit is responsible for obtaining user data from the data collection unit, constructing user portraits using data mining and machine learning techniques, and enriching user portraits in combination with external data sources during the construction process. The user portrait construction unit stores the constructed user portraits in the user portrait library for use by the recommendation algorithm unit. The user portrait construction unit interacts with the data collection unit and the user portrait library unit. The user portrait library unit is responsible for storing user portraits. The recommendation algorithm unit is responsible for obtaining user portraits from the user portrait library unit. The recommendation algorithm unit uses matrix decomposition to calculate which financial products the user is most interested in, and generates a personalized recommendation list, and passes the recommendation list to the result display unit for display. The recommendation algorithm unit interacts with the user portrait library unit, the data management module and the result display unit. The result display unit is responsible for receiving the recommendation list generated by the recommendation algorithm unit and displaying it in a certain format and layout. The result display unit provides a user interaction interface.

7. The business big data analysis system based on artificial intelligence according to claim 3 is characterized in that: The data type identification unit identifies the type and structure of data, specifically in the following steps: Data type identification: Numerical data: Identify numeric fields in the data, which are used to calculate statistics and perform numerical analysis; Text data: Identify text fields in the data, and perform text processing on the text fields; Date / time data: Identify date or time fields in the data, which are used for time series analysis or as part of features; Image / video data: Identify the image or video information in the data, and perform image preprocessing and feature extraction on the image or video information; Data structure identification: Tabular data: Identify whether the data is stored in a tabular format, with each field corresponding to a column and each record corresponding to a row. The tabular data is suitable for financial data analysis tasks; Time series data: Identify whether the data is arranged in time sequence. The time series data is used for time series analysis; Graph data: Identify whether the data is stored in the form of a graph, where nodes represent entities and edges represent relationships between entities. The graph data is used for graph theory analysis and network analysis; Data structure and type verification Use Python and Pandas to write code to verify whether the data type and structure meet expectations. For numerical data, calculate statistics and draw charts for visual analysis; for text data, perform experimental analysis; for time series data, draw charts for visual analysis.

8. The business big data analysis system based on artificial intelligence according to claim 4 is characterized in that: The specific steps of the algorithm and its configuration information selected by the selection and configuration unit are as follows: Receiving analysis target and data type information, which is transmitted by the target definition and data type identification submodule to the algorithm selection and configuration unit; The algorithm selection and configuration unit selects an algorithm from an algorithm library according to the analysis target and data type; When selecting an algorithm, the algorithm selection and configuration unit considers the accuracy, robustness, and computational efficiency of the algorithm. For time series data, the time series analysis algorithm or the deep learning algorithm is given priority. For numerical data, the machine learning algorithm is used, and a balance is made according to the specific data characteristics. After the algorithm is selected, the algorithm selection and configuration unit is configured, including setting hyperparameters; After the configuration is completed, the algorithm selection and configuration unit transmits the selected algorithm and its configuration information to the model training and optimization unit, and the model training and optimization unit uses this information to train the model.

9. The business big data analysis system based on artificial intelligence according to claim 6 is characterized in that: The recommendation algorithm unit uses matrix decomposition to complete and predict the user's historical behavior data. At the same time, it uses matrix decomposition to calculate which financial products the user is most interested in and generates a personalized recommendation list. The specific steps are: User-product interaction matrix: Construct a user-product interaction matrix R, where represents the rating or interaction behavior of user i on product j; Matrix decomposition: Decompose the user-product interaction matrix R into the product of two low-dimensional matrices U and V, namely , where U is the user feature matrix, V is the product feature matrix, and the matrix decomposition is performed by optimizing the loss function through gradient descent. The loss function measures the reconstruction matrix The difference between the original matrix R; Completion and prediction: Through matrix decomposition, missing values ​​in the user-product interaction matrix are completed and user-product interactions are predicted; Calculate the user's interest in financial products: After obtaining the user feature matrix U and the product feature matrix V, the user's interest in each financial product is obtained by calculating the dot product of the user feature vector and the product feature vector, expressed as , sort the financial products according to their interest scores and select the products with the highest scores as part of the recommendation list; Generate a personalized recommendation list: Generate a personalized recommendation list based on the calculated user interest score for financial products.

Citation Information

Patent Citations

  • Data analysis method based on big data operation analysis

    CN107908690A

  • Personalized insurance product recommendation engine system based on big data analysis and machine learning

    CN116823498A

  • Intelligent recommendation system and method based on deep learning and big data fusion

    CN118673220A

  • Photovoltaic operation fault diagnosis method and system

    CN119382613A