Digital financial management system based on big data

By designing a multi-module big data digital financial management system, the problem that existing systems are difficult to provide investment advice that is closely followed is solved, and the provision of personalized investment advice and market response is achieved, and the investment success rate and satisfaction are improved.

CN120182007AInactive Publication Date: 2025-06-20ZHEJIANG VOCATIONAL COLLEGE OF COMMERCE

Patent Information

Application Number
CN202510215076.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-20
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing financial investment risk management system based on big data is difficult to provide investment advice that keeps pace with current affairs and cannot meet users' financial management needs.

Method used

A digital financial management system based on big data is designed, including data collection, preprocessing, topic analysis, trend analysis, matching comparison, scoring mechanism and investment proposal generation, etc., to optimize system performance and accuracy through machine learning and deep learning algorithms.

Benefits of technology

It has realized the identification and trend analysis of hot topics, provided personalized investment advice, improved investor satisfaction and investment success rate, and was able to respond quickly to market changes and reduce investment risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182007A_ABST
    Figure CN120182007A_ABST
Patent Text Reader

Abstract

The invention discloses a digital financial management system based on big data, and relates to the technical field of financial management. Comprising a data acquisition module for acquiring financial related data through a web crawler unit and a data interface unit and storing the financial related data in an original database; the data preprocessing module is used for performing cleaning, word segmentation and labeling operation on the original data to obtain text data; and the topic analysis module is used for analyzing the preprocessed text data, identifying hot topics, judging sources of the hot topics and adding feature tags. The system covers a plurality of modules of data acquisition, preprocessing, topic analysis, trend analysis, matching comparison, scoring mechanism, investment suggestion generation and the like, and a complete digital financial management ecology is formed. The system can identify hot topics, judge sources of the hot topics and provide personalized investment suggestions according to risk preferences and market conditions of investors; and the satisfaction degree and the investment success rate of investors can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of financial management, and particularly to a digital financial management system based on big data. Background Art

[0002] With the continuous development and application of big data technology, the digital financial industry is facing more and more challenges and opportunities. How to extract valuable information from massive data and provide accurate and timely investment advice for investors has become an important issue.

[0003] After retrieval, the application solution with the Chinese patent application number CN202210715433.X discloses a financial investment risk management system and method based on big data, including a financial data collection module, an individual account fund management module, a permission management module, a financial investment risk assessment module, an investment income prediction module, an investment feasibility prediction module, and an investment management module; the financial data collection module is used to collect historical financial data and transmit the collected historical financial data to the financial investment risk assessment module; the individual account fund management module is used to collect personal fund movement information, predict personal short-term fund requirements and the personal account fund turnover time according to the collected information, and transmit the prediction information to the permission management module; the permission management module is used to connect the individual account fund management module and the investment feasibility prediction module to ensure that personal fund information is not maliciously leaked. The management system in the above patent has the following deficiencies: Although it can provide certain management services, it cannot provide investment advice that keeps up with current events and is difficult to meet the financial management needs of users. Summary of the Invention

[0004] The purpose of the present invention is to solve the deficiencies existing in the prior art and propose a digital financial management system based on big data.

[0005] To achieve the above purpose, the present invention adopts the following technical solutions:

[0006] A digital financial management system based on big data, comprising:

[0007] A data collection module that collects financial-related data through a web crawler unit and a data interface unit and stores it in a raw database;

[0008] A data preprocessing module that performs cleaning, word segmentation, and annotation operations on the raw data to obtain text data;

[0009] A topic analysis module that analyzes the preprocessed text data, identifies hot topics, determines their sources, and adds feature tags;

[0010] The trend analysis module divides the time series into multiple unit time intervals, analyzes the change in the discussion volume of hot topics in different time intervals, and determines their popularity trends;

[0011] The matching and comparison module searches for topics with the same or similar feature tags within a set time and compares their popularity values;

[0012] The scoring mechanism module scores the topic popularity and simultaneously scores the investment risks in combination with financial transaction data and market conditions;

[0013] The investment advice generation module generates personalized investment advice for investors based on the trend analysis results, the label matching and comparison results, and the scores provided by the scoring mechanism;

[0014] The intelligent algorithm optimization module uses machine learning and deep learning algorithms to learn and train historical data to optimize the performance and accuracy of the system.

[0015] Preferably, the data collection module includes:

[0016] The web crawler unit: collects financial-related text data from various channels; accurately captures content related to financial topics by setting specific keywords and screening rules, and stores it in the original database;

[0017] The data interface unit: establishes data interfaces with various financial data providers to obtain real-time financial transaction data, market data, and structured data.

[0018] Preferably, the data preprocessing module includes:

[0019] The data cleaning unit: cleans the collected original data to remove noise data, duplicate data, and invalid data;

[0020] The text tokenization and annotation unit: performs tokenization on the cleaned text data to break the text into meaningful words or phrases; at the same time, annotates the tokenized data according to the preset financial term library and semantic rules.

[0021] Among them, the feature tags include: category tags and content tags, where the category tags are divided according to the topic area, including local topics, domestic topics, or international topics.

[0022] Preferably, the topic analysis module includes:

[0023] The hot topic identification unit: statistically analyzes the preprocessed text data to calculate the occurrence frequency and discussion volume of different topics; identifies the current hot topics according to the set threshold and determines their popularity values;

[0024] Topic source judgment unit: Analyze the relevant text data of hot topics, and judge whether the topic is a local topic, a domestic topic or an international topic according to the geographical information and keyword features in it;

[0025] Category label adding unit: Add corresponding category labels to each hot topic according to the content and nature of the topic.

[0026] Preferably, the trend analysis module includes:

[0027] Time interval division unit: Divide the continuous time series into multiple unit time intervals; Count and analyze the discussion volume of hot topics within each time interval;

[0028] Heat trend judgment unit: Compare the change of the discussion volume of the same hot topic in different time intervals, and judge whether its heat trend is rising or falling;

[0029] The trend analysis module retrieves the quantity of each feature label within the set unit time interval according to the collected text data, and obtains the recent hot topics according to the quantity and proportion of the corresponding feature labels.

[0030] Preferably, the matching and comparison module includes:

[0031] Label matching unit: Search whether there are topics with the same content labels and category labels within the set time range; Quickly and accurately find the matching topics through the establishment of label indexes and similarity calculation methods;

[0032] Heat comparison unit: Compare the heat values of the matched topics and calculate the change of the two heat values.

[0033] Preferably, the trend analysis module searches whether there are the same content labels and category labels within the set time; If there is a match, analyze the heat trend of the hot topic, and the system provides corresponding investment suggestions according to the belonging trend;

[0034] The heat trends include: rising trend, stable trend, and falling trend. Among them, the rising trend includes the initial stage of rising, the middle stage of rising, and the final stage of rising. The stable trend includes the high-heat stable period and the low-heat stable period; The falling trend includes the initial stage of falling, the middle stage of falling, and the final stage of falling;

[0035] The stable trend is further divided into the rising relay stable period and the falling relay stable period according to the heat trend of the previous unit time interval;

[0036] The determination methods for the rising relay stable period and the falling relay stable period are as follows: The trend analysis module determines a stable trend when the change in the discussion volume of a hot topic within a unit time interval is between ±5% compared to the discussion volume in the previous unit time interval. Retrieve the previous trend of the current trend. If the previous trend is an upward trend, it is determined as the rising relay stable period; if the previous trend is a downward trend, it is determined as the falling relay stable period.

[0037] Preferably, the scoring mechanism module includes:

[0038] Topic heat scoring unit: Considering comprehensively the discussion volume, dissemination range, and influence factors of the topic, formulate a set of scoring systems for each hot topic;

[0039] Investment risk scoring unit: Combining financial transaction data and market condition factors, evaluate and score the investment risks related to different topics;

[0040] The investment advice generation module includes:

[0041] Trend investment advice unit: Provide trend-based investment advice for investors according to the result of the heat trend analysis of hot topics;

[0042] Matching and comparing investment advice unit: Provide targeted investment advice for investors according to the results of label matching and heat comparison;

[0043] Comprehensive investment advice unit: Combine the heat score and risk score provided by the scoring mechanism module, consider various factors comprehensively, and generate comprehensive investment advice for investors.

[0044] Preferably, the intelligent algorithm optimization module includes:

[0045] Machine learning unit: Use machine learning algorithms to learn and train historical data, establish a prediction model; continuously optimize the model parameters to improve the accuracy and reliability of hot topic recognition, heat trend prediction, and investment advice generation;

[0046] Among them, when learning and training historical data based on a decision tree, the specific formula is:

[0047]

[0048] Among them:

[0049] H(D) represents the entropy of the dataset D, and entropy is used to measure the uncertainty of data; the calculation formula of entropy is where p i is the probability of category i in the dataset D;

[0050] Values(A) represents the set of all values of attribute A;

[0051] D v represents the subset of the dataset D where the value of attribute A is v;

[0052] represents the proportion of the subset of the dataset D where the value of attribute A is v in the entire dataset D;

[0053] Among them, learning and training are performed on historical data based on the neural network algorithm, and the specific formula is:

[0054] Output formula of the output layer neuron:

[0055]

[0056] y represents the output of the neuron; σ represents the activation function, w i represents the connection weight between the i-th input and the neuron; x i represents the i-th input feature; b represents the bias term of the neuron;

[0057] Error backpropagation formula:

[0058] First, calculate the error of the output layer

[0059] y j represents the actual output of the j-th neuron; represents the expected output of the j-th neuron; f' represents the derivative of the activation function;

[0060] Then, calculate the error of the hidden layer

[0061] output represents the set of output layer neurons;

[0062] Weight update formula:

[0063] Δw ji =-ηδ j x i

[0064] Among them, η represents the learning rate, δ j represents the error of neuron j, x i represents the i-th input feature.

[0065] Preferably: The intelligent algorithm optimization module further includes:

[0066] Deep learning unit: Adopting deep learning technology, performing in-depth feature extraction and analysis on financial text data;

[0067] Feature selection is performed based on a sequence feature selection algorithm, and the specific formula is as follows:

[0068]

[0069] Where:

[0070] J(S) represents the score of the feature subset S;

[0071] P(S) represents the set of all binary subsets of the feature subset S;

[0072] d avg (C) represents the average distance between two features in the binary subset C;

[0073] a represents a regularization parameter used to control the complexity of the model.

[0074] The beneficial effects of the present invention are as follows:

[0075] 1. The system of the present invention covers multiple modules such as data collection, preprocessing, topic analysis, trend analysis, matching comparison, scoring mechanism, and investment advice generation, forming a complete digital financial management ecosystem; the system can identify hot topics, determine their sources, and provide personalized investment advice according to the risk preferences of investors and market conditions; this kind of accuracy and personalized service helps to improve the satisfaction and investment success rate of investors.

[0076] 2. The system of the present invention can collect and analyze data in real time, update hot topics and trend changes in a timely manner, and provide the latest information and advice for investors; this real-time and dynamic feature enables the system to quickly respond to market changes and provide timely investment opportunities for investors.

[0077] 3. Through the scoring mechanism module of the present invention, the system can comprehensively evaluate the topic popularity and investment risk, and provide risk warnings and control suggestions for investors; this helps to reduce investment risks and protect the interests of investors.

[0078] 4. Through the intelligent algorithm optimization module of the present invention, the system can use machine learning and deep learning technologies to learn and train historical data, and continuously optimize performance and accuracy; this not only improves the intelligent level of the system, but also reduces the need for manual intervention, improving efficiency and accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0079] Figure 1 FIG. is a system working flowchart of a digital financial management system based on big data proposed by the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0080] The technical solutions of the present invention will be further described in detail below in conjunction with specific embodiments.

[0081] Example 1:

[0082] A digital financial management system based on big data, comprising:

[0083] A data acquisition module that collects financial-related data from multiple channels through a web crawler unit and a data interface unit, and stores it in a raw database;

[0084] A data preprocessing module that performs operations such as cleaning, word segmentation, and annotation on the raw data to obtain high-quality text data, with each piece of text data as an independent data information;

[0085] A topic analysis module that analyzes the preprocessed text data, identifies hot topics, determines their sources, and adds feature tags;

[0086] A trend analysis module that divides the time series into multiple unit time intervals, analyzes the change in the discussion volume of hot topics in different time intervals, and determines their heat trends;

[0087] A matching and comparison module that searches for topics with the same or similar feature tags within a set time and conducts a comparison of heat values;

[0088] A scoring mechanism module that scores the topic heat according to factors such as the discussion volume, dissemination range, and influence of the topic, and at the same time scores the investment risk in combination with financial transaction data and market conditions;

[0089] An investment advice generation module that generates personalized investment advice for investors based on the trend analysis results, label matching and comparison results, and the scores provided by the scoring mechanism;

[0090] An intelligent algorithm optimization module that uses machine learning and deep learning algorithms to learn and train historical data to optimize the performance and accuracy of the system.

[0091] Among them, the data acquisition module includes:

[0092] A web crawler unit: responsible for collecting financial-related text data from multiple channels such as major news websites, social media platforms, and financial forums, including news reports, user comments, professional analysis articles, etc.; by setting specific keywords and screening rules, accurately capturing content related to financial topics and storing it in the raw database;

[0093] A data interface unit: establishes a data interface with various financial data providers to obtain real-time financial transaction data, market data, and other structured data to provide data support for subsequent analysis.

[0094] Among them, the data preprocessing module includes:

[0095] Data cleaning unit: Clean the collected raw data to remove noise data, duplicate data, and invalid data; for example, identify and delete advertising content, content with formatting errors, etc. in news reports to ensure data quality.

[0096] Text segmentation and annotation unit: Perform word segmentation on the cleaned text data to break the text into meaningful words or phrases; at the same time, according to the preset financial term library and semantic rules, annotate the segmented data for subsequent topic analysis and content tagging.

[0097] Among them, the feature tags include: category tags and content tags. Among them, the category tags are divided according to the topic area, including local topics, domestic topics, or international topics.

[0098] Among them, the topic analysis module includes:

[0099] Hot topic recognition unit: Through statistical analysis of the preprocessed text data, calculate the occurrence frequency and discussion volume of different topics; according to the set threshold, identify the current hot topics and determine their heat values; for example, the TF-IDF (term frequency-inverse document frequency) algorithm can be used to calculate the importance of words in the text to determine the heat of the topic.

[0100] Topic source judgment unit: Analyze the relevant text data of the hot topics and judge whether the topic is a local topic, a domestic topic, or an international topic according to the geographical information, keyword features, etc. in it; for example, if the text mentions a specific city name or regional event, it is judged as a local topic; if it involves national policies, events, etc., it is a domestic topic; if it contains relevant information of multiple countries or regions, it is an international topic.

[0101] Category tag adding unit: Add corresponding category tags to each hot topic according to the content and nature of the topic; for example, for topics involving the application of artificial intelligence technology in the financial field, category tags such as "AI" and "fintech" can be added.

[0102] Among them, the trend analysis module includes:

[0103] Time interval division unit: Divide the continuous time series into multiple unit time intervals, such as in days, weeks, or months; count and analyze the discussion volume of hot topics within each time interval.

[0104] Heat trend judgment unit: Compare the change in the discussion volume of the same hot topic in different time intervals to judge whether its heat trend is rising or falling; for example, a linear regression model or a moving average method can be used to analyze the change trend of the discussion volume.

[0105] The trend analysis module searches for the number of each feature tag within a set unit time interval based on the collected text data, and obtains recent hot topics according to the number and proportion of the corresponding feature tags.

[0106] Wherein, the matching and comparison module includes:

[0107] Tag matching unit: Search for topics with the same content tags and category tags within a set time range (e.g., 6 months); quickly and accurately find matching topics by establishing a tag index and similarity calculation method;

[0108] Heat comparison unit: compares the heat values ​​of the matched topics and calculates the change (increase or decay) between the two heat values; for example, the relative change rate can be used to measure the degree of change in the heat value.

[0109] Among them, the trend analysis module searches for the existence of the same content tags and category tags within a set time (such as three unit time intervals); if there is a match, the popularity trend of the hot topic is analyzed (for example, when the number and / or proportion of a certain content tag increases in the current unit time interval compared to the previous unit time interval, it is determined that the discussion volume of the hot topic is on an upward trend), and the system provides corresponding investment advice based on the corresponding trend.

[0110] For example, for two hot topics with the same content tags, in the A-th unit time interval, the noun of the hot topic is: Doubao (AI, large model, domestic topic); in the B-th unit time interval, the noun of the hot topic is: deepseek (AI, large model, international topic); among them, the category label is upgraded from domestic topic to international topic, which is determined as a topic upgrade. When the number and proportion of corresponding content tags increase and / or the proportion increases in the B-th unit time interval compared with that in the A-th unit time interval, it is determined that the discussion volume of the hot topic has increased; investment advice is provided based on the upgrade of the topic and the increase in the discussion volume.

[0111] In order to better judge the popularity of the current topic and the future trend development, the popularity trend includes: rising trend, stable trend, and falling trend. The rising trend includes the early rising period, the middle rising period, and the late rising period. The stable trend includes the high popularity stable period and the low popularity stable period. The falling trend includes the early falling period, the middle falling period, and the late falling period.

[0112] Among them, the stable trend is further divided into the rising relay stable period and the falling relay stable period according to the heat trend in the previous unit time interval. When providing investment advice, when the heat trend is in the initial rising stage and the mid-rising stage of the rising trend, as well as the rising relay stable period of the stable trend, more positive investment advice (such as starting to invest, increasing investment, etc.) is provided; when the heat trend is in the mid-falling stage of the falling trend, as well as the falling relay stable period of the stable trend, more negative investment advice (such as reducing investment, stopping investment, etc.) is provided.

[0113] Among them, the determination methods of the rising relay stable period and the falling relay stable period are as follows: the trend analysis module, based on the discussion volume of the hot topic, determines a stable trend when the change in the discussion volume within a unit time interval compared to the discussion volume in the previous unit time interval is between ±5%; retrieves the previous trend of the current trend. If the previous trend is an upward trend, it is determined as the rising relay stable period; if the previous trend is a downward trend, it is determined as the falling relay stable period.

[0114] Among them, the scoring mechanism module includes:

[0115] Topic heat scoring unit: Considering factors such as the discussion volume, dissemination scope, and influence of the topic comprehensively, formulating a set of scoring systems for each hot topic; for example, the heat score of the topic can be obtained by weighted calculation based on indicators such as the dissemination volume and citation times of the topic on different platforms;

[0116] Investment risk scoring unit: Combining factors such as financial transaction data and market conditions, evaluating and scoring the investment risks related to different topics; for example, for investment topics involving emerging technologies, due to their high uncertainty, a relatively high risk score can be given.

[0117] Among them, the investment advice generation module includes:

[0118] Trend investment advice unit: Based on the analysis results of the heat trend of the hot topic, providing trend-based investment advice for investors; for example, if the heat of a certain topic shows an upward trend and the related field has good development prospects, investors can be advised to pay attention to the investment opportunities in this field; conversely, if the heat decreases, investors are reminded to invest cautiously;

[0119] Matching and comparing investment advice unit: Based on the results of label matching and heat comparison, providing targeted investment advice for investors; for example, if the heat of a certain topic rises again in a short period of time and the content labels are similar to those before, investors can be prompted to pay attention to the possible reappearance of investment opportunities related to this topic; if the heat decays, investors are advised to avoid blindly following the trend for investment;

[0120] Comprehensive Investment Recommendation Unit: Combining the popularity score and risk score provided by the scoring mechanism module, considering various factors comprehensively, generate comprehensive investment recommendations for investors; for example, for investment projects with high popularity but also high risks, it can be recommended that investors make reasonable allocations according to their own risk tolerance.

[0121] Among them, the intelligent algorithm optimization module includes:

[0122] Machine Learning Unit: Using machine learning algorithms such as decision trees and neural networks to learn and train historical data, establish a prediction model. By continuously optimizing the model parameters, improve the accuracy and reliability of hot topic recognition, popularity trend prediction, and investment recommendation generation;

[0123] Among them, when learning and training historical data based on decision trees, the specific formula is:

[0124]

[0125] Among them:

[0126] H(D) represents the entropy of the dataset D, and entropy is used to measure the uncertainty of the data; the calculation formula of entropy is where p i is the probability of category i in the dataset D;

[0127] Values(A) represents the set of all values of the attribute A;

[0128] D v represents the subset of the dataset D where the value of the attribute A is v;

[0129] represents the proportion of the subset of the dataset D where the value of the attribute A is v in the entire dataset D;

[0130] Among them, when learning and training historical data based on neural network algorithms, the specific formula is:

[0131] Output formula of the output layer neuron:

[0132]

[0133] y represents the output of the neuron; σ represents the activation function, w i represents the connection weight between the i-th input and the neuron; x i represents the i-th input feature; b represents the bias term of the neuron;

[0134] Error backpropagation formula:

[0135] First calculate the error of the output layer

[0136] y j represents the actual output of the j-th neuron; represents the expected output of the j-th neuron; f' represents the derivative of the activation function;

[0137] Then calculate the error of the hidden layer

[0138] output represents the set of neurons in the output layer;

[0139] Weight update formula:

[0140] Δw ji = -ηδ j x i

[0141] where η represents the learning rate, and δ j represents the error of neuron j, and x i represents the i-th input feature;

[0142] Deep learning unit: Adopt deep learning technologies, such as convolutional neural network (CNN), recurrent neural network (RNN), etc., to perform in-depth feature extraction and analysis on financial text data; mine the potential information and laws hidden in the data, and provide more powerful support for investment decisions;

[0143] Based on the sequence feature selection algorithm, perform feature selection, and the specific formula is:

[0144]

[0145] where:

[0146] J(S) represents the score of the feature subset S;

[0147] P(S) represents the set of all binary subsets of the feature subset S;

[0148] d avg (C) represents the average distance between two features in the binary subset C;

[0149] a represents the regularization parameter, which is used to control the complexity of the model.

[0150] The working process of the system is as follows:

[0151] S1: The data acquisition module collects financial-related data from multiple channels through the web crawler unit and the data interface unit, and stores it in the original database;

[0152] S2: The data preprocessing module performs operations such as cleaning, word segmentation, and annotation on the original data to obtain high-quality text data, and each piece of text data serves as an independent data information;

[0153] S3: The topic analysis module analyzes the preprocessed text data, identifies hot topics, determines their sources, and adds feature tags;

[0154] S4: The trend analysis module divides the time series into multiple unit time intervals, analyzes the change in the discussion volume of hot topics in different time intervals, and determines their heat trends;

[0155] S5: The matching and comparison module searches for topics with the same or similar feature tags within a set time and conducts a comparison of heat values;

[0156] S6: The scoring mechanism module scores the topic heat based on factors such as the discussion volume, dissemination range, and influence of the topic, and at the same time scores the investment risk in combination with financial transaction data and market conditions;

[0157] S7: The investment advice generation module generates personalized investment advice for investors based on the trend analysis results, label matching and comparison results, and the scores provided by the scoring mechanism;

[0158] S8: The intelligent algorithm optimization module uses machine learning and deep learning algorithms to learn and train historical data, and continuously optimizes the performance and accuracy of the system.

[0159] As described above, it is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.

Claims

1. A digital financial management system based on big data, characterized in that: include: The data collection module collects financial related data through the web crawler unit and the data interface unit and stores it in the original database; The data preprocessing module cleans, segments and annotates the original data to obtain text data; Topic analysis module, which analyzes the preprocessed text data, identifies hot topics, determines their sources and adds feature tags; The trend analysis module divides the time series into multiple unit time intervals, analyzes the changes in the discussion volume of hot topics in different time intervals, and determines their popularity trends; The matching and comparison module searches for topics with the same or similar feature tags within a set time and compares their popularity values; The scoring mechanism module scores the popularity of topics and scores investment risks based on financial transaction data and market conditions; The investment advice generation module generates personalized investment advice for investors based on trend analysis results, tag matching comparison results, and the scores provided by the scoring mechanism; Intelligent algorithm optimization module uses machine learning and deep learning algorithms to learn and train historical data.

2. A digital financial management system based on big data according to claim 1, characterized in that: The data acquisition module comprises: Web crawler unit: collects financial-related text data from various channels; by setting specific keywords and screening rules, accurately captures content related to financial topics and stores it in the original database; Data interface unit: Establish data interfaces with various financial data providers to obtain real-time financial transaction data, market data and structured data.

3. A digital financial management system based on big data according to claim 1, characterized in that: The data preprocessing module comprises: Data cleaning unit: cleans the collected raw data to remove noise data, duplicate data and invalid data; Text segmentation and annotation unit: performs segmentation on the cleaned text data and decomposes the text into meaningful words or phrases. At the same time, the segmented data is annotated according to the preset financial term library and semantic rules. The feature tags include: category tags and content tags, wherein the category tags are divided according to the region of the topic, including local topics, domestic topics or international topics.

4. A digital financial management system based on big data according to claim 1, characterized in that: The topic analysis module includes: Hot topic identification unit: Calculate the frequency of occurrence and discussion volume of different topics by performing statistical analysis on the preprocessed text data; identify the current hot topic and determine its popularity value according to the set threshold; Topic source judgment unit: Analyze the relevant text data of hot topics and judge whether the topic is local, domestic or international based on the geographical information and keyword features; Category label adding unit: Add corresponding category labels to each hot topic according to the content and nature of the topic.

5. A digital financial management system based on big data according to claim 4, characterized in that: The trend analysis module includes: Time interval division unit: divide the continuous time series into multiple unit time intervals; count and analyze the discussion volume of hot topics in each time interval; Hot trend judgment unit: compares the changes in the discussion volume of the same hot topic in different time intervals to determine whether its hot trend is rising or falling; The trend analysis module searches for the number of each feature tag within a set unit time interval based on the collected text data, and obtains recent hot topics according to the number and proportion of the corresponding feature tags.

6. A digital financial management system based on big data according to claim 5, characterized in that: The matching and comparison module includes: Tag matching unit: Search for topics with the same content tags and category tags within a set time range; quickly and accurately find matching topics by establishing a tag index and similarity calculation method; Heat comparison unit: compares the heat values ​​of the matched topics and calculates the change between the two heat values.

7. A digital financial management system based on big data according to claim 6, characterized in that: The trend analysis module searches for the existence of the same content tags and category tags within a set period of time; If there is a match, the popularity trend of the hot topic will be analyzed, and the system will provide corresponding investment suggestions based on the trend; The heat trend includes: an upward trend, a stable trend, and a downward trend. The upward trend includes an early rising period, a middle rising period, and a late rising period. The stable trend includes a high heat stable period and a low heat stable period. The downward trend includes an early falling period, a middle falling period, and a late falling period. The stable trend is further divided into the rising relay stable period and the falling relay stable period according to the heat trend of the previous unit time interval; The method for determining the rising relay stable period and the falling relay stable period is that the trend analysis module determines it as a stable trend based on the discussion volume of the hot topic. When the discussion volume in a unit time interval changes from the discussion volume in the previous unit time interval by ±5%, the trend of the previous trend of the current trend is retrieved. If the previous trend is an upward trend, it is determined to be an rising relay stable period. If the previous trend is a downward trend, it is determined to be a falling relay stable period.

8. A digital financial management system based on big data according to claim 7, characterized in that: The scoring mechanism module includes: Topic popularity scoring unit: comprehensively consider the discussion volume, dissemination scope, and influence factors of the topic, and develop a scoring system for each hot topic; Investment risk scoring unit: Combine financial transaction data and market factors to evaluate and score investment risks related to different topics; The investment advice generating module comprises: Trend investment advice unit: Provide investors with trend-based investment advice based on the results of the trend analysis of hot topics; Matching and comparison investment advice unit: Provide investors with targeted investment advice based on the results of tag matching and popularity comparison; Comprehensive investment advice unit: Combines the heat score and risk score provided by the scoring mechanism module, comprehensively considers various factors, and generates comprehensive investment advice for investors.

9. A digital financial management system based on big data according to claim 1, characterized in that: The intelligent algorithm optimization module includes: Machine Learning Unit: Use machine learning algorithms to learn and train historical data and build prediction models; improve the accuracy and reliability of hot topic identification, popularity trend prediction, and investment advice generation by continuously optimizing model parameters; Among them, the historical data is learned and trained based on the decision tree, and the specific formula is: in: H(D) represents the entropy of the data set D. Entropy is used to measure the uncertainty of data. The calculation formula of entropy is where p i is the probability of category i in data set D; Values(A) represents the set of all values ​​of attribute A; D v It represents the subset of the data set D whose attribute A takes the value v; It represents the proportion of the subset of attribute A in data set D with value v in the entire data set D; Among them, the historical data is learned and trained based on the neural network algorithm. The specific formula is: The output formula of the output layer neurons is: y represents the output of the neuron; σ represents the activation function, w i represents the connection weight between the i-th input and the neuron; x i represents the i-th input feature; b represents the bias term of the neuron; Error back propagation formula: First, calculate the error of the output layer y j represents the actual output of the jth neuron; represents the expected output of the jth neuron; f' represents the derivative of the activation function; Then calculate the error of the hidden layer Output represents the set of neurons in the output layer; Weight update formula: Δw ji =-ηδ j x i Among them, η represents the learning rate, δ j represents the error of neuron j, x i represents the i-th input feature.

10. A digital financial management system based on big data according to claim 9, characterized in that: The intelligent algorithm optimization module also includes: Deep Learning Unit: Use deep learning technology to perform in-depth feature extraction and analysis of financial text data; Based on the sequence feature selection algorithm, feature selection is performed. The specific formula is: in: J(S) represents the score of feature subset S; P(S) represents the set of all binary subsets of feature subset S; d avg (C) represents the average distance between two features in the binary subset C; a represents the regularization parameter, which is used to control the complexity of the model.

Citation Information

Patent Citations

  • Financial investment risk management system and method based on big data

    CN115018354A

Cited By

  • Explanatable dynamic graph network income prediction model training method based on social media

    CN122472895A