Data prediction method and system based on large model, and storage medium
Through the data prediction method based on large models, the data prediction process is automatically processed, which solves the problem of low data prediction accuracy in the existing technology, and achieves higher prediction accuracy and automated strategy recommendation generation.
Patent Information
- Application Number
- CN202510126946.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-27
- Publication Date
- 2025-05-30
AI Technical Summary
In the prior art, the accuracy of data prediction is low, mainly because the prediction results are greatly affected by artificial subjectiveness.
The data prediction method based on the big model is adopted. By obtaining the object acquisition data of the object to be predicted, inputting the pre-trained large model for trend prediction, and obtaining the object prediction trend; then input the object prediction trend into the pre-trained agent for model simulation to predict risk prediction; at the same time, the search enhancement generation model for the pre-trained object input into the pre-trained object to be predicted is retrieved and generated to obtain object generation information; finally, the strategy suggestions are generated based on the object generation information and risk prediction results.
Through the automated data prediction process, the manual subjective impact is reduced, the accuracy of data prediction is improved, the strategy recommendations can be automatically generated, and the efficiency and accuracy of investment analysis are improved.
Smart Images

Figure CN120069193A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly to a data prediction method, system and storage medium based on a large model. Background Art
[0002] With the rise of the Internet and big data, many investment and financing platforms have emerged. The investment and financing platforms classify and arrange investment targets by industry, which is convenient for users to query and understand investment targets. To reduce investment risks, users need to conduct investment analysis and prediction on investment targets. Therefore, the problem of data prediction for investment targets has attracted more and more attention.
[0003] In the existing data prediction process, generally, the method of manual analysis and prediction is used to predict the risk or return of investment targets, and corresponding strategy suggestions are determined based on the prediction results. However, since the prediction results are greatly affected by human subjectivity, the accuracy of data prediction is low. Summary of the Invention
[0004] The purpose of the embodiments of the present invention is to provide a data prediction method, system and storage medium based on a large model to solve the problem of low accuracy of data prediction in the prior art.
[0005] The embodiments of the present invention are implemented as follows. A data prediction method based on a large model, the method includes:
[0006] Obtain the object acquisition data of the object to be predicted, and input the object acquisition data into the pre-trained large model for trend prediction to obtain the object prediction trend;
[0007] Input the object prediction trend into the pre-trained intelligent agent for model simulation to obtain the object simulation model, and perform risk prediction on the object to be predicted according to the object simulation model to obtain the risk prediction result;
[0008] Input the object to be predicted into the pre-trained retrieval-augmented generation model for retrieval and generation to obtain the object generation information;
[0009] Generate the strategy suggestion of the object to be predicted according to the object generation information and the risk prediction result.
[0010] Preferably, before inputting the object acquisition data into the pre-trained large model for trend prediction to obtain the object prediction trend, it further includes:
[0011] Obtain the object sample data, and input the object sample data into the large model for key data extraction to obtain key text, key images and key audio;
[0012] Feature extraction is performed on the key text, the key image, and the key audio to obtain text features, image features, and audio features, and the text features, the image features, and the audio features are feature - fused to obtain sample fusion features;
[0013] Trend prediction is performed based on the sample fusion features to obtain a sample prediction trend, and a model loss is determined based on the sample prediction trend, the key text, the key image, the key audio, the text features, the image features, the audio features, and the sample fusion features;
[0014] The parameters of the large model are updated according to the model loss until the large model converges, and the pre - trained large model is obtained.
[0015] Preferably, inputting the object sample data into the large model for key data extraction to obtain key text, key image, and key audio includes:
[0016] Obtaining the number of reposts and the number of comments of the text information, image information, and audio information in the object sample data according to the large model, and determining the information heat according to the number of reposts and the number of comments;
[0017] Performing heat sorting on the text information, the image information, and the audio information according to the information heat, and performing key data extraction on the text information, the image information, and the audio information according to the heat sorting result to obtain the key text, the key image, and the key audio.
[0018] Preferably, before inputting the object prediction trend into the pre - trained agent for model simulation to obtain an object simulation model, it further includes:
[0019] Obtaining an object trend sample, and inputting the object trend sample into the agent for model simulation to obtain a sample simulation model;
[0020] Performing risk prediction on the sample object corresponding to the object trend sample according to the sample simulation model to obtain a sample prediction result, and determining a learning loss according to the sample prediction result and the sample simulation model;
[0021] Updating the parameters of the agent according to the learning loss until convergence, and obtaining the pre - trained agent.
[0022] Preferably, before inputting the object to be predicted into the pre - trained retrieval - augmented generation model for retrieval and generation to obtain object generation information, it further includes:
[0023] Query the sample object in a preset database according to the retrieval enhancement generation model to obtain sample query information, and generate information from the sample query information to obtain sample generation information;
[0024] Determine a generation loss according to the sample generation information, and update the parameters of the retrieval enhancement generation model according to the generation loss until convergence to obtain the pre-trained retrieval enhancement generation model.
[0025] Preferably, obtaining object acquisition data of an object to be predicted includes:
[0026] Obtain the object address of the object to be predicted, and obtain local policy information by acquiring policy information according to the object address;
[0027] Obtain the object name of the object to be predicted, and obtain object news information, object litigation information, and object securities information by acquiring news information, litigation information, and securities information according to the object name;
[0028] Combine the object news information, the object litigation information, the object securities information, and the local policy information to obtain the object acquisition data.
[0029] Preferably, determining a model loss according to the sample prediction trend, the key text, the key image, the key audio, the text feature, the image feature, the audio feature, and the sample fusion feature includes:
[0030] Calculate the trend similarity between the sample prediction trend and the standard prediction trend, and determine a first loss according to the trend similarity;
[0031] Calculate the text similarity between the key text and the standard text, and determine a second loss according to the text similarity;
[0032] Calculate the image similarity between the key image and the standard image, and determine a third loss according to the image similarity;
[0033] Calculate the audio similarity between the key audio and the standard audio, and determine a fourth loss according to the audio similarity;
[0034] Calculate the feature similarity between the text feature and the first standard feature to obtain a first feature similarity, and determine a fifth loss according to the first feature similarity;
[0035] Calculate the feature similarity between the image feature and the second standard feature to obtain a second feature similarity, and determine a sixth loss according to the second feature similarity;
[0036] Calculate the feature similarity between the audio feature and the third standard feature to obtain a third feature similarity, and determine a seventh loss according to the third feature similarity;
[0037] Calculate the feature similarity between the sample fusion feature and the standard fusion feature to obtain a fourth feature similarity, and determine an eighth loss according to the fourth feature similarity;
[0038] Perform a weighted operation on the first loss, the second loss, the third loss, the fourth loss, the fifth loss, the sixth loss, the seventh loss, and the eighth loss to obtain the model loss.
[0039] Another object of the embodiments of the present invention is to provide a data prediction system based on a large model, and the system includes:
[0040] A trend prediction module, configured to obtain object acquisition data of an object to be predicted, and input the object acquisition data into a pre-trained large model for trend prediction to obtain an object prediction trend;
[0041] A risk prediction module, configured to input the object prediction trend into a pre-trained agent for model simulation to obtain an object simulation model, and perform risk prediction on the object to be predicted according to the object simulation model to obtain a risk prediction result;
[0042] A retrieval and generation module, configured to input the object to be predicted into a pre-trained retrieval-enhanced generation model for retrieval and generation to obtain object generation information;
[0043] A suggestion generation module, configured to generate a policy suggestion for the object to be predicted according to the object generation information and the risk prediction result.
[0044] Preferably, the trend prediction module is further configured to:
[0045] Obtain object sample data, and input the object sample data into the large model for key data extraction to obtain key text, key images, and key audio;
[0046] Extract features from the key text, the key images, and the key audio to obtain text features, image features, and audio features, and perform feature fusion on the text features, the image features, and the audio features to obtain a sample fusion feature;
[0047] Perform trend prediction according to the sample fusion feature to obtain a sample prediction trend, and determine a model loss according to the sample prediction trend, the key text, the key images, the key audio, the text features, the image features, the audio features, and the sample fusion feature;
[0048] Update the parameters of the large model according to the model loss until the large model converges, and obtain the pre-trained large model.
[0049] In an embodiment of the present invention, by inputting the object acquisition data into the pre-trained large model for trend prediction, the trend of the object to be predicted can be automatically predicted to obtain the object prediction trend. By inputting the object prediction trend into the pre-trained agent for model simulation, an object simulation model corresponding to the object to be predicted can be effectively generated. Based on the object simulation model, the risk of the object to be predicted can be effectively predicted to obtain the risk prediction result. By inputting the object to be predicted into the pre-trained retrieval-augmented generation model for retrieval generation, information of the object to be predicted can be effectively generated automatically. Based on the object-generated information and the risk prediction result, a policy recommendation for the object to be predicted can be automatically generated, without the need to manually process data of the investment object, preventing the phenomenon of low data prediction accuracy caused by artificial subjective influence. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 is a flowchart of a data prediction method based on a large model provided in the first embodiment of the present invention;
[0051] Figure 2 is a schematic structural diagram of a data prediction system based on a large model provided in the second embodiment of the present invention;
[0052] Figure 3 is a schematic framework diagram of a data prediction system based on a large model provided in the second embodiment of the present invention;
[0053] Figure 4 is a schematic structural diagram of a terminal device provided in the third embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0054] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, but not to limit the present invention.
[0055] In order to illustrate the technical solutions of the present invention, the following will be described through specific embodiments.
[0056] Embodiment 1
[0057] Please refer to Figure 1 , which is a flowchart of a data prediction method based on a large model provided in the first embodiment of the present invention. The data prediction method based on a large model can be applied to any device or system. The data prediction method based on a large model includes the following steps:
[0058] Step S10, obtain the object collection data of the object to be predicted, and input the object collection data into the pre-trained large model for trend prediction to obtain the object prediction trend;
[0059] Among them, the object collection data is obtained by collecting data on the object to be predicted. The object collection data includes information such as news information, litigation information, securities information, and local policies. By inputting the object collection data into the pre-trained large model for trend prediction, the income trend and risk trend of the object to be predicted can be effectively analyzed.
[0060] Optionally, before inputting the object collection data into the pre-trained large model for trend prediction to obtain the object prediction trend, it further includes:
[0061] Obtain object sample data, and input the object sample data into the large model for key data extraction to obtain key text, key images, and key audio; among them, data is widely collected from various data sources, including but not limited to enterprise financial reports, market research reports, policy documents, social media information, etc. The data collected has a wide range of sources, covering all aspects required for investment analysis. Through data collection technologies such as web crawlers and API interfaces, the collection data can be obtained efficiently and accurately. The collection data is integrated to obtain integrated data, and the integrated data is cleaned and formatted to remove redundant, incorrect, and irrelevant information to ensure the accuracy and consistency of the data. At the same time, investment experts are invited to annotate the formatted integrated data to obtain object sample data.
[0062] After completing the data annotation, construct a highly pre-trained large model. The large model integrates a variety of cutting-edge technologies such as natural language processing (NLP), image recognition, and time series analysis, and has excellent data processing and feature extraction capabilities. By deeply training on a large amount of market data, the model can automatically learn and understand the internal laws and correlations of the data, providing strong support for subsequent investment analysis.
[0063] Extract features from the key text, the key images, and the key audio to obtain text features, image features, and audio features, and fuse the text features, the image features, and the audio features to obtain sample fusion features; among them, integrate multi-modal data from different sources and different formats, such as text, images, videos, and audio. Through the application of advanced data fusion technologies, multi-modal data can be seamlessly docked and efficiently utilized. At the same time, use large model technologies to extract key features from the data, such as financial indicators, market trends, policy impacts, etc., providing strong support for subsequent intelligent agent analysis and investment strategy recommendations.
[0064] Perform trend prediction based on the sample fusion feature to obtain the sample prediction trend, and determine the model loss according to the sample prediction trend, the key text, the key image, the key audio, the text feature, the image feature, the audio feature, and the sample fusion feature;
[0065] Update the parameters of the large model according to the model loss until the large model converges to obtain the pre-trained large model.
[0066] Preferably, determining the model loss according to the sample prediction trend, the key text, the key image, the key audio, the text feature, the image feature, the audio feature, and the sample fusion feature includes:
[0067] Calculate the trend similarity between the sample prediction trend and the standard prediction trend, and determine the first loss according to the trend similarity; calculate the text similarity between the key text and the standard text, and determine the second loss according to the text similarity; calculate the image similarity between the key image and the standard image, and determine the third loss according to the image similarity; calculate the audio similarity between the key audio and the standard audio, and determine the fourth loss according to the audio similarity; calculate the feature similarity between the text feature and the first standard feature to obtain the first feature similarity, and determine the fifth loss according to the first feature similarity; calculate the feature similarity between the image feature and the second standard feature to obtain the second feature similarity, and determine the sixth loss according to the second feature similarity; calculate the feature similarity between the audio feature and the third standard feature to obtain the third feature similarity, and determine the seventh loss according to the third feature similarity; calculate the feature similarity between the sample fusion feature and the standard fusion feature to obtain the fourth feature similarity, and determine the eighth loss according to the fourth feature similarity; perform weighted operations on the first loss, the second loss, the third loss, the fourth loss, the fifth loss, the sixth loss, the seventh loss, and the eighth loss to obtain the model loss. Among them, in the weighted operation process, the weighting coefficients of the first loss, the second loss, the third loss, the fourth loss, the fifth loss, the sixth loss, the seventh loss, and the eighth loss can be set according to requirements.
[0068] Further, inputting the object sample data into the large model for key data extraction to obtain key text, key image, and key audio, includes:
[0069] Obtain the number of reposts and the number of comments on the text information, image information, and audio information in the object sample data according to the large model, and determine the information popularity based on the number of reposts and the number of comments; wherein, the number of reposts is the number of times the corresponding information is reposted on the network, and the number of comments is the number of times the corresponding information is commented by users on the network.
[0070] Sort the text information, the image information, and the audio information according to the information popularity, and perform key data extraction on the text information, the image information, and the audio information according to the popularity sorting result to obtain the key text, the key image, and the key audio.
[0071] Furthermore, obtain the object collection data of the object to be predicted, including:
[0072] Obtain the object address of the object to be predicted, and obtain local policy information according to the object address; wherein, match the object address with a policy query table to obtain local policy information, and the policy query table stores the corresponding relationship between different object addresses and the corresponding local policy information.
[0073] Obtain the object name of the object to be predicted, and obtain object news information, object litigation information, and object securities information according to the object name; wherein, match the object name with a news database, a litigation database, and a securities database respectively to obtain object news information, object litigation information, and object securities information.
[0074] Combine the object news information, the object litigation information, the object securities information, and the local policy information to obtain the object collection data.
[0075] Step S20, input the object prediction trend into the pre-trained intelligent agent for model simulation to obtain an object simulation model, and perform risk prediction on the object to be predicted according to the object simulation model to obtain a risk prediction result;
[0076] Among them, by inputting the object prediction trend into the pre-trained intelligent agent for model simulation, an object simulation model corresponding to the object to be predicted can be effectively generated, and based on the object simulation model, the object to be predicted can be effectively risk predicted to obtain a risk prediction result.
[0077] Optionally, before inputting the object prediction trend into the pre-trained intelligent agent for model simulation to obtain an object simulation model, it further includes:
[0078] Obtain an object trend sample and input the object trend sample into the agent for model simulation to obtain a sample simulation model; wherein, the object trend sample can be set according to requirements, and the sample simulation model is used for data simulation;
[0079] Perform risk prediction on the sample object corresponding to the object trend sample according to the sample simulation model to obtain a sample prediction result, and determine a learning loss according to the sample prediction result and the sample simulation model; wherein, by performing risk prediction on the sample object, the revenue risk or operating risk of the sample object is predicted;
[0080] Update the parameters of the agent according to the learning loss until convergence to obtain the pre-trained agent; wherein, after constructing a large model, an agent is constructed, including multiple agents with different functions, such as intelligent investment advisor agents, investment report agents, strategy generation agents, etc. The agents work together through cooperation to achieve the automation and intelligence of industrial investment analysis. At the same time, reinforcement learning technology is adopted to enable the agent to automatically learn and adapt to different investment environments and decision-making scenarios.
[0081] The agent can construct a simulation model of an industrial investment project based on the data analysis results of the large model. Through simulation analysis, the future revenue and risks of the project are predicted, providing intuitive and accurate project evaluation results for investors, helping investors deeply understand the potential value of the project, and also providing a scientific basis for the formulation of investment strategies.
[0082] Step S30, input the object to be predicted into the pre-trained retrieval-augmented generation model for retrieval and generation to obtain object generation information;
[0083] Wherein, in the data processing and analysis process, the retrieval-augmented generation technology is introduced. By combining information retrieval and generative models, relevant knowledge and information are actively retrieved to obtain object generation information.
[0084] Optionally, before inputting the object to be predicted into the pre-trained retrieval-augmented generation model for retrieval and generation to obtain object generation information, it further includes:
[0085] Query the information of the sample object in the preset database according to the retrieval-augmented generation model to obtain sample query information, and perform information generation on the sample query information to obtain sample generation information; wherein, the sample query information is obtained by querying the information of the sample object through the retrieval model in the retrieval-augmented generation model, and the sample generation information is obtained by performing information generation on the sample query information through the generation model in the retrieval-augmented generation model;
[0086] Determine the generation loss according to the sample generation information, and update the parameters of the retrieval-augmented generation model according to the generation loss until convergence to obtain the pre-trained retrieval-augmented generation model; wherein, calculate the similarity between the sample generation information and the standard generation information to obtain the generation similarity, and determine the generation loss based on the generation similarity.
[0087] Step S40, generate a policy recommendation for the object to be predicted according to the object generation information and the risk prediction result;
[0088] Among them, combine the object generation information and the risk prediction result to obtain a policy recommendation for the object to be predicted. At the same time, adopt the COT reasoning mechanism to imitate the thinking process of human experts, gradually disassemble and analyze complex problems by constructing intermediate thinking steps, enhance the knowledge richness and analysis depth, and make the output result closer to the actual needs.
[0089] In this embodiment, based on the powerful analysis capabilities of large models and agents, various forms of investment analysis and reports can be generated. First, it can automatically generate investment policy recommendations. The recommendations are based on big data analysis and intelligent reasoning, and can accurately reflect market trends and investment opportunities, providing scientific decision-making basis for investors. Second, it can also generate detailed investment reports. The reports cover multiple aspects such as the market background, industry analysis, financial status, and risk assessment of investment projects, providing comprehensive and in-depth investment analysis for investors. In addition, it also supports the customized generation of investment reports, providing personalized report content and formats according to the specific needs of investors.
[0090] Industrial investment analysis based on technologies such as large models and agents fundamentally solves multiple pain points existing in existing methods. First, it realizes all-round data integration, covering not only traditional financial data but also information resources from emerging channels such as news media and social platforms. Such multi-source data fusion provides a solid foundation for more comprehensive and in-depth market insights. Second, relying on advanced large language models and COT reasoning mechanisms, the system can conduct multi-level and multi-dimensional logical analysis of complex industrial investment scenarios, significantly improving the accuracy and reliability of predictions. Third, the agent architecture endows the system with high flexibility and scalability, enabling it to flexibly adjust analysis strategies according to real-time market dynamics to ensure the timeliness and practicality of prediction results. More importantly, it has a built-in continuous learning mechanism that can continuously absorb new market cases and user feedback over time, self-optimize and improve. It not only overcomes the problem of knowledge loss caused by personnel mobility but also ensures long-term effectiveness.
[0091] In this embodiment, by inputting the object acquisition data into the pre-trained large model for trend prediction, the trend of the object to be predicted can be automatically predicted to obtain the object prediction trend. By inputting the object prediction trend into the pre-trained agent for model simulation, an object simulation model corresponding to the object to be predicted can be effectively generated. Based on the object simulation model, the risk of the object to be predicted can be effectively predicted to obtain the risk prediction result. By inputting the object to be predicted into the pre-trained retrieval-augmented generation model for retrieval and generation, information of the object to be predicted can be effectively generated automatically. Based on the object-generated information and the risk prediction result, a policy recommendation for the object to be predicted can be automatically generated, without using manual methods to process data of the investment object, preventing the phenomenon of low data prediction accuracy caused by manual subjective influence.
[0092] Embodiment 2
[0093] Please refer to Figure 2 , which is a schematic structural diagram of the large model-based data prediction system 100 provided by the second embodiment of the present invention, including:
[0094] The trend prediction module 10 is used to obtain the object acquisition data of the object to be predicted, and input the object acquisition data into the pre-trained large model for trend prediction to obtain the object prediction trend.
[0095] Optionally, the trend prediction module 10 is further used to: obtain object sample data, and input the object sample data into the large model for key data extraction to obtain key text, key images, and key audio;
[0096] Extract features from the key text, the key images, and the key audio to obtain text features, image features, and audio features, and fuse the text features, the image features, and the audio features to obtain sample fusion features;
[0097] Perform trend prediction according to the sample fusion features to obtain sample prediction trends, and determine model losses according to the sample prediction trends, the key text, the key images, the key audio, the text features, the image features, the audio features, and the sample fusion features;
[0098] Update the parameters of the large model according to the model losses until the large model converges to obtain the pre-trained large model.
[0099] Preferably, the trend prediction module 10 is further used to: calculate the trend similarity between the sample prediction trend and the standard prediction trend, and determine the first loss according to the trend similarity;
[0100] Calculate the text similarity between the key text and the standard text, and determine the second loss according to the text similarity;
[0101] Calculate the image similarity between the key image and the standard image, and determine the third loss according to the image similarity;
[0102] Calculate the audio similarity between the key audio and the standard audio, and determine the fourth loss according to the audio similarity;
[0103] Calculate the feature similarity between the text feature and the first standard feature to obtain the first feature similarity, and determine the fifth loss according to the first feature similarity;
[0104] Calculate the feature similarity between the image feature and the second standard feature to obtain the second feature similarity, and determine the sixth loss according to the second feature similarity;
[0105] Calculate the feature similarity between the audio feature and the third standard feature to obtain the third feature similarity, and determine the seventh loss according to the third feature similarity;
[0106] Calculate the feature similarity between the sample fusion feature and the standard fusion feature to obtain the fourth feature similarity, and determine the eighth loss according to the fourth feature similarity;
[0107] Perform a weighted operation on the first loss, the second loss, the third loss, the fourth loss, the fifth loss, the sixth loss, the seventh loss, and the eighth loss to obtain the model loss.
[0108] Further, the trend prediction module 10 is further configured to: obtain the number of reposts and comments of the text information, image information, and audio information in the object sample data according to the large model, and determine the information popularity according to the number of reposts and the number of comments;
[0109] Sort the text information, the image information, and the audio information according to the information popularity, and perform key data extraction on the text information, the image information, and the audio information according to the popularity sorting result to obtain the key text, the key image, and the key audio.
[0110] Furthermore, the trend prediction module 10 is further configured to: obtain the object address of the object to be predicted, and obtain local policy information according to the object address;
[0111] Obtain the object name of the object to be predicted, and obtain object news information, object litigation information, and object securities information according to the object name by performing news information acquisition, litigation information acquisition, and securities information acquisition;
[0112] Combine the object news information, the object litigation information, the object securities information, and the local policy information to obtain the object acquisition data.
[0113] The risk prediction module 11 is configured to input the object prediction trend into a pre-trained agent for model simulation to obtain an object simulation model, and perform risk prediction on the object to be predicted according to the object simulation model to obtain a risk prediction result.
[0114] Optionally, the risk prediction module 11 is further configured to: obtain an object trend sample, and input the object trend sample into the agent for model simulation to obtain a sample simulation model;
[0115] Perform risk prediction on the sample object corresponding to the object trend sample according to the sample simulation model to obtain a sample prediction result, and determine a learning loss according to the sample prediction result and the sample simulation model;
[0116] Update the parameters of the agent according to the learning loss until convergence to obtain the pre-trained agent.
[0117] The retrieval and generation module 12 is configured to input the object to be predicted into a pre-trained retrieval-enhanced generation model for retrieval and generation to obtain object generation information.
[0118] Optionally, the retrieval and generation module 12 is further configured to: query information about a sample object in a preset database according to the retrieval-enhanced generation model to obtain sample query information, and perform information generation on the sample query information to obtain sample generation information;
[0119] Determine a generation loss according to the sample generation information, and update the parameters of the retrieval-enhanced generation model according to the generation loss until convergence to obtain the pre-trained retrieval-enhanced generation model.
[0120] The recommendation generation module 13 is configured to generate a policy recommendation for the object to be predicted according to the object generation information and the risk prediction result.
[0121] Please refer to Figure 3 , and the specific technical content is as follows:
[0122] I. Data Integration and Preprocessing
[0123] 1. Data Collection and Integration
[0124] First, the system is responsible for widely collecting data from various data sources, including but not limited to corporate financial reports, market research reports, policy documents, social media information, etc. The data sources are extensive, covering all aspects required for investment analysis. Through advanced data collection technologies such as web crawlers and API interfaces, the system can efficiently and accurately obtain data and integrate it to provide comprehensive data support for subsequent analysis.
[0125] 2. Data Preprocessing and Annotation
[0126] After data collection, the system cleans, integrates, and formats the data, removing redundant, incorrect, and irrelevant information to ensure data accuracy and consistency. At the same time, investment experts are invited to annotate the data to form a high-quality training sample set, providing a reliable data basis for subsequent model training and optimization.
[0127] II. Large Model Training and Optimization Technologies
[0128] 1. Construction of Pre-trained Large Model
[0129] After data preprocessing and annotation are completed, the system begins to construct a highly pre-trained large model. This model integrates various cutting-edge technologies such as natural language processing (NLP), image recognition, and time series analysis, and has excellent data processing and feature extraction capabilities. Through in-depth training on a large amount of market data, the model can automatically learn and understand the internal laws and correlations of the data, providing strong support for subsequent investment analysis.
[0130] 2. Multimodal Data Fusion and Feature Extraction
[0131] The system further integrates multimodal data from different sources and formats, such as text, images, videos, and audio. Through the application of advanced data fusion technologies, multimodal data can be seamlessly docked and efficiently utilized. At the same time, key features such as financial indicators, market trends, and policy impacts are extracted from the data using large model technologies, providing strong support for subsequent intelligent agent analysis and investment strategy recommendations.
[0132] III. Intelligent Agent Decision-making and Collaboration Technologies
[0133] 1. Intelligent Agent Construction
[0134] After constructing the large model, the system begins to build intelligent agents. The system includes multiple intelligent agents with different functions, such as intelligent investment advisor agents, investment report agents, and strategy generation agents. The intelligent agents work together collaboratively to achieve the automation and intelligence of industrial investment analysis. At the same time, reinforcement learning technology is adopted to enable the intelligent agents to automatically learn and adapt to different investment environments and decision-making scenarios.
[0135] 2. Simulation and Emulation Analysis
[0136] The system can build a simulation model of industrial investment projects based on the data analysis results of large models. Through emulation analysis, it predicts the future returns and risks of projects, providing intuitive and accurate project evaluation results for investors. This process not only helps investors deeply understand the potential value of projects but also provides a scientific basis for formulating investment strategies.
[0137] IV. Intelligent Analysis and Generation System Based on Large Models
[0138] 1. Application of RAG Technology and COT Reasoning Mechanism
[0139] In the process of data processing and analysis, the system introduces RAG technology. By combining information retrieval and generative models, when users ask questions according to specific topics, it actively retrieves relevant knowledge and information and integrates them into the generated analysis results. At the same time, it adopts the COT reasoning mechanism, imitating the thinking process of human experts, and gradually decomposing and analyzing complex problems by constructing intermediate thinking steps. The application of these technologies enhances the knowledge richness and analysis depth of the system, making the output results more in line with actual needs.
[0140] 2. Diversified Investment Analysis and Report Generation
[0141] Based on the powerful analysis capabilities of large models and intelligent agents, the system can generate various forms of investment analysis and reports. First, the system can automatically generate investment strategy suggestions. These suggestions are based on big data analysis and intelligent reasoning, accurately reflecting market trends and investment opportunities and providing a scientific decision-making basis for investors. Second, the system can also generate detailed investment reports, which cover multiple aspects such as the market background, industry analysis, financial status, and risk assessment of investment projects, providing comprehensive and in-depth investment analysis for investors. In addition, the system supports the customized generation of investment reports, providing personalized report content and formats according to the specific needs of investors.
[0142] The industrial investment analysis system based on technologies such as large models and agents fundamentally solves multiple pain points existing in existing methods. First of all, it realizes all-round data integration, covering not only traditional financial data but also information resources from emerging channels such as news media and social platforms. Such multi-source data fusion provides a solid foundation for more comprehensive and in-depth market insights. Secondly, relying on advanced large language models and COT reasoning mechanisms, the system can conduct multi-level and multi-dimensional logical analyses on complex industrial investment scenarios, significantly improving the accuracy and reliability of predictions. Thirdly, the agent architecture endows the system with high flexibility and scalability, enabling it to flexibly adjust analysis strategies according to real-time market dynamics and ensuring the timeliness and practicality of prediction results. More importantly, the system is built with a continuous learning mechanism that can continuously absorb new market cases and user feedback over time to self-optimize and improve. This not only overcomes the problem of knowledge loss caused by personnel mobility but also ensures the long-term effectiveness of the system.
[0143] In this embodiment, by inputting the object collection data into the pre-trained large model for trend prediction, the trend of the object to be predicted can be automatically predicted to obtain the object prediction trend. By inputting the object prediction trend into the pre-trained agent for model simulation, an object simulation model corresponding to the object to be predicted can be effectively generated. Based on the object simulation model, the risk of the object to be predicted can be effectively predicted to obtain the risk prediction result. By inputting the object to be predicted into the pre-trained retrieval-augmented generation model for retrieval generation, information of the object to be predicted can be automatically generated effectively. Based on the object-generated information and the risk prediction result, a strategy recommendation for the object to be predicted can be automatically generated, without the need to manually process data for the investment object, preventing the phenomenon of low data prediction accuracy caused by artificial subjective influence.
[0144] Embodiment III
[0145] Figure 4 It is a structural block diagram of a terminal device 2 provided in the third embodiment of the present application. As Figure 4 shown, the terminal device 2 of this embodiment includes: a processor 20, a memory 21, and a computer program 22 stored in the memory 21 and executable on the processor 20, such as a program for the data prediction method based on a large model. When the processor 20 executes the computer program 22, the steps in each of the above embodiments of the data prediction method based on a large model are implemented.
[0146] Exemplarily, the computer program 22 can be divided into one or more modules. The one or more modules are stored in the memory 21 and executed by the processor 20 to complete the present application. The one or more modules can be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program 22 in the terminal device 2. The terminal device may include, but is not limited to, a processor 20 and a memory 21.
[0147] The so-called processor 20 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0148] The memory 21 may be an internal storage unit of the terminal device 2, such as the hard disk or memory of the terminal device 2. The memory 21 may also be an external storage device of the terminal device 2, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the terminal device 2. Further, the memory 21 may also include both the internal storage unit and the external storage device of the terminal device 2. The memory 21 is used to store the computer program and other programs and data required by the terminal device. The memory 21 may also be used to temporarily store data that has been output or is to be output.
[0149] In addition, in each embodiment of the present application, the various functional modules may be integrated in one processing unit, or each unit may exist physically alone, or two or more units may be integrated in one unit. The above integrated unit may be implemented in the form of hardware or in the form of a software functional unit.
[0150] When an integrated module is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Among them, the computer-readable storage medium can be non-volatile or volatile. Based on such an understanding, to implement all or part of the processes in the above-mentioned embodiment methods of this application, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable storage medium can include: any entity or device that can carry computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable storage medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable storage medium does not include electrical carrier signals and telecommunication signals.
[0151] The above-described embodiments are only used to illustrate the technical solutions of this application, rather than to limit it; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of this application, and should all be included in the protection scope of this application.
Claims
1. A data prediction method based on a large model, characterized in that: The method comprises: Obtaining object acquisition data of an object to be predicted, and inputting the object acquisition data into a pre-trained large model for trend prediction to obtain an object prediction trend; Inputting the object prediction trend into the pre-trained intelligent agent for model simulation to obtain an object simulation model, and performing risk prediction on the object to be predicted based on the object simulation model to obtain a risk prediction result; Inputting the object to be predicted into the pre-trained retrieval enhancement generation model for retrieval generation to obtain object generation information; Generate a strategy recommendation for the object to be predicted based on the object generation information and the risk prediction result.
2. The data prediction method based on a large model as claimed in claim 1, characterized in that: Before inputting the object collection data into the pre-trained large model for trend prediction and obtaining the object prediction trend, the method further includes: Obtaining object sample data, and inputting the object sample data into the large model to extract key data, thereby obtaining key text, key image and key audio; Extracting features of the key text, the key image and the key audio to obtain text features, image features and audio features, and fusing the text features, the image features and the audio features to obtain sample fusion features; Perform trend prediction according to the sample fusion feature to obtain a sample prediction trend, and determine a model loss according to the sample prediction trend, the key text, the key image, the key audio, the text feature, the image feature, the audio feature, and the sample fusion feature; The parameters of the large model are updated according to the model loss until the large model converges to obtain the pre-trained large model.
3. The data prediction method based on a large model as claimed in claim 2, characterized in that: The object sample data is input into the large model to extract key data, and key text, key image and key audio are obtained, including: Acquire the number of reposts and comments on the text information, image information and audio information in the object sample data according to the large model, and determine the information popularity according to the number of reposts and the number of comments; The text information, the image information and the audio information are heat sorted according to the information heat, and key data of the text information, the image information and the audio information are extracted according to the heat sorting result to obtain the key text, the key image and the key audio.
4. The data prediction method based on a large model as claimed in claim 1, characterized in that: Before inputting the object prediction trend into the pre-trained intelligent agent for model simulation to obtain the object simulation model, the method further includes: Obtaining an object trend sample, and inputting the object trend sample into the agent for model simulation to obtain a sample simulation model; Performing risk prediction on the sample objects corresponding to the object trend samples according to the sample simulation model to obtain sample prediction results, and determining learning loss according to the sample prediction results and the sample simulation model; The parameters of the agent are updated according to the learning loss until convergence, thereby obtaining the pre-trained agent.
5. The data prediction method based on a large model as claimed in claim 1, characterized in that: The object to be predicted is input into the pre-trained retrieval enhancement generation model for retrieval generation, and before obtaining the object generation information, the method further includes: According to the retrieval enhancement generation model, information query is performed on the sample object in a preset database to obtain sample query information, and information generation is performed on the sample query information to obtain sample generation information; The generation loss is determined according to the sample generation information, and the parameters of the retrieval enhancement generation model are updated according to the generation loss until convergence, so as to obtain the pre-trained retrieval enhancement generation model.
6. The data prediction method based on a large model as claimed in claim 1, characterized in that: Obtain object collection data of the object to be predicted, including: Acquire the object address of the object to be predicted, and acquire policy information according to the object address to obtain local policy information; Obtaining the object name of the object to be predicted, and acquiring news information, litigation information, and securities information according to the object name to obtain object news information, object litigation information, and object securities information; The object news information, the object litigation information, the object securities information and the local policy information are combined to obtain the object collection data.
7. The data prediction method based on a large model as claimed in claim 2, characterized in that: Determining a model loss according to the sample prediction trend, the key text, the key image, the key audio, the text feature, the image feature, the audio feature, and the sample fusion feature includes: Calculating the trend similarity between the sample prediction trend and the standard prediction trend, and determining a first loss according to the trend similarity; Calculating the text similarity between the key text and the standard text, and determining the second loss according to the text similarity; Calculating an image similarity between the key image and the standard image, and determining a third loss according to the image similarity; Calculating an audio similarity between the key audio and the standard audio, and determining a fourth loss according to the audio similarity; Calculating a feature similarity between the text feature and a first standard feature to obtain a first feature similarity, and determining a fifth loss according to the first feature similarity; Calculating a feature similarity between the image feature and a second standard feature to obtain a second feature similarity, and determining a sixth loss according to the second feature similarity; Calculating a feature similarity between the audio feature and a third standard feature to obtain a third feature similarity, and determining a seventh loss according to the third feature similarity; Calculating feature similarity between the sample fusion feature and the standard fusion feature to obtain a fourth feature similarity, and determining an eighth loss according to the fourth feature similarity; The first loss, the second loss, the third loss, the fourth loss, the fifth loss, the sixth loss, the seventh loss and the eighth loss are weighted to obtain the model loss.
8. A data prediction system based on a large model, characterized in that: The system comprises: A trend prediction module, used to obtain object acquisition data of an object to be predicted, and input the object acquisition data into a pre-trained large model for trend prediction to obtain an object prediction trend; A risk prediction module is used to input the object prediction trend into the pre-trained intelligent agent for model simulation to obtain an object simulation model, and perform risk prediction on the object to be predicted based on the object simulation model to obtain a risk prediction result; A retrieval generation module, used for inputting the object to be predicted into the pre-trained retrieval enhancement generation model for retrieval generation to obtain object generation information; A suggestion generation module is used to generate a strategy suggestion for the object to be predicted based on the object generation information and the risk prediction result.
9. The data prediction system based on a large model as claimed in claim 8, characterized in that: The trend prediction module is also used for: Obtaining object sample data, and inputting the object sample data into the large model to extract key data, thereby obtaining key text, key image and key audio; Extracting features of the key text, the key image and the key audio to obtain text features, image features and audio features, and fusing the text features, the image features and the audio features to obtain sample fusion features; Perform trend prediction according to the sample fusion feature to obtain a sample prediction trend, and determine a model loss according to the sample prediction trend, the key text, the key image, the key audio, the text feature, the image feature, the audio feature, and the sample fusion feature; The parameters of the large model are updated according to the model loss until the large model converges to obtain the pre-trained large model.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.