Security field large model construction method and system

By building a large model in the securities field and adding compliance detection at its input and output ends, the problem of insufficient timeliness, professionalism and compliance in the application of general large models in the securities field is solved, and a more efficient and safer application of large models in the securities field is achieved.

CN120030278APending Publication Date: 2025-05-23SHANGHAI JIUFANGYUN INTELLIGENT TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510077441.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The application of existing general models in the securities field has problems of insufficient timeliness, professionalism and accuracy, and the securities industry has high requirements for data compliance and security.

Method used

By acquiring and preprocessing securities industry data, building a data compliance detection model, using compliance data for fine-tuning training of model instructions, building a professional large-scale language model focusing on the securities field, and adding compliance service detection at both the model input and output.

Benefits of technology

It realizes the professionalism, accuracy and timeliness of large models in the securities field, ensures the safety and compliance of model output, meets the complex financial market needs of securities practitioners and investors, and improves investment efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030278A_ABST
    Figure CN120030278A_ABST
Patent Text Reader

Abstract

The invention provides a security field large model construction method and system, and the method comprises the steps: S1, obtaining original security industry data, and carrying out the preprocessing of the obtained original security industry data; s2, constructing a data compliance detection model, and performing compliance detection on the preprocessed original security industry data by using the data compliance detection model to obtain compliance data; s3, building a model instruction fine tuning data set based on the obtained compliance data; and S4, performing supervised fine tuning training on the base large model by using the model instruction fine tuning data set to obtain a trained security field large model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a method and system for building a large model in the securities field. Background Art

[0002] In recent years, with the continuous development of financial markets and the rapid progress of information technology, investors and financial practitioners need more intelligent and convenient tools and methods to better understand market dynamics, assess risks and discover investment opportunities. Traditional securities investment advisory services can no longer meet the needs of the majority of stock investors. At present, artificial intelligence based on large models has demonstrated exciting new applications in many fields. However, although general large models perform well in general field tasks, their application in the financial securities investment industry has problems with timeliness, professionalism and accuracy. Therefore, the complexity and professionalism of the securities industry make it urgent for the industry to build a professional large language model specializing in the field of securities investment to make up for the shortcomings of the application of general large models in the securities field, provide more accurate, reliable and real-time securities decision support, meet the needs of securities practitioners and investors in complex financial markets, and improve investment efficiency. At the same time, the securities industry is a highly regulated industry, and the compliance and security of data have also become urgent problems to be solved by large models in the securities field.

[0003] Patent document CN117808596A (application number: 202410113285.3) discloses a large language model multi-agent system based on long short-term memory modules in the securities and futures industry. The system includes: a first agent module analyzes a current task based on the first information obtained to obtain a first analysis result; a second agent module analyzes the current task based on the second information obtained to obtain a second analysis result; the first agent module and the second agent module, based on the memory data obtained from the memory module, conduct a game on the first analysis result and the second analysis result obtained after analyzing the same task to obtain a game result; the third agent module performs adjustment analysis based on the game result to obtain a decision result. Summary of the invention

[0004] In view of the defects in the prior art, the purpose of the present invention is to provide a method and system for building a large model in the securities field.

[0005] A method for constructing a large model in the securities field provided by the present invention includes:

[0006] Step S1: obtaining original securities industry data, and preprocessing the obtained original securities industry data;

[0007] Step S2: construct a data compliance detection model, and use the data compliance detection model to perform compliance detection on the pre-processed original securities industry data to obtain compliance data;

[0008] Step S3: constructing a model instruction fine-tuning dataset based on the obtained compliance data;

[0009] Step S4: Use the model instruction fine-tuning data set to perform supervised fine-tuning training on the base large model to obtain a trained securities field large model.

[0010] Preferably, the step S1 comprises:

[0011] Step S1.1: Obtain original securities industry data, including: financial forum information, research reports, product / theme industry chain data, company financial report data, financial information announcements, industry user comments, and preset indicator data of financial research institutes;

[0012] Step S1.2: The obtained original securities industry data is cleaned, filtered and deduplicated to obtain the pre-processed original securities industry data.

[0013] Preferably, step S2 comprises:

[0014] Step S2.1: construct a compliance detection sensitive word library, and based on the constructed compliance detection sensitive word library, use the sensitive word filtering DFA algorithm to filter sensitive words on the pre-processed original securities industry data to obtain data containing sensitive words;

[0015] Step S2.2: Build a data compliance detection model based on mDeberta; use the data compliance detection model to perform compliance detection and screening on data containing sensitive words to obtain compliant data.

[0016] Preferably, the step S3 comprises:

[0017] Step S3.1: Set up model training tasks, including: securities industry-related questions and answers, user aspect-level / sentence-level sentiment analysis, research report opinion generation, financial report data interpretation, and domain-specific tasks for listed company questions and answers;

[0018] Step S3.2: according to different model training tasks, set instructions to fine-tune the formatting instance;

[0019] Step S3.3: Fine-tune the formatted instances according to the instructions of different tasks, and build the corresponding securities industry multi-task fine-tuning dataset based on the obtained compliance data.

[0020] Preferably, step S4 includes: selecting the chatglm series or the Baichuan series as the base large model, fine-tuning the data set using model instructions, and using the parameter fine-tuning technology LORA to perform supervised fine-tuning training of the large model.

[0021] Preferably, step S4 comprises:

[0022] Step S4.1: Initialize the pre-trained base model and freeze the bottom transformer layer;

[0023] Step S4.2: Update the parameters of the pre-trained base model. The formula is as follows:

[0024] W 0 +ΔW(1)

[0025] Among them, W 0 is the parameter initialized by the pre-trained base model, and ΔW is the parameter that needs to be updated;

[0026] Step S4.3: Assume that the matrix of the pre-trained base model is W 0 ∈R d×k , whose update is expressed as a low-rank decomposition:

[0027] W 0 +ΔW=W 0 +BA,B∈R d×r ,A∈R r×k (2)

[0028] Among them, the rank r<<min(d,k); R represents the weight matrix, including different d dimensions, r dimensions, and k dimensions;

[0029] Step S4.4: During the forward pass, W 0 and ΔW are multiplied by the same input x and finally added together:

[0030] h=W 0 x+ΔWx=W 0 x+BAx (3)

[0031] Step S4.5: During the training process, W 0 It remains unchanged and does not participate in gradient updates. It only trains parameter matrices A and B to obtain model update parameters ΔW.

[0032] Preferably, the method further comprises:

[0033] The compliance data is segmented to obtain multiple text segments; the text segments are vectorized through a text vectorization model; and the text vectors are stored in a vector database as a knowledge reserve for the securities industry in a large model.

[0034] Preferably, the method further comprises: performing compliance detection on the text content generated by the trained securities field big model through a data compliance detection model, thereby ensuring the compliance of the output content of the big model.

[0035] A securities field large model construction system provided by the present invention includes:

[0036] Module M1: obtain original securities industry data and pre-process the obtained original securities industry data;

[0037] Module M2: Build a data compliance detection model, use the data compliance detection model to perform compliance detection on the pre-processed original securities industry data, and obtain compliance data;

[0038] Module M3: Construct model instructions and fine-tune the dataset based on the obtained compliance data;

[0039] Module M4: Use the model instruction fine-tuning dataset to perform supervised fine-tuning training on the base large model to obtain the trained securities field large model.

[0040] Preferably, the module M1 comprises:

[0041] Module M1.1: Obtain original securities industry data, including: financial forum information, research reports, product / theme industry chain data, company financial report data, financial information announcements, industry user comments, and preset indicator data of financial research institutes;

[0042] Module M1.2: Perform cleaning, filtering and deduplication processing on the acquired original securities industry data to obtain the pre-processed original securities industry data;

[0043] The module M2 comprises:

[0044] Module M2.1: Build a compliance detection sensitive word library. Based on the built compliance detection sensitive word library, use the sensitive word filtering DFA algorithm to filter sensitive words on the pre-processed original securities industry data to obtain data containing sensitive words;

[0045] Module M2.2: Build a data compliance detection model based on mDeberta; use the data compliance detection model to perform compliance detection and screening on data containing sensitive words to obtain compliant data;

[0046] The module M3 comprises:

[0047] Module M3.1: Setting up model training tasks, including: securities industry-related Q&A, user aspect-level / sentence-level sentiment analysis, research report opinion generation, financial report data interpretation, and domain-specific tasks for listed company Q&A;

[0048] Module M3.2: Set instructions to fine-tune formatting instances according to different model training tasks;

[0049] Module M3.3: Fine-tune formatting instances according to the instructions of different tasks, and build corresponding securities industry multi-task fine-tuning datasets based on the acquired compliance data;

[0050] The module M4 includes: selecting the chatglm series or the Baichuan series as the base large model, using the model instructions to fine-tune the data set, and using the parameter fine-tuning technology LORA to perform supervised fine-tuning training of the large model.

[0051] Compared with the prior art, the present invention has the following beneficial effects:

[0052] 1. The present invention is based on cutting-edge artificial intelligence and deep learning technologies, adopts high-quality industry-specific data + supervised fine-tuning training methods, embeds large model knowledge memory components, takes into account the balance of training effect, efficiency and cost, builds a securities field large model, effectively injects securities field knowledge into the model, makes up for the lack of adaptability of general large models in securities field knowledge, and gives the large model field content generation professionalism, accuracy and timeliness;

[0053] 2. The present invention adds compliance service detection at both the input and output ends of the large model to effectively ensure the security compliance of the output content of the large model;

[0054] 3. Big models in the securities field will bring the full potential of artificial intelligence based on big models to the securities field, providing underlying AI capability support for common securities business scenarios such as investment consulting, customer service, investment research, risk control, compliance, and marketing;

[0055] 4. The securities field model constructed by the present invention based on the training method of continued pre-training in the securities field, supervised fine-tuning and human feedback reinforcement learning can complete the real-time and professional financial query function and financial logic analysis function in the financial field; the stock query function adopts the text2sql implementation of the domain model, and the stock logic analysis function realizes the domain model user thinking chain intention understanding, real-time tool library call, and financial response generation. Finally, a real-time human-computer interaction experience can be achieved in securities and financial scenarios such as comprehensive stock analysis, comprehensive market analysis, and comprehensive sector analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Other features, objects and advantages of the present invention will become more apparent from the detailed description of non-limiting embodiments made with reference to the following drawings:

[0057] Figure 1 Flowchart of the methodology for building large models in the securities domain. DETAILED DESCRIPTION

[0058] The present invention is described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those of ordinary skill in the art, several changes and improvements can also be made without departing from the concept of the present invention. These all belong to the protection scope of the present invention.

[0059] Example 1

[0060] According to a method for constructing a large model in the securities field provided by the present invention, as described above Figure 1 As shown, including:

[0061] Step 1: Obtain the original securities industry data and clean the obtained original securities industry data;

[0062] Specifically, the step 1 includes

[0063] Step 1.1: Collect and obtain data from various business lines in the securities industry, mainly including financial forum information, research reports, product / theme industry chain data, company financial report data, financial information announcements, industry user comments, and financial research institute characteristic indicator data.

[0064] Step 1.2: Perform pre-processing operations such as cleaning, filtering, and deduplication on the collected data.

[0065] Step 2: The data passes the model compliance detection peripheral system and performs compliance model service detection. If it passes, the instruction fine-tuning data set construction is performed.

[0066] Specifically, the step 2 includes:

[0067] Step 2.1: Build a model compliance detection peripheral system;

[0068] Specifically, the step 2.1 includes:

[0069] Step 2.1.1: Obtain training compliance data;

[0070] Step 2.1.2: Clean and label the training compliance data and build a compliance model training data set;

[0071] Step 2.1.3: Compliance experts build and maintain a set of compliance detection sensitive word libraries;

[0072] Step 2.1.4: Construct a sensitive word filtering DFA algorithm. The data submitted for inspection will be filtered for sensitive words first. Data containing sensitive words will enter the next step for model judgment.

[0073] Step 2.1.5: Use mDeberta as a pre-trained model to build a compliance detection model.

[0074] Step 2.2: The data is input into the model compliance detection peripheral system for compliance detection and screening.

[0075] Step 3: Construct a dataset for model instruction fine-tuning, aiming to align the content generation of the large model with the professional content of the securities industry;

[0076] Specifically, the said Step 3 includes:

[0077] Step 3.1: Domain experts design model training tasks. Such as: several domain tasks related to the securities industry, such as Q&A related to the securities industry, user aspect / sentence-level sentiment analysis, research report view generation, financial report data interpretation, listed company Q&A, etc.

[0078] Step 3.2: For different tasks, design instruction fine-tuning formatting instances;

[0079] For example: Q&A related to the securities industry

[0080] Instruction: "Please answer the following Q&A related to the securities industry knowledge"

[0081] Input: How to understand the stock market?

[0082] Output: The stock market refers to the place where stocks are issued and circulated, or it can also be said to be the place where the issued stocks are bought and transferred. The stock market can be divided into the primary market and the secondary market. The primary market is the market for raising funds through the issuance of stocks, and the secondary market is the market where stocks are traded on the stock exchange after they have been issued. The Shanghai Stock Exchange market is a typical representative of the Chinese stock market.

[0083] For example: Q&A about the general situation of listed companies

[0084] Instruction: Please answer the following company profile

[0085] Input: What is the main business of Pioneer Electronics?

[0086] Output: The main business of Pioneer Electronics is to provide intelligent, safe and efficient solutions and services for the urban gas industry.

[0087] For example: User sentence-level sentiment analysis

[0088] Instruction: Please judge the user's sentiment tendency (positive, negative, neutral) and the reason according to the following content, and return it in the following json format;

[0089] Input: I'm all green today;

[0090] Output: Negative, reason: In the field of stock investment, green represents a stock price decline, so it indicates that the stocks purchased by the user may be falling, resulting in losses in securities investment;

[0091] Example: User aspect-level sentiment analysis

[0092] Instructions: Please judge the user's emotional tendencies (positive, negative, neutral) and reasons in different aspects based on the following content, and return them in the following JSON format;

[0093] Input: I would like to take this opportunity to thank my account manager for being patient, responsible, and delivering information in a timely manner. She is my image and pride.

[0094] Output: Category: Service Evaluation, Sentiment: Positive, Reason: Account Manager Wang Jiachen was praised by customers for being patient, meticulous, serious and responsible, which shows that customers have a positive impression of her service.

[0095] For example: Research report opinion generation

[0096] Instructions: Extract the core ideas from the research report and list them in order.

[0097] Input: The current situation and development of corporate financial management. Corporate financial management products are oriented to institutional clients, and have the characteristics of diversified investment portfolios, risk and return balance, and professional financial management service needs. There are differences between corporate and private financial management in terms of scale, development, product characteristics, and investment strategies. Private financial management has a large scale and a long development time, while corporate financial management has a small scale and a short development time. Corporate financial management products face problems such as market uncertainty, low yields, and lack of flexibility. However, a comparative analysis of corporate financial management at home and abroad shows that China's corporate financial management market is expected to usher in a leapfrog development, but there is still a gap with mature markets.

[0098] Output: 1. The corporate wealth management market is in its early stages of development, with a small scale and a low proportion of institutional clients. 2. Fixed-income products account for a high proportion of corporate wealth management products, with a low risk level, mainly private placements, a high proportion of closed-end products, and a lower benchmark interest rate than private wealth management products.

[0099] For example: Interpretation of financial report data

[0100] Instructions: Please help me interpret the above financial data. Through the data, I can analyze the reasons behind it.

[0101] Input: Pudong Development Bank: The net profit in the 2022 annual report was 51.997 billion yuan, the net profit in the 2021 annual report was 53.766 billion yuan, the net profit in the 2020 annual report was 58.993 billion yuan, the net profit in the 2019 annual report was 59.506 billion yuan, and the net profit in the 2018 annual report was 56.515 billion yuan; the operating income in the 2022 annual report was 188.622 billion yuan, the operating income in the 2021 annual report was 190.982 billion yuan, the operating income in the 2020 annual report was 196.384 billion yuan, the operating income in the 2019 annual report was 190.688 billion yuan, and the operating income in the 2018 annual report was 170.865 billion yuan.

[0102] Output: Pudong Development Bank's net profit reached its highest point in 2020, and then declined in the following two years; you can also observe that Pudong Development Bank's operating income showed an increasing trend from 2018 to 2020, and then declined slightly in 2021. This may be related to the transformation of banking business and intensified competition, as banks need to continue to innovate and improve efficiency to remain competitive. In general, Pudong Development Bank's financial data reflects the challenges and opportunities facing the current banking industry.

[0103] Step 3.3: Fine-tune the formatted instances according to the instructions of different tasks, and build the corresponding securities industry multi-task fine-tuning dataset based on the obtained compliance data.

[0104] Specifically, the step 3.3 includes:

[0105] Step 3.3.1: Format the existing dataset.

[0106] Step 3.3.2: API collection + manual annotation.

[0107] Step 3.3.3: Tokenize the multi-task fine-tuning dataset to form a training dataset that can be directly used for model training.

[0108] Step 4: Evaluate based on usage scenarios and effects, select open source base large models such as the chatglm series and Baichuan series, and use the parameter efficient fine-tuning technology LORA for supervised fine-tuning training of large models.

[0109] Specifically, step 4 includes:

[0110] Step 4.1: Initialize the pre-trained base model and freeze the underlying transformer layers.

[0111] Step 4.2: Large model training requires updating the pre-trained model parameters. The formula is as follows:

[0112] W 0 +ΔW(1)

[0113] Among them, W 0 is the parameter initialized by the pre-trained model, and ΔW is the parameter that needs to be updated.

[0114] Step 4.3: Assume that the matrix of the pre-trained model is W 0 ∈R d×k , its update can be expressed as a low-rank decomposition:

[0115] W 0 +ΔW=W 0 +BA,B∈R d×r ,A∈R r×k (2)

[0116] Among them, the rank r<<min(d,k); R represents the weight matrix, including different d dimensions, r dimensions, and k dimensions,

[0117] Step 4.4: During the forward pass, W 0 and ΔW are multiplied by the same input x and finally added together:

[0118] h=W 0 x+ΔWx=W 0 x+BAx (3)

[0119] Step 4.5: During model training, W 0 It remains unchanged and does not participate in gradient updates. Only parameter matrices A and B are trained to obtain the model update parameter ΔW.

[0120] Step 5: Supervised fine-tuning training is completed to obtain a large model in the securities field.

[0121] Step 6: Add large model knowledge memory component.

[0122] Step 6.1: Divide the industry data obtained through the compliance testing service in step 2 into several text segments.

[0123] Step 6.2: The text segment is vectorized by a text vectorization model; in this embodiment, the text vectorization model is a bge vectorization model;

[0124] Step 6.3: Store the text vector in the vector database as a knowledge reserve of the securities industry of the big model to strengthen the knowledge memory ability of the big model. Provide corresponding reference knowledge for user questions, so that the content generated by the big model is more accurate, reliable and professional.

[0125] Step 7: After the big model generates content, it is also necessary to connect a layer of model compliance detection peripheral system in step 2 to further filter out the big model’s risky answers such as negative, unethical, illegal, and irregular, to ensure the compliance of the big model’s output content.

[0126] The present invention also provides a securities field big model construction system, which can be implemented by executing the process steps of the securities field big model construction method, that is, those skilled in the art can understand the securities field big model construction method as a preferred implementation of the securities field big model construction system.

[0127] Those skilled in the art know that, in addition to realizing the system and its various devices, modules, and units provided by the present invention in a purely computer-readable program code, it is entirely possible to realize the same functions in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, the system and its various devices, modules, and units provided by the present invention can be considered as a hardware component, and the devices, modules, and units included therein for realizing various functions can also be regarded as structures within the hardware component; the devices, modules, and units for realizing various functions can also be regarded as both software modules for realizing the method and structures within the hardware component.

[0128] The above describes the specific embodiments of the present invention. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which does not affect the essence of the present invention. In the absence of conflict, the embodiments of the present application and the features in the embodiments can be combined with each other arbitrarily.

Claims

1. A method for constructing a large model in the securities field, characterized in that: include: Step S1: obtaining original securities industry data, and preprocessing the obtained original securities industry data; Step S2: construct a data compliance detection model, and use the data compliance detection model to perform compliance detection on the pre-processed original securities industry data to obtain compliance data; Step S3: constructing a model instruction fine-tuning dataset based on the obtained compliance data; Step S4: Use the model instruction fine-tuning data set to perform supervised fine-tuning training on the base large model to obtain a trained securities field large model.

2. The method for constructing a large securities model according to claim 1, characterized in that: The step S1 comprises: Step S1.1: Obtain original securities industry data, including: financial forum information, research reports, product / theme industry chain data, company financial report data, financial information announcements, industry user comments, and preset indicator data of financial research institutes; Step S1.2: The obtained original securities industry data is cleaned, filtered and deduplicated to obtain the pre-processed original securities industry data.

3. The method for constructing a large securities model according to claim 1, characterized in that: The step S2 comprises: Step S2.1: construct a compliance detection sensitive word library, and based on the constructed compliance detection sensitive word library, use the sensitive word filtering DFA algorithm to filter sensitive words on the pre-processed original securities industry data to obtain data containing sensitive words; Step S2.2: Build a data compliance detection model based on mDeberta; use the data compliance detection model to perform compliance detection and screening on data containing sensitive words to obtain compliant data.

4. The method for constructing a large securities model according to claim 1, characterized in that: The step S3 comprises: Step S3.1: Set up model training tasks, including: securities industry-related questions and answers, user aspect-level / sentence-level sentiment analysis, research report opinion generation, financial report data interpretation, and domain-specific tasks for listed company questions and answers; Step S3.2: according to different model training tasks, set instructions to fine-tune the formatting instance; Step S3.3: Fine-tune the formatted instances according to the instructions of different tasks, and build the corresponding securities industry multi-task fine-tuning dataset based on the obtained compliance data.

5. The method for constructing a large model in the securities field according to claim 1, characterized in that: The step S4 includes: selecting the chatglm series or the Baichuan series as the base large model, fine-tuning the data set using model instructions, and using the parameter fine-tuning technology LORA to perform supervised fine-tuning training of the large model.

6. The method for constructing a large securities model according to claim 1, characterized in that: The step S4 comprises: Step S4.1: Initialize the pre-trained base model and freeze the bottom transformer layer; Step S4.2: Update the parameters of the pre-trained base model. The formula is as follows: W0+ΔW (1) Among them, W0 is the parameter initialized by the pre-trained base large model, and ΔW is the parameter that needs to be updated; Step S4.3: Assume that the matrix of the pre-trained base model is W0∈R d×k , whose update is expressed as a low-rank decomposition: W0+ΔW=W0+BA,B∈R d×r ,A∈R r×k (2) Among them, the rank r<<min(d,k); R represents the weight matrix, including different d dimensions, r dimensions, and k dimensions; Step S4.4: During the forward pass, both W0 and ΔW are multiplied by the same input x and then added together: h=W0x+ΔWx=W0x+BAx (3) Step S4.5: During the training process, W0 remains fixed and does not participate in gradient updates. Only the parameter matrices A and B are trained to obtain the model update parameters ΔW.

7. The method for constructing a large securities model according to claim 1, characterized in that: The method further comprises: The compliance data is segmented to obtain multiple text segments; the text segments are vectorized through a text vectorization model; and the text vectors are stored in a vector database as a knowledge reserve for the securities industry in a large model.

8. The method for constructing a large securities model according to claim 1, characterized in that: The method also includes: performing compliance detection on the text content generated by the trained securities field big model through a data compliance detection model, thereby ensuring the compliance of the output content of the big model.

9. A securities field large model construction system, characterized in that: include: Module M1: obtain original securities industry data and pre-process the obtained original securities industry data; Module M2: Build a data compliance detection model, use the data compliance detection model to perform compliance detection on the pre-processed original securities industry data, and obtain compliance data; Module M3: Construct model instructions and fine-tune the dataset based on the obtained compliance data; Module M4: Use the model instruction fine-tuning dataset to perform supervised fine-tuning training on the base large model to obtain the trained securities field large model.

10. The securities field large model construction system according to claim 9, characterized in that: The module M1 comprises: Module M1.1: Obtain original securities industry data, including: financial forum information, research reports, product / theme industry chain data, company financial report data, financial information announcements, industry user comments, and preset indicator data of financial research institutes; Module M1.2: Perform cleaning, filtering and deduplication processing on the acquired original securities industry data to obtain the pre-processed original securities industry data; The module M2 comprises: Module M2.1: Build a compliance detection sensitive word library. Based on the built compliance detection sensitive word library, use the sensitive word filtering DFA algorithm to filter sensitive words on the pre-processed original securities industry data to obtain data containing sensitive words; Module M2.2: Build a data compliance detection model based on mDeberta; use the data compliance detection model to perform compliance detection and screening on data containing sensitive words to obtain compliant data; The module M3 comprises: Module M3.1: Setting up model training tasks, including: securities industry-related Q&A, user aspect-level / sentence-level sentiment analysis, research report opinion generation, financial report data interpretation, and domain-specific tasks for listed company Q&A; Module M3.2: Set instructions to fine-tune formatting instances according to different model training tasks; Module M3.3: Fine-tune formatting instances according to the instructions of different tasks, and build corresponding securities industry multi-task fine-tuning datasets based on the acquired compliance data; The module M4 includes: selecting the chatglm series or the Baichuan series as the base large model, using the model instructions to fine-tune the data set, and using the parameter fine-tuning technology LORA to perform supervised fine-tuning training of the large model.

Citation Information

Patent Citations

  • Large language model multi-Agent system based on long and short term memory module in security and futures industry

    CN117808596A