Intelligent Question Answering Implementation Method and Device for Securities Compliance

Through an intelligent question-and-answer system, answers are generated using the securities compliance big model and internal and external regulations knowledge base, solving the problem of inefficient compliance issues in the securities industry and achieving more efficient compliance decision-making support.

CN119782491BActive Publication Date: 2025-06-27HUAAN SECURITIES CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510280068.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-27
Estimated Expiration
2045-03-11

AI Technical Summary

Technical Problem

The securities industry faces a large number of complex regulatory policies and laws and regulations, which makes it difficult for practitioners to respond efficiently to compliance issues in their daily business. The traditional regulatory research team model is inefficient and difficult to meet rapidly changing business needs.

Method used

A smart question-and-answer implementation method for securities compliance is adopted to determine the securities compliance model based on the target data set and the initial general model, and an internal and external regulatory knowledge base is formed, and the model and knowledge base are used to generate accurate answers to securities compliance issues.

Benefits of technology

It improves the accuracy and efficiency of responding to securities compliance issues, ensures that business decisions and operations keep up with the latest securities regulatory trends, and reduces securities compliance risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119782491B_ABST
    Figure CN119782491B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and device for implementing intelligent question answering for securities compliance. The method includes: determining a securities compliance large model based on a target data set and a pre-determined initial general large model; constructing an internal and external regulations knowledge base based on a regulations data set; when receiving a securities compliance question, determining a securities compliance answer corresponding to the securities compliance question based on the securities compliance question, the internal and external regulations knowledge base, and the securities compliance large model. It can be seen that the present invention can determine a securities compliance answer based on the obtained securities compliance question, the internal and external regulations knowledge base constructed from the regulations data set, and the securities compliance large model determined from the target data set, which is beneficial to improving the accuracy and efficiency of dealing with securities compliance questions, and further can ensure that securities business decisions and operations closely follow the latest securities regulatory dynamics and reduce securities compliance risks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of securities compliance, and particularly to a method and device for realizing intelligent question answering for securities compliance. Background Art

[0002] As a highly regulated industry, the securities industry faces a large number of complex regulatory policies, laws and regulations. Under a strict regulatory framework, in the process of daily business operations, securities practitioners will inevitably encounter various legal compliance issues frequently. These issues cover a wide range, from the compliance judgment of business operations to the application of regulations in complex trading scenarios. A slight mistake may lead to violations. This urgently requires securities companies to have a perfect and efficient compliance question answering function to provide accurate and timely legal guidance for practitioners and ensure that business activities are always operating on the track of compliance.

[0003] Currently, to address compliance issues, securities companies usually form a dedicated regulatory research team to deeply interpret regulations and provide legal support and compliance question answering services for company employees. However, this traditional model has high labor costs, is prone to errors and has low efficiency, and it is difficult to meet the needs of the rapid changes in the securities business and the massive compliance consultations. Therefore, it is particularly important to provide an intelligent question answering implementation solution for securities compliance to improve the accuracy and efficiency of dealing with securities compliance issues. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a method and device for realizing intelligent question answering for securities compliance, which can improve the accuracy and efficiency of dealing with securities compliance issues.

[0005] To solve the above technical problem, in the first aspect of the present invention, a method for realizing intelligent question answering for securities compliance is disclosed, and the method includes:

[0006] Determine a securities compliance large model based on a target data set and a pre-determined initial general large model; the target data set includes one or more of a securities industry question and answer data set, a legal knowledge question and answer data set, and a general encyclopedia question and answer data set;

[0007] Build an internal and external regulations knowledge base based on a regulations data set; the regulations data set contains a number of regulations data, and the regulations data includes external regulations data or internal regulations data;

[0008] When receiving a securities compliance question, determine a securities compliance answer corresponding to the securities compliance question based on the securities compliance question, the internal and external regulations knowledge base, and the securities compliance large model.

[0009] As an alternative implementation, in the first aspect of the present invention, determining the securities compliance large model based on the target data set and the pre-determined initial general large model includes:

[0010] Clean the target data set to remove non-target data in the target data set, obtaining a cleaned data set corresponding to the target data set; the non-target data includes duplicate data, irrelevant data, and data whose data quality does not meet the preset quality conditions in the target data set;

[0011] Annotate the cleaned data set corresponding to the target data set, obtaining a fine-tuning data set corresponding to the target data set;

[0012] Adjust the pre-determined initial general large model based on the fine-tuning data set, obtaining the securities compliance large model;

[0013] And, before constructing the internal and external regulations knowledge base based on the regulations data set, the method further includes:

[0014] Obtain the initial regulations data set and perform a correction operation on the initial regulations data set, obtaining a regulations data set corresponding to the initial regulations data set; the correction operation includes removing redundant information operation, revising format error operation, and unifying term expression operation.

[0015] As an alternative implementation, in the first aspect of the present invention, before determining the securities compliance answer corresponding to the securities compliance problem based on the securities compliance problem, the internal and external regulations knowledge base, and the securities compliance large model, the method further includes:

[0016] Perform chunking processing on each regulations data included in the internal and external regulations knowledge base, obtaining a number of text chunks corresponding to each regulations data; for each text chunk of each regulations data, perform a vectorization operation on the text chunk, obtaining a text chunk vector corresponding to the text chunk; based on the text chunk vectors corresponding to all the text chunks, construct a knowledge vector library corresponding to the internal and external regulations knowledge base.

[0017] As an alternative implementation, in the first aspect of the present invention, determining the securities compliance answer corresponding to the securities compliance problem based on the securities compliance problem, the internal and external regulations knowledge base, and the securities compliance large model includes:

[0018] Perform a vectorization operation on the securities compliance problem, obtaining a problem vector corresponding to the securities compliance problem; for each text chunk vector in the knowledge vector library corresponding to the internal and external regulations knowledge base, perform a similarity calculation based on the text chunk vector and the problem vector, obtaining a correlation parameter corresponding to the text chunk vector;

[0019] Determine target regulatory data from the internal and external regulations knowledge base based on the correlation parameters corresponding to all text block vectors in the knowledge vector library corresponding to the internal and external regulations knowledge base;

[0020] Input all the target regulatory data into the securities compliance large model, and generate a securities compliance answer corresponding to the securities compliance question based on the securities compliance large model.

[0021] As an alternative implementation, in the first aspect of the present invention, for each text block vector in the knowledge vector library corresponding to the internal and external regulations knowledge base, performing a similarity calculation based on the text block vector and the question vector to obtain the correlation parameter corresponding to the text block vector, including:

[0022] For each text block vector in the knowledge vector library corresponding to the internal and external regulations knowledge base, perform a similarity calculation on the text block vector and the question vector based on the cosine similarity method and the text matching algorithm to obtain the correlation parameter corresponding to the text block vector; the cosine similarity method is used to determine the cosine distance between the text block vector and the question vector; the text matching algorithm is used to determine the distance between the securities compliance question corresponding to the question vector and the text block corresponding to each text block vector.

[0023] As an alternative implementation, in the first aspect of the present invention, the determining target regulatory data from the internal and external regulations knowledge base based on the correlation parameters corresponding to all text block vectors in the knowledge vector library corresponding to the internal and external regulations knowledge base includes:

[0024] Based on the correlation parameters corresponding to all text block vectors in the knowledge vector library corresponding to the internal and external regulations knowledge base, determine several target text block vectors corresponding to the question vector from all text block vectors in the knowledge vector library;

[0025] For each target text block vector, perform a text backtracking operation on the target text block vector based on the internal and external regulations knowledge base corresponding to the knowledge vector library to obtain the regulatory data corresponding to the target text block vector;

[0026] Determine the regulatory data corresponding to all the target text block vectors as candidate regulatory data;

[0027] Based on the target attributes of each candidate regulatory data, determine at least one target regulatory data from all the candidate regulatory data; the target attributes include at least one of the document integrity, document length of the candidate regulatory data, and the correlation parameter of the corresponding target text block vector.

[0028] As an alternative implementation, in the first aspect of the present invention, adjusting the pre-determined initial general large model based on the fine-tuning dataset to obtain a securities compliance large model includes:

[0029] Injecting a low-rank matrix into the self-attention layer of the pre-determined initial general large model, and adding a prompt vector to each layer of the initial general large model; the prompt vector is a vector related to securities compliance and risks;

[0030] Based on the fine-tuning dataset, the low-rank matrix in the self-attention layer of the initial general large model, and the prompt vectors in each layer, adjusting the initial general large model by means of gradient descent and backpropagation to obtain a securities compliance large model.

[0031] The second aspect of the present invention discloses an intelligent question-answering implementation device for securities compliance, and the device includes:

[0032] A first determination module, configured to determine a securities compliance large model based on a target dataset and a pre-determined initial general large model; the target dataset includes one or more of a securities industry Q&A dataset, a legal knowledge Q&A dataset, and a general encyclopedia Q&A dataset;

[0033] A construction module, configured to construct an internal and external regulations knowledge base based on a regulations dataset; the regulations dataset contains a number of regulations data, and the regulations data includes external regulations data or internal regulations data;

[0034] A second determination module, configured to, when receiving a securities compliance question, determine a securities compliance answer corresponding to the securities compliance question based on the securities compliance question, the internal and external regulations knowledge base, and the securities compliance large model.

[0035] As an alternative implementation, in the second aspect of the present invention, the first determination module determines a securities compliance large model based on a target dataset and a pre-determined initial general large model, and the specific method includes:

[0036] Cleaning the target dataset to remove non-target data in the target dataset to obtain a cleaned dataset corresponding to the target dataset; the non-target data includes duplicate data, irrelevant data, and data whose data quality does not meet the preset quality conditions in the target dataset;

[0037] Labeling the cleaned dataset corresponding to the target dataset to obtain a fine-tuning dataset corresponding to the target dataset;

[0038] Adjusting the pre-determined initial general large model based on the fine-tuning dataset to obtain a securities compliance large model;

[0039] Further, the apparatus further includes:

[0040] A correction module, configured to obtain an initial regulation dataset and perform a correction operation on the initial regulation dataset before the composition module composes an internal and external regulation knowledge base based on the regulation dataset, so as to obtain a regulation dataset corresponding to the initial regulation dataset; the correction operation includes an operation of removing redundant information, an operation of revising format errors, and an operation of unifying term expressions.

[0041] As an alternative implementation manner, in the second aspect of the present invention, the apparatus further includes:

[0042] A chunking module, configured to perform chunking processing on each regulation data included in the internal and external regulation knowledge base before the second determination module determines a securities compliance answer corresponding to the securities compliance problem based on the securities compliance problem, the internal and external regulation knowledge base, and the securities compliance large model, so as to obtain a plurality of text chunks corresponding to each regulation data; for each text chunk of each regulation data, perform a vectorization operation on the text chunk to obtain a text chunk vector corresponding to the text chunk; and construct a knowledge vector library corresponding to the internal and external regulation knowledge base based on the text chunk vectors corresponding to all the text chunks.

[0043] As an alternative implementation manner, in the second aspect of the present invention, the second determination module determines a securities compliance answer corresponding to the securities compliance problem based on the securities compliance problem, the internal and external regulation knowledge base, and the securities compliance large model. The specific manner includes:

[0044] Perform a vectorization operation on the securities compliance problem to obtain a problem vector corresponding to the securities compliance problem; for each text chunk vector in the knowledge vector library corresponding to the internal and external regulation knowledge base, perform a similarity calculation based on the text chunk vector and the problem vector to obtain a correlation parameter corresponding to the text chunk vector;

[0045] Based on the correlation parameters corresponding to all the text chunk vectors in the knowledge vector library corresponding to the internal and external regulation knowledge base, determine target regulation data from the internal and external regulation knowledge base;

[0046] Input all the target regulation data into the securities compliance large model, and generate a securities compliance answer corresponding to the securities compliance problem based on the securities compliance large model.

[0047] As an alternative implementation manner, in the second aspect of the present invention, for each text chunk vector in the knowledge vector library corresponding to the internal and external regulation knowledge base, the second determination module performs a similarity calculation based on the text chunk vector and the problem vector to obtain a correlation parameter corresponding to the text chunk vector. The specific manner includes:

[0048] For each text block vector in the knowledge vector library corresponding to the internal and external regulations knowledge library, similarity calculation is performed on the text block vector and the problem vector based on the cosine similarity method and the text matching algorithm to obtain the correlation parameter corresponding to the text block vector; the cosine similarity method is used to determine the cosine distance between the text block vector and the problem vector; the text matching algorithm is used to determine the distance between the securities compliance problem corresponding to the problem vector and each text block corresponding to the text block vector.

[0049] As an optional implementation manner, in the second aspect of the present invention, the second determination module determines target regulation data from the internal and external regulations knowledge library based on the correlation parameters corresponding to all text block vectors in the knowledge vector library corresponding to the internal and external regulations knowledge library. The specific manner includes:

[0050] Based on the correlation parameters corresponding to all text block vectors in the knowledge vector library corresponding to the internal and external regulations knowledge library, several target text block vectors corresponding to the problem vector are determined from all text block vectors in the knowledge vector library;

[0051] For each target text block vector, a text backtracking operation is performed on the target text block vector based on the internal and external regulations knowledge library corresponding to the knowledge vector library to obtain the regulation data corresponding to the target text block vector;

[0052] The regulation data corresponding to all the target text block vectors is determined as candidate regulation data;

[0053] Based on the target attributes of each candidate regulation data, at least one target regulation data is determined from all the candidate regulation data; the target attributes include at least one of the document integrity, document length of the candidate regulation data, and the correlation parameter of the corresponding target text block vector.

[0054] As an optional implementation manner, in the second aspect of the present invention, the first determination module adjusts a pre-determined initial general large model based on the fine-tuning data set to obtain a securities compliance large model. The specific manner includes:

[0055] Inject a low-rank matrix into the self-attention layer of the pre-determined initial general large model, and add a prompt vector to each layer of the initial general large model; the prompt vector is a vector related to securities compliance and risks;

[0056] Based on the fine-tuning data set, the low-rank matrix in the self-attention layer of the initial general large model, and the prompt vector in each layer, the initial general large model is adjusted by the gradient descent and backpropagation methods to obtain a securities compliance large model.

[0057] The third aspect of the present invention discloses an intelligent question-answering implementation system for securities compliance, and the system includes:

[0058] A memory storing executable program codes;

[0059] A processor coupled to the memory;

[0060] The processor calls the executable program codes stored in the memory and executes the steps in the intelligent question-answering implementation method for securities compliance disclosed in the first aspect of the present invention.

[0061] The fourth aspect of the present invention discloses a computer storage medium. The computer storage medium stores computer instructions, which are used to execute the steps in the intelligent question-answering implementation method for securities compliance disclosed in the first aspect of the present invention when called.

[0062] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:

[0063] Based on the target data set and the pre-determined initial general large model, the present invention determines a large model for securities compliance; constructs an internal and external regulations knowledge base based on the regulations data set; when receiving a securities compliance question, based on the securities compliance question, the internal and external regulations knowledge base, and the large model for securities compliance, determines a securities compliance answer corresponding to the securities compliance question. It can be seen that the present invention can determine a securities compliance answer based on the obtained securities compliance question, the internal and external regulations knowledge base constructed from the regulations data set, and the large model for securities compliance determined from the target data set, which is beneficial to improving the accuracy and efficiency of dealing with securities compliance questions, and further can ensure that securities business decisions and operations closely follow the latest securities regulatory dynamics and reduce securities compliance risks. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0065] Figure 1 is a schematic flowchart of an intelligent question-answering implementation method for securities compliance disclosed in an embodiment of the present invention;

[0066] Figure 2 is a schematic flowchart framework of an intelligent question-answering implementation method for securities compliance disclosed in an embodiment of the present invention;

[0067] Figure 3 is a schematic structural diagram of an intelligent question-answering implementation device for securities compliance disclosed in an embodiment of the present invention;

[0068] Figure 4 It is a schematic structural diagram of an intelligent question - answering implementation system for securities compliance disclosed in an embodiment of the present invention. Specific implementation manners

[0069] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the scope of protection of the present invention.

[0070] The terms "first", "second", etc. in the specification and claims of the present invention and the above - mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "include" and "have" and any variations thereof are intended to cover non - exclusive inclusion. For example, a process, method, system, product or terminal including a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or terminals.

[0071] Referring to "embodiment" herein means that a specific feature, structure or characteristic described in connection with the embodiment may be included in at least one embodiment of the present invention. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described herein may be combined with other embodiments.

[0072] The present invention discloses an intelligent question - answering implementation method and device for securities compliance. Implementing the method described in the embodiments of the present invention can determine securities compliance answers based on the securities compliance questions obtained, the internal and external regulations knowledge base formed by the regulations data set, and the securities compliance large model determined by the target data set, which is beneficial to improving the accuracy and efficiency of dealing with securities compliance questions, and further can ensure that securities business decisions and operations keep up with the latest securities regulatory dynamics and reduce securities compliance risks. The following will be described in detail respectively.

[0073] Embodiment 1

[0074] Please refer to Figure 1 , Figure 1 It is a schematic flowchart of an intelligent question - answering implementation method for securities compliance disclosed in an embodiment of the present invention. Among them, Figure 1 The described method can be applied to the intelligent question - answering scenario for securities compliance, and the embodiments of the present invention do not make limitations. AsFigure 1 As shown in the figure, the intelligent question - answering implementation method for securities compliance includes the following operations:

[0075] 101. Determine a securities compliance large - model based on a target data set and a pre - determined initial general large - model;

[0076] In the embodiments of the present invention, the target data set includes one or more of a securities industry Q&A data set, a legal knowledge Q&A data set, and a general encyclopedia Q&A data set;

[0077] 102. Build an internal and external regulations knowledge base based on a regulations data set;

[0078] In the embodiments of the present invention, the regulations data set contains a number of regulations data, and the regulations data includes external regulations data or internal regulations data;

[0079] 103. When receiving a securities compliance question, determine a securities compliance answer corresponding to the securities compliance question based on the securities compliance question, the internal and external regulations knowledge base, and the securities compliance large - model.

[0080] In the embodiments of the present invention, in order to ensure that the large - language model can accurately understand and generate answers that conform to the characteristics of the securities industry, a fine - tuning data set is carefully selected and constructed from multiple sources. Among them, the securities industry Q&A data set covers questions and answers in aspects such as securities business operations, market trading rules, and risk management; the legal knowledge Q&A data set contains extensive interpretations of legal provisions, case analyses, and application examples to help the large - model understand complex legal logics; the general encyclopedia Q&A data set provides basic knowledge background support to enhance the generalization ability of the large - model so that it can reason in a broader knowledge background.

[0081] In the embodiments of the present invention, the internal and external regulations knowledge base is collected and sorted according to the special needs of the securities industry. So far, 9510 external regulations data have been accumulated and integrated with 873 company - internal regulations data to form a comprehensive and authoritative internal and external regulations knowledge base.

[0082] In the actual operation of the embodiments of the present invention, it can assist the compliance department in efficiently answering employees' regulations questions, ensuring that business decisions and operations keep up with the latest regulatory dynamics. When facing changes in regulatory policies, the technical solution of the present invention can quickly adapt, effectively reducing compliance risks caused by misunderstandings or neglect of new regulations, and ensuring the compliance of the company's business. During the product R & D or business model innovation stage, the technical solution of the present invention can provide accurate compliance assessments to ensure that every step of innovation is on the track of legality and compliance.

[0083] It can be seen that the present invention can determine the securities compliance answer based on the internal and external regulations knowledge base formed by the obtained securities compliance issues and regulations data set and the securities compliance large model determined by the target data set, which is beneficial to improving the accuracy and efficiency of dealing with securities compliance issues, and further can ensure that the securities business decision-making and operation keep up with the latest securities supervision dynamics and reduce the securities compliance risk.

[0084] In an alternative embodiment, the step of determining the securities compliance large model based on the target data set and the pre-determined initial general large model in step 101 may include:

[0085] Clean the target data set to remove the non-target data in the target data set, and obtain the cleaned data set corresponding to the target data set; the non-target data includes duplicate data, irrelevant data, and data whose data quality does not meet the preset quality conditions in the target data set;

[0086] Label the cleaned data set corresponding to the target data set to obtain the fine-tuning data set corresponding to the target data set;

[0087] Adjust the pre-determined initial general large model based on the fine-tuning data set to obtain the securities compliance large model;

[0088] And, before the step of forming the internal and external regulations knowledge base based on the regulations data set in step 102, the method may further include:

[0089] Obtain the initial regulations data set and perform a correction operation on the initial regulations data set to obtain the regulations data set corresponding to the initial regulations data set; the correction operation includes removing redundant information operation, revising format error operation, and unifying term expression operation.

[0090] In this alternative embodiment, it can be understood that the target data set undergoes a strict cleaning process to remove duplicate items, irrelevant or low-quality data points to obtain the cleaned data set to ensure the purity of the data set. Subsequently, the data can be manually labeled by a team of domain experts to clarify information such as problem types, keywords, and expected answers to obtain the fine-tuning data set, which provides high-quality learning materials for the fine-tuning of the large language model.

[0091] In this optional embodiment, the regulatory dataset used to build the internal and external regulations knowledge base also undergoes a rigorous manual cleaning process, including removing redundant information, correcting formatting errors, unifying term expressions, etc., to ensure the consistency and readability of regulatory provisions. In this way, not only a detailed regulatory index is established, but also a solid foundation is laid for subsequent retrieval enhancement generation techniques. Finally, this internal and external regulations knowledge base will become the core resource for compliance Q&A, supporting users to query the most relevant regulatory articles and their explanations, while tracing the information sources to ensure the authority and accuracy of the answers. Through the fine processing of the fine-tuning dataset and the internal and external regulations knowledge base, strong data support can be provided for the compliance Q&A of securities companies.

[0092] It can be seen that this optional embodiment can clean the target dataset to remove non-target data to obtain a fine-tuning dataset, and adjust the initial general large model based on the fine-tuning dataset to obtain a securities compliance large model, which is beneficial to improving the determination accuracy of the securities compliance large model; it can also obtain the initial regulatory dataset and perform correction operations on the initial regulatory dataset to obtain a regulatory dataset for constructing the internal and external regulations knowledge base, which is beneficial to improving the accuracy of the regulatory dataset, and further improving the determination accuracy of the internal and external regulations knowledge base, thereby improving the response accuracy and efficiency for securities compliance issues.

[0093] In another optional embodiment, before determining the securities compliance answer corresponding to the securities compliance question based on the securities compliance question, the internal and external regulations knowledge base, and the securities compliance large model in step 103, the method may further include:

[0094] Perform chunking processing on each regulatory data included in the internal and external regulations knowledge base to obtain a number of text chunks corresponding to each regulatory data; for each text chunk of each regulatory data, perform a vectorization operation on the text chunk to obtain a text chunk vector corresponding to the text chunk; based on the text chunk vectors corresponding to all text chunks, construct a knowledge vector library corresponding to the internal and external regulations knowledge base.

[0095] In this optional embodiment, it can be understood that the chunking processing process divides a large amount of text data into smaller, manageable and retrievable units. This process can be achieved through natural language processing techniques, such as using methods of clause segmentation and paragraph segmentation, or document chunking techniques based on topic models. Each chunked knowledge unit will be converted into a vector and stored in the vector library. Further, this process can use text embedding techniques such as BERT and TF-IDF to convert the text into vectors in a high-dimensional space.

[0096] It can be seen that this optional embodiment can perform chunking processing on each piece of regulatory data included in the internal and external regulatory knowledge base to obtain the text chunks corresponding to each piece of regulatory data, and further perform vectorization operations on the text chunks to obtain the text chunk vectors corresponding to the text chunks. Building a knowledge vector library based on the text chunk vectors is beneficial to improving the construction accuracy of the knowledge vector library, thereby improving the determination accuracy of the securities compliance answer determined based on the internal and external regulatory knowledge base, and thus improving the response accuracy and efficiency for securities compliance issues.

[0097] In another optional embodiment, the step 103 of determining the securities compliance answer corresponding to the securities compliance question based on the securities compliance question, the internal and external regulatory knowledge base, and the securities compliance large model may include:

[0098] Perform a vectorization operation on the securities compliance question to obtain the question vector corresponding to the securities compliance question; for each text chunk vector in the knowledge vector library corresponding to the internal and external regulatory knowledge base, perform a similarity calculation based on the text chunk vector and the question vector to obtain the correlation parameter corresponding to the text chunk vector;

[0099] Based on the correlation parameters corresponding to all text chunk vectors in the knowledge vector library corresponding to the internal and external regulatory knowledge base, determine the target regulatory data from the internal and external regulatory knowledge base;

[0100] Input all the target regulatory data into the securities compliance large model, and generate the securities compliance answer corresponding to the securities compliance question based on the securities compliance large model.

[0101] In this optional embodiment, it can be understood that inputting the original text of the regulatory data to the securities compliance large model instead of the text chunks to the large model is to avoid problems with the knowledge quality fed to the large model due to the lack of text chunking strategies in traditional Retrieval-Augmented Generation (RAG), which may lead to ineffective answering of questions. By directly retrieving the relevant original text with the powerful recall and retrieval ability of the large model, this problem can be effectively solved. Finally, input these selected target regulatory data into the securities compliance large model, and the securities compliance large model will generate the final answer by combining the document content of these target regulatory data.

[0102] It can be seen that this optional embodiment can perform a vectorization operation on the securities compliance question to obtain the question vector, perform a similarity calculation between each text chunk vector and the question vector to obtain the correlation parameter, determine the target regulatory data based on the correlation parameter, and input all the target regulatory data into the securities compliance large model to generate the securities compliance answer, which is beneficial to the determination accuracy of the target regulatory data, thereby improving the generation accuracy of the securities compliance answer, and thus improving the response accuracy and efficiency for securities compliance issues.

[0103] In yet another alternative embodiment, for each text block vector in the knowledge vector library corresponding to the internal and external regulations knowledge base, performing a similarity calculation based on the text block vector and the question vector to obtain the correlation parameter corresponding to the text block vector may include:

[0104] For each text block vector in the knowledge vector library corresponding to the internal and external regulations knowledge base, performing a similarity calculation on the text block vector and the question vector based on the cosine similarity method and the text matching algorithm to obtain the correlation parameter corresponding to the text block vector; the cosine similarity method is used to determine the cosine distance between the text block vector and the question vector; the text matching algorithm is used to determine the distance between the securities compliance problem corresponding to the question vector and the text block corresponding to each text block vector.

[0105] In this alternative embodiment, the essence of the similarity calculation process is to quickly retrieve the most relevant target regulation data from the knowledge vector library when the user asks a question. The text matching algorithm can be the BM25 algorithm, and the similarity calculation can refer to the following formula to achieve this goal by combining the cosine similarity and the BM25 text matching algorithm similarity.

[0106]

[0107] In the above formula, M represents the number of vectors in the knowledge vector library of the knowledge base, w represents the weight, represents the cosine similarity calculation, represents the BM25 similarity calculation, represents the text block vector i , represents the question vector, represents the text block vector i corresponding to the text block, represents the text block corresponding to the securities compliance problem corresponding to the question vector, Similarity i represents the text block vector i and the similarity calculation result between the question vector. The cosine similarity is used to measure the cosine distance between vectors, and it can quickly find the vector closest to the question vector in the vector space. The BM25 text matching algorithm is a probabilistic ranking function algorithm that considers the frequency of the query term in the document and the document length, and can effectively calculate the distance between the question and the text block.

[0108] It can be seen that this alternative embodiment can complete the calculation based on the cosine similarity method and the text matching algorithm when performing the similarity calculation based on the text block vector and the question vector, which is beneficial to improving the calculation accuracy of the correlation parameter, and then improving the determination accuracy of the target regulation data, thereby improving the response accuracy and efficiency for securities compliance problems.

[0109] In yet another alternative embodiment, determining the target regulatory data from the internal and external regulations knowledge base based on the correlation parameters corresponding to all text block vectors in the knowledge vector base corresponding to the internal and external regulations knowledge base may include:

[0110] Determining a plurality of target text block vectors corresponding to the problem vector from all text block vectors in the knowledge vector base based on the correlation parameters corresponding to all text block vectors in the knowledge vector base corresponding to the internal and external regulations knowledge base;

[0111] For each target text block vector, performing a text backtracking operation on the target text block vector based on the internal and external regulations knowledge base corresponding to the knowledge vector base to obtain the regulatory data corresponding to the target text block vector;

[0112] Determining the regulatory data corresponding to all target text block vectors as candidate regulatory data;

[0113] Determining at least one target regulatory data from all candidate regulatory data based on the target attributes of each candidate regulatory data; the target attributes include at least one of the document integrity, document length of the candidate regulatory data, and the correlation parameter of the corresponding target text block vector.

[0114] In this alternative embodiment, according to the correlation parameters calculated by similarity, the K (TOP-K) most relevant target text block vectors (K is generally set to 10 in actual operation) can be retrieved from the knowledge vector base; further, based on the target attributes of each candidate regulatory data, at least one (TOP-N) target regulatory data (generally set to 4 in actual operation) can be determined from all candidate regulatory data for input into the securities compliance large model for generation.

[0115] In this alternative embodiment, the text backtracking operation involves mapping the regulatory data from the knowledge vector base to the internal and external regulations knowledge base, which can be implemented by maintaining a mapping table when constructing the knowledge vector base. In this way, the original content corresponding to each vector can be quickly located, providing the necessary context information for the generation stage.

[0116] Based on the above alternative embodiment, the process framework diagram of the present invention can be as follows Figure 2As shown, first, the fine-tuned large model is used as the base, and then based on the constructed internal and external regulations knowledge base, the retrieval-augmented generation (RAG) technical route is adopted to implement the compliance question-answering function. The core of the technical route lies in combining two stages: retrieval and generation. First, the knowledge most relevant to the question is retrieved from the large-scale internal and external regulations knowledge base, and then the large model generates an answer based on this. The goal of the present invention is to improve the understanding and answering ability of securities compliance issues through the efficient retrieval and utilization of the internal and external regulations knowledge base.

[0117] It can be seen that this optional embodiment can determine a number of target text block vectors corresponding to the question vector based on the correlation parameters corresponding to all text block vectors; perform a text backtracking operation on the target text block vectors based on the internal and external regulations knowledge base to obtain the regulatory data corresponding to the target text block vectors and determine it as the candidate regulatory data, and determine at least one target regulatory data based on the target attributes of the candidate regulatory data, which is beneficial to improving the determination accuracy of the target regulatory data, thereby improving the generation accuracy of the securities compliance answer, and thus improving the response accuracy and efficiency for securities compliance issues.

[0118] In another optional embodiment, adjusting the pre-determined initial general large model based on the fine-tuning dataset to obtain a securities compliance large model includes:

[0119] Injecting a low-rank matrix into the self-attention layer of the pre-determined initial general large model, and adding a prompt vector to each layer of the initial general large model; the prompt vector is a vector related to securities compliance and risks;

[0120] Based on the fine-tuning dataset, the low-rank matrix in the self-attention layer of the initial general large model, and the prompt vector in each layer, adjusting the initial general large model through the gradient descent and backpropagation methods to obtain a securities compliance large model.

[0121] In this alternative embodiment, in terms of technical implementation, injecting a low-rank matrix into the self-attention layer of a pre-determined initial general large model can be based on the LoRA (Low-Rank Adaptation) fine-tuning method, while adding a prompt vector to each layer of the initial general large model can be based on the Prompt-tuning v2 fine-tuning method. Based on this, on the basis of the constructed large model fine-tuning dataset, the technical solution of the present invention can integrate the two fine-tuning methods of LoRA and Prompt-tuning v2 to train the large model. LoRA is a parameter-efficient fine-tuning technique that completes fine-tuning by adding a low-rank matrix between the key layers of a pre-trained model. The core of this method is that it does not require re-training the entire model, but adjusts the behavior of the model by introducing additional trainable parameters, thereby reducing memory requirements and training time. This means that the model can be efficiently adjusted to adapt to specific securities compliance tasks without sacrificing model performance. Prompt-tuning v2 is an efficient method based on deep prompt tuning, which achieves universality for different tasks and model scales by adding consecutive prompts to each layer of the pre-trained model. The key feature of this method is that prompts are added to each layer of the model, enabling the model to respond more directly to task-specific requirements. This means that securities compliance-related prompts can be embedded in each layer of the model to further improve the model's understanding and prediction ability in these specific fields.

[0122] In this alternative embodiment, it can be further understood that the initial general large model serving as the fine-tuning base model can be Alibaba's Tongyi Qianwen 2.5 open-source large model, or the open-source DeepSeek-R1, or other open-source large models, which are not limited in the embodiments of the present invention. Based on the initial general large model, the LoRA method is used to inject a low-rank matrix into the self-attention layer of the model to simulate the effect of full-parameter fine-tuning. At the same time, the Prompt-tuning v2 method is applied to add prompt vectors related to securities compliance and risks to each layer of the model. The overall fine-tuning architecture is as Figure 2 shown. In the subsequent fine-tuning process, the compliance Q&A pair dataset will be used to learn the low-rank matrix in LoRA and the prompt vectors in Prompt-tuning v2 through an optimization algorithm combining gradient descent and backpropagation. Such fine-tuning training will enable the model to have better adaptability to the compliance Q&A scenarios of securities companies while maintaining the performance of the base large model. Through this method, not only can the answer accuracy of the model be improved, but also it can be ensured that the decision-making process of the model complies with the compliance requirements of the securities industry.

[0123] It can be seen that this optional embodiment can adjust the initial general large model to obtain a securities compliance large model based on a fine-tuning dataset, a low-rank matrix in the self-attention layer of the initial general large model, and prompt vectors in each layer through gradient descent and backpropagation methods, which is beneficial to improving the determination accuracy of the securities compliance large model, and further improving the generation accuracy of securities compliance answers generated based on the securities compliance large model, thereby improving the response accuracy and efficiency for securities compliance issues.

[0124] Embodiment 2

[0125] Please refer to Figure 3 , Figure 3 which is a schematic structural diagram of an intelligent question-answering implementation device for securities compliance disclosed in an embodiment of the present invention. Among them, Figure 3 the described device can be applied to any intelligent question-answering scenario for securities compliance, and the embodiments of the present invention do not make limitations. As Figure 3 shown, the intelligent question-answering implementation device for securities compliance may include:

[0126] A first determination module 201, configured to determine a securities compliance large model based on a target dataset and a pre-determined initial general large model; the target dataset includes one or more of a securities industry question-answering dataset, a legal knowledge question-answering dataset, and a general encyclopedia question-answering dataset;

[0127] A construction module 202, configured to construct an internal and external regulations knowledge base based on a regulations dataset; the regulations dataset contains a number of regulations data, and the regulations data includes external regulations data or internal regulations data;

[0128] A second determination module 203, configured to determine a securities compliance answer corresponding to a securities compliance question based on the securities compliance question, the internal and external regulations knowledge base, and the securities compliance large model when receiving the securities compliance question.

[0129] It can be seen that implementing the device described in the embodiments of the present invention can determine a securities compliance answer based on the obtained securities compliance question, the internal and external regulations knowledge base constructed from the regulations dataset, and the securities compliance large model determined from the target dataset, which is beneficial to improving the response accuracy and efficiency for securities compliance questions, and further can ensure that securities business decisions and operations keep up with the latest securities regulatory dynamics and reduce securities compliance risks.

[0130] In an optional embodiment, the first determination module 201 determines the securities compliance large model based on the target dataset and the pre-determined initial general large model, and the specific method includes:

[0131] Clean the target data set to remove non-target data in the target data set, and obtain the cleaned data set corresponding to the target data set; the non-target data includes duplicate data, irrelevant data, and data whose data quality does not meet the preset quality conditions in the target data set;

[0132] Annotate the cleaned data set corresponding to the target data set to obtain the fine-tuning data set corresponding to the target data set;

[0133] Adjust the pre-determined initial general large model based on the fine-tuning data set to obtain a securities compliance large model;

[0134] In addition, the device may further include:

[0135] A correction module, configured to obtain an initial regulation data set and perform a correction operation on the initial regulation data set before the composition module 202 composes an internal and external regulation knowledge base based on the regulation data set, so as to obtain a regulation data set corresponding to the initial regulation data set; the correction operation includes removing redundant information operation, revising format error operation, and unifying term expression operation.

[0136] It can be seen that this optional embodiment can clean the target data set to remove non-target data to obtain a fine-tuning data set, and adjust the initial general large model based on the fine-tuning data set to obtain a securities compliance large model, which is beneficial to improving the determination accuracy of the securities compliance large model; it can also obtain an initial regulation data set and perform a correction operation on the initial regulation data set to obtain a regulation data set for composing an internal and external regulation knowledge base, which is beneficial to improving the accuracy of the regulation data set, and further improving the determination accuracy of the internal and external regulation knowledge base, thereby improving the response accuracy and efficiency for securities compliance issues.

[0137] In another optional embodiment, the device may further include:

[0138] A chunking module, configured to perform chunking processing on each regulation data included in the internal and external regulation knowledge base before the second determination module 203 determines a securities compliance answer corresponding to the securities compliance issue based on the securities compliance issue, the internal and external regulation knowledge base, and the securities compliance large model, so as to obtain a plurality of text chunks corresponding to each regulation data; for each text chunk of each regulation data, perform a vectorization operation on the text chunk to obtain a text chunk vector corresponding to the text chunk; based on the text chunk vectors corresponding to all text chunks, construct a knowledge vector library corresponding to the internal and external regulation knowledge base.

[0139] It can be seen that this optional embodiment can perform chunking processing on each piece of regulatory data included in the internal and external regulatory knowledge base to obtain the text chunks corresponding to each piece of regulatory data, and further perform vectorization operations on the text chunks to obtain the text chunk vectors corresponding to the text chunks. Building a knowledge vector base based on the text chunk vectors is beneficial to improving the construction accuracy of the knowledge vector base, and further improving the determination accuracy of the securities compliance answer determined based on the internal and external regulatory knowledge base, thereby improving the response accuracy and efficiency for securities compliance issues.

[0140] In yet another optional embodiment, the second determination module 203 determines the securities compliance answer corresponding to the securities compliance question based on the securities compliance question, the internal and external regulatory knowledge base, and the securities compliance large model. The specific method includes:

[0141] Perform a vectorization operation on the securities compliance question to obtain the question vector corresponding to the securities compliance question; for each text chunk vector in the knowledge vector base corresponding to the internal and external regulatory knowledge base, perform a similarity calculation based on the text chunk vector and the question vector to obtain the correlation parameter corresponding to the text chunk vector;

[0142] Based on the correlation parameters corresponding to all text chunk vectors in the knowledge vector base corresponding to the internal and external regulatory knowledge base, determine the target regulatory data from the internal and external regulatory knowledge base;

[0143] Input all the target regulatory data into the securities compliance large model, and generate the securities compliance answer corresponding to the securities compliance question based on the securities compliance large model.

[0144] It can be seen that this optional embodiment can perform a vectorization operation on the securities compliance question to obtain the question vector, perform a similarity calculation between each text chunk vector and the question vector to obtain the correlation parameter, determine the target regulatory data based on the correlation parameter, and input all the target regulatory data into the securities compliance large model to generate the securities compliance answer, which is beneficial to the determination accuracy of the target regulatory data, and further improves the generation accuracy of the securities compliance answer, thereby improving the response accuracy and efficiency for securities compliance issues.

[0145] In yet another optional embodiment, for each text chunk vector in the knowledge vector base corresponding to the internal and external regulatory knowledge base, the second determination module 203 performs a similarity calculation based on the text chunk vector and the question vector to obtain the correlation parameter corresponding to the text chunk vector. The specific method includes:

[0146] For each text block vector in the knowledge vector library corresponding to the internal and external regulations knowledge base, perform a similarity calculation on the text block vector and the question vector based on the cosine similarity method and the text matching algorithm in this article to obtain the correlation parameter corresponding to the text block vector; the cosine similarity method is used to determine the cosine distance between the text block vector and the question vector; the text matching algorithm in this article is used to determine the distance between the securities compliance question corresponding to the question vector and the text block corresponding to each text block vector.

[0147] It can be seen that this optional embodiment can complete the calculation based on the cosine similarity method and the text matching algorithm in this article when performing a similarity calculation based on the text block vector and the question vector, which is beneficial to improving the calculation accuracy of the correlation parameter, and then improving the determination accuracy of the target regulation data, thereby improving the response accuracy and efficiency for securities compliance issues.

[0148] In another optional embodiment, the second determination module 203 determines the target regulation data from the internal and external regulations knowledge base based on the correlation parameters corresponding to all text block vectors in the knowledge vector library corresponding to the internal and external regulations knowledge base. The specific method includes:

[0149] Based on the correlation parameters corresponding to all text block vectors in the knowledge vector library corresponding to the internal and external regulations knowledge base, determine several target text block vectors corresponding to the question vector from all text block vectors in the knowledge vector library;

[0150] For each target text block vector, perform a text backtracking operation on the target text block vector based on the internal and external regulations knowledge base corresponding to the knowledge vector library to obtain the regulation data corresponding to the target text block vector;

[0151] Determine the regulation data corresponding to all target text block vectors as candidate regulation data;

[0152] Based on the target attributes of each candidate regulation data, determine at least one target regulation data from all candidate regulation data; the target attributes include at least one of the document integrity, document length of the candidate regulation data, and the correlation parameter of the corresponding target text block vector.

[0153] It can be seen that this optional embodiment can determine several target text block vectors corresponding to the question vector based on the correlation parameters corresponding to all text block vectors; perform a text backtracking operation on the target text block vector based on the internal and external regulations knowledge base to obtain the regulation data corresponding to the target text block vector and determine it as candidate regulation data, and determine at least one target regulation data based on the target attributes of the candidate regulation data, which is beneficial to improving the determination accuracy of the target regulation data, and then improving the generation accuracy of the securities compliance answer, thereby improving the response accuracy and efficiency for securities compliance issues.

[0154] In yet another alternative embodiment, the first determination module 201 adjusts a pre-determined initial general large model based on a fine-tuning data set to obtain a securities compliance large model. The specific method includes:

[0155] Inject a low-rank matrix into the self-attention layer of the pre-determined initial general large model, and add a prompt vector to each layer of the initial general large model; the prompt vector is a vector related to securities compliance and risks;

[0156] Based on the fine-tuning data set, the low-rank matrix in the self-attention layer of the initial general large model, and the prompt vectors in each layer, adjust the initial general large model through gradient descent and backpropagation methods to obtain a securities compliance large model.

[0157] It can be seen that this alternative embodiment can adjust the initial general large model through gradient descent and backpropagation methods based on the fine-tuning data set, the low-rank matrix in the self-attention layer of the initial general large model, and the prompt vectors in each layer to obtain a securities compliance large model, which is beneficial to improving the determination accuracy of the securities compliance large model, and further improving the generation accuracy of generating securities compliance answers based on the securities compliance large model, thereby improving the response accuracy and efficiency for securities compliance issues.

[0158] Embodiment III

[0159] Please refer to Figure 4 , Figure 4 which is a schematic structural diagram of an intelligent question-answering implementation system for securities compliance disclosed in an embodiment of the present invention. As Figure 4 shown, the intelligent question-answering implementation system for securities compliance may include:

[0160] A memory 301 storing executable program code;

[0161] A processor 302 coupled to the memory 301;

[0162] The processor 302 calls the executable program code stored in the memory 301 and executes the steps in the method for implementing intelligent question-answering for securities compliance described in Embodiment I of the present invention.

[0163] Embodiment IV

[0164] An embodiment of the present invention discloses a computer storage medium, which stores computer instructions that, when called, are used to execute the steps in the method for implementing intelligent question-answering for securities compliance described in Embodiment I of the present invention.

[0165] Embodiment V

[0166] An embodiment of the present invention discloses a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to cause a computer to execute the steps in the intelligent question-answering implementation method for securities compliance described in Embodiment 1.

[0167] The system embodiments described above are merely illustrative. Modules described as separate components may or may not be physically separated, and components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.

[0168] Through the specific description of the above embodiments, those skilled in the art can clearly understand that each implementation manner can be realized by means of software plus a necessary general hardware platform, and of course, it can also be realized by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium. The storage medium includes read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc memories, magnetic disk memories, tape memories, or any other computer-readable medium capable of carrying or storing data.

[0169] Finally, it should be noted that: What is disclosed in an intelligent question-answering implementation method and device for securities compliance disclosed in the embodiments of the present invention is only the preferred embodiments of the present invention, which are only used to illustrate the technical solutions of the present invention, rather than to limit it; Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: They can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; And these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for implementing intelligent question and answer for securities compliance, characterized in that: The method comprises: Determine a securities compliance big model based on a target data set and a predetermined initial general big model; the target data set includes one or more of a securities industry question and answer data set, a legal knowledge question and answer data set, and a general encyclopedia question and answer data set; Building an internal and external regulation knowledge base based on a regulatory data set; the regulatory data set includes a number of regulatory data, and the regulatory data includes external regulatory data or internal regulatory data; Performing block processing on each regulatory data contained in the internal and external regulatory knowledge base to obtain a plurality of text blocks corresponding to each regulatory data; performing vectorization operation on each text block of each regulatory data to obtain a text block vector corresponding to the text block; and constructing a knowledge vector library corresponding to the internal and external regulatory knowledge base based on the text block vectors corresponding to all the text blocks; When a securities compliance question is received, a vectorization operation is performed on the securities compliance question to obtain a question vector corresponding to the securities compliance question; for each text block vector in the knowledge vector library corresponding to the internal and external regulations knowledge base, a similarity calculation is performed based on the text block vector and the question vector to obtain a correlation parameter corresponding to the text block vector; Determining target regulatory data from the internal and external regulations knowledge base based on correlation parameters corresponding to all text block vectors in the knowledge vector base corresponding to the internal and external regulations knowledge base; All the target regulatory data are input into the securities compliance big model, and securities compliance answers corresponding to the securities compliance questions are generated based on the securities compliance big model.

2. The intelligent question-answering method for securities compliance according to claim 1, characterized in that: The step of determining the securities compliance big model based on the target data set and the predetermined initial general big model includes: Cleaning the target data set to remove non-target data in the target data set, and obtaining a cleaned data set corresponding to the target data set; the non-target data includes duplicate data, irrelevant data, and data whose data quality does not meet a preset quality condition in the target data set; Annotating the cleaned data set corresponding to the target data set to obtain a fine-tuning data set corresponding to the target data set; Adjusting the predetermined initial general macro model based on the fine-tuning data set to obtain a securities compliance macro model; Furthermore, before the forming of the domestic and foreign regulations knowledge base based on the regulatory data set, the method further includes: An initial regulatory data set is obtained and a correction operation is performed on the initial regulatory data set to obtain a regulatory data set corresponding to the initial regulatory data set; the correction operation includes an operation of removing redundant information, an operation of correcting format errors, and an operation of unifying terminology expressions.

3. The intelligent question-answering method for securities compliance according to claim 1, characterized in that: For each text block vector in the knowledge vector library corresponding to the internal and external regulations knowledge library, similarity calculation is performed based on the text block vector and the question vector to obtain a correlation parameter corresponding to the text block vector, including: For each text block vector in the knowledge vector library corresponding to the internal and external regulations knowledge base, a similarity calculation is performed between the text block vector and the question vector based on the cosine similarity method and the matching algorithm in this paper to obtain the correlation parameter corresponding to the text block vector; the cosine similarity method is used to determine the cosine distance between the text block vector and the question vector; the matching algorithm in this paper is used to determine the distance between the securities compliance issue corresponding to the question vector and the text block corresponding to each of the text block vectors.

4. The intelligent question-answering method for securities compliance according to claim 1 or 3, characterized in that: The step of determining target regulatory data from the internal and external regulations knowledge base based on the correlation parameters corresponding to all text block vectors in the knowledge vector base corresponding to the internal and external regulations knowledge base includes: Based on the correlation parameters corresponding to all text block vectors in the knowledge vector library corresponding to the internal and external regulations knowledge library, a plurality of target text block vectors corresponding to the question vector are determined from all text block vectors in the knowledge vector library; For each of the target text block vectors, a text backtracking operation is performed on the target text block vector based on the internal and external regulations knowledge base corresponding to the knowledge vector library to obtain the regulatory data corresponding to the target text block vector; Determine the regulatory data corresponding to all the target text block vectors as candidate regulatory data; Based on the target attribute of each candidate regulatory data, at least one target regulatory data is determined from all the candidate regulatory data; the target attribute includes at least one of the document integrity, document length and relevance parameters of the corresponding target text block vector of the candidate regulatory data.

5. The intelligent question-answering method for securities compliance according to claim 2, characterized in that: The step of adjusting the predetermined initial general macro model based on the fine-tuning data set to obtain the securities compliance macro model includes: Injecting a low-rank matrix into a self-attention layer of a predetermined initial general large model, and adding a prompt vector in each layer of the initial general large model; the prompt vector is a vector related to securities compliance and risk; Based on the fine-tuning dataset, the low-rank matrix in the self-attention layer of the initial general large model and the prompt vector in each layer, the initial general large model is adjusted by gradient descent and back propagation methods to obtain a securities compliance large model.

6. An intelligent question-answering device for securities compliance, characterized in that: The device comprises: A first determination module is used to determine a securities compliance big model based on a target data set and a predetermined initial general big model; the target data set includes one or more of a securities industry question and answer data set, a legal knowledge question and answer data set, and a general encyclopedia question and answer data set; A building module, used to build an internal and external regulation knowledge base based on a regulatory data set; the regulatory data set includes a number of regulatory data, and the regulatory data includes external regulatory data or internal regulatory data; A second determination module is used to, when receiving a securities compliance question, determine a securities compliance answer corresponding to the securities compliance question based on the securities compliance question, the internal and external regulations knowledge base, and the securities compliance big model; The device also includes: a slicing module, configured to, before the second determining module determines the securities compliance answer corresponding to the securities compliance question based on the securities compliance question, the internal and external regulations knowledge base and the securities compliance big model, perform slicing processing on each regulatory data contained in the internal and external regulations knowledge base to obtain a plurality of text blocks corresponding to each regulatory data; for each text block of each regulatory data, perform vectorization operation on the text block to obtain a text block vector corresponding to the text block; and construct a knowledge vector library corresponding to the internal and external regulations knowledge base based on the text block vectors corresponding to all the text blocks; Furthermore, the second determination module determines the securities compliance answer corresponding to the securities compliance question based on the securities compliance question, the internal and external regulations knowledge base and the securities compliance big model, and the specific method includes: Performing a vectorization operation on the securities compliance issue to obtain a problem vector corresponding to the securities compliance issue; for each text block vector in the knowledge vector library corresponding to the internal and external regulations knowledge base, performing a similarity calculation based on the text block vector and the problem vector to obtain a correlation parameter corresponding to the text block vector; Determining target regulatory data from the internal and external regulations knowledge base based on correlation parameters corresponding to all text block vectors in the knowledge vector base corresponding to the internal and external regulations knowledge base; All the target regulatory data are input into the securities compliance big model, and securities compliance answers corresponding to the securities compliance questions are generated based on the securities compliance big model.

7. An intelligent question-answering system for securities compliance, characterized in that: The system comprises: A memory storing executable program code; a processor coupled to the memory; The processor calls the executable program code stored in the memory to execute the intelligent question-and-answer implementation method for securities compliance as described in any one of claims 1 to 5.

8. A computer storage medium, characterized in that: The computer storage medium stores computer instructions, which, when called, are used to execute the intelligent question-and-answer implementation method for securities compliance as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Method for constructing compliance risk control AI knowledge base field by using LLM

    CN118886489A

  • Financial knowledge question and answer method and system and storage medium

    CN119336887A