Interface calling code generation method and device, equipment and readable storage medium
The code generation method is called through interface training of large language models and data sets, and the problem of low efficiency and accuracy of calling code generation of financial indicator interfaces is solved, and efficient and accurate financial data acquisition and software development process is achieved.
Patent Information
- Application Number
- CN202510454978.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-18
AI Technical Summary
In the prior art, the financial index interface call code generation efficiency is low and the accuracy is low, the user operation takes a long time and requires high financial knowledge reserves.
By using the matching method to determine the target financial indicator interface document based on query instructions, combining the large language model and interface call code generation model, interface call code is generated, including the financial indicator interface document training data set, query instruction training data set and interface call code training data set data set data set, and precise matching and analysis are used using named entity recognition, financial information processing and vector models.
It improves the efficiency and accuracy of interface call code generation, simplifies user operation processes, reduces the requirements for financial knowledge, and is suitable for financial data acquisition and software development projects.
Smart Images

Figure CN120335796A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly relates to a method, device, equipment and readable storage medium for generating interface call codes. Background Art
[0002] In the current financial data retrieval process, a user needs to perform a series of operations on the user interface (UI), including selecting specific financial indicators and inputting relevant parameters. Thereafter, the system generates a financial indicator API (Application Programming Interface) call code based on the user's selection and input parameters to obtain corresponding market quotation data. However, this process has certain drawbacks. Firstly, the operations of the user on the UI are time-consuming and inefficient; secondly, it requires a high level of financial knowledge reserve and literacy of the user. As a result, the efficiency of generating the financial indicator API call code is low, and the accuracy is also low.
[0003] Therefore, how to improve the efficiency and accuracy of generating financial indicator interface call codes is a technical problem that those skilled in the art urgently need to solve. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a method, device, equipment and readable storage medium for generating interface call codes, which solves the technical problems of low efficiency and low accuracy in generating interface call codes in the prior art.
[0005] To solve the above technical problems, the present invention provides a method for generating interface call codes, including:
[0006] Determining a target financial indicator interface document based on a query instruction by using a matching method; the matching method is a method of matching based on the underlying assets in an indicator library or a method of inferring and querying by using a large language model;
[0007] Generating an interface call code by using an interface call code generation model based on the query instruction and the target financial indicator interface document; wherein, the interface call code generation model is a model obtained by training a large language model based on an interface call code generation instruction data set; the interface call code generation instruction data set includes a financial indicator interface document training data set, a generated query instruction training data set and a generated interface call code training data set.
[0008] Optionally, generating an interface call code by using an interface call code generation model based on the query instruction and the target financial indicator interface document includes:
[0009] Converting the time information in the query instruction into a set standard date and time, and identifying the underlying assets in the query instruction and the types of securities to which they belong;
[0010] Based on the query instruction, the target financial indicator interface document, the standard date and time, and the underlying object and the type of securities to which it belongs, use the interface call code generation model to generate the interface call code.
[0011] Optionally, converting the time information in the query instruction into a set standard date and time, and identifying the underlying object and the type of securities to which it belongs in the query instruction, includes:
[0012] Based on a time parsing tool, parse the time information in the query instruction to determine whether the parsing is successful;
[0013] When the parsing fails, use a time parsing large model to convert the vocabulary representing the time information in the query instruction into the standard date and time; the time parsing large model is a model trained on a large language model based on a time information extraction dataset;
[0014] Determine whether the underlying object expression in the query instruction matches the underlying object in the underlying object library;
[0015] If it matches, determine the underlying object and the type of securities to which it belongs based on the underlying object library;
[0016] If it does not match, infer the underlying object and the type of securities to which it belongs included in the query instruction based on the underlying object fuzzy matching large model; among them, the underlying object fuzzy matching large model is a model trained on a large language model based on an underlying object extraction dataset.
[0017] Optionally, determining the target financial indicator interface document based on the query instruction using a matching method, includes:
[0018] Use a named entity recognition model to identify the financial indicator-related vocabulary from the query instruction to obtain the financial indicator-related vocabulary;
[0019] Use a financial information processing large model to perform inference to determine the quantity and name of the financial indicators in the query instruction; the financial information processing large model is a model trained on a large language model based on an indicator inference dataset and a multi-indicator parsing dataset; the indicator inference dataset is a dataset for determining the indicator path and indicator name corresponding to the query instruction, and the multi-indicator parsing dataset is a dataset for determining multiple financial indicators in the query instruction;
[0020] Based on the financial indicator-related vocabulary, the quantity and name of the financial indicators, perform a match in the financial indicator library. If the match is successful, obtain the target financial indicator interface document;
[0021] If the match is not successful, based on the financial indicator-related vocabulary, the quantity of the financial indicator, and its name, use the financial information processing large model to perform a matching process to obtain a financial indicator path and a standard name;
[0022] Based on the financial indicator path and the standard name, use the best matching text similarity to retrieve in the financial indicator library, and screen out the top K financial indicators with high similarity to the query instruction to obtain the financial indicator document;
[0023] Use a financial vector model to determine the similarity between the financial indicator-related vocabulary and the financial indicator document, and select the document with the highest similarity as the target financial indicator interface document.
[0024] Optionally, before using the named entity recognition model to identify the vocabulary related to financial indicators from the query instruction to obtain the financial indicator-related vocabulary, it further includes:
[0025] Construct a named entity recognition dataset; the named entity recognition dataset is used to identify the vocabulary describing financial indicators in the query instruction;
[0026] Based on the named entity recognition dataset, train the named entity recognition model to be trained to obtain a trained named entity recognition model;
[0027] Combine the trained named entity recognition model and the large language model to obtain the named entity recognition model, and the named entity recognition model is used to identify the vocabulary used to describe indicator information in the instruction.
[0028] Optionally, before using the financial vector model to determine the similarity between the financial indicator-related vocabulary and the financial indicator document, and select the document with the highest similarity as the target financial indicator interface document, it further includes:
[0029] Based on the indicator inference dataset, use the similarity methods of vector cosine similarity and best matching text similarity to determine the correct indicator interface document corresponding to each query instruction, and construct a financial vector model training dataset; the financial vector model training dataset is used to distinguish the semantic differences between financial indicator interface documents with similar semantic content;
[0030] Based on the financial vector model training dataset, train the financial vector model to be trained to obtain the financial vector model; the financial vector model is used to determine the financial indicator corresponding to the query instruction.
[0031] Optionally, the construction process of the indicator inference dataset includes:
[0032] Generate multiple query instructions using a large language model based on the financial indicator interface document, and label the time information, financial indicator data, target information, and time information in each generated query instruction to obtain the labeled query instructions;
[0033] Based on the labeled query instructions and the financial indicator interface document, construct the indicator inference dataset, which is used to implement the inference from query instructions to financial indicators.
[0034] Optionally, the construction process of the multi - indicator parsing dataset includes:
[0035] Based on the financial indicator data and the multiple query instructions, construct the multi - indicator parsing dataset, which is used to correspond query instructions with financial indicator data.
[0036] Optionally, before generating the interface call code using the interface call code generation model based on the query instructions and the target financial indicator interface document, it further includes:
[0037] Construct an interface call code generation strategy based on the financial indicator interface document;
[0038] Generate callable interface call code using the large language model based on the interface call code generation strategy and the prompt template;
[0039] Execute each callable interface call code, obtain the return result of executing each interface call code, verify the return result, and retain the interface call code that passes the verification to obtain the verified interface call code;
[0040] According to the financial indicator interface document, the verified interface call code, and the target example, guide the large language model to generate query instructions that match the verified interface call code;
[0041] Based on the financial indicator interface document, the verified interface call code, and the matching query instructions, generate an interface call code generation instruction dataset;
[0042] Use the interface call code generation instruction dataset to train the large language model to obtain the interface call code generation model.
[0043] Optionally, after generating the interface call code using the interface call code generation model based on the query instructions and the target financial indicator interface document, it further includes:
[0044] Perform rule detection on the interface call code to obtain the final interface call code; where the rule detection is used to avoid generating unreasonable parameters.
[0045] The present application also provides an interface call code generation device, including:
[0046] A target financial indicator interface document determination module, configured to determine a target financial indicator interface document based on a query instruction by using a matching method; the matching method is a method of matching based on the targets in the indicator library or a method of performing inference query by using a large language model;
[0047] An interface call code generation module, configured to generate an interface call code based on the query instruction and the target financial indicator interface document by using an interface call code generation model; wherein, the interface call code generation model is a model obtained by training a large language model based on an interface call code generation instruction data set; the interface call code generation instruction data set includes a financial indicator interface document training data set, a generated query instruction training data set, and a generated interface call code training data set.
[0048] The present application also provides an interface call code generation device, including:
[0049] A memory, configured to store a computer program;
[0050] A processor, configured to execute the computer program to implement the steps of the interface call code generation method as described above.
[0051] The present application also provides a readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the interface call code generation method as described above are implemented.
[0052] A computer program product, including computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the interface call code generation method as described above are implemented.
[0053] It can be seen that the present invention determines the target financial indicator interface document by using the matching method based on the query instruction; the matching method is a method of matching based on the underlying objects in the indicator library or a method of performing inference queries using a large language model; based on the query instruction and the target financial indicator interface document, an interface call code generation model is used to generate an interface call code; wherein, the interface call code generation model is a model obtained by training a large language model based on an interface call code generation instruction dataset; the interface call code generation instruction dataset includes a financial indicator interface document training dataset, a generated query instruction training dataset, and a generated interface call code training dataset. Compared with the current manual operation on the interface to select financial indicators, and since financial indicators are relatively professional, the efficiency and accuracy of manual financial indicator selection are relatively low, resulting in low efficiency and low accuracy of the generated API call code (interface call code). In contrast, in this application, only by inputting a query instruction, the target financial indicator document corresponding to the query instruction can be obtained, and then based on the target financial indicator document and the query instruction, the interface call code generation model is directly used to generate the interface call code corresponding to the query instruction, thereby improving the efficiency and accuracy of interface call code generation.
[0054] In addition, the present invention also provides an interface call code generation device, equipment, and readable storage medium, which also have the above beneficial effects. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on the provided drawings.
[0056] Figure 1 It is a flowchart of an interface call code generation method provided by an embodiment of the present invention;
[0057] Figure 2 It is an example diagram of the construction process of a dataset provided by an embodiment of the present invention;
[0058] Figure 3 It is an example diagram of a financial indicator inference dataset provided by an embodiment of the present invention;
[0059] Figure 4 It is an example diagram of a multi - indicator parsing dataset provided by an embodiment of the present invention;
[0060] Figure 5 It is an example flowchart of a diversified query instruction generation method provided by an embodiment of the present invention;
[0061] FIG. 6(1) and FIG. 6(2) are schematic diagrams of a prompt template provided by an embodiment of the present invention;
[0062] Figure 7 FIG. is an example diagram of the generation process of a callable interface call code provided by an embodiment of the present invention;
[0063] Figure 8 FIG. is a schematic diagram of a prompt template for generating an interface call code provided by an embodiment of the present invention;
[0064] Figure 9 FIG. is an example diagram of a verified interface call code provided by an embodiment of the present invention;
[0065] Figure 10 FIG. is a schematic diagram of a query instruction corresponding to an interface call code provided by an embodiment of the present invention;
[0066] Figure 11 FIG. is an example diagram of a prompt template for a query instruction corresponding to an interface call code provided by an embodiment of the present invention;
[0067] Figure 12 FIG. is an example diagram of the construction process of a dataset for generating an interface call code instruction provided by an embodiment of the present invention;
[0068] Figure 13 FIG. is a sample diagram of a dataset for generating an interface call code instruction provided by an embodiment of the present invention;
[0069] Figure 14 FIG. is an example diagram of the process of a method for generating a dataset provided by an embodiment of the present invention;
[0070] Figure 15 FIG. is a sample diagram of a dataset for extracting target and time information provided by an embodiment of the present invention;
[0071] Figure 16 FIG. is a schematic diagram of the process of a method for training a large language model provided by an embodiment of the present invention;
[0072] Figure 17 FIG. is an example diagram of the process of a method for training a financial vector model provided by an embodiment of the present invention;
[0073] Figure 18 FIG. is an example diagram of the process of a method for constructing a training dataset for a financial vector model provided by an embodiment of the present invention;
[0074] Figure 19 FIG. is an example diagram of the process of training a financial vector model based on a training dataset for a financial vector model provided by an embodiment of the present invention;
[0075] Figure 20A flowchart example of a method for generating interface call code provided by an embodiment of the present invention;
[0076] Figure 21 A system for generating interface call code provided by an embodiment of the present invention;
[0077] Figure 22 A schematic structural diagram of a device for generating interface call code provided by an embodiment of the present invention;
[0078] Figure 23 A schematic structural diagram of a device for generating interface call code provided by an embodiment of the present invention. Detailed implementation manners
[0079] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part rather than all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0080] Some terms or concepts that appear during the description of the embodiments of the present application are applicable to the following explanations:
[0081] (1) Prompt In large language models (LLMs) such as GPT series models, "Prompt" refers to a piece of text or instruction provided to the model, which is used to guide the model to generate subsequent content. Simply put, Prompt is like a starting point or a question, and the model generates answers or continues to write text based on this starting point or question. In practical applications, Prompt can be very flexible, either a single sentence, a paragraph, or a more complex structure, such as a conversation with multiple parts, data input with a specific format, etc. By carefully designing Prompt, users can make the model perform various tasks, such as writing, translation, question answering, code generation, etc.
[0082] (2) NER, namely Named Entity Recognition, is a fundamental and important task in natural language processing. Named Entity Recognition refers to identifying entities with specific meanings in text. These entities usually belong to some predefined categories, such as person names, place names, organization names, time, dates, currencies, percentages, etc., and marking them. For example, in the sentence "XX Company released a new mobile phone in 2022", "XX Company" is an organization name and "2022" is a time, both of which are named entities that need to be identified by NER. In recent years, deep learning has achieved remarkable results in the NER task. The NER model based on LSTM can handle the sequential information in text well, capture long-distance dependencies, and improve the recognition accuracy of named entities.
[0083] (3) A meta-token can be understood as a kind of "meta-tag" or "meta-symbol". It is not a token that directly represents specific words or semantic units in the usual sense, but an abstract tag used to represent some meta-information about tokens or with special functions and meanings. It is a special element introduced during the process of text processing, model training, inference, etc., to help the model better understand and process information such as the structure and semantics of text. During model training and inference, meta-tokens can be used to control the behavior of the model. For example, some meta-tokens can be used as instruction tags to tell the model to perform specific tasks, such as abstract generation, sentiment analysis, etc. When the model encounters these meta-tokens, it will adjust the generation strategy or inference direction according to their meanings.
[0084] (4) A large language model (LLM for short) is an artificial intelligence model based on deep learning technology, aiming to understand and generate human language. These models learn the structure, meaning, and usage of language by training on a vast amount of text data, and thus can perform various natural language processing tasks.
[0085] Please refer to Figure 1 , Figure 1 which is a flowchart of a method for generating interface call code provided by an embodiment of the present invention. The method may include:
[0086] S101, determining a target financial metric interface document based on a query instruction using a matching method; the matching method is a method of matching based on the targets in the metric library or a method of performing inference queries using a large language model.
[0087] The execution entity in this embodiment is an electronic device. The electronic device in this embodiment can be a computer, a mobile phone, etc. The query instruction in this embodiment is a natural language query instruction input by the user according to needs. The query instruction is an instruction related to financial indicators, and the financial indicators to be queried can be determined based on this query instruction. In this embodiment, when using the matching method, if the financial indicators in the query instruction can be accurately matched with the financial indicators in the financial indicator library, the matching method is a method of matching based on the targets in the indicator library (financial indicator library). If the financial indicators in the query instruction cannot be accurately matched with the financial indicators in the financial indicator library, the method of using a large language model for inference query is adopted. The large language model in this embodiment can determine accurate financial indicator terms based on the nouns related to financial indicators in the query instruction. The target financial indicator interface document in this embodiment is a document related to financial indicators.
[0088] Further, based on any of the above embodiments, in order to improve the accuracy of determining the target financial indicator interface document, the above method of using the matching method based on the query instruction to determine the target financial indicator interface document may include:
[0089] S1011: Use a named entity recognition model to identify the terms related to financial indicators from the query instruction to obtain financial indicator related terms.
[0090] The named entity recognition model in this embodiment is a model trained based on a financial indicator named entity recognition data set constructed from the company's business data. Most traditional methods are trained based on data sets in life scenarios such as personal names and place names. The innovation lies in constructing a financial indicator named entity recognition data set, which can be directly applied in the financial field. The method of constructing the data set is the innovation point, enabling it to accurately identify the terms used to describe financial indicator information in the instruction. For example, the named entity recognition data set is ['Please O', 'Provide O', '3O', '0O', '0O', '5O', '9O', '6O', '.O', 'SO', 'ZO', 'From O', '2O', '0O', '2O', '4O', 'Year O', '1O', '1O', 'Month O', '1O', 'Day O', 'To O', '2O', '0O', '2O', '4O', 'Year O', '1O', '1O', 'Month O', '3O', '0O', 'Day O', 'Of O', 'Price B-INDI', 'Change I-INDI', 'Percentage I-INDI', '.O'].
[0091] S1012. Use the large financial information processing model for inference to determine the quantity and names of financial metrics in the query instruction. The large financial information processing model is a model obtained by training a large language model based on an index inference dataset and a multi-index parsing dataset. The index inference dataset is a dataset used to determine the index path and index name corresponding to the query instruction, and the multi-index parsing dataset is a dataset used to determine multiple financial metrics in the query instruction.
[0092] In this embodiment, the trained named entity recognition model is combined with the large financial information processing model to construct a financial metric inference system (large financial information processing model). This embodiment uses the large financial information processing model to infer the specific quantity and names of financial metrics involved in the user's query. Compared with the inference system based solely on the large language model, this fusion construction system can effectively avoid the probability of hallucinations and errors in the large model during the inference process.
[0093] S1013. Based on the financial metric-related vocabulary, the quantity and names of financial metrics, perform a match in the financial metric library. If the match is successful, obtain the target financial metric interface document.
[0094] This embodiment performs an exact match, an exact match in the financial metric library. If the user uses the standard metric name or the metric alias recorded in the library and the corresponding metric can be directly retrieved in the financial metric library, directly proceed to step S102.
[0095] S1014. If the match is unsuccessful, based on the financial metric-related vocabulary, the quantity and names of financial metrics, use the large financial information processing model to perform a matching process to obtain the financial metric path and standard name.
[0096] This embodiment performs a fuzzy matching process: If the exact match is unsuccessful, it means that the query instruction may use the alias or abbreviation of the metric. At this time, the system further infers the financial metric path and its standard name involved in the user's query through the large language model.
[0097] S1015. Based on the financial metric path and standard name, use the best-matching text similarity to retrieve in the financial metric library, and screen out the top K financial metrics with high similarity to the query instruction to obtain the financial metric document.
[0098] BM25 (Best Matching 25) in this embodiment is an algorithm widely used in the field of information retrieval. It is an improved version of TF-IDF and is mainly used to measure the relevance between a document and a query. The specific value of K is not limited in this embodiment. For example, K in this embodiment can be 3; or K in this embodiment can also be 5, etc. This embodiment uses the BM25 algorithm (Best Matching Text Similarity) to retrieve in the financial indicator library and screens out the top K financial indicators with the highest similarity to the user query (query instruction). The BM25 algorithm in this embodiment and the BM25 text similarity are different expressions of the same thing.
[0099] S1016. Use the financial vector model to determine the similarity between the financial indicator-related vocabulary and the financial indicator document, and select the document with the highest similarity as the target financial indicator interface document.
[0100] This embodiment uses the financial vector model to calculate the cosine similarity between the extracted financial indicator-related vocabulary and the financial indicator document, and selects the document with the highest similarity as the final indicator document. The financial vector model in this embodiment is an improved financial vector model, which is a model combined based on vector cosine similarity and BM25 text similarity (Best Matching Text Similarity). This embodiment can construct a dataset for training the financial vector model based on the indicator inference dataset, so as to train the financial vector model based on strong negative example contrast learning (merging the K indicator documents with the highest vector cosine similarity and the K indicator documents with the highest BM25 text similarity. During this process, duplicate indicator documents and labeled indicator documents are removed, and then strong negative example samples (i.e., indicator documents that are easy to be confused with the original query) are obtained), so as to accurately retrieve the corresponding financial indicator according to the user query. This dataset aims to guide the vector model to accurately distinguish the fine-grained semantic differences between indicator documents with similar semantic content, thereby improving the accuracy of vector retrieval.
[0101] Further, based on any of the above embodiments, in order to improve the accuracy of the named entity recognition model in recognizing financial indicators, before using the named entity recognition model to recognize the words related to financial indicators from the query instruction and obtaining the words related to financial indicators, it may further include: constructing a named entity recognition data set; the named entity recognition data set is used to recognize the words describing financial indicators in the query instruction; training the named entity recognition model to be trained based on the named entity recognition data set to obtain a trained named entity recognition model; combining the trained named entity recognition model with a large language model to obtain a named entity recognition model, and the named entity recognition model is used to recognize the words used to describe the indicator information in the instruction. Compared with the inference system based solely on the large language model, in the inference process, the named entity recognition model can effectively avoid the probability of the large model generating hallucinations and making mistakes.
[0102] It should be further noted that, in order to improve the accuracy of similarity determination, before using the financial vector model to determine the similarity between the words related to financial indicators and the financial indicator documents, and selecting the document with the highest similarity as the target financial indicator interface document, it may further include: based on the indicator inference data set, using the similarity methods of vector cosine similarity and best matching text similarity (BM25), determining the correct indicator interface document corresponding to each query instruction, and constructing a financial vector model training data set; the financial vector model training data set is used to distinguish the semantic differences between the financial indicator interface documents with similar semantic content; training the financial vector model to be trained based on the financial vector model training data set to obtain a financial vector model; the financial vector model is used to determine the financial indicators corresponding to the query instruction. In this embodiment, the correct indicator interface document refers to the correct indicator document corresponding to the query instruction, which can be considered as an indicator document marked manually and can be called a marked indicator document. The indicator inference data set in this embodiment is a data set that enables the model to accurately infer the path and specific name of the financial indicator from the query instruction. The best matching text similarity (BM25) in this embodiment can be combined with the vector cosine similarity to accurately judge the similarity from both the text and semantic aspects.
[0103] It should be further noted that, based on any of the above embodiments, in order to improve the accuracy of constructing the indicator inference dataset, the construction process of the above indicator inference dataset may include: generating multiple query instructions using a large language model based on the financial indicator interface document, and annotating the time information, financial indicator data, target information, and time information in each generated query instruction to obtain the annotated query instructions; constructing an indicator inference dataset based on the annotated query instructions and the financial indicator interface document, where the indicator inference dataset is used to implement the inference from the query instructions to the financial indicators. It should be noted that the large language model throughout the text is a deep learning model trained using a large amount of text data, capable of generating natural language text or understanding the meaning of language text. The large language model throughout the text can be the same or different as the basic model for training other models, and can be selected according to requirements. The indicator inference dataset can be denoted as, , where represents the i-th indicator document, and there are N indicator documents in the entire dataset; represents the j-th query for the i-th indicator, and each indicator has K queries. In the input part, introduce , which is used to direct the model to route to the indicator inference function. The context details the steps that the model needs to complete. The output part will cover the indicator path and indicator name corresponding to the user's query. Constructing this dataset aims to endow the model with the ability to accurately infer the correct financial indicators from the query instructions. Figure 2 FIG. is a flowchart example of constructing a dataset provided by an embodiment of the present invention, including the construction of a financial indicator inference dataset, the construction of a multi-indicator parsing dataset, and an indicator information NER dataset (named entity recognition dataset). Figure 3 FIG. is an example diagram of a financial indicator inference dataset provided by an embodiment of the present invention.
[0104] It should be further noted that, based on any of the above embodiments, the construction process of the above multi-indicator parsing dataset may include: constructing a multi-indicator parsing dataset based on the financial indicator data and multiple query instructions, where the multi-indicator parsing dataset is used to correspond the query instructions with the financial indicator data. The multi-indicator parsing dataset in this embodiment focuses on processing multiple financial indicators involved in complex queries and improving the model's multi-indicator parsing ability. The multi-indicator parsing dataset in this embodiment uses the query instructions as the input content, and the query instructions may involve the query requirements for single or multiple financial indicator data. And its output result is the financial indicator data that precisely corresponds to the input query instructions. For specific data examples of the multi-indicator parsing dataset, refer to Figure 4 , Figure 4 FIG. is an example diagram of a multi-indicator parsing dataset provided by an embodiment of the present invention.
[0105] S102. Based on the query instruction and the target financial indicator interface document, use the interface call code generation model to generate the interface call code. The interface call code generation model is a model obtained by training a large language model based on the interface call code generation instruction dataset. The interface call code generation instruction dataset includes the financial indicator interface document training dataset, the generated query instruction training dataset, and the generated interface call code training dataset.
[0106] The input of the interface call code generation model in this embodiment is the query instruction and the target financial indicator interface document, and the output is the interface call code. When training the interface call code generation model, it is necessary to use the financial indicator interface document training dataset, the generated query instruction training dataset, and the generated interface call code training dataset. The financial indicator interface document training dataset in this embodiment is a dataset including financial indicator interface documents. The generated query instruction training dataset in this embodiment can be based on a single or multiple indicator API documents, and by using prompt engineering techniques, guide the large language model to generate a series of rich and diverse queries that comprehensively cover various user usage habits and various application scenarios. At the same time, guide the large language model to accurately label the time information, indicator information, and target information in each query. In this process, strictly ensure that the indicator terms (financial indicators) (indicator_terms), time (time), and target (institution) terms marked by the model are only obtained from the corresponding user queries (Query) to prevent the model from having hallucination phenomena. The generation process of the interface call code training dataset in this embodiment can be: design a prompt template, which guides the large language model to generate reasonable API call codes, and this template can guide the large language model to generate reasonable API call codes. This prompt template will ensure that the generated API call codes comply with the regulations in terms of parameter selection and combination according to the requirements of the API description document, so as to achieve automated and efficient API call code generation. For ease of understanding, please refer to Figure 5 , Figure 5A flowchart of a method for generating a diversified query instruction provided by an embodiment of the present invention is provided. From the flowchart, it can be seen that the process is to guide the large language model to generate an API call code instruction (query instruction) matching the API call code based on the indicator API document (i.e., indicator document), the corresponding API call code, and the target example. It should be noted that the indicator API document and the indicator document in the full text express the same concept of financial indicator API document. Based on the online business data and the indicator document, an In-context prompt template for financial indicator query instructions is constructed. For details, see FIG6, which is a schematic diagram of a prompt template provided by an embodiment of the present invention. The data sample generated by the model based on this can be: [{"Query": "Query the increase or decrease of XX in December 2024","indicator_terms": ["increase or decrease"],"time": ["December 2024"],"institution": ["XX"]}.
[0107] It should be further explained that, based on any of the above embodiments, in order to generate an executable interface call code with compliant parameter settings, the above generation of the interface call code based on the query instruction and the target financial indicator interface document using the interface call code generation model may include:
[0108] S1021, converting the time information in the query instruction into a set standard date and time, and identifying the subject matter in the query instruction and the type of securities to which it belongs.
[0109] S1022, based on the query instruction, the target financial indicator interface document, the standard date and time, and the target and the type of securities to which it belongs, an interface call code generation model is used to generate an interface call code.
[0110] The standard date and time in this embodiment refers to the time that needs to follow the set format specification. For example, for the time information (date parameter), fill in a reasonable date before the current time. For example, when querying the closing price, make sure that the filled date is not a future date to prevent invalid queries. For the target information, select the corresponding target according to the securities category to which the financial indicator belongs, that is, ,in represents the name of the i-th label, Denote the code of the i-th underlying security. The process of determining the type of the underlying security in this embodiment can be as follows: ① Direct matching: Determine whether the underlying in the underlying library directly appears in the query instruction. If there is a matching item, directly determine the underlying and its affiliated security type. ② Inference matching: If no directly matching underlying is found in the underlying library, infer the underlying information and its affiliated security type implied in the query instruction with the help of a large language model. In this embodiment, by converting the time information in the query instruction into a set standard date and time, and identifying the underlying and its affiliated security type in the query instruction, the time, underlying, and security type in the generated interface call code are reasonable, providing the user with an executable interface call code with compliant parameter settings.
[0111] It should be further noted that, in order to improve the accuracy of time information and underlying parsing, the above conversion of the time information in the query instruction into a set standard date and time, and the identification of the underlying and its affiliated security type in the query instruction, may include:
[0112] S1: Parse the time information in the query instruction based on a time parsing tool to determine whether the parsing is successful.
[0113] S2: When the parsing fails, use a time parsing large model to convert the vocabulary representing the time information in the query instruction into a standard date and time; the time parsing large model is a model obtained by training a large language model based on a time information extraction dataset.
[0114] S3, Determine whether the underlying expression in the query instruction matches the underlying in the underlying library.
[0115] S4, If it matches, determine the underlying and its affiliated security type based on the underlying library.
[0116] S5, If it does not match, infer the underlying and its affiliated security type included in the query instruction based on a fuzzy underlying matching large model; among them, the fuzzy underlying matching large model is a model obtained by training a large language model based on an underlying extraction dataset.
[0117] The process of parsing time information in this embodiment may include: First, call the internal time parsing tool to extract and parse the time information in the query instruction. If the tool fails to parse successfully, the large language model will be enabled as an auxiliary means to identify and extract the vocabulary representing time information from the query instruction, and convert it into a standardized date format according to the following rules: ① If the time information only contains a date, it is converted into the YYYY-MM-DD format; ② If the time information contains both a date and a time, it is converted into the YYYY-MM-DD hh:mm:ss format. Target information matching (financial targets are converted into codes, and the codes are fixed rules): The system determines the target information in the query instruction through the following steps: ① Direct matching: Determine whether the target in the target library directly appears in the query instruction. If there is a matching item, directly determine the target and its affiliated security type. Inference matching: If no directly matching target is found in the target library, the large language model is used to infer the implied target information and its affiliated security type in the query instruction. This embodiment gives the determination methods for different situations of time information, target information, and security types, improving the accuracy of the determination of the above information.
[0118] It should be further noted that before generating the interface call code using the interface call code generation model based on the query instruction and the target financial indicator interface document, it may also include:
[0119] Step 1: Construct an interface call code generation strategy based on the financial indicator interface document.
[0120] The purpose of this step is to construct an API (interface) call code generation strategy based on the financial indicator interface document, and its core goal is to provide users with executable API call codes with compliant parameter settings. The specific requirements are as follows: (1) For the parameters whose value ranges are clearly specified in the indicator API document, strictly select parameter values according to the corresponding ranges. (2) For the parameters whose value ranges are not clearly specified, such as date or target information, the following measures are taken: ① For date parameters, fill in a reasonable date before the current time. For example, when querying the closing price, ensure that the filled date is not a future date to prevent invalid queries. ② For target information, select the corresponding target according to the security category to which the financial indicator belongs, that is , where represents the name of the i-th target, represents the security code of the i-th target.
[0121] Step 2: Use the large language model to generate a callable interface call code based on the interface call code generation strategy and the prompt template.
[0122] This embodiment designs a prompt template that guides the large language model to generate reasonable API call code. This template can guide the large language model to generate reasonable API call code. This prompt template will, according to the requirements of the API specification document, ensure that the generated API call code complies with the regulations in terms of parameter selection and combination, so as to achieve automated and efficient generation of API call code. Figure 7 This is an example diagram of the generation process of a callable interface call code provided by an embodiment of the present invention. Starting from Figure 7 it can be seen that generating a callable interface call code needs to be generated based on the metric API document, the target library, and the prompt template. Figure 8 This is a schematic diagram of a prompt template for generating interface call code provided by an embodiment of the present invention.
[0123] Step 3: Execute each callable interface call code, obtain the return result of executing each interface call code, verify the return result, and retain the interface call code that passes the verification to obtain the verified interface call code.
[0124] This embodiment sequentially executes each callable interface call code generated by the large language model. Check the return result of each callable interface call code, and only retain those callable interface call codes that return the correct flag code. In this way, it is ensured that the finally retained interface call codes are all reasonable and effective. Figure 9 This is an example diagram of a verified interface call code provided by an embodiment of the present invention.
[0125] Step 4: According to the financial metric interface document, the verified interface call code, and the target example, guide the large language model to generate a query instruction that matches the verified interface call code.
[0126] This embodiment is based on the metric API document (financial metric interface document), the corresponding verified interface call code, and the target example , and guides the large language model to generate a natural language query (query instruction) that matches the verified interface call code. It is required that this query instruction can cover each parameter in the verified interface call code, that is, the information provided by this query can correspond to each parameter in the verified interface call code. Among them, for the default parameters in the verified interface call code, they can be omitted in the query. Finally, obtain the query instruction corresponding to the verified interface call code, , where represents the i-th query of the verified interface call code. The API call code query generation process is as Figure 10 shown, Figure 10Schematic diagram of a query instruction corresponding to an interface call code provided by an embodiment of the present invention. From Figure 10 It can be seen that it is necessary to input the API call code, the metric API document, and the target example into the large language model to obtain the API call code instruction (query instruction). Figure 11 Schematic diagram of a prompt template example of a query instruction corresponding to an interface call code provided by an embodiment of the present invention. For example, an example of the query instruction for the interface call code is "[Please check the current closing prices of XX Information Security and XX Co., Ltd. and help me check the latest closing price situation of XX Information Security and XX Co., Ltd.]".
[0127] Step 5: Generate an interface call code generation instruction dataset based on the financial metric interface document, the verified interface call code, and the matching query instruction.
[0128] This step is used to construct an interface call code generation instruction dataset (API call code generation instruction fine-tuning dataset). Its input content is the user query (query instruction) and the corresponding financial metric interface document (this metric document covers the metric name, the metric API signature, and the API parameters). The output content is the metric name and the interface call code. This dataset is designed to guide the model to generate the interface call code according to the provided query instruction and the financial metric interface document. Figure 12 Schematic diagram of a construction process example of an interface call code generation instruction dataset provided by an embodiment of the present invention. From Figure 12 It can be seen that the input of the API call code and the metric API document, and the output of the API call code metric (query instruction) together constitute the API call code generation instruction fine-tuning dataset (interface call code generation instruction dataset). It should be noted that the metric document, the metric API document, the financial metric API document, and the financial metric document mentioned in the text refer to the financial metric interface document. Figure 13 Schematic diagram of an example of an interface call code generation instruction dataset provided by an embodiment of the present invention.
[0129] Step 6: Use the interface call code generation instruction dataset to train the large language model to obtain an interface call code generation model.
[0130] This embodiment can use the generated interface call code generation instructions to generate a dataset for training a large language model, and obtain an interface call code generation model. This embodiment constructs an interface call code generation instruction dataset that can comprehensively cover all financial metrics. This interface call code generation instruction dataset not only covers the value ranges of various parameters, but also contains rich and diverse instruction expressions. These instructions fully reflect the abbreviations and common names commonly used by professionals, and at the same time comprehensively cover various semantics of the parameter values in the interface call code generation instruction dataset, which can ensure that the interface call code generation model can accurately understand and generate interface call codes that meet the actual needs of the financial field.
[0131] It should be further noted that, based on any of the above embodiments, in order to improve the accuracy of interface call code generation, after generating the interface call code using the interface call code generation model based on the query instruction and the target financial metric interface document, it may further include: performing rule detection on the interface call code to obtain the final interface call code; where the rule detection is used to avoid generating unreasonable parameters. This embodiment does not limit the specific method of rule detection. For example, verifying whether the required parameters exist, and directly intercepting if a key parameter is missing; checking whether the parameter type meets the expectations (such as numerical values, strings, boolean values, etc.); verifying the value range of numerical parameters.
[0132] The flowchart of an interface call code generation method provided by an embodiment of the present invention may include: S101, determining the target financial metric interface document based on the query instruction using a matching method; the matching method is a method of matching based on the targets in the metric library, or a method of performing inference queries using a large language model; S102, generating an interface call code using the interface call code generation model based on the query instruction and the target financial metric interface document; where the interface call code generation model is a model obtained by training a large language model based on an interface call code generation instruction dataset; the interface call code generation instruction dataset includes a financial metric interface document training dataset, a generated query instruction training dataset, and a generated interface call code training dataset. Compared with the current manual operation on the interface to select financial metrics, and since the financial metrics are relatively professional, the efficiency and accuracy of manual financial metric selection are relatively low, resulting in low efficiency and low accuracy of the generated API call code (interface call code), this application only needs to input a query instruction to obtain the target financial metric document corresponding to the query instruction, and then based on the target financial metric document and the query instruction, directly generate the API call code corresponding to the query instruction using the API call code generation model, thereby improving the efficiency and accuracy of API call code generation.
[0133] In current technical practices, the industry and academia generally adopt methods of vector retrieval or combining named entity recognition (NER) with the BM25 algorithm to handle API selection tasks in life scenarios or light production applications. These methods can effectively capture the coarse-grained semantic information in natural language and achieve API selection through vector similarity matching or direct lexical matching. However, when faced with the selection of financial indicator APIs, these methods face significant challenges. Financial indicator APIs not only have a large number, exceeding 50,000 in scale, but also have extremely high professionalism, which makes it difficult for conventional natural language semantic analysis methods to effectively distinguish them. Usually, the professional knowledge of domain experts is required for in-depth reasoning and analysis. The specific manifestations are as follows: (1) High professionalism: The semantic meanings of professional terms in the financial field are rich and complex. For example, "ChinaBond valuation" corresponds to the "valuation net price - ChinaBond" indicator in the bond market, and "RMB / USD exchange rate" corresponds to the "closing price" indicator in the foreign exchange market. Existing technologies have great difficulties in accurately identifying and effectively distinguishing such subtle semantic differences. (2) Existence of homonymous indicators: There may be homonymous market indicators for different types of securities. For example, various security types such as stocks, bonds, funds, and foreign exchange all have the "closing price" indicator. In practical applications, users tend to use established professional terms to describe security types. For example, "closing price of the Shanghai and Shenzhen stock markets" specifically refers to the "closing price" indicator in the A-share market, and "RMB / USD exchange rate" corresponds to the "closing price" indicator in the foreign exchange market. This semantic complexity far exceeds the processing capacity of vector retrieval and the combination of named entity recognition with the BM25 algorithm. (3) Multi-indicator parsing: In the financial field, the need to obtain data on multiple financial indicators simultaneously is relatively common. For the fine-grained multi-indicator retrieval requirement, vector retrieval technology can only capture the coarse-grained semantic information of the instruction and is difficult to accurately identify multiple specific indicators. The traditional solution based on named entity recognition (NER) combined with BM25 can effectively identify standardized indicator names, but facing the complexity and diversity of natural language instructions, its generalization ability is seriously insufficient. The forms of natural language expressions are rich and variable, and a large number of variants are beyond the coverage of traditional NER models. For example, for the instruction "high, open, low, and closing prices of a certain stock", the above traditional methods cannot accurately map it to the four specific indicators of the highest price, lowest price, opening price, and closing price, and cannot meet the user's demand for refined multi-indicator data.
[0134] In summary, to solve the problem of financial indicator selection, it is necessary to deeply analyze semantic information and conduct more in-depth reasoning by combining the professional logic of the financial field. Existing technical means have obvious shortcomings in dealing with the professionalism and complexity of financial scenarios, and there is an urgent need for more advanced methods to solve them.
[0135] For the convenience of understanding the present invention, please specifically refer to Figure 14 , Figure 14A flowchart example of a dataset generation method provided by an embodiment of the present invention may specifically include:
[0136] S201: Construct an interface call code generation strategy based on a financial indicator interface document, and use a first large language model to generate callable interface call codes based on the interface call code generation strategy and a prompt template.
[0137] It should be noted that both the first large language model and the second large language model in this embodiment are large language models with a scale of hundreds of billions of parameters. For example, Tongyi Qianwen, the gpt series of openAI, claude3.7; and large language models with a scale of hundreds of billions of parameters such as deepseek.
[0138] S202: Execute each callable interface call code, obtain the return results of executing each interface call code, verify the return results, and retain the interface call codes that pass the verification to obtain verified interface call codes.
[0139] S203: According to the financial indicator interface document, the verified interface call codes, and the target examples, guide the second large language model to generate query instructions that match the verified interface call codes.
[0140] S204: Generate an interface call code generation instruction dataset based on the financial indicator interface document, the query instructions, and the verified interface call codes; the interface call code generation instruction dataset is used to guide a third large language model to generate interface call codes according to the provided natural language queries and the financial indicator interface document.
[0141] S205: Use a fourth large language model to generate multiple query instructions based on the financial indicator interface document, and label the time information, financial indicator data, and target information in each generated query instruction to obtain labeled query instructions.
[0142] It should be noted that the first large language model, the second large language model, the third large language model, and the fourth large language model in this embodiment are all large language models with hundreds of billions of parameters. The query instructions in this step are for constructing subsequent inference and named entity recognition datasets, and the purpose of the query instructions in S204 is to train the model to generate codes. This embodiment strictly ensures that the indicator vocabulary (financial indicators) (indicator_terms), time (time), and target (institution) vocabulary labeled by the model are only obtained from the corresponding user queries (Query) to prevent the model from having hallucination phenomena.
[0143] S206: Based on the obtained labeled query instructions and the financial indicator interface document, construct a financial indicator inference dataset, which is used to implement the inference from natural language queries to financial indicators.
[0144] S207: Based on the obtained financial indicator data and query instructions, construct a multi-indicator analysis dataset, which is used to correspond natural query instructions with financial indicator data.
[0145] S208: Based on the obtained time information, target information and query instructions, construct a target and time information extraction dataset, which is used to extract target and time information in natural language queries.
[0146] In this embodiment, the construction process of the target and time information extraction dataset is as Figure 2 shown. The purpose of this dataset is to improve the recognition accuracy of the large model for target information and time information in user queries. In the input prompt template, set the meta-tokens as <<<time>>> and <<<financial institution / product / market>>> respectively to instruct the model to perform the extraction operations of time and target information. Figure 15 This is a sample diagram of a target and time information extraction dataset provided by an embodiment of the present invention.
[0147] S209: Construct an instruction fine-tuning dataset for the large language model, including an interface call code generation instruction dataset, an indicator inference dataset, a multi-indicator analysis dataset, and a target and time information extraction dataset.
[0148] The interface call code generation instruction dataset in this embodiment is used to train the large model to accurately generate API call codes according to user queries and indicator API documents; the indicator inference dataset aims to train the model to accurately infer the path and specific name of financial indicators from user queries; multi-indicator analysis dataset: This dataset focuses on processing multiple financial indicators involved in complex queries and improves the model's multi-indicator analysis ability; the target and time information extraction dataset aims to train the model to accurately extract target information and its corresponding time information from user queries. The interface call code generation instruction dataset, the indicator inference dataset, the multi-indicator analysis dataset, and the target and time information extraction dataset in this embodiment can train the same model, so that the model can have functions such as interface call code generation, indicator inference, indicator analysis, and target and time information extraction, or each dataset can also be used to train a large language model respectively. It can be understood that when the above four parts of the datasets are integrated for instruction fine-tuning training of the large language model. Through this training, the accuracy and generalization ability of the model in financial information processing tasks are improved. Figure 16 This is a schematic flowchart of a large language model training method provided by an embodiment of the present invention, starting from Figure 16It can be seen that when training a large model, various datasets can be used to train the large language model, and at the same time, the parameters of the large language model can be adjusted based on the loss function of the next token prediction.
[0149] S210: Construct a named entity recognition dataset based on the annotated query instructions; the named entity recognition dataset is used to identify the vocabulary describing financial metric information in the instructions.
[0150] In S205, the metric information in the query instructions is confirmed, that is, the vocabulary describing financial metric information in the query instructions has been located. This step annotates the vocabulary in the query instructions based on the query instructions and the located vocabulary describing financial metric information to construct a named entity recognition dataset. Embodiments of the present invention consider existing academic datasets, which mainly focus on the generation of RESTful API call codes in life scenarios and the generation of call codes for Hugging Face AI models. These datasets exhibit the following characteristics: (1) Limited number of APIs: Usually cover some relatively common API interfaces. (2) Few parameters: The number of parameters involved in the APIs is relatively small. (3) Limited parameter value range: The value range of the parameters is small. (4) Clear parameter meaning: The meaning of the numbers is clear and easy to understand and process. Due to the above characteristics, the LLM after instruction fine-tuning can usually perform well in these scenarios in tasks of parameter selection and filling. However, the complexity of the API call code generation scenario in the financial field far exceeds the coverage of existing datasets, mainly reflected in the following aspects: (1) Numerous parameters and wide value range: Financial APIs often have a large number of parameters, and the value range of each parameter is extremely wide. The processing of some parameters requires professional financial knowledge to accurately convert the natural language expression into the parameters of the API call code. For example, when processing "the post-rights-adjusted closing price of a certain company", it is necessary to accurately convert "post-rights-adjusted" into the parameter "101". (2) Parsing of time information and underlying information: In the process of generating API call codes for financial metrics, the parsing of time information and underlying information is a key link. The time information must be converted into a standard format, such as YYYY-MM-DD or YYYY-MM-DD hh:mm:ss; while the underlying information needs to be converted into a unique underlying code, for example, XX Bank corresponds to 0000XX.SZ. This further increases the difficulty and complexity of generating API call codes for financial APIs. In view of the above challenges, embodiments of the present invention have constructed a dataset for generating interface call codes that can comprehensively cover all financial metrics, a dataset for generating interface call code instructions, a dataset for metric reasoning, a dataset for multi-metric parsing, and a dataset for extracting underlying and time information.
[0151] The embodiments of the present invention innovatively propose a dataset generation strategy, aiming to effectively solve the problem of the lack of high-quality datasets in the field of fetching data from financial indicator APIs. Through constructing a complete and systematic data synthesis process, this strategy can automatically generate four types of datasets, specifically including datasets for generating financial indicator API call codes, datasets for indicator reasoning, multi-indicator parsing datasets, and datasets for extracting time and underlying information. The construction of these datasets not only successfully fills the data gap in the field of fetching data from financial indicator APIs but also provides solid and reliable data support for the financial indicator fetching system based on large language models.
[0152] For a better understanding of the present invention, please specifically refer to Figure 17 , Figure 17 which is a flow example diagram of a financial vector model training method provided by the embodiments of the present invention, and specifically may include:
[0153] S301: Determine the query instruction and its corresponding annotated indicator document, and find out the indicator documents semantically similar to the query instruction through retrieval.
[0154] In this embodiment, for each indicator document, it is characterized in the way of "indicator name + indicator definition + value range of indicator API parameters".
[0155] S302: Calculate the similarity between each query instruction and all indicator documents respectively by using two methods: vector cosine similarity and BM25 text similarity.
[0156] S303: Merge the K indicator documents with the highest vector cosine similarity and the K indicator documents with the highest BM25 text similarity to obtain the financial vector model training dataset.
[0157] When merging in this embodiment, duplicate indicator documents need to be removed to obtain strong negative samples (i.e., indicator documents that are easily confused with the original query). The purpose of this step is to construct positive and negative samples for each query instruction. The positive sample is the annotated indicator document (which can also be called the correct indicator interface document); the negative samples are constructed by two methods respectively, and the negative samples obtained by these two methods may be repeated, so deduplication is required. The financial vector model training dataset in this embodiment is the indicator documents (strong negative samples) that are easily confused with the original query instruction. The financial vector model training dataset aims to guide the financial vector model to accurately distinguish the fine-grained semantic differences between indicator documents with similar semantic content, thereby improving the accuracy of financial vector retrieval. Given a query instruction and its corresponding annotated indicator document ,find out the indicator documents semantically similar to the query through retrieval. Calculate the query There are two ways to calculate the semantic similarity with the index document d: (1) Vector cosine similarity: Use the vector model EMB to embed and d to obtain the corresponding vector representations and . Then calculate the cosine similarity between the two, ; (2) BM25 text similarity: Calculate the BM25 text similarity between the query and the index document d, that is, . Using the above two methods, calculate the similarity between the query and all index documents . Combine the top K index documents with the highest vector cosine similarity and the top K index documents with the highest BM25 text similarity. During this process, remove duplicate index documents and labeled index documents, and then obtain strong negative example samples (i.e., index documents that are easily confused with the original query), that is, . The finally obtained financial vector model training data set can be expressed as , where represents the i-th query, represents the corresponding labeled index document, represents the strong negative example sample. For ease of understanding, please refer to Figure 18 , Figure 18 is a flow example diagram of a method for constructing a financial vector model training data set provided by an embodiment of the present invention. As can be seen from Figure 18 , the entire process mainly includes the calculation of similarity, and determining strong negative sample index documents based on similarity as the financial vector model training data set.
[0158] S304: Train the financial vector model to be trained based on the financial vector model training data set to obtain a financial vector model.
[0159] It can be understood that when training based on the financial vector model training data set, the financial vector model is trained in a way of contrast learning based on strong negative examples, so as to accurately retrieve the corresponding financial indicators according to the user's query. Given a data batch, denoted as , . Among them, is sampled from the financial vector model training data set (strong negative samples) . First, calculate the pairwise vector cosine similarity between the query instructions and the index documents in this batch to obtain a similarity matrix , where and correspond to the vector embeddings of the query and the index document respectively. The corresponding label is Next, calculate each query and the corresponding index document as well as the similarity matrix between the strongly negative samples obtained by sampling, , the corresponding annotation is . The training loss function of this model is . Figure 19 FIG. 0000422 is a flowchart example of training a financial vector model based on a training data set of a financial vector model provided by an embodiment of the present invention. It can be seen from this figure that this training process is based on contrastive learning, and optimizes the vector model by constructing positive and negative samples, so that it can accurately distinguish the semantic associations between financial indicators. The process is divided into three stages: data construction, vector encoding, and loss calculation.
[0160] In the current financial data retrieval process, users need to perform a series of operations on the user interface (UI), including selecting specific financial indicators and inputting relevant parameters. After that, the system generates a financial indicator API call code according to the user's selection and input parameters to obtain the corresponding market quotation data. However, this process has certain drawbacks. First, the operations of users on the UI take a long time and are inefficient; second, it requires a high level of financial knowledge reserve and literacy of users. To optimize the user experience, the present invention proposes a method for querying financial indicator data based on natural language input. Under this method, users only need to input natural language instructions, and the system can automatically identify and extract the financial indicators involved in the instructions, and then generate the corresponding API call code. The specific process of this method includes the following steps:
[0161] 1. Indicator selection: The system needs to accurately screen out the required market quotation indicators from a huge indicator library with a scale of 50,000 to 1,000,000 according to the natural language instructions input by the user.
[0162] 2. Parameter parsing: The system parses the parameters included in the natural language instructions and generates the corresponding financial indicator API call code accordingly. These parameters involve financial professional knowledge, such as complex concepts such as rights adjustment processing. In addition, the system has the ability to efficiently parse time parameters and underlying asset information, and can accurately convert "the first quarter of last year" into a specific date such as "2024.03.31", and can also accurately map a company name such as "China XX" to its stock code "000001.SZ".
[0163] With this improvement, users can query financial indicator data in a simpler way, significantly improving operational efficiency while lowering the threshold for users' financial knowledge. The current solution is mainly applicable to RestfulAPI and AI model call scenarios in the field of life services. Its operation process can be divided into the following two stages: 1. Vector retrieval: The system uses vector retrieval technology to calculate the similarity between the user instruction vector and the API document vector, and screens out the top K API documents with the highest similarity. These APIs will be regarded as the options that best meet the user's instruction requirements. 2. Generation of API call code: The system generates corresponding API call code based on K API documents and the user's specific instructions. This process is designed to simplify the user's operational process for obtaining financial market data through efficient vector matching technology, while ensuring the accuracy and relevance of the called API.
[0164] In order to make the present invention easier to understand, please refer to Figure 20 , Figure 20 An example flow chart of an interface call code generation method provided in an embodiment of the present invention may specifically include:
[0165] S401, calling a time parsing tool to extract and parse the time information in the query instruction.
[0166] S402, if the time parsing tool fails to parse successfully, use the time parsing big model to convert the words representing time information in the query instruction into standard date and time; the time parsing big model is a model obtained by training the big language model based on the time information extraction data set.
[0167] In this embodiment, if the time information only includes the date, it is converted into the YYYY-MM-DD format; if the time information includes both the date and time, it is converted into the YYYY-MM-DD hh:mm:ss format.
[0168] S403, determining whether the subject in the subject database directly appears in the query instruction, and if there is a match, directly determining the subject and the type of securities to which it belongs.
[0169] S404, if no directly matching target is found in the target library, the target fuzzy matching large model is used to infer the target information implicit in the query instruction and the type of securities to which it belongs; the target fuzzy matching large model is a model obtained by training the large language model based on the target extraction data set.
[0170] This step determines the target information in the query instruction.
[0171] S405, using a named entity recognition model, identifying words related to financial indicators from the query instruction to obtain words related to financial indicators.
[0172] S406. Use the large financial information processing model to infer the quantity and names of the financial metrics involved in the query instruction.
[0173] The reason why the large financial information processing model in this embodiment can identify the quantity and names of the financial metrics is that the large language model is trained based on the multi-metric analysis dataset.
[0174] S407. Based on the financial metric-related vocabulary, the quantity and names of the financial metrics, perform a match in the financial metric library. If the match is successful, obtain the target financial metric interface document.
[0175] S408. If the match fails, then based on the financial metric-related vocabulary, the quantity and names of the financial metrics, use the large financial information processing model to perform a matching process to obtain the financial metric path and standard name corresponding to the query instruction.
[0176] The reason why the large financial information processing model in this embodiment can determine the path and name is that it is trained using the metric inference dataset.
[0177] S409. Based on the financial metric path and standard name, use the BM25 algorithm to retrieve in the financial metric library, and screen out the top K financial metrics with high similarity to the query metric to obtain the financial metric document.
[0178] S410. Use the financial vector model to determine the cosine similarity between the financial metric-related vocabulary and the financial metric document, and select the document with the highest cosine similarity as the target financial metric interface document.
[0179] S411. According to the query instruction and the target financial metric interface document, use the interface call code generation large model to generate the interface call code.
[0180] S412. Perform rule detection on the interface call code to obtain the final API call code.
[0181] The system corresponding to this embodiment is Figure 21 , Figure 21A system for generating interface call codes provided in an embodiment of the present invention ensures accurate identification and processing of time information, target information and financial indicators in user queries through multi-step parsing, matching, reasoning and detection, and finally generates reasonable and effective interface call codes to provide users with accurate financial indicator data information services. The natural language instructions in the figure correspond to the query instructions in the previous text, and the time word parsing is to parse the time information. LLM multi-indicator reasoning refers to inferring the path and name of the financial indicator, NER extracting the indicator name refers to named entity recognition, and LLM indicator parsing refers to parsing multiple financial indicators. Institutional matching refers to the use of a large language model with hundreds of billions of parameters to annotate the query instructions, which essentially extracts the vocabulary that describes the time information, institutional information, and indicator information in the query instructions.
[0182] The present invention innovatively proposes a financial indicator data acquisition system based on a large language model and traditional natural language processing technology. The system is committed to accurately identifying and processing time information, target information and financial indicators in user queries through a multi-step parsing, matching, reasoning and detection process. Specifically, the system uses advanced technical means to ensure the accuracy of information processing in multiple links such as time information parsing, target information matching, financial indicator vocabulary recognition, indicator quantity and name reasoning, indicator precise matching and processing, indicator retrieval and screening, indicator document determination, indicator API call generation and API call hallucination detection. Through the above innovative measures, the financial indicator data acquisition system of the present invention can provide users with high-precision financial information services, effectively improve the accuracy and efficiency of financial indicator API data acquisition, and show significant advantages and application value in the field of financial information processing. The method proposed by the present invention can significantly improve the system's ability to understand natural language financial queries and enhance its ability to cope with complex business logic. The solution has broad application prospects in professional fields such as financial market analysis and can meet users' needs for efficient and accurate financial data acquisition.
[0183] 1. Technical problems that can be solved by the present invention: (1) Efficient modeling optimization of large-scale API call code generation for large language models. (2) Extraction of multiple financial indicator data from natural language. (3) Improvement and optimization of the accuracy of financial indicator data extracted from natural language. (4) Realization of automated operation to reduce manual intervention. (5) Configurability, strong controllability and high scalability.
[0184] 2. Business Value:
[0185] The present invention has achieved significant performance improvement in the field of natural language data acquisition of financial products and background data acquisition of financial software development projects, effectively reducing the threshold for product use and the difficulty of project development, and fully demonstrating its important business value and broad application prospects, which are mainly reflected in the following two aspects:
[0186] 1. iFinD A iFinD Natural Language Data Retrieval: (1) Introduce natural language query functionality in Excel add-ins, iFinD clients, and data interface products, enabling users to accurately express their specific data requirements in natural language. Compared with traditional query methods, this solution has the following outstanding advantages: (2) Simplify the operation process: Users no longer need to layer through the indicator directory via the graphical user interface, significantly streamlining the operation steps and significantly improving data acquisition efficiency. (3) Lower the usage threshold: This solution has relatively low requirements for users' professional financial knowledge reserves and operation proficiency. Even ordinary users can easily query the required data, greatly expanding the coverage of the user group.
[0187] 2. AIGC Market Data Retrieval - Extreme Project-Level Code Generation:
[0188] (1) This solution provides developers with the function of generating data retrieval code based on natural language queries. Developers only need to describe their needs in natural language to quickly generate the corresponding data retrieval code, without having to spend a lot of time and effort looking up numerous data source documents. This function effectively reduces the difficulty of development work and significantly improves the work efficiency of developers.
[0189] The following introduces the interface call code generation device provided by the embodiments of the present invention. The interface call code generation device described below can be correspondingly referred to the interface call code generation method described above.
[0190] For details, please refer to Figure 22 , Figure 22 which is a schematic structural diagram of an interface call code generation device provided by an embodiment of the present invention and may include:
[0191] A target financial indicator interface document determination module 100, configured to determine a target financial indicator interface document based on a query instruction using a matching method; the matching method is a method of matching based on the underlying assets in the indicator library or a method of performing inference queries using a large language model;
[0192] An interface call code generation module 200, configured to generate an interface call code based on the query instruction and the target financial indicator interface document using an interface call code generation model; wherein, the interface call code generation model is a model obtained by training a large language model based on an interface call code generation instruction data set; the interface call code generation instruction data set includes a financial indicator interface document training data set, a generated query instruction training data set, and a generated interface call code training data set.
[0193] Further, based on the above embodiments, the interface call code generation module 200 may include:
[0194] A time and target recognition unit, configured to convert the time information in the query instruction into a set standard date and time, and recognize the target in the query instruction and the type of securities to which it belongs;
[0195] An interface call code generation unit, configured to generate the interface call code by using the interface call code generation model based on the query instruction, the target financial indicator interface document, the standard date and time, and the target and the type of securities to which it belongs.
[0196] Further, based on the above embodiments, the above time and target recognition unit includes:
[0197] A time parsing and judgment sub-unit, configured to parse the time information in the query instruction based on a time parsing tool to determine whether the parsing is successful;
[0198] A parsing sub-unit based on a time parsing large model, configured to, when the parsing fails, use the time parsing large model to convert the vocabulary representing the time information in the query instruction into the standard date and time; the time parsing large model is a model obtained by training a large language model based on a time information extraction data set;
[0199] A target judgment sub-unit, configured to judge whether the target expression in the query instruction matches the target in the target library;
[0200] A target determination sub-unit, configured to, if it matches, determine the target and the type of securities to which it belongs based on the target library;
[0201] A single target determination sub-unit based on a large model, configured to, if it does not match, infer the target and the type of securities to which it belongs included in the query instruction based on the target fuzzy matching large model; wherein, the target fuzzy matching large model is a model obtained by training a large language model based on a target extraction data set.
[0202] Further, based on any of the above embodiments, the target financial indicator interface document determination module 100 may include:
[0203] A financial indicator related vocabulary determination unit, configured to recognize the vocabulary related to financial indicators from the query instruction by using a named entity recognition model to obtain the financial indicator related vocabulary;
[0204] A quantity and name determination unit for performing inference using a financial information processing large model to determine the quantity and name of financial indicators in the query instruction; the financial information processing large model is a model obtained by training a large language model based on an indicator inference dataset and a multi-indicator parsing dataset; the indicator inference dataset is a dataset for determining the indicator path and indicator name corresponding to the query instruction, and the multi-indicator parsing dataset is a dataset for determining multiple financial indicators in the query instruction;
[0205] A first target financial indicator interface document determination unit for performing matching in a financial indicator library based on the financial indicator-related vocabulary, the quantity and name of the financial indicators. If the matching is successful, the target financial indicator interface document is obtained;
[0206] A path and name determination unit for, if the matching is unsuccessful, performing matching processing using the financial information processing large model based on the financial indicator-related vocabulary, the quantity and name of the financial indicators to obtain a financial indicator path and a standard name;
[0207] A financial indicator document determination unit for retrieving in the financial indicator library based on the financial indicator path and the standard name using the best-matching text similarity, screening out the top K financial indicators with high similarity to the query instruction to obtain the financial indicator document;
[0208] A second target financial indicator interface document determination unit for determining the similarity between the financial indicator-related vocabulary and the financial indicator document using a financial vector model, and selecting the document with the highest similarity as the target financial indicator interface document.
[0209] Further, based on the above embodiment, the above interface call code generation device may further include:
[0210] A named entity recognition dataset construction module for constructing a named entity recognition dataset; the named entity recognition dataset is used to identify the vocabulary describing financial indicators in the query instruction;
[0211] A model training module for training a to-be-trained named entity recognition model based on the named entity recognition dataset to obtain a trained named entity recognition model;
[0212] A named entity recognition model construction module for combining the trained named entity recognition model and the large language model to obtain the named entity recognition model.
[0213] Further, based on the above embodiment, the above interface call code generation device may further include:
[0214] A financial vector model training dataset construction module, which is used to construct a financial vector model training dataset based on the metric inference dataset and the correct metric interface document corresponding to each query instruction determined by using similarity methods such as vector cosine similarity and best-matching text similarity; the financial vector model training dataset is used to distinguish semantic differences between financial metric interface documents with similar semantic content;
[0215] A financial vector model training module, which is used to train a financial vector model to be trained based on the financial vector model training dataset to obtain the financial vector model; the financial vector model is used to determine financial metrics corresponding to query instructions.
[0216] Further, based on any of the above embodiments, the above interface call code generation device may further include:
[0217] A labeled query instruction determination module, which is used to generate multiple query instructions based on financial metric interface documents by using a large language model, and label the time information, financial metric data, target information, and time information in each generated query instruction to obtain labeled query instructions;
[0218] A metric inference dataset construction module, which is used to construct the metric inference dataset based on the labeled query instructions and the financial metric interface documents, and the metric inference dataset is used to implement inference from query instructions to financial metrics.
[0219] Further, based on any of the above embodiments, the above interface call code generation device may further include:
[0220] A multi-metric parsing dataset construction module, which is used to construct the multi-metric parsing dataset based on the financial metric data and the multiple query instructions, and the multi-metric parsing dataset is used to correspond query instructions to financial metric data.
[0221] Further, based on any of the above embodiments, the above interface call code generation device may further include:
[0222] An interface call code generation strategy construction module, which is used to construct an interface call code generation strategy based on financial metric interface documents;
[0223] A callable interface call code determination module, which is used to generate callable interface call codes based on the interface call code generation strategy and a prompt template by using a large language model;
[0224] A verification module, which is used to execute each callable interface call code to obtain the return result of executing each interface call code, verify the return result, and retain the interface call codes that pass the verification to obtain verified interface call codes;
[0225] According to the financial indicator interface document, the verified interface call code, and the target example, guide the large language model to generate a query instruction that matches the verified interface call code;
[0226] An interface call code generation instruction dataset determination module, configured to generate an interface call code generation instruction dataset based on the financial indicator interface document, the verified interface call code, and the matching query instruction;
[0227] An interface call code generation model generation module, configured to train the large language model using the interface call code generation instruction dataset to obtain the interface call code generation model.
[0228] Further, based on any of the above embodiments, the above interface call code generation device may further include:
[0229] A rule verification module, configured to perform rule detection on the interface call code to obtain a final interface call code; wherein, the rule detection is used to avoid generating unreasonable parameters.
[0230] It should be noted that the order of the modules and units in the above interface call code generation device can be changed before and after without affecting the logic.
[0231] An interface call code generation device provided by an embodiment of the present invention may include: a target financial indicator interface document determination module 100, configured to determine a target financial indicator interface document based on a query instruction using a matching method; the matching method is a method of matching based on the targets in the indicator library or a method of performing inference queries using a large language model; an interface call code generation module 200, configured to generate an interface call code based on the query instruction and the target financial indicator interface document using an interface call code generation model; wherein, the interface call code generation model is a model obtained by training a large language model based on an interface call code generation instruction dataset; the interface call code generation instruction dataset includes a financial indicator interface document training dataset, a generated query instruction training dataset, and a generated interface call code training dataset. Compared with the current manual operation of selecting financial indicators on the interface, and since the financial indicators are relatively professional, the efficiency and accuracy of manual financial indicator selection are relatively low, resulting in low efficiency and low accuracy of the generated API call code (interface call code), in this application, only by inputting a query instruction, the target financial indicator document corresponding to the query instruction can be obtained, and then based on the target financial indicator document and the query instruction, the API call code corresponding to the query instruction can be directly generated using the API call code generation model, thereby improving the efficiency and accuracy of API call code generation.
[0232] The following introduces an interface call code generation device provided by an embodiment of the present invention. The interface call code generation device described below can be correspondingly referred to the interface call code generation method described above.
[0233] Please refer to Figure 23 , Figure 23 which is a schematic structural diagram of an interface call code generation device provided by an embodiment of the present invention, and may include:
[0234] A memory 10 for storing computer programs;
[0235] A processor 20 for executing the computer program to implement the above-mentioned interface call code generation method.
[0236] The memory 10, the processor 20, and the communication interface 30 all complete mutual communication through a communication bus 40.
[0237] In an embodiment of the present invention, the memory 10 is used to store one or more programs, and the program may include program code, and the program code includes computer operation instructions. In an embodiment of the present invention, the memory 10 may store a program for implementing the function of the interface call code generation method.
[0238] In a possible implementation manner, the memory 10 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function, etc.; the data storage area may store data created during use.
[0239] In addition, the memory 10 may include a read-only memory and a random access memory, and provide instructions and data to the processor. A part of the memory may also include NVRAM. The memory stores an operating system and operation instructions, executable modules, or data structures, or subsets thereof, or extended sets thereof. Among them, the operation instructions may include various operation instructions for implementing various operations. The operating system may include various system programs for implementing various basic tasks and processing hardware-based tasks.
[0240] The processor 20 may be a central processing unit (CPU), an application-specific integrated circuit, a digital signal processor, a field programmable gate array, or other programmable logic devices. The processor 20 may be a microprocessor or any conventional processor, etc. The processor 20 may call the program stored in the memory 10.
[0241] The communication interface 30 may be an interface of a communication module for connecting to other devices or systems.
[0242] Of course, it should be noted that Figure 23 the structure shown does not constitute a limitation on the interface call code generation device in the embodiments of the present invention. In actual applications, the interface call code generation device may include more or fewer components than Figure 23 those shown, or combine certain components.
[0243] Next, the computer-readable storage medium provided by the embodiments of the present invention will be introduced. The computer-readable storage medium described below can be correspondingly referred to the interface call code generation method described above.
[0244] The present invention also provides a readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned interface call code generation method are implemented.
[0245] The readable storage medium may include various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc.
[0246] In this specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple. For the relevant parts, reference can be made to the descriptions in the method part.
[0247] Those skilled in the art can further realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been generally described according to their functions in the above description. Whether these functions are executed in the form of hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0248] Finally, it should also be noted that in this article, relationships such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device.
[0249] The above has introduced in detail a method, apparatus, device and readable storage medium for generating interface call codes provided by the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A method for generating interface call code, characterized in that, Including: Determine the target financial indicator interface document using a matching method based on the query instruction; The matching method is a method of matching based on the targets in the indicator library or a method of inferring and querying using a large language model; Based on the query instruction and the target financial indicator interface document, use an interface call code generation model to generate interface call code; wherein, the interface call code generation model is a model obtained by training a large language model based on an interface call code generation instruction dataset; the interface call code generation instruction dataset includes a financial indicator interface document training dataset, a generated query instruction training dataset, and a generated interface call code training dataset.
2. The interface call code generation method according to claim 1, characterized in that Based on the query instruction and the target financial indicator interface document, using an interface call code generation model to generate interface call code includes: Convert the time information in the query instruction into a set standard date and time, and identify the target and its affiliated security type in the query instruction; Based on the query instruction, the target financial indicator interface document, the standard date and time, and the target and its affiliated security type, use the interface call code generation model to generate the interface call code.
3. The interface call code generation method according to claim 2, wherein, Converting the time information in the query instruction into a set standard date and time, and identifying the target and its affiliated security type in the query instruction includes: Based on a time parsing tool, parse the time information in the query instruction to determine whether the parsing is successful; When the parsing fails, use a time parsing large model to convert the vocabulary representing the time information in the query instruction into the standard date and time; the time parsing large model is a model obtained by training a large language model based on a time information extraction dataset; Judge whether the target expression in the query instruction matches the targets in the target library; If it matches, determine the target and its affiliated security type based on the target library; If it does not match, infer the target and its affiliated security type included in the query instruction based on the target fuzzy matching large model; wherein, the target fuzzy matching large model is a model obtained by training a large language model based on a target extraction dataset.
4. The interface call code generation method according to claim 1, wherein Determine the target financial indicator interface document using a matching method based on the query instruction, including: Use a named entity recognition model to identify the financial indicator-related vocabulary from the query instruction to obtain financial indicator-related vocabulary; Use a financial information processing large model to infer and determine the quantity and name of the financial indicators in the query instruction; the financial information processing large model is a model obtained by training a large language model based on an indicator inference dataset and a multi-indicator parsing dataset; the indicator inference dataset is a dataset for determining the indicator path and indicator name corresponding to the query instruction, and the multi-indicator parsing dataset is a dataset for determining multiple financial indicators in the query instruction; Based on the financial indicator-related vocabulary, the quantity and name of the financial indicators, perform a match in the financial indicator library. If the match is successful, obtain the target financial indicator interface document; If the match fails, based on the financial indicator-related vocabulary, the quantity of the financial indicator, and its name, use the large financial information processing model to perform a matching process to obtain a financial indicator path and a standard name; Based on the financial indicator path and the standard name, use the best-matching text similarity to retrieve in the financial indicator library, and screen out the top K financial indicators with high similarity to the query instruction to obtain a financial indicator document; Use the financial vector model to determine the similarity between the financial indicator-related vocabulary and the financial indicator document, and select the document with the highest similarity as the target financial indicator interface document.
5. The interface call code generation method according to claim 4, wherein, Before using the named entity recognition model to identify the vocabulary related to financial indicators from the query instruction to obtain the financial indicator-related vocabulary, it further includes: Construct a named entity recognition data set; the named entity recognition data set is used to identify the vocabulary describing financial indicators in the query instruction; Based on the named entity recognition data set, train the named entity recognition model to be trained to obtain a trained named entity recognition model; Combine the trained named entity recognition model with the large language model to obtain the named entity recognition model.
6. The interface call code generation method according to claim 4, wherein Before using the financial vector model to determine the similarity between the financial indicator-related vocabulary and the financial indicator document, and selecting the document with the highest similarity as the target financial indicator interface document, it further includes: Based on the indicator inference data set, use the similarity methods of vector cosine similarity and best-matching text similarity to determine the correct indicator interface document corresponding to each query instruction, and construct a financial vector model training data set; the financial vector model training data set is used to distinguish the semantic differences between financial indicator interface documents with similar semantic content; Based on the financial vector model training data set, train the financial vector model to be trained to obtain the financial vector model; the financial vector model is used to determine the financial indicator corresponding to the query instruction.
7. The interface call code generation method according to claim 4, wherein The construction process of the indicator inference data set includes: Based on the financial indicator interface document, use the large language model to generate multiple query instructions, and label the time information, financial indicator data, target information, and time information in each generated query instruction to obtain labeled query instructions; Based on the labeled query instructions and the financial indicator interface document, construct the indicator inference data set, and the indicator inference data set is used to realize the inference from the query instruction to the financial indicator.
8. The interface call code generation method according to claim 7, wherein The construction process of the multi-indicator analysis data set includes: Based on the financial indicator data and the multiple query instructions, construct the multi-indicator analysis data set, and the multi-indicator analysis data set is used to correspond the query instruction to the financial indicator data.
9. The interface call code generation method according to any one of claims 1 to 8, characterized in that Before generating the interface call code using the interface call code generation model based on the query instruction and the target financial indicator interface document, it further includes: Construct an interface call code generation strategy based on the financial indicator interface document; Based on the interface call code generation strategy and the prompt template, use the large language model to generate a callable interface call code; Execute each callable interface call code, obtain the return result of executing each interface call code, verify the return result, and retain the interface call code that passes the verification to obtain the verified interface call code; Based on the financial metric interface document, the verified interface call code, and the target example, guide the large language model to generate query instructions that match the verified interface call code; Generate an interface call code generation instruction dataset based on the financial metric interface document, the verified interface call code, and the matching query instructions; Use the interface call code generation instruction dataset to train the large language model to obtain the interface call code generation model.
10. The interface call code generation method according to claim 1, characterized in that, After generating the interface call code using the interface call code generation model based on the query instructions and the target financial metric interface document, it further includes: Perform rule detection on the interface call code to obtain the final interface call code; wherein, the rule detection is used to avoid generating unreasonable parameters.
11. An interface call code generation device, characterized in that It includes: A target financial metric interface document determination module for determining the target financial metric interface document based on the query instructions using a matching method; The matching method is a method of matching based on the targets in the metric library or a method of performing inference queries using a large language model; An interface call code generation module for generating interface call code based on the query instructions and the target financial metric interface document using the interface call code generation model; wherein, the interface call code generation model is a model obtained by training the large language model based on the interface call code generation instruction dataset; the interface call code generation instruction dataset includes a financial metric interface document training dataset, a generated query instruction training dataset, and a generated interface call code training dataset.
12. An interface call code generation device, characterized in that, It includes: A memory for storing computer programs; A processor for executing the computer program to implement the steps of the interface call code generation method according to any one of claims 1 to 10.
13. A readable storage medium, characterized in that, The computer program is stored on a readable storage medium, and when the computer program is executed by the processor, it implements the steps of the interface call code generation method according to any one of claims 1 to 10.