Query statement generation method and device based on large model, medium and equipment

By inputting business logic processes into the large language model (LLM), identifying and completing implicit steps and logical relationships, and generating query statements, the problem of unused SOP deep information in the existing technology is solved, and automated execution and efficiency improvement are achieved.

CN120541191AActive Publication Date: 2025-08-26ZHEJIANG ANT MISUAN TECHNOLOGY CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511036717.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2025-08-26
Estimated Expiration
2045-07-25

AI Technical Summary

Technical Problem

The application of standard operating procedures (SOP) in the prior art is still limited to primary text processing, and the inability to fully explore its deep semantic information, resulting in the inability to efficiently utilize the knowledge in SOP.

Method used

By inputting the business logic process described in natural language into a pre-trained large language model (LLM), identifying and completing the implicit steps and logical relationships in the business logic process, a query statement is generated for the business data structure.

Benefits of technology

It realizes the mining of deep semantic information of SOP, generates query statements that can be executed automatically, reduces manual intervention, and improves the efficiency and accuracy of business execution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120541191A_ABST
    Figure CN120541191A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a query statement generation method based on a large model, which comprises the following steps of: inputting a business logic process described by a natural language into a pre-trained LLM (Logical Language Model), and identifying each step included in the business logic process and a logic relationship among the steps through the LLM; according to the method, a business logic process is established, hidden steps in the business logic process are complemented, all the steps in the complete business logic process and the logic relation between the steps are determined, and finally a corresponding query statement is generated for a business data structure. According to the method, the business logic process is input into the LLM, so that the implicit steps in the business logic process are generated, the complete business logic process is obtained, corresponding query statements can be continuously generated through the LLM, utilization of the business logic process is not limited to a basic word processing stage any more, and the user experience is improved. However, implicit semantics can be further analyzed, and query statements which can be directly executed are generated to support business execution.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of computer technology, and in particular to a method, apparatus, storage medium, and device for generating query statements based on a large model. Background Art

[0002] A Standard Operating Procedure (SOP) is a document that details workflows and procedures. Typically, an SOP encompasses the detailed process for a specific business, product, or service, providing users with a reference and operational basis. SOPs, built through accumulated experience and knowledge gained through business operations, contain a wealth of valuable experience and knowledge. Therefore, effectively utilizing SOPs is a crucial aspect of their application.

[0003] In existing technologies, the application of SOPs is still at the primary stage of text processing. For example, by categorizing SOPs, the corresponding SOP can be quickly called up for different situations. Alternatively, keywords can be extracted from SOPs to automatically populate documents such as work orders and spreadsheets. Clearly, existing technologies do not fully exploit the deep semantic information contained in SOPs.

[0004] Based on this, how to more efficiently utilize the knowledge in SOP has become an urgent problem to be solved. Therefore, this manual provides a method for generating query statements based on a large model. Summary of the Invention

[0005] The embodiments of this specification provide a query statement generation method, device, storage medium and electronic device based on a large model to partially solve the problems existing in the above-mentioned prior art.

[0006] The embodiments of this specification adopt the following technical solutions: This specification provides a query statement generation method based on a large model, the method comprising: Obtain business logic processes described in natural language; Input the business logic process into a pre-trained large language model LLM; Identifying each step included in the business logic process through the LLM, determining the implicit steps in the business logic process, and determining the logical relationship between the obtained steps; Based on the business data structure of the business corresponding to the business logic process, a query statement executed on the business data structure is generated according to the obtained steps and the logical relationship, and the query statement is used to implement the business logic process.

[0007] This specification provides a query statement generation device based on a large model, the device comprising: An acquisition module is used to acquire the business logic process described in natural language; An input module, configured to input the business logic process into a pre-trained large language model LLM; an identification module, configured to identify the steps included in the business logic process through the LLM, determine the implicit steps in the business logic process, and determine the logical relationships between the obtained steps; A generation module generates a query statement executed on the business data structure based on the business data structure of the business corresponding to the business logic process and according to the obtained steps and the logical relationship, wherein the query statement is used to implement the business logic process.

[0008] This specification provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned query statement generation method based on the large model.

[0009] This specification provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the method for generating query statements based on a large model is implemented.

[0010] At least one of the above technical solutions adopted in the embodiments of this specification can achieve the following beneficial effects: The embodiments of this specification disclose a method for generating query statements based on a large model. This method inputs a business logic process described in natural language into a pre-trained LLM, uses the LLM to identify the steps contained in the business logic process and the logical relationships between the steps, completes the implicit steps in the business logic process, determines the steps in the complete business logic process and the logical relationships between the steps, and finally generates a corresponding query statement based on the business data structure targeted by the query. This method generates the implicit steps in the business logic process by inputting the business logic process into the LLM, obtaining a complete business logic process, allowing the corresponding query statement to be generated through the LLM. This makes it possible to further generate the corresponding query statement through the LLM, so that the utilization of the business logic process is no longer limited to the basic text processing stage, but can further analyze the implicit semantics therein and generate directly executable query statements to support business execution. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The drawings described herein are used to provide a further understanding of this specification and constitute a part of this specification. The exemplary embodiments and descriptions of this specification are used to explain this specification and do not constitute an improper limitation of this specification. In the drawings: Figure 1 A flowchart of query statement generation based on a large model provided in an embodiment of this specification; Figure 2 A schematic diagram provided in the embodiments of this specification to supplement the implicit steps; Figure 3 A schematic diagram of a query statement generation device based on a large model provided in an embodiment of this specification; Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this specification. DETAILED DESCRIPTION

[0012] To make the objectives, technical solutions, and advantages of this specification more clear, the following will clearly and completely describe the technical solutions of this specification in conjunction with the specific embodiments of this specification and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of this specification.

[0013] The technical solutions provided by the embodiments of this specification are described in detail below with reference to the accompanying drawings.

[0014] Figure 1 The flowchart for generating a query statement based on a large model provided in the embodiment of this specification specifically includes the following steps: S100: Acquire a business logic process described in natural language.

[0015] In the embodiments of this specification, Figure 1 The device for generating the query statement in the method shown can be any electronic device, such as a computer, a server, or a server cluster composed of multiple servers, etc. For the convenience of description, the following description only takes the server as an example.

[0016] In order to generate a query statement for executing a certain business, the server can first obtain the business logic process of the business. Specifically, the business logic process described in this specification includes a standard operating procedure (SOP). The SOP is a standard operating procedure for the business described in natural language. It is a document that describes in detail the work, workflow, and steps in the business. It is a standard process specification within the organization that provides the business for reference and operation by internal personnel of the organization. In addition to containing the standardized process of the business, the SOP also includes clear responsibilities and authorities of each role in the business, emergency measures, quality standards and assessments, document management, and other content.

[0017] Since the business logic process described in the natural language of this specification includes SOP, the business logic process includes the steps of business operations and the logical relationship between the steps. By processing the subsequent steps of the business logic process, a query statement for implementing the business logic process can be generated. By executing the query statement, the utilization of the business logic process and SOP can be improved, and the deep semantic information in the SOP can be fully utilized, rather than simply performing text processing. This makes the business logic process and SOP no longer just specification documents to guide users.

[0018] In this specification, the specific business logic process is not limited. For example, the business logic process is a marketing business, which includes the SOP of the marketing process for users, or the business logic process of project management, which includes the SOP of cross-departmental collaboration.

[0019] For the convenience of description, the following embodiments of this specification only take the business logic process of the risk control business as an example to illustrate. That is to say, the query statement generation method provided in this specification is described by taking the above-mentioned business as the risk control business as an example.

[0020] When the above-mentioned business is a risk control business based on the recorded risk log, the business logic process of the risk control business (that is, the risk control logic process) can be roughly illustrated as follows: The first step is to obtain the target field; The second step is to check which danger level the corresponding value of the target field falls into; The third step is to send corresponding attack alarm information according to the danger level.

[0021] Of course, the above is just an example. In actual business logic, the operations to be performed in each step are generally described in more specific natural language. For example, the second step may include checking the following fields: the frequency of operations on the same device / IP; if it is greater than 5 times / minute, it is classified as medium risk; the payment limit for new accounts; if the payment limit is greater than 5,000 yuan, it is classified as high risk; and checking the number of registered accounts for the same device / IP; if it is greater than 3, it is classified as low risk.

[0022] As can be seen, the process involving the first step of obtaining the target field and the second step of determining the risk level is equivalent to querying business data. When implemented through a query statement, this process may include multiple data query steps. For example, it is necessary to first obtain the operation data of each account to determine the active accounts, then obtain the account data of each active account to determine the recently registered active accounts, and then obtain other business data to determine the accounts with a transaction volume greater than 5,000 as the query results falling into the high-risk level.

[0023] The risk log is a risk log of other businesses recorded by the user when performing other businesses, and the risk control business is a business that requires risk control of the other businesses based on the risk log.

[0024] S102: Input the business logic process into a pre-trained large language model LLM.

[0025] After the server obtains the business logic process described above, it can input the natural language description of the business logic process into a pre-trained Large Language Model (LLM). This LLM can be pre-deployed on the server, or it can be deployed on other devices. If deployed on other devices, the server can send the business logic process to the other devices, allowing them to input the received business logic process into the LLM. The following description uses the LLM deployed on the server as an example.

[0026] Of course, since business logic processes are typically set based on manual experience, to avoid omissions or logical errors in the business logic process, in the embodiments of this specification, the server directly outputs the actual business logic process acquired in step S100 to the LLM. Based on this business logic process, the LLM identifies the implicit steps in the business logic process and then generates a query statement for implementing the business logic process. This prevents queries generated directly through the LLM from being executed or encountering conflicts during execution, thereby improving the accuracy and reliability of LLM-generated queries.

[0027] S104: Identify each step included in the business logic process through the LLM, determine the implicit steps in the business logic process, and determine the logical relationship between the obtained steps.

[0028] Generally, a business logic process consists of several steps following certain logical relationships. For example, in the example above, executing each step sequentially determines the type of attack alert to send. Translated into a data query process, this involves filtering business data using query statements based on different query conditions. More complex business logic processes may involve multiple steps with varying logical relationships. For example, high-frequency and high-value transactions for new users are further filtered based on the intermediate results of the new user query.

[0029] It can be seen from this that the data query process included in a business logic process can be regarded as consisting of the logical relationship between the steps of each query process and the specific content of executing each step. Therefore, to generate a query statement for implementing the business logic process, LLM needs to identify the steps included in the business logic process and the logical relationship between the steps based on the business logic process described in natural language.

[0030] Specifically, the LLM generates the query statements required to implement the business logic process in two stages. In the first stage, the LLM identifies the content related to data queries in the business logic process, determines the query steps directly described in the business logic process, and identifies the logical relationships between each step. In the second stage, based on each step and the logical relationships between them, the LLM identifies the implicit steps in the business logic process. Finally, the text of the implicit steps is generated, and the implicit steps and the logical relationships between them are determined.

[0031] Because business logic involves data querying, it may only describe the core content that needs to be executed, without directly including the implicit query steps or query conditions. This is because in some business scenarios, business logic processes serve as a standard for manual operations, and therefore are not designed to account for machine recognition. This may result in the omission of conditions that are common sense to users or necessary for business purposes.

[0032] To ensure the accuracy and reliability of generated query statements, the server identifies the data query content and determines the steps directly specified in the business logic process and the logical relationships between them. It then analyzes the inputs and outputs of the directly specified steps and the logical relationships between them to determine the steps implicit in those logical relationships. Finally, based on the directly identified steps and the identified implicit steps, it performs logical reasoning to determine the logical relationships between all steps.

[0033] For example, suppose the risk control SOP states, "Conduct a loan check on the applicant: If the number of inquiries in the past month exceeds five, the application will be automatically rejected." Clearly, this requires querying the applicant's loan check results. However, a potentially non-explicit element is the query of a third-party credit information cache pool. Based on common sense, in addition to "internal" loan inquiries, external credit information data is also required, not just based on internal business data. However, since business logic processes are generally executed "internally," routine external queries are not directly documented within the business logic process. The SOP for the business logic process identifies the following steps: Step 1: Query internal business data to determine the number of inquiries in a month; Step 2: Identify users with more than five inquiries; and the implicit step 4 could be: Query the external credit information of users with no more than five inquiries in the credit database; and finally, Step 3: Based on the two query results, determine the query results for the users who ultimately require risk control.

[0034] Figure 2 The schematic diagram provided for this specification is to supplement the implicit steps, wherein the SOP on the left indicates that the SOP in the business logic process is the object of step analysis, which contains the text content of the SOP. Steps 1 to 3 are the step processes and step logical relationships directly determined according to the SOP, and step 4 is a supplementary implicit step. This step requires the query result of step 1 as input, and after continued execution, it outputs the content that supports the final query result. Step 3 is executed based on the output results of steps 2 and 4. Step 4 and the logical relationship between step 4 and other steps are represented by dotted lines, indicating that the relevant content is supplemented through the reasoning of LLM. As shown in the figure, the arc on the left side of the LLM is marked with recognition, indicating that steps 1 to 3 are obtained by identifying the SOP, and the arc on the right side is marked with supplement, indicating that step 4 and the supplementary logical relationship are supplemented to the steps. In addition, since identifying the steps of the data query included in the business logic process and the logical relationships between the steps may also require a knowledge base corresponding to the business logic process, in the embodiments of this specification, the knowledge base corresponding to the business logic process can be pre-injected into the pre-trained LLM. The knowledge base may include at least one of the following: The knowledge graph corresponding to this business logic process includes the conceptual description of each entity required in the business logic process (used to describe what the entity is, such as the concepts of user and number of operations in the above example), the attributes of each entity (used to describe all possible attributes of the entity in this business logic), the relationship between each entity, etc.

[0035] After injecting the above-mentioned knowledge base containing the knowledge graph corresponding to the business logic process into the pre-trained LLM, when the server inputs the business logic process described in natural language to the LLM, the LLM can identify the semantics of the business logic process described in natural language based on its own reasoning ability, and map the semantics of the data query process involved in the business logic process to determine the concepts contained therein as entities in the knowledge graph, and then split the data query process into several specific query steps based on the semantics, and finally continue to infer the logical relationship between the steps based on the semantics of the data query process.

[0036] Of course, this manual only describes the process of generating query statements based on the business logic process. Therefore, it is obvious that some steps in the business logic process are not implemented through query statements. For example, the query results of high-risk accounts determined based on conditions should then be sent to the risk control department or the legal department. This is not a step implemented by the query statements generated in this manual.

[0037] It should be noted that the data query steps in the business logic process generated by LLM in step S104 are still expressed in natural language, that is, the data query steps and the relationships between the steps in the business logic process are supplemented and improved.

[0038] S106: Based on the business data structure of the business corresponding to the business logic process, according to the obtained steps and the logical relationship, generate a query statement executed on the business data structure, and the query statement is used to implement the business logic process.

[0039] In the embodiments of this specification, after the improved steps and the logical relationship between the steps are obtained through LLM, the server can generate an executable query statement through LLM based on the reasoning of natural language and the business data structure of the business data targeted by the query. The query statement is used to implement part of the data query involved in the business logic process. For example, the first and second steps in the example of step S100, and the last third step of sending the alarm information can be implemented manually. For example, according to the business logic process, after the query statement is generated, the query result is obtained and displayed by executing the query statement on the business data, so that the user can send an alarm message based on the query result. Of course, the third step can also be automatically executed by a pre-set program code. This specification does not limit this and can be set as needed.

[0040] Specifically, the server can use a pre-set prompt template to populate the business data structure, steps, and logical relationships into the template, using natural language input. The LLM then outputs an executable query statement corresponding to the business data structure. For example, the prompt template could be "Based on the steps recorded in ... and the logical relationships between the steps, generate a query statement to be executed on the data in the structure ...." The first ellipsis indicates the steps and logical relationships, and the second ellipsis indicates the business data structure. The pre-trained LLM can then output a query statement to be executed on the business data structure, implementing the data query process within the business logic process.

[0041] based on Figure 1 The large model-based query statement generation method shown here inputs a business logic process described in natural language into a pre-trained LLM. The LLM then identifies the steps within the business logic process and the logical relationships between them. It then completes the implicit steps within the business logic process, determining the steps within the complete business logic process and the logical relationships between them. Finally, based on the business data structure targeted by the query, it generates a corresponding query statement. By inputting the business logic process into the LLM, this method generates the implicit steps within the process, resulting in a complete business logic process. This allows the LLM to generate corresponding query statements, extending the use of the business logic process beyond basic text processing to further analyze the implicit semantics and generate directly executable query statements to support business execution.

[0042] The business logic process described by natural language in the existing technology or the SOP therein can be further utilized to obtain query statements that can be executed automatically, which greatly reduces the workload of manual intervention in the execution of such standards and further explores the value of SOP.

[0043] In order to facilitate further understanding of the method provided in this specification, an example corresponding to steps S100 to S106 is provided in the embodiments of this specification to illustrate how to supplement the implicit steps in the business logic process through step S104.

[0044] For example, assuming that in step S100, the server obtains an SOP, and the content of the SOP is: “##SOP rules Emails with the subject line "salary," "subsidy," or "allowance" and from non-company email addresses will be treated as phishing emails and handled as incidents. It can be seen that the SOP is an event processing method described in natural language. After the SOP rules are input into the LLM in step S102, the LLM can identify the steps included in the SOP and the logical relationships between the steps through a logical reasoning process.

[0045] Specifically, the LLM can identify the following results: the first step is to "query the email subject", "the subject includes salary, subsidy, and allowance", and "the sending email address is not company-related". The second step is to use the query results of the first step to uniformly handle them as phishing emails. The third step is to identify situations outside the collection of the query results of the first step as normal behavior and not take any action.

[0046] Continuing to use the LLM to perform logical reasoning based on the identification results, the LLM determines that the above process lacks specific information to determine company relevance. Specifically, the LLM can infer that the implicit fourth step in the above process is to determine whether the email address contains "@company.com." The logical relationship between the steps is then redefined as executing the first and fourth steps in parallel or sequentially. The second step is to determine the unified handling of the phishing email as an incident based on the query results after executing the first and fourth steps, and the third step remains unchanged. Therefore, the corresponding query statement in the SOP actually only involves the combined query results of the first, second, and third steps, that is, this part of the content requires the query statement to be generated.

[0047] Then, in the final step S106, the server may generate a query statement based on the business data structure and the result determined in step S104: GQL: MATCH (m:Mail)-[:SEND]->(p:Person) WHERE NOT m.sender CONTAINS '@company.com' AND m.title =~ '.(salary|subsidy|allowance).' RETURN m, p As can be seen, this query statement returns results that match the conditions, allowing subsequent business execution processes to execute the "transfer phishing email to an event" business operation based on the returned m and p. The conditions that meet this requirement include the SOP text containing "m.title =~ '.(Salary|Subsidy|Allowance)", which means the email subject contains salary, subsidy, and allowance; "m.senderCONTAINS '@company.com'", which means the sending email address is not company-related; and the matching result is the email and user "(m:Mail)-[:SEND]->(p:Person)", and the query results returned are also Mail and Person.

[0048] The third step of the subsequent SOP process can then handle the phishing email based on the query results. This can be done using the same risk control methods used for phishing emails. Of course, the above example is only one example of an implementation of the embodiments of this specification. Therefore, the specific subsequent handling of phishing emails is not a technical issue to be solved by the embodiments of this specification, and this specification does not limit the method used. It can be set as needed.

[0049] Furthermore, this specification does not restrict the specific structure of the business data; it can be configured as needed. Of course, in risk control business scenarios, risk control logs are generally used as business data. However, risk control logs contain a lot of noise and lack relationships between data. Therefore, in one or more embodiments of this specification, the server can process the risk control logs and convert them into graph data. The query statements generated by the LLM are graph query statements in the Graph Query Language (GQL).

[0050] Therefore, at least before the LLM generates a query statement, the server needs to process the risk control log first to obtain the graph data used for the query.

[0051] Specifically, the server may first obtain the business logs recorded in advance by the business. Taking the risk control business as an example, the risk control logs are obtained at this time.

[0052] Secondly, the pre-trained Large Language Model (LLM) is used to identify entities and their relationships within the business log. As previously mentioned, the business's knowledge graph is pre-imported into the LLM, enabling it to identify entities within that business scenario. For example, in the risk control business context, entities to be identified may include IP addresses, ports, account IDs, and user IDs. Similarly, the LLM can use the knowledge graph to identify time-dependent relationships between entities within risk control scenarios, such as access relationships and attack relationships.

[0053] Finally, with each entity as a node and the relationship between entities as an edge, the graph data corresponding to the business data is constructed.

[0054] Furthermore, in an embodiment of the present specification, the server can pre-process the risk control log before constructing the graph data, remove irrelevant information, and extract the content of key fields to avoid identifying unnecessary entities or relationships during the LLM reasoning process. The so-called unnecessary refers to the information that is unnecessary for the risk control business. Obviously, the irrelevant information that needs to be removed for different businesses is not exactly the same. Of course, for log cleaning, existing mature technologies can be used, such as filtering out necessary field content based on field names and keyword tables. This specification does not limit this and can be set as needed.

[0055] In the embodiments of this specification, before inputting the business logic process into the LLM, the server can also perform text cleaning and formatting on the business logic process to avoid irrelevant information in the business logic process or the SOP contained therein, while unifying the natural language description. This reduces ambiguity in the description and improves the accuracy and consistency of the generated query statements.

[0056] Furthermore, in the embodiments of this specification, to further improve the usability of generated query statements, the server can send the query statement as a verification statement to the user terminal before executing the query statement. The server then receives the verified statement returned by the user terminal and executes it within the established operating environment. This allows for a review of the query statement through human-computer interaction. Since operations and maintenance personnel do not need to write query statements from scratch, the efficiency of data queries based on business logic processes is greatly improved, reducing operating costs.

[0057] Furthermore, in the embodiments of this specification, after receiving the query statement, the server can execute the query statement on the business data of the corresponding business. Specifically, the server can first run the query statement's runtime environment and load the business data of the business corresponding to the business logic process, that is, the graph data. Then, the query statement is executed through the runtime environment based on the book data. Finally, the query statement's query result is determined, and the business execution continues based on the query result.

[0058] In the embodiments of this specification, when executing the query statement generation process for graph data in a risk control business scenario, the server may first execute the graph query statement generated in step S106 in the operating environment based on the graph data corresponding to the business data. In the risk control business scenario, the business logic process of step S100 includes the risk control SOP, so the graph query statement may query the graph data for elements required for risk control.

[0059] Therefore, after executing the graph query statement, the server can obtain the query result in the graph data, that is, at least one of the nodes and edges in the graph data.

[0060] Since nodes in the graph data correspond to entities such as users and activities, and edges represent temporal relationships between entities, typically business events between entities, after receiving a query result, the server can determine at least one of the risky entities and risky business events based on the nodes and edges in the graph data included in the query result. Of course, the specific risk depends on the content of the query result. If the query result only contains nodes, then the entity is at risk, but no direct risky relationship is identified. If the query result only contains edges, then a certain type of business event is directly at risk.

[0061] Then, based on the connections between nodes and edges in the graph data, the server can further determine risky elements in the graph data beyond the query results. Specifically, based on pre-set risk control rules, it can determine whether risky elements exist in the graph data beyond the query results. For example, the risk control rule could be "entities connected by risky edges are also risky," or "edges between risky nodes are also risky." This can be used to extend the determination of other risky elements not found in the query statement, that is, content that is not covered by the SOP but also poses risks.

[0062] Finally, based on the business events and / or entities with risks that are queried, corresponding risk control operations are executed. Of course, the specific risk control operations to be performed can also be determined by preset risk control rules, which will not be elaborated in this manual.

[0063] In addition, in the embodiments of this specification, the server can also generate security decision recommendations based on the query results, and display event information and response measures through visualization. For example, the display is based on the risk elements in the graph data determined by that SOP, and the query results are displayed in the graph data, to determine and display the risk control business that needs to be executed, to help the corresponding users, such as security engineers, make decisions based on the display content and take further action. Specifically, the server can determine the query results and the risk control business to be executed in the graph data, and send them to the user terminal, so that the user terminal displays the query results and the risk control business to be executed, and prompts the user to make a risk control decision. For example, execute the risk control business determined by the server, or select more nodes or edges with potential risks in the selected graph data, and so on.

[0064] In the embodiments of this specification, in order for the LLM to accurately convert the business logic process described in the input natural language into the above-mentioned query statement, it is not enough to simply inject the knowledge graph corresponding to the business logic process into the LLM. The LLM also needs to be fine-tuned and trained in the business scenario corresponding to the business logic process in advance.

[0065] Specifically, when fine-tuning the LLM, you can first obtain a sample logical process described in natural language. This sample logical process also contains several steps, referred to as sample steps below. This sample logical process and the aforementioned business logic process are business logic processes in the same business scenario. That is, the knowledge graph corresponding to the sample logical process is exactly the same as the knowledge graph corresponding to the aforementioned business logic process.

[0066] After obtaining the sample logic process, you can also Figure 1 In steps S102 to S106 shown in FIG, the sample logic process is input into the LLM to be trained, and the sample steps included in the sample logic process are identified by the LLM to be trained, and the sample steps implicit in the sample logic process are determined, as well as the logical relationship between the obtained sample steps. Based on the logical relationship between the sample steps and the business data structure, a statement to be optimized is generated for executing the business data structure. The statement to be optimized is used to implement the sample logic process. The process of generating the statement to be optimized is similar to the process of Figure 1 Steps S102 to S106 are identical and will not be described again here.

[0067] After the trained LLM generates a statement to be optimized, it can determine an optimized statement after adjusting the statement to be optimized. Since the statement to be optimized is also a query statement in form, it can be manually adjusted to ensure that the steps and the logical relationships between them fully conform to the data query process in the original sample logic process. Of course, the statement to be optimized can also be adjusted by other trained LLMs, and this embodiment of the present specification does not limit this.

[0068] After obtaining the optimized statement, the supervised fine-tuning training (SFT) of the LLM to be trained can be performed based on the optimized statement, that is, the optimized statement is used as the label corresponding to the sample logic process, and used as the supervision signal to adjust the model parameters of the LLM to be trained, so that the model parameters of the LLM to be trained are adjusted to make the LLM adapt to generate various business logic processes in the business scenario. Figure 1The device generating the query statement shown can be the same device or a different device. Furthermore, since the embodiments of this specification require that the original reasoning capabilities of the LLM be maintained as much as possible and adapted only to the business scenario, when performing SFT on the LLM, not all model parameters of the LLM are adjusted; instead, only some of the model parameters in the LLM are adjusted.

[0069] After the LLM is obtained through the above SFT training, the server can Figure 1 The method shown generates a query statement corresponding to the business logic process in this business scenario.

[0070] Furthermore, when applying the LLM to the SFT, adjustments are made in stages. Since the requirements for supplementing steps differ from those for improving the LLM's query output, separate training can avoid conflicts between the multiple tasks. Specifically, the process of supplementing the implicit steps still outputs the step content and step relationships described in natural language, while the process of generating the query statement outputs machine language (the query statement) based on the natural language text.

[0071] Similar to the SFT process of the LLM described above, during the process of training the LLM to supplement the implicit steps, the server can delete some data query content in the business logic process as a sample logic process, and use the original business logic process as the optimized business logic process.

[0072] Specifically, when fine-tuning the LLM, we first obtain the original logical process described in natural language. This original logical process also contains several steps, referred to as sample steps below. This original logical process and the aforementioned business logic process are business logic processes in the same business scenario. That is, the knowledge graph corresponding to the original logical process is exactly the same as the knowledge graph corresponding to the aforementioned business logic process.

[0073] After obtaining the original logic process, delete at least one sample step related to data query in the original logic process and use it as the input sample logic process. Figure 1 In steps S102 to S104, the LLM to be trained is used to identify the sample steps contained in the sample logic process, determine the sample steps implicit in the sample logic process (i.e., the deleted sample steps), and determine the logical relationships between the obtained sample steps as the logic process to be optimized. Figure 1 Steps S102 to S104 are identical and will not be described again here.

[0074] After the LLM to be trained generates the logic process to be optimized, the original logic process can be determined as the annotation of the sample logic process, and the LLM to be trained can be fine-tuned. Among them, the fine-tuning process not only needs to adjust the parameters of the sample steps supplemented by the LLM, but also needs to adjust the parameters of the LLM according to the logical relationship between the supplemented sample steps and other sample steps given by the LLM. Therefore, the logic process to be optimized obtained by the LLM to be trained can be regarded as a whole, rather than just performing SFT based on the supplemented implicit sample steps. In addition, since the logic process to be optimized is also a natural language in form, the logic process to be optimized can also be manually adjusted so that the steps and the logical relationship between the steps are completely consistent with the data query process in the original logic process, and used as annotations for SFT. Of course, the annotation of the sample logic process can also be determined in any other way. This embodiment of the present specification does not limit this and can be set as needed.

[0075] Continuing, no matter what method is used to obtain the optimized logic process corresponding to the sample logic process, the LLM to be trained can be subjected to supervised fine-tuning training (SFT) based on the optimized logic process. The optimized logic process is used as a supervisory signal to adjust the model parameters in the LLM to be trained, and the model parameters of the LLM to be trained are adjusted to make the LLM adapt to the implicit steps in the various business logic processes under the business scenario, as well as the logical relationships between the implicit steps. Among them, the device for training the LLM to be trained is the same as the one executing the above-mentioned Figure 1 The device generating the query statement shown can be the same device or a different device. Furthermore, since the embodiments of this specification require that the original reasoning capabilities of the LLM be maintained as much as possible and adapted only to the business scenario, when performing SFT on the LLM, not all model parameters of the LLM are adjusted; instead, only some of the model parameters in the LLM are adjusted.

[0076] Those skilled in the art will appreciate that the above description uses the risk control logic process as an example. In reality, the query statement generation methods provided in the embodiments of this specification can generate and execute queries corresponding to business logic processes in any business scenario. Furthermore, the query statements described above are only illustrated using graph queries as an example. The business data structure in the embodiments of this specification can also be in other forms, and the corresponding query statements will also be of the corresponding type.

[0077] In addition, it should be noted that the implicit steps supplemented by the LLM in the example of step S102 are not query statements executed against business data. This is only an example, and in fact, the embodiments of this specification do not limit whether the steps are directed at business data or external data. As long as the query statements required to support business execution can be supplemented by LLM. As described in the SFT process of LLM above, samples and annotations can be constructed by deleting the complete business logic process to achieve supervised training.

[0078] The above is a query statement generation method based on a large model provided in an embodiment of this specification. Based on the same idea, this specification also provides corresponding devices, storage media and electronic devices.

[0079] Figure 3 A schematic diagram of a query statement generation device based on a large model provided in an embodiment of this specification, the device comprising: An acquisition module 201 is used to acquire a business logic process described in a natural language; An input module 202 is used to input the business logic process into a pre-trained large language model LLM; Identification module 203, configured to identify the steps included in the business logic process through the LLM, determine the implicit steps in the business logic process, and determine the logical relationships between the obtained steps; The generating module 204 is used to generate a query statement executed on the business data structure based on the business data structure corresponding to the business logic process and according to the obtained steps and the logical relationship, wherein the query statement is used to implement the business logic process.

[0080] Optionally, the identification module 202 is used to identify the steps included in the business logic process and the logical relationships between the steps through the LLM, where the steps are texts described in natural language. According to the steps and the logical relationships between the steps, the implicit steps in the business logic process are identified, the text of the implicit steps is generated, and the implicit steps and the logical relationships between the steps are determined.

[0081] Optionally, the device further comprises: The running module 205 runs the running environment of the query statement and loads the business data of the business corresponding to the business logic process. Based on the business data, the query statement is run through the running environment to determine the query result of the query statement, and the business is continued to be executed based on the query result.

[0082] Optionally, the operation module 205 is also used to obtain the business logs recorded in advance by the business, identify the entities in the business logs and the relationships between the entities through a pre-trained large language model LLM, and construct graph data corresponding to the business data with the entities as nodes and the relationships between the entities as edges.

[0083] Optionally, the business logic process includes a risk control logic process; The business log is a risk log; The query statement includes a graph query statement.

[0084] Optionally, the operation module 205 is used to execute the graph query statement in the operation environment based on the graph data corresponding to the business data to obtain a query result of the graph data, wherein the query result includes at least one of the nodes and edges in the graph data; determine the business events and / or entities with risks based on the nodes and edges in the graph data contained in the query result, and execute corresponding risk control business according to preset risk control rules for the business events and / or entities with risks.

[0085] Optionally, the execution module 205 is further configured to send the query statement as a statement to be verified to a user terminal, and execute the verified statement through the execution environment in response to a verified statement returned by the user terminal.

[0086] Optionally, the device further comprises: The training module 206 is used to obtain a sample logical process described in natural language, wherein the sample logical process includes several sample steps; identify each sample step included in the sample logical process through the LLM to be trained, determine the sample steps implicit in the sample logical process, and determine the logical relationship between the obtained sample steps; based on the business data structure of the business corresponding to the sample logical process, generate a statement to be optimized for execution on the business data structure according to the logical relationship between each sample step, wherein the statement to be optimized is used to implement the sample logical process; determine an optimized statement after adjusting the statement to be optimized; and perform supervised fine-tuning training on the LLM to be trained according to the optimized statement.

[0087] This specification also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it can be used to execute the query statement generation method based on the large model provided above.

[0088] based on Figure 1 The query statement generation method based on the large model shown in the embodiment of this specification also provides Figure 4 The structural diagram of the electronic device shown in FIG. Figure 4 At the hardware level, the electronic device includes a processor, an internal bus, a network interface, memory, and non-volatile storage, and may also include other hardware required for its operations. The processor reads the corresponding computer program from the non-volatile storage into the memory and then runs it to implement the aforementioned large model-based query statement generation method.

[0089] The foregoing is merely an example of the present invention and is not intended to limit the present invention. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be included within the scope of the claims of the present invention.

Claims

1. A method for generating a query statement based on a large model, the method comprising: Obtain business logic processes described in natural language; Input the business logic process into a pre-trained large language model LLM; Identifying each step included in the business logic process through the LLM, determining the implicit steps in the business logic process, and determining the logical relationship between the obtained steps; Based on the business data structure of the business corresponding to the business logic process, a query statement executed on the business data structure is generated according to the obtained steps and the logical relationship, and the query statement is used to implement the business logic process.

2. The method according to claim 1, wherein the LLM is used to identify the steps included in the business logic process, determine the implicit steps in the business logic process, and determine the logical relationship between the obtained steps, specifically comprising: Identifying, by means of the LLM, each step included in the business logic process and the logical relationship between the steps, wherein each step is a text described in a natural language; Identify the implicit steps in the business logic process based on the steps and the logical relationships between the steps; The text of the implicit steps is generated, and the implicit steps and the logical relationships between the steps are determined.

3. The method of claim 1, further comprising: Run the query statement in the operating environment and load the business data of the business corresponding to the business logic process; Based on the business data, running the query statement through the operating environment; Determine a query result of the query statement, and continue to execute the business based on the query result.

4. The method according to claim 3, further comprising: Obtaining a pre-recorded business log of the business; Identify entities in the business log and the relationships between the entities using a pre-trained large language model (LLM); The graph data corresponding to the business data is constructed with the entities as nodes and the relationships between the entities as edges.

5. The method according to claim 4, wherein the business logic process includes a risk control logic process; The business log is a risk log; The query statement includes a graph query statement.

6. The method according to claim 5, wherein the query statement is executed in the execution environment based on the business data, specifically comprising: Based on the graph data corresponding to the business data, executing the graph query statement in the operating environment to obtain a query result of the graph data, the query result including at least one of a node and an edge in the graph data; Determining a query result of the query statement and continuing to execute the business based on the query result specifically includes: Determining risky business events and / or entities based on the nodes and edges in the graph data included in the query result; For business events and / or entities with risks, corresponding risk control operations are executed according to preset risk control rules.

7. The method according to claim 3, wherein executing the query statement through the execution environment comprises: Sending the query statement as a statement to be verified to the user terminal; In response to the verified statement returned by the user terminal, the verified statement is executed through the execution environment.

8. The method of claim 1, wherein the LLM is pre-trained, specifically comprising: Acquire a sample logical process described in natural language, where the sample logical process includes several sample steps; Identifying each sample step included in the sample logic process through the LLM to be trained, determining the sample steps implicit in the sample logic process, and determining the logical relationship between the obtained sample steps; Based on the business data structure of the business corresponding to the sample logical process and according to the logical relationship between each sample step, generating a statement to be optimized for execution on the business data structure, wherein the statement to be optimized is used to implement the sample logical process; Determine an optimized statement after adjusting the statement to be optimized; According to the optimized statement, supervised fine-tuning training is performed on the LLM to be trained.

9. A query statement generation device based on a large model, the device comprising: An acquisition module is used to acquire the business logic process described in natural language; An input module, configured to input the business logic process into a pre-trained large language model LLM; an identification module, configured to identify the steps included in the business logic process through the LLM, determine the implicit steps in the business logic process, and determine the logical relationships between the obtained steps; A generation module is used to generate a query statement executed on the business data structure based on the business data structure corresponding to the business logic process and according to the obtained steps and the logical relationship, wherein the query statement is used to implement the business logic process.

10. A computer-readable storage medium storing a computer program, wherein the computer program implements the method according to any one of claims 1 to 8 when executed by a processor.

11. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1 to 8 when executing the program.

Citation Information

Patent Citations

  • Business query statement construction method and device, equipment and storage medium

    CN117216094A

  • Data processing method, computing device, storage medium and computer program product

    CN119961280A

  • Query statement generation method and device and computing equipment

    CN120256456A

  • Data processing method and device, equipment and storage medium

    CN120353816A

  • Generating training examples for translation of natural language queries to executable database queries

    US20240378198A1