Risk control method and device, storage medium and electronic equipment
By automatically extracting risk control data features using pre-trained large language models in the risk control system, the problems of low efficiency and poor accuracy of manual search for data features in the prior art are solved, and more efficient and accurate risk control is achieved.
Patent Information
- Application Number
- CN202510652561.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-05-20
AI Technical Summary
In the prior art, the data characteristics required for risk control need to be manually searched in massive business data, which is inefficient and easy to introduce human errors, resulting in a decrease in risk control accuracy.
Using a pre-trained large language model (LLM), by inputting business data and risk types into LLM, data characteristics corresponding to risk types are automatically extracted and risk control is carried out based on these characteristics.
It improves risk control efficiency, reduces artificial errors, and enhances the accuracy of risk control.
Smart Images

Figure CN120180100A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of computer technology, and particularly to a risk control method, device, storage medium, and electronic device. Background Art
[0002] With the development of information technology, people are increasingly accustomed to conducting business online, and various risks arising in the online environment are emerging in an endless stream. Therefore, each service provider needs to perform risk control on the services it provides.
[0003] Risk control needs to be carried out based on certain data characteristics. For example, for the risk of account theft, it is necessary to determine the login addresses of an account within a certain period of time. If the login addresses of the account change frequently during this period, the account may be at risk of being stolen. That is to say, when performing risk control on the risk of account theft, the login addresses of the account within a certain period of time are the data characteristics for risk control.
[0004] However, in the prior art, the data characteristics required for risk control of different risks need to be found manually by experience in the massive business data of the service provider, which not only has extremely low efficiency but also easily introduces human errors, resulting in a reduction in the accuracy of risk control. Summary of the Invention
[0005] Embodiments of this specification provide a risk control method, device, storage medium, and electronic device to partially solve the problems existing in the above prior art.
[0006] Embodiments of this specification adopt the following technical solutions: A risk control method provided in this specification, the method includes: Obtain the business data of the service provider; Input the business data and the risk type to be controlled into a pre-trained large language model LLM; wherein, the LLM is pre-trained based on the data characteristics corresponding to the risk type that have been determined historically; Obtain the data characteristics corresponding to the risk type extracted by the LLM from the business data; Perform risk control on the risk of the risk type according to all the data characteristics corresponding to the risk type and the business data.
[0007] A risk control device provided in this specification, the device includes: An acquisition module, configured to obtain the business data of the service provider; An input module, configured to input the business data and the risk type to be controlled into a pre-trained large language model LLM; wherein, the LLM is pre-trained based on the data characteristics corresponding to the risk type that have been determined historically; An extraction module for obtaining the data features corresponding to the risk type extracted by the LLM from the service data. A risk control module for performing risk control on the risk of the risk type according to all the data features corresponding to the risk type and the service data.
[0008] A computer-readable storage medium provided in this specification, where the storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned risk control method is implemented.
[0009] An electronic device provided in this specification, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the above-mentioned risk control method is implemented.
[0010] At least one of the above-mentioned technical solutions adopted in the embodiments of this specification can achieve the following beneficial effects: The embodiments of this specification disclose a risk control method. The method pre-trains an LLM based on the data features corresponding to each risk type determined in history. During risk control, the service data of the service provider and the risk type to be risk-controlled are input into the pre-trained LLM, and the data features corresponding to the risk type extracted by the LLM from the service data are obtained. Then, according to the risk features corresponding to the risk type and the service data, risk control is performed on the risk corresponding to the risk type. This method does not require manual search for data features in a large amount of service data. Instead, it enables the LLM to learn the method of extracting data features from service data in history and uses the LLM to extract data features from service data, which can effectively improve the risk control efficiency and does not introduce human errors. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The drawings described herein are used to provide a further understanding of this specification, and constitute a part of this specification. The illustrative embodiments of this specification and their descriptions are used to explain this specification and do not constitute an improper limitation to this specification. In the drawings: Figure 1 A flowchart of a risk control method provided in an embodiment of this specification; Figure 2 A flowchart of a method for pre-training an LLM provided in an embodiment of this specification; Figure 3 A schematic diagram of a risk control device provided in an embodiment of this specification; Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of this specification. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0012] To make the objectives, technical solutions, and advantages of this specification clearer, the technical solutions of this specification will be clearly and completely described below in conjunction with specific embodiments of this specification and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope protected by this specification.
[0013] The following will detail the technical solutions provided by each embodiment of this specification in conjunction with the drawings.
[0014] Figure 1 The following is a flowchart of a risk control method provided by an embodiment of this specification, including the following steps: S100: Obtain the business data of the service provider.
[0015] In the embodiment of this specification, the device for performing risk control using the method as Figure 1 shown can be the server of the service provider. Specifically, an intelligent agent (Agent) with a large language model (LLM) as the core can be pre-deployed in the server, and risk control can be performed through this intelligent agent.
[0016] When the server performs risk control, it needs to be based on the business data of the service provider. Therefore, the server needs to first obtain the business data of the service provider. Specifically, the server can obtain the business data from the data warehouse storing the business data of the service provider. This data warehouse can be, for example, an Open Data Processing Service (ODPS) data warehouse. The business data of the service provider stored in this data warehouse can include the basic information of users (such as name, ID, ID card information, contact information, etc.) and the business behavior information of users (such as login behavior and its occurrence time and location, transaction behavior and its occurrence time, location, and the amount involved, etc.).
[0017] S102: Input the business data and the risk type to be risk-controlled into the pre-trained large language model LLM.
[0018] Since risk control based on business data requires extracting corresponding data features from a large amount of business data, in the embodiment of this specification, the method of manually finding data features based on experience is abandoned, and the process of steps S100 to S104 as Figure 1 shown is adopted to extract data features from business data using the LLM.
[0019] Specifically, the agent deployed on the server can input the business data obtained in step S100 and the risk types to be risk-controlled into a pre-trained LLM. Among them, the agent can generate a prompt message according to the business data and the risk types to be risk-controlled. This prompt message is used to enable the LLM to extract the data features corresponding to the risk types from the business data, and then input the generated prompt message into the LLM.
[0020] The LLM described in the embodiments of this specification is pre-trained based on the data features corresponding to each risk type that have been determined historically. The LLM can learn the method of manually extracting the data features corresponding to each risk type from business data historically, and then use the reasoning ability of the trained LLM to automatically extract the data features corresponding to the risk types to be risk-controlled from the business data obtained in step S100. The specific method of training the LLM will be described later.
[0021] S104: Obtain the data features corresponding to the risk types extracted by the LLM from the business data.
[0022] In the embodiments of this specification, the LLM is a knowledge-base-based LLM, such as a Retrieval-Augmented Generation (RAG) LLM, and its knowledge base at least contains the data features corresponding to each risk type that have been determined historically. After the agent obtains the data features corresponding to the risk types to be risk-controlled extracted by the LLM from the business data, it can also determine whether the data features extracted by the LLM already exist in the knowledge base. If they already exist, the data features extracted by the LLM can be directly discarded. If they do not exist, it means that the data features extracted by the LLM are new data features corresponding to the risk types, and the above knowledge base can be updated according to the data features corresponding to the risk types extracted by the LLM, that is, the data features corresponding to the risk types extracted by the LLM are saved to the above knowledge base.
[0023] In addition, the knowledge base on which the LLM is based may also include the description information of each field in the business data of the service provider and / or the description information of each risk type. The description information of each field is used to describe the meaning of each field in the business data in natural language, and the description information of each risk type is used to describe the meaning of each risk type in natural language, so that the LLM can better understand the business data obtained in step S100 and the risk types to be risk-controlled input in step S102, and enable the LLM to more accurately extract the data features corresponding to the risk types to be risk-controlled.
[0024] Correspondingly, in step S102, in addition to constructing data features corresponding to the risk types required for risk control that enable the LLM to extract from the input business data, the agent can also construct fields in the business data that the LLM itself cannot understand, as well as risk types that the LLM itself cannot understand. Then, in step S104, if the LLM outputs fields and / or risk types in the business data that it cannot understand, the description information of the fields and risk types that the LLM cannot understand can be added to the above knowledge base.
[0025] Of course, since the business data obtained by the agent from the data warehouse through step S100 is very complex, the business data may include business data with risk labels. If the risk label of the business data is the risk label corresponding to the risk type required for current risk control, then in step S102, the agent does not need to input the business data corresponding to the risk label of this risk type into the LLM again, but directly uses this business data as the data features corresponding to this risk type. If the business data does not have a risk label, or the risk label it has is not the risk label corresponding to the risk type required for current risk control, then the agent can input this business data into the LLM through step S102 to extract the data features corresponding to this risk type required for current risk control from this business data.
[0026] S106: Perform risk control on the risk of the risk type according to all the data features corresponding to the risk type and the business data.
[0027] Through the above steps S100 - S104, after the agent obtains the data features corresponding to the risk types required for risk control extracted by the LLM from the business data, it can screen out the business data with the risk of this risk type from the business data obtained in step S100 according to this data feature, and perform risk control based on the screened business data according to the preset risk control rules.
[0028] The above method pre - trains the LLM based on the data features corresponding to each risk type determined in history. During risk control, the business data of the service provider and the risk types required for risk control are input into the pre - trained LLM to obtain the data features corresponding to this risk type extracted by the LLM from the business data, and then perform risk control on the risk corresponding to this risk type according to the risk features and business data corresponding to this risk type. This method does not require manual search for data features in a large amount of business data, but enables the LLM to learn the method of extracting data features from business data in history and uses the LLM to extract data features from business data, which can effectively improve the risk control efficiency and will not introduce human errors.
[0029] The above Figure 1The method shown requires pre-training the LLM so that the LLM learns the extraction methods for extracting data features corresponding to each risk type from business data manually in history. Since the LLM is mainly a large model for processing natural language, in the embodiments of this specification, the extraction methods in history are converted into natural language to facilitate the training of the LLM, as Figure 2 shown.
[0030] Figure 2 is the flowchart of the method for pre-training the LLM provided by the embodiments of this specification, which specifically includes: S200: Determine the data features corresponding to the risk types that have been determined in history as historical data features.
[0031] When training the LLM, on the one hand, the agent needs to obtain the data features corresponding to each risk type that have been determined in history as historical data features, and on the other hand, it also needs to obtain the LLM to be trained. In the embodiments of this specification, in order to improve the training efficiency, the LLM to be trained can use the LLM that has been pre-trained.
[0032] S202: Query the extraction methods for extracting the historical data features from the business data in history.
[0033] In the embodiments of this specification, there are mainly two methods for manually extracting data features from business data in history. One is to query data with certain features from the data warehouse through a data query statement and use the queried data as data features. The other is to extract data with a certain risk through the extraction rules set in the risk control system as risk data. Therefore, the methods for extracting data features corresponding to each risk type from business data are actually hidden in the above data query statements and the extraction rules of the risk control system.
[0034] Thus, for the data query statement, the agent can query the data query statement for extracting historical data features from the database storing business data (such as the above-mentioned ODPS data warehouse) in history as the extraction method for extracting historical data features in history. Among them, the data query statement includes but is not limited to SQL statements. For the extraction rule, the agent can query the extraction rule used to extract the historical data features as the extraction method for extracting historical data features in history.
[0035] S204: Determine the natural language corresponding to the extraction method.
[0036] For the data query statement, the agent can input the queried data query statement for extracting historical data features into the pre-trained LLM to obtain the natural language corresponding to the data query statement output by the pre-trained LLM.
[0037] For the extraction rules queried from the risk control system, the agent can also input the queried extraction rules into the pre-trained LLM to obtain the description information corresponding to the extraction rules output by the pre-trained LLM. This description information is used to describe the extraction rules in the form of natural language.
[0038] Since the pre-trained LLM already has general natural language processing capabilities, the pre-trained LLM can convert data query statements such as SQL statements or extraction rules set in the risk control model into corresponding natural language for description.
[0039] S206: Fine-tune the pre-trained LLM according to the natural language.
[0040] In order to enable it to learn the extraction methods of data features corresponding to each risk type while maintaining the inference ability and generalization ability of the pre-trained LLM, a Low-Rank Adaptation (LoRA) can be added to the pre-trained LLM in the embodiments of this specification.
[0041] When training the LLM, several sample question-and-answer pairs, that is, QA pairs, can be determined as training samples according to the natural language descriptions of data query statements and / or extraction rules obtained in step S204. For example, question: What data features are data features with account theft risk? Answer: Data features with frequently changing login addresses within 3 days are data features with account theft risk.
[0042] After determining the training samples, all model parameters in the LLM except LoRA can be fixed, and the model parameters of LoRA can be fine-tuned according to the above training samples. That is, keep all model parameters in the LLM except LoRA unchanged and only adjust the model parameters of LoRA.
[0043] Through Figure 2 After training the LLM by the method shown, the trained LLM can be used to extract and mine data features corresponding to each risk type from a large amount of business data through Figure 1 the steps S100~S104 shown, and when a risk control request is received, the business data matching at least one data feature corresponding to the risk type to be risk-controlled in the business data can be determined through Figure 1 the step S106 shown as risk data, and the risk data can be risk-controlled.
[0044] The above is a risk control method provided by the embodiments of this specification. Based on the same idea, this specification also provides corresponding devices, storage media, and electronic devices.
[0045] Figure 3 Schematic diagram of a risk control device provided by an embodiment of this specification. The device includes: An acquisition module 301, configured to acquire business data of a service provider; An input module 302, configured to input the business data and the risk types to be risk-controlled into a pre-trained large language model LLM; wherein, the LLM is pre-trained based on data features corresponding to the risk types that have been determined historically; An extraction module 303, configured to acquire data features corresponding to the risk types extracted by the LLM from the business data; A risk control module 304, configured to perform risk control on the risks of the risk types according to all data features corresponding to the risk types and the business data.
[0046] Optionally, the device further includes: A training module 305, configured to determine data features corresponding to the risk types that have been determined historically as historical data features; query extraction methods for extracting the historical data features from the business data historically; determine natural languages corresponding to the extraction methods; and perform fine-tuning training on the pre-trained LLM according to the natural languages.
[0047] Optionally, the training module 305 is specifically configured to query data query statements for extracting the historical data features from a database storing the business data historically.
[0048] Optionally, the training module 305 is specifically configured to input the data query statements into the pre-trained LLM to obtain natural languages corresponding to the data query statements output by the pre-trained LLM.
[0049] Optionally, the training module 305 is specifically configured to query extraction rules for extracting the historical data features historically.
[0050] Optionally, the training module 305 is specifically configured to input the extraction rules into the pre-trained LLM to obtain description information corresponding to the extraction rules output by the pre-trained LLM.
[0051] Optionally, add a low-rank adapter LoRA to the pre-trained LLM; The training module 305 is specifically configured to determine a number of sample question-and-answer pairs according to the natural language; fix model parameters other than the LoRA in the LLM with the LoRA added, and perform fine-tuning training on the model parameters with the LoRA added according to the sample question-and-answer pairs.
[0052] Optionally, the pre-trained LLM is a knowledge-base based LLM; The knowledge base contains data features corresponding to various risk types that have been determined historically; After obtaining the data features corresponding to the risk type extracted by the LLM from the business data, the extraction module 303 is further configured to update the knowledge base according to the data features corresponding to the risk type extracted by the LLM if the data features corresponding to the risk type extracted by the LLM do not exist in the knowledge base.
[0053] This specification also provides a computer-readable storage medium storing a computer program, which when executed by a processor can be used to execute the above-provided risk control method.
[0054] Based on Figure 1 the risk control method shown, embodiments of this specification also provide Figure 4 a schematic structural diagram of the electronic device shown. As Figure 4 shown, at the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include other hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the above risk control method.
[0055] The above are only embodiments of this specification and are not intended to limit this specification. For those skilled in the art, this specification may have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification shall be included within the scope of the claims of this specification.
Claims
1. A risk control method, the method comprising: Obtain business data from service providers; Input the business data and the risk type required for risk control into a pre-trained large language model LLM; wherein the LLM is pre-trained based on data features corresponding to the risk type that have been determined historically; Obtaining data features corresponding to the risk type extracted by the LLM from the business data; Based on all data features corresponding to the risk type and the business data, risk control is performed on the risk of the risk type.
2. The method according to claim 1, wherein the LLM is pre-trained, specifically comprising: Determine data features corresponding to the risk types that have been determined historically as historical data features; Query the extraction method of extracting the historical data features from the business data in history; Determining a natural language corresponding to the extraction method; According to the natural language, fine-tune the pre-trained LLM.
3. The method according to claim 2, wherein the method for extracting the historical data features from the business data in the query history specifically comprises: The query is to extract historical data features from a database storing the business data.
4. The method according to claim 3, wherein determining the natural language corresponding to the extraction method comprises: The data query statement is input into a pre-trained LLM to obtain a natural language corresponding to the data query statement output by the pre-trained LLM.
5. The method according to claim 2, wherein the method for extracting the historical data features from the business data in the query history specifically comprises: The extraction rules historically used to extract the features of the historical data are queried.
6. The method according to claim 5, wherein determining the natural language corresponding to the extraction method comprises: The extraction rule is input into a pre-trained LLM to obtain description information corresponding to the extraction rule output by the pre-trained LLM.
7. The method according to claim 3 or 5, wherein a low-rank adapter LoRA is added to the pre-trained LLM; According to the natural language, the pre-trained LLM is fine-tuned, specifically including: Determining a number of sample question-answer pairs according to the natural language; The model parameters other than the LoRA in the LLM to which the LoRA is added are fixed, and the model parameters to which the LoRA is added are fine-tuned and trained according to the sample question-answer pairs.
8. The method of claim 1, wherein the pre-trained LLM is a knowledge base-based LLM; The knowledge base contains data features corresponding to each risk type that has been determined historically; After obtaining the data feature corresponding to the risk type extracted by the LLM from the business data, the method further includes: If the data feature corresponding to the risk type extracted by the LLM does not exist in the knowledge base, the knowledge base is updated according to the data feature corresponding to the risk type extracted by the LLM.
9. A wind control device, comprising: The acquisition module is used to obtain the business data of the service provider; An input module, used to input the business data and the risk type required for risk control into a pre-trained large language model LLM; wherein the LLM is pre-trained based on data features corresponding to the risk type that have been determined historically; An extraction module, used to obtain data features corresponding to the risk type extracted by the LLM from the business data; The risk control module is used to perform risk control on the risk of the risk type according to all data features corresponding to the risk type and the business data.
10. A computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the method according to any one of claims 1 to 8.
11. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1 to 8 when executing the program.
Citation Information
Patent Citations
Real-time data query method and device, computer equipment and readable storage medium
CN117076494A
Model training method and device, storage medium and electronic equipment
CN117593003A
Risk identification method and device, storage medium and electronic equipment
CN117787418A
User risk behavior perception method based on large language model and related equipment
CN118504586A
LLM-based business risk intelligence analysis method and apparatus, and storage medium
CN118504976A