Risk audit method and device, storage medium and electronic equipment
Patent Information
- Application Number
- CN202610750009.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-28
- Publication Date
- 2026-09-11
AI Technical Summary
[0003]有鉴于此,本发明实施例提供了一种风险审计方法、装置、存储介质及电子设备,以解决相关技术导致风险审计的准确性较低等问题;也就是说,本发明实施例可通过包含财务凭证数据和/或交互对话文本的训练样本集合对逻辑审计模型进行模型训练,从而可得到风险审计性能较高的审计应用模型,以提高推理输出结果的准确性,即可提高思考过程和风险等级结论的准确性,进而可有效提高风险审计的准确性
[0008]本发明实施例可获取训练样本集合,一个训练样本包括以下至少一种:一个审计对象的财务凭证数据和交互对话文本;并调用文本解析模型,分别对训练样本集合中的各个训练样本进行解析处理,得到各个训练样本的解析数据。然后,可调用逻辑审计模型,分别基于各个训练样本的解析数据,生成各个训练样本的G个推理输出结果,G为大于1的整数;其中,一个推理输出结果包括一个思考过程和一个风险等级结论,一个思考过程用于指示通过文本证据推理出一个风险等级结论的理由。基于此,可分别基于各个训练样本的G个推理输出结果,确定各个训练样本的G个推理输出结果中各个推理输出结果的综合奖励;并基于各个训练样本的各个推理输出结果的综合奖励,计算模型损失值,从而按照减小模型损失值的方向,优化逻辑审计模型中的模型参数,直至达到目标收敛条件,以将达到目标收敛条件的逻辑审计模型作为目标逻辑审计模型。进一步的,可基于目标逻辑审计模型,确定审计应用模型,审计应用模型支持用于进行风险审计。可见,本发明实施例可通过包含财务凭证数据和/或交互对话文本的训练样本集合对逻辑审计模型进行模型训练,从而可得到风险审计性能较高的审计应用模型,以提高推理输出结果的准确性,即可提高思考过程和风险等级结论的准确性,进而可有效提高风险审计的准确性。
Smart Images

Figure CN122736792A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and in particular to a risk auditing method, apparatus, storage medium, and electronic device. Background Technology
[0002] Currently, with the development of fintech, credit approval has shifted from manual to automated risk control models; however, these technologies often cannot explain the judgment logic and are prone to bias when handling rigorous financial logic, resulting in low accuracy in risk auditing. Therefore, there is currently no satisfactory solution to improve the accuracy of risk auditing. Summary of the Invention
[0003] In view of this, embodiments of the present invention provide a risk auditing method, apparatus, storage medium, and electronic device to solve the problem of low accuracy in risk auditing caused by related technologies. That is, embodiments of the present invention can train a logical auditing model by using a training sample set containing financial voucher data and / or interactive dialogue text, thereby obtaining an auditing application model with high risk auditing performance, improving the accuracy of reasoning output results, thus improving the accuracy of the thinking process and risk level conclusions, and effectively improving the accuracy of risk auditing.
[0004] According to one aspect of the present invention, a risk auditing method is provided, the method comprising: Obtain a training sample set, wherein a training sample includes at least one of the following: financial voucher data and interactive dialogue text of an audit object; The text parsing model is invoked to parse each training sample in the training sample set to obtain the parsed data of each training sample. The logical audit model is invoked to generate G inference outputs for each training sample based on the parsed data of each training sample, where G is an integer greater than 1. Each inference output includes a thought process and a risk level conclusion. The thought process is used to indicate the reasons for inferring a risk level conclusion from textual evidence. Based on the G inference outputs of each training sample, the comprehensive reward of each inference output among the G inference outputs of each training sample is determined; and based on the comprehensive reward of each inference output of each training sample, the model loss value is calculated, thereby optimizing the model parameters in the logic audit model in the direction of reducing the model loss value, until the target convergence condition is reached, so that the logic audit model that reaches the target convergence condition is taken as the target logic audit model; Based on the target logic audit model, an audit application model is determined, which supports risk auditing.
[0005] According to another aspect of the present invention, a risk auditing apparatus is provided, the apparatus comprising: The acquisition unit is used to acquire a training sample set, wherein a training sample includes at least one of the following: financial voucher data and interactive dialogue text of an audit object; The processing unit is used to call the text parsing model to parse each training sample in the training sample set to obtain the parsed data of each training sample. The processing unit is also used to call the logical audit model to generate G inference output results for each training sample based on the parsed data of each training sample, where G is an integer greater than 1; wherein, each inference output result includes a thought process and a risk level conclusion, and the thought process is used to indicate the reason for inferring a risk level conclusion through textual evidence. The processing unit is further configured to determine the comprehensive reward of each inference output result among the G inference output results of each training sample based on the G inference output results of each training sample; and calculate the model loss value based on the comprehensive reward of each inference output result of each training sample, thereby optimizing the model parameters in the logic auditing model in the direction of reducing the model loss value, until the target convergence condition is reached, so as to take the logic auditing model that reaches the target convergence condition as the target logic auditing model; The processing unit is further configured to determine an audit application model based on the target logical audit model, wherein the audit application model supports risk auditing.
[0006] According to another aspect of the present invention, an electronic device is provided, the electronic device including a processor and a memory storing a program, wherein the program includes instructions that, when executed by the processor, cause the processor to perform the methods mentioned above.
[0007] According to another aspect of the present invention, a non-transitory computer-readable storage medium storing computer instructions for causing a computer to perform the methods mentioned above is provided.
[0008] This invention provides an embodiment of a training sample set. Each training sample includes at least one of the following: financial voucher data and interactive dialogue text of an audit object. A text parsing model is invoked to parse each training sample in the training sample set, obtaining parsed data for each training sample. Then, a logical auditing model is invoked to generate G inference outputs for each training sample based on the parsed data, where G is an integer greater than 1. Each inference output includes a thought process and a risk level conclusion. The thought process indicates the reasoning for inferring a risk level conclusion from textual evidence. Based on this, the comprehensive reward of each inference output in each training sample can be determined based on the G inference outputs. The model loss value is calculated based on the comprehensive reward of each inference output, thereby optimizing the model parameters in the logical auditing model in the direction of reducing the model loss value until the target convergence condition is reached. The logical auditing model that reaches the target convergence condition is then used as the target logical auditing model. Furthermore, an audit application model can be determined based on the target logical auditing model, which supports risk auditing. As can be seen, embodiments of the present invention can train a logical audit model using a training sample set containing financial voucher data and / or interactive dialogue text, thereby obtaining an audit application model with high risk audit performance, which can improve the accuracy of reasoning output results, thereby improving the accuracy of the thinking process and risk level conclusions, and thus effectively improving the accuracy of risk audit. Attached Figure Description
[0009] Further details, features, and advantages of the invention are disclosed in the following description of exemplary embodiments in conjunction with the accompanying drawings, in which: Figure 1 A flowchart illustrating a risk audit method according to an exemplary embodiment of the present invention is shown; Figure 2 A schematic diagram of a risk audit system according to an exemplary embodiment of the present invention is shown; Figure 3 A flowchart illustrating another risk auditing method according to an exemplary embodiment of the present invention is shown; Figure 4 A schematic block diagram of a risk audit apparatus according to an exemplary embodiment of the present invention is shown; Figure 5 A structural block diagram of an exemplary electronic device that can be used to implement embodiments of the present invention is shown. Detailed Implementation
[0010] Embodiments of the present invention will now be described in more detail with reference to the accompanying drawings. While some embodiments of the invention are shown in the drawings, it should be understood that the invention can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the invention. It should be understood that the accompanying drawings and embodiments are for illustrative purposes only and are not intended to limit the scope of protection of the invention.
[0011] It should be understood that the various steps described in the method embodiments of the present invention may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this respect.
[0012] The term "comprising" and its variations as used herein are open-ended, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the following description. It should be noted that the concepts of "first", "second", etc., mentioned in this invention are used only to distinguish different devices, modules, or units, and are not intended to limit the order of functions performed by these devices, modules, or units or their interdependencies.
[0013] It should be noted that the terms "a" and "a plurality of" used in this invention are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0014] The names of the messages or information exchanged between the multiple devices in the embodiments of the present invention are for illustrative purposes only and are not intended to limit the scope of these messages or information.
[0015] It should be noted that the executing entity of the risk auditing method provided in this embodiment of the invention can be one or more electronic devices, and this embodiment of the invention does not limit this; wherein, the electronic device can be a terminal (i.e., a client) or a server. Therefore, when the executing entity includes multiple electronic devices, and among the multiple electronic devices includes at least one terminal and at least one server, the risk auditing method provided in this embodiment of the invention can be jointly executed by the terminal and the server. Accordingly, the terminal mentioned herein may include, but is not limited to: smartphones, tablets, laptops, desktop computers, intelligent voice interaction devices, etc. The server mentioned herein can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms, etc.
[0016] Optionally, embodiments of the present invention may also provide a risk audit system. In this case, the risk audit method provided by the embodiments of the present invention can be executed by the risk audit system. Optionally, the risk audit system may be installed on one or more electronic devices. Accordingly, the risk audit method proposed in the embodiments of the present invention may be executed by one or more electronic devices with the risk audit system installed. That is, the executing entity may be one or more electronic devices with the risk audit system installed. In other words, electronic devices can execute the risk audit method through the risk audit system, and so on.
[0017] Based on the above description, this embodiment of the invention proposes a risk auditing method, which can be executed by the aforementioned electronic device (terminal or server); or, the risk auditing method can be executed jointly by the terminal and the server, etc.; that is, the risk auditing method can be executed by a risk auditing system on an electronic device, etc. For ease of explanation, the following description will use the execution of the risk auditing method by an electronic device as an example; such as... Figure 1 As shown, this risk audit method may include the following steps S101-S105: S101, Obtain a training sample set, wherein a training sample includes at least one of the following: financial voucher data and interactive dialogue text of an audit object.
[0018] Optionally, an audit target can be any user, and this embodiment of the invention does not limit this; for example, an audit target can be any user applying for credit, etc.
[0019] Optionally, the training sample set can be obtained in ways including but not limited to the following: The first method of acquisition: The electronic device can store a set of training samples in its own storage space. In this case, the electronic device can obtain the set of training samples from its own storage space.
[0020] The second method of acquisition: The electronic device can obtain the training sample download link. In this case, the electronic device can use the training sample download link to download the training sample set in order to obtain the training sample set.
[0021] The third acquisition method: Electronic devices can also acquire an initial training sample set and perform data preprocessing on the initial training sample set to obtain a training sample set, etc. It should be noted that the specific implementation of data preprocessing in this embodiment of the invention is not limited; for example, data preprocessing may include, but is not limited to, at least one of the following: privacy desensitization (such as removing privacy information), data cleaning (such as cleaning duplicate content, photos, links, etc.), etc. Optionally, the database can input multi-source text data (such as unstructured logs, social security documents, interactive dialogue text, etc.) to store the initial training sample set.
[0022] Optionally, the electronic device may include a text perception and parsing layer (e.g., a risk audit system may include a text perception and parsing layer). In this case, the electronic device can obtain a training sample set through the text perception and parsing layer; for example, such as... Figure 2 As shown, the text perception and parsing layer may include a preprocessing module. Then, the electronic device can use the preprocessing module to preprocess the initial training sample set to obtain the training sample set, and so on.
[0023] S102, call the text parsing model to parse each training sample in the training sample set to obtain the parsed data of each training sample.
[0024] Optionally, the text parsing model can be any text semantic model, that is, any Large Language Model (LLM) used to process unstructured text; this embodiment of the invention does not limit this. Optionally, the text parsing model can be a deep text semantic model based on a Transformer structure (a deep learning sequence modeling architecture based on a self-attention mechanism) to process unstructured text. A deep text semantic model can be a large language model, that is, a text parsing model can be a large language model. In this embodiment of the invention, the text parsing model can also be referred to as a text parsing expert agent.
[0025] Optionally, financial voucher data may include, but is not limited to, at least one of the following: bank statements, supporting documents, social security documents, etc., and this embodiment of the invention does not limit this. Optionally, financial voucher data may be plain text data. Based on this, when parsing financial voucher data, the electronic device can use a text parsing model to capture the logical position of text in a paragraph (such as the amount and its corresponding transaction summary) using text semantic embedding and structural position embedding; correspondingly, through the defined credit entity field, the text parsing model can automatically extract the field content of each field in the credit entity field from long texts, and can output a standardized JSON format (JavaScript Object Notation, a lightweight, plain text, cross-language data format) financial feature object (also called financial feature object data) based on the field content of each field in the credit entity field. Optionally, when no field content is extracted, the field content of the corresponding field may be empty, and this embodiment of the invention does not limit this; optionally, the credit entity field may be set according to experience or actual needs, and this embodiment of the invention does not limit this; for example, the credit entity field may include, but is not limited to, at least one of the following: salary identifier, transaction date, and other key entity fields. Optionally, a financial feature object data may include the field content of each of the at least one financial parsing field; optionally, the at least one financial parsing field may be set according to experience or actual needs, such as at least one of the following: average monthly income, frequency of sensitive transactions, etc., which are not limited in this embodiment of the invention. Optionally, the field content of a financial parsing field may be empty, such as when the field content of the corresponding financial parsing field is not parsed. In this case, the embodiment of the invention can realize structured text parsing, that is, the financial voucher data can be parsed into structured financial feature object data.
[0026] Optionally, the interactive dialogue text of an audit object may include, but is not limited to, at least one of the following: dialogue text between the audit object and customer service, dialogue text with the approver, etc.; this embodiment of the invention does not limit this. In other words, an interactive dialogue text may include any dialogue text of an audit object. Optionally, the electronic device may call a text parsing model to parse and process the interactive dialogue text to extract the user's (i.e., the audit object, also known as the auditee) self-narrated text intent from the interactive dialogue text to obtain the dialogue extraction result. Optionally, a dialogue extraction result may include the field content of each dialogue extraction field in at least one dialogue extraction field; optionally, the field content of a dialogue extraction field may be empty, such as when no corresponding field content is extracted. Optionally, at least one dialogue extraction field may be set according to experience or actual needs, and this embodiment of the invention does not limit this; for example, at least one dialogue extraction field may include, but is not limited to, at least one of the following: self-narrated occupation, income declaration, and anti-fraud screening fields, etc.
[0027] Optionally, when a training sample includes financial voucher data and interactive dialogue text of an audit object, the electronic device can traverse each training sample in the training sample set and take the currently traversed training sample as the current training sample. Then, it can call a text parsing model to parse the financial voucher data in the current training sample to obtain the financial feature object data of the current training sample. It can also call a text parsing model to parse the interactive dialogue text in the current training sample based on preset dialogue extraction instructions to obtain the dialogue extraction result of the current training sample. Based on this, the financial feature object data and dialogue extraction result of the current training sample can be integrated to obtain the parsed data of the current training sample.
[0028] Optionally, the text parsing model may include a first text parsing model and a second text parsing model; optionally, the first text parsing model and the second text parsing model may be the same or different, and this embodiment of the invention does not limit this. Based on this, when calling the text parsing model to parse and process the financial voucher data in the current training sample to obtain the financial feature object data of the current training sample, the electronic device may call the first text parsing model to parse and process the financial voucher data in the current training sample to obtain the financial feature object data of the current training sample; correspondingly, when calling the text parsing model to parse and process the interactive dialogue text in the current training sample based on preset dialogue extraction instructions to obtain the dialogue extraction result of the current training sample, the second text parsing model may be called to parse and process the interactive dialogue text in the current training sample based on preset dialogue extraction instructions to obtain the dialogue extraction result of the current training sample.
[0029] Optionally, the preset dialogue extraction instructions can be set based on experience or actual needs, and this embodiment of the invention does not limit this. For example, the preset dialogue extraction instructions can be used to instruct parsing processing according to at least one dialogue extraction field (e.g., the preset dialogue extraction instructions can be dialogue extraction prompt text to prompt the extraction of the field content of at least one dialogue extraction field), to extract the field content of each dialogue extraction field from the current training sample, thereby obtaining the field content of each dialogue extraction field in the current training sample. That is, the dialogue extraction result of the current training sample may include the field content of each dialogue extraction field in the current training sample. Correspondingly, the financial feature object data of the current training sample may include the field content of each financial parsing field in the current training sample. Optionally, the electronic device can also parse the financial voucher data in the current training sample based on the preset financial parsing instructions. The preset financial parsing instructions can be used to instruct parsing processing according to at least one financial parsing field, or simply input the financial voucher data in the current training sample into the text parsing model, and this embodiment of the invention does not limit this. Optionally, the preset financial parsing instructions can be set based on experience or actual needs, and this embodiment of the invention does not limit this, such as the preset financial parsing instructions being financial parsing prompt text, etc.
[0030] Therefore, when a training sample includes the financial voucher data and interactive dialogue text of an audit object, the parsed data of the current training sample may include the financial feature object data and dialogue extraction results of the current training sample.
[0031] Optionally, the electronic device may also include a multi-agent collaborative reasoning layer (e.g., a risk audit system may include a multi-agent collaborative reasoning layer). This multi-agent collaborative reasoning layer may include, but is not limited to, at least one of the following: a text parsing model (also referred to here as a text parsing module, which can be used to call the text parsing model), a logic auditing model (also referred to here as a logic auditing module, which can be used to call the logic auditing model), and a cross-validation orchestrator, etc. This embodiment of the invention does not limit the specifics of these components. The multi-agent collaborative reasoning layer can be used to execute the entire process from text parsing to logical deduction. Based on this, the electronic device can use the multi-agent collaborative reasoning layer to call the text parsing model to parse and process each training sample in the training sample set, obtaining the parsed data for each training sample.
[0032] In summary, electronic devices can achieve feature extraction through text parsing models. This means they can parse financial voucher data and / or interactive dialogue text to output structured features; that is, the parsed data can be structured. For example, the parsed data may include, but is not limited to, the content of each field in at least one field. At least one field (such as at least one financial parsing field and / or at least one dialogue extraction field) may include, but is not limited to, at least one of the following: average monthly income, frequency of sensitive transactions, loan amount, occupation, social security status, and number of loans, etc.
[0033] S103, invoke the logical audit model, and generate G inference output results for each training sample based on the parsed data of each training sample, where G is an integer greater than 1; wherein, each inference output result includes a thought process and a risk level conclusion, and the thought process is used to indicate the reason for inferring a risk level conclusion through textual evidence.
[0034] Optionally, a logic auditing model can be any large language model; that is, the embodiment of this invention does not limit the model structure of the logic auditing model. In this embodiment, the logic auditing model can also be referred to as a logic auditing expert agent. Based on this, this embodiment can use policy sampling to call the logic auditing model to generate G inference output results in parallel.
[0035] Optionally, a thought process may also be referred to as thought process text, thought process data, or thought chain, etc., and this embodiment of the invention does not limit this. A thought process may include, but is not limited to, at least one of the following: textual evidence leading to a risk level result, reasoning process, etc., and this embodiment of the invention does not limit this. Optionally, the reasoning output result may also be referred to as an audit reasoning chain, etc.
[0036] In this embodiment of the invention, the logic auditing model can receive JSON features (i.e., interpretation data) from the text parsing model to generate inference output results. Optionally, for any training sample in the training sample set, the electronic device can call the logic auditing model to directly generate G inference output results for any training sample based on the parsed data of any training sample. For example, the parsed data of any training sample can be input into the logic auditing model to generate G inference output results for any training sample through the logic auditing model; or, the electronic device can call the logic auditing model to generate G inference output results for any training sample using the parsed data of any training sample and the original dialogue text (i.e., the interactive dialogue text in any training sample). For example, the parsed data of any training sample and the interactive dialogue text can be input into the logic auditing model to generate G inference output results for any training sample through the logic auditing model, and so on. This embodiment of the invention does not limit the specific content of the input logic auditing model. For any training sample, the data input into the logic auditing model may include, but is not limited to, at least one of the following: the parsed data of any training sample and the interactive dialogue text, etc. It should be noted that each inference output must follow a preset output format; optionally, the preset output format can be set according to experience or actual needs, and this embodiment of the invention does not limit this; for example, the preset output format can be in the format of <thinking process>...<conclusion>, so that an inference output includes, but is not limited to, a thinking process and a risk level conclusion, etc. Based on this, an inference output can demonstrate how a risk level conclusion is derived from textual evidence through the thinking process. It can be seen that this embodiment of the invention can construct a panoramic textual evidence chain, that is, through multi-agent collaboration, it can break down the semantic barriers between different text sources and achieve deep cross-verification of financial facts (i.e., financial voucher data) and self-narrated information (i.e., interactive dialogue text).
[0037] Optionally, when calling the logical audit model to generate G inference outputs for each training sample based on the parsed data of each training sample, the electronic device determines G temperature parameter values for any training sample in the training sample set; and calls the logical audit model to generate G inference outputs for any training sample based on the parsed data of any training sample and each of the G temperature parameter values; wherein, one inference output of any training sample can be generated based on the parsed data of any training sample and one of the G temperature parameter values, and the G inference outputs of any training sample can include the inference outputs of any training sample under each temperature parameter value. Optionally, the electronic device can input the parsed data of any training sample into the logic audit model to output the inference output results of any training sample under various temperature parameter values; or, it can input the parsed data of any training sample and the interactive dialogue text (i.e., the interactive dialogue text in any training sample) into the logic audit model to output the inference output results of any training sample under various temperature parameter values. In this case, an inference output result of any training sample can be generated based on the parsed data of any training sample, the interactive dialogue text in any training sample, and one of the G temperature parameter values, and so on.
[0038] Optionally, the G temperature parameter values can be set according to experience or actual needs, or they can be randomly generated. This embodiment of the invention does not limit this. Optionally, the G temperature parameter values corresponding to different training samples (the G temperature parameter values corresponding to a training sample can refer to the G temperature parameter values used to generate the G inference output results of the corresponding training sample) can be the same or different. This embodiment of the invention does not limit this.
[0039] As can be seen, the embodiments of the present invention can introduce appropriate temperature parameters during the reasoning generation process, so that the G reasoning output results are different on the reasoning path. For example, the focus can be on analyzing the stability of the flow, or on analyzing the contradictions in the dialogue, and so on.
[0040] Based on this, embodiments of the present invention can apply a scaling factor to the output probability distribution using a temperature parameter during the generation of the logic auditing model (i.e., the temperature parameter can be a scaling factor used to adjust the output probability distribution of the logic auditing model) to control the randomness and diversity of the generated results. This can be achieved by assigning different temperature parameter values to adjust the output probability distribution, thereby sampling the next word according to the adjusted probability, and so on. For example, the output probability can be adjusted to the ratio between the original logarithmic probability and the temperature parameter. In other words, embodiments of the present invention can scale the probability (such as the logarithmic probability) output by the logic auditing model using G temperature parameter values to generate G inference output results for any training sample, thereby controlling the randomness and / or diversity of the G inference output results for any training sample. For example, the smaller the temperature parameter value, the more stable the generated inference output results tend to be; the larger the temperature parameter value, the more random the generated inference output results, and so on.
[0041] Optionally, the electronic device can also invoke the logic auditing model through the multi-agent collaborative reasoning layer to generate G inference outputs for each training sample based on the parsed data of each training sample.
[0042] Optionally, when the multi-agent collaborative reasoning layer also includes a cross-validation orchestrator, the cross-validation orchestrator can adopt rule engine technology as the command center of the risk audit system. That is, the cross-validation orchestrator can be responsible for receiving text audit requests (i.e. risk audit requests) and can dynamically schedule the text parsing model to extract features, and then pass the features (i.e. parsed data) to the logical audit model for risk audit, or pass them to the audit application model below for risk audit. The audit application model can also be a logical audit model.
[0043] S104: Based on the G inference outputs of each training sample, determine the comprehensive reward of each inference output in the G inference outputs of each training sample; and based on the comprehensive reward of each inference output in each training sample, calculate the model loss value, thereby optimizing the model parameters in the logic audit model in the direction of reducing the model loss value, until the target convergence condition is reached, and take the logic audit model that reaches the target convergence condition as the target logic audit model.
[0044] Optionally, when calculating the model loss value based on the comprehensive reward of each inference output result of each training sample, for any training sample in the training sample set, the electronic device can calculate the relative advantage value of the G inference output results of any training sample based on the comprehensive reward of each inference output result of any training sample; after obtaining the relative advantage value of the G inference output results of each training sample, the model loss value is calculated based on the relative advantage value of the G inference output results of each training sample. For example, when calculating the model loss value based on the relative advantage values of the G inference outputs of each training sample, the sub-model loss value under any training sample can be calculated based on the relative advantage values of the G inference outputs of any training sample; after obtaining the sub-model loss values under each training sample, the model loss value can be calculated based on the sub-model loss values under each training sample, and so on; for example, the mean of the sub-model loss values under each training sample can be calculated, or the summation of the sub-model loss values under each training sample can be calculated, or the expected value between the sub-model loss values under each training sample can be negative to obtain the model loss value, and so on; the embodiments of the present invention do not limit this.
[0045] Accordingly, the electronic device can calculate the model loss value based on the relative advantage value and comprehensive reward of the G inference outputs of each training sample (the comprehensive reward of the G inference outputs of a training sample may include the comprehensive reward of each inference output of the corresponding training sample). Based on this, embodiments of the present invention can optimize the model parameters in the logic auditing model through the GRPO (Group Relative Policy Optimization) algorithm; in other words, the model loss value can be calculated based on the GRPO objective function, thereby enabling the updating of policy model parameters through the GRPO objective function, i.e., updating the policy model parameters in the logic auditing model; for example, the electronic device can calculate the model loss value based on the comprehensive reward of each inference output of each training sample using the GRPO algorithm.
[0046] Optionally, when calculating the relative advantage value of the G inference outputs of any training sample based on the comprehensive reward of each inference output of any training sample, the electronic device may use the comprehensive reward of each inference output of any training sample to calculate the mean of the within-group reward (i.e., the mean of the comprehensive reward of each inference output of any training sample) and the within-group standard deviation (i.e., the standard deviation between the comprehensive rewards of each inference output of any training sample); and may calculate the relative advantage value of the G inference outputs of any training sample based on the comprehensive reward of each inference output of any training sample, the mean of the within-group reward, and the within-group standard deviation of any training sample. The relative advantage value of the G inference outputs of any training sample may include the relative advantage value of each inference output of any training sample; correspondingly, for any inference output of any of the G inference outputs of any training sample, the relative advantage value of any inference output of any training sample may be calculated using the comprehensive reward of any inference output of any training sample, the mean of the within-group reward, and the within-group standard deviation of any training sample. For example, the relative advantage value of any inference output can be: (the total reward of any inference output - the mean of the within-group reward under any training sample) / the within-group standard deviation under any training sample, etc.
[0047] For example, an electronic device can use Formula 1.1 to calculate the GRPO objective function value: Formula 1.1 Among them, J GRPO (θ) can represent the GRPO objective function value, E can represent the expectation (i.e., the expectation of the sub-model loss value for each training sample), and π θ (o i |q) can represent the probability of the i-th inference output of any training sample under the current policy (i.e., the probability of the i-th inference output under the current policy), O i π can represent the i-th inference output, q can represent the model input (such as the parsed data of any training sample and / or interactive dialogue text), and π can represent the model input (such as the parsed data of any training sample and / or interactive dialogue text). old (o i |q) can represent the probability of the i-th inference output of any training sample under the old policy (i.e., the probability of the i-th inference output under the old policy), π θ (o i |q) / π old (o i |q) can represent the probability ratio of the i-th inference output, and clip(.) above can represent the clipping probability ratio of the i-th inference output. A iLet ε represent the relative advantage value of the i-th inference output, and let ε represent the pruning coefficient. Based on this, embodiments of the present invention can update the policy model parameters through the GRPO objective function. Specifically, after calculating the GRPO objective function value based on the relative advantage values of the G inference outputs for each training sample, the model loss value is determined based on the GRPO objective function value. For example, the negative value of the GRPO objective function can be used as the model loss value, or a penalty term (such as a KL divergence regularization term) can be determined, and the model loss value is determined using the GRPO objective function value and the penalty term (e.g., the sum of the negative value of the GRPO objective function and the penalty term), etc. Embodiments of the present invention do not limit this. Here, the new policy can refer to the model to be updated, and the old policy can refer to the old model before the update. For example, the sub-model loss value under any training sample can be the negative of the mean of the result loss values under each inference output result of any training sample. The result loss value under an inference output result can be the minimum of the product of the relative advantage value and probability ratio of the corresponding inference output result and the product of the relative advantage value and pruning probability ratio of the corresponding inference output result. Furthermore, the probability ratio of an inference output result can be the probability of the new strategy result of the corresponding inference output result (i.e., the probability of the corresponding inference output result under the new strategy) / the probability of the old strategy result of the corresponding inference output result. The pruning probability ratio of an inference output result can be the pruning result of the probability ratio of the corresponding inference output result (the probability ratio can be pruned to a set range). Optionally, the pruning coefficient can be set according to experience or according to actual needs, and this embodiment of the invention does not limit this.
[0048] Optionally, the electronic device can iteratively execute the aforementioned call logic auditing model, generating G inference outputs for each training sample based on the parsed data of each training sample, thereby continuously determining the model loss value to optimize the model parameters in the logic auditing model until the target convergence condition is reached. Optionally, the target convergence condition can be set based on experience or actual needs, and this embodiment of the invention does not limit this; for example, the target convergence condition can refer to the number of iterations reaching a preset iteration number threshold, or it can refer to the model loss value being less than a preset model loss threshold, etc. Optionally, both the preset iteration number threshold and the preset model loss threshold can be set based on experience or actual needs, and this embodiment of the invention does not limit this.
[0049] Optionally, the electronic device may also include a reinforcement learning optimization and closed-loop layer (e.g., a risk audit system may include a reinforcement learning optimization and closed-loop layer). In this case, the electronic device can optimize the logical audit model through the reinforcement learning optimization and closed-loop layer. Optionally, the electronic device may also use the target logical audit model as the initial logical audit model at preset intervals to continue optimizing the logical audit model, thereby updating the target logical audit model and subsequently updating the audit application model, etc. Optionally, the preset interval may be set based on experience or actual needs, and this embodiment of the invention does not limit this. Alternatively, the electronic device may also trigger the execution of the above-mentioned use of the target logical audit model as the initial logical audit model to continue optimizing the logical audit model when a logical audit model continues optimization instruction is detected; for example, the detection of a logical audit model continues optimization instruction may be determined at preset intervals, or when a logical audit model continues optimization operation executed by an administrator may be detected, etc.
[0050] S105, Based on the target logic audit model, determine the audit application model, which supports the use of risk auditing.
[0051] Optionally, the electronic device can also use a cross-validation orchestrator to summarize the outputs of multiple agents in a consistent manner, identify contradictions between the audit thought chain (i.e. the thinking process in the reasoning output) and textual facts (such as financial feature object data), and output the final approval recommendation (i.e. decision result).
[0052] Optionally, upon detecting a risk audit request, the electronic device can also acquire the data to be audited as indicated in the risk audit request. This data may include financial document data and / or interactive dialogue text of the target audit object indicated in the risk audit request. The target audit object can be any audit object. Then, the electronic device can invoke a text parsing model to parse the data to be audited, obtaining parsed data of the target audit object. This allows it to invoke an audit application model and, based on the parsed data of the target audit object, generate an inference output result for the target audit object. Optionally, the data to be audited may be data that has undergone data preprocessing. The inference output result of the target audit object can be used to guide credit approval for the target audit object, etc. For example, the risk level conclusion in the inference audit result can be used to guide credit approval; for instance, the risk level conclusion may be high risk, low risk, etc. This embodiment of the invention does not limit the specific representation of the risk level conclusion.
[0053] Optionally, the electronic device can also use a cross-validation orchestrator to perform cross-validation based on the thought process and financial characteristic object data (which can be the result of parsing and processing the financial voucher data of the target audit object) in the reasoning output of the target audit object, so as to obtain the approval suggestion instruction result. For example, the cross-validation orchestrator may include a cross-validation orchestration model (which can be any large language model). Then the electronic device can call the cross-validation orchestration model to perform cross-validation (such as verifying data consistency) on the thought process and financial characteristic object data in the reasoning output of the target audit object, and obtain the verification result of the target audit object. Based on the verification result (which can be used to indicate whether cross-validation is passed), the approval suggestion instruction result is determined. For example, the reasoning output and verification results of the target audit object can be added to the approval suggestion instruction result to obtain the approval suggestion instruction result; alternatively, when the verification result indicates that cross-validation has been passed (e.g., the thought process is consistent with the financial characteristic object data), the reasoning output result of the target audit object or the risk level conclusion in the reasoning output result of the target audit object can be added to the approval suggestion instruction result; alternatively, when the verification result indicates that cross-validation has not been passed, the failure verification instruction information can be used as the approval suggestion instruction result; alternatively, the decision result (e.g., pass or reject, which can be determined based on the risk level conclusion) can be added to the approval suggestion instruction result, and so on; the embodiments of the present invention do not limit this. Optionally, the failure verification instruction information can be set according to experience or actual needs, and the embodiments of the present invention do not limit this.
[0054] As can be seen, the embodiments of the present invention can realize an unstructured text credit risk audit architecture based on multi-agent collaboration. This architecture achieves an automated closed loop from extracting facts from multi-source text to logical deduction through the division of labor and cooperation among text parsing agents (i.e., text parsing models), logic auditing agents (i.e., logic auditing models), and cross-validation orchestrators. Furthermore, the embodiments of the present invention can implement a credit reasoning logic training method based on the GRPO algorithm. This method utilizes group relative advantage computation to replace traditional value networks, training models for pure text credit cases to generate highly interpretable thought chain audit reports, thus generating thought processes and effectively improving the accuracy of reasoning.
[0055] This invention provides an embodiment of a training sample set. Each training sample includes at least one of the following: financial voucher data and interactive dialogue text of an audit object. A text parsing model is invoked to parse each training sample in the training sample set, obtaining parsed data for each training sample. Then, a logical auditing model is invoked to generate G inference outputs for each training sample based on the parsed data, where G is an integer greater than 1. Each inference output includes a thought process and a risk level conclusion. The thought process indicates the reasoning for inferring a risk level conclusion from textual evidence. Based on this, the comprehensive reward of each inference output in each training sample can be determined based on the G inference outputs. The model loss value is calculated based on the comprehensive reward of each inference output, thereby optimizing the model parameters in the logical auditing model in the direction of reducing the model loss value until the target convergence condition is reached. The logical auditing model that reaches the target convergence condition is then used as the target logical auditing model. Furthermore, an audit application model can be determined based on the target logical auditing model, which supports risk auditing. As can be seen, embodiments of the present invention can train a logical audit model using a training sample set containing financial voucher data and / or interactive dialogue text, thereby obtaining an audit application model with high risk audit performance, which can improve the accuracy of reasoning output results, thereby improving the accuracy of the thinking process and risk level conclusions, and thus effectively improving the accuracy of risk audit.
[0056] Based on the above description, this embodiment of the invention also proposes another risk auditing method. Accordingly, this risk auditing method can be executed by the aforementioned electronic device (terminal or server); or, the risk auditing method can be executed jointly by the terminal and the server, etc.; that is, the risk auditing method can be executed by a risk auditing system on the electronic device, and so on. For ease of explanation, the following description will use the execution of this risk auditing method by an electronic device as an example; please refer to... Figure 3 This risk audit method may include the following steps S301-S307: S301, Obtain a training sample set, wherein a training sample includes at least one of the following: financial voucher data and interactive dialogue text of an audit object.
[0057] S302, call the text parsing model to parse each training sample in the training sample set to obtain the parsed data of each training sample.
[0058] S303, invoke the logic audit model to generate G inference output results for each training sample based on the parsed data of each training sample.
[0059] S304, for any training sample in the training sample set, and any inference output result among the G inference output results of any training sample, determine the inference content verification reward of any inference output result based on the parsed data of any training sample and any inference output result.
[0060] Among them, the reasoning content verification reward can also be called the consistency reward.
[0061] Optionally, the electronic device can utilize a pre-trained natural language reasoning model as a validator, taking the parsed "financial facts" (i.e., financial feature object data) as premises and the "statements" (i.e., thought processes) in the reasoning chain (i.e., reasoning output results) as assumptions, and calculating an implication score (i.e., the reasoning content verification reward for any reasoning output result). Based on this, if the data cited in the reasoning chain (i.e., evidence cited in the thought process) is consistent with the document evidence (i.e., financial feature object data), a positive reward can be given; if illusions occur (such as citing non-existent transaction records) or logical contradictions occur, penalties are given (e.g., the reward can be negative), and so on.
[0062] Accordingly, when determining the reasoning content verification reward for any reasoning output based on the parsed data of any training sample, the electronic device can invoke the reasoning verification model to determine the reasoning content verification reward for any reasoning output based on the thought process (i.e., thought process data) in any reasoning output and the financial feature object data of any training sample. Specifically, the thought process in any reasoning output and the financial feature object data of any training sample can be input into the reasoning verification model to output the reasoning content verification reward for any reasoning output. The reasoning content verification reward for any reasoning output can be used to indicate the consistency relationship between the thought process in any reasoning output and the financial feature object data of any training sample, i.e., to indicate the consistency relationship with the financial voucher data in any training sample. Optionally, the reasoning verification model can be a natural language reasoning model, i.e., any large language model; the embodiment of this invention does not limit the model structure of the reasoning verification model.
[0063] For example, when the thought process in any inference output is consistent with the financial feature data of any training sample, the inference content verification reward of any inference output can be positive; when the thought process in any inference output is inconsistent with the financial feature data of any training sample, the inference content verification reward of any inference output can be negative. Furthermore, the higher the consistency between the thought process in any inference output and the financial feature data of any training sample, the larger the inference content verification reward of any inference output can be; conversely, the lower the consistency between the thought process in any inference output and the financial feature data of any training sample, the smaller the inference content verification reward of any inference output can be (e.g., the larger the negative value), and so on. This embodiment of the invention does not limit the specific representation of the inference content verification reward.
[0064] S305, Based on the reasoning content verification reward of any reasoning output result, determine the comprehensive reward of any reasoning output result.
[0065] In one implementation, the electronic device may use the reasoning content verification reward of any reasoning output result as the comprehensive reward of any reasoning output result.
[0066] In another embodiment, the electronic device can also check any inference output result according to a preset checking rule to obtain a completeness reward for any inference output result; and determine a comprehensive reward for any inference output result based on the completeness reward and the inference content verification reward. For example, the completeness reward and the inference content verification reward for any inference output result can be weighted and summed to obtain the comprehensive reward for any inference output result. Optionally, the preset checking rule can be set based on experience or based on actual needs; this embodiment of the invention does not limit this.
[0067] Optionally, the preset inspection rules may include at least one inspection item. For example, at least one inspection item may include, but is not limited to, at least one of the following: thought process labeling check (e.g., checking whether the reasoning output includes a thought process), thought chain length check (e.g., checking whether the length of the thought process in the reasoning output is less than a preset thought chain length threshold), income verification check (e.g., checking whether income verification has been completed), liability calculation check (e.g., checking whether liability calculation has been completed), and anti-fraud screening (e.g., checking whether anti-fraud screening has been completed, such as checking whether preset anti-fraud information exists), etc.; this embodiment of the invention does not limit this. Optionally, the preset anti-fraud information may be set according to experience or actual needs, and this embodiment of the invention does not limit this; for example, the preset anti-fraud information may include words such as fraud. Optionally, the preset inspection rules may also include the pass conditions for each inspection item in at least one inspection item, and the pass conditions for an inspection item can be used to indicate the requirements for passing the corresponding inspection item.
[0068] Optionally, when checking any inference output result according to preset checking rules to obtain a completeness reward for any inference output result, for any check item among at least one check item, if any inference output result passes the check item, the first check reward can be used as the check reward for any inference output result under the check item; if any inference output result fails the check item, the second check reward can be used as the check reward for any inference output result under the check item. Based on this, the sum of the check rewards for any inference output result under each check item can be used as the completeness reward for any inference output result. Optionally, both the first check reward and the second check reward can be set according to experience or actual needs, and the embodiments of the present invention do not limit this; optionally, the first check reward can be greater than the second check reward, such as the first check reward can be 1, and the second check reward can be 0 or -1, etc.
[0069] In another implementation, the electronic device can also determine the posterior performance label (also known as the feedback label) of any training sample. The posterior performance label of a training sample can be used to indicate whether the audit object indicated by the corresponding training sample is a bad debt (i.e., whether the subsequent actual performance is a bad debt). In this case, historical cases (such as credit applications from 6-12 months ago) can be used as training samples, and the final performance (whether it is a bad debt) of these cases is known. Based on this, the electronic device can determine the risk posterior reward of any inference output result based on the posterior performance label of any training sample and the risk level conclusion in any inference output result. Thus, based on the inference content verification reward and the risk posterior reward of any inference output result, the comprehensive reward of any inference output result can be determined. For example, the inference content verification reward and the risk posterior reward of any inference output result can be weighted and summed to obtain the comprehensive reward of any inference output result.
[0070] Optionally, the electronic device can determine the risk posterior reward of any inference output result based on the posterior performance label of any training sample and the risk level conclusion in any inference output result, according to a preset risk posterior reward logic. Optionally, the preset risk posterior reward logic can be set according to experience or actual needs, and this embodiment of the invention does not limit it. For example, the preset risk posterior reward logic may include, but is not limited to, at least one of the following: bad debt and high-risk reward logic (i.e., if it is actually a bad debt (i.e., the posterior performance label indicates a bad debt), and the risk level conclusion is high risk or rejection is recommended, a first logical reward score is given, which can be large to give a high reward); bad debt but low-risk reward logic (i.e., if it is actually a bad debt, but the risk level conclusion is low risk, a second logical reward score is given, which can be small, such as negative or 0, to give a heavy penalty); no bad debt but high-risk reward logic (i.e., if it is actually not a bad debt (i.e., the posterior performance label indicates no bad debt or indicates a good debt), but the risk level conclusion is high risk, a third logical reward score is given, which can be small, such as 0 or negative, to give a small penalty), etc. Optionally, the first logical reward score, the second logical reward score, and the third logical reward score can all be set according to experience or actual needs, and this embodiment of the invention does not limit this.
[0071] In another implementation, the electronic device can perform a weighted summation of the completeness reward, the inference content verification reward, and the risk posterior reward for any inference output result to obtain the comprehensive reward for any inference output result, and so on. Based on this, when determining the comprehensive reward for any inference output result based on the inference content verification reward and the risk posterior reward, the electronic device can also check any inference output result according to preset checking rules to obtain the completeness reward for any inference output result, thereby performing a weighted summation of the completeness reward, the inference content verification reward, and the risk posterior reward for any inference output result to obtain the comprehensive reward for any inference output result.
[0072] For example, taking the i-th inference output among G inference outputs for any training sample as an example, the electronic device can use Formula 2.1 to determine the comprehensive reward of the given inference output: Equation 2.1 Among them, o i Let R represent the i-th inference output (which can represent any inference output), then R i The sum of the rewards for any inference output can be represented by w1, w2, and w3, which can represent the weights of each reward, respectively. con (o i) can represent the reasoning content verification reward for any reasoning output result, r fmt (o i ) can represent the completeness reward of any inference output, r risk (o i ,y) can represent the risk posterior reward of any inference output, and y can represent the posterior performance label of any training sample.
[0073] S306, based on the comprehensive reward of each inference output result of each training sample, calculate the model loss value, and then optimize the model parameters in the logic audit model in the direction of reducing the model loss value until the target convergence condition is reached, and take the logic audit model that has reached the target convergence condition as the target logic audit model.
[0074] In this embodiment of the invention, the model parameters in the logical audit model can be optimized through comprehensive rewards, thereby enabling the logical audit model to generate more rigorous and interpretable audit reports (i.e., inference output results) when processing complex textual logic such as "abnormal fund transactions". Furthermore, this embodiment of the invention can combine the logical consistency of textual facts with long-term bad debt results based on a multiple reward mechanism of textual consistency and posterior risk, using verifiable reward learning (RLVR, Reinforcement Learning from Verifiable Rewards), to guide the optimization of model strategies. In other words, this embodiment of the invention can transform long-term bad debt results into immediate reward signals that the model can learn, thereby effectively improving model performance.
[0075] S307, based on the target logic audit model, determines the audit application model, which supports the use of risk auditing.
[0076] In one implementation, the electronic device can use the target logic audit model as an audit application model.
[0077] In another implementation, the electronic device can also use the target logic audit model as a teacher model to distill a student model, and then use the student model as the audit application model. Based on this, the embodiments of the present invention can effectively balance the real-time requirements (low latency) of online inference with the depth requirements (high performance) of risk control logic.
[0078] In this embodiment of the invention, the teacher model has a huge number of parameters (e.g., 70B+), strong thinking ability, but slow reasoning and high cost; it mainly runs in an offline environment, using the GRPO algorithm to conduct in-depth exploration on massive historical data. The student model has a smaller number of parameters (e.g., 7B or 8B), and after quantization, its reasoning speed is fast; it is responsible for online real-time credit approval, which can realize real-time risk auditing. At this time, the audit application model can be used to conduct real-time risk auditing.
[0079] Accordingly, when determining the audit application model based on the target logic audit model, the electronic device can acquire a distillation analysis data set; and call the target logic audit model to generate multiple inference output results under each distillation analysis data set based on each distillation analysis data in the distillation analysis data set. Then, it can traverse each distillation analysis data in the distillation analysis data set, and take the currently traversed distillation analysis data as the current analysis data (also called the current distillation analysis data), and determine the comprehensive reward of each inference output result among the multiple inference output results under the current analysis data. Further, based on the comprehensive reward of each inference output result under the current analysis data, it can determine the inference output result label set from the multiple inference output results under the current analysis data; and combine the current analysis data and each inference output result label in the inference output result label set to form a distillation training data set to obtain the current distillation training data set, and then add the current distillation training data set to the target distillation training data set; after traversing each distillation analysis data in the distillation analysis data set, the audit application model is trained based on the target distillation training data set to determine the audit application model, which is a student model of the target logic audit model.
[0080] Optionally, the electronic device may store a distillation analysis data set in its own storage space. In this case, the electronic device can obtain the distillation analysis data set from its own storage space; or, the electronic device can obtain a download link for the distillation analysis data set and download the distillation analysis data set using the download link, etc.; this embodiment of the invention does not limit this. Optionally, the distillation analysis data in the distillation analysis data set may be the same as (i.e., include) the analysis data of each training sample, or may be different; this embodiment of the invention does not limit this. Correspondingly, a distillation analysis data set may include, but is not limited to, at least one of the following: a financial feature object data and a dialogue extraction result.
[0081] Optionally, the number of inference outputs in a set of multiple inference outputs under a single distillation analysis dataset can be G, or greater or less than G, and this embodiment of the invention does not limit this. Optionally, the distillation analysis dataset may include analysis data of samples whose risk level conclusions indicate high risk; based on this, this embodiment of the invention can use a teacher model to generate inference outputs containing detailed thought processes for historically high-risk complex cases.
[0082] It should be noted that the method for determining the comprehensive reward of an inference output result under the current parsed data can be the same as the method for determining the comprehensive reward of any inference output result of any training sample, and will not be repeated here in this embodiment of the invention. Based on this, the electronic device can score the inference output results generated by the teacher model through the above reward function to achieve quality filtering.
[0083] Optionally, when determining a set of inference output result labels (also referred to as the inference output result label set under the current parsed data) from multiple inference output results based on the comprehensive reward of each inference output result under the current parsed data, the electronic device can determine the inference output result with a comprehensive reward greater than a preset reward threshold from the multiple inference output results under the current parsed data, and add the determined inference output result to the inference output result label set to achieve the determination of the inference output result label set. At this time, the comprehensive reward of one inference output result label is greater than the preset reward threshold; or, the inference output result with the largest comprehensive reward among the multiple inference output results under the current parsed data can be added to the inference output result label set, and so on; the embodiments of the present invention do not limit this. Optionally, the preset reward threshold can be set according to experience or actual needs, and the embodiments of the present invention do not limit this. Based on this, the embodiments of the present invention can retain samples with extremely high comprehensive rewards (such as logically consistent and accurately predicting bad debts) to construct a "Golden Thinking Chain Dataset".
[0084] Optionally, when combining the current parsed data and each inference output result label in the inference output result label set to form a distillation training data set, for any inference output result label in the inference output result label set, the current parsed data and any inference output result label can be combined to form a distillation training data set, resulting in distillation training data under any inference output result label (which may include the current parsed data and any inference output result label). This distillation training data under any inference output result label is then added to the current distillation training data set (also referred to as the distillation training data set under the current parsed data). Based on this, the current distillation training data set may include distillation training data under each inference output result label; correspondingly, the target distillation training data set may include the distillation training data set under each distillation parsed data in the distillation parsed data set. Here, a target distillation training data set includes one parsed data and one inference output result label.
[0085] Furthermore, electronic devices can use the target distillation training dataset to fine-tune the student model, i.e., fine-tune the audit application model.
[0086] In one implementation, when training an audit application model based on a target distillation training dataset, the electronic device can invoke the audit application model to generate inference output results for each target distillation training dataset based on the parsed data in each target distillation training dataset. Based on the inference output results and inference output result labels (the inference output result label for one target distillation training dataset is the inference output result label in the corresponding target distillation training dataset), the model loss value of the audit application model is determined. This optimizes the model parameters of the audit application model in the direction of reducing the model loss value, until the audit application optimization convergence condition is met, thus completing the training of the audit application model and obtaining the final audit application model. Optionally, the audit application optimization convergence condition can be set based on experience or actual needs, and this embodiment of the invention does not limit this. For example, the audit application optimization convergence condition can refer to the number of iterations reaching a preset audit optimization iteration threshold, or it can refer to the model loss value of the audit application model being less than a preset audit optimization loss threshold, etc. Optionally, both the preset audit optimization iteration threshold and the preset audit optimization loss threshold can be set based on experience or actual needs, and this embodiment of the invention does not limit this.
[0087] Optionally, when determining the model loss value of the audit application model based on the inference output results and inference output result labels under each target distillation training data set, a consistency check can be performed on the inference output results and inference output result labels under any target distillation training data set. If the consistency check passes, the first check loss value can be used as the loss value under any target distillation training data set; if the consistency check fails, the second check loss value can be used as the loss value under any target distillation training data set. Then, the sum of the loss values under each target distillation training data set can be used as the model loss value of the audit application model. Optionally, the first check loss value can be less than the second check loss value; optionally, both the first check loss value and the second check loss value can be set according to experience or actual needs, and this embodiment of the invention does not limit this; for example, the first check loss value can be 0, the second check loss value can be 1, etc.
[0088] In another implementation, when training the audit application model based on the target distillation training dataset, the electronic device also invokes the audit application model for any target distillation training data in the target distillation training dataset. Based on the parsed data in each target distillation training data, the inference prediction probability under each target distillation training data is determined. The inference prediction probability under a target distillation training data may include the conditional probability of each word in the inference output result label of the corresponding target distillation training data. Based on this, for any target distillation training data in the target distillation training dataset, the electronic device can use the conditional probability of each word in the inference output result label of any target distillation training data to calculate the loss value under any target distillation training data. Correspondingly, the sum of the loss values under each target distillation training data can be used as the model loss value of the audit application model. This optimizes the model parameters of the audit application model in the direction of reducing the model loss value of the audit application model until the audit application optimization convergence condition is met, and so on. Optionally, when calculating the loss value under any target distillation training data using the conditional probability of each word in the inference output result label of any target distillation training data, the negative of the sum of the conditional probabilities of each word in the inference output result label of any target distillation training data can be used as the loss value under any target distillation training data; or, the negative of the sum of the logarithmic sum of the conditional probabilities of each word in the inference output result label of any target distillation training data can be used as the loss value under any target distillation training data, and so on; the embodiments of the present invention do not limit this.
[0089] Based on this, the student model can not only learn to output the final risk level conclusion, but more importantly, it can imitate the intermediate thinking process of the teacher model.
[0090] Optionally, when conducting risk audits through the audit application model, the electronic device may output corresponding reasoning results, or it may only output risk level conclusions; and / or, it may also determine the approval result (also known as the decision result) based on the risk level conclusions and output the approval result, such as the approval result being pass or rejection, etc.; the embodiments of the present invention do not limit this.
[0091] Optionally, audit logs and decision results generated online by the student model can be recorded. Over time, the actual repayment performance (true value) corresponding to these decisions gradually emerges (e.g., T+6 months can achieve T+N feedback, where N can be 6). Then, this new data with true labels (i.e., posterior performance labels) flows back into the training pool, which can become the fuel for the teacher model's next round of GRPO training, forming a positive reinforcement loop of "data-model-business" to achieve a data-driven closed loop.
[0092] In summary, this invention can efficiently train the logical reasoning ability of a logical auditing model using the GRPO algorithm, generating a text audit report (i.e., the reasoning output) containing a complete thought chain, thereby effectively improving the accuracy and interpretability of reasoning. Furthermore, based on the reinforcement learning concept of verifiable rewards, it utilizes bad debt facts in credit business as true labels and designs a composite reward function including text consistency components and risk posterior components, thus more accurately guiding model optimization. Additionally, this invention can leverage a teacher-student distillation-based self-evolving credit risk control model system, using the thought chain data produced by the strong reasoning model to drive the iteration of a lightweight model (i.e., this invention can distill to a lightweight model deployment), achieving low-cost, high-precision real-time auditing in complex text environments. This distillation architecture enables low-latency, low-cost online risk auditing, optimizing deployment efficiency. In other words, this invention can inject high-performance auditing thinking into a lightweight model through a "teacher-student" distillation mechanism, achieving a closed-loop data-driven iteration of the system. Based on this, this invention significantly improves the accuracy of risk identification in credit business and provides highly interpretable text audit reports.
[0093] This invention, after obtaining a training sample set, calls a text parsing model to parse each training sample in the set, obtaining parsed data for each training sample. Then, it calls a logic auditing model to generate G inference outputs for each training sample based on the parsed data. Next, for any training sample in the set and any one of the G inference outputs, a content verification reward is determined based on the inference output and the parsed data. A comprehensive reward is then determined based on the content verification reward. Furthermore, the model loss value is calculated based on the comprehensive reward of each inference output for each training sample. This allows for optimization of the model parameters in the logic auditing model in the direction of reducing the model loss value, until the target convergence condition is met. The logic auditing model that meets the target convergence condition is then used as the target logic auditing model. Finally, an audit application model is determined based on the target logic auditing model, which supports risk auditing. As can be seen, the embodiments of the present invention can continuously improve the risk identification capability of the intelligent agent based on business feedback through the GRPO algorithm, reward engine and knowledge distillation module, thereby continuously improving the accuracy of risk auditing; and, it can enable the audit application model to inherit the expert intuition and logical patterns of the target logical audit model, enabling it to quickly identify complex risks online without the need for expensive online reinforcement learning inference, thus effectively saving computing resources and providing real-time risk auditing capabilities.
[0094] Based on the description of the relevant embodiments of the above risk auditing method, this embodiment of the invention also proposes a risk auditing device, which can be a computer program (including program code) running in an electronic device; such as Figure 4 As shown, the risk audit device may include an acquisition unit 401 and a processing unit 402. The risk audit device can perform... Figure 1 or Figure 3 The risk audit method shown, i.e., the risk audit device can operate the above-mentioned unit: Acquisition unit 401 is used to acquire a training sample set, wherein a training sample includes at least one of the following: financial voucher data and interactive dialogue text of an audit object; The processing unit 402 is used to call the text parsing model to parse each training sample in the training sample set to obtain the parsed data of each training sample. The processing unit 402 is further configured to invoke the logical audit model to generate G inference output results for each training sample based on the parsed data of each training sample, where G is an integer greater than 1; wherein, each inference output result includes a thought process and a risk level conclusion, and the thought process is used to indicate the reasons for inferring a risk level conclusion through textual evidence. The processing unit 402 is further configured to determine the comprehensive reward of each inference output result among the G inference output results of each training sample based on the G inference output results of each training sample; and calculate the model loss value based on the comprehensive reward of each inference output result of each training sample, thereby optimizing the model parameters in the logic audit model in the direction of reducing the model loss value, until the target convergence condition is reached, so as to take the logic audit model that reaches the target convergence condition as the target logic audit model; The processing unit 402 is further configured to determine an audit application model based on the target logical audit model, wherein the audit application model supports risk auditing.
[0095] In one implementation, a training sample includes financial voucher data and interactive dialogue text of an audit object; when the processing unit 402 calls the text parsing model to parse each training sample in the training sample set to obtain the parsed data of each training sample, it can be specifically used for: Iterate through each training sample in the training sample set and take the currently iterated training sample as the current training sample; The text parsing model is invoked to parse and process the financial voucher data in the current training sample, thereby obtaining the financial feature object data of the current training sample. The text parsing model is invoked, and based on the preset dialogue extraction instructions, the interactive dialogue text in the current training sample is parsed and processed to obtain the dialogue extraction result of the current training sample. The financial feature object data and dialogue extraction results of the current training sample are integrated to obtain the parsed data of the current training sample.
[0096] In another implementation, when the processing unit 402 invokes the logic auditing model to generate G inference outputs for each training sample based on the parsed data of each training sample, it can specifically be used for: For any training sample in the training sample set, determine G temperature parameter values; The logical audit model is invoked to generate G inference output results for any training sample based on the parsed data of any training sample and each of the G temperature parameter values. The inference output of any training sample is generated based on the parsed data of the training sample and one of the G temperature parameter values.
[0097] In another implementation, when determining the comprehensive reward of each inference output result among the G inference output results of each training sample based on the G inference output results of each training sample, the processing unit 402 may specifically be used for: For any training sample in the training sample set, and any inference output result among the G inference output results of any training sample, the inference content verification reward of any inference output result is determined based on the parsed data of any training sample and the inference output result. Based on the reasoning content verification reward of any of the reasoning output results, determine the comprehensive reward of any of the reasoning output results.
[0098] In another implementation, when determining the comprehensive reward of any inference output result based on the inference content verification reward of any inference output result, the processing unit 402 may specifically be used to: Determine the posterior performance label for any of the training samples; Based on the posterior performance label of any training sample and the risk level conclusion in any inference output, determine the risk posterior reward of any inference output. The comprehensive reward for any given inference output is determined based on the inference content verification reward and risk posterior reward.
[0099] In another implementation, when calculating the model loss value based on the comprehensive reward of each inference output result of each training sample, the processing unit 402 may specifically be used for: For any training sample in the training sample set, the relative advantage value of the G inference outputs of any training sample is calculated based on the comprehensive reward of each inference output of the training sample. After obtaining the relative advantage values of the G inference outputs for each training sample, the model loss value is calculated based on the relative advantage values of the G inference outputs for each training sample.
[0100] In another implementation, when determining the audit application model based on the target logical audit model, the processing unit 402 may specifically be used to: Obtain the distillation analysis data set; The target logic audit model is invoked to generate multiple inference output results for each distillation analysis data in the distillation analysis data set. Traverse each distillation analysis data in the distillation analysis data set, and take the currently traversed distillation analysis data as the current analysis data, and determine the comprehensive reward of each inference output result among the multiple inference output results under the current analysis data; Based on the comprehensive reward of each inference output result under the current parsed data, a set of inference output result labels is determined from multiple inference output results under the current parsed data; and each inference output result label in the current parsed data and the set of inference output result labels is respectively composed of a distillation training data to obtain the current distillation training data set, and then the current distillation training data set is added to the target distillation training data set; After traversing through all the distillation analysis data in the distillation analysis data set, an audit application model is trained based on the target distillation training data set to determine the audit application model, which is a student model of the target logical audit model.
[0101] According to one embodiment of the present invention, Figure 4 Each unit in the risk audit device shown can be individually or entirely combined into one or more other units, or one or more of the units can be further divided into multiple functionally smaller units. This achieves the same operation without affecting the technical effect of the embodiments of the present invention. The above units are divided based on logical functions. In practical applications, the function of one unit can be implemented by multiple units, or the function of multiple units can be implemented by one unit. In other embodiments of the present invention, any risk audit device may also include other units. In practical applications, these functions can also be implemented with the assistance of other units, and can be implemented collaboratively by multiple units.
[0102] According to another embodiment of the present invention, it is possible to perform operations such as those described above by running on a general-purpose electronic device, such as a computer, which includes processing elements and storage elements such as a central processing unit (CPU), random access memory (RAM), and read-only memory (ROM). Figure 1 or Figure 3 The computer program (including program code) involved in each step of the corresponding method shown, to construct such... Figure 4 The risk audit apparatus shown herein, and the risk audit method for implementing embodiments of the present invention, are described. The computer program may be recorded on, for example, a computer storage medium, loaded onto the aforementioned electronic device via the computer storage medium, and run therein.
[0103] Based on the description of the method and apparatus embodiments above, an exemplary embodiment of the present invention also provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program executable by the at least one processor, which, when executed by the at least one processor, causes the electronic device to perform the method according to an embodiment of the present invention.
[0104] An exemplary embodiment of the present invention also provides a non-transitory computer-readable storage medium storing a computer program, wherein the computer program, when executed by a computer's processor, is used to cause the computer to perform a method according to an embodiment of the present invention.
[0105] An exemplary embodiment of the present invention also provides a computer program product, including a computer program, wherein, when executed by a computer's processor, the computer program is used to cause the computer to perform a method according to an embodiment of the present invention.
[0106] refer to Figure 5 The present invention will now be described in the form of a structural block diagram of an electronic device 500 that can serve as a server or client of the present invention, which is an example of a hardware device that can be applied to various aspects of the present invention. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0107] like Figure 5 As shown, the electronic device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. The RAM 503 may also store various programs and data required for the operation of the electronic device 500. The computing unit 501, ROM 502, and RAM 503 are interconnected via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0108] Multiple components in electronic device 500 are connected to I / O interface 505, including: input unit 506, output unit 507, storage unit 508, and communication unit 509. Input unit 506 can be any type of device capable of inputting information to electronic device 500. Input unit 506 can receive input digital or character information and generate key signal inputs related to user settings and / or function control of electronic device. Output unit 507 can be any type of device capable of presenting information and may include, but is not limited to, a display, speaker, video / audio output terminal, vibrator, and / or printer. Storage unit 508 may include, but is not limited to, disk and optical disk. Communication unit 509 allows electronic device 500 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks, and may include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers, and / or chipsets, such as Bluetooth™ devices, WiFi devices, WiMax devices, cellular communication devices, and / or the like.
[0109] The computing unit 501 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 501 performs the various methods and processes described above. For example, in some embodiments, the risk auditing method can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 500 via ROM 502 and / or communication unit 509. In some embodiments, the computing unit 501 can be configured to perform the risk auditing method by any other suitable means (e.g., by means of firmware).
[0110] The program code used to implement the methods of the present invention can be written in any combination of one or more programming languages. This program code can be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code can be executed entirely on the machine, partially on the machine, as a standalone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0111] In the context of this invention, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0112] As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, device, and / or apparatus (e.g., disk, optical disk, memory, programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including machine-readable media that receive machine instructions as machine-readable signals. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.
[0113] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0114] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0115] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other.
[0116] Furthermore, it should be understood that the above-disclosed embodiments are merely preferred embodiments of the present invention and should not be construed as limiting the scope of the present invention. Therefore, any equivalent variations made in accordance with the claims of the present invention are still within the scope of the present invention.
Claims
1. A risk auditing method, characterized in that, include: Obtain a training sample set, wherein a training sample includes at least one of the following: financial voucher data and interactive dialogue text of an audit object; The text parsing model is invoked to parse each training sample in the training sample set to obtain the parsed data of each training sample. The logical audit model is invoked to generate G inference outputs for each training sample based on the parsed data of each training sample, where G is an integer greater than 1. Each inference output includes a thought process and a risk level conclusion. The thought process is used to indicate the reasons for inferring a risk level conclusion from textual evidence. Based on the G inference outputs of each training sample, the comprehensive reward of each inference output among the G inference outputs of each training sample is determined; and based on the comprehensive reward of each inference output of each training sample, the model loss value is calculated, thereby optimizing the model parameters in the logic audit model in the direction of reducing the model loss value, until the target convergence condition is reached, so that the logic audit model that reaches the target convergence condition is taken as the target logic audit model; Based on the target logic audit model, an audit application model is determined, which supports risk auditing.
2. The method according to claim 1, characterized in that, A training sample includes financial voucher data and interactive dialogue text of an audit object; the text parsing model is invoked to parse each training sample in the training sample set to obtain parsed data for each training sample, including: Iterate through each training sample in the training sample set and take the currently iterated training sample as the current training sample; The text parsing model is invoked to parse and process the financial voucher data in the current training sample, thereby obtaining the financial feature object data of the current training sample. The text parsing model is invoked, and based on the preset dialogue extraction instructions, the interactive dialogue text in the current training sample is parsed and processed to obtain the dialogue extraction result of the current training sample. The financial feature object data and dialogue extraction results of the current training sample are integrated to obtain the parsed data of the current training sample.
3. The method according to claim 1 or 2, characterized in that, The call logic auditing model generates G inference outputs for each training sample based on the parsed data of each training sample, including: For any training sample in the training sample set, determine G temperature parameter values; The logical audit model is invoked to generate G inference output results for any training sample based on the parsed data of any training sample and each of the G temperature parameter values. The inference output of any training sample is generated based on the parsed data of the training sample and one of the G temperature parameter values.
4. The method according to claim 1 or 2, characterized in that, The step of determining the comprehensive reward of each inference output result among the G inference output results of each training sample, based on the G inference output results of each training sample respectively, includes: For any training sample in the training sample set, and any inference output result among the G inference output results of any training sample, the inference content verification reward of any inference output result is determined based on the parsed data of any training sample and the inference output result. Based on the reasoning content verification reward of any of the reasoning output results, determine the comprehensive reward of any of the reasoning output results.
5. The method according to claim 4, characterized in that, The reasoning content verification reward based on any of the reasoning output results, determining the comprehensive reward for any of the reasoning output results, includes: Determine the posterior performance label for any of the training samples; Based on the posterior performance label of any training sample and the risk level conclusion in any inference output, determine the risk posterior reward of any inference output. The comprehensive reward for any given inference output is determined based on the inference content verification reward and risk posterior reward.
6. The method according to claim 1 or 2, characterized in that, The calculation of the model loss value based on the comprehensive reward of each inference output result of each training sample includes: For any training sample in the training sample set, the relative advantage value of the G inference outputs of any training sample is calculated based on the comprehensive reward of each inference output of the training sample. After obtaining the relative advantage values of the G inference outputs for each training sample, the model loss value is calculated based on the relative advantage values of the G inference outputs for each training sample.
7. The method according to claim 1 or 2, characterized in that, The determination of the audit application model based on the target logical audit model includes: Obtain the distillation analysis data set; The target logic audit model is invoked to generate multiple inference output results for each distillation analysis data in the distillation analysis data set. Traverse each distillation analysis data in the distillation analysis data set, and take the currently traversed distillation analysis data as the current analysis data, and determine the comprehensive reward of each inference output result among the multiple inference output results under the current analysis data; Based on the comprehensive reward of each inference output result under the current parsed data, a set of inference output result labels is determined from multiple inference output results under the current parsed data; and each inference output result label in the current parsed data and the set of inference output result labels is respectively composed of a distillation training data to obtain the current distillation training data set, and then the current distillation training data set is added to the target distillation training data set; After traversing through all the distillation analysis data in the distillation analysis data set, an audit application model is trained based on the target distillation training data set to determine the audit application model, which is a student model of the target logical audit model.
8. A risk auditing device, characterized in that, The device includes: The acquisition unit is used to acquire a training sample set, wherein a training sample includes at least one of the following: financial voucher data and interactive dialogue text of an audit object; The processing unit is used to call the text parsing model to parse each training sample in the training sample set to obtain the parsed data of each training sample. The processing unit is also used to call the logical audit model to generate G inference output results for each training sample based on the parsed data of each training sample, where G is an integer greater than 1; wherein, each inference output result includes a thought process and a risk level conclusion, and the thought process is used to indicate the reason for inferring a risk level conclusion through textual evidence. The processing unit is further configured to determine the comprehensive reward of each inference output result among the G inference output results of each training sample based on the G inference output results of each training sample; and calculate the model loss value based on the comprehensive reward of each inference output result of each training sample, thereby optimizing the model parameters in the logic auditing model in the direction of reducing the model loss value, until the target convergence condition is reached, so as to take the logic auditing model that reaches the target convergence condition as the target logic auditing model; The processing unit is further configured to determine an audit application model based on the target logical audit model, wherein the audit application model supports risk auditing.
9. An electronic device, characterized in that, include: processor; as well as Stored program memory, The program includes instructions that, when executed by the processor, cause the processor to perform the method according to any one of claims 1-7.
10. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-7.