Service message compliance risk identification method and device, equipment, medium and product

By combining large-scale model technology and a multimodal compliance knowledge base, the problems of low efficiency and high rate of missed detection in cross-border clearing compliance reviews have been solved, enabling efficient and accurate identification and management of compliance risks.

CN122069313APending Publication Date: 2026-05-19INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INDUSTRIAL AND COMMERCIAL BANK OF CHINA
Filing Date
2026-02-13
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Current technologies rely on manual configuration of rule bases for monitoring compliance risks in cross-border clearing, resulting in low efficiency and high rate of missed detections in compliance reviews, making it difficult to adapt to complex and ever-changing regulatory requirements.

Method used

By introducing large-scale model technology, a multimodal compliance knowledge base is constructed. Combined with a large language model for semantic parsing and risk identification, a closed-loop process of "preliminary identification by large-scale model - secondary manual verification - model fine-tuning and optimization" is formed, realizing automated collaboration between automatic rule generation and model optimization.

Benefits of technology

It has significantly improved the efficiency and accuracy of compliance review for cross-border clearing business, reduced the cost of manual intervention, and enhanced the ability to dynamically adapt to new compliance risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122069313A_ABST
    Figure CN122069313A_ABST
Patent Text Reader

Abstract

The invention provides a compliance risk identification method and device of a service message, equipment, a medium and a product, and relates to the field of data processing. The method comprises the following steps: obtaining a target service message and a pre-constructed compliance knowledge base; performing formatting analysis and field splitting processing on the target service message to obtain a structured message content item; according to a preset compliance risk check rule set, executing multiple compliance risk identification operations on the message content item; for each compliance risk identification operation, retrieving from a compliance knowledge base to obtain related knowledge fragments; and jointly inputting the message content item and the knowledge fragment into a large language model to obtain a compliance risk identification result corresponding to the target service message. According to the method provided by the invention, comprehensiveness, accuracy and real-time performance of compliance risk identification are realized, and compliance management efficiency of cross-border liquidation business is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing, and in particular to a method, apparatus, equipment, medium and product for identifying compliance risks in business messages. Background Technology

[0002] In cross-border clearing scenarios, banks need to process a large number of cross-border payment transaction messages. They need to conduct compliance reviews of the content of these transaction messages to ensure that the clearing process itself complies with all applicable laws, regulations, regulatory requirements, international standards and internal rules and regulations, and can effectively identify, assess, monitor and control the risks that may occur in this process.

[0003] In existing technologies, compliance risk monitoring in cross-border clearing primarily relies on manually configured rule bases. That is, business personnel manually extract compliance rules based on regulatory documents and internal policies, convert these rules into matching logic for message fields (e.g., regular expressions or keyword matching rules), and deploy them into the clearing system. When the system receives cross-border business messages, it uses a rule engine to compare each message field. If a violation rule is found, it is marked as a risk and pushed to the manual review process. Manual reviewers then need to consider contextual semantics, transaction background, and other information to make a secondary judgment and determine whether to take measures such as refunds or freezes.

[0004] However, manually configuring the rule base has problems such as delayed updates, insufficient semantic understanding, and incomplete rule coverage, resulting in low efficiency of compliance review, high rate of missed detection, and difficulty in adapting to complex and ever-changing regulatory requirements. Summary of the Invention

[0005] This application provides a method, apparatus, equipment, medium, and product for identifying compliance risks in business messages, in order to solve the technical problems of low efficiency and high rate of missed detection in compliance reviews.

[0006] Firstly, this application provides a method for identifying compliance risks in business messages, including:

[0007] Obtain the target business message and the pre-built compliance knowledge base;

[0008] The target business message is formatted, parsed, and its fields are split to obtain structured message content items;

[0009] Based on a pre-defined set of compliance risk checks, multiple compliance risk identification operations are performed on the message content items;

[0010] For each compliance risk identification operation, relevant knowledge fragments are retrieved from the compliance knowledge base;

[0011] The message content items and knowledge fragments are input into the large language model to obtain the compliance risk identification results corresponding to the target business message.

[0012] Secondly, this application provides a compliance risk identification device for business messages, comprising:

[0013] The acquisition module is used to acquire target business messages and a pre-built compliance knowledge base;

[0014] The processing module is used to format and parse the target business message and split its fields to obtain structured message content items;

[0015] The inspection module is used to perform multiple compliance risk identification operations on message content items according to a preset set of compliance risk inspection rules;

[0016] The retrieval module is used to retrieve relevant knowledge fragments from the compliance knowledge base for each compliance risk identification operation;

[0017] The identification module is used to input message content items and knowledge fragments into the large language model to obtain compliance risk identification results corresponding to the target business message.

[0018] Thirdly, this application provides an electronic device, including: a processor, and a memory communicatively connected to the processor;

[0019] The memory stores instructions that the computer executes;

[0020] The processor executes computer-executable instructions stored in memory to implement the method as described in the first aspect.

[0021] Fourthly, this application provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the method as described in the first aspect.

[0022] Fifthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the method as described in the first aspect.

[0023] The compliance risk identification method, apparatus, equipment, medium, and product for business messages provided in this application acquires the target business message and a pre-built compliance knowledge base. The target business message undergoes formatted parsing and field splitting to obtain structured message content items. Based on a preset set of compliance risk inspection rules, multiple compliance risk identification operations are performed on the message content items. For each compliance risk identification operation, relevant knowledge fragments are retrieved from the compliance knowledge base. The message content items and knowledge fragments are input into a large language model to obtain the compliance risk identification result corresponding to the target business message. The method in this application achieves comprehensiveness, accuracy, and real-time compliance risk identification, significantly improving the efficiency of compliance management in cross-border clearing business. Attached Figure Description

[0024] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0025] Figure 1 A schematic diagram of the structure of the business message compliance risk identification system provided in this application embodiment;

[0026] Figure 2 A flowchart illustrating the compliance risk identification method for business messages provided in this application embodiment;

[0027] Figure 3 A flowchart illustrating a compliance risk identification method for business messages provided in another embodiment of this application;

[0028] Figure 4 A flowchart illustrating a compliance risk identification method for business messages provided in yet another embodiment of this application;

[0029] Figure 5 A schematic diagram of the structure of the compliance risk identification device for business messages provided in the embodiments of this application;

[0030] Figure 6 A schematic diagram of the structure of the electronic device provided in this application.

[0031] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation

[0032] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0033] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of the relevant data all comply with the relevant laws, regulations, and standards of the relevant countries and regions, have taken necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation access points for users to choose to authorize or refuse.

[0034] In cross-border clearing scenarios, banks need to process a large number of cross-border payment transaction messages, which typically contain key fields such as payee and payee information, transaction amount, currency type, transaction path, and intermediary bank information.

[0035] In existing technologies, compliance risk monitoring in cross-border clearing primarily relies on manually configured rule bases. That is, business personnel manually extract compliance rules based on regulatory documents and internal policies, convert these rules into matching logic for message fields (e.g., regular expressions or keyword matching rules), and deploy them into the clearing system. When the system receives cross-border business messages, it uses a rule engine to compare each message field. If a violation rule is found, it is marked as a risk and pushed to the manual review process. Manual reviewers then need to consider contextual semantics, transaction background, and other information to make a secondary judgment and determine whether to take measures such as refunds or freezes.

[0036] However, manually configuring the rule base has problems such as delayed updates, insufficient semantic understanding, and incomplete rule coverage, resulting in low efficiency of compliance review, high rate of missed detection, and difficulty in adapting to complex and ever-changing regulatory requirements.

[0037] Based on this, this application combines the interpretation of compliance rules, risk identification, and rule iteration optimization by introducing large-scale model technology, forming a closed-loop process of "preliminary identification by large-scale model - secondary manual verification - fine-tuning and optimization of model". This significantly improves the efficiency and accuracy of compliance review for cross-border clearing business, reduces the cost of manual intervention, and enhances the dynamic adaptability to new compliance risks.

[0038] Specifically, this application constructs a closed-loop cross-border clearing compliance risk identification system based on large-scale model technology, encompassing "automatic rule generation - intelligent risk identification - manual feedback optimization - dynamic model iteration." By integrating cross-border clearing message knowledge, compliance requirement knowledge, and business Q&A pairs, a multimodal compliance knowledge base is built. The large-scale model is used for semantic analysis and risk identification of message content, and LoRA fine-tuning is performed in conjunction with manual review results, achieving automated collaboration between rule generation and model optimization. The core innovation of this concept lies in transforming the traditional static model of manually configured rules into a dynamic closed-loop system of "dynamic interpretation of the large-scale model - manual feedback-driven model iteration," thereby solving problems such as low rule generation efficiency, insufficient semantic understanding, and fragmented human-machine collaboration in existing technologies.

[0039] It should be noted that the methods, apparatus, equipment, media and products for identifying compliance risks of business messages provided in this application can be used in the field of data processing, or in any field other than data processing. The application fields of the methods, apparatus, equipment, media and products for identifying compliance risks of business messages in this application are not limited.

[0040] The specific application scenario of this application is applicable to cross-border payment and settlement scenarios of banks and financial institutions, and is specifically deployed in the business message compliance risk identification system. Figure 1 This is a schematic diagram of the structure of the business message compliance risk identification system provided in this application embodiment. Figure 1 As shown, the knowledge base module 4 is divided into three modules based on the type of input knowledge content: cross-border clearing message knowledge module 1, clearing business compliance requirements knowledge module 2, and clearing business knowledge Q&A module 3. The system also includes a data receiving and processing module 5, a risk identification module 6, a risk management module 7, a data collection and processing module 8, and a large model fine-tuning module 9.

[0041] Module 1, the knowledge module for cross-border clearing messages, stores the different types of cross-border clearing messages and their corresponding message formats as knowledge. Module 2, the knowledge module for clearing business compliance requirements, stores all applicable laws, regulations, regulatory requirements, international standards, and internal rules and regulations related to the clearing process. Module 3, the knowledge module for clearing business Q&A pairs, stores common business knowledge Q&A pairs in clearing business, such as terminology. These constitute the content of Module 4 of the knowledge base. This part is used during the risk identification phase to retrieve relevant knowledge as the basis for the overall model to assess the compliance risks of clearing messages.

[0042] Data receiving and processing module 5 receives clearing messages and formats and splits them according to message type and business type (receive / send, transfer, etc.), generating formatted clearing messages that clearly define the specific content of each message item. Risk identification module 6 identifies risks in clearing messages, including relevant knowledge retrieval and a large-scale model for risk identification. Risk handling module 7 allows business personnel to manually conduct a secondary review and handle compliance risks in clearing messages identified by the large-scale model in risk identification module 6.

[0043] Data collection and processing module 8 collects the assessment opinions and corresponding handling measures from business personnel regarding the liquidation compliance risks identified by the large model in risk identification module 6, and labels the risks as hits. Large model fine-tuning module 9 uses the liquidation message risk identification and handling details data from data collection and processing module 8 as the fine-tuning dataset for LoRA fine-tuning. After fine-tuning, this module synchronously receives and identifies the compliance risks of liquidation messages and the manual assessment results, calculates the accuracy rate, and evaluates the fine-tuning effect.

[0044] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will be described below with reference to the accompanying drawings.

[0045] Figure 2 This is a flowchart illustrating the compliance risk identification method for business messages provided in this application embodiment. Figure 2 As shown, the compliance risk identification method for business messages in this embodiment may include the following steps:

[0046] Step 2100: Obtain the target business message and the pre-built compliance knowledge base.

[0047] Step 2200: Perform format parsing and field splitting on the target business message to obtain structured message content items.

[0048] Step 2300: Perform multiple compliance risk identification operations on message content items according to the preset set of compliance risk inspection rules.

[0049] Step 2400: For each compliance risk identification operation, retrieve relevant knowledge fragments from the compliance knowledge base.

[0050] Step 2500: Input the message content items and knowledge fragments into the large language model to obtain the compliance risk identification results corresponding to the target business message.

[0051] The target business message can be a cross-border clearing message, such as a Cross-Border Interbank Payment System (CIPS) message or a customer credit transfer message (pacs.008), which can be obtained through API or file upload.

[0052] The steps involved in building a compliance knowledge base may include: acquiring business message knowledge, business compliance requirement knowledge, and business knowledge question-and-answer pairs. Based on this knowledge, the compliance knowledge base is then constructed.

[0053] It should be noted that in practical applications, a dynamic update mechanism can also be introduced into the compliance knowledge base. By monitoring changes in business compliance requirements in real time, and combining natural language processing (NLP) technology to perform semantic analysis on the new content, compliance rules can be automatically generated and synchronized to the compliance knowledge base.

[0054] Specifically, compliance requirement knowledge files can be subscribed to via API. When updates are detected, natural language processing models are used to extract keywords from the new content and map them to relevant rules. This ensures that the compliance knowledge base is always synchronized with the latest requirements, avoiding rule gaps caused by lagging manual interpretation. Furthermore, natural language processing-driven rule mapping reduces human intervention and improves the accuracy of rule generation.

[0055] After obtaining the target service message, the system formats and splits the message according to its type and service function to obtain a formatted service message containing all its components. For example, for the pacs.008 message, the system parses... <ntrydtls>The transaction details fields under the tags, such as payee and payee names, addresses, and country codes, are converted into structured data in JSON format.

[0056] Subsequently, based on a pre-defined set of compliance risk check rules, multiple compliance risk identification operations are performed on the message content items. The compliance risk check rule set is a set of configurable logical rules, such as screening rules, large transaction rules, sensitive word rules, and outlier rules. By executing multiple compliance risk identification operations on the message content items in parallel, several potential risk signals can be generated, thereby quickly filtering out obviously non-compliant transactions and improving identification efficiency.

[0057] When a rule triggers a risk signal, knowledge retrieval technology is used to retrieve relevant knowledge fragments from the compliance knowledge base. Then, large language modeling technology is used to combine these retrieved knowledge fragments with risk identification, comprehensively assessing the compliance risks inherent in the target business message. Using large language models enables faster and more comprehensive interpretation of regulatory requirements and rule generation.

[0058] It should be noted that the compliance risks identified in this embodiment mainly include, but are not limited to, cases where both the payer and payee are located overseas, agency transfer transactions, transfer transactions involving high-risk countries or regions, and the determination of the completeness of the six elements of the payee. The six elements of the payee and payee include: name or title, account number, country or region of residence, type of identification document, identification document number, and occupation or industry category.

[0059] The compliance risk identification method for business messages provided in this application obtains the target business message and a pre-built compliance knowledge base. The target business message undergoes formatted parsing and field splitting to obtain structured message content items. Based on a pre-set set of compliance risk inspection rules, multiple compliance risk identification operations are performed on the message content items. For each compliance risk identification operation, relevant knowledge fragments are retrieved from the compliance knowledge base. The message content items and knowledge fragments are then input into a large language model to obtain the compliance risk identification result corresponding to the target business message. This method achieves comprehensiveness, accuracy, and real-time compliance risk identification, significantly improving the efficiency of compliance management in cross-border clearing operations.

[0060] Figure 3 A flowchart illustrating a compliance risk identification method for business messages provided in another embodiment of this application. For example... Figure 3 As shown, based on the above embodiments, after step 2500, the method of this embodiment may further include the following steps:

[0061] Step 3100: In response to the risk assessment opinions input by business personnel, generate the final risk determination result.

[0062] Step 3200: Based on the final risk assessment result, perform corresponding risk handling operations on the target business messages that are assessed as having risks; the risk handling operations include at least one of the following: resending for review, redoing the business, or refunding the payment.

[0063] The risk assessment opinion is obtained in response to the manual review of the compliance risk identification results by business personnel. In this embodiment, business personnel conduct a secondary manual review of the compliance risk identification results identified by the large language model and input the risk assessment opinion. After receiving the risk assessment opinion, the system can generate a final risk judgment result. Based on the final risk judgment result, the system identifies the target business messages with risks and takes risk handling actions such as resubmission for review, business rework, or refund for these risky target business messages according to their risk level.

[0064] In practical applications, before generating the final risk assessment result, anomaly detection algorithms can be used to statistically analyze the misclassified cases of the large language model and determine whether the number of misclassified cases exceeds a preset threshold. When the number of misclassified cases exceeds the preset threshold, an update to the large language model is triggered. Specifically, the update to the large language model includes generating a fine-tuned dataset based on the risk assessment opinions and updating the parameters of the large language model using LoRA technology.

[0065] The technical solution in this embodiment effectively integrates the efficiency of artificial intelligence with the professionalism of human judgment by introducing business personnel to manually review the compliance risk identification results generated by the large language model and combining them with their input risk assessment opinions to generate the final risk judgment result. This significantly improves the accuracy and reliability of risk decision-making in highly sensitive business scenarios such as cross-border payments. At the same time, the model's misjudgment rate is continuously monitored through anomaly detection algorithms. When misjudged cases exceed a preset threshold, a lightweight fine-tuning mechanism based on LoRA is automatically triggered. High-quality fine-tuning datasets are constructed using human assessment opinions, realizing closed-loop optimization and adaptive evolution of the large language model. This not only reduces the risk of compliance omissions and false alarms but also enhances the system's intelligence level and long-term stability.

[0066] Figure 4 This is a flowchart illustrating a method for identifying compliance risks in business messages, as provided in another embodiment of this application. Figure 4 As shown, based on the above embodiment, after step 3200, the method of this embodiment may further include the following steps:

[0067] Step 4100: Filter the compliance risk identification results, final risk assessment results, and risk handling operations to obtain sample data that meets the preset retention conditions. The preset retention conditions include: data with accurate risk identification and business assessment as risk, or data with inaccurate risk identification and business assessment as risk.

[0068] Step 4200: Perform structured processing on the sample data.

[0069] Step 4300: Based on the structured sample data, construct a fine-tuning training set.

[0070] Step 4400: Update the large language model based on the fine-tuned training set.

[0071] To specifically enhance the model's risk identification capabilities in real-world business scenarios and achieve intelligent evolution that improves accuracy with use, this embodiment periodically constructs a fine-tuning training set for fine-tuning the large language model. Specifically, valuable sample data is selected from historical processing records, retaining two types of key cases: first, "positive examples" where the model correctly identifies risks and business personnel ultimately confirm them as risks; and second, "negative examples" where the model fails to accurately identify risks but are still judged as risks after manual review. These two types of data are crucial for improving the model's risk sensitivity and discrimination accuracy.

[0072] Next, the selected samples underwent structured processing, transforming heterogeneous information such as original messages, model outputs, human evaluation opinions, and handling results into standardized input-label pairs. A high-quality fine-tuning training set was constructed according to a preset ratio, which included: 20% positive risk identification examples, 20% negative risk identification examples, and 60% basic knowledge question-answering and conventional logical semantic question-answering pairs.

[0073] Based on the fine-tuned training set, the parameters of the large language model are adjusted using the low-rank adaptive LoRA method to obtain the updated large language model.

[0074] After fine-tuning the large language model, it enters the parallel operation phase, that is, the testing phase in which the updated large language model and the original large language model run synchronously.

[0075] Specifically, during the parallel execution phase, the large language model and the updated large language model run in parallel. At this time, the large language model and the updated large language model simultaneously perform risk identification on newly received target business messages and output the compliance risk identification results for the target business messages respectively.

[0076] The system separately calculates the accuracy of compliance risk identification results output by the large language model and the updated large language model. When the parallel operation period ends and the accuracy of the updated large language model is higher than that of the large language model, the updated large language model is used for subsequent risk identification.

[0077] It should be noted that during the parallel verification period, the compliance risk identification results actually sent to the business side are still generated based on the identification results of the original large language model.

[0078] The technical solution in this embodiment generates a fine-tuned dataset (such as misjudged cases) based on the results of manual review and updates the model parameters using LoRA technology. During the parallel validation period, the fine-tuned model runs synchronously with the original model, and the system selects the version with better performance by comparing accuracy. For example, if the fine-tuned model has a higher accuracy rate in identifying "high-risk country transfers" scenarios, the system will replace the original model.

[0079] The technical solution provided in this application constructs a structured fine-tuning training set for real business scenarios by screening high-value risk samples (including risk cases accurately identified by the model and cases that were missed but manually confirmed as risks). Based on this, the large language model is continuously optimized, which effectively improves the accuracy and generalization ability of the model in complex compliance scenarios such as cross-border payments. At the same time, this mechanism realizes the efficient transformation of human expert experience into model capabilities, forming a closed loop of "identification-review-learning-evolution". This not only significantly reduces the false positive and false negative rates, but also enhances the intelligence level and adaptability of the system in long-term operation.

[0080] Figure 5 This is a schematic diagram of the structure of the compliance risk identification device for business messages provided in the embodiments of this application. Figure 5 As shown, the compliance risk identification device 500 for business messages provided in this embodiment may include: an acquisition module 510, a processing module 520, an inspection module 530, a retrieval module 540, and an identification module 550.

[0081] Among them, the acquisition module 510 is used to acquire the target business message and the pre-built compliance knowledge base.

[0082] The processing module 520 is used to perform format parsing and field splitting on the target business message to obtain structured message content items.

[0083] The inspection module 530 is used to perform multiple compliance risk identification operations on message content items according to a preset set of compliance risk inspection rules.

[0084] The retrieval module 540 is used to retrieve relevant knowledge fragments from the compliance knowledge base for each compliance risk identification operation.

[0085] The identification module 550 is used to input the message content items and knowledge fragments into the large language model to obtain the compliance risk identification results corresponding to the target business message.

[0086] In one feasible implementation, the device may further include a generation module for generating a final risk assessment result in response to risk assessment opinions input by business personnel. An execution module is used to perform corresponding risk handling operations on the target business messages assessed as having risks, based on the final risk assessment result; the risk handling operations include at least one of remanding for review, redoing the business transaction, or refunding the payment.

[0087] In one feasible implementation, the device may further include a screening module for screening compliance risk identification results, final risk assessment results, and risk handling operations to obtain sample data that meets preset retention conditions. These preset retention conditions include: data where risk identification is accurate and the business assessment classifies it as risk, or data where risk identification is inaccurate and the business assessment classifies it as risk. The processing module 520 may also be used to perform structured processing on the sample data. A construction module is used to construct a fine-tuned training set based on the structured sample data. An update module is used to update the large language model based on the fine-tuned training set.

[0088] In one feasible implementation, the update module can be used to fine-tune the training set and adjust the parameters of the large language model using the low-rank adaptive LoRA method to obtain the updated large language model.

[0089] In one feasible implementation, the device may further include a running module for running the large language model and the updated large language model in parallel; an output module for synchronously performing risk identification on newly received target business messages during the parallel running period, and outputting the compliance risk identification results for the target business messages respectively; a statistics module for separately calculating the identification accuracy of the compliance risk identification results output by the large language model and the updated large language model; and a usage module for using the updated large language model for subsequent risk identification when the parallel running period ends and the identification accuracy of the updated large language model is higher than that of the large language model.

[0090] In one feasible implementation, the construction module can specifically be used to acquire business message knowledge, business compliance requirement knowledge, and business knowledge question-and-answer pairs. Based on the business message knowledge, business compliance requirement knowledge, and business knowledge question-and-answer pairs, a compliance knowledge base is constructed.

[0091] In one feasible implementation, the statistics module can also be used to count misjudged cases in the large language model using anomaly detection algorithms. When the number of misjudged cases exceeds a preset threshold, an update to the large language model is triggered.

[0092] In one feasible implementation, the update module can be used to generate a fine-tuned dataset based on risk assessment opinions; and to update the parameters of the large language model using LoRA technology.

[0093] The apparatus in this embodiment can be used to execute the technical solutions of the above method embodiments. Its implementation principle and technical effects are similar, and will not be described again here.

[0094] Figure 6 A schematic diagram of the structure of the electronic device provided in this application. Figure 6 As shown, the electronic device 60 provided in this embodiment includes at least one processor 601 and a memory 602. Optionally, the device 60 further includes a communication component 603. The processor 601, memory 602, and communication component 603 are connected via a bus 604.

[0095] In a specific implementation, at least one processor 601 executes computer execution instructions stored in memory 602, causing at least one processor 601 to perform the above-described method.

[0096] The specific implementation process of processor 601 can be found in the above method embodiments, and its implementation principle and technical effect are similar. It will not be repeated here.

[0097] In the above embodiments, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor.

[0098] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage device.

[0099] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.

[0100] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.

[0101] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the above-described method.

[0102] The aforementioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium can be any available medium accessible to a general-purpose or special-purpose computer.

[0103] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can reside in an Application Specific Integrated Circuit (ASIC). Alternatively, the processor and the readable storage medium can exist as discrete components in the device.

[0104] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to this application.

[0105] It should be further noted that although the steps in the flowchart are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowchart may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.

[0106] It should be understood that the above-described device embodiments are merely illustrative, and the device of this application can also be implemented in other ways. For example, the division of units / modules in the above embodiments is only a logical functional division, and there may be other division methods in actual implementation. For example, multiple units, modules, or components may be combined, or integrated into another system, or some features may be ignored or not executed.

[0107] Furthermore, unless otherwise specified, the functional units / modules in the various embodiments of this application can be integrated into one unit / module, or each unit / module can exist physically separately, or two or more units / modules can be integrated together. The integrated units / modules described above can be implemented in hardware or as software program modules.

[0108] When integrated units / modules are implemented in hardware, the hardware can be digital circuits, analog circuits, etc. The physical implementation of the hardware structure includes, but is not limited to, transistors, memristors, etc. Unless otherwise specified, the processor can be any suitable hardware processor, such as a CPU, GPU, FPGA, DSP, and ASIC, etc. Unless otherwise specified, the storage unit can be any suitable magnetic or magneto-optical storage medium, such as Resistive Random Access Memory (RRAM), Dynamic Random Access Memory (DRAM), Static Random Access Memory (SRAM), Enhanced Dynamic Random Access Memory (EDRAM), High-Bandwidth Memory (HBM), Hybrid Memory Cube (HMC), etc.

[0109] If the integrated unit / module is implemented as a software program module and sold or used as an independent product, it can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned memory includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.

[0110] In the above embodiments, the descriptions of each embodiment have their own emphasis. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments. The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combination of these technical features does not contradict each other, it should be considered within the scope of this specification.

[0111] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the following claims.

[0112] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.< / ntrydtls>

Claims

1. A method for identifying compliance risks in business messages, characterized in that, The method includes: Obtain the target business message and the pre-built compliance knowledge base; The target service message is formatted, parsed, and its fields are split to obtain structured message content items; Based on a preset set of compliance risk inspection rules, multiple compliance risk identification operations are performed on the message content items; For each of the aforementioned compliance risk identification operations, relevant knowledge fragments are retrieved from the compliance knowledge base; The message content items and the knowledge fragments are input into the large language model to obtain the compliance risk identification results corresponding to the target business message.

2. The method according to claim 1, characterized in that, After obtaining the compliance risk identification result corresponding to the target business message, the method further includes: The system generates a final risk assessment result based on the risk assessment opinions input by business personnel. Based on the final risk assessment result, corresponding risk handling operations are performed on the target business messages that are assessed as having risks; the risk handling operations include at least one of resubmission for review, business redoing, or refund.

3. The method according to claim 2, characterized in that, After performing corresponding risk handling operations on the target service messages assessed as having risks based on the final risk assessment result, the method further includes: The compliance risk identification results, the final risk assessment results, and the risk handling operations are screened to obtain sample data that meets preset retention conditions. The preset retention conditions include: data that is accurately identified as a risk and is assessed as a risk by the business, or data that is inaccurately identified as a risk and is assessed as a risk by the business. The sample data is then processed in a structured manner. Based on the structured sample data, a fine-tuning training set is constructed; The large language model is updated based on the fine-tuned training set.

4. The method according to claim 3, characterized in that, The update of the large language model based on the fine-tuned training set includes: Using the fine-tuned training set, the parameters of the large language model are adjusted using the low-rank adaptive LoRA method to obtain the updated large language model.

5. The method according to claim 3, characterized in that, After updating the large language model based on the fine-tuned training set, the method further includes: The large language model and the updated large language model are run in parallel. During parallel operation, the large language model and the updated large language model synchronously perform risk identification on newly received target service messages and output compliance risk identification results for the target service messages respectively. The accuracy rates of the compliance risk identification results output by the large language model and the updated large language model are calculated separately. When the parallel operation period ends and the recognition accuracy of the updated large language model is higher than the accuracy of the large language model, the updated large language model is used for subsequent risk identification.

6. The method according to claim 1, characterized in that, The steps for constructing the compliance knowledge base include: Acquire knowledge of business messages, business compliance requirements, and business knowledge Q&A pairs; The compliance knowledge base is constructed based on the business message knowledge, the business compliance requirement knowledge, and the business knowledge question-and-answer pairs.

7. The method according to claim 2, characterized in that, Before generating the final risk determination result in response to the risk assessment opinions input by business personnel, the method further includes: The misjudgment cases of the large language model are statistically analyzed using anomaly detection algorithms; When the number of misjudged cases exceeds a preset threshold, an update to the large language model is triggered.

8. The method according to claim 7, characterized in that, The update to the large language model includes: A fine-tuning dataset is generated based on the aforementioned risk assessment opinions; The parameters of the large language model are updated using LoRA technology.

9. A compliance risk identification device for business messages, characterized in that, The device includes: The acquisition module is used to acquire target business messages and a pre-built compliance knowledge base; The processing module is used to perform format parsing and field splitting on the target service message to obtain structured message content items; The inspection module is used to perform multiple compliance risk identification operations on the message content items according to a preset set of compliance risk inspection rules; The retrieval module is used to retrieve relevant knowledge fragments from the compliance knowledge base for each of the aforementioned compliance risk identification operations; The identification module is used to input the message content items and the knowledge fragments into the large language model to obtain the compliance risk identification result corresponding to the target business message.

10. An electronic device, characterized in that, include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory to implement the method as described in any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1 to 8.

12. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method of any one of claims 1 to 8.