Risk early warning method and device based on large model and secure multi-party computation

By employing large-scale models and secure multi-party computation methods, the conflict between data security and privacy protection in traditional financial risk early warning systems has been resolved, enabling efficient and secure cross-institutional risk identification and improving the accuracy and real-time nature of risk early warning.

CN122133148APending Publication Date: 2026-06-02INDUSTRIAL AND COMMERCIAL BANK OF CHINA

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INDUSTRIAL AND COMMERCIAL BANK OF CHINA
Filing Date
2025-08-14
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Traditional financial risk early warning systems rely on data from a single institution, resulting in insufficient risk capture across institutions and markets. Furthermore, there is a conflict between data security and privacy protection, and existing technologies struggle to achieve efficient and secure cross-institutional data sharing and model optimization.

Method used

The method employs a large model and secure multi-party computation approach. Multimodal data is collected with user authorization, and then anonymized and feature-generated. Secure multi-party computation is used to encrypt highly sensitive data, and differential privacy noise is combined to process moderately sensitive data. Finally, a large model is used for risk warning.

Benefits of technology

It enables efficient integration of data from multiple institutions while ensuring data privacy, improving the accuracy and real-time nature of risk warnings, meeting the risk identification needs across institutions and markets, and reducing computing resource requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122133148A_ABST
    Figure CN122133148A_ABST
Patent Text Reader

Abstract

This application provides a risk warning method and apparatus based on a large model and secure multi-party computation, which can be applied to the fields of artificial intelligence, big data, and privacy computing. The method includes: collecting multimodal data with user authorization or consent, wherein the multimodal data comes from multiple data providers; performing de-identification and feature generation processing on the multimodal data to obtain multimodal feature data, wherein the de-identification and feature generation processing at least includes using secure multi-party computation to perform de-identification and feature generation processing on at least some data in the multimodal data, and the multimodal feature data includes at least two of structured data, text data, time-series data, and graph-structured data; fusing the multimodal feature data to obtain fused feature data with semantics; and processing the fused feature data using a large model to obtain a risk warning result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of artificial intelligence, big data and privacy computing technologies, specifically to a risk warning method, device, electronic device, storage medium and program product based on large models and secure multi-party computation. Background Technology

[0002] Traditional financial risk early warning systems typically rely on data from a single institution to build models. From a technical perspective, this approach has significant drawbacks. Data from a single institution can only reflect a localized business scenario, resulting in insufficient coverage of features and difficulty in capturing potential risks across institutions and markets, thus affecting the accuracy of risk warnings. Furthermore, the data distribution of a single institution is limited by its customer base and business scope, making the model prone to overfitting to localized data. This leads to a significant drop in generalization performance when migrating to different scenarios. Improving model adaptability often requires investing more computing resources in model adjustments and optimizations, yet achieving the desired results is often difficult.

[0003] Multi-institutional data sharing can compensate for the shortcomings of single-institutional data, but it faces numerous technical obstacles. First, there is a conflict between data security and privacy protection. Directly sharing raw data can lead to the leakage of client privacy, thus violating data security requirements. Existing encryption technologies are inefficient when processing high-dimensional unstructured data, consuming significant computational resources and causing severe feature loss, affecting data usability. Second, different institutions have significantly different data formats and feature definitions. Traditional data cleaning and alignment methods rely on manual rules, making it difficult to achieve automated cross-modal semantic mapping. This not only increases the time cost of data processing and reduces computational efficiency but also affects model accuracy due to poor feature consistency. Third, multi-institutional joint training requires collaborative updating of model parameters without data being stored locally. Existing federated learning technologies, when processing large models, involve large parameter transmission volumes, consuming excessive communication resources and incurring high communication costs. They also pose a risk of gradient leakage, failing to guarantee data security and struggling to support the joint optimization of complex risk warning models, resulting in low model training efficiency.

[0004] In recent years, the development of large-scale models and secure multi-party computation (MPC) technologies has offered new possibilities for solving the aforementioned problems, but there are still technological gaps in their integrated application. Large-scale models have shown great potential in natural language understanding and cross-modal data fusion, but their fine-tuning and inference rely on large-scale labeled data. In cross-institutional scenarios, due to data privacy restrictions, models cannot directly utilize multi-source data for optimization, hindering their advantages. Furthermore, large-scale models themselves are computationally intensive, demanding extremely high computer resources; failure to efficiently utilize multi-source data leads to significant resource waste. While secure MPC technologies can achieve joint computation in an encrypted state, the encryption process results in the loss of some feature information. Current technologies lack alignment mechanisms between encrypted data and plaintext semantics, reducing the risk feature representation capability of joint models and affecting the accuracy of early warnings. Meanwhile, cross-agency risks have dynamic evolution characteristics, requiring early warning systems to have real-time data processing and model update capabilities. However, the joint computing latency under the existing secure multi-party computing framework is relatively high, especially in multi-round interaction scenarios. The inference speed of large models is also difficult to meet the needs of high-frequency risk monitoring. The real-time optimization of the two lacks mature technical solutions, resulting in delayed risk warnings and an inability to respond to potential risks in a timely manner. Summary of the Invention

[0005] In view of at least one aspect of the above problems, embodiments of this application provide a risk warning method, apparatus, electronic device, storage medium, and program product based on large models and secure multi-party computation.

[0006] According to a first aspect of this application, a risk warning method based on a large model and secure multi-party computation is provided. The method includes: collecting multimodal data with user authorization or consent, wherein the multimodal data comes from multiple data providers; performing desensitization and feature generation processing on the multimodal data to obtain multimodal feature data, wherein the desensitization and feature generation processing includes at least using secure multi-party computation to perform desensitization and feature generation processing on at least some data in the multimodal data, and the multimodal feature data includes at least two of structured data, text data, time-series data, and graph structured data; fusing the multimodal feature data to obtain fused feature data with semantics of the fused multimodal feature data; and processing the fused feature data using a large model to obtain a risk warning result.

[0007] A second aspect of this application provides a risk warning device based on a large model and secure multi-party computation. The device includes: a multimodal data acquisition module for acquiring multimodal data upon obtaining user authorization or consent, wherein the multimodal data comes from multiple data providers; a multimodal feature data acquisition module for performing desensitization and feature generation processing on the multimodal data to obtain multimodal feature data, wherein the desensitization and feature generation processing includes at least using secure multi-party computation to perform desensitization and feature generation processing on at least some data in the multimodal data, and the multimodal feature data includes at least two of structured data, text data, time-series data, and graph-structured data; a fusion feature data acquisition module for fusing the multimodal feature data to obtain fusion feature data with semantic fusion of the multimodal feature data; and a risk warning result acquisition module for processing the fusion feature data using a large model to obtain a risk warning result.

[0008] A third aspect of this application provides an electronic device comprising: one or more processors; and a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the method described above.

[0009] A fourth aspect of this application also provides a computer-readable storage medium having a computer program or instructions stored thereon, which, when executed by a processor, implement the steps of the above-described method.

[0010] The fifth aspect of this application also provides a computer program product, including a computer program or instructions that, when executed by a processor, implement the steps of the above-described method. Attached Figure Description

[0011] The above-mentioned contents, other objects, features and advantages of this application will become clearer from the following description of embodiments with reference to the accompanying drawings, in which:

[0012] Figure 1 This illustration schematically depicts an application scenario of a risk warning method based on a large model and secure multi-party computation according to an embodiment of this application.

[0013] Figure 2 This is a flowchart of a risk warning method based on a large model and secure multi-party computation according to an embodiment of this application;

[0014] Figure 3 This is a flowchart illustrating the method of using secure multi-party computation to de-identify and generate features for highly sensitive data according to an embodiment of this application;

[0015] Figure 4This is a flowchart illustrating the method of desensitizing and generating features for sensitive data using differential privacy noise matching in an embodiment of this application;

[0016] Figure 5 This is a detailed flowchart of the method for fusing multimodal feature data according to an embodiment of this application;

[0017] Figure 6 This is a detailed flowchart of the method for embedding multimodal data into a unified semantic space according to an embodiment of this application;

[0018] Figure 7 This is a detailed flowchart of the method for generating risk warning results according to an embodiment of this application;

[0019] Figure 8 This schematically illustrates a structural block diagram of a risk warning device based on a large model and secure multi-party computation according to an embodiment of this application; and

[0020] Figure 9 The diagram illustrates an electronic device suitable for implementing a risk warning method based on a large model and secure multi-party computation, according to an embodiment of this application. Detailed Implementation

[0021] The embodiments of this application will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of this application. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of this application for ease of explanation. However, it will be apparent that one or more embodiments may be implemented without these specific details. Furthermore, descriptions of well-known structures and technologies are omitted in the following description to avoid unnecessarily obscuring the concepts of this application.

[0022] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of this application. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0023] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.

[0024] In the technical solution of this application, the user information (including but not limited to user personal information, user image information, user device information, such as location information) and data (including but not limited to data used for analysis, stored data, and displayed data) involved are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of related data all comply with relevant laws, regulations, and standards, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entry points for users to choose to authorize or refuse.

[0025] In scenarios involving automated decision-making using personal information, the methods, devices, and systems provided in this application all offer users corresponding entry points for choosing to agree to or reject the automated decision-making results. If the user chooses to reject, the process proceeds to the expert decision-making stage. Here, "automated decision-making" refers to the activity of automatically analyzing and evaluating an individual's behavioral habits, interests, or economic, health, and credit status through computer programs, and then making a decision. Here, "expert decision-making" refers to the activity of making decisions by personnel who specialize in a particular field, possess specialized experience, knowledge, and skills, and have reached a certain level of professional expertise.

[0026] Some exemplary embodiments of this application provide a risk warning method and apparatus based on a large model and secure multi-party computation. The method and apparatus can be applied to enterprise-level risk control modeling in a multimodal data environment, achieving high-precision and highly scalable cross-institutional risk identification tasks while ensuring privacy. The method includes: collecting multimodal data with user authorization or consent, wherein the multimodal data comes from multiple data providers; performing desensitization and feature generation processing on the multimodal data to obtain multimodal feature data, wherein the desensitization and feature generation processing at least includes using secure multi-party computation to perform desensitization and feature generation processing on at least some data in the multimodal data, and the multimodal feature data includes at least two of structured data, text data, time-series data, and graph-structured data; fusing the multimodal feature data to obtain fused feature data with semantically fused multimodal feature data; and processing the fused feature data using a large model to obtain a risk warning result.

[0027] In the embodiments of this application, the method can achieve the fusion of multimodal data and achieve an adaptive balance between privacy protection and model accuracy through a dynamic privacy control mechanism. Specifically, in terms of data security, the method performs desensitization and feature generation processing on multimodal data through secure multi-party computation, enabling multiple data providers to complete collaborative processing without exposing the original data. This technically blocks the risk of original data leakage. Furthermore, combined with the premise of user authorization or consent, it not only meets the requirements of data security regulations but also builds a reliable privacy protection barrier in cross-institutional data collaboration, effectively solving the problem of balancing privacy and security in traditional data sharing. In terms of computational real-time performance and efficiency, the method performs desensitization and feature generation simultaneously, avoiding the redundant steps of desensitization followed by separate feature generation in traditional processes. This reduces intermediate steps in data processing, thereby improving the overall process efficiency. The direct processing of fused feature data by the large model achieves end-to-end risk warning, eliminating the feature transfer and transformation losses in traditional multi-model step-by-step processing. This allows for faster output of warning results, meeting the needs of real-time risk monitoring, and is particularly suitable for financial risk warning scenarios with high timeliness requirements.

[0028] Figure 1 The diagram illustrates an application scenario of a risk warning method based on a large model and secure multi-party computation according to an embodiment of this application.

[0029] like Figure 1 As shown, application scenario 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, and a server 105. Network 104 serves as a medium for providing a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. Network 104 may include various connection types, such as wired or wireless communication links or fiber optic cables, etc.

[0030] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 via the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social media platform software, etc. (for example only).

[0031] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with displays and support web browsing, including but not limited to smartphones, tablets, laptops, and desktop computers.

[0032] Server 105 can be a server that provides various services, such as a backend management server that supports websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (this is just an example). The backend management server can analyze and process data such as received user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.

[0033] It should be noted that the risk warning method based on large models and secure multi-party computation provided in this application embodiment can generally be executed by server 105. Correspondingly, the risk warning device based on large models and secure multi-party computation provided in this application embodiment can generally be located in server 105. The risk warning method based on large models and secure multi-party computation provided in this application embodiment can also be executed by a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105. Correspondingly, the risk warning device based on large models and secure multi-party computation provided in this application embodiment can also be located in a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105.

[0034] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0035] The following will be based on Figure 1 The described scene, through Figures 2-7 A risk warning method based on large models and secure multi-party computation according to embodiments of this application is described in detail.

[0036] Figure 2 This is a flowchart of a risk warning method based on a large model and secure multi-party computation according to an embodiment of this application, with reference to... Figure 2 The risk warning method based on large model and secure multi-party computation according to the embodiments of this application may include the following steps S210~S240.

[0037] In step S210, with the user's authorization or consent obtained, multimodal data is collected, wherein the multimodal data comes from multiple data providers.

[0038] For example, the term "multimodal data" can include various types and formats of data sets used for risk warning. For instance, it can include, but is not limited to, the following modalities of data: structured data, such as credit records from financial institutions (including numerical fields such as loan amount, repayment period, and number of overdue payments), and balance sheets in corporate financial statements (including quantitative indicators such as total assets and total liabilities); unstructured data, such as corporate annual reports or financial statements, risk warning paragraphs in corporate annual reports, industry regulatory updates published by news media, and unstructured text information such as public opinion comments about a company on social media; and semi-structured data, such as data in JSON (JavaScript) format. Loan contracts written using the ObjectNotation (ONFC) technical specification are characterized by the structured storage of contractual content such as loan terms, borrower and lender information, and repayment plans in the form of key-value pairs, nested objects, or arrays. Time-series data, such as real-time price fluctuation sequences, time series of fund inflows and outflows for a specific account, and exchange rate curves over continuous periods, reflects data that changes over time. Graph-structured data, such as investment relationship networks between enterprises (nodes can be enterprises, edges represent investment connections), cooperative relationship graphs of upstream and downstream supply chains, and credit relationship graphs between financial institutions and customers, reflects data that demonstrates entity relationships. It should be understood that these different modalities of data can reflect risk-related information from multiple dimensions, collectively forming a multi-dimensional basis for risk assessment and early warning.

[0039] For example, the phrase "multiple data providers" indicates the cross-domain nature of the data sources, which can encompass various institutions or entities holding relevant risk information. These can include, but are not limited to, commercial banks, securities companies, credit reporting agencies, news media platforms, and core enterprises in the supply chain. For instance, commercial banks can provide corporate credit data and account transaction records; securities companies can provide stock trading data and listed company financial reports; credit reporting agencies can provide structured data such as credit scores for enterprises or individuals; news media platforms can provide textual data such as industry trends; and core enterprises in the supply chain can provide graph-structured data such as upstream and downstream cooperative relationships. By collecting data from these different entities, the limitations of a single institution's data coverage can be overcome, providing a data foundation for capturing potential risks across institutions and markets.

[0040] Continue to refer to Figure 2 In step S220, the multimodal data is subjected to desensitization and feature generation processing to obtain multimodal feature data. The desensitization and feature generation processing includes at least using secure multi-party computation to perform desensitization and feature generation processing on at least some data in the multimodal data. The multimodal feature data includes at least two of structured data, text data, time-series data, and graph structure data.

[0041] In some embodiments of this application, the multimodal data may include a first type of data, a second type of data, and a third type of data. The first type of data, the second type of data, and the third type of data may be divided according to their sensitivity. For example, the first type of data may be more sensitive than the second type of data, and the second type of data may be more sensitive than the third type of data. That is, the first type of data may be highly sensitive data, the second type of data may be moderately sensitive data, and the third type of data may be low-sensitivity data.

[0042] For example, the first category of data (i.e., highly sensitive data) can include data directly related to the core privacy of individuals or enterprises, which may pose significant security risks if leaked. Examples include, at the individual level, ID numbers, complete bank account information, biometric data, salary slips, and detailed repayment records in credit history; and at the enterprise level, undisclosed core financial data (such as internal fund flows, undisclosed major debt contracts), key customer lists, and transaction details. The second category of data (i.e., moderately sensitive data) can include information that, while not directly revealing identity, can be indirectly linked to the entity when combined with other information, or information that is business-sensitive but not core. Examples include, at the individual level, anonymized transaction amount ranges, credit scores, occupation types, and income ranges; and at the enterprise level, accounts payable, industry classification codes, summary indicators in publicly available financial reports, and types of cooperative relationships with upstream and downstream enterprises. The third category of data (i.e., low-sensitivity data) can include information that does not involve the privacy of individuals or enterprises, is publicly available, or has been processed to be non-identifiable. For example, at the individual level, there are anonymous group statistics and consumer trend tags from public channels; at the enterprise level, there are publicly available industry reports, interpretations of industry policies published by news media, and publicly available supply chain topology maps.

[0043] In step S220, different desensitization and feature generation strategies can be adopted for different types of data. In some embodiments, secure multi-party computation can be used to desensitize and generate features for the first type of data, differential privacy noise can be used to desensitize and generate features for the second type of data, and static desensitization or plaintext processing can be used to desensitize and generate features for the third type of data. It should be understood that static desensitization or plaintext processing can be described as ordinary desensitization. That is, secure multi-party computation, differential privacy noise, and ordinary desensitization are used respectively to desensitize and generate features for highly sensitive data, moderately sensitive data, and low-sensitivity data. In this way, data privacy protection can be effectively balanced with processing efficiency. By matching highly sensitive data with more secure processing methods, the risk of leakage of core privacy information can be minimized. For moderately sensitive data, differential privacy noise addition is selected to balance privacy and efficiency, which can meet basic privacy requirements and avoid efficiency loss caused by over-processing. For low-sensitivity data, ordinary de-identification or plaintext processing is used to improve processing speed while ensuring compliance. Ultimately, a precise balance between the strength of privacy protection and processing efficiency is achieved for data of different sensitivity levels.

[0044] In some embodiments, secure multi-party computation (MPC) can be used to de-identify and generate features for highly sensitive data. For example, MPC can be applied to highly sensitive structured data to extract risk feature vectors usable under encryption, thus protecting the data privacy of each participant. MPC allows multiple participants to collaboratively complete computational tasks through encryption protocols without mutual trust, and the original data of each party remains encrypted or segmented, preventing access by other participants. It can ensure that data is "usable but not visible" during computation through protocols such as secret sharing, homomorphic encryption, or obfuscated circuits. Mathematical proofs demonstrate that even with malicious parties, it is impossible to deduce the data of other participants in secure MPC. By processing highly sensitive data through MPC, participants can collaboratively complete computations without sharing original data, ensuring that data remains encrypted throughout the process. For example, when jointly calculating corporate financial risk indicators, the core financial data of each party remains in encrypted form, and only authorized intermediate results can be released, mathematically preventing data leakage.

[0045] Figure 3 This is a flowchart illustrating the method of using secure multi-party computation to de-identify and generate features for highly sensitive data according to embodiments of this application, with reference to... Figure 3 The method according to the embodiments of this application, which employs secure multi-party computation to desensitize and generate features for highly sensitive data, may include the following steps S310 to S330.

[0046] In step S310, the first-type data from multiple data providers is encoded into multiple first vectors. For example, first-type data from different data providers such as commercial banks and credit reporting agencies (e.g., highly sensitive data such as corporate account statements and personal salary details) can be filtered through key fields to identify core fields directly related to risk assessment, such as "average daily account balance" reflecting liquidity, "number of high-frequency large-amount transfers" reflecting transaction anomalies, and "default records in the past three years" related to credit history. Subsequently, different types of fields are specifically encoded. For example, continuous numerical fields (e.g., balance amount) can be normalized and mapped to the [0,1] interval, and categorical fields (e.g., default status) can be converted into binary vectors using one-hot encoding. Then, the encoded fields can be integrated into a first vector of uniform length (e.g., 512 dimensions) to ensure that data from different sources can achieve format compatibility in subsequent encryption calculations.

[0047] In step S320, multiple first vectors are encrypted using an encryption protocol to generate multiple first encrypted encoded vectors. In this step S320, the first vectors are securely protected using an encryption protocol, generating encrypted encoded vectors that can be used for multi-party collaborative computation. For example, mainstream secure multi-party computation protocols can be used. Homomorphic encryption is first performed on the first vectors of each data provider, converting the original vectors into ciphertext. For instance, an encryption scheme can be used to process numerical vectors, ensuring that basic operations such as addition and multiplication can still be performed directly in the ciphertext state. Based on this, a secret sharing mechanism is used to split each ciphertext vector into multiple "shares." Different data providers only hold a portion of these shares, and a single share cannot be used to deduce the original data content. Taking joint processing by two institutions as an example, institution A's vector, after encryption, will be split into two shares, A1 and A2. Institution A retains A1 and transmits A2 to institution B. Simultaneously, institution B transmits its vector share B1 to institution A and retains B2 itself, ensuring that no single party can obtain the complete encrypted information, thus blocking the possibility of original data leakage at the underlying protocol level.

[0048] In step S330, based on pre-built risk features, secure multi-party computation is performed on multiple first encrypted encoding vectors from multiple data providers to generate first feature data. That is, in step S330, secure multi-party computation can be performed on first encrypted encoding vectors from multiple sources based on a pre-built risk feature library to generate usable first feature data. For example, the pre-built risk feature library may include basic statistical features, such as "flow volatility" calculated using the mean and variance in encrypted state, and "proportion of high-frequency large-amount transfers" obtained through share-based collaborative computation. It may also include more complex time-series and correlation features, such as "credit default cycle distribution" generated based on time-series analysis in encrypted state, and "cross-institutional fund correlation" calculated using the cosine similarity of encrypted vectors from multiple institutions. Throughout the computation process, the original data is always distributed among the participants in the form of encrypted shares. The aggregated feature result is only output when the preset computation rules are met. For example, when multiple banks jointly calculate the comprehensive default probability of a company, each party only needs to contribute its local encrypted share to participate in the computation, ultimately obtaining only an integrated probability value without leaking any single institution's original data.

[0049] For example, parallel secure multiplication pre-computation can be used to generate and store a large number of triples satisfying c=a×b during system idle periods. During actual computation, the pre-computation results are directly called, significantly reducing the communication rounds of multiplication operations from the traditional O(n) to O(1). Based on a multi-party joint optimization graph scheduling algorithm, dependency analysis is performed on complex feature computation processes, merging redundant intermediate computation paths, transforming matrix operations that originally required serial execution into parallel processing task flows. Simultaneously, with the help of a secure multi-party computation platform, encrypted matrix operations are deployed on GPU (Graphics Processing Unit) clusters or customized FPGA (Field-Programmable Gate Array) chips, further improving processing efficiency through hardware acceleration. These technical means can improve or solve the bottleneck problem of encrypted computation efficiency, effectively supporting the real-time requirements of risk warning systems.

[0050] Figure 4 This is a flowchart illustrating the method of desensitizing and generating features for sensitive data using differential privacy noise reduction according to embodiments of this application, see below. Figure 4 The method according to the embodiments of this application, which uses differential privacy noise addition to desensitize and generate features for sensitive data, may include the following steps S410-S420.

[0051] In step S410, based on a preset differential privacy perturbation mechanism, noise is added to the second type of data from multiple data providers to generate multiple second encryption encoding vectors. The preset differential privacy perturbation mechanism includes adding Laplace noise to numerical data and adding hot encoding perturbation to categorical data.

[0052] In step S420, based on the pre-constructed risk features, a secure multi-party computation is performed on multiple second cryptographic encoding vectors from multiple data providers to generate second feature data.

[0053] In some embodiments, differential privacy noise can be used to de-identify and generate features on moderately sensitive data. Differential privacy noise can include adding controlled noise (e.g., Laplace noise or Gaussian noise) to moderately sensitive data, making it impossible for attackers to infer sensitive information from statistical results. For example, mathematical perturbations can be used to ensure that the addition or removal of a single record has a negligible impact on the overall result. Applying differential privacy noise to moderately sensitive data adds designed noise to the data, preventing attackers from inferring the specific value of a single sample from the model output, thus preserving the statistical properties of the data while avoiding the leakage of precise numerical values.

[0054] See Table 1 for examples of differential privacy noise-adding strategies used in some embodiments of this application.

[0055] Table 1 Examples of Differential Privacy Noise Addition Strategies

[0056]

[0057] Differential privacy noise-adding strategies can be designed with different processing methods for different types of data fields to preserve data usability while ensuring privacy. For example, for numerical fields, a sub-strategy of adding Laplace noise can be used. Specifically, random noise conforming to a Laplace distribution is added to the original value to make the processing result meet the ε-DP (i.e., ε-differential privacy) requirements. For fields such as loan overdue days and tax amounts, this method ensures that individual information cannot be inferred from group data. For categorical fields, a sub-strategy of hot-encoding perturbation can be used. After hot-encoding the field, data that may be used to infer individual identity is scrambled or discarded. Taking credit rating as an example, this operation makes low-frequency categories difficult to identify, avoiding privacy leaks. For scenarios that need to balance accuracy and privacy, a sub-strategy of interval mapping fuzzification can be used to convert specific precise values ​​into interval values. For example, "12 days overdue" can be converted into "10-15 days". After processing, such fields can retain a certain degree of usability to support statistical analysis while preventing the inference of individual transaction details from precise values.

[0058] Optionally, based on the aforementioned differential privacy-preserving noise enhancement, various optional enhancement methods can be used to further improve the processing effect. For example, multiple noisy versions of the same data under different privacy budget values ​​ε (e.g., 0.1, 1.0, 10.0) can be generated through multi-scale perturbation generation combined with contrastive learning. These versions can then be used as "different views of the same semantics" input into the model for contrastive learning training. In this way, the model can be optimized to gradually learn to "ignore" the noise generated by perturbation during the learning process, thereby outputting a representation that is robust to semantic information and can accurately capture the core meaning of the data even in the face of noise interference of different intensities. For example, pseudo-sample amplification mechanisms can be used to construct pseudo-samples based on the original field values ​​by floating them up or down by a certain proportion or adding Gaussian perturbations. These pseudo-samples can participate in model training in conjunction with the actual labels of the data, effectively alleviating the data sparsity problem caused by insufficient sample size of sensitive data, while helping large models to more comprehensively understand the data distribution characteristics of sensitive fields, improving the model's adaptability and generalization performance for this type of data. For example, an adversarial semantic consistency mechanism can be used to feed the original field and the noisy field together into a shared large model encoder, while a discriminator is introduced to determine whether the input data has been noisy. During training, the encoder aims to "deceive" the discriminator by optimizing its own parameters, making it unable to accurately distinguish between the original data and the noisy data. This ensures that the semantic feature vector obtained after encoding the noisy field is as close as possible to the semantic representation of the original field, thus maximizing the preservation of semantic consistency of the data while ensuring privacy.

[0059] In the embodiments of this application, data perturbation processing based on differential privacy mechanisms is performed on sensitive data to ensure that this data participates in secure multi-party computation and large-scale model semantic parsing while meeting privacy protection compliance requirements. Furthermore, multiple feature enhancement mechanisms can be integrated to mitigate information distortion caused by noise and improve the modeling and expressive capabilities of subsequent models for this type of feature.

[0060] In some embodiments, common desensitization techniques can be used to desensitize and generate features for low-sensitivity data. Common desensitization can include directly modifying or deleting sensitive information using static techniques (such as replacement, masking, generalization, etc.) to prevent the data from being directly associated with individuals. For example, sensitive fields can be replaced or masked based on rules, such as replacing ID numbers with random strings or partial masks, or replacing company names in publicly available industry reports with anonymous ones. This protects company identities while allowing models to directly analyze industry trends. In the embodiments of this application, the low-sensitivity data can be statically desensitized or processed in plaintext, which can improve processing efficiency while ensuring that the data does not contain sensitive information.

[0061] For example, large models can be used to embed, extract semantics, and mine entity relationships from external data such as financial statements and news sentiment to extract relevant features. For instance, the input text "Accounts receivable increased by 200% compared to the same period last year" corresponds to the feature "Receivable Risk: High"; the input text "The company has undisclosed related-party transactions with third parties" corresponds to the feature "Fraud Signal: Strong"; and the input text "Chairman's personal funds made large advances" corresponds to the feature "Abnormal Governance Structure: Medium".

[0062] In step S220, after the multimodal data is desensitized and feature generated using methods such as secure multi-party computation, differential privacy noise addition, and ordinary desensitization processing, multimodal feature data can be obtained. The multimodal feature data may include at least two of structured data, text data, time-series data, and graph structure data.

[0063] Table 2 provides examples of structured data included in the multimodal feature data. Referring to Table 2, structured data can include numerical and Boolean feature data. For example, it can originate from multiple channels such as financial statements, annual reports, profit statement notes, and public opinion data, providing direct quantitative or qualitative judgment criteria for risk warning. For instance, an "abnormal growth rate of accounts receivable" of 0.23 can signal financial strain for the company; a "decline of 12.5% ​​in the proportion of main business revenue" can reflect a potential decline in core business; and a "non-operating profit and loss ratio of 37.8%" can indicate low profitability. Boolean feature data, such as "financial fraud risk indicator" being True, can indicate that the model has detected implicit risk expressions from the semantics of financial statements; "involvement in undisclosed related-party transactions" being True can indicate weaknesses in the company's internal controls; and a "negative online public opinion intensity index" of 0.86, highly correlated with the probability of default, can be directly used for risk probability calculations, etc.

[0064] Table 2. Examples of Structured Data

[0065]

[0066] Table 3 provides examples of the text data included in the multimodal feature data. Referring to Table 3, the text data can include multi-dimensional semantic information extracted from various types of text data to support in-depth risk modeling and semantic relationship analysis. For example, "financial statement semantic classification tags," such as "business contraction" and "implicit guarantee," can be obtained from financial statement text. Multi-tag classification can more meticulously characterize corporate risks. "Public opinion event themes," such as "litigation," extracted from news and social media data, can serve as node annotations for event graphs. "Public opinion time feature tags," such as "frequent negative news," can provide a time dimension reference for time-series modeling. "Main body behavioral verb phrase sets," such as "pledge," "guarantee," and "backdoor listing," extracted from financial statements or news, can form an important foundation for constructing semantic relationship graphs. The generation of this text data depends on the processing of texts such as financial statements, public opinion news, and court judgments. Through large-scale model extraction of events, entities, and emotions, combined with command-based dialogue prompts to extract potential risk statements, a structured semantic identifier is ultimately formed.

[0067] Table 3. Examples of Text Data

[0068]

[0069] Table 4 provides examples of time-series data included in the multimodal feature data. Referring to Table 4, time-series data can focus on the dynamic changes of data over time, providing a trend reference in the time dimension for risk warning. For example, when the "Public Opinion Change Slope (7 days)" is +0.12, it indicates that risk sentiment is accelerating in the past 7 days; when the "Financial Indicator Variation Time Window" is 2, it indicates that the company has experienced a surge in financial indicators within two quarters, suggesting abnormal fluctuations. Time-series data can be extracted from time-dimensional data such as public opinion time series and comparisons of financial indicators across multiple periods, effectively capturing the evolution of risk over time and enhancing the model's sensitivity to dynamic changes in risk and its predictive accuracy.

[0070] Table 4. Examples of Time Series Data

[0071]

[0072] Table 5 provides examples of graph-structured data included in the multimodal feature data. Referring to Table 5, graph-structured data can be used to construct relationships between enterprises and various entities and events. For example, "Enterprise and Parent Company Equity Edge," such as "A→B (40% shareholding)," reflects the equity control relationship between enterprises; "Enterprise and Negative Event Node," such as "A→'Labor Dispute'→Public Opinion Heat = 0.9," establishes the association and impact degree between the enterprise and specific negative events, providing a basis for graph neural networks to extract subgraph features; "Enterprise and Key Person Edge," such as "A→'Zhang San' (involved in violations)," helps assess the path of risk transmission from key figures to the enterprise. Graph-structured data, constructed through entity relationships along the company-behavior-funding path within the text and combined with access to publicly available external equity data, can further enrich graph information such as guarantee chains and shareholding chains, enhancing the ability to capture complex associated risks.

[0073] Table 5. Examples of graph structure data

[0074]

[0075] In the embodiments of this application, by implementing differentiated desensitization and feature generation strategies for data with different sensitivity levels, the overall effectiveness of risk warning methods can be improved while ensuring data security. Specifically, it can improve the accuracy of resource allocation, dynamically adjust the processing complexity according to the sensitivity of the data, and avoid over-encryption of low-risk data. For example, plaintext or lightweight desensitization can be directly applied to low-sensitivity data, saving more than 90% of computing resources compared to secure multi-party computation, significantly improving data processing speed. It can improve the feasibility of parallel data processing, allowing desensitization and feature generation to be performed in parallel for data with different sensitivity levels. For example, while high-sensitivity data undergoes encrypted computation through secure multi-party computation, low-sensitivity data can simultaneously complete plaintext feature extraction, thereby improving overall processing efficiency and meeting the timeliness requirements of real-time risk warning. Differential privacy only requires adding local noise to process moderately sensitive data, without the need for multi-round interactions like secure multi-party computation, which can significantly reduce cross-institutional communication. In other words, the method achieves a good balance between data security, computational efficiency, and model accuracy through a "sensitivity classification-technology adaptation" strategy.

[0076] Return to reference Figure 2 In step S230, the multimodal feature data is fused to obtain fused feature data with semantics of the multimodal feature data.

[0077] Figure 5 This is a detailed flowchart of the method for fusing multimodal feature data according to embodiments of this application. (See reference...) Figure 2 and Figure 5 Step S230 may further include sub-steps S510~S530.

[0078] In sub-step S510, at least two of the structured data, text data, temporal data, and graph structure data are embedded into a unified semantic space by a multimodal encoder to generate multiple embedding vectors in the unified semantic space.

[0079] In this sub-step, a multimodal encoder embeds data of different sensitivities and modalities, such as structured data, text data, temporal data, and graph structure data obtained in step S220, into a unified semantic space. That is, in sub-step S510, different modalities are transformed into a unified semantic space R^d (d being the spatial dimension) using a specific encoder corresponding to each modality and a shared mapping layer. For example, structured data can be encoded using a dedicated representation encoder; text data can be encoded using a pre-trained large language model; temporal data can have local and global dynamic features extracted using neural networks; graph structure data can have node semantics extracted using graph neural networks; and the input data from all modalities is aligned to a unified spatial dimension via a linear layer or multilayer perceptron, outputting a data set in a shared semantic space. For example, each modality can be projected into the unified semantic space R^d through a modality-specific encoder and a general linear mapping layer.

[0080] It should be understood that a dedicated representation encoder can be understood as an encoder designed for a specific type of data or task, capable of transforming input data into a feature representation suitable for model processing. It can extract key features in a customized manner based on the characteristics of the data and task requirements, and compared to a general encoder, it can process specific data more efficiently and accurately, improving the model's performance on relevant tasks.

[0081] Figure 6 This is a detailed flowchart of the method for embedding multimodal data into a unified semantic space according to embodiments of this application. (See reference...) Figure 2 , Figure 5 and Figure 6 Sub-step S510 may further include sub-steps S610~S650.

[0082] In sub-step S610, the structured data is encoded using a dedicated characterization encoder to obtain a first encoded feature.

[0083] In sub-step S620, the text data is encoded using a pre-trained large language model to obtain a second encoded feature.

[0084] In sub-step S630, a pre-trained deep learning model is used to extract local and global dynamic features from the time-series data to obtain the third encoded features.

[0085] In sub-step S640, node semantics are extracted from the graph structure data using a pre-trained graph neural network to obtain the fourth encoded feature.

[0086] In sub-step S650, the first coding feature, the second coding feature, the third coding feature and the fourth coding feature are embedded into a unified semantic space to generate multiple embedding vectors in the unified semantic space.

[0087] For example, for the structured data X_smpc generated by secure multi-party computation, in sub-step S610, the structured data X_smpc is encoded using a dedicated representation encoder to obtain a first encoded feature, such as an embedding vector E_smpc(x). For example, the embedding vector E_smpc(x) = W_smpc·Enc_smpc(x), where X_smpc can be the structured data generated by secure multi-party computation, such as an encrypted numerical matrix; Enc_smpc(x) can be the intermediate feature extraction result after the table encoder processes the structured data X_smpc; and W_smpc can be a learnable linear mapping matrix used to project the intermediate features onto R^d.

[0088] For example, for the structured data X_dp after differential privacy processing, in sub-step S610, the structured data X_dp is encoded using a dedicated representation encoder to obtain a first encoded feature, such as an embedding vector E_dp(x). For example, the embedding vector E_dp(x) = W_dp·Enc_dp(x), where X_dp is the sensitive data after differential privacy processing; Enc_dp(x) is the encoder suitable for differential privacy data; and W_dp is a linear mapping parameter shared with other modalities.

[0089] For example, for text data X_text, in substep S620, the text data X_text is encoded using a pre-trained large language model to obtain a second encoded feature, such as an embedding vector E_text(x). For example, the embedding vector E_text(x) = W_text·LLM(x), where LLM(x) is the text feature extracted by the large language model, and W_text is used to compress the high-dimensional text feature to R^d.

[0090] For example, for time-series data X_ts (e.g., transaction records, price fluctuations, etc.), in sub-step S630, a pre-trained deep learning model is used to extract local and global dynamic features from the time-series data X_ts to obtain a third encoded feature, such as an embedding vector E_ts(x). For example, the embedding vector E_ts(x) = W_ts·TSEnc(x), where TSEnc(x) is the output of the time encoder, and W_ts is the shared mapping parameter.

[0091] For example, for graph-structured data X_graph (e.g., transaction networks, relationships, etc.), in sub-step S640, a pre-trained graph neural network is used to extract node semantics from the graph-structured data X_graph to obtain a fourth encoded feature, such as an embedding vector E_graph(x). For example, the embedding vector E_graph(x) = W_g·GCN(x), where GCN(x) is the node or graph-level feature extracted by the graph convolutional network, and W_g is the corresponding linear mapping parameter.

[0092] It should be noted that in the above formulas, Enc_·(x) represents the modality-specific encoder, and W_· represents the learnable parameters of the unified mapping layer. Their function is to eliminate representational differences between different encoders, ensuring that semantically similar data from different modalities are close in distance within R^d, and vice versa. In other words, this process can include two steps: "modality-specific feature extraction (h=Enc_·(x))" and "cross-modality alignment (E(x)=W·h)". It employs a modular design—each modality is equipped with an independent encoder, while the shared mapping space R^d ensures consistent embedding vector dimensions, balancing modal characteristics with cross-modal comparability.

[0093] In some embodiments of this application, to address the semantic drift problem that may occur after mapping partial modal data (e.g., data processed by secure multi-party computation or differential privacy noise enhancement), a modal residual correction mechanism can be introduced. In this mechanism, modal pseudo-labels (e.g., corporate financial health scores) can be constructed as cross-modal consistency targets to assist in correcting embedding bias; cross-modal KL divergence loss (e.g., measuring the distribution difference between E_smpc and E_text) can also be used to reduce the distribution differences between different modal vectors; and trainable residual correction terms can be added to the vectors from secure multi-party computation to further correct semantic shifts.

[0094] It should be noted that the KL divergence loss is a loss function built based on the KL divergence, primarily used to measure the difference between two probability distributions.

[0095] In some embodiments of this application, a cross-modal attention mechanism can be introduced to enhance semantic associations between different modalities. For example, a shared structure can be utilized, using text data as the query entry point and embedding vectors of modalities such as structured data, time-series data, and graph-structured data as keys / values, to guide the alignment of data from different modalities and model semantic associations. For instance, when the keyword "related transaction" appears in the text, the attention mechanism can guide attention to the embedding of the "hidden collateral amount" field in the structured data, which can both generate an interpretable attention graph and help downstream risk reasoning focus on key areas.

[0096] In sub-step S520, the distribution of the multiple embedding vectors is aligned through adversarial training between the modality discriminator and the multimodal encoder.

[0097] In this sub-step S520, the distribution alignment of multiple embedding vectors can be achieved through adversarial training between the modality discriminator and the multimodal encoder, thereby improving the consistency of the embedding vector representation of different modal data, strengthening the semantic expression capability of encrypted modalities, and enhancing the robustness and discriminative ability of the model in privacy protection scenarios.

[0098] For example, in adversarial training, a secure multi-party computation modality encoder (representing the structured data modality after secure multi-party computation processing) (which can be understood as an imitator) and a modality discriminator (which can be understood as a referee) are introduced. Both are optimized together through adversarial learning. The goal of the secure multi-party computation modality encoder is to encode the structured data after secure multi-party computation processing into an embedding vector E_smpc(x), and to make this embedding vector as "disguised" as possible as the text modality embedding vector E_text(x), so that the discriminator cannot distinguish its source. The modality discriminator receives the embedding vector E_smpc(x) and the text embedding vector E_text(x), and determines whether the input vector comes from the secure multi-party computation modality or the text modality (outputting the corresponding probability), with the goal of accurately distinguishing between the two. In the initial stage, the distributions of the embedding vector E_smpc(x) and the text embedding vector E_text(x) differ significantly, making it easy for the discriminator to distinguish them. As training progresses, the secure multi-party computation modality encoder adjusts its parameters based on the discriminator's judgment results, making the embeddings closer to the text semantic features. The discriminator also improves its discrimination ability to avoid being "deceived." When the Nash equilibrium is finally reached, the discriminator can no longer distinguish between the two (e.g., the accuracy is close to 50%). At this point, the distribution of the embedding vector E_smpc(x) is close to the text semantic distribution, and the adversarial training objective is achieved.

[0099] In adversarial training for multimodal semantic fusion, the distribution of embedding vectors from different modalities can be forced to align, improving cross-modal consistency. Different modalities have significant inherent differences (e.g., secure multi-party computation (MPC) embedding vectors are encrypted structured data, while text is natural language). Even when projected into the same dimensional space, their semantic distributions may still be disjointed (e.g., the value of "revenue growth" in an MPC embedding vector is far from the text's "significant improvement in company performance"). Adversarial training, by having MPC embedding vectors mimic the text distribution, forces both to speak the same language in a unified semantic space, ensuring high similarity in both dimensionality and semantic probability distribution. This can compensate for the information loss in encrypted data and enhance the semantic expressive power of the MPC modality. Data processed by MPC loses some information after encryption or privacy processing, potentially resulting in semantically impoverished embedding vectors. Text modality, however, has a rich semantic distribution. By mimicking this distribution, MPC embedding vectors can indirectly supplement missing information, enabling encrypted data embeddings to express complete semantics. This can also enhance the model's robustness to modal differences and improve its discriminative ability in privacy-preserving scenarios. In practical applications, data quality may be unstable (e.g., changes in encryption methods for secure multi-party computation, textual ambiguity, etc.). Adversarial training, through the game between the discriminator and encoder, exposes the model to a large number of modal differences during training. Ultimately, this makes the secure multi-party computation embedding vector less sensitive to its own noise and changes in encryption methods, and the text embedding vector is more inclusive of the special characteristics of the secure multi-party computation embedding vector, ensuring that even if the data is encrypted, the model can still stably extract consistent risk features.

[0100] For example, the input modality representation can be defined first. For instance, the three main types of modality embedding vectors are as follows: E_smpc∈R^{n×d}, which represents the embedding vector of highly sensitive structured data after secure multi-party computation, where n is the number of samples and d is the embedding dimension; E_dp∈R^{m×d}, which represents the embedding vector of moderately sensitive fields after differential privacy noise addition, where m is the number of samples; and E_text / ts∈R^{t×d}, which represents the embedding vector of unencrypted modal data such as text and time series, where t is the number of samples. Secondly, an adversarial perturbation generator is constructed. This generator adds a perturbation δ to E_smpc and E_dp to make them closer to the distribution of plaintext semantic vectors such as E_text. For example, this can be achieved using δ=G_θ(E_enc) and E_adv=E_enc+ε*Normalize(δ), where E_enc∈{E_smpc,E_dp} is the mode to be perturbed, G_θ is a small multilayer perceptron perturbation network, ε controls the perturbation strength, and Normalize(⋅) ensures that the perturbation amplitude is limited (to prevent privacy leakage). The design requirements of this adversarial perturbation generator are that δ cannot be used to reverse-engineer the original data (to protect privacy), and E_adv must remain semantically unchanged but allow for mode confusion (to enhance alignment). Then, a modality discriminator is constructed. A multi-class classifier D_φ(E) can be designed, for example, D_φ(E)∈{smpc,dp,text,ts}, to identify the source modality of the embedded vector. smpc represents the data modality after secure multi-party computation, dp represents the data modality after differential privacy noise processing, text represents the modality of text data, and ts represents the modality of time series data.For example, the optimization loss of the modality discriminator is L_disc=CrossEntropy(D_φ(E),modal_label), where L_disc represents the loss function of the modality discriminator, used to measure the discriminator's judgment error of the source modality of the embedding vector; CrossEntropy represents the cross-entropy loss function, used to measure the difference between the model's prediction and the true label. Here, it quantifies the discriminator's classification error by calculating the distance between the predicted probability distribution of the discriminator's output and the label distribution of the true source modality of the embedding vector; D_φ(E) represents the output of the modality discriminator, where D is the modality discriminator model and φ is the discriminator's learnable parameters (such as weights, biases, etc.). E is the embedding vector input to the discriminator (which can be an embedding of any modality such as E_smpc, E_dp, E_text / ts, etc.). The output of D_φ(E) is a probability distribution representing the probability that the discriminator judges the input embedding vector to come from each modality (e.g., smpc, dp, text, ts). modal_label represents the true source modality label of the embedding vector. For example, if the input embedding vector comes from the smpc modality, then its label is smpc; if it comes from the text modality, the label is text. modal_label can be a one-hot encoded or class-indexed true label, used to compare with the discriminator's prediction result D_φ(E) to calculate the cross-entropy loss. The encoder's optimization objective is L_adv = -L_disc, that is, to force different modalities to converge to the same distribution by "deceiving" the discriminator. Here, L_adv represents the adversarial loss, which is the optimization objective of the multimodal encoder in adversarial training.

[0101] Optionally, in sub-step S530, the distribution of the plurality of embedded vectors is compared and optimized using a temperature-scaled contrastive loss function.

[0102] In this sub-step S530, by introducing modality contrast training, a modality contrast loss function can be designed to enhance the consistency of the embedding space, thereby alleviating the problem of missing information in encrypted data.

[0103] For example, after the above steps, an embedded vector representation encoded by a multimodal encoder can be obtained, such as E_smpc, E_dp, E_text, E_ts, etc.

[0104] Then, positive and negative sample pairs can be constructed. Positive sample pairs can bring multimodal data with the same semantics closer together, while negative sample pairs can distance multimodal data with different semantics, thereby making the distribution of multimodal embedding vectors in a unified space tend to be uniform. Table 6 shows an example of a constructed positive sample pair, for example, the positive sample pair can be multimodal source data from the same enterprise; Table 7 shows an example of a constructed negative sample pair, for example, the negative sample pair can be data from different enterprises or mismatched combinations of different modalities.

[0105] Table 6 Examples of positive sample pairs

[0106]

[0107] Table 7 Examples of Negative Sample Pairs

[0108]

[0109] Next, a loss function can be constructed. In embodiments of this application, a contrastive loss function incorporating temperature scaling can be constructed. For example, an example of a contrastive loss function incorporating temperature scaling is as follows:

[0110]

[0111] It should be noted that the denominator includes all negative sample embeddings in the current batch. In practical applications, the total loss is the average over all positive sample pairs, which can be expressed as follows:

[0112]

[0113] In some exemplary embodiments, various modality-aware enhancement strategies can be employed to improve the robustness and generalization ability of multimodal contrastive training. For example, a modality matching masking strategy can be used, randomly masking one modality in each training batch and allowing only the remaining modalities to undergo contrast consistency training. This avoids the model's over-reliance on modalities with strong expressive power (e.g., text), forcing the model to autonomously learn semantic associations from other modalities, thereby enhancing its adaptability to scenarios where a single modality is missing. For example, an intra-modal contrastive strategy can be used, focusing on feature differentiation within the same modality. For instance, comparing and training two embedding vectors E_smpc_i and E_smpc_j from a secure multi-party computation modality strengthens the ability to distinguish different samples within the encrypted modality, helping the model accurately identify differences between different risk levels even when relying solely on the secure multi-party computation field, thus improving semantic discrimination accuracy in a single modality. For example, a temperature adaptive mechanism can be used to dynamically adjust the temperature parameter τ according to the characteristics of different modalities. For example, for encrypted modalities, a higher temperature parameter τ can be used to enhance the tolerance during training and prevent the model from overfitting weak semantic signals due to noise or information loss that may exist in encrypted data; for plaintext modalities (such as text and time series data), a lower temperature parameter τ can be used to improve the model's accuracy in capturing and expressing its fine semantics, allowing for clearer distinction of subtle semantic differences in plaintext modalities.

[0114] In the embodiments of this application, an overall joint optimization objective can also be constructed. For example, the overall joint optimization objective L can be achieved by combining multiple loss functions. total :L total =L task +α*L adv +β*L contrast +γ*L reg L task Loss for primary tasks (e.g., risk prediction, scoring), L adv To combat modal indistinguishable loss, L contrast For cross-modal alignment consistency loss, L reg The perturbation regularization term (used to prevent overfitting and privacy leaks) uses α, β, and γ as adjustable weight parameters. For example, the main task loss L... task This can represent the supervised loss relevant to a specific application, such as the cross-entropy loss used in risk prediction tasks or the mean squared error used in regression tasks. Its role is to ensure that the model can learn features directly related to business objectives. Adversarial loss L adv This can represent the loss of the modality discriminator, used to measure the distributional differences in embeddings of different modes. Contrastive loss L contrastBy forcing semantically related multimodal pairs to move closer to each other in space, cross-modal semantic consistency is enhanced, playing a crucial role, especially in information completion for encrypted data. The perturbation regularization term L... reg This is a stability loss that adds a small perturbation to the input data. Its purpose is twofold: firstly, to prevent overfitting and make the model more robust to input perturbations; and secondly, to protect privacy by masking sensitive features through perturbation, preventing attackers from using the model to reverse engineer the original data. The weight parameters α, β, and γ each play a different regulatory role: α controls the intensity of adversarial training; a larger value indicates a higher level of focus on modality distribution alignment. β controls the intensity of contrastive learning; a larger value indicates a stronger constraint on semantically relevant pairs. γ is used to balance the model's robustness and privacy protection.

[0115] Return to reference Figure 2 In step S240, the fused feature data is processed using a large model to obtain risk warning results.

[0116] For example, in some embodiments, step S240 may include: a semantic enhancement sub-step, which uses a large model to identify fused feature data in order to identify implicit risk patterns in financial statements; a time-series anomaly detection sub-step, which uses a large model to model the fused feature data and identify high-frequency abnormal behaviors within a period; and a graph risk scoring sub-step, which uses structured graphs such as equity, guarantee chains, and supply chains to construct a risk transmission graph and calculates the probability of default through path scoring methods such as attenuation propagation.

[0117] Figure 7 This is a detailed flowchart illustrating the method for generating risk warning results according to embodiments of this application. (See reference...) Figure 2 and Figure 7 Step S240 may further include sub-steps S710 to S730.

[0118] In sub-step S710, the fused feature data is identified using a large model to obtain a credit risk score. For example, the credit risk score can achieve a probability mapping from multimodal features to "whether or not to default" through weight learning.

[0119] In sub-step S720, the large model is used to identify the similarity between the fused feature data and the current time-series behavior in order to perform anomaly detection on the current time-series behavior. For example, anomaly detection can use the similarity between the feature vector obtained by multimodal fusion and the most recent time-series behavior to determine whether it deviates from the cluster center of normal transaction behavior.

[0120] In sub-step S730, based on the fused feature data, a risk transmission graph is constructed using a graph neural network, and the default probability is calculated through path scoring. For example, in the risk extrapolation of the guarantee chain, the feature vector obtained by multimodal fusion can be fed into the graph neural network node initialization to predict the probability of on-chain risk contagion.

[0121] In this embodiment, the technical solution of using a large model to process fused feature data to obtain risk warning results achieves comprehensiveness and accuracy of risk assessment through multi-dimensional analysis. It covers the risk quantification of the subject itself and the monitoring of abnormal behavior in real time, while also taking into account the risk transmission analysis of related networks. This comprehensively improves the accuracy, timeliness and depth of risk warning, and provides multi-level and three-dimensional support for risk control decision-making.

[0122] Optionally, in some exemplary embodiments, step S240 may include: a risk scoring sub-step, which outputs a risk probability or score value as a direct quantitative basis for risk control decisions, with a unified risk semantic vector as input and a continuous risk score or multi-level risk level label as output; an abnormal behavior identification sub-step, which identifies atypical high-risk behaviors such as potential fraud, abnormal operations, and data manipulation, and can be based on a self-supervised training anomaly detection model, by comparing score thresholds and using a unified risk vector and the original transaction sequence in parallel encoding to perform behavioral anomaly detection, with output including anomaly label, type label, anomaly occurrence probability, and explanatory signal; and a rule fusion sub-step, which integrates... The fusion of expert rules and model results enhances transparency and control in regulatory scenarios. It allows for dual-channel integration of model results and rule systems, and adds a post-screening mechanism through rule logic trees (e.g., a model with a moderate score but "frequent large-amount transactions at night" raises the risk level). The output is a final risk control judgment label that integrates machine learning and rules. The decision explanation and visualization sub-step provides an interpretable path for risk control decisions, meeting audit and compliance requirements. It can use explanation tools to display the feature weights, original fields, time series segments, and semantic segments on which the model decision depends, constructing a "causal inference path," and outputting a decision report, a traceable causal chain, and a structured audit report.

[0123] Optionally, in some embodiments of this application, before performing desensitization and feature generation processing on the multimodal data, the method may further include: determining a desensitization and feature generation strategy for the multimodal data based on the permission level and importance level of the multimodal data, wherein the permission level is used to characterize the sensitivity of the multimodal data, the importance level is used to characterize the degree of influence of the multimodal data on the risk warning result, and the desensitization and feature generation strategy includes secure multi-party computation, differential privacy noise addition, static desensitization processing, or plaintext processing.

[0124] For example, data can be managed hierarchically, such as assigning three permission levels and one importance level to each data point. The permission levels characterize the sensitivity of the multimodal data, and the importance level characterizes the impact of the multimodal data on the risk warning results. The combined effect of the permission levels and importance levels controls which desensitization and feature generation strategy is used: secure multi-party computation, differential privacy noise addition, static desensitization, or plaintext processing. In other words, dynamic privacy control is implemented, and the corresponding desensitization and feature generation strategy is determined based on the results of this dynamic privacy control.

[0125] Table 8 provides examples of dynamic privacy control. For instance, referring to Table 8, data A could be tax revenue, with a high access level (high sensitivity) and a high impact level (high influence), meaning it has a high impact on the risk warning result. The recommended strategy is to perform desensitization and feature generation through secure multi-party computation. Through dynamic privacy control, the final processing strategy is to perform desensitization and feature generation through secure multi-party computation. Data B could be the company's debt ratio, with a medium access level (medium sensitivity) and a high impact level, meaning it has a high influence on the risk warning result. The recommended strategy is to perform desensitization and feature generation through differential privacy noise addition. Through dynamic privacy control, considering its high impact level, the final processing strategy is to perform desensitization and feature generation through secure multi-party computation. Data C could be the business term, with a low access level (low sensitivity) and a weak impact level, meaning it has a low influence on the risk warning result. The recommended strategy is plaintext processing. Through dynamic privacy control, the final processing strategy is plaintext processing. Data D can be the province of registration, with a low permission level (low sensitivity) and a weak impact level (low influence on the risk warning result). The recommended strategy is plaintext processing, with dynamic privacy control, and the final processing strategy is to ignore it.

[0126] Table 8 Examples of Dynamic Privacy Controls

[0127]

[0128] In the embodiments of this application, a dynamic privacy control mechanism can be used to dynamically balance the strength of data privacy protection with the accuracy required for modeling, thereby avoiding feature distortion or over-encryption.

[0129] Based on the above method, embodiments of this application also provide a risk warning device based on a large model and secure multi-party computation. The following will combine... Figure 8 The device is described in detail.

[0130] Figure 8A schematic diagram illustrates the structural block diagram of a risk warning device based on a large model and secure multi-party computation according to an embodiment of this application. Figure 8 As shown, the risk warning device 800 based on large model and secure multi-party computation in this embodiment may include a behavioral multimodal data acquisition module 810, a multimodal feature data acquisition module 820, a fusion feature data acquisition module 830, and a risk warning result acquisition module 840.

[0131] The multimodal data acquisition module 810 is used to acquire multimodal data upon obtaining user authorization or consent, wherein the multimodal data comes from multiple data providers. In one embodiment, the multimodal data acquisition module 810 can be used to execute step S210 and its sub-steps described above, which will not be repeated here.

[0132] The multimodal feature data acquisition module 820 is used to perform desensitization and feature generation processing on the multimodal data to obtain multimodal feature data. The desensitization and feature generation processing includes at least using secure multi-party computation to perform desensitization and feature generation processing on at least some data in the multimodal data. The multimodal feature data includes at least two of structured data, text data, time-series data, and graph-structured data. In one embodiment, the multimodal feature data acquisition module 820 can be used to execute step S220 and its sub-steps described above, which will not be repeated here.

[0133] The fusion feature data acquisition module 830 is used to fuse the multimodal feature data to obtain fused feature data with semantics of the fused multimodal feature data. In one embodiment, the fusion feature data acquisition module 830 can be used to execute step S230 and its sub-steps described above, which will not be repeated here.

[0134] The risk warning result acquisition module 840 is used to process the fused feature data using a large model to obtain risk warning results. In one embodiment, the risk warning result acquisition module 840 can be used to execute step S240 and its sub-steps described above, which will not be repeated here.

[0135] According to embodiments of this application, any multiple modules among the multimodal data acquisition module 810, multimodal feature data acquisition module 820, fused feature data acquisition module 830, and risk warning result acquisition module 840 can be combined into one module, or any one of these modules can be split into multiple modules. Alternatively, at least some of the functions of one or more of these modules can be combined with at least some of the functions of other modules and implemented in one module. According to embodiments of this application, at least one of the multimodal data acquisition module 810, multimodal feature data acquisition module 820, fused feature data acquisition module 830, and risk warning result acquisition module 840 can be at least partially implemented as hardware circuitry, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-a-chip, a system-on-a-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or implemented in hardware or firmware by any other reasonable means of integrating or packaging the circuitry, or implemented in any one of the three implementation methods of software, hardware, and firmware, or in a suitable combination of any of these. Alternatively, at least one of the multimodal data acquisition module 810, multimodal feature data acquisition module 820, fusion feature data acquisition module 830, and risk warning result acquisition module 840 can be at least partially implemented as a computer program module, which can perform corresponding functions when the computer program module is run.

[0136] Figure 9 The diagram illustrates an electronic device suitable for implementing a risk warning method based on a large model and secure multi-party computation, according to an embodiment of this application.

[0137] like Figure 9 As shown, an electronic device 900 according to an embodiment of this application includes a processor 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage portion 908 into a random access memory (RAM) 903. The processor 901 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 901 may also include onboard memory for caching purposes. The processor 901 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of this application.

[0138] RAM 903 stores various programs and data required for the operation of electronic device 900. Processor 901, ROM 902, and RAM 903 are interconnected via bus 904. Processor 901 executes various operations of the method flow according to embodiments of this application by executing programs in ROM 902 and / or RAM 903. It should be noted that the programs may also be stored in one or more memories other than ROM 902 and RAM 903. Processor 901 may also execute various operations of the method flow according to embodiments of this application by executing programs stored in said one or more memories.

[0139] According to embodiments of this application, the electronic device 900 may further include an input / output (I / O) interface 905, which is also connected to a bus 904. The electronic device 900 may also include one or more of the following components connected to the input / output (I / O) interface 905: an input section 906 including a keyboard, mouse, etc.; an output section 907 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 908 including a hard disk, etc.; and a communication section 909 including a network interface card such as a LAN card, modem, etc. The communication section 909 performs communication processing via a network such as the Internet. A drive 910 is also connected to the input / output (I / O) interface 905 as needed. A removable medium 911, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 910 as needed so that computer programs read from it can be installed into the storage section 908 as needed.

[0140] This application also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs, which, when executed, implement the method according to the embodiments of this application.

[0141] According to embodiments of this application, the computer-readable storage medium can be a non-volatile computer-readable storage medium, such as including but not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this application, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to embodiments of this application, the computer-readable storage medium may include ROM 902 and / or RAM 903 and / or one or more memories other than ROM 902 and RAM 903 described above.

[0142] Embodiments of this application also include a computer program product comprising a computer program containing program code for performing the methods shown in the flowchart. When the computer program product is run on a computer system, the program code enables the computer system to implement the risk warning method based on large models and secure multi-party computation provided in the embodiments of this application.

[0143] When the computer program is executed by the processor 901, it performs the functions defined in the system / apparatus of this application embodiment. According to the embodiments of this application, the systems, apparatuses, modules, units, etc., described above can be implemented by computer program modules.

[0144] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0145] Those skilled in the art will understand that the features described in the various embodiments of this application can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in this application. In particular, the features described in the various embodiments of this application can be combined and / or combined in various ways without departing from the spirit and teachings of this application. All such combinations and / or combinations fall within the scope of this application.

Claims

1. A risk early warning method based on large models and secure multi-party computation, characterized in that, The method includes: Multimodal data is collected with the user's authorization or consent, wherein the multimodal data comes from multiple data providers; The multimodal data is subjected to desensitization and feature generation processing to obtain multimodal feature data. The desensitization and feature generation processing includes at least some data in the multimodal data being desensitized and feature generated using secure multi-party computation. The multimodal feature data includes at least two of structured data, text data, time-series data, and graph structure data. The multimodal feature data is fused to obtain fused feature data with semantic meaning. The fused feature data is processed using a large model to obtain risk warning results.

2. The method according to claim 1, characterized in that, The multimodal data includes a first type of data, a second type of data, and a third type of data. The sensitivity of the first type of data is higher than that of the second type of data, and the sensitivity of the second type of data is higher than that of the third type of data. The desensitization and feature generation processes include: Secure multi-party computation is used to de-identify and generate features for the first type of data; Differential privacy noise reduction is used to de-identify and generate features for the second type of data; and The third type of data is desensitized and feature generated using static desensitization or plaintext processing methods.

3. The method according to claim 2, characterized in that, The process of desensitizing and generating features for the first type of data using secure multi-party computation includes: encoding the first type of data from multiple data providers into multiple first vectors respectively; The plurality of first vectors are encrypted and encoded using an encryption protocol to generate a plurality of first encrypted encoded vectors; and Based on pre-constructed risk characteristics, secure multi-party computation is performed on multiple first cryptographic encoding vectors from multiple data providers to generate first feature data; and / or, The process of desensitizing and generating features for the second type of data using differential privacy noise addition includes: Based on a preset differential privacy perturbation mechanism, noise is added to the second type of data from multiple data providers to generate multiple second encrypted encoding vectors. The preset differential privacy perturbation mechanism includes: adding Laplace noise to numerical data and adding hot encoding perturbation to categorical data; and Based on pre-built risk features, secure multi-party computation is performed on multiple second cryptographic encoding vectors from multiple data providers to generate second feature data.

4. The method according to claim 3, characterized in that, The process of fusing the multimodal feature data to obtain semantically fused feature data includes: A multimodal encoder is used to embed at least two of the structured data, text data, temporal data, and graph structure data into a unified semantic space to generate multiple embedding vectors in the unified semantic space; and The distribution of the multiple embedding vectors is aligned through adversarial training between the modality discriminator and the multimodal encoder.

5. The method according to any one of claims 1-4, characterized in that, The process of fusing the multimodal feature data to obtain fused feature data with semantic meaning further includes: The distribution of the multiple embedding vectors is optimized by using a temperature-scaled contrastive loss function.

6. The method according to claim 5, characterized in that, The step of embedding at least two of the structured data, text data, temporal data, and graph structure data into a unified semantic space using a multimodal encoder to generate multiple embedding vectors in the unified semantic space includes: The structured data is encoded using a dedicated characterization encoder to obtain a first encoded feature; The text data is encoded using a pre-trained large language model to obtain a second encoding feature; The time-series data is subjected to local and global dynamic feature extraction using a pre-trained deep learning model to obtain a third encoded feature; The graph structure data is subjected to node semantic extraction using a pre-trained graph neural network to obtain a fourth encoded feature; and The first encoding feature, the second encoding feature, the third encoding feature, and the fourth encoding feature are embedded into a unified semantic space to generate multiple embedding vectors in the unified semantic space.

7. The method according to any one of claims 1-4 and 6, characterized in that, The process of using a large model to process the fused feature data to obtain risk warning results includes: The fused feature data is identified using a large model to obtain a credit risk score; The similarity between the fused feature data and the current temporal behavior is identified using a large model to perform anomaly detection on the current temporal behavior; and Based on the fused feature data, a risk transmission map is constructed using a graph neural network, and the probability of default is calculated through path scoring.

8. The method according to claim 7, characterized in that, Before performing desensitization and feature generation processing on the multimodal data, the method further includes: Based on the permission level and importance level of the multimodal data, a privacy label is determined for the multimodal data, wherein the permission level characterizes the sensitivity of the multimodal data, and the importance level characterizes the impact of the multimodal data on the risk warning result; and Based on the pre-built mapping relationship between privacy labels and desensitization and feature generation strategies, the desensitization and feature generation strategies of the multimodal data are determined according to the privacy labels. The desensitization and feature generation strategies include secure multi-party computation, differential privacy noise addition, static desensitization processing, or plaintext processing.

9. A risk early warning device based on a large model and secure multi-party computation, characterized in that, The device includes: A multimodal data acquisition module is used to acquire multimodal data with the user's authorization or consent, wherein the multimodal data comes from multiple data providers; A multimodal feature data acquisition module is used to perform desensitization and feature generation processing on the multimodal data to obtain multimodal feature data. The desensitization and feature generation processing includes at least using secure multi-party computation to perform desensitization and feature generation processing on at least some data in the multimodal data. The multimodal feature data includes at least two of structured data, text data, time-series data, and graph structure data. A fusion feature data acquisition module is used to fuse the multimodal feature data to obtain fusion feature data with semantic meaning derived from the fusion of the multimodal feature data; and The risk warning result acquisition module is used to process the fused feature data using a large model to obtain risk warning results.

10. An electronic device, comprising: One or more processors; Memory, used to store one or more computer programs. The characteristic feature is that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 8.

11. A computer-readable storage medium having a computer program or instructions stored thereon, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the method according to any one of claims 1 to 8.

12. A computer program product, comprising a computer program or instructions, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the method according to any one of claims 1 to 8.