Digital RMB related data processing method and device based on large model technology
Through large-scale model technology, digital RMB transaction data is processed, semantic alignment and entity relationship network construction is carried out, and an exception path is identified by a graph-enhanced inference model, which solves the multi-dimensional semantic correlation-intensive data mining and intelligent analysis requirements in digital RMB data processing, and realizes efficient anomaly detection and compliance monitoring.
Patent Information
- Application Number
- CN202510751381.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-06
AI Technical Summary
The existing technology is difficult to meet the multi-dimensional semantic correlation-intensive data mining and intelligent analysis requirements of digital RMB-related data, and there are problems such as slow response, weak semantic understanding ability, and low data fusion.
Using a method based on big model technology, we use multimodal data of digital RMB transaction event streams to obtain semantic alignment and build a unified semantic space, extract target triplets to build an entity relationship network, and combine the graph to enhance the inference model and risk control rules to identify abnormal paths.
It has realized the effective mining and analysis of digital RMB-related data, improved the accuracy and efficiency of abnormal behavior detection, adapted to the rapid changes in the financial environment and policies, and enhanced the security and compliance of the system.
Smart Images

Figure CN120256646A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of financial data processing. Specifically, it relates to a method and device for processing digital RMB-related data based on large model technology. Background Art
[0002] With the rapid development of the digital economy, digital RMB has gradually entered the actual application stage and plays an important role in multiple scenarios such as retail payment, cross-border payment, and fiscal subsidy distribution. The data types it involves are complex, including user behavior data, transaction data, payment channel data, contract execution data, etc., showing the characteristics of coexistence of highly structured and unstructured data and rapid dynamic evolution.
[0003] In this context, how to efficiently process and intelligently analyze digital RMB-related data has become an urgent problem to be solved. Currently, most of the data processing methods in related technologies rely on rule engines and manually set analysis models, which have problems such as slow response, weak semantic understanding ability, and low data fusion degree, and are difficult to meet the multi-dimensional and semantic association-intensive data mining and intelligent decision-making needs of digital RMB.
[0004] For the above problems, no effective solution has been proposed yet. Summary of the Invention
[0005] The embodiments of this application provide a method and device for processing digital RMB-related data based on large model technology to at least solve the technical problem that the data processing methods in related technologies are difficult to meet the multi-dimensional semantic association-intensive data mining and intelligent analysis needs of digital RMB.
[0006] According to one aspect of the embodiments of this application, a method for processing digital RMB-related data based on large model technology is provided, including: obtaining multi-modal data involved in the transaction event stream of digital RMB, and mapping the multi-modal data to a unified semantic space for semantic alignment to obtain target data, where in the unified semantic space, the closer the distance between the vectors corresponding to data with more similar semantics, and the farther the distance between the vectors corresponding to data with greater semantic differences; extracting target triples from the target data, and constructing an entity relationship network based on the target triples to obtain a target knowledge graph, where the target triples contain two entity elements and one relationship element, and the target triples are used to represent the association relationship between two entities; using a graph-enhanced reasoning model, combined with risk control rules, to identify abnormal paths in the target knowledge graph, where the abnormal paths correspond to abnormal processing behaviors of digital RMB, and the risk control rules are used to indicate the characteristics and patterns of the rules that should be followed for the processing behaviors of digital RMB, and the graph-enhanced reasoning model is an inference framework combining graph neural network and large model technology.
[0007] Optionally, the multimodal data includes: structured data and unstructured data of different modalities. Among them, the structured data includes at least one of the following: transaction records, account information, device binding records, and the unstructured data includes at least one of the following: policy texts, user behavior logs, contract descriptions, image vouchers. Mapping the multimodal data into a unified semantic space for semantic alignment, the obtained target data includes: using the modality encoder in the semantic alignment model to extract the semantic features corresponding to the data of different modalities, where the ways of feature extraction for the data of different modalities are different, and the modalities include at least one of the following: text, speech, image, video; using the connector in the semantic alignment model to perform transformation processing on the semantic features corresponding to different modalities, where the transformation processing is used to eliminate the differences in distribution, scale, and noise of the semantic features of different modalities, and obtain semantic features in a unified representation form to ensure that the data of all modalities can be represented in the same semantic space; using the generator in the semantic alignment model to map the transformed semantic features into a unified semantic space to obtain the target data.
[0008] Optionally, the training steps of the semantic alignment model include: obtaining a sample data set, where the sample data set contains multiple sample data, and each sample data contains data of different modalities for describing the same content; performing feature extraction on the data of different modalities in the sample data to obtain feature vectors corresponding to different modalities; pairing the feature vectors of different modalities corresponding to the same sample data to obtain positive sample pairs, and randomly selecting feature vectors of different modalities corresponding to different sample data for pairing to obtain negative sample pairs; forming a training data set with the positive samples and negative sample pairs, and on the basis of the training data set, combining with a contrast loss function to train an initial model to obtain a semantic alignment model, where the contrast loss function is used to make the feature vectors in the positive sample pairs output by the initial model closer in the semantic space during the training process, while the feature vectors in the negative sample pairs are farther away in the semantic space.
[0009] Optionally, the method further includes: determining the data source of the multimodal data; in the case where the data source indicates that the text data in the multimodal data comes from different financial institutions, extracting the target keywords in the text data of different financial institutions, where the target keywords include at least one of the following: term expressions within each financial institution, specific encodings, behavior descriptions; according to the standard expression mapping table, converting the keywords corresponding to different financial institutions into unified standard vocabulary, where the standard expression mapping table is a mapping table from the special terms of different financial institutions to unified standard vocabulary, and the special terms of different financial institutions include the respective data expressions, feature encodings, and behavior descriptions of each financial institution.
[0010] Optionally, construct an entity relationship network based on the target triples to obtain a target knowledge graph, including: using a stream processing engine and a message middleware to listen to the transaction event stream, and when the multimodal data corresponding to the transaction event stream is updated, extract the latest target triples from the target data corresponding to the updated multimodal data; based on the latest target triples, perform entity change awareness and relationship evolution detection to obtain the change information corresponding to the constructed target knowledge graph, where entity change awareness and relationship evolution detection are used to determine the entity elements and / or relationship elements that have changed in the latest target triples compared to the constructed target knowledge graph; according to the change information, perform an update operation on the constructed target knowledge graph to obtain the latest target knowledge graph, where the update operation includes: node insertion, edge creation, and attribute update.
[0011] Optionally, use a graph-enhanced reasoning model and combine risk control rules to identify abnormal paths in the target knowledge graph, including: in the graph neural network of the graph-enhanced reasoning model, use a neighbor aggregation mechanism and a relationship perception mechanism to analyze the feature information of entity nodes and the edges between entity nodes in the target knowledge graph, and generate node representations containing context information corresponding to each entity node, where the neighbor aggregation mechanism is used to integrate the feature information of entity nodes and the neighbor nodes corresponding to the entity nodes, and the relationship perception mechanism is used to model the influence of different types of edges on information transmission, and the feature information of entity nodes includes at least one of the following: account attributes, transaction behavior characteristics, contract status information, and the feature information of edges includes at least one of the following: transaction amount, time interval, call relationship; according to the node representations corresponding to each entity node in the target knowledge graph, identify the behavior patterns corresponding to each node path in the target knowledge graph, where the behavior patterns include at least one of the following: account fund jump, multi-contract call chain, multi-account association control; use the large language model in the graph-enhanced reasoning model, and according to the risk control rules corresponding to each behavior pattern, perform logical constraint verification and conflict detection on the node paths to obtain the risk probability corresponding to each node path, and determine the node paths with risk probability exceeding the preset risk probability as abnormal paths.
[0012] Optionally, the method further includes: when it is detected that the policy text related to digital currency in the information source is updated, obtain the latest policy text; use a large language model to analyze the policy text to identify the constrained entities in the policy text, as well as the restriction conditions and business logics corresponding to the constrained entities; according to the restriction conditions and business logics, generate a syntax constraint graph structure corresponding to the constrained entities, and convert the syntax constraint graph structure into the corresponding structured risk control rules.
[0013] According to another aspect of the embodiments of the present application, there is also provided an apparatus for processing digital RMB-related data based on large model technology, including: a semantic alignment module, configured to obtain multimodal data involved in the transaction event stream of digital RMB, and map the multimodal data into a unified semantic space for semantic alignment to obtain target data, wherein in the unified semantic space, the closer the distance between the vectors corresponding to data with more similar semantics, and the farther the distance between the vectors corresponding to data with greater semantic differences; a network construction module, configured to extract target triples from the target data and construct an entity relationship network based on the target triples to obtain a target knowledge graph, wherein the target triples include two entity elements and one relationship element, and the target triples are used to represent the association relationship between two entities; a risk identification module, configured to adopt a graph-enhanced inference model and combine risk control rules to identify abnormal paths in the target knowledge graph, wherein the abnormal paths correspond to abnormal processing behaviors of digital RMB, and the risk control rules are used to indicate the characteristics and patterns of the rules that should be followed for the processing behaviors of digital RMB, and the graph-enhanced inference model is an inference framework that combines graph neural network and large model technology.
[0014] According to yet another aspect of the embodiments of the present application, there is also provided an electronic device, including: a memory and a processor, and the processor is configured to run a program stored in the memory, wherein when the program runs, it executes a method for processing digital RMB-related data based on large model technology.
[0015] According to still another aspect of the embodiments of the present application, there is also provided a non-volatile storage medium, and the non-volatile storage medium includes a stored computer program, wherein the device where the non-volatile storage medium is located executes a method for processing digital RMB-related data based on large model technology by running the computer program.
[0016] In the embodiments of the present application, multimodal data involved in the transaction event stream of digital RMB is acquired, and the multimodal data is mapped into a unified semantic space for semantic alignment to obtain target data. Among them, in the unified semantic space, the closer the distance between the vectors corresponding to data with more similar semantics, the farther the distance between the vectors corresponding to data with greater semantic differences; target triples are extracted from the target data, and an entity relationship network is constructed based on the target triples to obtain a target knowledge graph. Among them, the target triples contain two entity elements and one relationship element, and the target triples are used to represent the association relationship between two entities; a graph-enhanced reasoning model is used to identify abnormal paths in the target knowledge graph in combination with risk control rules. Among them, the abnormal paths correspond to abnormal handling behaviors of digital RMB, and the risk control rules are used to indicate the characteristics and patterns of the rules that should be followed for the handling behaviors of digital RMB. The graph-enhanced reasoning model is a way of combining a graph neural network and a large model technology inference framework. By introducing a cross-modal embedding model and a unified semantic space construction mechanism, semantic alignment and knowledge extraction of multi-source data are carried out, and by constructing a large model-driven semantic perception risk control engine, context modeling and intention reasoning of transaction behaviors are realized, achieving the purpose of effectively mining and analyzing digital RMB-related data, and further solving the technical problem that the data processing methods in the related technologies are difficult to meet the data mining and intelligent analysis requirements of the multi-dimensional semantic association-intensive data of digital RMB. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application, and do not constitute an improper limitation to the present application. In the drawings:
[0018] Figure 1 is a hardware structure block diagram of a computer terminal (or electronic device) for implementing a method for processing digital RMB-related data based on large model technology according to an embodiment of the present application;
[0019] Figure 2 is a schematic diagram of a method flow for processing digital RMB-related data based on large model technology according to an embodiment of the present application;
[0020] Figure 3 is a schematic diagram of risk identification evolutionary modeling for the financial field according to an embodiment of the present application;
[0021] Figure 4 is a schematic diagram of the structure of a device for processing digital RMB-related data based on large model technology according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0022] To enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the scope of protection of this application.
[0023] It should be noted that the terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0024] In the related art, in the risk monitoring and data processing technology for digital RMB, there are still structural technical shortcomings in key dimensions such as intelligence, real-time performance, semantic understanding, multi-modal fusion, and system scalability, making it difficult to meet the regulatory and operational requirements of financial institutions for compliance, security, and high availability, specifically including:
[0025] 1) The problem of semantic-level fusion and unified modeling of multi-source heterogeneous financial data;
[0026] The data involved in the digital RMB ecosystem has significant "4V" characteristics: massive (Volume), high-speed (Velocity), diverse (Variety), and high veracity (Veracity). The data types cover structured transaction records, account information, device identifiers, etc., and multi-modal data sources such as unstructured policy texts, user behavior logs, contract descriptions, and image vouchers. In addition, due to the inconsistent data coding standards across financial institutions and platforms, there is strong heterogeneity in the data structure, resulting in problems such as low information extraction accuracy, poor semantic alignment ability, and high information loss rate in traditional data integration solutions based on, severely restricting the unified expression and in-depth analysis of data value.
[0027] 2) The problem of dynamic identification and real-time response to new financial risk models;
[0028] The programmability and account - de - coupling characteristics of digital RMB make it difficult for traditional risk control measures that rely on static rule libraries to capture potential complex risk behaviors. According to the latest monitoring data, the evolution cycle of new attack vectors has been compressed from 36 months in the traditional financial field to 715 days. The proportion of complex risk paths such as cross - chain transactions and contract jumps has increased to 38%, and the proportion of zero - day attacks among all high - risk behaviors has reached 21%. The response latency of existing batch - based monitoring systems is generally between 4 and 6 hours, resulting in about 19% of high - risk transactions not being detected in a timely manner.
[0029] 3) The digital RMB policy system presents the characteristics of "high update frequency, deep structural hierarchy, and complex semantic constraints". The regulatory policies are updated an average of 23 times per month, and cover multiple authoritative institutions. The text content of the policies is semantically complex, with an average of 57 implicit business logic constraints per policy. The traditional rule - modeling method based on manual transformation is extremely inefficient.
[0030] To solve the above problems, relevant solutions are provided in the embodiments of this application, which are described in detail below.
[0031] According to the embodiments of this application, a method embodiment for processing digital RMB - related data based on large - model technology is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer - executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0032] The method embodiments provided by the embodiments of this application can be executed on a mobile terminal, a computer terminal, or a similar computing device. Figure 1 A hardware structure block diagram of a computer terminal (or electronic device) for implementing a method for processing digital RMB - related data based on large - model technology is shown. As Figure 1 shown, the computer terminal 10 (or electronic device) may include one or more (shown as 102a, 102b,..., 102n in the figure) processors 102 (the processor 102 may include, but is not limited to, a processing device such as a micro - processor MCU or a field - programmable gate array FPGA), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may further include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which can be included as one of the ports of the BUS bus), a network interface, a power supply, and / or a camera. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above - mentioned electronic device. For example, the computer terminal 10 may further include more Figure 1more or fewer components shown, or having a configuration different from that shown in Figure 1 that shown.
[0033] It should be noted that one or more of the above-mentioned processors 102 and / or other data processing circuits can generally be referred to as "data processing circuits" herein. The data processing circuit can be embodied in software, hardware, firmware, or any combination thereof, in whole or in part. In addition, the data processing circuit can be a single independent processing module, or be incorporated in whole or in part into any one of other elements in the computer terminal 10 (or electronic device). As involved in the embodiments of the present application, the data processing circuit is a processor control (such as the selection of a variable resistor terminal path connected to an interface).
[0034] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage devices corresponding to the processing method of digital RMB-related data based on large model technology in the embodiments of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implements the above-mentioned processing method of digital RMB-related data based on large model technology. The memory 104 can include high-speed random access memory, and can also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 104 can further include a memory remotely set relative to the processor 102, and these remote memories can be connected to the computer terminal 10 through a network. Examples of the above-mentioned network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and combinations thereof.
[0035] The transmission device 106 is used to receive or send data via a network. Specific examples of the above-mentioned network can include the wireless network provided by the communication provider of the computer terminal 10. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one instance, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0036] The display can be, for example, a touch-screen liquid crystal display (LCD), which enables the user to interact with the user interface of the computer terminal 10 (or electronic device).
[0037] Under the above operating environment, the embodiments of the present application provide a method for processing digital RMB-related data based on large model technology. Figure 2It is a schematic diagram of a method flow for processing digital RMB-related data based on large model technology provided by an embodiment of the present application. As Figure 2 shown, the method includes the following steps:
[0038] Step S202: Obtain multi-modal data involved in the transaction event stream of digital RMB, and map the multi-modal data to a unified semantic space for semantic alignment to obtain target data. Among them, in the unified semantic space, the closer the distance between the vectors corresponding to data with more similar semantics, and the farther the distance between the vectors corresponding to data with greater semantic differences;
[0039] Step S204: Extract target triples from the target data, and construct an entity relationship network based on the target triples to obtain a target knowledge graph. Among them, the target triple contains two entity elements and one relationship element, and the target triple is used to represent the association relationship between two entities;
[0040] Step S206: Use a graph-enhanced reasoning model, combined with risk control rules, to identify abnormal paths in the target knowledge graph. Among them, the abnormal path corresponds to the abnormal handling behavior of digital RMB, and the risk control rules are used to indicate the characteristics and patterns of the rules that should be followed for the handling behavior of digital RMB. The graph-enhanced reasoning model is an inference framework that combines graph neural networks and large model technology.
[0041] Through the above steps, by introducing a cross-modal embedding model and a unified semantic space construction mechanism, semantic alignment and knowledge extraction of multi-source data are carried out, and by constructing a large model-driven semantic perception risk control engine, context modeling and intention reasoning of transaction behaviors are realized, achieving the purpose of effectively mining and analyzing digital RMB-related data, and further solving the technical problem that the data processing method in the related technology is difficult to meet the data mining and intelligent analysis requirements of the multi-dimensional semantic association-intensive data of digital RMB.
[0042] Next, the method for processing digital RMB-related data based on large model technology in steps S202 to S206 of the embodiment of the present application will be further introduced.
[0043] The embodiment of the present application focuses on the data diversity and semantic complexity problems in the digital RMB ecosystem, and proposes a new multi-modal semantic governance system with cross-structure, multi-modal, cross-system semantic connection and knowledge extraction capabilities. The following is a specific introduction.
[0044] First, obtain the multimodal data involved in the transaction event stream of digital RMB. For example, through a highly adaptable data access layer, structured data (such as account information, fund details, device binding records) and unstructured data (such as central bank policy documents, scenario protocol contract texts, transaction voucher images, device behavior logs) can be collected from the core systems of financial institutions, regulatory middle platforms, smart contract execution platforms, etc. All types of data are quasi-real-time collected through the data bus normalization modeling and message queue asynchronous processing mechanism, with high throughput and high-reliability collection capabilities.
[0045] Then, cross-modal representation learning can be performed on the collected multimodal data related to digital RMB. By introducing a unified financial semantic embedding space, data of different modalities can be made comparable and vector-aligned in the semantic dimension. The specific steps are as follows.
[0046] In some embodiments of this application, the multimodal data includes: structured data and unstructured data of different modalities. Among them, the structured data includes at least one of the following: transaction records, account information, device binding records, and the unstructured data includes at least one of the following: policy texts, user behavior logs, contract descriptions, image vouchers; mapping the multimodal data into a unified semantic space for semantic alignment, and obtaining the target data includes the following steps: using the modality encoder in the semantic alignment model to extract the semantic features corresponding to the data of different modalities, where the ways of feature extraction for the data of different modalities are different, and the modalities include at least one of the following: text, speech, image, video; using the connector in the semantic alignment model to perform conversion processing on the semantic features corresponding to different modalities, where the conversion processing is used to eliminate the differences in distribution, scale, and noise of the semantic features of different modalities, and obtain semantic features in a unified representation form to ensure that data of all modalities can be represented in the same semantic space; using the generator in the semantic alignment model to map the semantic features after conversion processing into the unified semantic space to obtain the target data.
[0047] Specifically, the original data of different modalities (images, audio, video, and text) can be converted into feature vectors (semantic features) through the modality encoder of the semantic alignment model. The data of each modality can be converted into a format that can be processed by the model through a specific encoder. For example, images and videos may be encoded through a convolutional neural network, while text may be encoded through language models such as word embeddings. After the feature vectors (semantic features) are extracted, these different modality feature vectors can be converted into a unified representation through a connector to ensure that the data of all modalities can be represented in the same semantic space. Finally, the generator is responsible for generating the final output, mapping the processed semantic features to the unified semantic space to obtain the target data.
[0048] For example, taking three types of heterogeneous modality data, namely text, images, and behavior logs, as an example, a method combining cross-modal representation learning and contrastive learning can be adopted to achieve semantic consistency modeling and vector space alignment. Specifically, text data extracts semantic vectors through a pre-trained language model (such as BERT, RoBERTa, etc.), image data is encoded into visual feature vectors through a vision Transformer, and behavior log data extracts behavior pattern features through a sequence modeling network (such as Transformer-XL). To ensure the semantic consistency of different modality data in the unified space, the system introduces a multi-modal contrastive loss function based on positive and negative sample contrast, prompting the text, image, and log vectors originating from the same financial entity or behavior event to approach each other in the space, while samples with different semantics remain far apart. At the same time, a modality-specific normalization strategy and scale alignment technology are adopted to quantitatively process the differences in distribution, scale, and noise level of different modality features, ensuring the consistency and stability of the semantic space.
[0049] The training process of the above semantic alignment model is introduced below, as follows.
[0050] In some embodiments of the present application, the training steps of the semantic alignment model include: obtaining a sample data set, where the sample data set contains multiple pieces of sample data, and each piece of sample data contains data of different modalities for describing the same content; extracting features from the data of different modalities in the sample data to obtain feature vectors corresponding to different modalities; pairing the feature vectors of different modalities corresponding to the same piece of sample data to obtain positive sample pairs, and randomly selecting feature vectors of different modalities corresponding to different sample data for pairing to obtain negative sample pairs; forming a training data set with the positive samples and negative sample pairs, and training an initial model on the basis of the training data set in combination with a contrast loss function to obtain a semantic alignment model, where the contrast loss function is used to make the feature vectors in the positive sample pairs output by the initial model closer in the semantic space during the training process, while the feature vectors in the negative sample pairs are farther away in the semantic space.
[0051] For example, taking the data of two modalities, text and image, as an example to illustrate the training method. First, collect a sample data set containing text and image pairs, and these data pairs should describe the same content or scene. Use a pre-trained text model to extract feature vectors from the text, and use a pre-trained image model to extract feature vectors from the image; for the feature vectors of each pair of text-image data, label them as positive sample pairs (similar), and randomly select other texts or images to pair with the current text or image to form negative sample pairs (dissimilar); determine to use a contrast loss function, which encourages the feature vectors of the positive sample pairs to be closer in the space, while the feature vectors of the negative sample pairs are farther away; driven by the contrast loss, train the model to optimize the feature representation, so that similar text-image pairs are closer in the feature space. Through training, the model learns a feature representation that can map text and images to the same semantic space, thus achieving cross-modal alignment.
[0052] On the other hand, for multi-source data across institutions, the embodiments of the present application can utilize a multi-institution semantic bridging model (Multi-Entity Alignment Module) and use contrast learning and synonymous semantic clustering techniques to achieve unified alignment of coding specifications, term expressions, and behavior granularities among different financial institutions. The specific steps are as follows.
[0053] In some embodiments of the present application, the method further includes: determining the data source of the multimodal data; when the data source indicates that the text data in the multimodal data comes from different financial institutions, extracting the target keywords in the text data of different financial institutions, where the target keywords include at least one of the following: term expressions within each financial institution, specific codes, and behavior descriptions; according to the standard expression mapping table, converting the keywords corresponding to different financial institutions into unified standard vocabulary, where the standard expression mapping table is a mapping table containing the mapping from the special terms of different financial institutions to the unified standard vocabulary, and the special terms of different financial institutions include the respective data expressions, feature codes, and behavior descriptions of each financial institution.
[0054] Specifically, collect the original data from each financial institution, including coding specifications, term expressions, behavior descriptions, etc., clean the collected data to remove invalid or incorrect data to ensure data quality, and label the cleaned data, including term standardization, behavior classification, etc., for subsequent processing. Then, based on the standardized data, a term mapping table (standard expression mapping table) can be established to map the terms of different institutions into a unified term system. During this process, for polysemous words or synonyms, disambiguation can be performed through context analysis to ensure the accuracy of the terms.
[0055] After obtaining new multi-source data, natural language processing techniques or keyword extraction algorithms based on deep learning can be used to automatically identify and extract keywords from the text data, and context analysis can be performed on the extracted keywords to ensure the accuracy and relevance of the keywords; the extracted keywords can be converted into unified standard vocabulary using the term mapping table (standard expression mapping table). And as the internal terms of financial institutions are updated and the terms of the financial industry develop, the embodiments of the present application can continuously update the mapping table to maintain its timeliness and accuracy.
[0056] By constructing the standard expression mapping table, the unique expression methods of each financial institution are uniformly converted, eliminating the understanding obstacles caused by language differences, and ensuring the consistency and reliability of the construction of the knowledge graph. This method not only improves the efficiency of data processing but also enhances the ability of cross-institutional cooperation, laying a foundation for building a more comprehensive and accurate transaction monitoring system.
[0057] The embodiments of the present application can effectively perform semantic alignment on digital RMB transaction data from different sources and in different modalities, and construct a dynamic knowledge graph that comprehensively reflects transaction activities. This process not only covers structured data, but also deeply analyzes unstructured data, thereby achieving a comprehensive understanding and monitoring of digital RMB transaction events. By mapping this data to a unified semantic space, the semantic gap between heterogeneous data is eliminated, enabling data from different modalities to be compared and analyzed within the same framework, significantly improving the accuracy and efficiency of abnormal behavior detection.
[0058] After obtaining the target data with semantic alignment, target triples can be continuously extracted from the target data to automatically construct an entity relationship network (for example, automatically generate a multi-dimensional entity relationship network such as user-account-device-behavior-contract, etc.), providing dynamic input for financial knowledge modeling driven by the target knowledge graph. The specific steps are as follows.
[0059] In some embodiments of the present application, constructing an entity relationship network based on the target triples to obtain the target knowledge graph includes: using a stream processing engine and a message middleware to listen to the transaction event stream, and when the multi-modal data corresponding to the transaction event stream is updated, extracting the latest target triples from the target data corresponding to the updated multi-modal data; based on the latest target triples, performing entity change awareness and relationship evolution detection to obtain the change information corresponding to the constructed target knowledge graph, where entity change awareness and relationship evolution detection are used to determine the entity elements and / or relationship elements that have changed in the latest target triples compared to the constructed target knowledge graph; according to the change information, performing an update operation on the constructed target knowledge graph to obtain the latest target knowledge graph, where the update operation includes: node insertion, edge creation, and attribute update.
[0060] This graph construction mechanism does not rely on a static rule library, but supports continuous learning based on the semantic space, and can automatically identify new concepts, supplement new relationships, and update node weights following the evolution of the scenario, thereby forming a highly dynamic and strongly evolvable knowledge base, providing structural guarantees for upper-layer tasks such as intelligent contract traceability, risk control modeling, and compliance auditing.
[0061] Specifically, the embodiments of the present application design a knowledge graph construction and evolution mechanism triggered by an event stream to achieve real-time mapping of various changes such as transaction behaviors, account structures, device changes, and contract executions in the digital RMB ecosystem. The system combines a graph database (such as Neo4j) with a stream processing platform such as Flink to construct a second-level graph update channel, supporting version backtracking, sub-graph differential annotation, and graph expansion. At the same time, this module is combined with the policy structure parsed by the large model, converts the policy semantics into graph constraint nodes, and constructs an inferable multi-fusion knowledge graph system.
[0062] For example, in the knowledge graph construction and evolution mechanism triggered by event streams, we can first integrate a stream processing engine (such as Apache Flink) and a message middleware (such as Kafka) to achieve real-time capture and parsing of multi-source event streams from trading systems, account systems, and contract execution platforms, and perform preliminary filtering and classification based on event types, timestamps, and entity identifiers. In terms of incremental updates to the graph, a dual-track update strategy based on entity change awareness and relationship evolution detection is designed. When new entities, new relationships, or changes in existing entity attributes are identified, the system automatically completes node insertion, edge establishment, or attribute update in the graph database to ensure that the graph structure is synchronized with the financial behavior state in real time.
[0063] At the same time, a multi-granularity version control mechanism can be introduced to record graph snapshots at time windows such as minutes, hours, and days, and cooperate with a dynamic evolution algorithm for behavior labels based on node activity and evolution frequency to continuously optimize the semantic expression and risk annotation of entity nodes and relationship chains. For the policy structure parsed by the large model, the system extracts the constraint entities, conditional rules, and applicable behaviors involved in the policy through a semantic parsing engine, and automatically generates corresponding policy nodes and constraint edges, which are injected into the graph to form a "policy - entity - behavior" ternary relationship structure. Further, combined with sub-graph expansion and policy impact propagation algorithms, it excavates the entity links and behavior paths triggered by policies, constructs an inferable and traceable policy propagation chain, and realizes dynamic perception and intelligent reasoning support for regulatory policies in the financial behavior network.
[0064] Real-time monitoring and rapid response are achieved through the stream processing engine and the message middleware, ensuring the dynamic update ability of the knowledge graph. In the monitoring of digital currency transactions, this means that the system can promptly capture any changes in transaction events. Whether it is the opening of a new account, the change in transaction amount, or the update of contract status, it can be quickly reflected in the knowledge graph, providing a solid foundation for the real-time detection of abnormal behaviors. The entity change awareness and relationship evolution detection mechanism can automatically identify the update requirements of the knowledge graph, reduce manual intervention, and improve the automation level and response speed of the system.
[0065] Furthermore, in the embodiments of this application, graph neural network and large model technologies can also be integrated to establish a graph-enhanced semantic reasoning framework, which fully utilizes graph-structured data to enhance knowledge relationships and entity features, and combines risk control rules to perform intelligent risk identification on the target knowledge graph. The specific steps are as follows.
[0066] In some embodiments of the present application, the steps of identifying abnormal paths in the target knowledge graph by using a graph-enhanced inference model in combination with risk control rules are as follows: In the graph neural network of the graph-enhanced inference model, a neighbor aggregation mechanism and a relationship perception mechanism are used to analyze the feature information of entity nodes and the edges between entity nodes in the target knowledge graph, and generate node representations containing context information corresponding to each entity node. Among them, the neighbor aggregation mechanism is used to integrate the feature information of entity nodes and the neighbor nodes corresponding to the entity nodes, and the relationship perception mechanism is used to model the influence of different types of edges on information transmission. The feature information of entity nodes includes at least one of the following: account attributes, transaction behavior characteristics, contract status information. The feature information of edges includes at least one of the following: transaction amount, time interval, call relationship; According to the node representations corresponding to each entity node in the target knowledge graph, identify the behavior patterns corresponding to each node path in the target knowledge graph, where the behavior patterns include at least one of the following: account fund transfer, multi-contract call chain, multi-account association control; Use the large language model in the graph-enhanced inference model to perform logical constraint verification and conflict detection on the node paths according to the risk control rules corresponding to each behavior pattern, obtain the risk probability corresponding to each node path, and determine the node paths with risk probabilities exceeding the preset risk probability as abnormal paths.
[0067] In this embodiment, a graph-enhanced inference model can be used to identify complex high-risk behaviors including cross-chain fund transfer, multi-account association control, and smart contract abnormal paths, such as Figure 3As shown in the figure, a model inference logic that supports the three-step identification chain of "path clustering - feature attribution - risk confidence" is constructed, which is applicable to multi-dimensional, non-linear, and dynamically evolving financial risk control scenarios. Specifically, first, through entity feature encoding and relationship structure modeling, financial entity nodes such as accounts, transactions, and contracts and their associated relationships are mapped to a heterogeneous graph. Node features include account attributes, transaction behavior features, contract status information, etc., and edge features include multi-dimensional attributes such as transaction amount, time interval, and call relationship. During the graph neural network processing, neighbor aggregation and relation-aware aggregation mechanisms can be adopted to combine the features of the node itself with the semantic information of adjacent nodes and edges, and dynamically update the node representation through multi-hop context propagation, effectively capturing complex behavior patterns such as cross-account fund transfers and multi-contract call chains, thereby enhancing the perception and reasoning ability of potential risk behaviors. Then, using graph search algorithms and large language models, combined with risk control rules, potential abnormal paths are modeled, logical constraint verification and conflict detection are performed based on the graph rule engine, and rule trigger judgments based on paths and context conditions are made to achieve dynamic evaluation of risk control strategies to identify risk behaviors including "contract fund splitting to avoid supervision", "cross-contract jump to hide the controlling party", "function re-entry vulnerability triggering fund migration", etc.
[0068] In addition, the embodiments of the present application can achieve automatic update of risk control rules, and can continuously iterate and update risk identification features and patterns according to real-time data flow, user feedback, and changes in expert rules or policies, effectively adapting to the rapid evolution of new risk behaviors and ensuring the continuous and efficient risk identification ability of the system. The specific steps are as follows.
[0069] In some embodiments of the present application, the method further includes the following steps: when it is detected that the policy text related to digital RMB in the information source is updated, obtain the latest policy text; use a large language model to analyze the policy text to identify the constrained entities, as well as the restricted conditions and business logics corresponding to the constrained entities; according to the restricted conditions and business logics, generate a syntactic constraint graph structure corresponding to the constrained entities, and convert the syntactic constraint graph structure into a corresponding structured risk control rule.
[0070] For example, it is possible to conduct real-time monitoring on policy documents issued by financial regulatory agencies. After detecting policy updates through monitoring, multi-layer semantic understanding is performed on the newly issued policy documents to automatically identify the constraint entities (such as "corporate wallet", "offshore collection account") in the policy terms, the limiting conditions (such as "daily limit", "cross-border operations require permission application"), and the business logic (such as "contract functions cannot be activated before user real-name authentication"). Based on the tree structure parsing model, a syntax graph is generated, and then the graph is transformed by a rule engine for policy structured expression. The parsed policy rules can be directly transformed into control structures such as risk control rules and graph constraints in the system, and can be automatically injected according to business modules (payment system, contract engine, account system).
[0071] Through the automatic parsing of policy texts by large language models, changes in policies and regulations can be captured in a timely manner, and updated risk control rules can be automatically generated, which is of great significance for the compliant operation of the digital RMB system. This method not only reduces the workload of manually interpreting policy texts but also ensures the timeliness and accuracy of system rules, enabling quick adaptation to the new regulatory environment and avoiding security vulnerabilities caused by lagging rules. By transforming the syntax constraint graph structure into structured risk control rules, the system can adjust its monitoring strategy more flexibly and improve the recognition rate of abnormal transactions.
[0072] Meanwhile, to enhance the interpretability of risk identification results, after the graph neural network inference is completed, the system can also introduce the SHAP (SHapley Additive exPlanations) mechanism. Using the graph visualization engine and the large model language generation ability, based on the game theory feature attribution method, it quantifies the influence degree of each node feature, neighbor contribution, and path structure on the final risk determination, automatically outputs the risk path, graph structure evidence chain, and reasoning explanation text to support the output of the risk propagation chain visualization graph and the feature attribution report, ensuring that the risk decision-making process has the capabilities of being auditable, traceable, and interpretable, thereby enhancing the interpretability of the risk control decision-making process to address the problem of the lack of auditability in related technologies.
[0073] For example, the content of the structured risk control report automatically generated by the system can include but is not limited to: 1) Risk assessment level (high / medium / low risk and its confidence level); 2) Trigger path graph (contract node call chain and visible abnormal trigger path); 3) Explanation of risk factor identification (trigger function, variable value, call context, etc.); 4) Policy hit clauses (corresponding relationship entries with regulatory policies); 5) Generated natural language explanation (for auditors to read).
[0074] It should be noted that considering the characteristics of the multi-institutional participation in digital RMB (such as the central bank, commercial banks, payment institutions, and clearing platforms), the platform or system implementing the method in the embodiments of this application introduces a tenant isolation mechanism. Through Namespace partition management, policy-level access control (RBAC), API rate limiting, and resource quota policies, it ensures data security, policy independence, and service isolation among business institutions. At the same time, it supports different deployment modes (centralized, edge, and hybrid cloud deployment) to meet the regulatory and compliance requirements of different financial institutions.
[0075] The method for processing digital RMB-related data based on large model technology proposed in this application realizes the effective fusion and semantic alignment of multi-modal data through a cross-modal embedding model and a unified semantic space construction mechanism, constructs a dynamically updated knowledge graph, and then combines a graph-enhanced reasoning model with risk control rules to be able to monitor and identify abnormal behaviors in digital RMB transactions in real time. This method not only improves the accuracy and efficiency of transaction monitoring, but also can adapt to the changing financial environment and policy requirements, enhancing the security and compliance of the digital RMB system. In addition, by standardizing the special terms of different financial institutions, it promotes information sharing and understanding among institutions, further improving the collaborative efficiency of the entire financial ecosystem.
[0076] According to the embodiments of this application, an embodiment of a processing device for digital RMB-related data based on large model technology is also provided. Figure 4 It is a schematic structural diagram of a processing device for digital RMB-related data based on large model technology provided according to the embodiments of this application. As Figure 4 shown, the device includes:
[0077] A semantic alignment module 40, configured to obtain the multi-modal data involved in the transaction event stream of digital RMB, and map the multi-modal data into a unified semantic space for semantic alignment to obtain target data. Among them, in the unified semantic space, the closer the distance between the vectors corresponding to data with more similar semantics, and the farther the distance between the vectors corresponding to data with greater semantic differences;
[0078] A network construction module 42, configured to extract target triples from the target data, and construct an entity relationship network based on the target triples to obtain a target knowledge graph, where the target triples include two entity elements and one relationship element, and the target triples are used to represent the association relationship between two entities;
[0079] A risk identification module 44, which is used to adopt a graph-enhanced inference model and combine risk control rules to identify abnormal paths in the target knowledge graph. Among them, the abnormal paths correspond to abnormal handling behaviors of digital RMB. The risk control rules are used to indicate the characteristics and patterns of the rules that should be followed for the handling behaviors of digital RMB. The graph-enhanced inference model is an inference framework that combines graph neural network and large model technologies.
[0080] Optionally, the multimodal data includes: structured data and unstructured data of different modalities. Among them, the structured data includes at least one of the following: transaction records, account information, device binding records. The unstructured data includes at least one of the following: policy texts, user behavior logs, contract descriptions, image vouchers; mapping the multimodal data into a unified semantic space for semantic alignment, and the obtained target data includes: using the modality encoder in the semantic alignment model to extract the semantic features corresponding to the data of different modalities. Among them, the feature extraction methods corresponding to the data of different modalities are different, and the modalities include at least one of the following: text, speech, image, video; using the connector in the semantic alignment model to perform transformation processing on the semantic features corresponding to different modalities. Among them, the transformation processing is used to eliminate the differences in distribution, scale, and noise of the semantic features of different modalities, and obtain semantic features in a unified representation form to ensure that the data of all modalities can be represented in the same semantic space; using the generator in the semantic alignment model to map the semantic features after the transformation processing into a unified semantic space to obtain the target data.
[0081] Optionally, the training steps of the semantic alignment model include: obtaining a sample data set, where the sample data set contains multiple sample data, and each sample data contains data of different modalities for describing the same content; performing feature extraction on the data of different modalities in the sample data to obtain feature vectors corresponding to different modalities; pairing the feature vectors of different modalities corresponding to the same sample data to obtain positive sample pairs, and randomly selecting feature vectors of different modalities corresponding to different sample data for pairing to obtain negative sample pairs; forming a training data set with the positive samples and negative sample pairs, and on the basis of the training data set, combining with a contrast loss function to train an initial model to obtain a semantic alignment model. Among them, the contrast loss function is used to make the feature vectors in the positive sample pairs output by the initial model closer in the semantic space during the training process, while the feature vectors in the negative sample pairs are farther away in the semantic space.
[0082] Optionally, the semantic alignment module 40 is further configured to: determine the data source of the multimodal data; in the case that the data source represents that the text data in the multimodal data comes from different financial institutions, extract the target keywords in the text data of different financial institutions, where the target keywords include at least one of the following: term expressions within each financial institution, specific codes, and behavior descriptions; according to the standard expression mapping table, convert the keywords corresponding to different financial institutions into unified standard vocabulary, where the standard expression mapping table is a mapping table containing the mapping from the special terms of different financial institutions to the unified standard vocabulary, and the special terms of different financial institutions include the respective data expressions, feature codes, and behavior descriptions of each financial institution.
[0083] Optionally, constructing an entity relationship network based on the target triples to obtain the target knowledge graph includes: using a stream processing engine and a message middleware to listen to the transaction event stream, and in the case that the multimodal data corresponding to the transaction event stream is updated, extracting the latest target triples from the target data corresponding to the updated multimodal data; based on the latest target triples, performing entity change perception and relationship evolution detection to obtain the change information corresponding to the constructed target knowledge graph, where the entity change perception and relationship evolution detection are used to determine the changed entity elements and / or relationship elements that appear in the latest target triples compared to the constructed target knowledge graph; according to the change information, performing an update operation on the constructed target knowledge graph to obtain the latest target knowledge graph, where the update operation includes: node insertion, edge creation, and attribute update.
[0084] Optionally, the graph augmentation inference model is adopted, combined with risk control rules, to identify abnormal paths in the target knowledge graph, including: in the graph neural network of the graph augmentation inference model, the neighbor aggregation mechanism and the relationship perception mechanism are adopted to analyze the feature information of entity nodes and the edges between entity nodes in the target knowledge graph, and generate node representations containing context information corresponding to each entity node. Among them, the neighbor aggregation mechanism is used to integrate the feature information of entity nodes and their corresponding neighbor nodes, and the relationship perception mechanism is used to model the influence of different types of edges on information transmission. The feature information of entity nodes includes at least one of the following: account attributes, transaction behavior characteristics, contract status information, and the feature information of edges includes at least one of the following: transaction amount, time interval, call relationship; according to the node representations corresponding to each entity node in the target knowledge graph, identify the behavior patterns corresponding to each node path in the target knowledge graph, where the behavior patterns include at least one of the following: account fund transfer, multi-contract call chain, multi-account association control; use the large language model in the graph augmentation inference model to perform logical constraint verification and conflict detection on the node paths according to the risk control rules corresponding to each behavior pattern, obtain the risk probability corresponding to each node path, and determine the node paths with risk probability exceeding the preset risk probability as abnormal paths.
[0085] Optionally, the risk identification module 44 is further configured to: in the case of detecting an update of the policy text related to the digital currency in the information source, obtain the latest policy text; use the large language model to analyze the policy text to identify the constraint entities, as well as the restriction conditions and business logics corresponding to the constraint entities; generate a syntax constraint graph structure corresponding to the constraint entities according to the restriction conditions and business logics, and convert the syntax constraint graph structure into a corresponding structured risk control rule.
[0086] It should be noted that each module in the above digital currency-related data processing device based on the large model technology can be a program module (for example, a set of program instructions for implementing a specific function), or a hardware module. For the latter, it can be presented in the following forms, but not limited to: the manifestation form of each of the above modules is a processor, or the functions of each of the above modules are implemented by a processor.
[0087] It should be noted that the digital currency-related data processing device based on the large model technology provided in this embodiment can be used to execute Figure 2 the digital currency-related data processing method based on the large model technology shown, therefore, the relevant explanations of the above digital currency-related data processing method based on the large model technology also apply to the embodiments of this application, and will not be repeated here.
[0088] The embodiments of the present application also provide a non-volatile storage medium. The non-volatile storage medium includes a stored computer program. Wherein, the device where the non-volatile storage medium is located executes the following method for processing digital RMB-related data based on large model technology by running the computer program: obtaining multi-modal data involved in the transaction event stream of digital RMB, and mapping the multi-modal data into a unified semantic space for semantic alignment to obtain target data. Wherein, in the unified semantic space, the closer the distance between the vectors corresponding to the data with more similar semantics, and the farther the distance between the vectors corresponding to the data with greater semantic differences; extracting target triples from the target data, and constructing an entity relationship network based on the target triples to obtain a target knowledge graph. Wherein, the target triple contains two entity elements and one relationship element, and the target triple is used to represent the association relationship between two entities; using a graph-enhanced reasoning model, combined with risk control rules, to identify abnormal paths in the target knowledge graph. Wherein, the abnormal path corresponds to the abnormal processing behavior of digital RMB, and the risk control rules are used to indicate the characteristics and patterns of the rules that should be followed for the processing behavior of digital RMB. The graph-enhanced reasoning model is a reasoning framework that combines graph neural network and large model technology.
[0089] The embodiments of the present application also provide a computer program product, including a computer program. When the computer program is executed by a processor, it implements the steps of the method for processing digital RMB-related data based on large model technology described in each embodiment of the present application: obtaining multi-modal data involved in the transaction event stream of digital RMB, and mapping the multi-modal data into a unified semantic space for semantic alignment to obtain target data. Wherein, in the unified semantic space, the closer the distance between the vectors corresponding to the data with more similar semantics, and the farther the distance between the vectors corresponding to the data with greater semantic differences; extracting target triples from the target data, and constructing an entity relationship network based on the target triples to obtain a target knowledge graph. Wherein, the target triple contains two entity elements and one relationship element, and the target triple is used to represent the association relationship between two entities; using a graph-enhanced reasoning model, combined with risk control rules, to identify abnormal paths in the target knowledge graph. Wherein, the abnormal path corresponds to the abnormal processing behavior of digital RMB, and the risk control rules are used to indicate the characteristics and patterns of the rules that should be followed for the processing behavior of digital RMB. The graph-enhanced reasoning model is a reasoning framework that combines graph neural network and large model technology.
[0090] The serial numbers of the above embodiments of the present application are only for description and do not represent the advantages and disadvantages of the embodiments.
[0091] In the above embodiments of the present application, the descriptions of each embodiment have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0092] In several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of the units can be a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection shown or discussed with each other can be through some interfaces. The indirect coupling or communication connection of units or modules can be in an electrical or other form.
[0093] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0094] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0095] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks, or optical discs and other various media that can store program codes.
[0096] The above is only the preferred embodiment of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.
Claims
1. A method for processing digital RMB-related data based on large model technology, characterized in that Including: Obtain multimodal data involved in the transaction event stream of digital RMB, and map the multimodal data to a unified semantic space for semantic alignment to obtain target data. In the unified semantic space, the closer the distance between the vectors corresponding to data with more similar semantics, and the farther the distance between the vectors corresponding to data with greater semantic differences; Extract target triples from the target data, and construct an entity relationship network based on the target triples to obtain a target knowledge graph. The target triples include two entity elements and one relationship element, and the target triples are used to represent the association relationship between two entities; Adopt a graph-enhanced reasoning model, combined with risk control rules, to identify abnormal paths in the target knowledge graph. The abnormal paths correspond to abnormal handling behaviors of digital RMB, and the risk control rules are used to indicate the characteristics and patterns of the rules that should be followed for the handling behaviors of digital RMB. The graph-enhanced reasoning model is an inference framework combining graph neural network and large model technology.
2. The method for processing digital RMB-related data based on large model technology according to claim 1, wherein The multimodal data includes: structured data and unstructured data of different modalities. The structured data includes at least one of the following: transaction records, account information, device binding records. The unstructured data includes at least one of the following: policy texts, user behavior logs, contract descriptions, image vouchers. Mapping the multimodal data to a unified semantic space for semantic alignment to obtain target data includes: Adopt a modality encoder in the semantic alignment model to extract semantic features corresponding to data of different modalities. The ways of feature extraction corresponding to data of different modalities are different. The modalities include at least one of the following: text, speech, image, video; Adopt a connector in the semantic alignment model to perform transformation processing on the semantic features corresponding to different modalities. The transformation processing is used to eliminate the differences in distribution, scale, and noise of the semantic features of different modalities to obtain the semantic features in a unified representation form, so as to ensure that data of all modalities can be represented in the same semantic space; Adopt a generator in the semantic alignment model to map the semantic features after the transformation processing to the unified semantic space to obtain the target data.
3. The method for processing digital RMB-related data based on large model technology according to claim 2, wherein, The training steps of the semantic alignment model include: Obtain a sample data set, where the sample data set contains multiple sample data, and each sample data contains data of different modalities for describing the same content; Extract features of data of different modalities in the sample data to obtain feature vectors corresponding to different modalities; Pair the feature vectors of different modalities corresponding to the same sample data to obtain positive sample pairs, and randomly select feature vectors of different modalities corresponding to different sample data for pairing to obtain negative sample pairs; The positive samples and the negative sample pairs are formed into a training data set, and based on the training data set, combined with a contrastive loss function, an initial model is trained to obtain the semantic alignment model, where the contrastive loss function is used to make the feature vectors in the positive sample pairs closer in the semantic space and the feature vectors in the negative sample pairs farther away in the semantic space in the process of training.
4. The method for processing digital RMB-related data based on large model technology according to claim 2, wherein, The method further includes: Determining the data source of the multimodal data; When the data source represents that the text data in the multimodal data comes from different financial institutions, extracting target keywords in the text data of different financial institutions, where the target keywords include at least one of the following: term expressions within each financial institution, specific encodings, and behavior descriptions; According to a standard expression mapping table, converting the keywords corresponding to different financial institutions into unified standard vocabulary, where the standard expression mapping table is a mapping table containing the mapping from the special terms of different financial institutions to unified standard vocabulary, and the special terms of different financial institutions include the respective data expressions, feature encodings, and behavior descriptions of each financial institution.
5. The method for processing digital RMB-related data based on large model technology according to claim 1, wherein Constructing an entity relationship network based on the target triples to obtain a target knowledge graph, including: Using a stream processing engine and a message middleware to monitor the transaction event stream, and when the multimodal data corresponding to the transaction event stream is updated, extracting the latest target triples from the target data corresponding to the updated multimodal data; Based on the latest target triples, performing entity change perception and relationship evolution detection to obtain change information corresponding to the constructed target knowledge graph, where the entity change perception and relationship evolution detection are used to determine the entity elements and / or relationship elements that have changed in the latest target triples compared to the constructed target knowledge graph; According to the change information, performing an update operation on the constructed target knowledge graph to obtain the latest target knowledge graph, where the update operation includes: node insertion, edge creation, and attribute update.
6. The method for processing digital RMB-related data based on large model technology according to claim 1, wherein Using a graph-enhanced inference model and combining risk control rules to identify abnormal paths in the target knowledge graph, including: In the graph neural network of the graph-enhanced inference model, using a neighbor aggregation mechanism and a relationship perception mechanism to analyze the feature information of entity nodes and the edges between entity nodes in the target knowledge graph, generating node representations containing context information corresponding to each entity node, where the neighbor aggregation mechanism is used to integrate the feature information of entity nodes and the neighbor nodes corresponding to the entity nodes, the relationship perception mechanism is used to model the influence of different types of edges on information transmission, the feature information of the entity nodes includes at least one of the following: account attributes, transaction behavior characteristics, and contract status information, and the feature information of the edges includes at least one of the following: transaction amount, time interval, and call relationship; Based on the node representations corresponding to each entity node in the target knowledge graph, identify the behavior patterns corresponding to each node path in the target knowledge graph, where the behavior patterns include at least one of the following: account fund jump, multi-contract call chain, multi-account association control; Use the large language model in the graph-enhanced reasoning model to perform logical constraint verification and conflict detection on the node paths according to the risk control rules corresponding to each behavior pattern, obtain the risk probabilities corresponding to each node path, and determine the node paths with risk probabilities exceeding the preset risk probability as the abnormal paths.
7. The method for processing digital RMB-related data based on large model technology according to claim 6, wherein The method further includes: When it is detected that the policy text related to digital currency in the information source is updated, obtain the latest policy text; Use the large language model to analyze the policy text to identify the constrained entities in the policy text, as well as the limiting conditions and business logics corresponding to the constrained entities; According to the limiting conditions and the business logics, generate a syntax constraint graph structure corresponding to the constrained entities, and convert the syntax constraint graph structure into the corresponding structured risk control rules.
8. A processing device for digital RMB-related data based on large model technology, characterized in that, It includes: A semantic alignment module, configured to obtain the multi-modal data involved in the transaction event stream of digital currency, and map the multi-modal data to a unified semantic space for semantic alignment to obtain target data, where in the unified semantic space, the closer the distance between the vectors corresponding to data with more similar semantics, and the farther the distance between the vectors corresponding to data with greater semantic differences; A network construction module, configured to extract target triples from the target data and construct an entity relationship network based on the target triples to obtain a target knowledge graph, where the target triples contain two entity elements and one relationship element, and the target triples are used to represent the association relationship between two entities; A risk identification module, configured to use a graph-enhanced reasoning model and combine risk control rules to identify abnormal paths in the target knowledge graph, where the abnormal paths correspond to abnormal handling behaviors of digital currency, the risk control rules are used to indicate the characteristics and patterns of the rules that should be followed for the handling behaviors of digital currency, and the graph-enhanced reasoning model is an inference framework combining graph neural network and large model technology.
9. An electronic device, characterized in that, It includes: A memory and a processor, where the processor is configured to run the program stored in the memory, and when the program runs, it executes the method for processing digital currency-related data based on large model technology according to any one of claims 1 to 7.
10. A non-volatile storage medium, characterized in that, The non-volatile storage medium includes a stored computer program, and the device where the non-volatile storage medium is located executes the method for processing digital currency-related data based on large model technology according to any one of claims 1 to 7 by running the computer program.
Citation Information
Patent Citations
Intelligent search method and system based on multi-source heterogeneous data
CN116049454A
Data asset identification and risk early warning system and method based on financial knowledge graph and large language model
CN118247057A
Financial knowledge graph reasoning method and system based on multilayer path semantic modeling
CN119294495A
Digital base fusion system and electronic equipment
CN119760007A
Cross-modal knowledge reasoning method and device for industrial quality inspection and medium
CN120069096A
Cited By
Enterprise customer risk control management method and system based on AI intelligence
CN120634279A
Data weaving relation reasoning method and system based on graph neural network
CN121146049A
Abnormal transaction behavior analysis method and device, equipment and storage medium
CN121961725A