Cross-modal transaction feature self-generation and anomaly detection system and method fused with large language model

By integrating the cross-modal transaction feature autogeneration and anomaly detection system with large language models, the problem of insufficient utilization of unstructured data in traditional methods is solved, comprehensive and accurate identification and timely warning of financial risks are achieved, adapting to market changes, and improving the efficiency and accuracy of financial transaction monitoring.

CN120258990APending Publication Date: 2025-07-04SHENZHEN YSSTECH INFORMATION TECH CO LTD

Patent Information

Application Number
CN202510733677.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

Traditional financial risk assessment and transaction monitoring methods are difficult to effectively utilize unstructured data, and cannot keep up with market changes in time, resulting in system performance degradation and unable to meet actual business needs.

Method used

A cross-modal transaction feature autogeneration and anomaly detection system that integrates large language models is adopted. Multi-source data is collected through the data integration module. The cross-modal feature extraction module uses the large language model to extract cross-modal transaction features, combines the abnormal transaction feature to identify abnormal transaction features, and automatically updates through the feature auto-update module to form a closed-loop feedback mechanism.

Benefits of technology

It realizes the effective use of unstructured data, improves the comprehensiveness and accuracy of risk identification, has good adaptability and timeliness, can respond to market changes and new risk characteristics in a timely manner, and provides strong risk warning and monitoring methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120258990A_ABST
    Figure CN120258990A_ABST
Patent Text Reader

Abstract

The invention discloses a cross-modal transaction feature self-generation and anomaly detection system and method fused with a large language model, and the system comprises a data integration module which is used for connecting a multi-source data platform and collecting structured and unstructured data; the cross-modal feature extraction module is used for processing the collected data by using a large language model and extracting cross-modal transaction features; the anomaly detection module is used for analyzing the cross-modal transaction feature vector and identifying an abnormal transaction feature; the feature self-updating module is used for automatically updating cross-modal transaction features; and the feedback receiving module is used for receiving the anomaly detection result from the anomaly detection module and feedback information of an external user or system, and sending the feedback information to the feature self-updating module to form a closed-loop feedback mechanism. Automatic generation and updating of cross-modal transaction features are realized by utilizing semantic understanding and feature extraction capabilities of a large language model, and abnormal transaction features are found in time in combination with an anomaly detection algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technology in the field of financial risk assessment and transaction monitoring, and in particular to a cross-modal transaction feature self-generation and anomaly detection system and method integrating a large language model. Background Art

[0002] With the continuous development and innovation of the financial market, transaction data presents the characteristics of multi-source, heterogeneous, and massive, covering structured data (such as transaction records, financial statements) and unstructured data (such as news reports, social media content). Traditional risk assessment and transaction monitoring methods mainly rely on the analysis of structured data and are difficult to effectively utilize the rich information contained in unstructured data, resulting in certain limitations in the perception and identification of market risks.

[0003] At the same time, the financial market environment is complex and changeable, and risk characteristics are also constantly evolving, which requires risk assessment and transaction monitoring systems to have the ability to automatically update and adapt to new risk characteristics. However, most existing systems rely on manual setting and adjustment of features and cannot keep up with the rhythm of market changes in a timely manner, resulting in a decline in system performance and difficulty in meeting actual business needs. Summary of the Invention

[0004] In view of this, in view of the deficiencies of the existing technology, the main purpose of the present invention is to provide a cross-modal transaction feature self-generation and anomaly detection system and method integrating a large language model, which realizes the automatic generation and update of cross-modal transaction features by using the powerful semantic understanding and feature extraction capabilities of the large language model, and combines anomaly detection algorithms to timely discover abnormal transaction features.

[0005] To achieve the above object, the present invention adopts the following technical solutions: A cross-modal transaction feature self-generation and anomaly detection system integrating a large language model, comprising: A data integration module for docking with multi-source data platforms and collecting structured and unstructured data; a cross-modal feature extraction module for processing the collected data by using a large language model and extracting cross-modal transaction features; An anomaly detection module for analyzing cross-modal transaction feature vectors based on a preset anomaly detection algorithm to identify abnormal transaction features; A feature self-update module for automatically updating cross-modal transaction features according to system feedback and data changes; A feedback receiving module for receiving anomaly detection results from the anomaly detection module and feedback information from external users or systems, and sending the feedback information to the feature self-update module to form a closed-loop feedback mechanism; The data integration module, cross-modal feature extraction module, anomaly detection module, and feature self-update module are connected in sequence, and the feedback receiving module is connected to the anomaly detection module and the feature self-update module.

[0006] As a preferred solution: The data integration module supports docking with financial data platforms such as Caihui, DM, Wind, and YY, flexibly expands the data ecosystem according to requirements, and realizes data collection by using standardized interface technology.

[0007] As a preferred solution: The cross-modal feature extraction module specifically is: inputting unstructured data into a pre-trained large language model, outputting a semantic vector representation after being processed by the model encoding layer, and then fusing the semantic vector with the structured data features through a feature fusion sub-module to generate a cross-modal transaction feature vector.

[0008] As a preferred solution: The feature fusion sub-module uses the following formula for feature fusion: ; where F represents the cross-modal transaction feature vector, F text represents the semantic vector output by the large language model, F struct represents the structured data feature vector, and α and β are respectively the weight coefficients of the semantic feature and the structured feature, used to balance the contribution degrees of the two features in the fusion process, and their value ranges are between [0,1], and α + β = 1.

[0009] As a preferred solution: The feature self-update module includes a feature evaluation sub-module and a feature optimization sub-module. The feature evaluation sub-module is used to evaluate the effectiveness and relevance of existing features, and the feature optimization sub-module optimizes and adjusts the features according to the evaluation results.

[0010] As a preferred solution: The anomaly detection algorithm adopts one or a combination of multiple methods based on statistics, machine learning, or deep learning, and is used to flexibly detect abnormal transaction features according to different types of transaction data and business scenarios.

[0011] As a preferred solution: When the anomaly detection module adopts a method based on statistics, the following formula is used to calculate the anomaly score of the data: ; where S anomaly represents the anomaly score, x represents the current cross-modal transaction feature vector, μ represents the mean of historical data, σ represents the standard deviation of historical data, and by comparing the anomaly score with a preset threshold, it is judged whether the transaction feature is abnormal.

[0012] As a preferred solution: The following formula is used by the feature evaluation sub-module to evaluate the effectiveness of existing features: ; where E represents feature effectiveness, TP represents true positive, TN represents true negative, FP represents false positive, and FN represents false negative. The performance of features in anomaly detection is evaluated by calculating feature effectiveness, providing a basis for feature optimization.

[0013] A detection method applied to the system includes the following steps: S1. Connect to multi-source data platforms through a data integration module to collect structured and unstructured data; S2. Use a large language model to process the collected data and extract cross-modal transaction features; S3. Analyze the cross-modal transaction feature vectors based on a preset anomaly detection algorithm to identify abnormal transaction features; S4. Receive the anomaly detection results and external feedback information through a feedback receiving module and send them to the feature self-update module; S5. According to the received feedback information, the feature self-update module automatically updates the cross-modal transaction features to achieve the automatic generation, anomaly detection, and automatic update and optimization of cross-modal transaction features.

[0014] As a preferred solution: In step S2, the specific process of extracting cross-modal transaction features is as follows: Input the unstructured data into a pre-trained large language model, and after being processed by the model encoding layer, output the semantic vector representation. Then, through the feature fusion sub-module, fuse the semantic vector with the structured data features to generate a cross-modal transaction feature vector.

[0015] Compared with the prior art, the present invention has obvious advantages and beneficial effects. Specifically, as can be seen from the above technical solutions, First, the use of a large language model realizes the automatic generation of cross-modal transaction features, fully excavates the potential risk information in unstructured data, enriches the feature dimensions of risk assessment, and improves the comprehensiveness and accuracy of risk identification.

[0016] Second, the feature self-update module can automatically optimize and update transaction features according to system feedback and data changes, and continuously adjust and optimize through a closed-loop feedback mechanism, enabling the system to have good adaptability and timeliness, and being able to respond to changes in the market environment and the emergence of new risk features in a timely manner.

[0017] Third, the combination of the anomaly detection algorithm and cross-modal transaction features realizes the rapid and accurate identification of abnormal transaction behaviors, provides a powerful means of risk warning and monitoring for financial institutions, and helps to reduce transaction risks and losses.

[0018] To more clearly elaborate the structural features and functions of the present invention, the following will be described in detail in combination with the drawings and specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 Schematic diagram of the detection system architecture of the present invention; Figure 2 Schematic diagram of the detection method flow of the present invention. Detailed implementation manners

[0020] The present invention is as Figure 1 and Figure 2 shown, a cross-modal transaction feature self-generation and anomaly detection system and method integrating large language models. The system includes a data integration module, a cross-modal feature extraction module, an anomaly detection module, a feature self-update module, and a feedback receiving module, where: The data integration module supports docking mainstream financial data platforms such as Caihui, DM, Wind, and YY, and uses standardized interface technology to implement data collection. The collected data includes structured data (such as transaction records, financial statements, etc.) and unstructured data (such as news reports, social media content). For example, in stock trading monitoring, this module can collect multi-source data such as stock trading records, company financial statements, and relevant news reports in real time. It can be flexibly adapted according to the different formats and protocols of data sources to ensure stable data transmission and efficient collection. During the collection process, this module also performs preliminary cleaning and preprocessing on the data to remove obvious noise and incorrect data to improve the efficiency and quality of subsequent processing.

[0021] The core of the cross-modal feature extraction module is a pre-trained large language model. The unstructured data is input into the large language model, and after being processed by the model encoding layer, a semantic vector representation is output. Then, through the feature fusion sub-module, the semantic vector is fused with the structured data features to generate a cross-modal transaction feature vector. Taking news reports as an example, the large language model can extract the semantic vector of the news text and then fuse it with the structured data features (such as transaction price, trading volume, etc.) in the stock trading record to form a cross-modal transaction feature vector containing text semantic information and transaction data features.

[0022] The feature fusion sub-module uses the following formula for feature fusion: ; where F represents the cross-modal transaction feature vector, F text represents the semantic vector output by the large language model, F struct represents the structured data feature vector, and α and β are the weight coefficients of the semantic feature and the structured feature respectively, which are used to balance the contribution degrees of the two features in the fusion process. Their value ranges are between [0,1], and α + β = 1. In practical applications, the weight coefficients can be adjusted according to data characteristics and business requirements to achieve the best fusion effect. Specifically, the weight coefficients α and β are dynamically adjusted according to the data information entropy, and the formula is: ; Among them, H text represents the text information entropy, which measures the information density or uncertainty of unstructured text data and reflects the richness of key information in the text.

[0023] ; Text preprocessing: Remove stop words (such as "de", "shi"), special symbols, and H TML tags. Extract financial entities (such as company names, stock codes) and keywords (such as "financial report", "merger and acquisition").

[0024] Word frequency statistics: Statistically analyze the frequency distribution of the remaining words ; Calculate the entropy value according to the word frequency distribution. The higher the entropy value, the more dispersed the text information; the lower the entropy value, the more concentrated the information.

[0025] For example, news text: "Company X releases its financial report, with a 20% increase in net profit" → Extract keywords "financial report", "net profit", and calculate the entropy value. If the word frequency distribution is uniform (such as multiple keywords appearing evenly), H text it is higher; if it is concentrated in a few keywords (such as "sharp drop" repeating), H text it is lower.

[0026] H struct represents the structured information entropy, which measures the volatility or complexity of structured data (such as transaction records) and reflects the abnormality degree of quantitative features.

[0027] ; Data selection: Extract key indicators of structured data, such as price (price), trading volume (volume).

[0028] Volatility calculation: Price volatility: Calculate the standard deviation σ of the price price .

[0029] Volume volatility: Calculate the standard deviation σ of the trading volume volume .

[0030] Entropy value synthesis: Add the two standard deviations as the comprehensive entropy value of the structured data.

[0031] For example, the standard deviation σ of a certain stock price price = 15%, and the standard deviation σ of the trading volume volume = 30% → H struct = 45%; The higher the entropy value, the greater the volatility of the structured data, which may indicate abnormal trading behavior.

[0032] If the text information entropy H text is high (such as dense breaking news), increase the weight of text features (α↑).

[0033] If the structured information entropy H struct is high (such as sharp price fluctuations), increase the weight of structured features (β↑).

[0034] The anomaly detection module analyzes the cross-modal transaction feature vectors based on a preset anomaly detection algorithm to identify abnormal transaction features. The anomaly detection algorithm can adopt one or a combination of multiple methods including statistical-based methods, machine learning-based methods, or deep learning-based methods. For example, in the statistical-based method, the anomaly score of the data is calculated using the following formula: ; where S anomaly represents the anomaly score, x represents the current cross-modal transaction feature vector, μ represents the mean of historical data, σ represents the standard deviation of historical data, and by comparing the anomaly score with a preset threshold, it is determined whether the transaction feature is abnormal.

[0035] The anomaly threshold T is dynamically adjusted based on market volatility, and the formula is T = μ + k.σ, where k is optimized in real time according to the risk tolerance (conservative strategy k = 3, aggressive strategy k = 2), and the market volatility, trading volume and other indicators are monitored in real time through reinforcement learning to dynamically adjust the value of k.

[0036] In the machine learning or deep learning-based methods, a trained classification model or clustering model can be used to identify abnormal transaction features. For example, a support vector machine (SVM) is used to classify normal and abnormal transaction features, or a deep neural network (DNN) is used to automatically learn complex patterns and abnormal features in the data. The anomaly detection module also sets different warning levels according to business requirements and risk levels. When abnormal transaction features are detected, warning signals are sent in a timely manner, and the anomaly detection results are sent to the feedback receiving module.

[0037] The feature self-update module includes a feature evaluation sub-module and a feature optimization sub-module. The feature evaluation sub-module is used to evaluate the effectiveness and relevance of existing features, and the following formula is used for evaluation: ; where E represents feature effectiveness, TP represents true positives, TN represents true negatives, FP represents false positives, and FN represents false negatives.

[0038] E min = 0.7, when E < E min trigger weight adjustment, feature replacement or combination optimization strategies. For example, in the stock trading scenario, if the news sentiment score feature decreases in effectiveness due to noise data, the system automatically reduces its weight and introduces the change in analyst ratings as a new feature.

[0039] The performance of features in anomaly detection is evaluated by calculating feature effectiveness, providing a basis for feature optimization. The feature optimization sub-module optimizes and adjusts features according to the evaluation results. For example, it improves the quality and effectiveness of features by adjusting feature weights, adding or deleting features, etc. In addition, this module dynamically adjusts the feature update strategy according to market changes and business requirements, such as adding new features and optimizing feature combinations, to ensure that the system can continuously adapt to new risk features and market environments. Additionally, the feature evaluation cycle is dynamically adjusted according to business requirements. Real-time monitoring is adopted in high-frequency trading scenarios, and evaluation is conducted on a daily or weekly basis in conventional scenarios.

[0040] The feedback receiving module is used to receive the anomaly detection results from the anomaly detection module and the feedback information from external users or systems, and send the feedback information to the feature self-update module, forming a closed-loop feedback mechanism. For example, when the anomaly detection module detects an anomaly in a certain transaction, it sends the anomaly detection result to the feedback receiving module. At the same time, external users or systems can also confirm or correct the detection result according to the actual situation and send the feedback information to the feedback receiving module. The feedback receiving module integrates this feedback information and sends it to the feature self-update module so that the feature self-update module can optimize and update the cross-modal transaction features according to the feedback information. The closed-loop feedback mechanism enables the system to continuously learn and improve, enhancing the accuracy and timeliness of anomaly detection.

[0041] In addition, the system has a real-time processing framework. The real-time processing framework uses Apache Flink or Kafka Streams to implement stream computing, ensuring that the latency of feature extraction and anomaly detection is less than 100 ms. Model lightweighting is achieved through knowledge distillation (such as TinyBERT) to reduce resource consumption. H The detection method of the cross-modal transaction feature self-generation and anomaly detection system applied to the fusion large language model includes the following steps:

[0042] S1. Connect to multi-source data platforms through the data integration module to collect structured and unstructured data. For example, in financial transaction monitoring, multi-source data such as stock trading records, company financial statements, and news reports are collected. The data integration module fetches data from each data source regularly or in real time according to the preset collection strategy and frequency, and conducts preliminary sorting and preprocessing to ensure the integrity and consistency of the data. S1. Connect to multi-source data platforms through the data integration module to collect structured and unstructured data. For example, in financial transaction monitoring, multi-source data such as stock trading records, company financial statements, and news reports are collected. The data integration module fetches data from each data source regularly or in real time according to the preset collection strategy and frequency, and conducts preliminary sorting and preprocessing to ensure the integrity and consistency of the data.

[0043] S2. Use a large language model to process the collected data and extract cross-modal transaction features. Specifically, input unstructured data into a pre-trained large language model. After being processed by the model's encoding layer, semantic vector representations are output. Then, through a feature fusion sub-module, the semantic vectors are fused with structured data features to generate cross-modal transaction feature vectors. For example, perform semantic analysis on news report texts, extract semantic features related to stock trading, and fuse them with the structured data features in stock trading records to form cross-modal transaction feature vectors. In this process, the large language model can deeply understand the semantic information of the text, capture potential factors in the news report that may affect stock trading, such as company performance and industry dynamics, thus providing richer feature information for anomaly detection.

[0044] S3. Analyze the cross-modal transaction feature vectors based on a preset anomaly detection algorithm to identify abnormal transaction features. For example, use a statistics-based method to calculate anomaly scores, or use a machine learning-based classification model to classify transaction features to identify abnormal transaction features. The anomaly detection algorithm sets thresholds according to historical data and business rules. When the anomaly score of a transaction feature exceeds the threshold, it is determined that the transaction is abnormal and a warning mechanism is triggered. At the same time, the anomaly detection module records the anomaly detection results, including the feature information of abnormal transactions, anomaly scores, warning levels, etc., for subsequent analysis and processing.

[0045] S4. Receive anomaly detection results and external feedback information through a feedback receiving module. For example, receive anomaly transaction alerts from the anomaly detection module and confirmation or correction information of the alerts from external users. The feedback receiving module classifies and organizes the received feedback information, extracts valuable information, such as the actual situation of abnormal transactions, the permissions and roles of feedback users, etc., to ensure the accuracy and reliability of the feedback information.

[0046] S5. According to the received feedback information, the feature self-update module automatically updates the cross-modal transaction features. This includes evaluating the effectiveness and relevance of existing features, and optimizing and adjusting the features according to the evaluation results to achieve the automatic update and optimization of cross-modal transaction features. For example, adjust feature weights and optimize feature combinations according to feedback information to improve the accuracy and timeliness of anomaly detection. The feature self-update module regularly or real-time evaluates and updates the features to ensure that the system can promptly adapt to market changes and the emergence of new risk features, and maintain good performance and adaptability.

[0047] The core goal of this system is to achieve the automatic generation of cross-modal transaction features, anomaly detection, and automatic update and optimization of features. Its overall operation process is as follows: Data Acquisition Phase: The data integration module serves as the data entry point for the entire system and is responsible for connecting to multi-source data platforms, including but not limited to mainstream financial data platforms such as Caihui, DM, Wind, and YY. It uses standardized interface technology to flexibly adapt to the formats and protocols of different data sources and collects structured data (such as transaction records, financial statements) and unstructured data (such as news reports, social media content) in real-time or at regular intervals. During the collection process, this module will perform preliminary cleaning and preprocessing on the data to remove noise and incorrect data to ensure data quality. And for unstructured data, it is necessary to remove H TML tags, special characters, and stop words (such as "de", "shi"), and extract financial entities (such as company names, stock codes) through regular expressions. Missing values in structured data are filled using time series interpolation or industry averages. Sensitive data (such as trading accounts) needs to be anonymized, and data transmission is encrypted using the TLS protocol to ensure security.

[0048] Feature Extraction Phase: The collected data is transmitted to the cross-modal feature extraction module. Here, the pre-trained large language model plays a key role in converting unstructured data (such as news text) into semantic vector representations. This process realizes in-depth semantic understanding with the help of the encoding layer of the large language model and mines potential risk information in the text. Subsequently, the semantic vectors are fused with the structured data features to generate comprehensive cross-modal transaction feature vectors, providing rich and comprehensive feature inputs for subsequent anomaly detection.

[0049] Anomaly Detection Phase: The cross-modal transaction feature vectors flow into the anomaly detection module, which analyzes based on preset anomaly detection algorithms (covering methods based on statistics, machine learning, or deep learning). Taking the statistical-based method as an example, an anomaly score is calculated and compared with the preset threshold to accurately identify abnormal transaction features. Once an anomaly is detected, the module immediately generates a warning signal, sends the anomaly detection results to the feedback receiving module, and records detailed anomaly information for subsequent analysis and traceability.

[0050] Feedback and Feature Update Phase: The feedback receiving module receives warning information from the anomaly detection module and feedback from external users or systems (such as confirmation and correction of warnings). These feedback messages are the optimization basis for the feature self-update module. The feature evaluation sub-module in the feature self-update module evaluates the effectiveness of existing features, and the feature optimization sub-module adjusts the feature weights and filters features accordingly to achieve iterative feature updates, ensuring that the system feature set always conforms to the market changes and the evolution law of risk characteristics, forming a closed-loop optimization.

[0051] Iterative Optimization Phase: The updated features are re-integrated into the data processing process to guide subsequent cross-modal feature extraction and anomaly detection work, continuously improving the system performance. The system accumulates experience during operation, adapts to the complex and dynamic changes in the financial market, and strengthens the risk identification and warning effectiveness.

[0052] Examples of actual application scenarios: Data collection: The data integration module collects trading records on the stock trading platform in real time (structured data, including trading time, price, quantity, etc.), and synchronously collects financial news and discussion content on social media stock forums (unstructured data). It interfaces with the stock trading platform through the API, reads and preliminarily organizes data according to the platform's data format specifications; for news websites and social media, it uses web crawler technology to capture text data at a set frequency and performs simple text cleaning and filtering.

[0053] Feature extraction and fusion: After receiving the data, the cross-modal feature extraction module inputs news texts and discussion content on stock forums into a pre-trained large language model (such as BERT). The model's encoding layer performs in-depth semantic analysis on the texts and outputs semantic vectors reflecting market information. At the same time, structured features such as price fluctuations and trading volume changes in stock trading records are extracted, and are weighted and fused by the feature fusion sub-module to generate a comprehensive cross-modal trading feature vector, which contains both semantic information and quantitative indicators of trading data.

[0054] The pre-trained large language model (such as BERT) needs to be fine-tuned through a financial domain corpus. The fine-tuning process includes: Domain corpus construction: Collect financial reports, research reports, and news to construct a corpus of millions of tokens; Task design: Masked Language Modeling (MLM), for predicting and training financial terms (such as "EBITDA", "quantitative easing"); Sentiment classification task, annotating the sentiment polarity of texts to improve the ability to capture market sentiment; -5 , batch size 32, train for 3 - 5 rounds, freeze the underlying parameters and only fine-tune the top encoding layer.

[0055] Abnormal trading identification: The anomaly detection module uses a statistics-based anomaly detection algorithm to calculate the anomaly score of the current cross-modal trading feature vector in real time based on the mean and standard deviation calculated from historical trading data. If the anomaly score exceeds the preset threshold, it is determined as an abnormal trading, such as a sudden large-scale sell order. The system quickly triggers an alarm, indicates a possible market manipulation behavior, and sends the alarm details to the feedback receiving module.

[0056] Feedback and Feature Optimization: The feedback receiving module receives warnings from the anomaly detection module and simultaneously receives feedback from traders on the accuracy of the warnings (such as false alarm and missed alarm marks). Based on this, the feature self-update module initiates an evaluation and optimization process: the feature evaluation sub-module quantitatively analyzes the effectiveness of existing features. If it is found that the original feature combination fails to capture sufficient information and leads to missed alarms, the feature optimization sub-module adjusts the weights of semantic features and structured features, increases the feature dimensions, and the optimized features are put back into operation to improve the accuracy of subsequent anomaly detection.

[0057] Continuous Optimization: The system continuously processes newly incoming data and accumulates trading cases. The feature self-update module regularly reviews the effectiveness of features, and in combination with changes in market styles (such as switching from value investment to growth investment), dynamically introduces new features (such as analyst rating adjustments, keywords for industry technological breakthroughs, etc.), and eliminates redundant and outdated features to ensure that the system always operates efficiently with the optimal feature set, providing a strong technical support for stock trading monitoring.

[0058] The overall operation principle of this system embodies the closed-loop logic of "data-driven - feature extraction - anomaly detection - feedback optimization". Each module collaborates closely, with the large language model as the technical core, to achieve automated processing and intelligent optimization of cross-modal trading features, precisely serving the needs of financial risk monitoring, efficiently coping with the complex and ever-changing financial market environment, and providing strong guarantees for financial institutions' risk prevention and control and decision-making support.

[0059] Taking insider trading detection in high-frequency stock trading as an example: Data Collection and Feature Fusion: The data integration module is connected to the NASDAQ trading platform and Bloomberg news interface in real time, collecting trading records of a certain technology stock (structured data: price volatility of 25%, volume surge of 400%) and social media discussions (unstructured data: "insiders sold in advance"). The unstructured data is processed by the BERT model fine-tuned in the financial field to generate semantic vectors (negative sentiment, confidence level 0.91), and at the same time, the price standard deviation (σ price = 18%) and volume standard deviation (σ volume = 35%) of the structured data are extracted, and the structured information entropy H struct = 53% is calculated.

[0060] According to the dynamic weight formula: , the text information entropy H text = 0.6, and the weight is adjusted to α = 0.7 to generate a cross-modal feature vector.

[0061] Anomaly Detection and Real-time Warning: The anomaly detection module adopts a statistical threshold strategy to calculate the anomaly score S of the current feature vector anomaly= |30 - 10| / 5 = 4; exceeding the dynamic threshold T = μ + 2.5σ = 22.5. The system immediately triggers a red alert, marks it as "suspected insider trading", and pushes the trading time and associated accounts (such as "abnormal trading in the accounts of senior management's relatives") to the regulatory platform through the API.

[0062] Feedback optimization and system iteration: The regulatory agency confirms that the transaction is insider trading (TP), and the feedback receiving module passes the result to the feature self-update module. The feature evaluation sub-module calculates the current feature effectiveness E = 0.83 (threshold 0.7), but the recall rate Recall = 0.62Recall = 0.62. The feature optimization sub-module is launched: ① Introduce a new feature of "transaction period concentration" to capture multi-account collaborative operations within a short period; ② Adjust the weight coefficient α = 0.75 to strengthen text features; ③ Update the fine-tuned BERT model to enhance the semantic capture of keywords such as "inside information". After optimization, the recall rate of the system in the next cycle is increased to 0.85, and the false alarm rate is stabilized at 2.8%, forming a "data-driven - detection - optimization" closed loop.

[0063] The design focus of the present invention is that through the above system and method, the automatic generation and update of cross-modal transaction features are realized, the information in multi-source data is effectively utilized, the efficiency and accuracy of financial risk assessment and transaction monitoring are improved, and strong technical support is provided for financial institutions. In the specific implementation process, according to different business requirements and data characteristics, the parameters and algorithms of each module can be flexibly configured and adjusted to adapt to diverse application scenarios. For example, when processing different types of financial transaction data, an appropriate large language model and anomaly detection algorithm can be selected, the weight coefficient of feature fusion can be adjusted, and the strategies and frequencies of feature update can be optimized, etc., to ensure that the system can fully exert its advantages and meet actual business needs. At the same time, the closed-loop feedback mechanism of the present invention can continuously optimize the feature extraction and anomaly detection processes, enabling the system to maintain good performance and adaptability during long-term operation, promptly discover and warn of potential financial risks, and provide guarantee for the stable operation of financial institutions. In practical applications, the system can help financial institutions discover abnormal trading behaviors in advance, such as insider trading and market manipulation, and take timely measures for risk prevention and control to reduce economic losses. At the same time, through in-depth analysis of cross-modal transaction features, financial institutions can better understand market dynamics and customer behaviors, provide data support for business decisions, and enhance market competitiveness.

[0064] The above is only a preferred embodiment of the present invention, and does not impose any limitation on the technical scope of the present invention. Therefore, any minor modifications, same changes, and decorations made to the above embodiments based on the technical essence of the present invention still fall within the scope of the technical solution of the present invention.

Claims

1. A cross-modal transaction feature self-generation and anomaly detection system integrating large language models, characterized in that including: A data integration module for connecting to multi-source data platforms and collecting structured and unstructured data; A cross-modal feature extraction module that uses a large language model to process the collected data and extract cross-modal transaction features; An anomaly detection module that analyzes cross-modal transaction feature vectors based on a preset anomaly detection algorithm to identify abnormal transaction features; A feature self-update module that automatically updates cross-modal transaction features according to system feedback and data changes; A feedback receiving module for receiving anomaly detection results from the anomaly detection module and feedback information from external users or systems, and sending the feedback information to the feature self-update module to form a closed-loop feedback mechanism; The data integration module, cross-modal feature extraction module, anomaly detection module, and feature self-update module are connected in sequence, and the feedback receiving module is connected to the anomaly detection module and the feature self-update module.

2. The cross-modal transaction feature self-generation and anomaly detection system integrating a large language model according to claim 1, characterized in that The data integration module supports connecting to financial data platforms such as Caixin, DM, Wind, and YY, flexibly expands the data ecosystem according to requirements, and uses standardized interface technology to achieve data collection.

3. The cross-modal transaction feature self-generation and anomaly detection system integrating a large language model according to claim 1, characterized in that, The cross-modal feature extraction module specifically: inputs unstructured data into a pre-trained large language model, outputs a semantic vector representation after being processed by the model encoding layer, and then fuses the semantic vector with structured data features through a feature fusion sub-module to generate a cross-modal transaction feature vector.

4. The cross-modal transaction feature self-generation and anomaly detection system integrating a large language model according to claim 3, characterized in that, The feature fusion sub-module uses the following formula for feature fusion: ; Among them, F represents the cross-modal transaction feature vector, and F text represents the semantic vector output by the large language model, and F struct represents the structured data feature vector. α and β are the weight coefficients of semantic features and structured features respectively, which are used to balance the contributions of the two features in the fusion process. Their value ranges are between [0, 1], and α + β = 1.

5. The cross-modal transaction feature self-generation and anomaly detection system integrating a large language model according to claim 1, characterized in that, The feature self-update module includes a feature evaluation sub-module and a feature optimization sub-module. The feature evaluation sub-module is used to evaluate the effectiveness and relevance of existing features, and the feature optimization sub-module optimizes and adjusts the features according to the evaluation results.

6. The cross-modal transaction feature self-generation and anomaly detection system integrating a large language model according to claim 1, characterized in that The anomaly detection algorithm uses one or a combination of methods based on statistics, machine learning, or deep learning to flexibly detect abnormal transaction features according to different types of transaction data and business scenarios.

7. The cross-modal transaction feature self-generation and anomaly detection system integrating a large language model according to claim 6, wherein, When the anomaly detection module uses a method based on statistics, the following formula is used to calculate the anomaly score of the data: ; Among them, S anomaly represents the anomaly score, x represents the current cross-modal transaction feature vector, μ represents the mean of historical data, σ represents the standard deviation of historical data, and the transaction feature is judged whether it is abnormal by comparing the anomaly score with a preset threshold.

8. The cross-modal transaction feature self-generation and anomaly detection system integrating a large language model according to claim 5, characterized in that The feature evaluation sub-module evaluates the effectiveness of existing features using the following formula: ; where E represents feature effectiveness, TP represents true positives, TN represents true negatives, FP represents false positives, and FN represents false negatives. By calculating feature effectiveness, the performance of features in anomaly detection is evaluated, providing a basis for feature optimization.

9. A method applied to the cross-modal transaction feature self-generation and anomaly detection system of the integrated large language model as described in any one of claims 1-8, characterized in that, including the following steps: S1. Connect to multi-source data platforms through the data integration module and collect structured and unstructured data; S2. Use a large language model to process the collected data and extract cross-modal transaction features; S3. Analyze cross-modal transaction feature vectors based on a preset anomaly detection algorithm to identify abnormal transaction features; S4. Receive anomaly detection results and external feedback information through the feedback receiving module and send them to the feature self-update module; S5. According to the received feedback information, the feature self-update module automatically updates cross-modal transaction features to achieve the automatic generation, anomaly detection, and automatic update and optimization of cross-modal transaction features.

10. The method according to claim 9, wherein In step S2, the specific process of extracting cross-modal transaction features is as follows: The unstructured data is input into a pre-trained large language model, and after being processed by the model's encoding layer, a semantic vector representation is output. Then, through the feature fusion sub-module, the semantic vector is fused with the structured data features to generate a cross-modal transaction feature vector.

Citation Information

Patent Citations

  • Power grid business data feature extraction method

    CN119669722A

  • Multi-source heterogeneous data fusion and processing method based on big data

    CN119783037A

  • Quantitative transaction method and system fusing multi-source information data

    CN119904306A

  • Training method of financial risk identification model based on big data

    CN120013672A

Cited By

  • Multi-source heterogeneous data fusion pre-training method for unmanned aerial vehicle cluster task planning

    CN122548661A