Server for multimodal-based prediction model for detecting abnormal transaction and driving method thereof

The multimodal prediction model improves anomaly detection in financial transactions by integrating image and audio data with transaction history, addressing the limitations of conventional systems that rely solely on quantitative data.

WO2026049174A1PCT designated stage Publication Date: 2026-03-05ONCLEV INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/000878
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-29
Filing Date
2025-01-15
Publication Date
2026-03-05

AI Technical Summary

Technical Problem

Conventional anomaly detection systems in financial transactions rely solely on quantitative financial data, failing to consider qualitative contexts, leading to low detection accuracy and potential significant financial losses.

Method used

A server-based multimodal prediction model that integrates image, audio, and transaction history data to analyze transaction environments, extracting keywords and vectors to determine anomalies using a trained prediction model with threshold values.

Benefits of technology

Enhances detection accuracy by considering both quantitative and qualitative contexts, enabling precise identification of abnormal transactions and reducing financial losses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025000878_05032026_PF_FP_ABST
    Figure KR2025000878_05032026_PF_FP_ABST
Patent Text Reader

Abstract

According to the present invention, a server for a multimodal-based prediction model for detecting an abnormal transaction comprises: a communication interface; and a processor, wherein the processor is configured to receive at least one piece of image data obtained by capturing a transaction environment, at least one piece of voice data obtained by capturing the transaction environment, and at least one piece of transaction history data from an external electronic device through the communication interface, extract at least one piece of first analysis data obtained by analyzing the at least one piece of image data and the at least one piece of voice data, extract at least one piece of second analysis data obtained by analyzing the at least one piece of transaction history data, input the at least one piece of first analysis data and the at least one piece of second analysis data into the prediction model to output at least one piece of transaction prediction data composed of at least one piece of first analysis data and second analysis data, identify the at least one piece of transaction prediction data as an abnormal transaction when the at least one piece of transaction prediction data exceeds a preset first threshold, and transmit the at least one piece of identified transaction prediction data to the external electronic device through the communication interface, and the prediction model is trained on the basis of a plurality of pieces of image data obtained by capturing a transaction environment, a plurality of pieces of voice data recorded in the transaction environment, a plurality of pieces of transaction history data, a plurality of pieces of first analysis data and a plurality of pieces of second analysis data, and a plurality of pieces of abnormal transaction detection data and a plurality of pieces of normal transaction detection data.
Need to check novelty before this filing date? Find Prior Art

Description

Server and its operation method for a multimodal-based prediction model for anomaly detection

[0001] Various embodiments of the present invention relate to a server and a method for operating the same for a multimodal-based prediction model for detecting abnormal transactions.

[0002] According to the Financial Supervisory Service, the amount of damage caused by abnormal transactions has exceeded 800 billion won over the past five years, which can result in enormous financial losses for financial companies and customers. Therefore, FDS, which can identify and detect abnormal transaction patterns, is being widely adopted by many banks and insurance companies to prevent financial crime and protect customers' assets.

[0003] Meanwhile, according to the "2023 Unfair Trade Investigation Results" released by the Korea Exchange Market Surveillance Committee, the scale of unfair profits resulting from unfair trading reached approximately KRW 766.3 billion, the highest in the past five years. While various surveillance systems, such as circuit breakers and trading suspensions, have been introduced to ensure stock market stability and investor protection, a system for real-time monitoring of abnormal trading and signs of abnormalities resulting from complex factors has yet to be established, despite the growing need for such a system. In particular, the designation of stocks under caution is often carried out without explanation, potentially resulting in significant losses.

[0004] Meanwhile, conventional analysis systems primarily use only financial transaction data as input to their anomaly detection algorithms. Using only this data to detect anomalies has limitations, such as failing to consider the qualitative context of the transaction.

[0005] For example, a customer who previously visited alone might one day find himself accompanied by a mysterious man, or might fail to properly answer a bank teller's question about the purpose of a withdrawal. These qualitative contexts are difficult to capture solely through transaction history data. This inability to fully understand the circumstances surrounding a transaction can lead to low detection accuracy.

[0006] The present invention can receive at least one contract data, extract keywords and analysis data, distinguish them into vectors, input each data into a prediction model, and provide output transaction prediction data.

[0007] A server for a multimodal-based prediction model for abnormal transaction detection according to the present invention includes a communication interface; and a processor. The processor is configured to receive, through the communication interface, at least one image data photographing a transaction environment, at least one audio data photographing the transaction environment, and at least one transaction history data from an external electronic device, extract at least one first analysis data analyzed by analyzing the at least one image data and the at least one audio data, extract at least one second analysis data analyzed by analyzing the at least one transaction history data, input the at least one first analysis data and the at least one second analysis data into the prediction model, output at least one transaction prediction data composed of the at least one first analysis data and the second analysis data, determine the at least one transaction prediction data as an abnormal transaction when the at least one transaction prediction data exceeds a preset first threshold value, and transmit the determined at least one transaction prediction data to the external electronic device through the communication interface, wherein the prediction model is trained based on a plurality of image data photographing the transaction environment, a plurality of audio data recorded in the transaction environment, a plurality of transaction history data, a plurality of first analysis data, a plurality of second analysis data, a plurality of abnormal transaction detection data, and a plurality of normal transaction detection data.

[0008] According to various embodiments of the present invention, by receiving at least one contract data, extracting keywords and analysis data, distinguishing them into vectors, inputting each data into a prediction model, and providing output transaction prediction data, it is possible to easily and accurately predict and respond to intentional unfair transactions by a contractor during an investment transaction, and it is possible to easily check anomalies in vectors extracted with preset threshold values, so that a module capable of grasping both quantitative and qualitative contexts can be applied, and furthermore, there is an advantage in that it can be added to and applied to FDS used by actual financial institutions to more precisely detect abnormal transactions.

[0009] FIG. 1 illustrates a block diagram of a server and a network according to various embodiments of the present invention.

[0010] FIG. 2 is a flowchart illustrating how a server operates according to various embodiments.

[0011] FIG. 3 is a structural diagram for explaining the basic operation of a server according to various embodiments.

[0012] A server for a multimodal-based prediction model for detecting abnormal transactions according to the present invention can receive one contract data, extract keywords and analysis data, distinguish them into vectors, input each data into a prediction model, and provide output transaction prediction data. Specifically, the server for a multimodal-based prediction model for detecting abnormal transactions according to the present invention includes a communication interface; and a processor; wherein the processor (1) receives at least one image data photographing a transaction environment from an external electronic device, at least one audio data photographing the transaction environment, and at least one transaction history data; (2) extracting at least one first analysis data by analyzing the at least one image data and the at least one voice data, extracting at least one second analysis data by analyzing the at least one transaction history data, (3) inputting the at least one first analysis data and the at least one second analysis data into the prediction model to output at least one transaction prediction data composed of the at least one first analysis data and the second analysis data, and (4) determining the at least one transaction prediction data as an abnormal transaction when the at least one transaction prediction data exceeds a preset first threshold value, and transmitting the determined at least one transaction prediction data to the external electronic device through the communication interface. Meanwhile, the prediction model is trained based on a plurality of image data photographing a transaction environment, a plurality of voice data recorded in a transaction environment, a plurality of transaction history data, a plurality of first analysis data, and a plurality of second analysis data, a plurality of abnormal transaction detection data, and a plurality of normal transaction detection data.

[0013] Hereinafter, various embodiments of the present document will be described with reference to the attached drawings. It should be understood that the embodiments and the terms used therein are not intended to limit the technology described in the present document to a specific embodiment, but rather include various modifications, equivalents, and / or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar components. The singular expression may include plural expressions unless the context clearly indicates otherwise. In this document, expressions such as "A or B" or "at least one of A and / or B" may include all possible combinations of the items listed together. Expressions such as "first," "second," "first," or "second," may modify the corresponding components regardless of order or importance, and are only used to distinguish one component from another, but do not limit the corresponding components. When it is said that a component (e.g., a first component) is “(functionally or communicatively) connected” or “connected” to another component (e.g., a second component), said component may be directly connected to said other component, or may be connected via another component (e.g., a third component).

[0014] In this document, "configured to" may be used interchangeably with, for example, "suitable for," "capable of," "modified to," "made to," "capable of," or "designed to," either in hardware or software. In some contexts, the phrase "a device configured to" may mean that the device is "capable of" doing something together with other devices or components. For example, the phrase "a processor configured to perform A, B, and C" may mean a dedicated processor (e.g., an embedded processor) for performing the operations, or a general-purpose processor (e.g., a CPU or application processor) that can perform the operations by executing one or more software programs stored in a memory device.

[0015] An electronic device according to various embodiments of the present document may include, for example, at least one of a smartphone, a tablet PC, a desktop PC, a laptop PC, a netbook computer, a workstation, and a server.

[0016] Referring to FIG. 1, an electronic device (101) within a network environment (100) according to various embodiments is described. The electronic device (101) may include a bus (110), a processor (120), a memory (130), an input / output interface (140), a display (150), and a communication interface (160). In some embodiments, the electronic device (101) may omit at least one of the components or additionally include other components. The bus (110) may include a circuit that connects the components (110-170) to each other and transmits communication (e.g., control messages or data) between the components. The processor (120) may include one or more of a central processing unit, an application processor, or a communication processor (CP). The processor (120) may, for example, execute operations or data processing related to control and / or communication of at least one other component of the electronic device (101).

[0017] The memory (130) may include volatile and / or non-volatile memory. The memory (130) may store, for example, commands or data related to at least one other component of the electronic device (101). According to one embodiment, the memory (130) may store software and / or programs (140).

[0018] The input / output interface (140) can, for example, transmit commands or data input from a patient or other external device to other component(s) of the electronic device (101), or output commands or data received from other component(s) of the electronic device (101) to the patient or other external device.

[0019] The display (150) may include, for example, a liquid crystal display (LCD), a light emitting diode (LED) display, an organic light emitting diode (OLED) display, a micro electro mechanical systems (MEMS) display, or an electronic paper display. The display (150) may, for example, display various contents (e.g., text, images, videos, icons, and / or symbols) to the patient. The display (150) may include a touch screen and may receive, for example, touch, gesture, proximity, or hovering input using an electronic pen or a part of the patient's body. The communication interface (160) may, for example, establish communication between the electronic device (101) and an external device (e.g., the first external electronic device (102), the second external electronic device (104), or the server (108)). For example, the communication interface (160) can be connected to a network (162) via wireless communication or wired communication to communicate with an external device (e.g., a second external electronic device (104) or a server (108)).

[0020] The wireless communication may include, for example, cellular communication using at least one of LTE, LTE-A (LTE Advance), CDMA (code division multiple access), WCDMA (wideband CDMA), UMTS (universal mobile telecommunications system), WiBro (Wireless Broadband), or GSM (Global System for Mobile Communications). In one embodiment, the wireless communication may include, for example, at least one of WiFi (wireless fidelity), Bluetooth, Bluetooth low energy (BLE), Zigbee, near field communication (NFC), Magnetic Secure Transmission, radio frequency (RF), or body area network (BAN). In one embodiment, the wireless communication may include GNSS. The GNSS may be, for example, GPS (Global Positioning System), Glonass (Global Navigation Satellite System), Beidou Navigation Satellite System (hereinafter "Beidou"), or Galileo, the European global satellite-based navigation system. Hereinafter, in this document, "GPS" may be used interchangeably with "GNSS." Wired communication may include at least one of, for example, USB (universal serial bus), HDMI (high definition multimedia interface), RS-232 (recommended standard 232), power line communication, or POTS (plain old telephone service).The network (162) may include at least one of a telecommunications network, for example, a computer network (e.g., a LAN or WAN), the Internet, or a telephone network.

[0021] Each of the first and second external electronic devices (102, 104, 106) may be the same or a different type of device as the electronic device (101). According to various embodiments, all or part of the operations executed in the electronic device (101) may be executed in another one or more electronic devices (e.g., electronic devices (102, 104, 106), or server (108). According to one embodiment, when the electronic device (101) is to perform a certain function or service automatically or upon request, the electronic device (101) may request at least some functions related thereto from another device (e.g., electronic devices (102, 104, 106), or server (108)) instead of executing the function or service by itself or in addition. The other electronic device (e.g., electronic devices (102, 104, 106), or server (108)) may execute the requested function or additional function and transmit the result to the electronic device (101). The electronic device (101) may process the received result as is or additionally to provide the requested function or service. For this purpose, for example, cloud computing, distributed computing, or client-server computing technology may be used.

[0022] FIG. 2 is a flowchart for explaining how a server operates according to various embodiments, and FIG. 3 is a structural diagram for explaining the basic operation of a server according to various embodiments.

[0023] In operation 201, the server (108) (e.g., the processor (120) of FIG. 1) may receive, from an external electronic device (101, 102, 104, 106) at least one image data photographing a transaction environment, at least one audio data photographing the transaction environment, and at least one transaction history data through a communication interface (e.g., the communication interface (160) of FIG. 1). According to one embodiment, the at least one image data may include at least one RPG image data and infrared image data photographed from an external camera module installed in the transaction environment, the at least one audio data may include at least one voice data and recording data acquired from an external microphone module installed in the transaction environment, and the at least one transaction history data may include at least one text data and at least one numeric data including a transaction history used in the transaction environment. According to another embodiment, at least one image data may be composed of RPG image data and RPG video data captured and stored through an RPG camera module linked to an external electronic device (101, 102, 104, 106), and infrared image data and infrared video data captured and stored through an infrared camera module. This allows the server (108) to confirm whether the actual customer visited alone and performed a contract or visited with a companion and performed a contract, and to obtain infrared image data and infrared video data to determine the customer's emotional state. According to another embodiment, at least one voice data may be composed of voice data and recording data recorded and stored through a microphone module linked to an external electronic device (101, 102, 104, 106). This allows the server (108) to obtain voice data to determine the customer's emotional state based on the trembling of the customer's voice when the actual customer performs a contract.According to another embodiment, at least one transaction history data may be composed of text data in the form of text and numerical data in the form of numerical values ​​among image data captured and / or recorded and stored through an RPG camera module and / or application linked to an external electronic device (101, 102, 104, 106). This allows the server (108) to obtain text data and numerical data for a contract actually signed by a customer and determine whether there is an abnormality in the contract. Thereafter, the server (108) may store at least one image data, at least one audio data, and at least one transaction history data in a memory (e.g., memory (130) of FIG. 1).

[0024] In operation 203, the server (108) (e.g., the processor (120) of FIG. 1) may extract at least one first analysis data by analyzing at least one image data and at least one voice data, as illustrated in FIG. 3, and may extract at least one second analysis data by analyzing at least one transaction history data. According to one embodiment, the at least one first analysis data may include headcount data, heart rate data, voice tremor data, and question-and-answer delay time data, and the at least one second analysis data may include at least one transaction keyword data and at least one numerical data. According to another embodiment, the server (108) may extract headcount data captured from at least one RPG image data, heart rate data of headcount data captured from at least one infrared image data, voice tremor data recorded from at least one voice data, and at least one transaction keyword data directly associated with a cost for a transaction from at least one text data, and at least one numerical data directly associated with a cost used for the transaction from at least one numerical data. This means that the server (108) can determine whether the person is a passenger of the contractor by checking the number of people among at least one RPG video data which is the original data, can determine whether the contractor is lying by checking the psychological changes of the contractor among the infrared video data, can determine whether the contractor is lying based on the voice and delay among the voice data, can determine whether the contractor is fraudulent by checking the data directly related to the contract among the text data, and can determine the amount and period of the contractor by checking the contract and the contract amount and contract time directly related to the contract among the numerical data.

[0025] In operation 205, the server (108) (e.g., the processor (120) of FIG. 1) may input at least one first analysis data and at least one second analysis data into a prediction model, as illustrated in FIG. 3, and output at least one transaction prediction data composed of at least one first analysis data and the second analysis data. According to one embodiment, the prediction model may be trained by inputting a plurality of image data capturing a transaction environment, a plurality of audio data capturing a transaction environment, a plurality of transaction history data, a plurality of first analysis data, and a plurality of second analysis data, and a plurality of first analysis data and second analysis data, and outputting a plurality of abnormal transaction detection data and a plurality of normal transaction detection data.

[0026] Specifically, the predictive model for predicting whether a transaction is abnormal and outputs transaction prediction data can be driven by an AI neural network model trained through unsupervised learning from basic data. The predictive model can be configured to increase the ease of data collection and as output values ​​of various data. The predictive model has a structure that outputs binary (normal or abnormal) data from image-audio-text data, and at least one of BigScience's bloom and T0pp, EleutherAI's GPT series, Tsinghua UNIV's GLM series, GOOGLE's UL and T5 series, and META AI's OPT series can be used. According to one embodiment, the predictive model can be implemented using the structure of OpenAI's Chat GPT, Google's BARD series, and NVIDIA's translation service-based Transformer model using multiple cloud foundation models. For example, the predictive model can be implemented by training a multi-modal model based on at least one of text data, image data, and audio data. In one embodiment, this can reduce the amount of labeled task-specific training data compared to existing deep learning approaches, and once built, it can be trained on a variety of tasks with a small amount of training data, making data collection and labeling easier and improving accuracy.

[0027] According to another embodiment, the server (108) may perform a learning process of a prediction model for predicting an investment prediction report based on a plurality of image data capturing a transaction environment, a plurality of audio data capturing a transaction environment, a plurality of transaction history data, a plurality of first analysis data, a plurality of second analysis data, a plurality of first analysis data and a plurality of second analysis data, a plurality of abnormal transaction detection data, and a plurality of normal transaction detection data, by obtaining a result value (output data) using a prediction model to which arbitrary weights are assigned, comparing the obtained result value with labeled data of the learning data, and performing backpropagation according to the error, thereby optimizing the weights. Specifically, the learning of the prediction model refers to a process of training the prediction model based on the learning data and labeled data or unlabeled data, so that the prediction model can determine output data for input data. In other words, the prediction model forms rules for the data and makes a judgment. According to one embodiment, the server (108) may use a plurality of learning algorithms among a plurality of learning algorithms for calculating a predicted value. For example, an ensemble method can be used for a predictive model, which can achieve better predictive performance than when learning algorithms are used separately. Training a predictive model can mean adjusting the weights of the model. In one embodiment, various learning methods can be used, such as supervised learning, unsupervised learning, reinforcement learning, imitation learning, and federated learning.

[0028] Although not shown, the server (108) may include an evaluation step for evaluating the performance of the prediction model during the learning process of the prediction model. In the evaluation step, the prediction model may be evaluated using an evaluation data set. The evaluation of the prediction model may be a step of evaluating the prediction model learned through the learning step and using the prediction model to make predictions on new data. Specifically, the evaluation step may be a step of measuring whether the learned prediction model is capable of generalizing to new data.

[0029] According to another embodiment, the server (108) may extract an image feature vector for an image and an audio feature vector for an audio based on at least one first analysis data, and may extract a transaction feature vector for a transaction based on at least one second analysis data. This is because the server (108) may convert past transaction history data and current transaction history data for a plurality of customers into a dynamic graph form and input the graph into a prediction model to extract an account feature vector, a transaction feature vector, and a neighboring account feature vector, and may extract a plurality of image feature vectors, a plurality of audio feature vectors, and a plurality of transaction feature vectors, and may input the classified data based on the result values ​​into the prediction model and train the model. Thereafter, the server (108) may convert at least one first analysis data and at least one second analysis data into a graph form and input the graph into the prediction model to extract each of an image feature vector, an audio feature vector, and a transaction feature vector, and may classify the data based on previously trained data.

[0030] According to another embodiment, the server (108) may input data into a learnable aggregator based on at least one first analysis data and at least one second analysis data, and extract a vector image-audio feature vector in which a portion related to an image feature vector and an audio feature vector is reset as output data, a vector image-transaction feature vector in which a portion related to an image feature vector and an audio feature vector is reset, and a vector audio-transaction feature vector in which a portion related to an audio feature vector and an audio feature vector is reset as output data.

[0031] This means that the server (108) can extract even the adjacent vectors between the respective vectors based on a plurality of image feature vectors, a plurality of audio feature vectors, and a plurality of transaction feature vectors, and based on the result values ​​thereof, can input the classified data into a prediction model and train it, and further, can concatenate the extracted vector image-audio feature vectors, image-transaction feature vectors, and audio-transaction feature vectors and utilize them as inputs of a feedforward neural network to be used for abnormal transaction detection. In addition, the server (108) can be extracted through a model other than the prediction model. This can be extracted with data trained by a different model rather than increasing the amount of training data by granting all functions to the prediction model. Thereafter, the server (108) can combine the output image feature vectors, audio feature vectors, and transaction feature vectors, convert them into a graph form, and input them into a prediction model to extract each vector image-audio feature vector, image-transaction feature vector, and audio-transaction feature vector, and can classify them based on the previously trained data.

[0032] In another embodiment, the prediction model may additionally learn feature vectors for multiple images, feature vectors for multiple audio, feature vectors for multiple transactions, and multiple video-audio feature vectors, multiple video-transaction feature vectors, and multiple audio-transaction feature vectors. This allows the server (108) to embed vectors for each feature based on the prediction model, embed vectors combining similar content between each feature, and determine whether the transaction is a fraudulent transaction by considering both quantitative and qualitative contexts.

[0033] In operation 207, the server (108) (e.g., the processor (120) of FIG. 1) may determine whether at least one transaction prediction data exceeds a first threshold value. According to one embodiment, the first threshold value may be set based on whether at least two or more vectors among the image-audio feature vector, the image-transaction feature vector, and the voice-transaction feature vector detect an anomaly, and whether at least two or more vectors among the image vector, the voice vector, and the transaction vector detect an anomaly. Specifically, the server (108) may determine that the video vector is an anomaly if, among the video vectors, the contractor side is verified to have multiple people (e.g., if the number of people on the contractor side is determined to be multiple in an RPG image) and data showing emotional fluctuations (e.g., if there is a large fluctuation per second in an infrared image), and if, among the voice vectors, the contractor side has a voice tremor level higher than the average value (e.g., if the vibration value compared to static is 10 Hz or higher) and a question-and-response delay time higher than the average (e.g., if it is determined that there was silence for more than 3 seconds in response to an employee's question) at least once, the server may determine that the voice vector is an anomaly. In addition, the server (108) may determine that, among the transaction vectors, the text written by the contractor side is inaccurate (e.g., the signature of the signature fluctuates a lot, and it is not a stamp of the same name) and the numerical value is abnormally high (e.g., the acquisition cost is very high compared to the membership fee) at least once, the server may determine that the transaction vector is an anomaly.In addition, the server (108) can determine that it is an abnormal phenomenon if it is found that the image data with emotional fluctuations and the data with shaking sounds match even once among the video-audio feature vectors, and if it is found that the image data with emotional fluctuations and the data with inaccurate signatures and abnormally high values ​​match even once among the video-transaction feature vectors, it can determine that it is an abnormal phenomenon, and if it is found that the data with inaccurate signatures and abnormally high values ​​match even once among the audio-transaction feature vectors, it can determine that it is an abnormal phenomenon. Thereafter, the server (108) can determine whether it is an abnormal transaction by comparing it with at least one transaction prediction data based on the first threshold value and the criteria of the six vectors.

[0034] In operation 209, if the server (108) (e.g., the processor (120) of FIG. 1) determines that at least one transaction prediction data exceeds the first threshold, the server (108) may determine that at least one transaction prediction data is an abnormal transaction. In one embodiment, the server (108) may determine that an abnormal phenomenon occurs when, among the video-audio feature vectors, image data with emotional fluctuations and trembling sound data are found to be in a state of matching at least once, and may determine that an abnormal phenomenon occurs when, among the video-transaction feature vectors, image data with emotional fluctuations and data with abnormally high values ​​that are not accurate and have signatures are found to be in a state of matching at least once. In addition, the server (108) can determine that the image vector is an anomaly if it is verified at least once that the contractor side is a human being with two emotions and that the infrared image fluctuates greatly among the image vectors, and can determine that the voice vector is an anomaly if it is verified at least once that the contractor side has a voice tremor of 14 Hz or higher and a question-and-answer delay time of 4.5 seconds or higher among the voice vectors. In addition, the server (108) can determine that the transaction vector is an anomaly if it is verified that the signature of the text written by the contractor side is inaccurate and the numerical value is abnormally high among the transaction vectors. Thereafter, the server (108) can determine that at least two or more vectors among the image-audio feature vector, the image-transaction feature vector, and the voice-transaction feature vector detect an anomaly, and at least two or more vectors among the image vector, the voice vector, and the transaction vector detect an anomaly, and therefore at least one transaction prediction data exceeds the first threshold value.

[0035] Meanwhile, in operation 207, if the server (108) (e.g., the processor (120) of FIG. 1) determines that at least one transaction prediction data does not exceed the first threshold value, in operation 211, the server (108) (e.g., the processor (120) of FIG. 1) may determine whether at least one transaction prediction data exceeds the second threshold value. The second threshold value, which is set to be smaller than the first threshold value, may be set based on whether at least one of the video-audio feature vector, the video-transaction feature vector, and the voice-transaction feature vector detects an anomaly, and whether at least one of the video vector, the voice vector, and the transaction vector detects an anomaly. Here, the video vector, the voice vector, and the transaction vector, and the video-audio feature vector, the video-transaction feature vector, and the voice-transaction feature vector may be the same as the anomaly verification above. Thereafter, the server (108) can determine whether a transaction is abnormal by comparing it with at least one transaction prediction data based on the first threshold value and the six vector criteria. Thereafter, the server (108) can determine whether a transaction is likely abnormal by comparing it with at least one transaction prediction data based on the second threshold value and the six vector criteria.

[0036] In operation 213, if the server (108) (e.g., the processor (120) of FIG. 1) determines that at least one transaction prediction data exceeds the second threshold, the server (108) may determine that at least one transaction data has a possibility of being an abnormal transaction. In one embodiment, the server (108) may determine that an abnormal phenomenon is present if, among the image-audio feature vectors, image data with emotional fluctuations and trembling sound data are found to be consistent at least once, and if, among the image vectors, it is verified that the contractor side is two humans and that the emotionally fluctuating infrared image fluctuates at least once, the image vector may be determined to be an abnormal phenomenon. If no other abnormalities are found, the server may determine that at least one transaction data has a possibility of being an abnormal transaction.

[0037] Meanwhile, in operation 211, if the server (108) (e.g., the processor (120) of FIG. 1) determines that at least one transaction prediction data does not exceed the second threshold value, then in operation 215, the server (108) (e.g., the processor (120) of FIG. 1) may determine that at least one transaction data is a normal transaction. In one embodiment, the server (108) may determine that at least one transaction data is a normal transaction if no abnormality is found in the entirety of the image vector, the audio vector, and the transaction vector, and the image-audio feature vector, the image-transaction feature vector, and the audio-transaction feature vector.

[0038] In operation 217, the server (108) (e.g., the processor (120) of FIG. 1) may transmit at least one determined transaction prediction data to an external electronic device (101, 102, 104, 106) via a communication interface (e.g., the communication interface (160) of FIG. 1). In one embodiment, the server (108) may determine whether at least one transaction data is an abnormal transaction, has a possibility of being an abnormal transaction, or is normal through a first threshold value and / or a second threshold value, and transmit the determined transaction data together with the data to the external electronic device (101, 102, 104, 106).

[0039]

[0040] According to various embodiments of the present invention, by receiving at least one contract data, extracting keywords and analysis data, distinguishing them into vectors, inputting each data into a prediction model, and providing output transaction prediction data, it is possible to easily and accurately predict and respond to intentional unfair transactions by a contractor during an investment transaction, and it is possible to easily check anomalies in vectors extracted with preset threshold values, so that a module capable of grasping both quantitative and qualitative contexts can be applied, and furthermore, there is an advantage in that it can be added to and applied to FDS used by actual financial institutions to more precisely detect abnormal transactions.

[0041]

[0042] According to various embodiments, a server for a multimodal based prediction model for anomaly detection comprises: a communication interface; a processor; Including, the processor is configured to receive, through the communication interface, at least one image data photographing a transaction environment, at least one voice data photographing the transaction environment, and at least one transaction history data from an external electronic device, extract at least one first analysis data analyzed by analyzing the at least one image data and the at least one voice data, extract at least one second analysis data analyzed by analyzing the at least one transaction history data, input the at least one first analysis data and the at least one second analysis data into the prediction model, output at least one transaction prediction data composed of the at least one first analysis data and the second analysis data, determine the at least one transaction prediction data as an abnormal transaction when the at least one transaction prediction data exceeds a preset first threshold value, and transmit the determined at least one transaction prediction data to the external electronic device through the communication interface, wherein the prediction model is trained based on a plurality of image data photographing a transaction environment, a plurality of voice data recorded in the transaction environment, a plurality of transaction history data, a plurality of first analysis data, and a plurality of second analysis data, a plurality of abnormal transaction detection data, and a plurality of normal transaction detection data.

[0043] According to various embodiments, the at least one image data includes at least one RPG image data and infrared image data captured from an external camera module installed in the transaction environment, the at least one voice data includes at least one voice data and recording data acquired from an external microphone module installed in the transaction environment, and the at least one transaction history data includes at least one text data and at least one numeric data containing a transaction history used in the transaction environment.

[0044] According to various embodiments, the processor is configured to extract headcount data captured from the at least one RPG image data, heart rate data of the headcount data captured from the at least one infrared image data, voice tremor degree data and question-and-answer delay time data recorded from the at least one voice data, and extract at least one transaction keyword data directly associated with a cost of the transaction from the at least one text data, and at least one numerical data directly associated with a cost used in the transaction from the at least one numerical data, wherein the at least one first analysis data includes the headcount data, the heart rate data, the voice tremor degree data, and the question-and-answer delay time data, and the at least one second analysis data includes the at least one transaction keyword data and the at least one numerical data.

[0045] According to various embodiments, the processor is configured to extract an image feature vector for an image and an audio feature vector for an audio based on the at least one first analysis data, respectively, extract a transaction feature vector for a transaction based on the at least one second analysis data, and extract a vector image-audio feature vector in which a portion related to the image feature vector and the audio feature vector is reset as output data by inputting the data into a learnable aggregator based on the at least one first analysis data and the at least one second analysis data, a video-transaction feature vector in which a portion related to the image feature vector and the audio feature vector is reset, and a audio-transaction feature vector in which a portion related to the audio feature vector and the audio feature vector is reset, and the prediction model is further trained with feature vectors for a plurality of images, feature vectors for a plurality of audios, feature vectors for a plurality of transactions, a plurality of image-audio feature vectors, a plurality of video-transaction feature vectors, and a plurality of audio-transaction feature vectors.

[0046] According to various embodiments, the first threshold value is set based on the detection of an anomaly by at least two vectors among the video-audio feature vector, the video-transaction feature vector, and the voice-transaction feature vector, and the detection of an anomaly by at least two vectors among the video vector, the voice vector, and the transaction vector, and the second threshold value, which is set smaller than the first threshold value, is set based on the detection of an anomaly by at least one vector among the video-audio feature vector, the video-transaction feature vector, and the voice-transaction feature vector, and the detection of an anomaly by at least one vector among the video vector, the voice vector, and the transaction vector, and the processor performs anomaly detection by connecting the video vector, the voice vector, the transaction vector, the video-audio feature vector, the video-transaction feature vector, and the voice-transaction feature vector, and when the at least one transaction prediction data does not exceed the first threshold value but exceeds the second threshold value, it is determined that the at least one transaction data has a possibility of being an abnormal transaction, and If the at least one transaction prediction data does not exceed the second threshold value, the at least one transaction data is determined to be a normal transaction.

[0047]

[0048] The term "module" or "part" used in this document includes a unit composed of hardware, software, or firmware, and can be used interchangeably with terms such as logic, logic block, component, or circuit, for example. The "module" or "part" can be an integrally configured component or a minimum unit or a part thereof that performs one or more functions. The "module" or "part" can be implemented mechanically or electronically, and can include, for example, an ASIC (application-specific integrated circuit) chip, FPGAs (field-programmable gate arrays), or a programmable logic device, known or to be developed in the future, that performs certain operations, and can be executed by the processor (120). At least a part of the device (e.g., modules or functions thereof) or method (e.g., operations) according to various embodiments can be implemented as instructions stored in a computer-readable storage medium (e.g., memory (130)) in the form of a program module. When the above command is executed by a processor (e.g., processor (120)), the processor can perform a function corresponding to the command. The computer-readable recording medium may include a hard disk, a floppy disk, a magnetic medium (e.g., a magnetic tape), an optical recording medium (e.g., a CD-ROM, a DVD, a magneto-optical medium (e.g., a floptical disk), a built-in memory, etc. The command may include a code generated by a compiler or a code executable by an interpreter. A module or program module according to various embodiments may include at least one or more of the above-described components, some of which may be omitted, or other components may be further included. Operations performed by a module, a program module, or other components according to various embodiments may be executed sequentially, in parallel, iteratively, or heuristically, or at least some operations may be executed in a different order, omitted, or other operations may be added.

[0049] The embodiments disclosed in this document are presented for the purpose of explaining and understanding the disclosed technical content, and do not limit the scope of the present disclosure. Therefore, the scope of the present disclosure should be interpreted to include all modifications or various other embodiments based on the technical concepts of the present disclosure.

Claims

1. In the server for the multimodal-based prediction model for abnormal transaction detection, communication interface; Processor; including, The above processor, Through the above communication interface, at least one image data photographing the transaction environment, at least one voice data photographing the transaction environment, and at least one transaction history data are received from an external electronic device, Extracting at least one first analysis data by analyzing at least one of the above video data and at least one of the above audio data, and extracting at least one second analysis data by analyzing at least one of the above transaction history data, By inputting at least one first analysis data and at least one second analysis data into the prediction model, at least one transaction prediction data composed of at least one first analysis data and the second analysis data is output, If at least one of the above transaction prediction data exceeds a preset first threshold value, the at least one transaction prediction data is judged as an abnormal transaction, Through the communication interface, the determined at least one transaction prediction data is set to be transmitted to the external electronic device, The above prediction model is, It is learned based on multiple video data capturing the transaction environment, multiple audio data capturing the transaction environment, multiple transaction history data, multiple first analysis data, multiple second analysis data, multiple abnormal transaction detection data, and multiple normal transaction detection data. Server for multimodal based predictive model for anomaly detection.

2. In paragraph 1, At least one of the above image data, Contains at least one RPG image data and infrared image data captured from an external camera module installed in the above transaction environment, At least one of the above voice data, Contains at least one voice data and recording data obtained from an external microphone module installed in the above transaction environment, At least one of the above transaction history data, Containing at least one text data and at least one numeric data containing transaction details used in the above transaction environment, Server for multimodal based predictive model for anomaly detection.

3. In paragraph 2, The processor is set to extract the number of people data captured from the at least one RPG image data, the heart rate data of the number of people data captured from the at least one infrared image data, and the voice tremor degree data and question-and-answer delay time data recorded from the at least one voice data. At least one transaction keyword data directly related to a cost of the transaction is extracted from the at least one text data, and at least one numerical data directly related to a cost used in the transaction is extracted from the at least one numerical data. Server for multimodal based predictive model for anomaly detection.

4. In paragraph 3, At least one of the first analysis data, Including the above-mentioned number of people data, the above-mentioned heart rate data, the above-mentioned voice tremor degree data, and the above-mentioned question and answer delay time data, At least one of the second analysis data, comprising at least one transaction keyword data and at least one numerical data, Server for multimodal based predictive model for anomaly detection.

5. In paragraph 4, The above processor, Extracting an image feature vector for an image and an audio feature vector for an audio based on at least one of the first analysis data, Extracting a transaction feature vector for a transaction based on at least one of the second analysis data, and A vector image-audio feature vector that resets the portion related to the image feature vector and the audio feature vector based on the at least one first analysis data and the at least one second analysis data and is output as data input to a learnable aggregator, a video-transaction feature vector that resets the portion related to the image feature vector and the audio feature vector, and a voice-transaction feature vector that resets the portion related to the audio feature vector and the audio feature vector. Server for multimodal based predictive model for anomaly detection.

6. In paragraph 5, The above prediction model is, Additional feature vectors for multiple images, feature vectors for multiple voices, feature vectors for multiple transactions, and multiple image-audio feature vectors, multiple image-transaction feature vectors, and multiple voice-transaction feature vectors are learned. Server for multimodal based predictive model for anomaly detection.

7. In paragraph 6, The above first threshold value is, At least two of the above image-audio feature vectors, the image-transaction feature vectors, and the voice-transaction feature vectors are set based on detecting an anomaly, and at least two of the above image vectors, the voice vectors, and the transaction vectors are set based on detecting an anomaly, The second threshold value set to be smaller than the first threshold value is At least one of the above image-audio feature vectors, the image-transaction feature vectors, and the voice-transaction feature vectors is set based on detecting an abnormality, and at least one of the above image vectors, the voice vectors, and the transaction vectors is set based on detecting an abnormality, The above processor, Performing abnormal transaction detection by connecting the above image vector, the above audio vector, the above transaction vector, the above image-audio feature vector, the above image-transaction feature vector, and the above audio-transaction feature vector, If the at least one transaction prediction data does not exceed the first threshold value but exceeds the second threshold value, the at least one transaction data is determined to have a possibility of being an abnormal transaction, and If the at least one transaction prediction data does not exceed the second threshold value, the at least one transaction data is set to be determined as a normal transaction. Server for multimodal based predictive model for anomaly detection.

Citation Information

Patent Citations

  • Illegal financial transaction detection program

    JP2021144356A

  • Device for inspecting LCD assembly status and PCB board and methods thereof

    KR1020240147624A

  • System and method for risk assessment of financial transactions and computer program for the same

    KR102409019B1

  • System for detecting abnormal prepaid financial transaction and method of operation of the same

    KR102529027B1

  • Trench assembly

    KR102671553B1