Deep fraud content detection system based on multi-modal data fusion
The deep fraud content detection system, which integrates multimodal data fusion, uses chat analysis and historical analysis modules to identify AI-forged videos and voice recordings. This solves the problem of difficulty in identifying deep fraud in existing technologies and enables accurate identification and prevention of fraudulent content.
Patent Information
- Application Number
- CN202610071186.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-20
- Publication Date
- 2026-02-17
AI Technical Summary
Existing technologies struggle to effectively identify and prevent deepfake content generated by artificial intelligence, especially AI-forged videos and audio, leading to a high success rate for scams.
The deep fraud content detection system employs multimodal data fusion, including chat analysis, historical analysis, and fraud identification modules. By analyzing chat content, historical records, and memory time limits, it determines whether the requester is suspicious and verifies whether the suspicious parts are fraudulent content.
It can accurately identify AI-generated fake videos and audio, prevent fraud, reduce economic losses, and improve recognition accuracy.
Smart Images

Figure CN121543041A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of information processing technology, specifically to a deep fraud content detection system based on multimodal data fusion. Background Technology
[0002] "Deepfake" refers to fraudulent activities using deepfake technology. This technology is typically based on artificial intelligence, especially generative adversarial networks (GANs), to synthesize realistic fake images, audio, or video to impersonate others and commit fraud. Deepfake content can impersonate familiar people, gaining their trust to complete the scam. As technology advances, deepfake content becomes increasingly sophisticated, with images and audio so realistic that they are difficult to distinguish from genuine images or audio by the naked eye or computer. Consequently, deepfake content can successfully swindle large sums of money, and current technology is not well-suited to address this issue. Summary of the Invention
[0003] To address the aforementioned technical problems, a deep fraud content detection system based on multimodal data fusion is provided. This technical solution solves the problems mentioned in the background section.
[0004] To achieve the above objectives, the technical solution adopted by the present invention is as follows: A deep fraud content detection system based on multimodal data fusion includes: The chat analysis module acquires the chat content to be detected, identifies the two parties in the conversation, determines the demander and the demanded party based on the chat content, and analyzes the actual demand information of the demander based on the chat content. The historical analysis module acquires the historical chat records of the requester and the requestee, analyzes the historical request information of the requester based on the historical chat records, and analyzes the memory limit of the requester. The fraud identification module determines whether the requester is suspicious based on historical chat records and memory time limits. If not, no action is taken. If so, based on historical request information, at least one suspicious part of the actual request information is analyzed and verified. The economic loss caused to the requester by the suspicious part is verified, and based on the economic loss, it is determined whether the suspicious part is fraudulent content.
[0005] Preferably, obtaining the chat content to be detected includes the following steps: If the content to be detected is speech, then based on speech recognition, the content to be detected is converted into first text, and the first text is used as the chat content to be detected. If the content to be detected is a video, then the audio of the video is extracted to obtain the target audio. Based on speech recognition, the target audio is converted into second text, and the second text is used as the chat content to be detected. If the content to be detected is text, then the content to be detected will be used as the chat content to be detected.
[0006] Preferably, determining the requester and the requested party in the conversation based on the chat content to be detected includes the following steps:
[0007] The two parties in the chat content to be tested are designated as the first party and the second party, respectively. The portion of the chat content to be tested generated by the first party is designated as the first part, and the portion of the chat content to be tested generated by the second party is designated as the second part.
[0008] Based on big data, at least one type of text requesting money is obtained, and the text requesting money is divided into monetary and non-monetary parts;
[0009] Based on a Chinese database, at least one approximate word is obtained. Using the approximate word, words in the non-monetary part of the request text are replaced. Each replacement method is used as an approximate request text.
[0010] At least one number appearing in the chat content to be detected is used as a feature value. The feature value is used to replace the amount part in the approximate request text to obtain the feature request text.
[0011] The feature request text appearing in the first part is summarized as the first content, and the feature request text appearing in the second part is summarized as the second content;
[0012] If the first dialogue party has more content than the second dialogue party, then the first dialogue party is the requester and the second dialogue party is the requestee; otherwise, the second dialogue party is the requester and the first dialogue party is the requestee.
[0013] Preferably, the step of analyzing the actual needs of the requester based on the chat content to be detected includes the following steps: The portion of the chat content to be detected that is generated by the requester is taken as the target portion, and the feature-requested text that appears in the target portion is taken as the actual request information of the requester.
[0014] Preferably, the analysis to obtain the historical demand information of the demand side includes the following steps: Take at least one number that appears in the historical chat history as the target value, and use the target value to replace the amount part in the approximate request text to obtain the target request text; The portion of historical chat logs generated by the requester will be considered as the historical request portion. The target request text that appears in the historical requirements section will be used as the historical requirements information of the requester.
[0015] Preferably, the analysis to obtain the demander's memory duration includes the following steps: In historical chat logs, obtain at least one feature information, where the feature information is the reply information received from the requester's question; The words in the feature information are replaced, and each replacement method is used as the feature approximation information. The feature approximation information generated by the same feature information is summarized into a feature approximation information set. In historical chat records, the time when the client asks for similar information is used as the feature time. The feature times generated by similar information with the same feature are arranged from smallest to largest to obtain the feature time series. The difference between adjacent feature times in the feature time series is taken to obtain at least one feature difference. The minimum value of the feature differences of all feature information is taken as the memory limit of the demand side.
[0016] Preferably, the step of determining whether the requester is a suspicious person based on historical chat records and memory time limits includes the following steps: The location where the approximate feature information first appears in the chat content to be detected is taken as the target location, and the time when the target location appears is taken as the target time. The last location of the approximate feature information is identified in the historical demand information and used as a reference location. The time when the reference location appears is used as a reference time. If the approximate feature information for generating the reference time and the approximate feature information for generating the target time come from the same set of approximate feature information, then the reference time is mapped to the target time, and the corresponding reference time is subtracted from the target time to obtain the forgetting time. If there is a forgetting time that is less than the memory limit, then the requester is judged to be a suspicious person; otherwise, the requester is judged not to be a suspicious person.
[0017] Preferably, the analysis to obtain at least one suspicious portion of the actual demand information includes the following steps: If the demander agrees to meet the demander's target request in the historical demand information, the maximum value of the monetary portion of the met target request will be used as the baseline value. If the monetary value in the feature request text of the actual demand information of the demander exceeds the benchmark value, the feature request text will be considered a suspicious part.
[0018] Preferably, the analysis and verification of the suspicious portion, and the verification of the economic losses caused to the demander by the suspicious portion, includes the following steps: In the chat content to be tested, obtain the response of the requested party to the suspicious part as the part to be analyzed. The last part of the part to be analyzed is the amount of money that appears at the end, which is considered as the economic loss caused to the requesting party by the suspicious part.
[0019] Preferably, the determination of whether a suspicious portion is fraudulent based on economic loss includes the following steps: If the economic loss is zero, the suspicious part is determined not to be fraudulent; otherwise, the suspicious part is determined to be fraudulent.
[0020] Compared with the prior art, the beneficial effects of the present invention are as follows: By setting up chat analysis, historical analysis, and fraud identification modules, information is generated to verify the identity of the requester based on the historical chat records of the requester and the requestee. This avoids judging the authenticity of AI-generated images and voices based on whether the person's behavior or voice is unnatural. Furthermore, it can more accurately identify AI-generated fake videos, voices, or chat texts based on the requester's habits, and determine the valid fraudulent content based on the actual reaction of the requestee, while removing invalid fraudulent content. This allows for more accurate identification of valid fraudulent content and targeted prevention. Attached Figure Description
[0021] Figure 1 This is a flowchart illustrating the deep fraud content detection system based on multimodal data fusion of the present invention.
[0022] Figure 2 This is a schematic diagram of the process for obtaining the chat content to be detected according to the present invention;
[0023] Figure 3 This is a flowchart illustrating the process of determining the requester and the requested party in a conversation based on the chat content to be detected, according to the present invention.
[0024] Figure 4 This is a schematic diagram illustrating the process of obtaining historical demand information from the demand side in the analysis of this invention;
[0025] Figure 5 This is a schematic diagram illustrating the process of obtaining the demander's memory through analysis according to the present invention;
[0026] Figure 6 This is a schematic diagram of the process of determining whether a requester is a suspicious person based on historical chat records and memory time limits according to the present invention.
[0027] Figure 7 This is a schematic diagram of the process by which the analysis of the present invention yields at least one suspicious portion of the actual demand information. Detailed Implementation
[0028] The following description is intended to disclose the invention and enable those skilled in the art to implement it. The preferred embodiments described below are merely examples, and other obvious variations will occur to those skilled in the art.
[0029] Reference Figure 1 As shown, a deep fraud content detection system based on multimodal data fusion includes: The chat analysis module acquires the chat content to be detected, identifies the two parties in the conversation, determines the demander and the demanded party based on the chat content, and analyzes the actual demand information of the demander based on the chat content. The historical analysis module acquires the historical chat records of the requester and the requestee, analyzes the historical request information of the requester based on the historical chat records, and analyzes the memory limit of the requester. The fraud identification module determines whether the requester is suspicious based on historical chat records and memory time limits. If not, no action is taken. If so, based on historical request information, at least one suspicious part of the actual request information is analyzed and verified. The economic loss caused to the requester by the suspicious part is verified, and based on the economic loss, it is determined whether the suspicious part is fraudulent content.
[0030] Deepfake scams are generated by AI. They typically impersonate acquaintances and use deceptive tactics to trick victims into transferring money, thus achieving their fraudulent goals. Previously, due to the imperfections of AI technology, the videos or audio generated by this technology might appear unnatural in some cases. However, with the development of AI technology, these issues have become increasingly difficult to detect, even by the naked eye or computer. Therefore, to achieve accurate identification, it is necessary to approach the problem from a different angle, avoiding the angles that AI technology excels at and identifying from angles that it does not currently prioritize. In this solution, a series of steps are set up to address this issue.
[0031] Reference Figure 2 As shown, obtaining the chat content to be detected includes the following steps: If the content to be detected is speech, then based on speech recognition, the content to be detected is converted into first text, and the first text is used as the chat content to be detected. If the content to be detected is a video, then the audio of the video is extracted to obtain the target audio. Based on speech recognition, the target audio is converted into second text, and the second text is used as the chat content to be detected. If the content to be detected is text, then the content to be detected will be used as the chat content to be detected.
[0032] The content to be detected comes in various forms. For convenience, all of it is converted into text format during detection. Existing speech recognition technology can accomplish this.
[0033] Reference Figure 3 As shown, determining the requester and the requestee in a conversation based on the content of the chat to be detected includes the following steps: The two parties in the chat content to be tested are designated as the first party and the second party, respectively. The portion of the chat content to be tested generated by the first party is designated as the first part, and the portion of the chat content to be tested generated by the second party is designated as the second part. Based on big data, at least one type of text requesting money is obtained, and the text requesting money is divided into monetary and non-monetary parts; Based on a Chinese database, at least one approximate word is obtained. Using the approximate word, words in the non-monetary part of the request text are replaced. Each replacement method is used as an approximate request text. At least one number appearing in the chat content to be detected is used as a feature value. The feature value is used to replace the amount part in the approximate request text to obtain the feature request text. The feature request text appearing in the first part is summarized as the first content, and the feature request text appearing in the second part is summarized as the second content; If the first dialogue party has more content than the second dialogue party, then the first dialogue party is the requester and the second dialogue party is the requestee; otherwise, the second dialogue party is the requester and the first dialogue party is the requestee.
[0034] When two people are chatting, it's necessary to identify the requester and the recipient. The requester is the one seeking money, and the recipient is the one transferring the money. If the requester is not a real person but an AI-generated acquaintance, then the transfer constitutes fraud. Therefore, it's crucial to determine if the requester is a real person. This requires identifying the requester and recipient in the chat content to be tested. Requesters typically borrow money, thus demanding payment, generating request text. However, the amount requested is not fixed, and non-amount portions can be expressed in various ways. Therefore, based on the actual situation, at least one characteristic request text is generated. This characteristic request text can then be compared to identify the requester, as the requester will demand money, while the recipient will not.
[0035] Based on the chat content to be detected, the actual needs of the requester are analyzed and obtained through the following steps: The portion of the chat content to be detected that is generated by the requester is taken as the target portion, and the feature-requested text that appears in the target portion is taken as the actual request information of the requester.
[0036] In order to verify the identity of the requester, that is, to determine whether they are an AI-generated fake, it is necessary to obtain their current chat content, that is, the actual requester's information.
[0037] Reference Figure 4 As shown, the analysis to obtain historical demand information from the demand side includes the following steps: Take at least one number that appears in the historical chat history as the target value, and use the target value to replace the amount part in the approximate request text to obtain the target request text; The portion of historical chat logs generated by the requester will be considered as the historical request portion. The target request text that appears in the historical requirements section will be used as the historical requirements information of the requester.
[0038] When verifying the identity of the requester, it is necessary to refer to the requester's historical request information. Therefore, it is also necessary to extract the required content from the historical chat records of the requester and the requested party. That is, it is necessary to obtain the target request text generated by the requester in the historical chat records. In this way, based on their historical request habits, possible anomalies in the current request can be identified.
[0039] Reference Figure 5 As shown, the analysis of the demander's memory duration includes the following steps: In historical chat logs, obtain at least one feature information, where the feature information is the reply information received from the requester's question; The words in the feature information are replaced, and each replacement method is used as the feature approximation information. The feature approximation information generated by the same feature information is summarized into a feature approximation information set. In historical chat records, the time when the client asks for similar information is used as the feature time. The feature times generated by similar information with the same feature are arranged from smallest to largest to obtain the feature time series. The difference between adjacent feature times in the feature time series is taken to obtain at least one feature difference. The minimum value of the feature differences of all feature information is taken as the memory limit of the demand side.
[0040] When identifying whether a requester is suspicious, the primary method is to assess their ability to recall information. A person's memory capacity is generally fixed, and the duration of their recall for certain information is also relatively fixed. When they forget, they will ask again to retrieve the relevant information. Therefore, by identifying characteristic information in historical chat logs, the time it takes to retrieve this information after it has been forgotten can be determined. This allows us to determine the recall time of the information. Since the time of each message in the chat log is fixed, the time of the characteristic information can be determined. However, it's important to note that approximation is necessary during characteristic information identification, as the information may be expressed in different ways. Therefore, approximate characteristic information is used for identification, resulting in at least one characteristic difference. The minimum value of this characteristic difference represents the requester's recall time. Typically, their recall time for known information will not be shorter than this recall time. Therefore, if their recall time is shorter than this recall time, it is highly likely that they are an AI-generated fake acquaintance, requiring further analysis of the chat content.
[0041] Reference Figure 6 As shown, determining whether a requester is suspicious, based on historical chat logs and memory time limits, includes the following steps: The location where the approximate feature information first appears in the chat content to be detected is taken as the target location, and the time when the target location appears is taken as the target time. The last location of the approximate feature information is identified in the historical demand information and used as a reference location. The time when the reference location appears is used as a reference time.
[0042] If the approximate feature information for generating the reference time and the approximate feature information for generating the target time come from the same set of approximate feature information, then the reference time is mapped to the target time, and the corresponding reference time is subtracted from the target time to obtain the forgetting time. If there is a forgetting time that is less than the memory limit, then the requester is judged to be a suspicious person; otherwise, the requester is judged not to be a suspicious person.
[0043] Reference Figure 7 As shown, analyzing at least one suspicious portion of the actual demand information includes the following steps: If the demander agrees to meet the demander's target request in the historical demand information, the maximum value of the monetary portion of the met target request will be used as the baseline value. If the monetary value in the feature request text of the actual demand information of the demander exceeds the benchmark value, the feature request text will be considered a suspicious part.
[0044] Based on historical information, there is an upper limit to the amount a borrower can lend. Therefore, this limit can be used to identify suspicious aspects of their behavior and to determine whether it constitutes fraud. When one party borrows money, it is necessary to determine how much the other party will lend in order to ascertain the amount of fraud. If the other party does not lend money, it indicates that they can identify the fraudulent content, and thus the fraud is deemed invalid and no further intervention is needed. However, if the borrower is suspicious and their suspicious behavior can obtain money, then they are considered to have a high risk of fraud, and therefore, their behavior should be prevented.
[0045] The analysis and verification of the suspicious parts, and the determination of the economic losses caused to the demand side by the suspicious parts, include the following steps: In the chat content to be tested, obtain the response of the requested party to the suspicious part as the part to be analyzed. The last part of the part to be analyzed is the amount of money that appears at the end, which is considered as the economic loss caused to the requesting party by the suspicious part.
[0046] In the dialogue between the requester and the recipient, the amount of money borrowed will constantly change because even if both are real people, when one party borrows money, the other party may not lend the full amount. Therefore, it is necessary to determine the final amount lent.
[0047] Based on the economic loss, determining whether a suspicious portion is fraudulent involves the following steps: If the economic loss is zero, the suspicious part is determined not to be fraudulent; otherwise, the suspicious part is determined to be fraudulent.
[0048] Furthermore, this solution also proposes a storage medium on which a computer-readable program is stored, which, when invoked, executes the aforementioned deep fraud content detection system based on multimodal data fusion.
[0049] It is understandable that the storage medium can be a magnetic medium, such as a floppy disk, hard disk, or magnetic tape; an optical medium, such as a DVD; or a semiconductor medium, such as a solid-state drive (SSD).
[0050] In summary, the advantages of this invention are as follows: by setting up a chat analysis module, a history analysis module, and a fraud identification module, information for verifying the identity of the requester is generated based on the historical chat records of the requester and the requestee. This avoids relying on unnatural behavior or voice quality to determine the authenticity of AI-generated images and voices. Furthermore, based on the requester's habits, it can more accurately identify AI-generated fake videos, voice recordings, or chat texts. Based on the actual reaction of the requestee, it determines valid fraudulent content and removes invalid fraudulent content, thus enabling more precise identification of valid fraudulent content and targeted prevention.
[0051] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the claimed invention. The scope of protection claimed by the appended claims and their equivalents is defined.
Claims
1. A deep fraud content detection system based on multimodal data fusion, characterized in that, include: The chat analysis module acquires the chat content to be detected, identifies the two parties in the conversation, determines the demander and the demanded party based on the chat content, and analyzes the actual demand information of the demander based on the chat content. The historical analysis module acquires the historical chat records of the requester and the requestee, analyzes the historical request information of the requester based on the historical chat records, and analyzes the memory limit of the requester. The fraud identification module determines whether the requester is suspicious based on historical chat records and memory time limits. If not, no action is taken. If so, based on historical request information, at least one suspicious part of the actual request information is analyzed and verified. The economic loss caused to the requester by the suspicious part is verified, and based on the economic loss, it is determined whether the suspicious part is fraudulent content.
2. The deep fraud content detection system based on multimodal data fusion according to claim 1, characterized in that, The process of obtaining the chat content to be detected includes the following steps: If the content to be detected is speech, then based on speech recognition, the content to be detected is converted into first text, and the first text is used as the chat content to be detected. If the content to be detected is a video, then the audio of the video is extracted to obtain the target audio. Based on speech recognition, the target audio is converted into second text, and the second text is used as the chat content to be detected. If the content to be detected is text, then the content to be detected will be used as the chat content to be detected.
3. The deep fraud content detection system based on multimodal data fusion according to claim 2, characterized in that, The process of determining the requester and the requested party in a conversation based on the chat content to be detected includes the following steps: The two parties in the chat content to be tested are designated as the first party and the second party, respectively. The portion of the chat content to be tested generated by the first party is designated as the first part, and the portion of the chat content to be tested generated by the second party is designated as the second part. Based on big data, at least one type of text requesting money is obtained, and the text requesting money is divided into monetary and non-monetary parts; Based on a Chinese database, at least one approximate word is obtained. Using the approximate word, words in the non-monetary part of the request text are replaced. Each replacement method is used as an approximate request text. At least one number appearing in the chat content to be detected is used as a feature value. The feature value is used to replace the amount part in the approximate request text to obtain the feature request text. The feature request text appearing in the first part is summarized as the first content, and the feature request text appearing in the second part is summarized as the second content; If the first dialogue party has more content than the second dialogue party, then the first dialogue party is the requester and the second dialogue party is the requestee; otherwise, the second dialogue party is the requester and the first dialogue party is the requestee.
4. The deep fraud content detection system based on multimodal data fusion according to claim 3, characterized in that, The process of analyzing the actual needs of the requester based on the chat content to be detected includes the following steps: The portion of the chat content to be detected that is generated by the requester is taken as the target portion, and the feature-requested text that appears in the target portion is taken as the actual request information of the requester.
5. The deep fraud content detection system based on multimodal data fusion according to claim 4, characterized in that, The analysis to obtain historical demand information from the demand side includes the following steps: Take at least one number that appears in the historical chat history as the target value, and use the target value to replace the amount part in the approximate request text to obtain the target request text; The portion of historical chat logs generated by the requester will be considered as the historical request portion. The target request text that appears in the historical requirements section will be used as the historical requirements information of the requester.
6. The deep fraud content detection system based on multimodal data fusion according to claim 5, characterized in that, The analysis to determine the demander's memory duration includes the following steps: In historical chat logs, obtain at least one feature information, where the feature information is the reply information received from the requester's question; The words in the feature information are replaced, and each replacement method is used as the feature approximation information. The feature approximation information generated by the same feature information is summarized into a feature approximation information set. In historical chat records, the time when the client asks for similar information is used as the feature time. The feature times generated by similar information with the same feature are arranged from smallest to largest to obtain the feature time series. The difference between adjacent feature times in the feature time series is taken to obtain at least one feature difference. The minimum value of the feature differences of all feature information is taken as the memory limit of the demand side.
7. The deep fraud content detection system based on multimodal data fusion according to claim 6, characterized in that, The process of determining whether a requester is suspicious based on historical chat records and memory time limits includes the following steps: The location where the approximate feature information first appears in the chat content to be detected is taken as the target location, and the time when the target location appears is taken as the target time. The last location of the approximate feature information is identified in the historical demand information and used as a reference location. The time when the reference location appears is used as a reference time. If the approximate feature information for generating the reference time and the approximate feature information for generating the target time come from the same set of approximate feature information, then the reference time is mapped to the target time, and the corresponding reference time is subtracted from the target time to obtain the forgetting time. If there is a forgetting time that is less than the memory limit, then the requester is judged to be a suspicious person; otherwise, the requester is judged not to be a suspicious person.
8. A deep fraud content detection system based on multimodal data fusion according to claim 7, characterized in that, The analysis, which yields at least one suspicious portion of the actual demand information, includes the following steps: If the demander agrees to meet the demander's target request in the historical demand information, the maximum value of the monetary portion of the met target request will be used as the baseline value. If the monetary value in the feature request text of the actual demand information of the demander exceeds the benchmark value, the feature request text will be considered a suspicious part.
9. A deep fraud content detection system based on multimodal data fusion according to claim 8, characterized in that, The analysis and verification of the suspicious parts, and the verification of the economic losses caused to the demand side by the suspicious parts, includes the following steps: In the chat content to be tested, obtain the response of the requested party to the suspicious part as the part to be analyzed. The last part of the part to be analyzed is the amount of money that appears at the end, which is considered as the economic loss caused to the requesting party by the suspicious part.
10. A deep fraud content detection system based on multimodal data fusion according to claim 9, characterized in that, The determination of whether a suspicious portion is fraudulent based on economic loss includes the following steps: If the economic loss is zero, the suspicious part is determined not to be fraudulent; otherwise, the suspicious part is determined to be fraudulent.
Citation Information
Patent Citations
Information verification method and electronic equipment
CN107508834A
Prompting method and device, and computer storage medium
CN109547651A
Risk prediction model training method and device
CN112200382A
User chat anti-fraud automatic judgment method and system
CN115203695A
Fraud information identification method and device, electronic equipment and storage medium
CN118820829A