Illegal activity participant intelligent identification method based on deep context analysis

Through deep context analysis and large-scale language model fine-tuning, combined with a blackword knowledge base and chat content standardization, the problems of implicit communication and context understanding in the identification of illegal online activities are solved, and efficient and accurate identification of illegal activity participants is achieved.

CN120632095APending Publication Date: 2025-09-12ZHEJIANG POLICE COLLEGE +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510657605.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing technologies are unable to effectively identify the implicit communications of participants in illegal online activities, are unable to conduct in-depth contextual understanding, static rules are difficult to adapt to dynamic changes, have a single analysis dimension, and have a high false alarm rate, and are unable to fully reflect the true intentions and behavioral habits of illegal actors.

Method used

Adopting a method based on deep context analysis, using large language models for fine-tuning and data correction, we build a real-time dynamic blackword knowledge base. Combined with chat content standardization and feature extraction, we conduct multi-dimensional analysis and summary of long conversations, build a complete event chain, and deeply analyze context-related information.

Benefits of technology

Accurately capture relevant information involved in the case and deeply analyze context-related information clues, which improves the accuracy and adaptability of illegal activity identification, reduces the false alarm rate, and comprehensively reflects the behavioral habits and intentions of illegal actors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120632095A_ABST
    Figure CN120632095A_ABST
Patent Text Reader

Abstract

An illegal activity participant intelligent identification method based on deep context analysis comprises the following steps: 1) fine tuning based on a language large model: 1.1) carrying out dialogue multi-dimensional analysis and summarization by using an online large teacher model, 1.2) carrying out manual verification and data correction, 1.3) constructing an instruction fine tuning SFT, 1.4) carrying out instruction fine tuning on a student model, 1.5) constructing an RAG black word knowledge base; 2) chat content analysis agent construction: 2.1) chat file format standardization, and 2.2) feature information extraction tools; 2.3) structuring the unstructured data; 2.4) cleaning the fund data; 2.5) standardizing message content; 2.6) integrating the fine tuning and enlarging model in the first stage; and 3) comprehensively analyzing the dialogue context. According to the method, case-related information is accurately captured, context-related information clues are deeply analyzed, and illegal activity participants are mined.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of network security and relates to a method for intelligently identifying participants in illegal network activities. Background Art

[0002] With the continuous development of technology, participants in online illegal activities can use more and more different social media platforms to conduct illegal activities. Currently, the conversation context recognition technology for participants in illegal activities mainly includes:

[0003] Based on keyword or regular expression matching, for example, a database of sensitive terms involved in a case is set up to match conversation records. However, this method has limited coverage and cannot handle sensitive word variations, new code words, typos, etc. It also ignores contextual information, resulting in a high false positive rate.

[0004] Based on traditional machine learning methods: using support vector machines (SVM), naive Bayes, and other methods, combined with bag-of-words models and TF-IDF feature engineering to classify conversation fragments and determine whether they are suspected of being illegal. However, feature engineering relies too much on expert experience and has difficulty capturing the deep semantics of long conversations. It also has weak generalization capabilities. If very complex and diverse conversations or other records appear, the accuracy will be greatly reduced.

[0005] Directly applying a general-purpose large language model: This approach uses a general-purpose large model for conversation and behavior analysis without specific data training, fine-tuning, or optimization. Because the model lacks specific domain knowledge about illegal activities and doesn't understand cryptic expressions or codewords, despite its high generality, large-scale models can still experience significant hallucinations and misinterpret the meaning of some illicit terms. General-purpose models are typically large, making direct inference expensive.

[0006] Existing technologies for analyzing offenders' activities are often relatively simple (for example, relying solely on simple chat content keywords). This has the following problems:

[0007] 1) It is difficult to deal with the obscure language used by criminals: Participants in illegal activities often use very obscure codes, black words, and code names to circumvent them. Methods based on traditional keywords or simple rules can easily fail.

[0008] 2) Inability to perform deep contextual understanding: The recognition of a single message or a small number of messages is often unclear. Long-term (e.g., days or months) conversation history and context are needed to more accurately determine the intentions and relationships of illegal actors.

[0009] 3) Static rules are difficult to adapt to dynamic changes: The language and patterns of communication in illegal activities are constantly evolving and updating, and fixed keywords or rules are difficult to maintain long-term effectiveness.

[0010] 4) Single analysis dimension: Relying solely on keyword analysis of conversation content, the information summarized is incomplete and can be easily circumvented by anti-detection actions of illegal actors. It cannot fully reflect the true intentions and behavioral habits of illegal activity participants. Summary of the Invention

[0011] In order to overcome the shortcomings of existing technologies, the present invention provides an intelligent identification method for illegal activity participants based on deep context analysis. It uses artificial intelligence technology to efficiently analyze the context of users' online conversations, accurately capture relevant information involved in the case, deeply analyze context-related information clues, and explore illegal activity participants. It provides an artificial intelligence solution for automatically identifying, analyzing, summarizing, and automatically recommending information clues involved in the case for the conversation information of illegal activity participants.

[0012] The technical solution adopted by the present invention to solve its technical problem is:

[0013] A method for intelligently identifying participants in illegal activities based on deep context analysis, comprising the following steps:

[0014] 1) Fine-tuning based on the large language model:

[0015] 1.1) Using a large online "teacher" model to analyze and summarize conversations from multiple dimensions,

[0016] 1.2) Manual verification and data correction,

[0017] 1.3) Construct the instruction fine-tuning SFT. Pair the original conversation records with the analysis report after manual data correction to construct an instruction fine-tuning dataset that conforms to the large model fine-tuning format. Each sample typically contains:

[0018] Commands: General analysis and summary prompts to extract multi-dimensional information from conversation records;

[0019] Input: original conversation record text;

[0020] Output: Structured analysis summary after manual data correction;

[0021] All processed conversations are combined in the above format, aggregated in JSON format and form the SFT dataset for fine-tuning training;

[0022] 1.4) Fine-tune the instructions for the "student" model,

[0023] Use the SFT dataset constructed in 1.3) to fine-tune the "student" model, allowing it to learn to imitate the high-quality analysis summary after manual correction, generate structured data, and use fine-tuning techniques to simulate parameter changes through low-rank decomposition. Set the learning rate, training cycle, and batch size for fine-tuning training;

[0024] 1.5) Build RAG blackword knowledge base,

[0025] Establish and maintain a real-time, dynamically updated Milvus-based blackword knowledge base to store obscure words and phrases and their corresponding case types;

[0026] 2) Chat content analysis agent construction process is as follows:

[0027] 2.1) Chat file format standardization,

[0028] 2.2) Feature Information Extraction Tools: Build a series of tools to extract and classify feature chat content;

[0029] 2.3) Structuring unstructured data;

[0030] 2.4) Funds data cleaning;

[0031] 2.5) Standardize message content. Unify chat data into a text-structured format and store chat records in JSON format. Fields include timestamp, recipient, chat content, and message category. Regular expressions or JSONSchema techniques are used to ensure data format consistency.

[0032] 2.6) Integrate the fine-tuned large model from the first stage

[0033] After the user inputs the chat log file, the chat file format is verified, the content is standardized, the message type is classified, and the unstructured data is structured. This results in chat conversation data in a standard format that is more in line with the input format accepted by the fine-tuned large model.

[0034] 3) Comprehensive analysis of the conversation context. The process is as follows:

[0035] For long conversation records received, the text block size is set and the long conversation is split, while retaining the overlapping parts between adjacent blocks to ensure the accuracy of contextual information at the boundaries of the text blocks; after using the fine-tuned model to analyze the conversation block by text block, multiple summaries are combined in chronological order of the conversation, and prompt words are designed to refine multi-block summaries. The large model allows for a comprehensive summary of the contents of multiple sequentially arranged summaries, fully linking the contextual events and clue information of the long conversation to build a complete event chain. At the same time, the emotional changes of different text blocks are analyzed to better understand the emotional direction of the conversation and better associate it with actual evidence of illegal activities.

[0036] The beneficial effects of the present invention are mainly manifested in: accurately capturing relevant information involved in the case, deeply analyzing context-related information clues, and discovering participants in illegal activities. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 This is a block diagram of the principle of an intelligent identification method for illegal activity participants based on deep context analysis. DETAILED DESCRIPTION

[0038] The present invention will be further described below with reference to the accompanying drawings.

[0039] Reference Figure 1 , an intelligent identification method for illegal activity participants based on deep context analysis, comprising the following steps:

[0040] 1) Fine-tuning based on the large language model:

[0041] 1.1) Using a large online "teacher" model to analyze and summarize conversations from multiple dimensions,

[0042] First, we prepare a large-scale original conversation dataset. This dataset is derived from real-world cases and from mobile devices such as the phones of participants in real illegal activities. This data will serve as the input for the subsequent fine-tuning of the training model.

[0043] Subsequently, we carefully designed multi-dimensional analysis prompts to meet the needs of illegal activity identification. These prompts summarize the conversation's brief topic, full topic, case type (e.g., suspected pornography, gambling, drug-related, not involved), and the relationship between the two parties in a fixed format. At the same time, we extract various types of clues from the original conversation, including time, location, person, price, and blacklisted words. These prompts guide a large "teacher" model (such as Qwen2.5-72B or equivalent) to compile deeply analyzed clues. We chose large models like the 72B model primarily because they excel at understanding complex contexts, following diverse instructions, and performing complex reasoning. They can generate high-quality preliminary analysis summaries without any targeted fine-tuning.

[0044] 1.2) Manual verification and data correction,

[0045] The analysis summary generated by the "teacher" model in 1.1) is submitted to a trained or experienced reviewer. The reviewer carefully checks the information of each dimension by comparing the original conversation with the analysis summary, especially focusing on verifying the accuracy of the black word clue information. If the generated analysis is completely accurate, it is directly adopted and recorded in the database. If the analysis contains certain errors or omissions (such as errors in the case type involved, missing key case information, etc.), the reviewer will manually make detailed corrections to ensure that the final analysis summary is completely accurate and consistent with the domain knowledge information.

[0046] 1.3) Supervised Fine-Tuning (SFT)

[0047] The original conversation records are paired with the analysis reports after manual data correction to construct an instruction fine-tuning dataset that conforms to the large model fine-tuning format. Each sample typically contains:

[0048] Instruction: A general analysis and summary prompt to extract multi-dimensional information from conversation records;

[0049] Input: original conversation record text;

[0050] Output: Structured analysis summary after manual data correction;

[0051] All processed conversations are combined in the above format, aggregated in JSON format and form the SFT dataset for fine-tuning training;

[0052] 1.4) Fine-tune the instructions for the "student" model,

[0053] Select a "student" model (such as Qwen-7B) with low computing resource requirements and suitable for deployment, and use the SFT dataset constructed in 1.3) to fine-tune the "student" model to learn to imitate the high-quality analysis summary after manual correction, generate structured data, and use efficient fine-tuning technology (Low-Rank Adaptation, LoRA) to simulate the change of parameters through low-rank decomposition, thereby achieving indirect training of large models with extremely small parameters. Set appropriate learning rate, training cycle and batch size for fine-tuning training.

[0054] 1.5) Build RAG blackword knowledge base,

[0055] To ensure that the fine-tuned model can better utilize the latest and most accurate knowledge of the case-related fields during the reasoning process, a real-time and dynamically updated Milvus-based blackword knowledge base is established and maintained to store obscure words, code words, and other obscure vocabulary and sentences and their corresponding types of cases (such as points -> online gambling, methamphetamine -> drug trafficking). This module deeply integrates the Retrieval Enhancement Generation (RAG) module implemented based on the Llama-Index framework. According to the input dialogue context, it efficiently retrieves relevant information from the blackword knowledge base and injects these retrieval information corresponding to the case type markers into the instructions, thereby guiding the fine-tuned model to make more accurate judgments and greatly improving the determination of the case type. This module uses embedding models and vector indexes for semantic retrieval. Even if the words are not exactly the same or the blackword variants are not identical, it can effectively discover blackwords that are related to the input context in meaning, and assist in determining the case type.

[0056] 2) Chat content analysis agent construction process is as follows:

[0057] 2.1) Chat file format standardization,

[0058] Due to the variety of chat file formats received, it is necessary to perform input chat file format verification and format conversion. The format is converted to predefined specifications to facilitate Agents to understand and process. At the same time, the integrity of the data (whether it contains fields such as timestamp, sender, and message content) is verified. If the verification fails, the user is prompted with missing data, but the program operation is not blocked.

[0059] 2.2) Feature information extraction tools,

[0060] A series of tools were built to extract and classify chat content, as shown in Table 1. For example, transfers / red envelopes, emoticons, images, audio, video, IP / domain names, URLs, system information, and sharing and forwarding are distinguished from text information. This facilitates feature extraction and identification of specific content that may be helpful in solving cases, such as special transaction amounts, images / audio / video containing special information, and websites involved in the case.

[0061] Table 1 shows the feature information extraction tools:

[0062]

[0063] 2.3) Structuring unstructured data,

[0064] Convert non-text information (images, audio) into parseable text. For images, use the OCR engine to extract the text in the image, use the ResNet image classification model to identify the image subject, and merge the results to fill the content field. For speech, use a speech recognition tool to transcribe the text and write it to the content field.

[0065] 2.4) Fund data cleaning,

[0066] Funding data must be relatively accurate. A transfer or red envelope transaction requires both sending and receiving information, or only receiving information. Red envelopes in group chats may involve multiple redemptions. After processing, data is stored in a MySQL database for subsequent analysis.

[0067] 2.5) Standardization of message content,

[0068] After feature extraction and structuring of unstructured data, the chat data needs to be unified into a text structure format (e.g., automatic classification, labeling, and text extraction of unstructured data such as images, voice, and video) to facilitate subsequent model understanding. Chat records are usually stored in JSON format, with fields including timestamp, recipient, chat content, message category, etc. Regular expressions or JSON Schema technology are applied to ensure data format consistency.

[0069] 2.6) Integrate the fine-tuned large model from the first stage

[0070] After the user inputs the chat record file, the chat file format verification, content standardization, message type classification,

[0071] During the unstructured data structuring stage, chat conversation data in a standard format is obtained, which is more in line with the input format accepted by the fine-tuned large model, so that the large model can more intuitively and easily understand the conversation context and ensure the accuracy of conversation analysis.

[0072] 3) Comprehensive analysis of the conversation context. The process is as follows:

[0073] For the long conversation records received, considering that the input length that the model can accept is limited, an appropriate text block size is set to split the long conversation, while retaining the overlapping parts between adjacent blocks to ensure the accuracy of the context information at the text block boundary. After using the fine-tuned model to analyze the conversation block by text block, multiple summaries are combined in chronological order of the conversation, and prompt words are designed specifically for refining multi-block summaries. This allows the large model to comprehensively summarize the contents of multiple sequentially arranged summaries, fully associate the contextual events and clue information of the long conversation, and build a complete event chain. At the same time, the emotional changes of different text blocks are analyzed to better understand the emotional direction of the conversation and better associate it with actual evidence of illegal activities.

[0074] The first stage of local summarization of each text block can focus more on the retrieval, extraction and analysis of black words, accurately capture relevant information related to the case, and at the same time extract details such as information about the people involved, address information, and money information in a more granular manner; the second stage of global summary can focus more on the integration of contextual information, and can more accurately extract key clues and conversation topics of conversations in different time periods, analyze the emotional trends and case development trends of different time periods, so as to deeply analyze the relevance of contextual information and more accurately identify information about participants in illegal activities.

[0075] The solution of this embodiment uses artificial intelligence technology to efficiently analyze the context of users' online conversations, accurately capture relevant information involved in the case, deeply analyze context-related information clues, and explore intelligent identification methods for participants in illegal activities, providing an artificial intelligence solution for automatically identifying, analyzing, summarizing, and automatically recommending information clues involved in the case for conversation information of illegal activity participants.

[0076] The embodiments of this specification are merely examples of implementations of the invention and are provided for illustrative purposes only. The scope of protection of the present invention should not be considered limited to the specific embodiments described in these embodiments. The scope of protection of the present invention also extends to equivalent technical means that can be conceived by a person of ordinary skill in the art based on the invention.

Claims

1. A method for intelligently identifying participants in illegal activities based on deep context analysis, characterized in that: The method comprises the following steps: 1) Fine-tuning based on the large language model: 1.1) Using a large online "teacher" model to conduct multi-dimensional analysis and summary of conversations; 1.2) Manual verification and data correction; 1.3) Construct the instruction fine-tuning SFT. Pair the original conversation records with the analysis report after manual data correction to construct an instruction fine-tuning dataset that conforms to the large model fine-tuning format. Each sample typically contains: Commands: General analysis and summary prompts to extract multi-dimensional information from conversation records; Input: original conversation record text; Output: Structured analysis summary after manual data correction; All processed conversations are combined in the above format, aggregated in JSON format and form the SFT dataset for fine-tuning training; 1.4) Fine-tune the instructions for the "student" model, Use the SFT dataset constructed in 1.3) to fine-tune the "student" model, allowing it to learn to imitate the high-quality analysis and summary after manual correction, generate structured data, and use fine-tuning techniques to simulate parameter changes through low-rank decomposition. Set the learning rate, training cycle, and batch size for fine-tuning training; 1.5) Build RAG blackword knowledge base, Establish and maintain a real-time, dynamically updated Milvus-based blackword knowledge base to store obscure words and phrases and their corresponding case types; 2) Chat content analysis agent construction process is as follows: 2.1) Chat file format standardization, 2.2) Feature Information Extraction Tools: Build a series of tools to extract and classify feature chat content; 2.3) Structuring unstructured data; 2.4) Funds data cleaning; 2.5) Standardize message content. Chat data is unified into a text-structured format. Chat records are stored in JSON format, with fields including timestamp, recipient, chat content, and message category. Regular expressions or JSON Schema techniques are used to ensure data format consistency. 2.6) Integrate the fine-tuned large model from the first stage After the user inputs the chat log file, the chat file format is verified, the content is standardized, the message type is classified, and the unstructured data is structured. This results in chat conversation data in a standard format that is more in line with the input format accepted by the fine-tuned large model. 3) Comprehensive analysis of the conversation context. The process is as follows: For long conversation records received, the text block size is set and the long conversation is split, while retaining the overlapping parts between adjacent blocks to ensure the accuracy of contextual information at the boundaries of the text blocks; after using the fine-tuned model to analyze the conversation block by text block, multiple summaries are combined in chronological order of the conversation, and prompt words are designed to refine multi-block summaries. The large model allows for a comprehensive summary of the contents of multiple sequentially arranged summaries, fully linking the contextual events and clue information of the long conversation to build a complete event chain. At the same time, the emotional changes of different text blocks are analyzed to better understand the emotional direction of the conversation and better associate it with actual evidence of illegal activities.

2. The method for intelligently identifying participants in illegal activities based on deep context analysis according to claim 1, characterized in that: In 1.1) above, we first prepare a large-scale original conversation dataset. This dataset is derived from real-world cases and original conversation data extracted from the mobile devices of participants in real illegal activities. This data will serve as the input for the training dataset required for subsequent fine-tuning of the training model. Subsequently, multi-dimensional analysis prompts were designed to meet the needs of illegal activity identification. The short topic, complete topic, type of case involved, and relationship between the two parties in the conversation were summarized in a fixed format. At the same time, various types of clue information contained in the original conversation were extracted, including time clue information, location clue information, character clue information, price clue information or black word clue information. These prompts guided the large-scale "teacher" model to sort out the clue information after in-depth analysis.

3. The method for intelligently identifying participants in illegal activities based on deep context analysis according to claim 2, characterized in that: In 1.2), the analysis summary generated by the "teacher" model in 1.1) is submitted to a trained or experienced reviewer. The reviewer carefully checks the information of each dimension by comparing the original conversation with the analysis summary, and verifies the accuracy of the black word clue information. If the generated analysis is completely accurate, it is directly adopted and recorded in the database. If there are errors or omissions in the analysis, the reviewer will perform manual and detailed corrections to ensure that the final analysis summary is completely accurate and consistent with the domain knowledge information.

4. The method for intelligently identifying participants in illegal activities based on deep context analysis according to claim 3, characterized in that: In the above 1.5), a retrieval enhancement generation RAG module based on the Llama-Index framework is integrated. According to the input dialogue context, relevant information is retrieved from the black word knowledge base, and these retrieval information corresponding to the case type mark are injected into the instruction, thereby guiding the fine-tuned model to make more accurate judgments, greatly improving the determination of the case type; semantic retrieval is performed using the embedding model and vector index. Even if the words are not exactly the same or the variants of the black words, the black words that are related to the input context in meaning can be effectively discovered to assist in determining the case type.

5. The method for intelligently identifying participants in illegal activities based on deep context analysis according to any one of claims 1 to 4, characterized in that: In the above 2.1), the input chat file format is verified and converted to a predefined format to facilitate Agents to understand and process it. At the same time, the integrity of the data is verified. If the verification fails, the user is prompted that the data is missing, but the program operation is not blocked.

6. The method for intelligently identifying participants in illegal activities based on deep context analysis according to any one of claims 1 to 4, characterized in that: In 2.3), the non-text information is converted into parsable text information. If it is an image, the OCR engine is used to extract the text content in the image, and the ResNet image classification model is used to identify the image subject. The results are combined to supplement the content field. For speech, the speech recognition tool is used to transcribe the text and write it into the content field.

7. The method for intelligently identifying participants in illegal activities based on deep context analysis according to any one of claims 1 to 4, characterized in that: In 2.4), the financial data must be relatively accurate. Transfers or red envelopes require both sending and receiving information, or only receiving information can constitute a financial transaction. In group chat messages, red envelopes involve multiple redemptions. After processing, they are stored in the MySQL database for subsequent analysis.