LLM-based electronic forensics data analysis method and system

By combining convolutional neural network decision tree and large language model to deeply process electronic forensic data and using vector databases for mixed retrieval, the complex and time-consuming problem of traditional electronic forensic data analysis process is solved, and efficient and accurate data analysis and report generation are achieved.

CN120067906APending Publication Date: 2025-05-30XIAMEN MEIYABAIKE INFORMATION SECURITY RES INST CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510063558.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-12-27
Filing Date
2025-01-15
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The traditional electronic evidence forensics analysis process is complex, time-consuming and easy to lead to the omission of key information, making it difficult to efficiently process massive data.

Method used

The method of fusion convolutional neural network decision tree and large language model (LLM) is used to pre-process electronic forensic data, data cleaning, text information extraction and keyword extraction, and mixed searches are combined with vector databases and specialized databases to finally generate intelligent reports.

Benefits of technology

The data analysis process of case handlers has been simplified, the data processing speed and analysis efficiency have been significantly improved, and the precise extraction and positioning ability of the data involved in the case has been enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067906A_ABST
    Figure CN120067906A_ABST
Patent Text Reader

Abstract

The invention discloses an LLM-based electronic forensics data analysis method and system. The method comprises the steps of 1, collecting to-be-detected data; 2, off-line preprocessing is carried out on the to-be-detected data regularly, and the to-be-detected data with labels are output; step 3, performing data cleaning on the to-be-detected data with the labels based on an LLM model, unifying the sensitive data, the normal data and the abnormal data after desensitization processing to a format suitable for analysis, and forming standardized to-be-detected data; step 4, performing information extraction and vectorization conversion on the standardized to-be-detected data based on an LLM model; 5, storing the text data subjected to vectorization conversion in the step 4 into a vector database, and storing a non-text file into a special database; and step 6, based on the vector database and the special database, according to the time range, the retrieval mode and the content input by the user, outputting a retrieval result. According to the invention, based on the LLM model, intelligent and accurate information extraction of case-related data is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to data analysis technology in the field of electronic data forensics, and specifically relates to an electronic forensics data analysis method based on LLM. Background Art

[0002] In recent years, with the rapid development of the Internet and the rapid progress of artificial intelligence technology, especially the official release of OpenAI's ChatGPT in 2022, which marks the arrival of the era of large AI models. This wave has not only affected the underlying large models, infrastructure, and machine learning operations, but also deeply influenced all aspects of consumer applications, promoting the initial construction of the generative AI ecosystem. In the field of electronic data forensics, as an important tool for combating information technology crimes, electronic data forensics searches electronic systems comprehensively by using cutting-edge technologies and following strict procedures, extracts and analyzes evidence related to crimes, and provides strong support for litigation. However, the traditional electronic forensics data analysis process is complex. Case handlers need to perform business modeling on electronic forensics data, use modeling tools to extract involved data, including structured and unstructured data, and then manually analyze, extract, and summarize information in depth to form a final forensics report. In this process, challenges such as large data volume, scattered information, and easy omission of key information are often faced, consuming a large amount of manpower and time. Therefore, aiming at this pain point in electronic forensics data analysis, simplifying the case handling process, improving the ability to process a large amount of electronic forensics data, achieving efficient and accurate extraction of involved information, obtaining clues, and forming a summary report based on this has become a key problem to be solved in this field.

[0003] In the prior art, a patent for invention with a publication number of CN118689870A discloses a system and method for analyzing the results of electronic data forensics of involved suspects. The system includes a forensics module, a data parsing and cleaning module, and a real-time case data analysis module. Through the real-time case data analysis module, case handling police officers can preview the data parsing progress in the case research and judgment room in real time, restore the case situation based on the parsed data, uniformly screen key clues, automatically match the subscribed screening conditions, and notify the corresponding case handling police officers, solving the problem of the large number of involved electronic evidences in current mass cases, and the information data islands among electronic physical evidence laboratory engineers, case analysis personnel, and data administrators, achieving the purpose of accelerating the detection speed of various cases, shortening the case handling cycle, and helping the masses recover losses.

[0004] In addition, in the prior art, the article "ChatGPT for Digital Forensic Investigation: The Good, The Bad, and The Unknown." by M. Scanlon, F. Breitinger et al. on ArXiv (2023).. Mark Scanlon, Frank Breitinger, Christopher Hargreaves, Jan-Niclas Hilgert, John Sheppard. explored the application of ChatGPT (especially its latest version GPT-4) in the field of digital forensics, including its advantages, potential risks, and unknown areas. The article evaluated the capabilities of ChatGPT in multiple digital forensics use cases through a series of experiments, such as understanding evidence, searching for evidence, code generation, anomaly detection, incident response, and education. The study found that while ChatGPT may be useful in some low-risk applications, many applications are currently either not applicable (because evidence needs to be uploaded to the service), or require users to have sufficient knowledge of the subject being questioned to identify incorrect assumptions, inaccuracies, and errors.

[0005] The above prior art helps case handlers efficiently process data through a real-time analysis module, solves the problem of information silos, and uses an asynchronous loading data parsing strategy based on a distributed file system to save computing resources and improve parsing performance, or uses the application of an open-source AI large model (ChatGPT) in digital forensics. However, a large amount of new data to be detected flows in every day, and although the asynchronous loading data parsing strategy based on the distributed file system and the data processing of the AI large model reduce part of the computational time cost, it is still difficult to meet the requirements of real-time analysis. Summary of the Invention

[0006] A brief overview of the embodiments of the present invention is given below to provide a basic understanding of certain aspects of the present invention. It should be understood that the following overview is not an exhaustive overview of the present invention. It is not intended to identify the key or important parts of the present invention, nor is it intended to limit the scope of the present invention. Its purpose is merely to present certain concepts in a simplified form as a prelude to the more detailed description to be discussed later.

[0007] The present invention aims to optimize the process for case handlers to analyze, extract, and summarize data to form an evidentiary report in a massive electronic forensics data environment. Innovatively, a decision tree integrated with a convolutional neural network first preprocesses a large amount of data to be detected flowing in every day, and then uses the excellent language understanding and reasoning ability of a large language model (hereinafter referred to as LLM) to deeply process the electronic forensics data, including steps such as data cleaning, data annotation, text information extraction, format conversion, and keyword extraction.

[0008] According to one aspect of the present application, there is provided a method for analyzing electronic forensics data based on LLM, including:

[0009] Step 1, collect data to be detected; the data to be detected includes text data, image data, and video data;

[0010] Step 2, perform offline preprocessing on the data to be detected at regular intervals, and output the data to be detected with labels; the data to be detected with labels includes normal data without sensitive information and abnormal information, sensitive data containing sensitive information, and abnormal data containing abnormal information; desensitize the sensitive data containing sensitive information, and retain the source data and the mapping relationship between the source data and the desensitized sensitive data in the database; through desensitization processing, the leakage of sensitive information can be prevented, ensuring the security of the data;

[0011] Step 3, perform data cleaning on the data to be detected with labels based on the LLM model, unify the desensitized sensitive data, normal data, and abnormal data into a format suitable for analysis to form standardized data to be detected; and further add different annotations and metadata to the desensitized sensitive data, normal data, and abnormal data in the standardized data to be detected according to different preset rules; this step of adding different annotations and metadata can be used to identify whether an image, video, or text has been tampered with;

[0012] Step 4, extract text information and keywords from the data output in Step 3 based on the LLM model to obtain sensitive data files, normal data files, and abnormal data files; scan the sensitive data files, normal data files, and abnormal data files, and extract the chat data and text data therein respectively to form sensitive data text files, sensitive data non-text files, normal data text files, normal data non-text files, abnormal data text files, and abnormal data non-text files, segment the sensitive data text files, normal data text files, and abnormal data text files, and perform vectorization conversion; in the prior art, only text is processed, and through the offline preprocessing, data classification, data annotation, and extraction of text files and non-text files in Steps 2, 3, and 4 of the present application, the comprehensive processing of text data and video image data can be conveniently achieved;

[0013] Step 5: Store the sensitive data text file, normal data text file, and abnormal data text file after the vectorization conversion in step 4 into a vector database. In this application, Elasticsearch 8 (abbreviated as ES8) is used as the vector database; store the sensitive data non-text file, normal data non-text file, and abnormal data non-text file output in step 4 into a dedicated database;

[0014] Step 6: During retrieval, based on the vector database and the dedicated database, output the retrieval results according to the time range, retrieval method, and content input by the user.

[0015] As a further implementation solution, it also includes steps for intelligent report summary and generation. Specifically, it includes: Based on the LLM model, combined with the user's keyword information, write a detailed intelligent report summary on the basis of the retrieval results.

[0016] Among them, the data to be detected collected in step 1 includes text data, image data, and video data. Generally, for smartphones, personal computers, enterprise servers, and other data storage devices, professional forensic software in the industry is used to extract electronic data, and the content covers various forms such as text messages, emails, social media records, web browsing history, photos, and videos.

[0017] As an implementation solution, the offline preprocessing of the data to be detected at regular intervals in step 2 is specifically carried out through a decision tree integrated with a convolutional neural network. The decision tree includes a root node and leaf nodes, and each node (root node and leaf nodes) includes a convolutional neural network; the root node is used to classify the data to be detected into normal data without sensitive information and abnormal information, sensitive data containing sensitive information, and abnormal data (whether it is abnormal is determined according to the pre-established criteria for different forensic purposes), and then pass it to the leaf nodes in the next layer to further identify the classified data to be detected, and at the same time remove redundant or irrelevant features of the data to be detected. Finally, the labeled data to be detected is output at the leaf nodes, and the label is normal information, sensitive information, or abnormal information; the convolutional neural network includes an input layer, a convolutional layer, a pooling layer, a fully connected layer, and an output layer. The pooling layer uses the max pooling method, and the others are implemented using general existing convolutional neural network structures, which will not be elaborated here. This decision tree is used offline. When used, a large number of new data to be detected flowing in every day first pass through the decision tree for rough processing and classification. Therefore, the data processing speed of the entire forensic process is greatly accelerated. At the same time, to avoid large changes in the structure of the decision tree caused by a large amount of data to be detected, the depth of the decision tree is set to 3 or 4.

[0018] As an implementation solution, in step 5, the sensitive data text file, normal data text file, and abnormal data text file after vectorized conversion in step 4 are stored in a vector database. Based on its vector and scalar retrieval functions, the vector database implements a dual storage strategy for text files: First, perform inverted index word segmentation processing on the text file, disassemble the text information of the text file into words, and store them in the document set of the vector database; then, store the disassembled words in a vectorized manner so that the disassembled words are converted into matrix form through a vector model and stored in the vector set of the vector database.

[0019] As a further implementation solution, in particular, for data containing context relationships in the disassembled words, such as user chat records of instant messaging software such as WeChat and QQ, to ensure that the context semantics are not ignored during vector retrieval, the sliding window technique is adopted to divide the disassembled words into multiple overlapping segments, and each segment generates a vector. For example, the disassembled words are vectorized and stored in segments of every 500 characters, where 400 characters represent the current content and 100 characters are the ending content of the previous segment. At the same time, corresponding scalar information is supplemented for this vector (vector content) so that during the retrieval process, filtering is first performed through the scalar information to achieve more accurate query results.

[0020] Furthermore, step 6 realizes the hybrid retrieval of user vectors and scalars based on the vector database and a dedicated database, specifically including:

[0021] The user inputs the time range of the data to be retrieved for evidence collection, selects a suitable retrieval method and keywords;

[0022] Perform scalar filtering on the vector database and the dedicated database based on the time range;

[0023] Perform full-text index word segmentation retrieval on the content obtained after scalar filtering based on the keywords and retrieval method, and output the full-text retrieval results;

[0024] Convert the full-text retrieval results into vector form through a vector model, and use the vector recall algorithm to perform recall calculation on the full-text retrieval results in vector form. The recall algorithm is a technique used to find the data most relevant to a given query from a large amount of data. In this process, multiple vector recall algorithms can be used or different vector recall algorithms can be selected according to the scenario or data characteristics, such as the Euclidean algorithm and the cosine distance algorithm, etc. The Euclidean algorithm evaluates the similarity between two vectors by calculating the straight-line distance between them, and the smaller the distance, the higher the similarity. The cosine distance algorithm evaluates the similarity between two vectors by calculating the cosine value of the angle between them, and the closer the cosine value is to 1, the higher the similarity. The selection of these algorithms depends on the specific application scenario and data characteristics.

[0025] The recall algorithm is used to perform recall calculations on the full-text retrieval results in vector form to determine semantically similar content and obtain vector retrieval results;

[0026] The full-text retrieval results and the vector retrieval results are merged, and the merged results are rearranged according to the similarity scores to obtain the final retrieval results.

[0027] Through the above series of complex processing processes and combined with the application of the recall algorithm, this data analysis method can finally provide an accurate and efficient retrieval result.

[0028] According to another aspect of the present application, there is provided an electronic forensics data analysis system based on LLM, including:

[0029] A data collection module to be detected, used to collect data to be detected, and the data to be detected includes text data, image data, and video data;

[0030] An offline preprocessing module, used to perform offline preprocessing on the data to be detected regularly and output the data to be detected with labels; the data to be detected with labels includes normal data without sensitive information and abnormal information, sensitive data containing sensitive information, and abnormal data containing abnormal information; desensitize the sensitive data containing sensitive information, and retain the source data and the mapping relationship between the source data and the desensitized sensitive data in the database;

[0031] A data cleaning module, used to clean the data to be detected with labels based on the LLM model, unify the desensitized sensitive data, normal data, and abnormal data into a format suitable for analysis to form standardized data to be detected; and further add different annotations and metadata to the desensitized sensitive data, normal data, and abnormal data in the standardized data to be detected according to different preset rules;

[0032] A data processing module, based on the LLM model, extracts text information and keywords from the data output by the data cleaning module to obtain sensitive data files, normal data files, and abnormal data files; scans the sensitive data files, normal data files, and abnormal data files, extracts the chat data and text data respectively from them to form sensitive data text files, sensitive data non-text files, normal data text files, normal data non-text files, abnormal data text files, and abnormal data non-text files, segments the sensitive data text files, normal data text files, and abnormal data text files, and performs vectorization conversion;

[0033] A storage module that stores the sensitive data text files, normal data text files, and abnormal data text files after the vectorized conversion by the data processing module into a vector database; and stores the sensitive data non-text files, normal data non-text files, and abnormal data non-text files output by the data processing module into a dedicated database;

[0034] A retrieval module that, based on the vector database and the dedicated database, outputs retrieval results according to the time range, retrieval method, and content input by the user;

[0035] This electronic forensics data analysis system is used to execute the above-mentioned electronic forensics data analysis method based on LLM.

[0036] Through the above solution, compared with the prior art, this application has the following advantages:

[0037] 1. First, preprocessing is performed through an improved decision tree. A large amount of new data to be detected flowing in every day is first roughly processed and classified by the decision tree, thus greatly reducing the computational time cost and significantly accelerating the data processing speed of the entire forensics process, meeting the requirements of real-time analysis;

[0038] 2. By performing desensitization processing during the preprocessing process, the leakage of sensitive information can be prevented, ensuring the security of the data;

[0039] 3. It cleverly combines the comprehension and reasoning capabilities of the LLM model and the advanced vector recall technology to achieve efficient extraction and accurate summarization of massive electronic forensics data.

[0040] In summary, currently, the systems related to electronic data forensics analysis on the market mainly rely on case handlers to perform business modeling on electronic forensics data, use modeling tools to extract relevant data in the case, and then manual in-depth analysis, extraction, and induction of information are required to form the final forensics report. This application is based on the LLM model and applies the LLM intelligent large model to analyze in all steps of electronic forensics data analysis. Combining the actual situation of existing electronic data analysis, through intelligent segmentation processing of electronic forensics data and vectorized conversion, and then storing it in a database that is compatible with both vector and vector data, case handlers can interact efficiently through the LLM model, extract information from the relevant data in the case, summarize and form a report. This process greatly simplifies the research and judgment and summarization process of case handlers when facing massive data, effectively reduces the work complexity, significantly improves the case handling efficiency, and enhances the ability to accurately extract and locate relevant data in the case. It has been effectively implemented in the product, and the entire method has achieved good results in actual application and promotion. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] The present invention can be better understood by referring to the descriptions given below in conjunction with the accompanying drawings, in which the same or similar reference numerals are used throughout the drawings to denote the same or similar components. The accompanying drawings, together with the following detailed description, are included in this specification and form a part of this specification, and are further used to illustrate the preferred embodiments of the present invention and to explain the principles and advantages of the present invention. In the attached

[0042] In the figures:

[0043] Figure 1 is an implementation block diagram of the LLM-based electronic forensics data analysis method of the present invention;

[0044] Figure 2 is a schematic diagram of the dual storage strategy of the present invention;

[0045] Figure 3 is a schematic diagram of the hybrid retrieval of the present invention. Detailed implementation manners

[0046] Embodiments of the present invention will be described below with reference to the accompanying drawings. Elements and features described in one drawing or one embodiment of the present invention can be combined with elements and features shown in one or more other drawings or embodiments. It should be noted that for the sake of clarity, representations and descriptions of components and processes irrelevant to the present invention and known to those of ordinary skill in the art are omitted in the drawings and the description.

[0047] The LLM model, i.e., the large language model, is an advanced neural network system built based on Transformer technology. Through deep learning and large-scale data training, this model has powerful reasoning capabilities and text generation capabilities. When processing natural language text, the LLM model can understand complex semantic relationships and generate coherent and logically strong text content according to the given prompt words. The core technology behind it - Transformer, is an architecture specifically designed for processing sequence data, which effectively captures long-range dependencies in the text through the self-attention mechanism, thus performing excellently in various natural language processing tasks. The reasoning capabilities of the LLM model are not limited to simple text filling or continuation, but also include in-depth understanding and logical reasoning of text content, making it widely used in multiple fields such as question-answering systems, text summarization, and machine translation. In addition, the text generation capabilities of the LLM large language model are particularly prominent. It can automatically generate rich and diverse text content with different styles according to the prompt words or topics input by users. This generation process is not just a simple stacking of words, but is based on the model's profound understanding and grasp of language rules, as well as the analysis and learning of a large amount of text data. Therefore, the text generated by the LLM model often has high readability and fluency, and can well meet the writing needs of users.

[0048] In the processing of prompt words, the LLM model also demonstrates extremely high flexibility and adaptability. Whether it is simple keywords, phrases, or complex sentences and paragraphs, the LLM model can accurately capture their semantic information and integrate it into the generated text. This ability enables the LLM model to have broad application prospects in fields such as creative writing and content creation, helping users quickly generate high-quality text content and improve work efficiency and creative levels.

[0049] In summary, as an advanced neural network system based on Transformer technology, the LLM large language model plays a crucial role in the field of natural language processing with its reasoning ability and text generation ability.

[0050] Vector retrieval: In the fields of mathematics and computer science, vectors, matrices, and linear algebra form the core of basic concepts. A vector is defined as a quantity with a specific direction and magnitude, usually represented as a point or arrow in a multi-dimensional space. A matrix is a rectangular array composed of numbers or symbols, which plays an important role in representing linear transformations or systems of equations. Linear algebra, as a branch of mathematics, focuses on the study of vector spaces and linear mappings and is indispensable for solving numerous scientific and engineering problems.

[0051] The Euclidean distance is a method used to measure the distance between two points in a multi-dimensional space. Its principle is based on the Pythagorean theorem and is widely used in the fields of geometry and data analysis. The cosine distance evaluates the similarity between two vectors by calculating the cosine value of the angle between them, and this method is particularly popular in the fields of text analysis and recommendation systems.

[0052] Similarity calculation refers to the process of evaluating the similarity degree of two objects in specific features. This process has important applications in multiple fields such as information retrieval, data mining, and machine learning. Scalar-vector hybrid retrieval involves comprehensively using scalar and vector information in the retrieval process, aiming to improve the accuracy and efficiency of retrieval. Vector recall refers to the process in information retrieval where, by calculating the similarity between the query vector and each vector in the database, the data items most relevant to the query are selected.

[0053] RAG, that is, Retrieval-Augmented Generation, is a cutting-edge technical means that cleverly combines information retrieval and the generation ability of large language models (LLMs). By introducing an external knowledge base, the RAG technology can effectively make up for the deficiencies of large language models in specific domain knowledge and solve the problem that these large models cannot obtain the latest knowledge in real time.

[0054] When generating text, the RAG technology first retrieves relevant information from a pre - constructed knowledge base, which is usually achieved by a method of hybrid retrieval of scalars and vectors. Specifically, the combined use of scalars and vectors can greatly improve the efficiency and accuracy of retrieval. Scalar retrieval mainly focuses on single numerical features, while vector retrieval can handle multi - dimensional data features. By combining these two methods, the respective advantages can be fully utilized, so as to quickly find the required information in large - scale datasets. This hybrid retrieval method has extensive applications in many fields, such as image recognition, natural language processing, and recommendation systems, etc. By optimizing the hybrid ratio of scalars and vectors and the retrieval algorithm, the performance of retrieval can be further improved to meet the requirements in different scenarios. After in - depth analysis of the data and information recall, the screened and sorted information is used as part of the input. By closely associating the prompt words with the recalled information and combining the reasoning ability of the language model, more rich and accurate text content can be generated. This method not only makes full use of the powerful ability of the language model in language processing but also ensures the accuracy and timeliness of the generated content. In this way, more detailed and accurate information can be provided in a specific field to meet the needs of users in specific scenarios.

[0055] The core advantage of the RAG technology lies in its ability to combine the information in the external knowledge base with the generation ability of the language model, thus generating more accurate and timely text. This technology can not only improve the performance of the language model in a specific field but also expand its application scope, enabling it to handle more complex and professional problems. In this way, the RAG technology can provide users with more rich and accurate information to meet their needs in a specific field. In addition, the RAG technology can also update the knowledge base in real - time to ensure that the generated text content is the latest, thereby improving the timeliness and reliability of the information. Generally speaking, the RAG technology provides users with an efficient, accurate, and timely information generation method by combining information retrieval and the generation ability of the language model.

[0056] Embodiments of the present invention aim to optimize the process in which case-handling personnel analyze, extract, and summarize data to form a forensic report in a massive electronic forensics data environment. First, a decision tree integrated with a convolutional neural network preprocesses a large amount of data to be detected flowing in every day. Then, by leveraging the excellent language understanding and reasoning capabilities of a large language model (LLM for short), in-depth processing of electronic forensics data is carried out, including steps such as data cleaning, data annotation, text information extraction, format conversion, and keyword extraction. In the data storage phase, this method specifically performs intelligent segmentation processing on chat data and text data (such as articles, transcripts, etc.), and implements vectorization conversion, and then stores them in a database that is compatible with both vector and vector data, such as Elasticsearch 8 (ES8 for short). When case-handling personnel need to extract specific case-related information from massive electronic forensics data, the present invention allows for efficient interaction through the LLM model. Specifically, it tokenizes the user query and converts it into a vector form, and then accurately retrieves relevant case-related information in the ES8 database. The retrieved data is summarized and then input into the LLM model again for intelligent summarization, and finally a complete report is output according to the format specified by the case-handling personnel. The core value of the present invention lies in its ingenious combination of the understanding and reasoning capabilities of the LLM model and advanced vector retrieval technology, achieving efficient extraction, accurate summarization, and automated report generation of massive electronic forensics data. This process greatly simplifies the research and judgment and summarization processes of case-handling personnel in the face of massive data, effectively reduces the work complexity, significantly improves the case-handling efficiency, and enhances the ability to accurately extract and locate case-related data.

[0057] Embodiments of the present invention provide an LLM-based method for analyzing electronic forensics data, which specifically includes the following processes:

[0058] 1. Data extraction

[0059] Smartphones, personal computers, enterprise servers, and other data storage devices store rich information resources. With the help of professional forensics software in the industry, electronic data is extracted, including various forms such as text messages, emails, social media records, web browsing history, photos, and videos, to form data to be detected.

[0060] 2. Offline preprocessing

[0061] The data to be detected is periodically preprocessed offline to output data to be detected with labels; the data to be detected with labels includes normal data without sensitive information and abnormal information, sensitive data containing sensitive information, and abnormal data containing abnormal information; the sensitive data containing sensitive information is desensitized, and the source data and the mapping relationship between the source data and the desensitized sensitive data are retained in the database.

[0062] Offline preprocessing is performed on the data to be detected at regular intervals, such as preprocessing at a preset time every day. The preprocessing is specifically carried out through a decision tree integrated with a convolutional neural network. The decision tree includes a root node and leaf nodes, and each node (the root node and leaf nodes) includes a convolutional neural network; the root node is used to classify the data to be detected into normal data without sensitive information and abnormal information, sensitive data containing sensitive information, and abnormal data (whether it is abnormal is determined according to the pre-established criteria for different forensic purposes), and then it is passed to the leaf nodes in the next layer for further identification of the classified data to be detected. At the same time, redundant or irrelevant features of the data to be detected are removed, and finally, the data to be detected with labels is output at the leaf nodes, and the labels are normal information, sensitive information, or abnormal information; the convolutional neural network includes an input layer, a convolutional layer, a pooling layer, a fully connected layer, and an output layer. The pooling layer uses the max pooling method, and the others are implemented using the general existing convolutional neural network structure, which will not be elaborated here. This decision tree is used offline. When in use, a large number of newly incoming data to be detected every day first undergoes rough processing and classification by the decision tree. Therefore, the data processing speed of the entire forensic process is greatly accelerated. At the same time, to avoid large changes in the structure of the decision tree caused by a large amount of data to be detected, the depth of the decision tree is set to 3 or 4.

[0063] 3. Data cleaning

[0064] Data cleaning involves multiple aspects, including but not limited to removing duplicate data to ensure the uniqueness and accuracy of the data; performing format conversion to unify the data into a format suitable for analysis; and data annotation, attaching necessary labels and metadata to the data for subsequent processing and analysis. Through these steps, the data is transformed into a structured and standardized form, facilitating more in-depth analysis. This not only improves the quality of the data but also ensures the accuracy and reliability of the analysis results. At the data warehousing stage, a large language model is used to extract specific information from the data, such as ID numbers, dates, bank card information, special holidays, and user key information, and combined with the reasoning and summarization capabilities of a large prediction model to achieve information extraction, annotation, and filling into relevant fields.

[0065] 4. Data warehousing

[0066] After completing the data cleaning work, according to the nature differences of the data, data governance is carried out and the data is stored in the database. During the data storage process, the vector and scalar retrieval functions of ES8 are utilized to implement a dual storage strategy for the forensics data. First, the data is subjected to inverted index word segmentation processing, and the text information is disassembled into words and then stored in the document collection of ES. Then, the data is stored in a vectorized manner so that the data can be converted into a matrix form through a vector model and stored in the vector collection of ES. In particular, for data with context relationships, such as the user chat records of instant messaging software like WeChat and QQ, in order to ensure that the context semantics are not ignored during vector retrieval, the sliding window technique is adopted to divide the text into multiple overlapping segments, and each segment generates a vector. The data is vectorized and stored in segments of every 500 characters, where 400 characters represent the current content and 100 characters are the ending content of the previous segment. At the same time, corresponding scalar information is supplemented for the vector content so that during the retrieval process, the scalar information is first used for screening to achieve more accurate query results.

[0067] In this embodiment, the retrieval methods for pictures and videos are different from those for text. Therefore, the vectorized converted sensitive data text files, normal data text files, and abnormal data text files are stored in the vector database; the sensitive data non-text files, normal data non-text files, and abnormal data non-text files are stored in a dedicated database.

[0068] 5. User Vector and Scalar Hybrid Retrieval

[0069] The user inputs the time range of the forensics data to be retrieved, and this time range determines the start and end time points for the system to retrieve the data. The user also needs to select a suitable retrieval method and content (keywords), and this step is equally important because it will affect the specific methods and scope for the system to retrieve the data. At the same time, the user can also choose to require the system to introduce the recall algorithm in detail here in order to understand the technical details in the retrieval process more deeply. After the user completes these inputs, the system will perform scalar filtering on the forensics data according to the time range information specified by the user to ensure that the retrieved data range meets the user's requirements.

[0070] During the retrieval process, the system will perform full-text index word segmentation retrieval on the keywords input by the user. In this way, the system can quickly obtain content similar to the keywords.

[0071] Subsequently, the system will convert the retrieved content into vector form through a vector model. The purpose of this is to use vector recall algorithms to perform recall calculations on vector data. Recall algorithms are techniques for finding the data most relevant to a given query from a large amount of data. In this process, the system may use multiple vector recall algorithms, such as the Euclidean algorithm and the cosine distance algorithm. The Euclidean algorithm evaluates the similarity between two vectors by calculating the straight-line distance between them. The smaller the distance, the higher the similarity. The cosine distance algorithm, on the other hand, evaluates the similarity between two vectors by calculating the cosine value of the angle between them. The closer the cosine value is to 1, the higher the similarity. The choice of these algorithms depends on the specific application scenario and data characteristics. The system uses these recall algorithms to perform recall calculations on vector data to determine semantically similar content. Finally, the system will merge the full-text retrieval results with the vector retrieval results. This step is to combine the advantages of the two retrieval methods to obtain more comprehensive and accurate retrieval results. After the merging is completed, the system will re-rank the results according to the similarity scores. The purpose of this is to ensure that the final retrieval results can be sorted according to the similarity to the retrieved content, so that users can more easily find the information they need.

[0072] Through this series of complex processing procedures and combined with the use of recall algorithms, the system can finally provide an accurate and efficient retrieval result.

[0073] 6. Intelligent Report Summary and Generation

[0074] Based on the retrieval results, combined with the keywords of the questions submitted by the user and the prompting words generated by the LLM model, a detailed report is written. This process requires combining the retrieved information with the user's questions to produce a logically coherent and accurate response. In this link, the LLM model will use the retrieved information to enhance its ability to generate responses, ensuring that the responses not only come from its training data but also incorporate the latest and relevant data. The model integrates the retrieved information with the user's questions. The core of this step lies in how to effectively combine this information to generate a comprehensive and accurate response. The model needs to identify the key elements in the retrieved information and combine these elements with the context of the user's questions.

[0075] To achieve this goal, prompt words are used to guide the generation of responses. Prompt words are phrases or sentences that the model relies on during the response generation process. They may be a restatement of the question, a summary of the retrieved information, or a direct indication of the response. Through carefully designed prompt words, the LLM model can more accurately focus on the key parts of the retrieved information and incorporate this information into the final response. For example, when the user asks the question "How to treat a cold", the model may retrieve relevant information about cold treatment methods. Subsequently, it will use prompt words such as "Common cold treatment methods include" to guide the generation of the response. The model will integrate the retrieved treatment methods after the prompt words to form a complete response. Finally, the model will use its language generation ability to transform the integrated information into the content required by the prompt words. This process may involve optimizing the fluency, coherence, and grammar of the language to ensure that the response is both accurate and easy to understand.

[0076] An embodiment of the present invention further provides an LLM-based electronic forensics data analysis system, which includes:

[0077] A data collection module to be detected, configured to collect data to be detected, where the data to be detected includes text data, image data, and video data;

[0078] An offline preprocessing module, configured to perform offline preprocessing on the data to be detected regularly and output the data to be detected with labels; the data to be detected with labels includes normal data without sensitive information and abnormal information, sensitive data containing sensitive information, and abnormal data containing abnormal information; desensitize the sensitive data containing sensitive information, and retain the source data and the mapping relationship between the source data and the desensitized sensitive data in the database;

[0079] A data cleaning module, configured to perform data cleaning on the data to be detected with labels based on the LLM model, unify the desensitized sensitive data, normal data, and abnormal data into a format suitable for analysis to form standardized data to be detected; and further add different annotations and metadata to the desensitized sensitive data, normal data, and abnormal data in the standardized data to be detected according to different preset rules;

[0080] The data processing module extracts text information and keywords from the data output by the data cleaning module based on the LLM model, obtaining sensitive data files, normal data files, and abnormal data files; scans the sensitive data files, normal data files, and abnormal data files, extracts the chat data and text data therein respectively, forms sensitive data text files, sensitive data non-text files, normal data text files, normal data non-text files, abnormal data text files, and abnormal data non-text files, segments the sensitive data text files, normal data text files, and abnormal data text files, and performs vectorization conversion;

[0081] The storage module stores the sensitive data text files, normal data text files, and abnormal data text files after vectorization conversion by the data processing module into the vector database; stores the sensitive data non-text files, normal data non-text files, and abnormal data non-text files output by the data processing module into a dedicated database;

[0082] The retrieval module outputs retrieval results based on the vector database and the dedicated database according to the time range, retrieval method, and content input by the user;

[0083] This electronic forensics data analysis system is used to execute the above-mentioned electronic forensics data analysis method based on LLM.

[0084] Currently, the systems related to electronic data forensics analysis on the market mainly rely on case-handling personnel to perform business modeling on electronic forensics data, use modeling tools to extract relevant case data, and then manual in-depth analysis, extraction, and induction of information are required to form the final forensics report. The electronic forensics data analysis method and system based on LLM of the present invention, in combination with the actual situation of existing electronic data analysis, simplify the electronic forensics data through intelligent segmentation processing, perform vectorization conversion, and then store it in a database that is compatible with both vector and vector data. Case-handling personnel can interact efficiently through the LLM model, extract information from the relevant case data, summarize and form a report. This process greatly simplifies the research and judgment and summary process of case-handling personnel when facing a large amount of data, effectively reduces the work complexity, significantly improves the case-handling efficiency, and enhances the ability to accurately extract and locate relevant case data. It has been effectively implemented in the product, and the entire method has achieved good results in actual application and promotion.

[0085] The method of the present invention is not limited to being executed in the time sequence described in the specification, and can also be executed in other time sequences, in parallel, or independently. Therefore, the execution sequence of the method described in this specification does not limit the technical scope of the present invention.

[0086] Although the present invention has been disclosed above by the description of specific embodiments of the present invention, it should be understood that all the above embodiments and examples are exemplary, rather than restrictive. Those skilled in the art can design various modifications, improvements or equivalents to the present invention within the spirit and scope of the appended claims. These modifications, improvements or equivalents should also be considered to be included within the protection scope of the present invention.

Claims

1. An electronic forensic data analysis method based on LLM, characterized by: include: Step 1, collecting data to be detected; the data to be detected includes text data, image data and video data; Step 2: Regularly pre-process the data to be detected offline and output the data to be detected with labels; the data to be detected with labels include normal data without sensitive information and abnormal information, sensitive data with sensitive information, and abnormal data with abnormal information; Desensitize sensitive data containing sensitive information, and retain the source data and the mapping relationship between the source data and the desensitized sensitive data in the database; Step 3: Based on the LLM model, the labeled data to be tested is cleaned, and the sensitive data, normal data, and abnormal data after desensitization are unified into a format suitable for analysis to form standardized data to be tested; And according to different preset rules, different annotations and metadata are further added to the desensitized sensitive data, normal data and abnormal data in the standardized data to be detected; Step 4: Extract text information and keywords from the data output in step 3 based on the LLM model to obtain sensitive data files, normal data files, and abnormal data files; Scan sensitive data files, normal data files, and abnormal data files, extract chat data and text data therein respectively, form sensitive data text files, sensitive data non-text files, normal data text files, normal data non-text files, abnormal data text files, and abnormal data non-text files, segment the sensitive data text files, normal data text files, and abnormal data text files, and perform vectorization conversion; Step 5, storing the sensitive data text file, normal data text file and abnormal data text file after vectorization conversion in step 4 into a vector database; storing the sensitive data non-text file, normal data non-text file and abnormal data non-text file outputted in step 4 into a special database; Step 6: During the search, based on the vector database and the specialized database, the search results are output according to the time range, search method and content input by the user.

2. The electronic forensic data analysis method according to claim 1, characterized in that: It also includes the steps of summarizing and generating intelligent reports, including: based on the LLM model, writing on the basis of the search results and combining the user's keyword information to form a detailed intelligent report summary.

3. The electronic forensic data analysis method according to claim 1, characterized in that: In the step 2, offline preprocessing is performed on the data to be detected at regular intervals, specifically, offline preprocessing is performed by a decision tree integrated with a convolutional neural network, the decision tree includes a root node and a leaf node, and each node includes a convolutional neural network; the root node is used to classify the data to be detected into normal data without sensitive information and abnormal information, sensitive data with sensitive information, and abnormal data, and then the data is passed to the leaf node of the next layer to further identify the classified data to be detected, and at the same time, redundant or irrelevant features of the data to be detected are removed, and finally the data to be detected with a label is output at the leaf node, and the label is normal information, sensitive information or abnormal information.

4. The electronic forensic data analysis method according to claim 1, characterized in that: In the step 5, the sensitive data text file, normal data text file and abnormal data text file vectorized and converted in step 4 are stored in the vector database. The vector database implements a dual storage strategy for the text files based on its vector and scalar retrieval functions: first, an inverted index word segmentation process is performed on the text file, and the text information of the text file is decomposed into words and then stored in the document collection of the vector database; then, the decomposed words are vectorized and stored, so that the decomposed words are converted into a matrix form through a vector model and stored in the vector collection of the vector database.

5. The electronic forensic data analysis method according to claim 4, characterized in that: For data containing contextual relationships in the disassembled vocabulary, the sliding window technology is used to divide the disassembled vocabulary into multiple overlapping segments, each segment generates a vector; and the corresponding scalar information is supplemented for the vector.

6. The electronic forensic data analysis method according to claim 1, characterized in that: The step 6 implements user vector and scalar mixed retrieval based on the vector database and the special database, specifically including: The user enters the time range of the forensic data to be retrieved, selects the appropriate search method and keywords; scalar filtering is performed on the vector database and the specialized database based on the time range; Perform full-text indexing and word segmentation retrieval on the content obtained after scalar filtering based on keywords and retrieval methods, and output full-text retrieval results; The full-text search results are converted into vector form through a vector model, and the vector recall algorithm is used to perform recall calculation on the full-text search results in vector form; The full-text search results in vector form are recalled by using a recall algorithm to determine semantically similar content to obtain vector search results; The full-text search results are combined with the vector search results, and the combined results are rearranged according to the similarity score to obtain the final search results.

7. An electronic forensic data analysis system based on LLM, characterized by: include: The data collection module is used to collect the data to be detected, and the data to be detected includes text data, image data and video data; The offline preprocessing module is used to periodically perform offline preprocessing on the data to be detected and output the data to be detected with labels; the data to be detected with labels include normal data without sensitive information and abnormal information, sensitive data with sensitive information, and abnormal data with abnormal information; Desensitize sensitive data containing sensitive information, and retain the source data and the mapping relationship between the source data and the desensitized sensitive data in the database; The data cleaning module is used to clean the labeled data to be tested based on the LLM model, unify the sensitive data, normal data and abnormal data after desensitization processing into a format suitable for analysis, and form standardized data to be tested; And according to different preset rules, different annotations and metadata are further added to the desensitized sensitive data, normal data and abnormal data in the standardized data to be detected; The data processing module extracts text information and keywords from the data output by the data cleaning module based on the LLM model to obtain sensitive data files, normal data files, and abnormal data files; Scan sensitive data files, normal data files, and abnormal data files, extract chat data and text data therein respectively, form sensitive data text files, sensitive data non-text files, normal data text files, normal data non-text files, abnormal data text files, and abnormal data non-text files, segment the sensitive data text files, normal data text files, and abnormal data text files, and perform vectorization conversion; A storage module stores the sensitive data text files, normal data text files and abnormal data text files converted by the data processing module into a vector database; and stores the sensitive data non-text files, normal data non-text files and abnormal data non-text files output by the data processing module into a special database; The search module, based on the vector database and the specialized database, outputs the search results according to the time range, search method and content input by the user; The electronic forensic data analysis system is used to execute the LLM-based electronic forensic data analysis method described in any one of claims 1-6.

Citation Information

Patent Citations

  • Electronic data evidence obtaining result analysis system and method for criminal suspects involved in case

    CN118689870A