Mobile phone information automatic investigation system and method for network-related cases

By employing automatic privilege escalation and kernel-level interception technologies within a wireless local area network, the convenience and data extraction issues of traditional mobile phone evidence collection methods have been resolved. This enables wireless and automated mobile phone information investigation, improving the efficiency of investigations and evidence acquisition capabilities in internet-related cases.

CN122496820APending Publication Date: 2026-07-31LIAONING POLICE ACAD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
LIAONING POLICE ACAD
Filing Date
2026-04-29
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing mobile phone electronic data forensics technology relies on wired connections, which is cumbersome, time-consuming, and makes it difficult to quickly and selectively extract core data involved in a case. In particular, it cannot effectively intercept encrypted voice evidence, which affects the efficiency of case investigation.

Method used

By automatically elevating privileges at the operating system level within a wireless local area network, establishing wireless communication links, and enabling targeted traversal and selective extraction of mobile phone communication flow, network flow, and financial flow data, and by intercepting native voice files through kernel-level interception technology, a wireless and automated mobile phone information investigation system can be constructed.

Benefits of technology

It enables wireless and automated mobile phone information investigation, improving the convenience and efficiency of evidence collection. It can quickly and accurately extract case-related data, especially encrypted voice evidence, supporting efficient case investigation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122496820A_ABST
    Figure CN122496820A_ABST
Patent Text Reader

Abstract

This application discloses an automatic mobile phone information investigation system and method for internet-related cases, belonging to the field of automatic investigation. This solution completely eliminates the reliance on data cables, manual activation of debugging modes, and installation of specific drivers by constructing a wireless, automatically privileged acquisition environment, thus solving the technical problems of poor mobility and complex operation. Furthermore, the solution abandons the time-consuming full-disk backup mode, instead performing targeted traversal and selective extraction of key data such as communication, network, and financial flow based on investigation instructions, significantly improving evidence collection efficiency. Crucially, by employing kernel-level interception and redirection technology, it can directly intercept native voice files before the chat software performs encryption operations, effectively overcoming the difficulty of obtaining voiceprint evidence using traditional methods, and providing key technical support for case investigation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of automated investigation, and more specifically, to an automated mobile phone information investigation system and method for internet-related cases. Background Technology

[0002] With the widespread adoption of smartphones in daily life, they are increasingly becoming a key tool for cybercrime. Criminals extensively utilize smartphones for telecommunications fraud, illegal transactions, and invasions of privacy, making the rapid and accurate extraction of electronic data from involved mobile phones a crucial step in combating crime and securing evidence. Therefore, developing an efficient and intelligent automated mobile phone information investigation solution is of paramount practical significance for improving the efficiency of investigating cybercrime cases.

[0003] However, existing mobile phone electronic data forensics technologies have shown numerous shortcomings in addressing the current complex cybercrime situation. Traditional forensics methods heavily rely on wired connections, typically requiring a USB cable to connect the phone to a dedicated forensics terminal. This method not only limits the mobility and flexibility of forensics work, hindering rapid on-site investigations of mobile suspects, but also places high technical demands on operators. Forensic personnel must be familiar with the system settings of different brands and models of mobile phones, manually enable developer modes such as USB debugging, and find and install specific drivers. The entire process is cumbersome and time-consuming, making it difficult to promote on a large scale in grassroots law enforcement units. In addition, traditional forensics methods often employ full disk image backups. With modern smartphones often boasting hundreds of gigabytes of storage space, full disk backups are time-consuming, severely impacting the crucial time for case investigation. More importantly, this indiscriminate collection method lacks specificity and makes it difficult to quickly and targetedly extract core data related to cases, such as communication flows, network flows, and fund flows. Especially when dealing with cases such as telecommunications fraud that require voiceprint comparison, it is unable to effectively intercept the native voice files of applications such as WeChat and QQ that are encrypted in real time during transmission, thus leading to the loss of key evidence.

[0004] Therefore, how to break through the technical bottlenecks of traditional wired connections and manual settings, and develop a mobile phone information investigation method that requires no complicated operation, can access wirelessly, and can automatically, accurately, and deeply extract core case-related data, especially solving the problem of effectively intercepting dynamic evidence such as encrypted voice, has become an urgent technical challenge in this field. Summary of the Invention

[0005] To address the aforementioned problems in existing technologies, this application provides an automatic mobile phone information investigation method for internet-related cases, comprising: Step 1, after connecting the mobile phone to be investigated (with a data collection application installed) and the evidence collection terminal (with background management software installed) to the same wireless local area network, performing automatic privilege escalation at the operating system level on the data collection application and binding the privilege escalation state to the established wireless communication link to obtain an privilege escalation data collection environment; Step 2, based on the data collection range instructions input by the investigators, performing targeted traversal and selective extraction of communication flow data, network flow data, and financial flow data of the mobile phone to be investigated in the privilege escalation data collection environment to obtain conventional case-related data packets; Step 3, performing kernel-level interception and redirection copying of file transfer instructions triggered when synchronizing audio files in network chat software in the privilege escalation data collection environment, intercepting and outputting the original voice file before the encryption operation is executed; Step 4, uniformly encoding and encapsulating the conventional case-related data packets and the original voice file to obtain a standard data exchange packet; Step 5, performing structured parsing and storage of the standard data exchange packet to extract multi-dimensional feature entity sets, and performing association mining and in-depth analysis on the multi-dimensional feature entity sets to obtain a visualized clue topology map.

[0006] This application also provides an automatic mobile phone information investigation system for internet-related cases, comprising: an privilege escalation state and link binding module, used to automatically escalate privileges at the operating system level for the mobile phone to be investigated with a data collection application installed and an evidence collection terminal with a background management software installed after connecting them to the same wireless local area network, and binding the privilege escalation state to the established wireless communication link to obtain an privilege escalation data collection environment; and a conventional case-related data packet generation module, used to perform targeted traversal and selective processing of communication flow data, network flow data, and financial flow data of the mobile phone to be investigated in the privilege escalation data collection environment based on the data collection range instructions input by the investigators. The system extracts and obtains standard data packets related to the case; a native voice file output module is used to intercept and redirect file transfer instructions triggered when synchronizing audio files with network chat software in the privilege escalation acquisition environment at the kernel level, intercepting and outputting native voice files before the encryption operation is executed; a unified encoding and encapsulation module is used to uniformly encapsulate the standard data packets related to the case and the native voice files to obtain standard data exchange packets; and a visual clue topology map generation module is used to perform structured parsing and storage of the standard data exchange packets to extract multi-dimensional feature entity sets, and to perform association mining and in-depth analysis on the multi-dimensional feature entity sets to obtain a visual clue topology map.

[0007] Compared with existing technologies, this application provides an automated mobile phone information investigation system and method for internet-related cases, aiming to address the core pain points of existing evidence collection methods, such as strong reliance on physical cables, cumbersome operation procedures, lack of targeted data extraction, and inability to effectively obtain dynamically encrypted evidence. This solution completely eliminates the dependence on data cables, manual activation of debugging modes, and installation of specific drivers by constructing a wireless, automatically privileged acquisition environment, thus solving the technical problems of poor mobility and complex operation. Furthermore, the solution abandons the time-consuming full-disk backup mode, instead performing targeted traversal and selective extraction of key data such as communication, network, and financial flows based on investigation instructions, significantly improving evidence collection efficiency. Crucially, by employing kernel-level interception and redirection technology, it can directly intercept native voice files before the chat software performs encryption operations, effectively overcoming the difficulty of obtaining voiceprint evidence using traditional methods, and providing key technical support for case investigation. Attached Figure Description

[0008] The above and other objects, features and advantages of this application will become more apparent from the more detailed description of the embodiments of this application in conjunction with the accompanying drawings.

[0009] Figure 1 This is a flowchart of an automatic mobile phone information investigation method for internet-related cases according to an embodiment of this application.

[0010] Figure 2 This is a data flow diagram of an automatic mobile phone information investigation method for internet-related cases according to an embodiment of this application.

[0011] Figure 3 This is a flowchart of step 3 in the automatic mobile phone information investigation method for internet-related cases according to an embodiment of this application.

[0012] Figure 4 This is a block diagram of an automatic mobile phone information investigation system for internet-related cases, according to an embodiment of this application. Detailed Implementation

[0013] The embodiments of this application will now be described in more detail with reference to the accompanying drawings. It should be understood that the drawings and embodiments of this application are for illustrative purposes only and are not intended to limit the scope of protection of this application.

[0014] To address the limitations of the aforementioned background technology, this application proposes an automatic mobile phone information investigation method for internet-related cases. For example... Figure 1 and Figure 2 As shown, Figure 1 This is a flowchart of an automatic mobile phone information investigation method for internet-related cases according to an embodiment of this application. Figure 2 This is a data flow diagram of an automatic mobile phone information investigation method for internet-related cases according to an embodiment of this application.

[0015] In step 1, after connecting the mobile phone to be investigated (with the data collection application installed) and the evidence collection terminal (with the backend management software installed) to the same wireless local area network, an automatic privilege escalation operation is performed on the data collection application at the operating system level, and the privilege escalation status is bound to the established wireless communication link to obtain an privileged data collection environment. It should be understood that in existing mobile electronic data forensics practices, wired connections are commonly used. This not only severely restricts the mobility and immediacy of investigation work but also constitutes a high technical barrier for frontline law enforcement personnel, requiring them to perform tedious manual debugging and driver configuration for different devices. More importantly, the openness and uncertainty of the data transmission link pose potential risks to the integrity and originality of evidence. Therefore, step 1 aims to construct a wireless data channel that requires no physical contact, can automatically penetrate system protection, and establish a stable, reliable, and highest-privilege wireless data channel. This fundamentally solves the inherent defects of traditional evidence collection methods in terms of convenience, universality, and security, laying the foundation for subsequent accurate and efficient automated investigations.

[0016] In one embodiment, step 1 includes: broadcasting a data packet carrying a handshake feature code in the forensic terminal within the same wireless local area network and receiving a device fingerprint confirmation response returned by the mobile phone to be investigated; performing link discovery and a three-way handshake connection between the mobile phone to be investigated and the forensic terminal to establish a basic communication link; sending a vulnerability exploit payload to the acquisition application through the basic communication link to obtain the highest system management privileges; and conducting an environmental stability assessment of the basic communication link based on the privilege level quantification value and network bandwidth, transmission delay, and packet loss rate to obtain a communication link with privilege escalation status; performing a joint digital signature on the session identifier and privilege escalation status identifier in the communication link with privilege escalation status to obtain an encrypted binding token; and deploying the encrypted binding token at both ends of the communication link with privilege escalation status and locking the input and output port permissions to generate a privilege escalation acquisition environment.

[0017] The relevant operational details are as follows: First, the basic communication link is discovered and established. This process begins by connecting the evidence collection terminal with the backend management software installed and the mobile phone to be investigated, which has the data collection application pre-installed, to the same wireless local area network. In actual police applications, the police data collection software APP is directly installed on the victim's mobile phone that needs evidence collection, while the backend management software is installed on the police officer's dedicated evidence collection computer, and both share the same WIFI. Using wireless WIFI for communication not only completely eliminates the physical constraints of data cables, but also solves the pain points of evidence collection operations that require connecting the mobile phone with a data cable, manually enabling USB debugging mode, and installing specific mobile phone drivers during conventional investigations. At the same time, the police data collection APP has powerful update compatibility technology, supporting different mainstream smart operating system versions. For different iterations of operating systems such as Android 10 to 13 or iOS, the R&D and maintenance personnel only need to maintain the successful installation and execution of the APP on different system versions, without having to constantly update it according to the complex mobile phone brands and models on the market, fundamentally solving the problem of traditional evidence collection software requiring frequent firmware updates. The forensic terminal is a portable computer or dedicated server with powerful computing and storage capabilities. The background management software running on it serves as the command and control center for the entire exploration mission. The data acquisition application installed on the mobile phone to be explored is a lightweight, low-power proxy service program whose core function is to listen for commands, execute low-level operations, and transmit data back. Within the same network segment, such as the forensic terminal with IP addresses 192.168.1.10 and the mobile phone to be explored with IP addresses 192.168.1.15, the forensic terminal continuously sends specially crafted UDP packets to the network broadcast address (192.168.1.255). The payload of this packet contains a predefined, highly unique handshake signature, such as an identifier consisting of the hexadecimal byte sequence 0xA1B2C3D4E5F67890. The data acquisition application on the mobile phone to be explored runs silently in the background, specifically listening for UDP broadcasts on a specific port. When the received packet payload perfectly matches the built-in handshake signature, a response mechanism is triggered. In response, the data collection application will construct an acknowledgment packet containing its own device fingerprint and unicast it to the IP address of the forensic terminal. This device fingerprint information is a structured data object, such as a JSON string: {"device_model":"Model-X","os_version":"Android 14","mac_address":"AA:BB:CC:DD:EE:FF","app_instance_id":"UUID-12345"}, used to explicitly identify the device.After receiving this confirmation response and verifying its legitimacy, the forensic terminal initiates a standard TCP three-way handshake to establish a stable, reliable, connection-oriented TCP long connection with the mobile phone to be investigated on a designated port. At this point, a basic communication link for subsequent command exchange and data transmission is established.

[0018] Secondly, automatic privilege escalation and environmental stability assessment are performed at the operating system level. After the basic communication link is stably established, the forensic terminal sends a privilege escalation command encapsulating an exploit payload to the data collection application through this link. This exploit payload is carefully designed binary code targeting known security vulnerabilities in a specific version of the operating system of the mobile phone being investigated. After receiving the command, the data collection application executes this payload in the operating system kernel mode or through a high-privilege system service, exploiting the vulnerability to bypass system security restrictions, thereby obtaining the highest administrative privileges that are normally inaccessible to users, such as root privileges in the Android system. After successful privilege escalation, the data collection application generates a privilege escalation status identifier to accurately record and represent the current privilege status. This identifier is a structured data object, such as a JSON object, whose specific content is {"status":"root","level":2,"timestamp":"2026-04-23T10:00:00Z"}. In this object, "status":"root" clearly indicates that the highest privileges have been acquired; "timestamp" records the time of successful privilege escalation for traceability and auditing; and "level":2 is a quantitative representation of the privilege level. For example, setting a quantitative value for the privilege level of a regular user. The value is 0, while the value for semi-administrative privileges obtained through system application vulnerabilities is 1, and the value for obtaining full root privileges is 0. The value is 2. In this embodiment, root privileges were successfully obtained, therefore... =2.

[0019] Meanwhile, to ensure the stability and reliability of subsequent data transmission, the forensic terminal will send a series of probe data packets bidirectionally through the basic communication link to measure and accurately calculate the key performance indicators of the current wireless network in real time, such as network bandwidth. 95Mbps, average round-trip latency The response time was 12ms, and the packet loss rate was... The value is 0.1%, or 0.001. After obtaining the quantified value of the permission level and network performance parameters, the backend management software will score the data based on a comprehensive environmental stability assessment model. The calculation formula is as follows: In one embodiment, the environmental stability assessment of the basic communication link is performed based on the quantified value of the permission level, network bandwidth, transmission delay, and packet loss rate, including: the environmental stability assessment of the basic communication link is performed using the following formula: in, Quantification value for permission level, For network bandwidth, For transmission delay, For packet loss rate, This is the permission weight coefficient. These are the network weight coefficients. This is the attenuation control coefficient. Environmental stability is scored. Among them, This is the authority weighting coefficient, and its value is relatively high, such as 0.7, which is intended to emphasize the importance of obtaining high authority for the exploration environment; This refers to network weighting coefficients, such as 0.3, used to balance the impact of network quality; while This is the attenuation control coefficient, a relatively large positive number such as 100. Its function is to non-linearly amplify and penalize the packet loss rate through an exponential decay function, because the loss of any data packet can have a fatal impact on the integrity of electronic evidence. These three weighting coefficients were determined by domain experts based on extensive experience data from real cases, combined with the stringent standards of evidence stability in judicial investigations, and after repeated optimization through multiple rounds of experimental testing and statistical analysis. They aim to reasonably balance the relative importance of access level, network quality, and packet loss risk in the comprehensive score. Substitute the above example values ​​into the formula for calculation: =0.7×2+0.3×(95 / 12)×e^{-100×0.001}=3.548. The backend management software has a preset security evidence-gathering threshold, for example, 2.8. This threshold is preset by evidence-gathering experts based on practical experience and judicial evidence requirements, representing the minimum acceptable quality of the investigation environment. Because the calculated... (3.548) is significantly higher than this threshold (2.8), and the system determines that the current environment meets the requirements of high stability and high privilege. Therefore, the acquisition application injects the aforementioned privilege escalation status identifier, that is, an object containing information such as {"status":"root","timestamp":"..."}, into the context of the current TCP session, forming a communication link with an explicit privilege escalation status.

[0020] Finally, the privilege escalation state is encrypted and bound to the communication link. To prevent the privilege escalation session from being maliciously hijacked, tampered with, or replayed by a third party, the acquired privilege escalation state is strongly bound to the current unique communication link. This process first extracts the session identifier (e.g., the session identifier SID-XYZ789 generated by the TCP protocol stack and the privilege escalation state identifier generated earlier) from the communication link with the privilege escalation state. Subsequently, the background management software concatenates these two core pieces of information into a single string and calls its built-in asymmetric encryption module. This module uses the RSA encryption algorithm and digitally signs the SHA-256 hash digest of the string using a pre-generated and securely stored forensic terminal private key, such as a 2048-bit key. The result of this signing process is a unique and unforgeable encrypted binding token. This token is immediately transmitted through the secure link to the data collection application on the mobile phone to be investigated. The two ends of the link—forensics... The terminal's backend management software and the data collection application on the mobile phone being investigated load this token into the memory of their respective processes. Before each data exchange, both parties verify this token using the forensic terminal's public key to ensure the session has not been tampered with. Furthermore, the data collection application, leveraging its newly acquired highest administrative privileges, calls the operating system's underlying interfaces to lock the network input / output ports associated with this TCP connection and set firewall rules to allow communication only with the forensic terminal's IP address, thereby isolating the link at the operating system level and preventing interference from any other unauthorized processes. Through this series of operations, a completely isolated, authenticated, state-locked, and privilege-elevated data collection environment with the highest operational privileges is finally constructed.

[0021] In step 2, based on the collection scope instructions input by the investigators, the communication flow data, network flow data, and fund flow data of the mobile phone under investigation are selectively extracted and traversed in an elevated collection environment to obtain the conventional case-related data packets. Correspondingly, in traditional mobile phone electronic data forensics, a method of full-disk mirroring or indiscriminate batch data export is often used. This not only consumes a lot of investigation time but also generates massive amounts of redundant information, placing a huge burden on subsequent data analysis and clue assessment. This indiscriminate collection method is severely lagging behind the high timeliness and accuracy required for investigating cybercrime cases. Therefore, step 2 aims to abandon this inefficient model by introducing an intelligent, targeted data extraction mechanism based on investigator instructions. This enables precise screening and rapid extraction of high-value information such as communication flow, network flow, and fund flow from massive amounts of data, thereby greatly improving evidence collection efficiency and ensuring that the obtained evidence is highly relevant to the core of the case.

[0022] In one embodiment, step 2 includes: extracting and mapping the time window, target application package name, domain name involved in the case, and fund account keywords in the collection scope instruction to obtain an extraction feature rule set; performing targeted traversal of the target application database according to the extraction feature rule set in the privileged collection environment, and performing information relevance scoring and threshold screening on candidate data records based on time proximity, keyword matching density, and interaction frequency to obtain multi-dimensional case-related information; performing timestamp standardization and deduplication cleaning on the multi-dimensional case-related information, constructing a hierarchical structure object according to communication flow, network flow, and fund flow categories, encrypting it, and attaching a hash verification file header to obtain a regular case-related data packet.

[0023] The relevant operational details are as follows: First, the collection scope instruction is parsed and the feature extraction rule set is generated. This process is carried out in the privileged collection environment constructed in the aforementioned steps. Investigators input a structured collection scope instruction on the backend management software interface of the evidence collection terminal, based on the specific circumstances of the case. Based on this instruction mapping mechanism, users can directly skip redundant data unrelated to the case, achieving selective and targeted collection. The user-selectable extraction scope broadly includes address books, call logs, call recordings, online chat software such as WeChat and QQ, relevant websites, relevant apps, relevant mini-programs, and fund transfer records. This precise division perfectly covers and responds to the three core evidence collection contents of communication flow, network flow, and fund flow required by the Ministry of Public Security for investigating internet-related cases. Simultaneously, thanks to the concurrent processing capabilities of the privileged environment, the collected data can be transmitted in real time to the police evidence collection software in the management backend, enabling rapid comprehensive comparison and cross-correlation of the basic information of the person being collected with local data from the collected mobile phone and cloud data in actual combat. This targeted traversal and real-time feedback mechanism completely solves the problem of the previous method of collecting evidence by physically mirroring the entire machine, which consumed a lot of valuable case-handling time. The instruction is a JSON object containing multi-dimensional constraints, the content of which is set by the investigators according to the needs of the case, for example: {"time_window":["2026-04-20T00:00:00Z","2026-04-23T12:00:00Z"],"target_apps":["com.tencent.mm","com.eg.android.AlipayGphone"],"domains":["suspicious-trade.com"],"keywords":["Zhang San","139********","Specific bank card number 6228..."]}. This instruction clearly defines the scope of a targeted data extraction task. Its core requirement is to conduct in-depth investigation of the data from the target applications WeChat (com.tencent.mm) and Alipay (com.eg.android.AlipayGphone) installed on the device within a precise time window from 00:00 on April 20, 2026 to 12:00 noon on April 23, 2026. The investigation will focus on all network activity records related to the suspicious domain "suspicious-trade.com," as well as all information containing any one or more keywords such as "Zhang San," the phone number "139********," or "specific bank card number 6228..." appearing in the application data. This instruction is sent to the data collection application on the target mobile phone via an established encrypted communication link. The data collection application has a built-in JSON deserialization module, which can instantly parse the instruction into an operable data structure within the program.

[0024] After parsing, the data acquisition application starts the rule mapping engine to convert these high-level semantic constraints into low-level, executable extraction rules. The specific mapping process is as follows: 1) The time window (time_window) is converted into a timestamp range suitable for database queries. For example, the start time 2026-04-20T00:00:00Z and the end time 2026-04-23T12:00:00Z will be converted into Unix timestamps 1776528000 and 1776816000, and used in the WHERE clause of the SQL statement, such as msg.createTime BETWEEN 1776528000 AND 1776816000. 2) The target application package name (target_apps), such as com.tencent.mm for WeChat and com.eg.android.AlipayGphone for Alipay, is precisely mapped to their application data root directory in the phone's storage, such as / data / data / com.tencent.mm / . Within this directory, the target application database storing core data is further located. For example, WeChat chat logs are typically located in encrypted database files such as EnMicroMsg.db under the MicroMsg / subdirectory. 3) Domains involved in the case, such as suspicious-trade.com, are converted into SQL query conditions for fuzzy matching in the browser history database, such as urlLIKE '%suspicious-trade.com%'. 4) Fund account keywords, such as name, phone number, or bank card number, are compiled into a series of regular expressions or SQL LIKE clauses for searching in fields such as message content, contact notes, or transaction details. All these converted low-level rules are encapsulated and aggregated to form a structured set of feature extraction rules.

[0025] Secondly, perform targeted data extraction and information relevance scoring in the privilege escalation collection environment. The collection application, leveraging its acquired highest administrative privilege, can bypass the sandbox isolation mechanism of the Android operating system and directly access the target application database file located in the previous step. It starts a targeted traversal of these databases according to the instructions in the extraction feature rule set. For example, it may execute the following SQL statement to query the WeChat database: SELECT talker, content, createTime FROM message WHERE (createTime BETWEEN... AND...) AND (content LIKE '%张三%' OR content LIKE '%specific bank card number%'). By executing a series of such queries, the application obtains a preliminary candidate data record set containing a large number of potentially relevant records.

[0026] To refine truly valuable information from these candidate records, the application performs an information relevance scoring for each record. This scoring aims to quantify the closeness of the association between each piece of data and the case, and its calculation formula is: , in this formula: is the final information relevance score. is the time proximity factor, which measures the closeness of the record occurrence time to the key time point of the case. This factor is obtained by comparing the record time with the center point of the time window set by the investigators in the instruction or the critical moment of the case occurrence. Its value can be set as a function that non-linearly decreases from 1 to 0 as the time difference increases, such as a Gaussian function, to ensure that the record closest to the key time point obtains the highest proximity score. is the keyword matching density, which represents the ratio of the number of keywords matched in the record to the total amount of information. Its calculation method is to divide the total number of keywords clearly hit in the record by the total length of the text of the record (such as the total number of words or total number of characters). For example, if 2 keywords are hit in a chat record containing 4 words, then its matching density is 2 divided by 4, resulting in 0.5. is the interaction frequency factor, which reflects the historical interaction frequency between the communication object associated with the record (such as a certain contact) and the person involved in the case. Obtaining this factor requires a pre-scan of the entire communication database (such as call records, message databases) to count the total number of interactions between the device and all contacts, and then compare and normalize the total number of interactions of the communication object associated with the current record with the highest number of interactions among all contacts, thereby obtaining a frequency score between 0 and 1. , , These are preset weighting coefficients, pre-configured in the backend management software by evidence experts based on case types such as financial fraud and drug trafficking. For example, in a financial fraud case, the weight of keywords related to funds... It may be set to a higher value, such as 0.5. This is a confidence penalty coefficient for the data type, with a value less than or equal to 1. It is used to adjust the initial importance of different data sources, such as direct transfer records. It can be set to 1.0, while the normal web browsing history can be set to 0.8.

[0027] For example, in a WeChat chat message containing the message "Zhang San, final payment transferred to card 6228," the message matches the keywords "Zhang San" and "6228." The total character count for these keywords is 6, with "Zhang San" accounting for 2 and "6228" for 4. The total message length (excluding punctuation) is 11 characters. Therefore, the keyword matching density is... =6 / 11=0.55. This is the time proximity factor of the record. The calculated value is 0.9, and the interaction frequency factor with the other party is... =0.7. If the weight is set to... =0.2, =0.5, =0.3, data type is chat history =0.9, then its score is: =(0.2×0.9+0.5×0.55+0.3×0.7)×0.9=0.5985. After calculating the scores of all candidate records, a dynamic threshold is applied for filtering. This dynamic threshold can be set to a preset fixed value such as 0.55, or it can be adaptively adjusted according to the total number of returned results, such as always retaining the top 20% of records with the highest scores. In this embodiment, the threshold is set to 0.55, so the records with a score of 0.5985 will be retained, while all records with scores below 0.55 will be filtered out. Through this precise filtering, a high-value, low-redundancy, multi-dimensional set of case-related information is finally obtained.

[0028] In particular, during the actual investigation of internet-related cases, especially telecommunications fraud cases, criminal suspects often deliberately fragment a complete investigative interaction into multiple seemingly unrelated applications to circumvent the risk control strategies and tracking of a single platform. Taking a typical telecommunications fraud scenario as an example, the same key investigative entity identifier, such as the mobile phone number 139-XXXX-1234, might only appear as a brief, contentless call in call logs, a few casual greetings without any high-risk sensitive words in WeChat chats, and a small transaction in Alipay transfer records that is unlikely to trigger risk control. When these three records belonging to different applications are examined in isolation, their respective time proximity, keyword matching density, and interaction frequency are all at a very low level. If an atomic, context-aware linear scoring mechanism is used, that is, each record is independently scored for its information relevance, then all three records will be judged as below the screening threshold due to insufficient individual scores, and thus incorrectly filtered out one by one during the data cleaning stage. However, the simultaneous appearance of the same entity identifier in multiple different applications constitutes a crucial and high-value cross-application co-occurrence correlation in cybercrime cases. This behavioral pattern is a powerful characteristic signal of criminal suspects conducting counter-investigation through communications. The blind spot of the original scoring model in perceiving such correlations directly leads to the erroneous discarding of fragmented evidence chains with high investigative value but low individual scores, severely affecting the parallel analysis of cases and the accurate tracing of suspects. Therefore, this application proposes a preferred implementation method to construct a mechanism that can accurately identify, quantify, and utilize such cross-application co-occurrence correlations to non-linearly correct information relevance scores, thereby re-aggregating and recalling key evidence that has been deliberately scattered by criminal suspects.

[0029] Specifically, in a preferred embodiment, in an elevated data collection environment, a targeted traversal of the target application database is performed according to an extraction feature rule set. Candidate data records are then scored for information relevance and screened based on temporal proximity, keyword matching density, and interaction frequency to obtain multi-dimensional case-related information, including: In the privileged data collection environment, the target application database is traversed according to the feature extraction rule set. The read contact list, call logs, web browsing history, and transfer records are matched and extracted to obtain a candidate data record set. First, from the massive amount of raw data, an efficient preliminary screening is performed based on the feature extraction rule set converted from the investigator's instructions. This aims to quickly narrow the analysis scope from the entire mobile phone data to a controllable subset highly relevant to the case, providing high-quality foundational data for subsequent, more refined and complex in-depth correlation analysis. This forms a superset containing all potentially involved information, namely the candidate data record set. In specific implementation, this process follows the previous steps. The collection application, utilizing its acquired highest administrative privileges, can ignore the operating system's sandbox isolation mechanism and directly mount and access these databases according to the precisely mapped target application database file paths within the rule set (e.g., / data / data / com.tencent.mm / MicroMsg / EnMicroMsg.db). The constraints in the rule set, such as time windows and identifiers of the entities involved, have been transformed into specific, executable SQL WHERE clauses. The data collection application iterates through the rule set, executing corresponding query commands for each target database. For example, in the call log database, it executes `SELECT * FROM call_logs WHERE number='139-XXXX-1234' AND date BETWEEN 'start_timestamp' AND 'end_timestamp'`; in the WeChat database, it executes `SELECT * FROM message WHERE content LIKE '%139-XXXX-1234%' AND createTime BETWEEN ...`. All the result records returned by these queries across different applications, no matter how fragmented or seemingly unrelated their content, are uniformly formatted, collected, and merged. For example, for entity 139-XXXX-1234, the application executes the above series of targeted queries to accurately find and extract a matching record from the databases of com.android.dialer (call application), com.tencent.mm (WeChat), and com.eg.android.AlipayGphone (Alipay), and then incorporates these three records together to form the candidate data record set.

[0030] For each record in the candidate data set, the entity identifier is extracted and labeled with the source application number. A binary assignment matrix is ​​constructed using the entity identifier as the row index and the source application number as the column index to generate an entity co-occurrence adjacency matrix. Next, the unstructured, discrete list of records is transformed into a structured mathematical expression that intuitively reflects the occurrence relationship between entities and applications. This is a crucial step in making implicit co-occurrence relationships explicit and computable. The effect of this step is to generate an entity co-occurrence adjacency matrix. It projects all scattered recorded information into a standardized two-dimensional space. The specific implementation process is as follows, and the formula is: The purpose of this formula is to define the assignment rules for each element in the matrix. Its input is the entity identifier. and source applications The output is matrix elements. The value of is either 0 or 1. Let be the value of the element in the k-th row and j-th column of the entity co-occurrence adjacency matrix; This is the identifier for the k-th entity; For the j-th source application; In application The set of all candidate data records contained therein; This is a function that extracts the entity identifier from record r. Using the example above, the entity... For 139-XXXX-1234, the application... These correspond to phone calls, WeChat, and Alipay, respectively. Since the number is recorded in all three applications, it is included in the matrix. The corresponding row vector is [1,1,1].

[0031] The cross-application co-occurrence degree of each entity identifier is calculated by summing the rows of the entity co-occurrence adjacency matrix. Then, the application distribution entropy is calculated based on the proportion of records for each entity in different applications to obtain the cross-application co-occurrence degree and the application distribution entropy. Subsequently, the strength of cross-application co-occurrence is quantified from two complementary dimensions: breadth and morphology. Knowing only that an entity appears in multiple applications is insufficient; it is also necessary to characterize the uniformity of its record distribution across these applications. This allows for the calculation of two key quantitative indicators for each entity, providing a two-dimensional input for constructing a comprehensive enhancement factor. The calculation process for cross-application co-occurrence degree is expressed as follows: The purpose of this formula is to calculate how many different applications an entity spans. Its input is the entity co-occurrence adjacency matrix. medium entity The corresponding row vector. The output is the cross-application co-occurrence of this entity. .in, Let k be the cross-application co-occurrence of the k-th entity; This represents the total number of applications involved in this survey. For entity 139-XXXX-1234, its... =1+1+1=3. The calculation process of applied distribution entropy can be expressed as: The proportion is: The purpose of this formula is to quantify the uniformity of record distribution using information entropy theory. Its input is entities. Number of records in each application The output is the application distribution entropy of this entity. . A higher value indicates a more uniform distribution and better reflects the fragmented nature of communication. In the example, entity 139-XXXX-1234 has one record in each of the three applications, meaning... =1, =1, =1, the total number of records is 3. Therefore, = = =1 / 3. Its application distribution entropy =1.0986, reaching the maximum value in this scenario, indicating that its distribution is extremely uniform.

[0032] Based on cross-application co-occurrence and application distribution entropy, a co-occurrence enhancement factor is calculated for each entity identifier. This enhancement factor is then adjusted in conjunction with the information relevance scores of each record in the candidate data record set to obtain an enhanced candidate data record set. This step is the core of the entire optimization scheme. Its implementation stems from merging the two quantitative indicators calculated in the preceding steps into a single enhancement factor with clear physical meaning that directly impacts the original score, thereby performing a non-linear correction to the scoring system. In this way, the scores of all records associated with high-value cross-application co-occurrence entities will be significantly improved, while the scores of entity records active only in a single application will remain unchanged. The calculation logic for the co-occurrence enhancement factor is as follows: This formula aims to comprehensively calculate the enhancement factor. Its input is... and and two preset parameters and . The maximum enhancement factor is set by experts based on practical experience. For example, setting it to 2.0 means that the score can be increased up to 3 times the original score. This is the entropy sensitivity adjustment coefficient, for example, set to 1.5, used to adjust the dependence of the enhancement effect on the uniformity of distribution. The output is the co-occurrence enhancement factor. For entity 139-XXXX-1234, its enhancement factor is... =1+2.0×((3-1) / (3-1))×(1-e^{-1.5×1.0986})=2.616. After obtaining the entity-specific enhancement factor, the application will initiate a global score correction process to generate the final enhanced candidate data record set. This process is not a simple replacement, but a mapping operation that preserves the original information and only performs non-linear enhancement on the score dimension. The core reason for performing multiplicative non-linear correction on the co-occurrence enhancement factor and the original information relevance score of each record in the candidate data record set is that multiplicative correction can proportionally amplify the scores of all records related to cross-application co-occurrence entities without changing the score of a single application entity record (whose enhancement factor is always 1). This can effectively recall fragmented evidence while maintaining the basic structure of relative importance within the original score system. The score correction process is as follows: , The relevance score of the original information for the i-th candidate data record is calculated by the aforementioned basic implementation process. , The application scores the relevance of the i-th candidate data record after co-occurrence enhancement. In practice, the application iterates through each record in the candidate data record set. First, identify the entity identifier associated with the record. Then, it queries and retrieves the entity corresponding to the pre-calculated set of enhancement factors, with the entity identifier as the key. The values ​​are then multiplied. The final product of this series of calculations is a set with the exact same structure as the original candidate set, but with each record's score field updated to the enhanced score. The new dataset is the augmentation candidate data record set. As mentioned above, the three original records of entity 139-XXXX-1234 have very low scores due to weak features: 0.22 (call), 0.18 (WeChat), and 0.25 (Alipay). In the global correction process, these three records will all be multiplied by an augmentation factor as high as 2.616 because they are associated with entity 139-XXXX-1234. After augmentation, their new scores become: 0.22×2.616=0.57552, 0.18×2.616=0.47088, and 0.25×2.616=0.654.

[0033] An adaptive threshold is applied to the enhanced candidate data record set to obtain multidimensional case-related information. Due to the structural changes in the scoring system, the original fixed or simple dynamic thresholds may no longer be applicable; a new threshold that can automatically adapt to the enhanced scoring distribution characteristics is needed to ensure the accuracy of the screening. The effect of this step is to accurately retain all high-value case-related information under the new scoring distribution while filtering out noise. The calculation logic for the adaptive threshold is as follows: The purpose of this formula is to dynamically set the selection boundaries based on the global statistical characteristics of the enhanced scores. Its input is the arithmetic mean of the enhanced scores for all records. and standard deviation and a preset screening sensitivity coefficient. For example, setting it to 1.0. The output is an adaptive dynamic filtering threshold. This threshold is based on the global mean, shifted downwards by a range controlled by the standard deviation, automatically adapting to the score distribution across different case scenarios. After enhancement, scores for a large number of records associated with cross-application co-existing entities are significantly boosted, pushing up the overall mean and standard deviation, and the threshold is dynamically adjusted accordingly. For example, calculations show that in this scenario... If the threshold is 0.45, then the three records that originally had scores below 0.3 now have new scores of 0.57552, 0.47088, and 0.654, all of which are higher than this adaptive threshold. As a result, they are successfully retained and ultimately output to the multi-dimensional case information set, thus achieving the goal of recalling fragmented evidence.

[0034] Finally, the multi-dimensional case-related information is encapsulated to generate a standard case-related data package. The data collection application receives the multi-dimensional case-related information set after scoring and filtering, and first performs data cleaning and standardization operations. This process includes: removing incomplete or corrupted fields from data records; converting all timestamps in all records (which may come from different applications and have different formats) to Coordinated Universal Time (UTC) and conforming to the ISO 8601 standard format; and deduplicating duplicate records with identical content using methods such as hash comparison. After cleaning and standardization, the application constructs a JSON object with a clear tree-like hierarchical structure according to the inherent logical attributes of the data—communication flows such as call logs, SMS messages, and social media messages; network flows such as browser history and application access records; and financial flows such as payment application transfer records and bank application transaction records. To ensure the confidentiality and integrity of the evidence, the application then calls a high-strength symmetric encryption algorithm such as AES-256 to encrypt the entire JSON object, generating an encrypted data body. The encryption key is dynamically generated by the forensic terminal at the start of the investigation task and distributed to the data collection application through an established secure channel. A structured header is appended to the front of the encrypted data body. This header is a metadata block containing crucial information for verification and tracing, primarily including: device identifiers for the forensic terminal and the mobile phone under investigation, precise timestamps of the start and end of the investigation, a declaration of the hash algorithm type (e.g., MD5 or SHA-256), and a hash checksum calculated from the original unencrypted JSON object. Finally, this hash checksum header is merged with the encrypted data body into a binary stream, and optionally, lossless compression (e.g., Gzip) is performed to generate a single, complete, secure, and evidence-compliant standard data packet.

[0035] In step 3, the file transfer instruction triggered by the synchronized audio files of the online chat software in the privilege escalation acquisition environment is intercepted and redirected at the kernel level, allowing the original audio file to be intercepted and output before the encryption operation is executed. It is understandable that in current cybercrime investigation practices, instant messaging tools such as WeChat and QQ have become key channels for transmitting information related to cases, with voice messages being frequently used due to their intuitiveness and convenience. However, to protect user privacy, these applications typically encrypt the audio immediately after recording and before storing it locally or sending it over the network. This mechanism means that traditional, file system-level forensic tools can only obtain the encrypted ciphertext file when extracting data, and cannot reconstruct the original audio content. This constitutes a fatal evidentiary barrier for cases requiring voiceprint comparison to identify suspects, such as telecommunications fraud. Therefore, step 3 uses deep-level operating system kernel technology to intercept and copy the original plaintext audio data directly from memory in a seamless manner within a tiny time window before the encryption program starts, thereby ensuring that unencrypted original voice evidence that can be used for forensic identification and voiceprint analysis is obtained.

[0036] Figure 3 This is a flowchart of step 3 in the automatic mobile phone information investigation method for internet-related cases according to an embodiment of this application. Figure 3 As shown, in one embodiment, step 3 includes: step 31, locating the original system call service number corresponding to the file transfer instruction in the system service scheduling table and constructing a hook monitoring descriptor containing the original system call service number and the callback function entry address; step 32, when the network chat software triggers the audio transfer instruction, intercepting the register index value based on the hook monitoring descriptor, copying the parameter stream to the kernel address space, and obtaining the audio memory buffer based on the process identifier matching degree and the file header magic number similarity; step 33, asynchronously snapshotting the audio memory buffer before the encryption operation is executed and writing the plaintext audio data to a secure isolation directory to generate a native voice file.

[0037] The relevant operational details are as follows: First, locate the target system call in the operating system kernel and construct a hook monitoring descriptor. This process is executed in the privileged acquisition environment with the highest privileges constructed in steps 1 and 2. The acquisition application utilizes the acquired root privileges to execute privileged instructions to jump from user mode (exception level EL0) to kernel mode (exception level EL1), gaining direct read and write capabilities to the core memory regions of the Android operating system (based on the Linux kernel). Upon entering kernel mode, the primary task of the acquisition application is to locate and parse the system call table. This table is a core data structure in the operating system kernel, maintaining the mapping relationship between all system API functions from user mode to kernel mode, essentially acting as an address book for kernel functions. When saving voice files, regardless of how the upper-level code is encapsulated, online chat software must ultimately call the underlying file writing functions provided by the operating system. In the Android environment, this typically corresponds to core system calls such as `sys_write` or `vfs_write` in the underlying Virtual File System (VFS). The goal of the acquisition application is to intercept calls to these file transfer instructions. By traversing the sys_call_table, you can find the unique system call service number corresponding to the sys_write function. For example, in the Android system based on the ARM64 architecture, this service number is usually 64, i.e., 0x40.

[0038] After locating the service number, the data collection application dynamically constructs a hook monitoring descriptor in kernel memory. This is a data structure defined by this invention for precise monitoring and triggering interception logic. Its core components include: 1) the target system call service number, i.e., the previously located 64; 2) the identifiers of one or more target processes, such as the application package name of WeChat, com.tencent.mm, to ensure that the interception operation only applies to specific applications and avoids interfering with the normal operation of the operating system; 3) an entry address pointing to a callback function (CallbackFunction) written by the data collection application itself and also located in kernel space. After construction, the data collection application will redirect the function pointer corresponding to service number 64 in sys_call_table to this newly created callback function entry address by modifying it. At the same time, it must safely save the original sys_write function pointer for subsequent restoration of the call chain.

[0039] Secondly, comparison, interception, and confidence scoring are performed when the target instruction is triggered. When a user completes a voice recording in WeChat and prepares to send it, the WeChat application calls the file write API, attempting to write the raw audio data in memory, encoded in AMR or SILK formats, to a local temporary file for subsequent encryption and network transmission. This operation triggers the kernel-level hook set up in the previous step, and the operating system's control is instantly transferred to the preset callback function. At this time, according to the system call convention of the ARM architecture, the CPU's general-purpose registers, such as the X8 register, contain key information about this call, such as the system service number of this call. The file callback function first compares the value in the X8 register to confirm whether it is the set target service number 64. If the match is successful, it indicates that a file write operation of the target process has been successfully intercepted. Next, the callback function parses the parameters of this sys_write call from the parameter general-purpose registers such as X1 and X2 or the kernel stack. The most crucial parameters are the memory address pointer to the data source to be written (usually conventionally in the X1 register) and the length of the data (usually conventionally in the X2 register). The callback function immediately copies this memory region, namely the buffer containing the raw audio data, completely to a separate kernel-secure address space managed by the acquisition application.

[0040] To ensure that the captured data is indeed the expected audio file and not irrelevant data, the capture application needs to perform a quick and accurate capture confidence score. The scoring logic is based on the following formula: In this formula, This represents the final intercept confidence score. This is the process identifier matching score. Since the hook itself has already partially limited the process, it can be further verified here. If it is a complete match, it is 1; otherwise, it is 0. It is the file header magic number similarity. The callback function reads the first few bytes of the data buffer copied to the kernel space and compares them with a built-in dictionary containing magic number features of various audio formats such as AMR 2321414D52 and SILK. It gives a similarity score between 0 and 1 based on the degree of matching. It is the size (in bytes) of the currently captured buffer, and It is the average audio buffer size of the application as captured in history. This is an empirical value that can be obtained by conducting multiple tests on the target application in the early stage. For example, it was found through testing that the size of voice messages sent by WeChat is generally concentrated around 3KB. and These are the feature weights corresponding to process matching and magic number matching, respectively, preset by experts, for example... =0.6, =0.4. It is a penalty coefficient for buffer size deviation, such as 0.1, used to reduce the score of data with abnormally large or small sizes.

[0041] If, during an interception, the process is confirmed to be com.tencent.mm, that is... =1, the header byte of the buffer matches the AMR magic number exactly, that is =1, current buffer size It is 3.2KB, while the historical average is 3.2KB. It is 3.0KB. Substituting into the formula, we get: =(0.6×1)+(0.4×1)-0.1×|3.2-3.0| / 3.0=0.9933. This score will be compared with a preset security interception threshold, which is set by experts to balance precision and recall, such as 0.85. Since 0.9933 is much greater than 0.85, the data stream is identified as the original audio to be encrypted and is officially marked as the intercepted audio memory buffer.

[0042] Finally, an asynchronous snapshot is taken and the original call is restored before encryption execution. After confirming that the intercepted audio memory buffer is a high-confidence target, the callback function needs to complete the data saving within a very short time. This time window exists between the file write API being called and WeChat's own encryption function being executed. To ensure efficiency, the callback function will initiate an asynchronous snapshot read operation, that is, initiate a background task that does not block the current process, write the clean plaintext audio bitstream in the audio memory buffer directly to the secure isolation directory specified by the forensic terminal's backend management software through the previously established secure wireless communication link, and persist it in the format of [application name]_[timestamp]_[session object].amr, generating a native voice file with complete voiceprint characteristics. After initiating the asynchronous write instruction, the core task of the callback function is to restore the context to prevent the target application or operating system from crashing. It will retrieve the original sys_write function pointer saved at the beginning of step 1 and return the CPU control and register state to the original API process intact.

[0043] In step 4, the standard data packets and raw voice files involved in the case are uniformly encoded and encapsulated to obtain a standard data exchange packet. It should be understood that in the mobile phone investigation process of internet-related cases, the previous steps have successfully extracted two core but heterogeneous types of electronic data: one is the standard data packets containing structured text information, and the other is the raw voice files existing as unstructured binary streams. If these two types of data sources are processed and transmitted in isolation, not only will their inherent correlation be unreflected, such as which chat session a specific voice message belongs to, but it will also fail to meet the normative requirements of the public security industry for electronic evidence to follow unified standards, facilitate cross-platform exchange, and be accepted by the judiciary. Furthermore, directly transmitting raw data in an unstable wireless transmission environment is highly susceptible to data corruption or loss due to network fluctuations. Therefore, step 4 integrates and transforms these heterogeneous and fragmented raw data into a standardized evidentiary entity that conforms to official specifications, possesses inherent logical connections, and is suitable for reliable transmission in complex network environments—that is, a standard data exchange packet.

[0044] In one embodiment, step 4 includes: mapping the attribute fields of the regular data packets involved in the case to a data dictionary to construct an index description file, and associating the original audio file as a binary attachment to the node path of the index description file to obtain a normalized dataset; calculating the optimal fragment volume and performing binary cyclic segmentation on the normalized dataset based on the current available bandwidth, theoretical maximum bandwidth, and historical retransmission rate, and attaching an auto-incrementing sequence number and cyclic redundancy check code to each data block to obtain a fragment transmission queue; pushing the fragment transmission queue sequentially to the background management software through an encrypted secure channel and performing sequence number reassembly and double hash verification; after the verification is passed, persistently writing the complete data stream to the storage area to obtain a standard data exchange packet.

[0045] The relevant operational details are as follows: First, the standard case-related data packets and native audio files are normalized. This process is executed on the data collection application of the mobile phone to be investigated. The data collection application first loads the encrypted standard case-related data packets generated in the previous step, which contain information such as communication flow, network flow, and fund flow, and decrypts them using a pre-shared key to restore the internal JSON structured data. At the same time, it also reads the native audio files that have been intercepted and stored in a secure isolation directory. The core of the processing lies in mapping and associating these two different data formats according to a preset data dictionary table to construct a unified index description file. This data dictionary table is pre-configured in the data collection application according to industry standards such as the public security information code set or BCP (Data Interchange Standard Format), and it defines the mapping rules from application internal field names to standard intelligence field names. For example, the dictionary stipulates that the "talker" field in WeChat chat records must be mapped to a standard XML tag. <contactid type="wechat">The "amount" field in Alipay transfer records is mapped to... <transactionamount currency="CNY">.

[0046] The data acquisition application iterates through the decrypted JSON object and dynamically generates a structured XML index description file according to the rules of the data dictionary table. For each raw audio file, the application calculates its content's SHA-256 hash value, obtaining a unique 64-character hexadecimal string such as e3b0c442... as the unique identifier for that audio file. Then, in the XML index description file, under the chat history node associated with that audio message, a... <attachment>Or similar tags that contain reference information for the audio file, such as <attachment type="audio / amr" ref="hash:sha256:e3b0c442..." / > In this way, unstructured speech files are treated as binary large objects (BLOBs), logically linked to structured text information through their hash values. Finally, the generated XML index description file is packaged with all referenced native speech files to form a logically unified and formatted normalized dataset.

[0047] Secondly, the optimal fragment size is dynamically calculated based on network status, and binary fragmentation is performed. Before data transmission, to address the challenge of unstable wireless network channels, the previously generated, potentially large, normalized dataset needs to be intelligently fragmented. The data acquisition application has a built-in network status monitoring module, which continuously measures and updates three key network metrics in real time by sending probe data packets: 1) Current available bandwidth. That is, the actual data transmission rate that the current link can achieve; 2) the theoretical maximum bandwidth. 3) Historical retransmission rate This records the proportion of data packets that needed to be retransmitted due to network errors in a recent period. These parameters collectively reflect the current health of the network. Based on these real-time parameters, the data acquisition application uses the following formula to dynamically calculate the optimal fragmentation volume threshold under the current transmission environment. : In this formula, It is a preset system base fragment size, a conservative value that ensures a high success rate of transmission even under extremely poor network conditions, such as 512KB. The ratio reflects the current network utilization rate, and its value is between 0 and 1. It is a bandwidth scaling factor, such as 0.8, used to adjust the sensitivity of fragment size to changes in network bandwidth. The item is a reliability adjustment factor. If the historical retransmission rate is high, the value of this factor will decrease, thereby actively reducing the fragment size to adapt to unstable links.

[0048] For example, if the theoretical maximum bandwidth of the current network Available bandwidth is 300Mbps, measured in real time. 180Mbps, historical retransmission rate It is 2%, which is 0.02. Substitute it into the formula to calculate: Therefore, the optimal fragment size is currently 742.6 KB. The acquisition application will then use this size as a standard to cyclically divide the entire normalized dataset's binary stream. For each data block, a header will be appended, containing a strictly auto-incrementing cryptographic sequence number starting from 0 (for reassembly at the receiver) and a cyclic redundancy check (CRC32) code calculated for that data block (for rapid error detection). All fragmented data blocks are then organized sequentially to form a fragment transmission queue to be sent.

[0049] Finally, data is pushed, reassembled, and double-hash verified via an encrypted channel. The acquisition application utilizes the secure communication channel established in the first step, reinforced with an encrypted binding token, to push data blocks from the fragmented transmission queue one by one to the backend management software of the forensics terminal via application layer protocols such as HTTPS or FTPS, according to their auto-incrementing sequence numbers. The backend management software's listening service receives these data slices on a designated port. Upon receiving each slice, it first verifies its CRC code. If the verification fails, it immediately requests the sender to retransmit the data block with that sequence number. If the verification succeeds, it stores the data block in a memory buffer according to the auto-incrementing sequence number in the data block header and reassembles it in the correct order. Upon receiving the last data block marked with an end-of-line character, the backend management software reconstructs a complete data stream in memory that is completely identical to the one before the sender's segmentation.

[0050] To ensure the absolute integrity and tamper-free nature of the evidence during transmission, a rigorous double hash verification is performed. The backend management software calculates both the MD5 and SHA-256 values ​​for this reconstructed complete data stream. These two calculated hash values ​​are precisely compared with the original MD5 and SHA-256 hash values ​​pre-calculated by the sender and sent along with metadata before the fragmented transmission begins. Only when both hash values ​​match exactly can it be confirmed that the entire data stream was secure, lossless, and untampered during transmission. After successful verification, this verified complete data stream is persistently written from memory to the local disk storage area of ​​the forensic terminal by the backend management software and registered in the backend database as an independent evidence entity, including recording metadata such as its source device, acquisition time, and hash value. At this point, a fully encapsulated, uniformly structured, and rigorously verified standard data exchange packet is officially generated.

[0051] In step 5, the standard data exchange package is structured and parsed and stored to extract a multi-dimensional feature entity set. This multi-dimensional feature entity set is then subjected to association mining and in-depth analysis to obtain a visualized clue topology map. Specifically, the multi-dimensional feature entity set includes characteristics such as personal attributes, social relationships, activity trajectories, and fund flow. In other words, in practical applications of mobile phone electronic data forensics, although the standard data exchange package generated in the preceding steps solves the problems of unified data encapsulation and reliable transmission, it is essentially still a structured raw data set. For investigators, directly facing massive amounts of text-based chat logs, call logs, and transfer records to clarify relationships, discover financial flows, and gain a comprehensive understanding of the case is akin to finding a needle in a haystack—inefficient and prone to missing crucial clues. Therefore, step 5 applies artificial intelligence and data mining technology to conduct in-depth and automated knowledge discovery on this standardized data, automatically extracting discrete data points into interrelated feature entities, and further quantifying the strength of their correlation. Finally, in a highly intuitive and human-cognitive visual topology map, the criminal network and clues hidden behind the data are clearly presented to the investigators.

[0052] In one embodiment, step 5 includes: extracting and standardizing entities from standard data exchange packets based on named entity recognition models and public security domain data dictionaries to obtain a multidimensional feature entity set by extracting and standardizing entities based on person attributes, social relationships, activity trajectories, and fund flows; performing cross-pairing traversal on nodes in the multidimensional feature entity set, calculating topological edge weight values ​​based on interaction frequency, total fund transaction amount, and number of common contacts, and removing weakly related edges below the isolation threshold to obtain a relational topological relationship matrix; performing optimal coordinate allocation on nodes in the relational topological relationship matrix, mapping topological edge weight values ​​to the line width and color saturation of connecting edges, and rendering business icons and attribute labels for various types of nodes to obtain a visualized clue topology map.

[0053] The relevant operational details are as follows: First, multi-dimensional feature entity extraction is performed based on the Named Entity Recognition (NER) model and a public security domain data dictionary. This process is executed in the background management software of the evidence collection terminal. The background management software first unpacks the received standard data exchange packet, separating the XML index description file containing core text information and binary data such as the original audio file as an attachment. Next, the software takes all the text content in the XML file, such as chat logs, contact notes, and transaction remarks, as input and feeds it into a pre-trained Named Entity Recognition (NER) model. This NER model adopts the currently mainstream BERT-BiLSTM-CRF architecture, where the BERT layer is responsible for learning deep word vector representations from massive corpora, the BiLSTM layer captures contextual word order information through a bidirectional long short-term memory network, and the CRF layer ensures the logical validity of the labeled sequence at the model output, such as the start label of an address entity cannot be directly followed by the internal label of a name entity. The training process of this model is based on a massive amount of anonymized real case documents and a corpus manually annotated by experts. The network weight parameters are continuously optimized through the backpropagation algorithm, enabling it to accurately identify and annotate predefined entity categories from unstructured text.

[0054] Working in parallel with the NER model is a public security-related data dictionary. This is a meticulously maintained, structured knowledge base containing a large number of proprietary terms, slang, codes, and patterns specific to the public security industry. For example, it stores the correspondence between bank card BIN numbers and issuing banks nationwide, the domain name formats of common virtual currency trading platforms, and keyword templates for various online fraud tactics. When the NER model performs identification, it simultaneously queries this data dictionary. For instance, the NER model might identify Zhang San as PER (person's name), while the data dictionary can use regular expression matching to identify 139xxxxxxxx as PHONE_NUMBER (phone number) and 622848... as BANK_ACCOUNT (bank account). Through this combination of model and rules, four core entities can be efficiently and accurately extracted from standard data exchange packets: 1) Person attributes (name, ID number, phone number); 2) Social relationships (WeChat / QQ number, caller); 3) Activity trajectory (base station location, IP address, login location); 4) Fund flow (bank account, third-party payment account, transaction amount). All extracted entities undergo a rigorous normalization process, including unifying time zones, removing duplicates, and aligning data formats (e.g., converting all amounts to floating-point numbers in yuan). This process ultimately forms a set of discrete node data of various types, namely a multidimensional feature entity set.

[0055] Secondly, association mining and topological edge weight calculation are performed on the multidimensional feature entity set. The backend management software abstracts each entity in the multidimensional feature entity set, such as a specific person or a specific bank account, as a node in mathematical graph theory. Subsequently, the software starts the association mining engine to perform cross-pairing traversal on all nodes in the node set. Between any two nodes, as long as there is a direct or indirect interaction record in the original data, a candidate connection edge is established between them. For example, if there is a call record between the mobile phone number of entity Zhang San and the mobile phone number of entity Li Si, then an edge is established between the nodes representing these two people; if entity account A has transferred money to entity account B, an edge is also established.

[0056] To differentiate the importance of different relationships, a topological edge weight value needs to be calculated for each edge. This weight value is a comprehensive measure that reflects the tightness of the relationship between two nodes, and its calculation formula is as follows: In this formula, This represents the weight value of the edge connecting node i and node j. This refers to the interaction frequency between these two nodes, specifically the total number of calls, messages, or transfers within a defined time window. (This is used here.) Logarithmic smoothing can effectively suppress the excessive influence of extremely high-frequency interactions on the total weight, making the increase in interactions from 1 to 10 times contribute more to the weight than the increase from 101 to 110 times. It represents the total amount of funds transacted between nodes i and j. It is a normalized capital base used to scale the absolute transaction amount to a range comparable to other dimensions. This value can be preset by experts according to the type of case, for example, it can be set to RMB 1 million for general economic cases. It is the number of common contacts that nodes i and j have, while That is the total number of contacts after deduplication of these two nodes, therefore This essentially calculates the Jaccard similarity coefficient between the two individuals' social circles. , , These are weighting factors corresponding to interaction frequency, financial transactions, and social relationships, respectively, with a total of 1. These factors can be dynamically adjusted by investigators based on the case's focus. For example, in investigating telecommunications fraud cases, the weighting of financial transactions can be adjusted. Increase to 0.6, frequency weighting Set to 0.3, social weight Set it to 0.1.

[0057] For example, node Zhang San and node Li Si communicated a total of 20 times during the exploration period, that is... =20, the total amount of funds transferred is 50,000 yuan, that is =50000, Zhang San has 50 contacts, and Li Si has 80, of which 15 are mutual contacts. =15, =50 + 80 - 15 = 115. Substitute into the formula to calculate: After calculating the weights of all node pairs, an adjacency matrix in graph theory is formed. The software then performs graph pruning on this matrix, removing all weights below a preset isolation threshold, such as 0.1. This threshold is used to filter out weakly related edges that are accidental or have weak connections, ultimately generating a sparse but core relational topology matrix.

[0058] Finally, visualization rendering is performed to generate a clue topology graph. The backend management software imports the purified relational topology matrix into its front-end visualization graphics rendering engine. This engine can be built based on WebGL technology or mature graph visualization libraries (such as D3.js or G6) to ensure high-performance real-time rendering of large-scale relational networks. The engine first calls and executes a parameter-optimized force-directed layout algorithm, which distributes all entity nodes with random coordinates in a two-dimensional or three-dimensional virtual canvas space during the initialization phase. Subsequently, the algorithm enters an iterative calculation cycle. In each iteration, it simulates two core physical forces: one is a repulsive force based on Coulomb's law, acting between all node pairs, causing them to push each other away to avoid visual overlap and confusion; the other is an attractive force based on Hooke's law, acting only between node pairs connected by edges defined in the matrix, pulling them closer like a spring. The spring constant is strictly proportional to its topological edge weight, meaning that the closer the node pair is, the greater the attractive force it experiences. The algorithm continues to iterate until the total energy of the entire network system converges to a minimum value, that is, the resultant force on all nodes tends to balance. After this calculation, the entire network will automatically expand to form a final stable structure with a reasonable layout and closely related nodes automatically aggregating into clusters. Each node is assigned the optimal display coordinates on the canvas.

[0059] Next, the engine transforms the edge weights into specific visual elements using a non-linear (e.g., logarithmic) mapping function to adapt to human visual perception. This mapping ensures that even small differences in weight values ​​produce visually recognizable changes. Edges with higher weights have thicker geometric linewidths, and the increase in width follows a preset mapping curve. For example, the linewidth change from a weight of 0.1 to 0.5 may be much greater than the change from 0.5 to 0.9, highlighting the establishment of key connections. Edges also have higher color saturation and are rendered along a precisely defined color gradient (e.g., a smooth transition from light gray, representing a weak correlation (hexadecimal value #CCCCCC), to deep red, representing a strong correlation (#B71C1C)). To further enhance information delivery, for relationships with clear directionality (such as fund flows), the engine also renders an arrow at the target node of the connecting edge to clearly indicate the flow direction.

[0060] Simultaneously, the engine matches and renders Scalable Vector Graphics (SVG) format business icons for different types of nodes to ensure icon clarity at any scaling level. For example, a human icon represents a person, a bank icon represents a fund account, and a geographic location icon represents an activity trajectory. Furthermore, the visual size of a node can be dynamically set to be proportional to a core metric of the node (such as the node's degree centrality, i.e., the number of connected edges, or the total transaction amount of fund-related nodes), making the core hub nodes in the network readily apparent. When an investigator's mouse hovers over any node or edge, an event listener is triggered, instantly displaying a richly formatted floating information box. This information box structurally displays all extracted attribute tags of the entity (such as name, card number, transaction time, transaction details, etc.) in key-value pairs, and can embed hyperlinks, allowing users to directly trace back to the corresponding original evidence item after clicking. In addition, the rendering engine provides a complete set of interactive tools, including infinite zooming via the mouse wheel, panning and dragging on the canvas, and a search bar that supports keyword and attribute filtering, allowing investigators to quickly locate and focus on targets of interest in complex networks. Clicking on a node highlights it and fixes it in its current position, or selectively expands its first- or second-order related nodes, temporarily hiding other irrelevant entities, thereby enabling layer-by-layer in-depth analysis of complex networks. Such a static data matrix is ​​transformed into a dynamic, interactive, and information-rich visual topology map of clues, intuitively presented to investigators for efficient case linking and in-depth source tracing and analysis.

[0061] In summary, the automatic mobile phone information investigation method for internet-related cases based on the embodiments of this application is explained, aiming to solve the core pain points of existing evidence collection methods, such as strong dependence on physical cables, cumbersome operation procedures, lack of targeted data extraction, and inability to effectively obtain dynamically encrypted evidence. This solution completely eliminates the dependence on data cables, manual activation of debugging mode, and installation of specific drivers by constructing a wireless, automatically privileged acquisition environment, thus solving the technical problems of poor mobility and complex operation. Furthermore, the solution abandons the time-consuming full-disk backup mode, instead performing targeted traversal and selective extraction of key data such as communication, network, and financial flows based on investigation instructions, significantly improving evidence collection efficiency. Crucially, by employing kernel-level interception and redirection technology, it can directly intercept native voice files before the chat software performs encryption operations, effectively overcoming the difficulty of obtaining voiceprint evidence using traditional methods, and providing key technical support for case investigation.

[0062] Figure 4 This is a block diagram of an automatic mobile phone information investigation system for internet-related cases, according to an embodiment of this application. Figure 4 As shown, the automatic mobile phone information investigation system 100 for internet-related cases according to an embodiment of this application includes: an escalation status and link binding module 110, used to automatically escalate the privileges of the mobile phone to be investigated with the collection application installed and the evidence collection terminal with the background management software installed after connecting them to the same wireless local area network, and bind the escalation status with the established wireless communication link to obtain an escalation collection environment; and a conventional case-related data packet generation module 120, used to perform targeted traversal and selection of communication flow data, network flow data and fund flow data of the mobile phone to be investigated in the escalation collection environment based on the collection range instructions input by the investigators. The system includes a data exchange module 130 for extracting data from the data exchange data and the native audio file. The module is used to intercept and redirect file transfer instructions triggered by the synchronized audio files of the network chat software in the privilege escalation acquisition environment at the kernel level, intercepting and outputting the native audio file before the encryption operation is executed. A unified encoding and encapsulation module 140 is used to uniformly encapsulate the data exchange data and the native audio file to obtain a standard data exchange packet. A visual clue topology generation module 150 is used to perform structured parsing and storage of the standard data exchange packet to extract the multi-dimensional feature entity set, and to perform association mining and in-depth analysis on the multi-dimensional feature entity set to obtain a visual clue topology map.

[0063] Here, those skilled in the art will understand that the specific operations of each step in the aforementioned automated mobile phone information investigation system for internet-related cases have been referenced above. Figures 1 to 3 The method for automatic investigation of mobile phone information in internet-related cases is described in detail here, and therefore, its repeated description will be omitted.< / attachment> < / transactionamount> < / contactid>

Claims

1. A method for automatic investigation of mobile phone information in internet-related cases, characterized in that, include: Step 1: After connecting the mobile phone to be investigated with the data collection application and the evidence collection terminal with the background management software to the same wireless local area network, perform automatic privilege escalation operation at the operating system level on the data collection application and bind the privilege escalation status with the established wireless communication link to obtain the privilege escalation data collection environment. Step 2: Based on the collection range instructions input by the survey personnel, perform targeted traversal and selective extraction of the communication flow data, network flow data, and fund flow data of the mobile phone to be surveyed in the privileged collection environment to obtain the conventional case-related data packets. Step 3: Intercept and redirect the file transfer instruction triggered by the network chat software when synchronizing audio files in the privilege escalation acquisition environment at the kernel level, and intercept and output the original voice file before the encryption operation is executed; Step 4: Perform unified encoding and encapsulation on the regular case-related data packets and the original voice files to obtain standard data exchange packets; Step 5: Perform structured parsing and storage of the standard data exchange package to extract the multi-dimensional feature entity set, and perform association mining and in-depth analysis on the multi-dimensional feature entity set to obtain a visualized clue topology map.

2. The method for automatic mobile phone information investigation in internet-related cases according to claim 1, characterized in that, The multidimensional feature entity set includes features such as personality attributes, social relationships, activity trajectories, and fund flow.

3. The automatic mobile phone information investigation method for internet-related cases according to claim 1, characterized in that, Step 1 includes: By broadcasting data packets carrying handshake signatures in the evidence collection terminal within the same wireless LAN and receiving the device fingerprint confirmation response returned by the mobile phone to be investigated, a link discovery and three-way handshake connection are performed between the mobile phone to be investigated and the evidence collection terminal to establish a basic communication link. The exploit payload is sent to the collection application through the basic communication link to obtain the highest management privileges of the system. The environmental stability of the basic communication link is evaluated based on the privilege level quantification value, network bandwidth, transmission delay and packet loss rate to obtain the communication link with privilege escalation status. A joint digital signature is performed on the session identifier and the privilege escalation status identifier in the communication link with privilege escalation status to obtain an encrypted binding token. The encrypted binding token is then deployed at both ends of the communication link with privilege escalation status and the input and output port permissions are locked to generate a privilege escalation acquisition environment.

4. The automatic mobile phone information investigation method for internet-related cases according to claim 3, characterized in that, An environmental stability assessment of the basic communication link is conducted based on the quantified value of the permission level and network bandwidth, transmission delay, and packet loss rate. This includes assessing the environmental stability of the basic communication link using the following formula: , in, Quantification value for permission level, For network bandwidth, For transmission delay, For packet loss rate, This is the permission weight coefficient. These are the network weight coefficients. This is the attenuation control coefficient. Assess environmental stability.

5. The method for automatic mobile phone information investigation in internet-related cases according to claim 1, characterized in that, Step 2 includes: Extract and map the time window, target application package name, domain name involved in the case, and fund account keywords in the collection scope instruction to obtain the feature rule set; In the privileged data collection environment, the target application database is traversed in a targeted manner according to the feature extraction rule set, and the candidate data records are scored for information relevance and screened by threshold based on time proximity, keyword matching density and interaction frequency to obtain multi-dimensional case-related information. After performing timestamp standardization and deduplication cleaning on multidimensional case-related information, hierarchical structure objects are constructed according to communication flow, network flow, and fund flow categories. These objects are then encrypted and a hash verification file header is attached to obtain regular case-related data packets.

6. The method for automatic mobile phone information investigation in internet-related cases according to claim 1, characterized in that, Step 3 includes: Locate and transfer the original system call service number corresponding to the file instruction in the system service scheduling table, and construct a hook monitoring descriptor containing the original system call service number and the callback function entry address; When the online chat software triggers the audio transfer instruction, the register index value is compared and intercepted based on the hook monitoring descriptor. The parameter stream is copied to the kernel address space and the interception confidence score is calculated based on the process identifier matching degree and the file header magic number similarity to obtain the audio memory buffer. Before the encryption operation is performed, an asynchronous snapshot of the audio memory buffer is read and plaintext audio data is written to a secure isolated directory to generate a native voice file.

7. The method for automatic mobile phone information investigation in internet-related cases according to claim 1, characterized in that, Step 4 includes: Data dictionary mapping is performed on the attribute fields of the regular data packets involved in the case to construct an index description file, and the original voice file is associated as a binary attachment to the node path of the index description file to obtain a normalized dataset; Based on the current available bandwidth, theoretical maximum bandwidth and historical retransmission rate, the optimal fragmentation volume is calculated and binary cyclic segmentation is performed on the normalized dataset. An auto-incrementing sequence number and cyclic redundancy check code are added to each data block to obtain the fragmentation transmission queue. The fragmented transmission queue is pushed to the background management software in sequence through an encrypted secure channel, and the sequence number is reassembled and double hash verification is performed. After the verification is successful, the complete data stream is persistently written to the storage area to obtain the standard data exchange packet.

8. The method for automatic mobile phone information investigation in internet-related cases according to claim 2, characterized in that, Step 5 includes: Based on the named entity recognition model and the public security domain data dictionary, entity extraction and normalization cleaning of standard data exchange packets are performed to obtain a multi-dimensional feature entity set by extracting and normalizing the attributes of the person, social relationships, activity trajectory and fund flow. Cross-pairing traversal is performed on the nodes in the multidimensional feature entity set. The topological edge weight values ​​are calculated based on the interaction frequency, total amount of fund transactions and number of common contacts. Weakly associated edges below the isolation threshold are removed to obtain the associated topological relationship matrix. Optimal coordinate allocation is performed on the nodes in the association topology matrix, the topology edge weight values ​​are mapped to the line width and color saturation of the connecting edges, and business icons and attribute labels are rendered for various types of nodes to obtain a visual clue topology map.

9. The method for automatic mobile phone information investigation in internet-related cases according to claim 5, characterized in that, In the privileged data collection environment, the target application database is traversed according to the feature extraction rule set. Candidate data records are then scored for relevance based on temporal proximity, keyword matching density, and interaction frequency, and thresholded to obtain multi-dimensional case-related information, including: In the privileged data collection environment, the target application database is traversed according to the feature extraction rule set, and the read address book, call records, web browsing history and transfer records are matched and extracted to obtain a candidate data record set; Extract the entity identifiers contained in each record in the candidate data record set and label the source application number. Construct a binary assignment matrix with the entity identifier as the row index and the source application number as the column index to generate the entity co-occurrence adjacency matrix. The cross-application co-occurrence degree of each entity identifier is calculated by summing the rows of the entity co-occurrence adjacency matrix, and the application distribution entropy is calculated based on the proportion of records of each entity in different applications, so as to obtain the cross-application co-occurrence degree and the application distribution entropy. Based on cross-application co-occurrence and application distribution entropy, co-occurrence enhancement factors are calculated for each entity identifier. The co-occurrence enhancement factors are then corrected with the information relevance scores of each record in the candidate data record set to obtain an enhanced candidate data record set. Adaptive threshold filtering is performed on the enhanced candidate data record set to obtain multidimensional case-related information.

10. An automated mobile phone information investigation system for internet-related cases, characterized in that, include: The privilege escalation status and link binding module is used to automatically escalate the privilege of the mobile phone to be investigated with the data collection application installed and the evidence collection terminal with the background management software installed after they are connected to the same wireless local area network. The module binds the privilege escalation status to the established wireless communication link to obtain the privilege escalation data collection environment. The standard case data packet generation module is used to perform targeted traversal and selective extraction of communication flow data, network flow data and fund flow data of the mobile phone to be investigated in an elevated collection environment based on the collection range instructions input by the investigators to obtain standard case data packets. The native voice file output module is used to perform kernel-level interception and redirection copying of file transfer instructions triggered when synchronizing audio files in network chat software in the privilege escalation acquisition environment, and to intercept and output the native voice file before the encryption operation is executed; The unified encoding and encapsulation module is used to encapsulate and encapsulate regular data packets involved in cases and native audio files in a unified manner to obtain standard data exchange packets. The visualization clue topology generation module is used to perform structured parsing and storage of standard data exchange packets to extract multi-dimensional feature entity sets, and to perform association mining and in-depth analysis on the multi-dimensional feature entity sets to obtain a visualization clue topology map.