A network traceability method, device, storage medium and electronic equipment
By constructing a network tracing method, analyzing enterprise network flow data and entity characteristics, and generating multi-dimensional digital profiles, the problem of locating individuals in network economic crimes has been solved. This has enabled the tracing of behind-the-scenes controllers and the automated consolidation of evidence, thereby improving the efficiency of case analysis and the reliability of evidence.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING HONGTENG INTELLIGENT TECH CO LTD
- Filing Date
- 2026-04-20
- Publication Date
- 2026-07-21
Smart Images

Figure CN122432402A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of computer technology, and in particular to a network tracing method, apparatus, storage medium, and electronic device. Background Technology
[0002] In recent years, with the rapid development of internet technology, economic crime cases have shown a trend of involving huge sums of money, being highly concealed, and expanding internationally. Criminal gangs often conceal their true internet activity, posing a significant obstacle to investigators in locating the devices used in the crimes and analyzing the movements of the perpetrators.
[0003] Currently, the investigation and tracing of such cyber economic crimes mainly rely on traditional fund flow analysis and business registration and equity penetration. However, facing the highly sophisticated modus operandi of human-machine separation, the relevant technologies lack effective tools for in-depth analysis of online clues, often only able to locate shell companies or unsuspecting individuals acting as cover. The investigative methods in these technologies struggle to effectively link initial fragmented clues with the physical terminals, real identities, and modus operandi of the actual controllers behind the scenes, leading to technical bottlenecks such as long case investigation cycles, difficulty in locating the actual controllers, and challenges in securing core electronic evidence. Summary of the Invention
[0004] This specification provides a network tracing method, apparatus, storage medium, and electronic device, the technical solutions of which are as follows: Firstly, embodiments of this specification provide a method for tracing network origins, the method comprising: In response to the enterprise identification information entered by the client, the enterprise network flow tracing process is triggered in the network feature resource cluster based on the enterprise identification information to extract the target enterprise network flow data and the enterprise-related entity feature data. Based on the target enterprise network flow data and the enterprise associated entity feature data, entity interaction features and underlying hardware fingerprints are extracted and the associated topology information of the entity interaction features and the underlying hardware fingerprints is established to identify the physical terminal that generates enterprise business traffic and the business personnel associated with the physical terminal. The storage medium information of the physical terminal and the network behavior trajectory of the personnel in the business position are extracted from the network feature resource cluster. The underlying hardware fingerprint, the storage medium information and the network behavior trajectory are aggregated to obtain a multi-dimensional digital profile. Based on the multi-dimensional digital profile, a suspicion assessment result for the personnel in the business position and an evidence guidance list for the physical terminal are generated.
[0005] Secondly, embodiments of this specification provide a network tracing device, the device comprising: The tracing module is used to respond to the enterprise identification information entered by the client, and trigger the enterprise network flow tracing process in the network feature resource cluster based on the enterprise identification information to extract the target enterprise network flow data and the enterprise related entity feature data. The parsing module is used to parse and extract entity interaction features and underlying hardware fingerprints based on the target enterprise network flow data and the enterprise associated entity feature data, and to establish the association topology information of the entity interaction features and the underlying hardware fingerprints in order to locate the physical terminal that generates enterprise business traffic and the business personnel associated with the physical terminal. The aggregation module is used to extract the storage medium information of the physical terminal and the network behavior trajectory of the personnel in the business position from the network feature resource cluster, and aggregate the underlying hardware fingerprint, the storage medium information and the network behavior trajectory to obtain a multi-dimensional digital profile; The generation module is used to generate a suspicion assessment result for the personnel in the business position and an evidence guidance list for the physical terminal based on the multi-dimensional digital profile.
[0006] Optionally, the device is also used for: Acquire multi-source heterogeneous underlying data for the task of tracing the source of economic crimes online. The multi-source heterogeneous underlying data includes raw network flow data, open source intelligence data, and third-party government and enterprise basic data. Feature modeling is performed on the multi-source heterogeneous underlying data to construct a network feature resource cluster representing the association of network entities. The network feature resource cluster includes a basic data resource library, an identity archive library, a black and gray industry tag library, and a computer-based intelligence database.
[0007] Optionally, the step of performing feature modeling on the multi-source heterogeneous underlying data to construct a network feature resource cluster representing the association of network entities includes: The enterprise identifier and legal person association information are extracted from the aforementioned third-party government and enterprise basic data and stored in the basic data resource library; The original network flow data is parsed to obtain service interaction messages, and digital signatures and underlying hardware fingerprints that identify the real operator are extracted from the service interaction messages. Using the enterprise identifier as a mapping benchmark, the business interaction message, the digital signature, the underlying hardware fingerprint, and the open-source intelligence data are cross-domain associated and fused to generate entity profile data representing the mapping relationship between virtual identity and physical enterprise and store it in the identity profile database. Based on a preset black and gray industry rule model, the entity file data is risk-identified to generate corresponding black and gray industry tags and stored in the black and gray industry tag library. Based on the original network flow data and the open-source intelligence data, the terminal operation trajectory data corresponding to the underlying hardware fingerprint is extracted, and the underlying hardware fingerprint and the terminal operation trajectory data are associated and stored in the computer-side intelligence database. Based on the aforementioned basic data resource library, identity archive library, black and gray industry tag library, and computer-based intelligence database, a network feature resource cluster representing the association of network entities is constructed.
[0008] Optionally, establishing the association topology information between the entity interaction features and the underlying hardware fingerprint to identify the physical terminal generating enterprise business traffic and the business personnel associated with the physical terminal includes: Based on the entity interaction features, calculate the abnormal interaction frequency of the physical terminal logging into different enterprise entity accounts within a preset period, and extract the IP drift frequency and operation timing features corresponding to the physical terminal based on the entity interaction features. Based on the abnormal interaction frequency, the IP drift frequency and the operation timing characteristics, multi-dimensional data classification processing is performed to obtain the business job role corresponding to the physical terminal. Establish a mapping topology relationship between the business role, the underlying hardware fingerprint, and the entity interaction features to identify the physical terminal that generates enterprise business traffic and the business personnel associated with the physical terminal.
[0009] Optionally, the step of aggregating the underlying hardware fingerprint, the storage medium information, and the network behavior trajectory to obtain a multi-dimensional digital profile, and generating a suspicion assessment result for the personnel in the business position and an evidence guidance list for the physical terminal based on the multi-dimensional digital profile, includes: Based on the underlying hardware fingerprint, the storage medium information, and the network behavior trajectory, the computer device configuration parameters and software static operation data corresponding to the physical terminal are determined, as well as the virtual identity identifiers and enterprise mapping information associated with the personnel in the business positions are determined. The storage medium information, the computer device configuration parameters, and the software static operation data are aggregated to generate documents and software static behavior features. The network behavior trajectory and the network download interaction records of the physical terminal are aggregated to generate network and trajectory dynamic behavior features. The virtual identity identifier and the enterprise mapping information are aggregated to generate virtual identity and associated enterprise features. The underlying hardware fingerprint is fused with the static behavioral features of the document and software, the dynamic behavioral features of the network and trajectory, and the virtual identity and associated enterprise features to construct a multi-dimensional digital profile for the personnel in the business positions. Based on the multidimensional digital profile, abnormal behavior of the personnel in the business positions is evaluated to obtain the suspicion assessment result, and the distribution path of electronic evidence is analyzed from the multidimensional digital profile to generate an evidence guidance list.
[0010] Optionally, the step of analyzing the distribution path of electronic evidence from the multi-dimensional digital profile to generate an evidence guidance list includes: Based on the static behavioral characteristics of documents and software in the multi-dimensional digital profile, the target storage medium range of the physical terminal is determined. Perform a recursive file system scan on the target storage medium range to match the absolute path of the target file including preset business keywords; Extract the file hash value and last modified time corresponding to the absolute path of the target file, and generate an evidence guidance list including the file hash value, the absolute path of the target file, and the last modified time.
[0011] Optionally, extracting the storage medium information of the physical terminal and the network behavior trajectory of the personnel in the business position from the network feature resource cluster includes: Based on the physical terminal and the personnel in the business position, fuzzy semantic diffusion retrieval is performed on the file metadata archived in the network feature resource cluster to obtain fuzzy semantic diffusion retrieval results. Based on the fuzzy semantic diffusion retrieval results, the associated directory of case-related document materials and image materials in the physical terminal is determined, and the storage medium information is determined based on the associated directory. The system retrieves Internet Protocol (IP) access point data corresponding to the personnel in the business positions from the network feature resource cluster. Based on the IP access point data and the geographical location data in the network feature resource cluster, it draws a map to obtain a historical spatiotemporal map of the personnel in the business positions. Based on the historical spatiotemporal map, it extracts the physical coordinate area of the office as the network behavior trajectory.
[0012] Thirdly, embodiments of this specification provide a computer storage medium storing a plurality of instructions adapted for loading by a processor and executing the above-described method steps.
[0013] Fourthly, this specification provides a computer program product storing at least one instruction adapted to be loaded by a processor and to execute the method steps of one or more embodiments of this specification.
[0014] Fifthly, embodiments of this specification provide an electronic device that may include: a processor and a memory; wherein the memory stores a computer program adapted to be loaded by the processor and to execute the above-described method steps.
[0015] The beneficial effects of the technical solutions provided in some embodiments of this specification include at least the following: In one or more embodiments of this specification, the above-described method addresses the pain points in the investigation of cyber economic crimes, such as difficulty in tracing and locating leads due to counter-investigation methods, the inability to effectively bind fragmented network clues to physical entities, blind searching of massive amounts of evidence, and the extreme difficulty in preventing tampering and securing evidence. By constructing a strong correlation topology mapping from virtual enterprise clues to application-layer interaction features and then to the underlying physical hardware fingerprints, cross-domain penetration and locking of the hidden behind-the-scenes actual controllers and their physical operating terminals can be achieved. At the same time, through the fusion technology of multi-dimensional digital profiling, the originally isolated and chaotic massive network spatiotemporal trajectories, hardware environment configurations, and local file directories are reconstructed into a three-dimensional archive. This not only achieves objective and automated suspicion warnings but also extends the system output directly to the physical evidence securing stage, constructing a seamless physical data closed loop from front-end clue tracing to back-end physical evidence seizure, thereby improving the ability and accuracy of network tracing analysis and the rigor of judicial evidence securing. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in the embodiments or prior art of this specification, the drawings used in the description of the embodiments or prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is a schematic diagram of a network tracing system provided in the embodiments of this specification; Figure 2 This is a flowchart illustrating a network tracing method provided in the embodiments of this specification; Figure 3 This is a schematic diagram illustrating the process of constructing a network feature resource cluster according to an embodiment of this specification; Figure 4 This is a schematic diagram of a process for establishing associated topology information provided in the embodiments of this specification; Figure 5 This is a schematic diagram of an abnormal behavior evaluation process provided in the embodiments of this specification; Figure 6 This is a schematic diagram of the structure of a network tracing device provided in the embodiments of this specification; Figure 7 This is a schematic diagram of the structure of an electronic device provided in the embodiments of this specification. Detailed Implementation
[0018] The technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this specification.
[0019] In the description of this specification, it should be understood that the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance. In the description of this specification, it should be noted that, unless otherwise expressly specified and limited, "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or devices. Those skilled in the art can understand the specific meaning of the above terms in this specification based on the specific circumstances. Furthermore, in the description of this specification, unless otherwise stated, "multiple" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship.
[0020] Please see Figure 1 This is a schematic diagram of a network tracing system provided in this specification. Figure 1 As shown, the network tracing system may include at least a client cluster and a service platform 100.
[0021] The client cluster may include at least one client, such as Figure 1 As shown, it specifically includes client 1 corresponding to user 1, client 2 corresponding to user 2, ..., client n corresponding to user n, where n is an integer greater than 0.
[0022] Each client in a client cluster can be an electronic device with communication capabilities, including but not limited to: wearable devices, handheld devices, personal computers, tablets, in-vehicle devices, smartphones, computing devices, or other processing devices connected to a wireless modem. Electronic devices may have different names in different networks, such as: user equipment, access terminal, user unit, user station, mobile station, mobile station, remote station, remote terminal, mobile device, user terminal, electronic device, wireless communication device, user agent or user device, cellular phone, cordless phone, personal digital assistant (PDA), and electronic devices in 5G networks or future evolved networks.
[0023] The service platform 100 can be a standalone server device, such as a rack-mount, blade, tower, or cabinet-type server device, or a workstation, mainframe, or other hardware device with strong computing power; or it can be a server cluster composed of multiple servers. The servers in the service cluster can be composed in a symmetrical manner, wherein each server is functionally and hierarchically equivalent in the transaction chain, and each server can provide services independently. The independent provision of services can be understood as not requiring the assistance of other servers.
[0024] In one or more embodiments of this specification, the service platform 100 can establish a communication connection with at least one client in the client cluster. Based on this communication connection, data interaction is completed during the network tracing process, such as online transaction data interaction. For example, a client typically refers to a browser or terminal used by a case-handling user. The service platform 100 is the initiating end of the interaction, used to input investigative clues (such as the name of the company involved) and receive the analysis results. The server is the backend big data tracing and analysis system that provides network tracing. For example, it includes a network feature resource cluster (data foundation), an artificial intelligence analysis engine (computing core), and a business logic layer (processing flow).
[0025] It should be noted that the service platform 100 establishes a communication connection with at least one client in the client cluster via a network for interactive communication. This network can be a wireless network or a wired network. Wireless networks include, but are not limited to, cellular networks, wireless LANs, infrared networks, or Bluetooth networks. Wired networks include, but are not limited to, Ethernet, universal serial bus (USB), or controller area networks. In one or more embodiments of the specification, technologies and / or formats including Hyper Text Markup Language (HTML), Extensible Markup Language (XML), etc., are used to represent data exchanged over the network (such as target compressed packets). Furthermore, conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), and Internet Protocol Security (IPsec) can be used to encrypt all or some links. In other embodiments, customized and / or dedicated data communication technologies can be used to replace or supplement the aforementioned data communication technologies.
[0026] The network tracing system embodiments provided in this specification and the network tracing methods described in one or more embodiments belong to the same concept. The execution entity corresponding to the network tracing methods involved in one or more embodiments of this specification can be the aforementioned service platform 100. based on Figure 1 The following is a detailed description of the scenario illustrated in the diagram, with reference to specific embodiments.
[0027] In one embodiment, such as Figure 2 As shown, a network tracing method is proposed. This method can be implemented using a computer program and can run on a network tracing device based on the von Neumann architecture. The computer program can be integrated into an application or run as a standalone utility application. The network tracing device can be a server, such as a service platform or server cluster.
[0028] Specifically, the network tracing method includes: S102: In response to the enterprise identification information entered by the client, the enterprise network flow tracing process is triggered in the network feature resource cluster based on the enterprise identification information to extract the target enterprise network flow data and enterprise associated entity feature data. Enterprise identification information refers to a combination of characters that can uniquely identify or vaguely point to a real enterprise entity in a business system. Its forms include, but are not limited to, the enterprise's unified social credit code, business registration name, taxpayer identification number, or legal representative's name.
[0029] A network feature resource cluster refers to a distributed computing and storage platform that pre-aggregates multi-source underlying data. This cluster typically integrates multi-dimensional data warehouses, including but not limited to raw traffic databases, government and enterprise basic information databases, and open-source intelligence databases.
[0030] Target enterprise network flow data refers to data packet records, traffic logs, session metadata (such as 5-tuple information), and application layer business messages that have direct or indirect communication and interaction with the target enterprise in the network communication protocol stack.
[0031] Enterprise-related entity characteristic data refers to entity description data that is physically or logically bound to the target enterprise, such as the internet access IP address, domain name, legal person identity file, and terminal hardware feature code of the enterprise.
[0032] Optionally, the user sends a query request containing enterprise identification information to the server via an application programming interface (API) using a client. Upon receiving the request, the server converts the enterprise identification information into a graph query language (such as a Cypher statement). In the network feature resource cluster with a knowledge graph as its underlying architecture, the server uses the enterprise identification information as the central root node for a breadth-first search. It then performs a topological traversal along a predefined "enterprise-network asset-communication record" relationship, extracting network flow data nodes with network connections to the root node as the target enterprise's network flow data. Simultaneously, it extracts personnel nodes and device nodes with subordinate relationships to the root node as the enterprise's associated entity feature data. This method can rapidly penetrate multiple layers of hidden proxy networks to discover deep, hidden relationships.
[0033] Optionally, after the client submits the enterprise identification information, the server uses a text search engine (such as a search engine component based on inverted index technology) to perform precise or fuzzy matching on a preset government and enterprise basic information table to obtain a standardized globally unique enterprise number. This triggers a cascading query command in the database. Using this globally unique enterprise number as a foreign key, network packet records carrying this number feature are filtered and extracted from the traffic log detail table as the target enterprise's network flow data. Simultaneously, the corresponding entity registration data is extracted from the identity profile dimension table and the asset dimension table as the enterprise's associated entity feature data. This approach has low system overhead and strong compatibility with traditional structured databases.
[0034] S104: Based on the target enterprise network flow data and the enterprise associated entity feature data, analyze and extract entity interaction features and underlying hardware fingerprints, and establish the association topology information of the entity interaction features and the underlying hardware fingerprints to lock the physical terminal that generates enterprise business traffic and the business personnel associated with the physical terminal. Entity interaction characteristics refer to the data payloads generated by a target enterprise during its daily business network communications, which can characterize the attributes of a specific business system or the usage habits of operators. These include, but are not limited to: Hypertext Transfer Protocol header fields, session identifiers, and Secure Sockets Layer fingerprints in application layer interaction messages, as well as digital signatures and encryption feature offsets of specific business software (such as tax filing software and customs declaration systems).
[0035] Hardware fingerprints refer to unique physical identifiers embedded in the underlying physical device, which are extremely difficult to tamper with by conventional software. These include, but are not limited to, network card physical addresses (MAC addresses), hard drive volume label serial numbers, motherboard universal unique identifiers, and central processing unit serial numbers.
[0036] Association topology information refers to a data structure established using graph theory or multidimensional mapping techniques, used to intuitively or logically represent a strongly bound network relationship model between "virtual business accounts - network interaction messages - underlying hardware devices - real operators".
[0037] Business personnel refer to specific roles that perform actual business operations behind physical terminals, inferred from network interaction characteristics. Examples include finance personnel responsible for bulk invoicing, customs declarants responsible for import and export declarations, and business personnel responsible for fund allocation.
[0038] In a schematic manner, the server-side system loads a pre-defined feature library of specific business software communication protocols (such as the message feature rule set of a known tax invoicing system), and uses regular expressions or string matching algorithms to perform high-speed scanning of the target enterprise's network flow data. When a rule is matched, offset truncation technology is used to directly extract the digital certificate signature from the corresponding message as an entity interaction feature, and the device feature code in the same session is extracted as the underlying hardware fingerprint. Subsequently, the rule decision tree is invoked, and a table lookup mapping is performed based on the digital certificate type in the entity interaction features (such as the tax UKey signature corresponding to financial personnel) to determine the business role. Finally, in the relational database, a joint view table containing a hardware fingerprint foreign key, a certificate signature foreign key, and a role primary key is constructed. A many-to-one mapping relationship topology is established at the table level, and the corresponding physical terminal entity and business personnel are located directly based on the primary and foreign key relationships of the data table.
[0039] S106: Extract the storage medium information of the physical terminal and the network behavior trajectory of the personnel in the business position from the network feature resource cluster, and aggregate the underlying hardware fingerprint, the storage medium information and the network behavior trajectory to obtain a multi-dimensional digital profile; Storage medium information refers to structured or semi-structured metadata that reflects the internal file storage status of the physical terminal, obtained through methods such as network stream parsing, terminal probe backhaul, or data snapshots. Specifically, it includes, but is not limited to: the directory hierarchy tree of the file system, hidden disk partition information, the absolute physical path of specific format files (such as reports and vouchers), file size, and characteristic hash values.
[0040] Network behavior trajectories refer to the concrete spatiotemporal activity records left by personnel in the aforementioned business positions when crossing the Internet boundary within a preset time window. Specifically, this includes, but is not limited to: Internet Protocol (IP) address drift records, geographical coordinates of base station access, start and end times of Virtual Private Network (VPN) connections, and access frequency and timing data for specific business interfaces.
[0041] A multidimensional digital profile refers to a highly structured and comprehensive data model (such as a wide table, JSON document, or knowledge graph node set) generated by a computer system in memory or a persistent database. This multidimensional digital profile objectively quantifies the static hardware attributes, dynamic network trajectory, and local storage ecosystem of an entity.
[0042] In a schematic manner, a distributed computing engine (such as Spark / Flink components) extracts corresponding historical storage logs and internet access time-series logs concurrently from the distributed file system of the network feature resource cluster, based on the underlying hardware fingerprint and business personnel identification, using mechanisms such as MapReduce. Natural language processing and directory tree reconstruction are performed on the extracted storage media information, and the execution results are converted into local asset feature vectors. The extracted network behavior trajectory is mapped into a dynamic trajectory vector in a spatiotemporal coordinate system through a Geographic Information System (GIS) service. Then, using the physical terminal entity corresponding to the underlying hardware fingerprint as the core anchor point, feature splicing is used to align the hardware configuration vector, asset feature vector, and dynamic trajectory vector of the underlying hardware fingerprint in spatiotemporal dimensions and fuse them into a multidimensional matrix. The fused multidimensional matrix is persistently stored as a multidimensional digital profile entity archive in a document-oriented NoSQL database (such as MongoDB).
[0043] S108: Based on the multi-dimensional digital profile, generate a suspicion assessment result for the personnel in the business position and an evidence guidance list for the physical terminal.
[0044] Suspicion assessment results refer to the quantitative judgment output on whether personnel in business positions are suspected of participating in economic crimes. It can be expressed as a continuous risk probability value (such as a score of 0-100) or a discrete risk level label (such as high risk, suspected, normal), and is a weighted index representing the degree of actual control they have over the companies involved.
[0045] The evidence guidance list is used to quickly locate and verify the technical metadata of the original physical evidence, including but not limited to: the hard disk partition where the target file is located, the absolute directory path, the file's last modification time, and the file's characteristic hash value.
[0046] In a schematic manner, the multi-dimensional digital profile is processed into feature vectors and input into a pre-trained large-scale model for suspicion assessment. A quantitative score between 0 and 100 is calculated through forward propagation, serving as the suspicion assessment result for the personnel in the aforementioned business position, while simultaneously determining the evidence guidance list for physical targets. When the suspicion assessment result exceeds a preset high-risk threshold, the large-scale model triggers an evidence-gathering script. The evidence-gathering script performs regular expression filtering on the static behavioral feature dimensions of documents and software in the multi-dimensional digital profile, extracting the absolute physical paths of files that conform to the naming rules of the files involved in the case (such as including internal accounts, real transaction records, and proxy opening records). A hash algorithm function is called to calculate and summarize the extracted file paths, automatically rendering and generating a report containing "terminal MAC address, absolute physical path of the file, file size, and hash signature" as the evidence guidance list for the physical terminal.
[0047] Optionally, the following example illustrates the model training process for generating a large model for suspicion assessment: In some embodiments, a pre-trained basic large language model can be obtained and adapted to a knowledge base-based suspicion assessment generation scenario to obtain a large suspicion assessment generation model. However, directly applying the basic large language model to a suspicion assessment generation scenario is often difficult to adapt to new scenarios. Therefore, the basic large language model is first obtained to create an initial large suspicion assessment generation model, and sample data for the new suspicion assessment generation scenario is obtained. This sample data includes suspicion assessments. Since the basic large language model is usually a pre-trained open-source AIGC model with content generation capabilities, this specification only requires adapting it to the suspicion assessment generation scenario. Specifically, the sample data can be used to fine-tune the initial large suspicion assessment generation model. After the model fine-tuning training is completed, a large suspicion assessment generation model adapted to the suspicion assessment generation scenario is obtained.
[0048] Model creation: Obtain the basic large language model, create an initial suspect assessment generation scenario plugin model for the suspect assessment generation scenario, and form an initial suspect assessment generation large model based on the basic large language model and the initial suspect assessment generation scenario plugin model; the basic large language model (MLLM) includes, but is not limited to, the DeepSeek large model, the GPT series large models, etc.
[0049] Sample data acquisition: Acquire sample data for new suspect assessment generation scenarios. This sample data includes the underlying hardware fingerprint, storage medium information, and network behavior trajectory of the suspect assessment generation scenario.
[0050] Sample data annotation: Based on the suspicion assessment generation scenario, the suspicion assessment result label and evidence guidance list label are labeled accordingly. Optionally, in some cases, the evidence guidance list label is empty when the suspicion assessment result label is of the no suspicion type.
[0051] Model training process: Input sample data into the initial suspicion assessment to generate a large model for at least one round of model training. During the model forward training process: Based on the sample data, use the initial suspicion assessment to generate a large model to obtain the predicted suspicion assessment results and a list of predicted evidence guidelines for physical terminals. During the reverse training of the model, the first model loss value is determined by the model loss function based on the predicted suspicion assessment result and the suspicion assessment result label. The second model loss value is determined by the model loss function based on the predicted evidence guidance list and the evidence guidance list label. The combined model loss is the comprehensive model loss of the first model loss value and the second model loss value. Based on this comprehensive model loss, the model parameters are adjusted in the initial large model for generating suspicion assessment to obtain the trained large model for generating suspicion assessment.
[0052] As an illustration, the initial suspect assessment generation scenario plugin model can be created based on a machine learning model.
[0053] Optionally, the model's training termination conditions may include, for example, the loss function value being less than or equal to a preset loss function threshold, or the number of iterations reaching a preset threshold. Specific training termination conditions can be determined based on actual circumstances and are not specifically limited here.
[0054] It should be noted that the machine learning models involved in one or more embodiments of this specification include, but are not limited to, fitting of one or more of the following machine learning models: Convolutional Neural Network (CNN) model, Deep Neural Network (DNN) model, Recurrent Neural Networks (RNN) model, embedding model, Gradient Boosting Decision Tree (GBDT) model, Logistic Regression (LR) model, etc.
[0055] In the embodiments of this specification, the methods described in S102-S108 above address the pain points in the investigation of cyber economic crimes, namely, the difficulty in tracing and locating the source due to counter-investigation methods, the inability to effectively bind fragmented network clues to physical entities, the blind search of massive amounts of evidence, and the extreme difficulty in preventing tampering and securing evidence. By constructing a strong correlation topology mapping from virtual enterprise clues to application-layer interaction features and then to the underlying physical hardware fingerprints, cross-domain penetration and locking of the hidden behind-the-scenes actual controllers and their physical operating terminals can be achieved. At the same time, through the fusion technology of multi-dimensional digital profiling, the originally isolated and chaotic massive network spatiotemporal trajectories, hardware environment configurations, and local file directories are reconstructed into a three-dimensional archive. This not only achieves objective and automated suspicion warning, but also extends the system output directly to the physical evidence securing stage, constructing a seamless physical data closed loop from front-end clue tracing to back-end physical evidence seizure, improving the ability and accuracy of network tracing analysis and the rigor of judicial evidence securing.
[0056] Please see Figure 3 , Figure 3 This is a schematic diagram illustrating the process of constructing a network feature resource cluster as proposed in this specification. Specifically: S202: Obtain multi-source heterogeneous underlying data for the task of tracing the source of economic crimes online.
[0057] The multi-source heterogeneous underlying data includes raw network flow data, open-source intelligence data, and third-party government and enterprise basic data. The anti-economic crime network tracing task refers to the investigative process by which public security, tax, or auditing departments, in response to economic crimes such as money laundering, illegal fundraising, tax evasion, and false invoicing, start from virtual clues in cyberspace, trace back to and locate the real perpetrators and equipment used in the physical space.
[0058] Multi-source heterogeneous underlying data refers to a collection of data with diverse storage formats collected from the network layer, physical device layer, and third-party business application layer.
[0059] Raw network flow data refers to unfiltered or only partially anonymized underlying network communication records, including but not limited to: PCAP format network packets, NetFlow / sFlow format traffic statistics logs, and five-tuple (source IP, destination IP, source port, destination port, transport protocol) session records.
[0060] Open-source intelligence data refers to threat intelligence or contextual data that is automatically collected from the open internet space through legal and compliant means and is helpful in the assessment of economic crimes. This includes, but is not limited to: publicly available dark web forum records of black and gray market transactions, leaked malicious proxy IP pools from open-source communities such as GitHub, and historical DNS records from publicly available domain name systems.
[0061] Third-party government and enterprise basic data refers to the basic physical archive data obtained from government agencies or large enterprises through dedicated line connections or authorized interfaces.
[0062] S204: Perform feature modeling on the multi-source heterogeneous underlying data to construct a network feature resource cluster representing the network entity association; The network feature resource cluster includes a basic data resource library, an identity archive library, a black and gray industry tag library, and a computer-based intelligence database.
[0063] The basic data resource repository serves as the data foundation of the cluster, storing detailed factual data that has undergone preliminary cleaning (such as deduplication, anonymization, and format alignment). The data is presented as flat, wide tables or documents, preserving the maximum information entropy of the original network flow and intelligence, providing a safety net for subsequent complex queries and tracing.
[0064] The identity archive is used to store entity association mapping data that breaks the physical and virtual boundaries of "human-machine-enterprise" (such as binding the ID card of a business legal person, the MAC address of a specific network card, and the virtual account involved in the case under the same entity ID), representing the homogeneity of cross-domain entities.
[0065] The black and gray industry tag library is used to store a set of tags with clear characteristics of illegal and criminal activities after evaluation by a rule engine or machine learning model, as well as the corresponding warning thresholds and hit rule metadata of these tags.
[0066] The computer-based intelligence database focuses on the micro-level intelligence repository of physical hardware and terminal software ecosystems. It specifically stores device-level parameters extracted from underlying traffic, such as operating system fingerprints, specific hacking tools, encrypted communication fingerprints of the business software involved in the case, and directory tree characteristics of local storage media.
[0067] As an illustration, the AI natural language processing engine is invoked to perform named entity recognition on government and enterprise basic data and open source intelligence. At the same time, the deep packet inspection engine is used to extract the application layer payload from the raw network stream data. All the extracted detailed factual data is stored in a distributed search engine in a unified JSON format to build a basic data resource library.
[0068] Using graph embedding algorithms (such as Node2Vec and TransE), IP addresses, enterprise credit codes, and MAC addresses in the basic resource database are mapped into high-dimensional feature vectors. The cosine similarity between different feature vectors is calculated, and nodes with similarity exceeding the same source threshold are connected by edges. This generates a graph network in the graph database containing relationship types such as control, login, and affiliation, thereby constructing an identity archive.
[0069] Pre-trained graph convolutional neural networks (GCNs) are used to classify and predict the subgraph topology in the identity archive (such as identifying star network clustering features), automatically assigning criminal qualitative labels to high-risk entities and related edges, and writing the label data and weights into a key-value database (such as Redis) to build a black and gray industry label library.
[0070] Finally, using session reassembly technology, feature parameters related to terminal device heartbeats and driver API calls are extracted from the network flow to construct device profile families for terminal entities, which are then persisted to a distributed columnar database (such as HBase) to build a computer-side intelligence database.
[0071] In this specification, the above-mentioned methods S202-S204 are used to break down the cross-domain physical and virtual barriers between "human-machine-enterprise", and realize the deep cross-binding and highly structured integration of massive fragmented data. This not only provides an extremely solid and highly consistent underlying data foundation for subsequent intelligent source tracing and multi-dimensional profiling, but also fundamentally improves the system's early perception sensitivity to complex and hidden economic crime networks and the breadth and depth of underlying correlation mining.
[0072] In one feasible implementation, specifically performing feature modeling on the multi-source heterogeneous underlying data as described in S204 to construct a network feature resource cluster representing network entity associations can be done in the following manner: Step A2: Extract enterprise identifiers and legal entity association information from the third-party government and enterprise basic data and store them in the basic data resource library; This example illustrates receiving basic data streams from third-party government and enterprise entities via a distributed message queue. A named entity recognition algorithm is then used to perform context-based semantic sequence labeling on the data stream text. This automatically identifies and extracts strings representing the enterprise's name and specific code format as the enterprise identifier, while simultaneously extracting entity fragments such as names and identification numbers from the surrounding context as legal entity association information.
[0073] Then, the data cleaning component uses regular expressions to verify the enterprise identifier (such as verifying the 18-digit unified social credit code checksum) and performs hash-based anonymization processing on the legal representative's ID number. The extracted and verified enterprise identifier and legal representative association information are serialized into standard data objects and written in batches to the basic data resource library built on the distributed file system using a columnar storage format.
[0074] Step A4: Parse the original network flow data to obtain the service interaction message, and extract the digital signature and underlying hardware fingerprint that identify the real operator from the service interaction message; Business interaction messages refer to the pure application layer payload data obtained by disassembling the original network flow data using network and transport layer protocols, removing underlying addressing and routing encapsulation, and restoring it. These messages contain specific business logic semantics.
[0075] Step A6: Using the enterprise identifier as the mapping benchmark, cross-domain association and fusion of the business interaction message, the digital signature, the underlying hardware fingerprint and the open source intelligence data are performed to generate entity profile data representing the mapping relationship between virtual identity and entity enterprise and store it in the identity profile database. As an illustration, a wide table containing dozens to hundreds of columns of entity files is predefined in a distributed relational data warehouse. Next, large-scale data join table extraction, transformation, and loading operations are performed. Using the enterprise identifier as the global primary key, a left join or full outer join operation is performed, as follows: 1. Write the static business registration information corresponding to the enterprise identifier in the basic resource database into the base column of the wide table; 2. By associating with the network log table, the extracted digital signatures and underlying hardware fingerprints are filled into the virtual credential column and hardware feature column of the wide table; 3. Using the IP address previously used by the hardware fingerprint as a foreign key, risk characteristics (such as the IP belonging to a well-known botnet) are matched in the open-source intelligence table. The matched intelligence tags are then concatenated into a string array and filled into the open-source intelligence extension column of the wide table. Through the row and column concatenation and aggregation dimensionality reduction of the above multiple underlying data tables, the fragmented records originally scattered across multiple databases are transformed into a single database row record containing all mapping dimensions (i.e., the entity profile data). Then, a data commit operation is performed to store this record in a specific schema that serves as the identity profile database.
[0076] Step A8: Based on the preset black and gray industry rule model, perform risk identification on the entity file data, generate corresponding black and gray industry tags, and store them in the black and gray industry tag library; Black and gray market rule models refer to sets of computer logic deployed in the memory of computing nodes to assess the risk of an entity being suspected of economic crimes. Black and gray market rule models include two forms: one is a deterministic Boolean logic rule engine based on the experience of business experts (such as the Drools rule tree); the other is a probabilistic machine learning classifier (such as the random forest model and the graph neural network model) trained based on historical crime sample data.
[0077] Black and gray market labels refer to machine-readable identifiers (usually standardized strings, enumerated dictionary values, or high-dimensional feature vectors) generated by computers and used to qualitatively describe the risk characteristics of entities. Examples include cross-provincial high-frequency fraudulent invoicing, underground money transfer channels, and malicious proxy IP pools.
[0078] The black and gray industry tag library refers to a dedicated data warehouse used to persistently store the aforementioned tags and their metadata (such as tag timestamps, hit rule IDs, and confidence scores).
[0079] Indicatively, a feature engineering script is invoked to perform feature vectorization processing on the input high-dimensional nested JSON format entity archive data, converting it into a graph feature matrix or one-dimensional tensor containing node attributes and edge weights. This feature matrix is then input into a pre-trained graph convolutional neural network or gradient boosting tree model to perform forward propagation calculations.
[0080] The control model outputs a floating-point probability of risk based on the topological relationships between nodes (such as one machine controlling multiple devices) and time-series characteristics (such as high-frequency operations late at night). This floating-point probability is compared with preset confidence thresholds for each crime type. If the probability value of a certain feature dimension exceeds the threshold, a label mapping function is invoked to automatically generate the corresponding black and gray industry label.
[0081] Finally, using message middleware or high-speed caching interfaces, the entity ID and the generated black and gray industry tag object are persistently stored in a high-speed database such as Redis or HBase in key-value pair format to build a black and gray industry tag library.
[0082] Step A10: Extract terminal operation trajectory data corresponding to the underlying hardware fingerprint based on the original network flow data and the open source intelligence data, and associate the underlying hardware fingerprint with the terminal operation trajectory data and store it in the computer-side intelligence database; Terminal operation trajectory data refers to an ordered sequence of network interaction actions and spatial displacements formed on a continuous timeline, with a specific physical terminal (i.e., the underlying hardware fingerprint) as the main body. It not only includes basic timestamps, source IPs, and destination IPs, but also covers higher-order attributes after being augmented by open-source intelligence, such as IP drift frequency, VPN node jump links, dark web access footprints, and detection timing of specific dangerous ports.
[0083] In a schematic manner, using the underlying hardware fingerprint as a filtering condition, the original network flow data containing the fingerprint is intercepted in real time in a stream processing engine (such as Flink), and the access timestamp and public network exit IP address are extracted from the packets. Open-source intelligence data (such as the globally publicly available Tor exit node library, IDC data center IP segment dictionary, and malicious proxy pool) is introduced, and a high-speed collision is performed using a hash table in memory. If the extracted public network exit IP matches the open-source intelligence database, an intelligence tag is added to the IP. The data points with timestamps, original IPs, and intelligence tags are serialized and reassembled in chronological order to generate a multi-dimensional terminal operation trajectory data (e.g., represented as a Markov chain containing jumper topology relationships). Finally, using the underlying hardware fingerprint as a tag index, the reassembled trajectory sequence is directly written into a time-series database specifically optimized for time series analysis, thereby constructing and updating the computer-side intelligence database.
[0084] Step A12: Construct a network feature resource cluster representing the association of network entities based on the basic data resource library, identity archive library, black and gray industry tag library and computer-based intelligence database.
[0085] In the embodiments described in this specification, the aforementioned method is used to accurately transform massive, disordered, and highly disguised raw bitstreams and text into structured holographic entity mapping maps and multi-dimensional temporal trajectory archives. This not only achieves a qualitative leap from fragmented and isolated clues to a three-dimensional chain of evidence, making deeply hidden criminal network topologies impossible to hide, but also provides a data foundation for intelligent source tracing, precise profiling and analysis, and automated physical evidence preservation in upper-level systems.
[0086] Optional, please see Figure 4 , Figure 4 This is a flowchart illustrating a process for establishing association topology information as proposed in this specification. Specifically, the process of establishing the association topology information between the entity interaction features and the underlying hardware fingerprint to identify the physical terminal generating enterprise business traffic and the business personnel associated with that physical terminal can be referenced as follows: Step 302: Calculate the abnormal interaction frequency of the physical terminal logging into different enterprise entity accounts within a preset period based on the entity interaction features, and extract the IP drift frequency and operation timing features corresponding to the physical terminal based on the entity interaction features; Abnormal interaction frequency: refers to the ratio of the number of times the same physical terminal (underlying hardware fingerprint) frequently switches and successfully logs into different corporate accounts with no obvious equity relationship within a preset time window (such as 1 hour or 1 calendar day). This indicator is a quantitative parameter for measuring black market models such as one person controlling multiple enterprises.
[0087] IP drift frequency refers to the rate at which the public IP address bound to a physical terminal changes across network segments or regions within a continuous network session. Extremely high IP drift frequencies typically indicate that the terminal is using a dial-up IP pool or a dynamic VPN network proxy.
[0088] Operation timing characteristics refer to the distribution pattern of network interaction behavior along the time axis. At the computer level, this is represented by a time series matrix, used to characterize abnormal time periods such as high-frequency concurrency late at night and early morning, automated script requests at fixed time intervals, and regular tax filing on non-working days.
[0089] As an illustration, the sliding window function of a stream processing engine (such as Spark Streaming) can be used to count the total number of enterprise account switching of the physical terminal in the past 24 hours in real time, and calculate the abnormal interaction frequency; at the same time, the number of source IP changes in the packets can be counted to obtain the IP drift frequency, and the time distribution waveform of the request can be extracted as the operation timing feature using Fourier transform or time series binning technology.
[0090] Step 304: Perform multi-dimensional data classification processing based on the abnormal interaction frequency, the IP drift frequency and the operation timing characteristics to obtain the business role corresponding to the physical terminal; The extracted abnormal interaction frequency, IP drift frequency, and operation timing features are input into a pre-trained classification model. The classification model uses a pre-calculated classification hyperplane in a high-dimensional feature space to determine the category cluster to which the feature vector falls, thereby outputting the corresponding business role.
[0091] Optionally, the model training process for the classification model is explained below: Model creation: Create an initial classification model for the classification scenario based on the machine learning model; Sample data acquisition: Acquire a large amount of sample data, which includes the frequency of abnormal interactions, the IP drift frequency, and the operation timing characteristics.
[0092] Sample data labeling: Based on the needs of classification scenarios, an expert service is introduced to manually label the sample data with the corresponding business role labels.
[0093] Model training process: Input sample data into the initial classification model for at least one round of model training to obtain the predicted business job roles. Based on the predicted business job roles and business job role labels, use the model loss function to determine the model loss value. Based on the model loss value, adjust the model parameters of the initial classification model until the model training termination condition is met to obtain the classification model.
[0094] Optionally, the model's training termination conditions may include, for example, the loss function value being less than or equal to a preset loss function threshold, or the number of iterations reaching a preset threshold. Specific training termination conditions can be determined based on actual circumstances and are not specifically limited here.
[0095] Step 306: Establish a mapping topology relationship between the business role, the underlying hardware fingerprint, and the entity interaction features to identify the physical terminal that generates enterprise business traffic and the business personnel associated with the physical terminal.
[0096] Schematic illustration: In a graph database based on attribute graphs, an endpoint representing the underlying hardware fingerprint is created or updated. The business role derived in step 304 is assigned as a label to this endpoint, and entity interaction characteristics (such as a specific UKey certificate number) are instantiated as virtual credential nodes. A relationship edge representing the actual control relationship is forcibly generated between these two.
[0097] This specification employs the aforementioned method to transform abstract network counter-surveillance behaviors into multi-dimensional mathematical features that can be objectively calculated by computers. This not only enables rapid and accurate machine-level qualitative profiling of covert operators but also, through topological mapping technology, links and locks down criminal roles (persons), tools (hardware devices), and virtual evidence (interactive features) at the data level. This significantly improves the efficiency of network tracing and identification.
[0098] Optional, please see Figure 5 , Figure 5 This is a flowchart illustrating an abnormal behavior assessment proposed in this specification. Specifically, it involves aggregating the underlying hardware fingerprint, the storage medium information, and the network behavior trajectory to obtain a multi-dimensional digital profile. Based on this multi-dimensional digital profile, it generates a suspicion assessment result for the personnel in the business position and an evidence guidance list for the physical terminal, including: Step 402: Based on the underlying hardware fingerprint, the storage medium information and the network behavior trajectory, determine the computer device configuration parameters and software static operation data corresponding to the physical terminal, and determine the virtual identity identifier and enterprise mapping information associated with the personnel in the business position; Computer equipment configuration parameters refer to structured data that reflects the hardware performance and model characteristics of a physical terminal, including but not limited to: operating system version number, kernel version, CPU model and number of cores, total memory capacity, motherboard brand, and monitor resolution.
[0099] Static software operation data refers to the software environment fingerprint parsed from the terminal storage medium, including but not limited to: a list of specific business software that has been installed, the software running path in the registry, the last modification timestamp in the configuration file, and whether there are installation remnants of specific debugging tools or agent software.
[0100] Virtual identity refers to the collection of various authentication payloads used by business personnel in the digital space, including but not limited to: business system login accounts, encrypted certificate serial numbers, fingerprints associated with third-party social accounts, and token fingerprints for specific sessions.
[0101] Enterprise mapping information refers to the file information of legal entities that are identified through entity association discovery technology and have a subordinate, employment, or actual control relationship with the aforementioned virtual identity, including the enterprise name, taxpayer identification number, and its position in the organizational chart.
[0102] In a schematic manner, the underlying hardware fingerprint is used as an index to perform heuristic scanning on raw network flow data. By comparing with a pre-set fingerprint database, the configuration parameters of the computer device corresponding to the physical terminal are inferred and determined. Deep packet inspection is performed on file directory snapshots in the storage medium information. Feature matching is performed on the directory tree using the file system signatures of known malware or business software to determine the static operation data of the software. Correlation analysis is performed on business request packets in the network behavior trajectory. By intercepting application-layer authentication fields, the virtual identity strongly bound to the hardware fingerprint is determined. Finally, the determined virtual identity is used as the query key to perform a topological traversal in the identity archive, extracting the enterprise entities with which there is a "login-control" logical edge, thereby determining the specific enterprise mapping information.
[0103] Step 404: Aggregate the storage medium information, the computer device configuration parameters, and the software static operation data to generate documents and software static behavior features; aggregate the network behavior trajectory and the network download interaction records of the physical terminal to generate network and trajectory dynamic behavior features; and aggregate the virtual identity identifier and the enterprise mapping information to generate virtual identity and associated enterprise features. Document and software static behavior characteristics refer to the behavioral characteristics that reflect the terminal's local static state, such as determining whether the device belongs to a development machine, a tax return machine, or a black market automated control machine.
[0104] Network and trajectory dynamic behavior characteristics refer to features that reflect the movement state of a terminal in cyberspace. By aggregating trajectory and traffic, it is possible to identify operational behaviors such as abnormally frequent downloads and rapid switching between cross-border nodes.
[0105] Virtual identity and associated enterprise characteristics refer to features that reflect the "person-enterprise relationship" dimension. By aggregating identity identifiers and enterprise mappings, it is possible to characterize the concentration of control power and business coverage of the operator.
[0106] In a schematic manner, the feature engineering component in the distributed computing framework is invoked to obtain the file directory depth and sensitive suffix distribution recorded in the storage medium information. This data is then standardized and normalized with the determined computer device configuration parameters and software static operation data. A matrix concatenation algorithm is used to generate a qualitative tensor representing the physical attributes of the device, i.e., the static behavioral features of documents and software. The time series coordinates in the network behavior trajectory are obtained and spatiotemporally aligned and aggregated with the traffic peaks and protocol distribution in the network download interaction records to generate a dynamic vector reflecting the frequency of behavior and spatial displacement, i.e., the dynamic behavioral features of the network and trajectory. Finally, the locked virtual identity identifier is associated with the corresponding enterprise mapping information through Cartesian product operation or primary key mapping to generate a many-to-many mapping record representing the person and the enterprise entity behind them, i.e., the virtual identity and associated enterprise features.
[0107] Step 406: The underlying hardware fingerprint is fused with the document and software static behavior features, the network and trajectory dynamic behavior features, and the virtual identity and associated enterprise features to construct a multi-dimensional digital profile for the personnel in the business positions; Indicatively, the distributed database's aggregated index engine is invoked, using the underlying hardware fingerprint as the globally unique primary key, to allocate a specific entity storage space in memory. Multi-table joins or mounting operations are performed, using the document and software static behavior features generated in step 404 as the device's static attribute columns, the network and trajectory dynamic behavior features as the device's time-series dynamic list, and the virtual identity and associated enterprise features as the device's organizational attribute tags, all of which are then populated into the entity storage space. Finally, a consistency protocol is used to logically verify and deduplicate all heterogeneous features within this space, generating a highly structured holographic data object, thereby constructing a multi-dimensional digital profile for business personnel.
[0108] Step 408: Based on the multi-dimensional digital profile, perform abnormal behavior assessment on the personnel in the business positions to obtain the suspicion assessment result, and analyze the distribution path of electronic evidence from the multi-dimensional digital profile to generate an evidence guidance list.
[0109] The distribution path of electronic evidence refers to the specific address of electronic data that is directly related to the facts and is hidden in the physical terminal storage medium, including but not limited to: absolute file path, hidden partition identifier, database table entry location or system underlying log sector.
[0110] In step 408, abnormal behavior assessment of the personnel in the business positions is performed based on the multi-dimensional digital profile to obtain the suspicion assessment result, which can be referred to in step S108.
[0111] In one feasible implementation, the specific execution of parsing the electronic evidence distribution path from the multi-dimensional digital profile to generate an evidence guidance list can refer to the following method: Step B2: Based on the static behavioral characteristics of the documents and software in the multi-dimensional digital profile, determine the target storage medium range of the physical terminal; The target storage medium range refers to a set of logical disk partitions or specific directories that are narrowed down based on the characteristics of the profile and are likely to contain evidence related to the case.
[0112] In a schematic way, the static behavioral characteristics of documents and software in the multi-dimensional digital profile are analyzed to extract the installation path prefix of the tools commonly used by the personnel in this business position (such as a specific invoicing software), and the target storage medium range of the physical terminal is determined accordingly (such as locking it to the cache network associated with the software).
[0113] Step B4: Perform a recursive file system scan on the target storage medium range to match the absolute path of the target file that includes preset business keywords; To illustrate, a high-performance file traversal engine is launched within the target storage medium. Using the file system recursive scanning mechanism, all filenames or file header metadata that match preset business keywords (such as fund flow, UKey records) are retrieved, thereby matching and recording the corresponding absolute path of the target file.
[0114] Step B6: Extract the file hash value and last modification time corresponding to the absolute path of the target file, and generate an evidence guidance list including the file hash value, the absolute path of the target file, and the last modification time.
[0115] To illustrate, the kernel-level file reading interface is invoked to quickly calculate the file hash value without changing the file timestamp, and the last modification time in the file metadata is read synchronously. Finally, the above fields are mapped to a structured template to generate an evidence guidance list containing path + hash + time.
[0116] In this specification, a closed-loop chain from virtual clues to physical evidence is constructed using the above method, ensuring the integrity and immutability of electronic data.
[0117] In one feasible implementation, the extraction of the storage medium information of the physical terminal and the network behavior trajectory of the personnel in the business position from the network feature resource cluster can be carried out in the following manner: Step C2: Based on the physical terminal and the business personnel, perform fuzzy semantic diffusion retrieval on the file metadata archived in the network feature resource cluster to obtain fuzzy semantic diffusion retrieval results. Based on the fuzzy semantic diffusion retrieval results, determine the associated directory of case-related document materials and image materials in the physical terminal. Based on the associated directory, determine the storage medium information. Fuzzy semantic diffusion retrieval refers to using word embedding techniques or knowledge graphs to extend the initial search features (such as ledgers) in a semantic vector space. For example, it can expand from ledgers to semantically related terms or file extensions such as transaction history, income and expenditure, xls, pdf, and screenshots, thereby achieving full coverage matching of unstructured data.
[0118] Fuzzy semantic diffusion retrieval refers to the set of file metadata records with high relevance scores matched in the network feature resource cluster after the above semantic extension, including but not limited to file path, file name, file size and feature summary.
[0119] The associated directory refers to the specific storage path that is identified by statistical analysis of the distribution density of the above search results in the physical terminal file system, and that has the highest file clustering degree and has the characteristics of being involved in the case.
[0120] This example illustrates how to obtain core keywords for case-related searches. One approach is to automatically read risk tags preset for the current physical terminal or business personnel from the black and gray market tag library using an automatic tag conversion method, and map these tags to a set of initial core keywords using a preset business ontology library. Another approach is to retrieve business scope descriptions or open-source intelligence text associated with the enterprise identifier from the network feature resource cluster using an intelligence association extraction method, and extract top-ranked terms as core keywords using the TF-IDF algorithm. A third approach is to respond to client input commands using an interactive command parsing method, segment the command text, and extract noun entities as core keywords.
[0121] The system performs fuzzy semantic diffusion retrieval, mapping the core keywords of the case obtained above to a pre-trained high-dimensional semantic vector space (such as Word2Vec or BERT encoding space). It then uses a cosine similarity algorithm to calculate an extended feature word set whose semantic distance from the core keywords is less than a preset threshold (e.g., extending from ledger to ledger, xls, financial statements, screenshots, etc.). Subsequently, the system uses the extended feature word set as a retrieval operator to perform multimodal matching on the file metadata of the archived files belonging to the target physical terminal in the network feature resource cluster, thereby obtaining fuzzy semantic diffusion retrieval results containing file paths, file names, and file types.
[0122] Based on the search results, the associated directories are determined, and hierarchical clustering analysis is performed on the file paths in the fuzzy semantic diffusion search results to calculate the "involvement confidence weight" of each file system directory node. The weight is obtained by comprehensively weighting the number, size and generation frequency of the involved document materials (such as Excel and PDF) and image materials (such as JPG and PNG) in the directory. The physical path with the highest weight and exceeding the preset threshold value is determined as the associated directory.
[0123] Finally, the logical partition identifier (partition number), physical disk index, and mount point attributes of the associated directory are parsed, and these underlying physical distribution parameters are encapsulated into structured storage medium information as static feature inputs for subsequent construction of multi-dimensional digital profiles.
[0124] Step C4: Retrieve the Internet Protocol IP access point data corresponding to the business personnel from the network feature resource cluster. Based on the IP access point data and the geographical location data in the network feature resource cluster, draw a map to obtain the historical access spatiotemporal map of the business personnel. Extract the office physical coordinate area as the network behavior trajectory based on the historical access spatiotemporal map.
[0125] IP access point data refers to the record of the source IP address, port number, and corresponding precise timestamp sequence generated when business personnel conduct network interactions.
[0126] Geographic location data refers to base station location data or broadband access point location database that is pre-stored in the cluster database and maps IP address ranges to physical geographic coordinates (latitude and longitude).
[0127] Historical access spatiotemporal maps refer to digital topology models that reflect the activity range and frequency distribution of target objects by serializing and rendering spatial coordinate points with a time dimension through a geographic information system.
[0128] The physical coordinate area of the office refers to the real physical office address automatically clustered and identified by algorithms based on access frequency, dwell time and business activity (such as activity during high-frequency office hours).
[0129] In a schematic manner, based on the virtual identity of the personnel in the aforementioned business positions, the full IP access point data for the preset monitoring period is retrieved from the network feature resource cluster. Each IP record is then rapidly correlated and matched with the geographical location data in the cluster. A coordinate transformation algorithm is used to convert the IP address into a latitude and longitude point in a coordinate system, and a four-dimensional spatiotemporal coordinate sequence (longitude, latitude, altitude, and time) is generated by combining it with a timestamp. This spatiotemporal coordinate sequence is then input into a GIS rendering engine to perform map drawing. During this process, points are marked on a two-dimensional map, and a time-series connection algorithm is used to connect the points in chronological order to form a trajectory line. Furthermore, a kernel density estimation (KDE) algorithm is used to calculate the access heat value of each area in the space, thereby generating a historical spatiotemporal map that intuitively reflects the intensity of personnel activity.
[0130] This manual introduces fuzzy semantic diffusion retrieval and GIS spatiotemporal clustering algorithms to achieve a breakthrough in accurately extracting core evidence clues from massive amounts of messy raw data. On the one hand, by utilizing the dimensional expansion of the semantic vector space, it solves the problem of criminals evading technical investigation by tampering with file names or hiding directory storage, ensuring comprehensive location and media archiving of case-related documents and images within physical terminals. On the other hand, by performing spatiotemporal topological mapping between virtual IP access point data and multidimensional geographic location metadata, a dynamic behavioral map is constructed and the real physical coordinates of the office are automatically clustered and identified, completely penetrating the virtual disguise layer of cyberspace and achieving accurate tracing of the physical trajectory of crime dens and office environments.
[0131] The following will combine Figure 6 This specification provides a detailed description of the network tracing device provided in the embodiments. It should be noted that... Figure 6 The network tracing device shown is used to execute this specification. Figures 1-6 The methods shown in the embodiments are illustrated for ease of explanation, showing only the parts related to the embodiments of this specification. For specific technical details not disclosed, please refer to this specification. Figures 1-5 The example shown.
[0132] Please see Figure 6 This diagram illustrates the structure of a network tracing device according to an embodiment of this specification. The network tracing device 1 can be implemented as all or part of a device through software, hardware, or a combination of both. According to some embodiments, the network tracing device 1 includes a tracing module 11, a parsing module 12, an aggregation module 13, and a generation module 14, specifically used for: The tracing module 11 is used to respond to the enterprise identification information entered by the client, trigger the enterprise network flow tracing process in the network feature resource cluster based on the enterprise identification information, and extract the target enterprise network flow data and enterprise related entity feature data. The parsing module 12 is used to parse and extract entity interaction features and underlying hardware fingerprints based on the target enterprise network flow data and the enterprise associated entity feature data, and to establish the association topology information of the entity interaction features and the underlying hardware fingerprints, so as to lock the physical terminal that generates enterprise business traffic and the business personnel associated with the physical terminal. The aggregation module 13 is used to extract the storage medium information of the physical terminal and the network behavior trajectory of the business personnel from the network feature resource cluster, and aggregate the underlying hardware fingerprint, the storage medium information and the network behavior trajectory to obtain a multi-dimensional digital profile. The generation module 14 is used to generate a suspicion assessment result for the personnel in the business position and an evidence guidance list for the physical terminal based on the multi-dimensional digital profile.
[0133] It should be noted that the network tracing device provided in the above embodiments is only illustrated by the division of the above functional modules when executing the network tracing method. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the network tracing device and the network tracing method embodiments provided in the above embodiments belong to the same concept, and the implementation process is detailed in the method embodiments, which will not be repeated here.
[0134] The example numbers in this specification are for descriptive purposes only and do not represent the superiority or inferiority of the examples.
[0135] This specification also provides a computer storage medium that can store multiple instructions adapted to be loaded and executed by a processor as described above. Figures 1-5 The network tracing method described in the illustrated embodiment can be found in the following documentation for its specific execution process. Figures 1-5 The specific details of the illustrated embodiments will not be elaborated here.
[0136] This specification also provides a computer program product that stores at least one instruction, said at least one instruction being loaded and executed by the processor as described above. Figures 1-5 The network tracing method described in the illustrated embodiment can be found in the following documentation for its specific execution process. Figures 1-5 The specific details of the illustrated embodiments will not be elaborated here.
[0137] Please refer to Figure 7 This is a structural block diagram of an electronic device provided in an embodiment of this specification. The electronic device in this specification may include one or more of the following components: a processor 1010, a memory 1020, an input device 1030, an output device 1040, and a bus 1050. The processor 1010, memory 1020, input device 1030, and output device 1040 may be connected to each other via the bus 1050.
[0138] Processor 1010 may include one or more processing cores. Processor 1010 connects to various parts of the electronic device using various interfaces and lines, and performs various functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in memory 1020, and by calling data stored in memory 1020. Optionally, processor 1010 may be implemented using at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). Processor 1010 may integrate one or more of a central processing unit (CPU), graphics processing unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the displayed content; and the modem handles wireless communication. It is understood that the modem may also not be integrated into processor 1010 and may be implemented separately through a communication chip.
[0139] The memory 1020 may include random access memory (RAM) or read-only memory (ROM). Optionally, the memory 1020 may include non-transitory computer-readable storage medium. The memory 1020 may be used to store instructions, programs, code, code sets, or instruction sets.
[0140] The input device 1030 is used to receive input instructions or data, and includes, but is not limited to, a keyboard, mouse, camera, microphone, or touch device. The output device 1040 is used to output instructions or data, and includes, but is not limited to, a display device and a speaker. In this embodiment, the input device 1030 can be a temperature sensor for acquiring the operating temperature of the electronic device. The output device 1040 can be a speaker for outputting audio signals.
[0141] In addition, those skilled in the art will understand that the structure of the electronic device shown in the above figures does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements. For example, the electronic device may also include radio frequency circuits, input units, sensors, audio circuits, wireless fidelity (WIFI) modules, power supplies, Bluetooth modules, etc., which will not be described in detail here.
[0142] In the embodiments of this specification, the executing entity for each step can be the electronic device described above. Optionally, the executing entity for each step can be the operating system of the electronic device. The operating system can be Android, iOS, or other operating systems; this specification does not limit this.
[0143] exist Figure 7 In the electronic device, the processor 1010 can be used to call a program stored in the memory 1020 and execute it to implement the network tracing method as described in the various method embodiments of this specification.
[0144] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory, or random access memory, etc.
[0145] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in the embodiments of this specification are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the features, data, and information involved in this specification were all obtained under full authorization.
[0146] The above-disclosed embodiments are merely preferred embodiments of this specification and should not be construed as limiting the scope of this specification. Therefore, any equivalent variations made in accordance with the claims of this specification shall still fall within the scope of this specification.
Claims
1. A network tracing method, characterized in that, The method includes: In response to the enterprise identification information entered by the client, the enterprise network flow tracing process is triggered in the network feature resource cluster based on the enterprise identification information to extract the target enterprise network flow data and the enterprise-related entity feature data. Based on the target enterprise network flow data and the enterprise associated entity feature data, entity interaction features and underlying hardware fingerprints are extracted and the associated topology information of the entity interaction features and the underlying hardware fingerprints is established to identify the physical terminal that generates enterprise business traffic and the business personnel associated with the physical terminal. The storage medium information of the physical terminal and the network behavior trajectory of the personnel in the business position are extracted from the network feature resource cluster. The underlying hardware fingerprint, the storage medium information and the network behavior trajectory are aggregated to obtain a multi-dimensional digital profile. Based on the multi-dimensional digital profile, a suspicion assessment result for the personnel in the business position and an evidence guidance list for the physical terminal are generated.
2. The method according to claim 1, characterized in that, The method further includes: Acquire multi-source, heterogeneous underlying data for the task of tracing the source of economic crimes online; Feature modeling is performed on the multi-source heterogeneous underlying data to construct a network feature resource cluster that represents the association of network entities.
3. The method according to claim 2, characterized in that, The multi-source heterogeneous underlying data includes raw network flow data, open-source intelligence data, and third-party government and enterprise basic data. The network feature resource cluster includes a basic data resource library, an identity archive library, a black and gray industry tag library, and a computer-based intelligence database. The step of performing feature modeling on the multi-source heterogeneous underlying data to construct a network feature resource cluster representing network entity associations includes: The enterprise identifier and legal person association information are extracted from the aforementioned third-party government and enterprise basic data and stored in the basic data resource library; The original network flow data is parsed to obtain service interaction messages, and digital signatures and underlying hardware fingerprints that identify the real operator are extracted from the service interaction messages. Using the enterprise identifier as a mapping benchmark, the business interaction message, the digital signature, the underlying hardware fingerprint, and the open-source intelligence data are cross-domain associated and fused to generate entity profile data representing the mapping relationship between virtual identity and physical enterprise and store it in the identity profile database. Based on a preset black and gray industry rule model, the entity file data is risk-identified to generate corresponding black and gray industry tags and stored in the black and gray industry tag library. Based on the original network flow data and the open-source intelligence data, the terminal operation trajectory data corresponding to the underlying hardware fingerprint is extracted, and the underlying hardware fingerprint and the terminal operation trajectory data are associated and stored in the computer-side intelligence database. Based on the aforementioned basic data resource library, identity archive library, black and gray industry tag library, and computer-based intelligence database, a network feature resource cluster representing the association of network entities is constructed.
4. The method according to claim 1, characterized in that, The step of establishing the association topology information between the entity interaction features and the underlying hardware fingerprint to identify the physical terminal that generates enterprise business traffic and the business personnel associated with the physical terminal includes: Based on the entity interaction features, calculate the abnormal interaction frequency of the physical terminal logging into different enterprise entity accounts within a preset period, and extract the IP drift frequency and operation timing features corresponding to the physical terminal based on the entity interaction features. Based on the abnormal interaction frequency, the IP drift frequency and the operation timing characteristics, multi-dimensional data classification processing is performed to obtain the business job role corresponding to the physical terminal. Establish a mapping topology relationship between the business role, the underlying hardware fingerprint, and the entity interaction features to identify the physical terminal that generates enterprise business traffic and the business personnel associated with the physical terminal.
5. The method according to claim 1, characterized in that, The process involves aggregating the underlying hardware fingerprint, the storage medium information, and the network behavior trajectory to obtain a multi-dimensional digital profile. Based on this multi-dimensional digital profile, a suspicion assessment result for the personnel in the business position and an evidence guidance list for the physical terminal are generated, including: Based on the underlying hardware fingerprint, the storage medium information, and the network behavior trajectory, the computer device configuration parameters and software static operation data corresponding to the physical terminal are determined, as well as the virtual identity identifiers and enterprise mapping information associated with the personnel in the business positions are determined. The storage medium information, the computer device configuration parameters, and the software static operation data are aggregated to generate documents and software static behavior features. The network behavior trajectory and the network download interaction records of the physical terminal are aggregated to generate network and trajectory dynamic behavior features. The virtual identity identifier and the enterprise mapping information are aggregated to generate virtual identity and associated enterprise features. The underlying hardware fingerprint is fused with the static behavioral features of the document and software, the dynamic behavioral features of the network and trajectory, and the virtual identity and associated enterprise features to construct a multi-dimensional digital profile for the personnel in the business positions. Based on the multidimensional digital profile, abnormal behavior of the personnel in the business positions is evaluated to obtain the suspicion assessment result, and the distribution path of electronic evidence is analyzed from the multidimensional digital profile to generate an evidence guidance list.
6. The method according to claim 5, characterized in that, The step of analyzing the distribution path of electronic evidence from the multi-dimensional digital profile to generate an evidence guidance list includes: Based on the static behavioral characteristics of documents and software in the multi-dimensional digital profile, the target storage medium range of the physical terminal is determined. Perform a recursive file system scan on the target storage medium range to match the absolute path of the target file including preset business keywords; Extract the file hash value and last modified time corresponding to the absolute path of the target file, and generate an evidence guidance list including the file hash value, the absolute path of the target file, and the last modified time.
7. The method according to claim 1, characterized in that, The step of extracting the storage medium information of the physical terminal and the network behavior trajectory of the personnel in the business position from the network feature resource cluster includes: Based on the physical terminal and the personnel in the business position, fuzzy semantic diffusion retrieval is performed on the file metadata archived in the network feature resource cluster to obtain fuzzy semantic diffusion retrieval results. Based on the fuzzy semantic diffusion retrieval results, the associated directory of case-related document materials and image materials in the physical terminal is determined, and the storage medium information is determined based on the associated directory. The system retrieves Internet Protocol (IP) access point data corresponding to the personnel in the business positions from the network feature resource cluster. Based on the IP access point data and the geographical location data in the network feature resource cluster, it draws a map to obtain a historical spatiotemporal map of the personnel in the business positions. Based on the historical spatiotemporal map, it extracts the physical coordinate area of the office as the network behavior trajectory.
8. A network traceability device, characterized in that, The device includes: The tracing module is used to respond to the enterprise identification information entered by the client, and trigger the enterprise network flow tracing process in the network feature resource cluster based on the enterprise identification information to extract the target enterprise network flow data and the enterprise related entity feature data. The parsing module is used to parse and extract entity interaction features and underlying hardware fingerprints based on the target enterprise network flow data and the enterprise associated entity feature data, and to establish the association topology information of the entity interaction features and the underlying hardware fingerprints in order to locate the physical terminal that generates enterprise business traffic and the business personnel associated with the physical terminal. The aggregation module is used to extract the storage medium information of the physical terminal and the network behavior trajectory of the personnel in the business position from the network feature resource cluster, and aggregate the underlying hardware fingerprint, the storage medium information and the network behavior trajectory to obtain a multi-dimensional digital profile; The generation module is used to generate a suspicion assessment result for the personnel in the business position and an evidence guidance list for the physical terminal based on the multi-dimensional digital profile.
9. A computer storage medium, characterized in that, The computer storage medium stores a plurality of instructions adapted for loading by a processor and executing the steps of the method as described in any one of claims 1 to 7.
10. An electronic device, characterized in that, include: A processor and a memory; wherein the memory stores a computer program adapted to be loaded by the processor and to execute the steps of the method as described in any one of claims 1 to 7.