File circulation tracing method and device, equipment and medium
By obtaining information from asset and personnel databases, processing log data, entity mapping, and updating the graph database using the knowledge graph model, the problem that file traceability and circulation methods in the existing technology are difficult to cope with complex security needs, and efficient and flexible file traceability and circulation monitoring are achieved.
Patent Information
- Application Number
- CN202510392822.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-06-20
AI Technical Summary
Existing file traceability and circulation methods are difficult to effectively respond to the increasingly complex file flow security needs, especially in the face of security threats such as unauthorized access, tampering and leakage. They lack global view and context information, lack of query flexibility, and lack of performance when handling massive data and complex relationship queries.
By obtaining entity attribute information from the preset asset database and personnel database, receiving log data sent by the target client, preprocessing and identifying, entity mapping, and using the knowledge graph model to write the data into the graph database, continuously tracking the status changes of the file's entire life cycle and updating the graph database, so as to obtain query results related to the traceability file based on query instructions.
It realizes the monitoring of the global view and context information of the file flow process, improves the flexibility and efficiency of query, can efficiently track and trace the file flow process, and optimizes the file traceability and circulation method to cope with complex security needs.
Smart Images

Figure CN120179867A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data security, and particularly relates to a method, device, equipment and medium for tracing the flow of files. Background Art
[0002] Sensitive files usually face security threats such as unauthorized access, tampering and leakage during the flow process. In enterprise compliance management, it is necessary to strictly monitor the access and use of files. The file tracing and flow technology aims to record and monitor the key links of files from creation to deletion, so that enterprises can conduct internal compliance audits to ensure the transparency and security of information during the flow process.
[0003] Existing file tracing and flow methods often adopt digital watermarking or log auditing technology, etc. Among them, digital watermarking technology embeds specific identification information into electronic documents. The embedding and extraction processes of watermarks are complex and the tracing accuracy is limited. It is often used to trace file leakers and is difficult to perform fine-grained tracing for complex flow processes and cannot monitor the file operation status; log auditing technology conducts comprehensive tracing by analyzing the generated audit logs, but usually provides isolated event records, lacks a global view and context information, is difficult to achieve complete tracing, and is limited in query flexibility. Facing dynamically adjusted query requirements, it is necessary to rewrite the query script. When facing complex multi-dimensional relational data, it is usually difficult to trace efficiently and in real time. In addition, most existing systems rely on relational databases and often have insufficient performance when processing massive data and complex relational queries. Therefore, the existing technology is difficult to efficiently meet the increasingly complex file flow security requirements.
[0004] In summary, how to optimize the file tracing and flow method to efficiently meet the increasingly complex file flow security requirements is an urgent problem to be solved at present. Summary of the Invention
[0005] In view of this, the purpose of the present invention is to provide a method, device, equipment and medium for tracing the flow of files, which can optimize the file tracing and flow method to efficiently meet the increasingly complex file flow security requirements. The specific solutions are as follows:
[0006] In the first aspect, the present application provides a method for tracing the flow of files, which is applied to a server and includes:
[0007] Obtain first entity attribute information from a preset asset database and a preset personnel database, and receive log data sent by a target client; the log data is data collected by the target client from a number of target security devices;
[0008] Preprocess the log data, and identify the processed data to obtain second entity attribute information related to sensitive files;
[0009] Perform entity mapping using the first entity attribute information and the second entity attribute information to obtain mapped data;
[0010] Write the mapped data into a graph database according to a preset knowledge graph model, continuously track the change of the full life cycle state of files in the graph database, and update the graph database using the tracked state change information, so as to obtain query results related to the file to be traced from the current latest graph database based on the received query instruction, and feedback the query results accordingly.
[0011] Optionally, before obtaining the first entity attribute information from the preset asset database and the preset personnel database, it further includes:
[0012] Use an asset automatic detection tool to scan the network environment of several target assets, and determine the first asset information of the target assets based on the obtained scan results;
[0013] Receive manually input second asset information through a preset interface, and unify the formats of the first asset information and the second asset information to obtain processed asset information;
[0014] Structurally store the processed asset information in the preset asset database according to the asset type of the processed asset information.
[0015] Optionally, the preprocessing of the log data includes:
[0016] Filter the log data based on a custom filtering rule to obtain target log data related to the sensitive file;
[0017] Parse the target log data into a unified data format, and classify the target log data to obtain classified data;
[0018] Complete the relevant information of the classified data based on the preset asset database and the preset personnel database.
[0019] Optionally, the identification of the obtained processed data to obtain the second entity attribute information related to the sensitive file includes:
[0020] Input the obtained processed data into a queue so that the processed data is used as the real-time data of the queue;
[0021] Read corresponding data to be identified from the real-time data of the queue, and identify the read data to be identified to obtain corresponding identification results;
[0022] Normalize the recognition result according to the log type, and eliminate the ambiguity of the same-name attributes in the recognition result to obtain the second entity attribute information related to the sensitive file.
[0023] Optionally, writing the mapped data into the graph database according to the preset knowledge graph model, and continuously tracking the change of the full life cycle state of the file in the graph database and updating the graph database with the tracked state change information includes:
[0024] Write the mapped data into the graph database in the form of multi-threaded concurrency according to the preset knowledge graph model, and assign a unique file identifier to the sensitive file in the mapped data;
[0025] Continuously track the change of the full life cycle state of the file in the graph database based on the file identifier, and when the full life cycle state of the file in the graph database changes, automatically record the corresponding state change information and update it to the graph database.
[0026] Optionally, obtaining the query result related to the file to be traced from the current latest graph database based on the received query instruction, and making a corresponding feedback on the query result includes:
[0027] Determine the corresponding file to be traced based on the received query instruction, and construct a query condition;
[0028] According to the query condition, perform association analysis on the file to be traced starting from the target node of the current latest graph database to obtain the query result related to the file to be traced, and make a corresponding feedback on the query result.
[0029] Optionally, the file transfer and traceability method further includes:
[0030] Synchronously monitor the event change status of the preset asset database and the preset personnel database based on the preset interface;
[0031] When the event change status is monitored, update the first entity attribute information based on the event change status, so as to update the graph database with the updated first entity attribute information.
[0032] In a second aspect, the present application provides a file transfer and traceability device, which is applied to a server and includes:
[0033] A data receiving module, configured to obtain first entity attribute information from a preset asset database and a preset personnel database, and receive log data sent by a target client; the log data is data collected by the client from a plurality of target security devices;
[0034] A data recognition module, configured to preprocess the log data and recognize the processed data to obtain second entity attribute information related to sensitive files;
[0035] An entity mapping module, configured to perform entity mapping using the first entity attribute information and the second entity attribute information to obtain mapped data;
[0036] A change tracking module, configured to write the mapped data into a graph database according to a preset knowledge graph model, continuously track the full life cycle state changes of files in the graph database, and update the graph database using the tracked state change information, so as to obtain a query result related to a file to be traced from the current latest graph database based on a received query instruction, and feedback the query result accordingly.
[0037] In a third aspect, the present application provides an electronic device, including:
[0038] A memory, configured to store a computer program;
[0039] A processor, configured to execute the computer program to implement the foregoing file transfer and traceability method.
[0040] In a fourth aspect, the present application provides a computer-readable storage medium, configured to store a computer program; wherein, when the computer program is executed by a processor, the foregoing file transfer and traceability method is implemented.
[0041] In this embodiment, first entity attribute information is obtained from a preset asset database and a preset personnel database, and log data sent by a target client is received; the log data is data collected by the client from a number of target security devices; the log data is preprocessed, and the obtained processed data is identified to obtain second entity attribute information related to sensitive files; the first entity attribute information and the second entity attribute information are used for entity mapping to obtain mapped data; the mapped data is written into a graph database according to a preset knowledge graph model, and the full life cycle state change of files in the graph database is continuously tracked, and the graph database is updated using the tracked state change information, so as to obtain a query result related to a file to be traced from the current latest graph database based on a received query instruction, and the query result is fed back accordingly. As can be seen from the above, in this application, first entity attribute information is obtained from a preset asset database and a preset personnel database, and log data sent by a target client is received to preprocess the log data and perform identification to obtain second entity attribute information related to sensitive files. The first entity attribute information and the second entity attribute information are used for entity mapping, and the mapped data is written into a graph database according to a preset knowledge graph model. The full life cycle state change of files in the graph database is continuously tracked, and the graph database is updated using the tracked state change information, so as to obtain a query result related to a file to be traced from the current latest graph database based on a received query instruction, and the query result is fed back accordingly. In this way, through the above process of this application, the graphic structure of the graph database is used to store data, the log data of different security devices can be collected and integrated, a globally associated view is constructed, the integrity of file association is improved, the query result can be fed back according to the received query instruction, the flexibility of the query is improved, and the graph database is generated based on the log data, the operation state of files can be monitored, and then the file tracing and transfer method is optimized to efficiently meet the increasingly complex security requirements of file transfer. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0043] Figure 1 It is a flowchart of a file transfer and tracing method disclosed in this application;
[0044] Figure 2It is a process timing diagram of a file transfer traceability method disclosed in this application;
[0045] Figure 3 It is a schematic diagram of the mapping relationship between entities disclosed in this application;
[0046] Figure 4 It is a schematic diagram of an example of the internal structure of the "person - file" relationship disclosed in this application;
[0047] Figure 5 It is a schematic diagram of the structure of a file transfer traceability device disclosed in this application;
[0048] Figure 6 It is a structural diagram of an electronic device disclosed in this application. Detailed implementation manners
[0049] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0050] Existing file traceability transfer methods often use digital watermarking or log auditing technologies, etc. Among them, digital watermarking technology embeds specific identification information into electronic documents. The process of embedding and extracting watermarks is complex and the tracking accuracy is limited. It is often used to trace the file leaker. It is difficult to perform fine-grained tracking for complex transfer processes and cannot monitor the file operation status; log auditing technology conducts comprehensive traceability by analyzing the generated audit logs, but usually provides isolated event records, lacks a global view and context information, is difficult to achieve complete traceability, and its query flexibility is limited. Facing dynamically adjusted query requirements, it is necessary to rewrite the query script. When facing complex multi-dimensional relational data, it is usually difficult to perform traceability efficiently and in real time. In addition, most existing systems rely on relational databases, and their performance is often insufficient when dealing with massive data and complex relational queries. Therefore, the existing technology is difficult to efficiently meet the increasingly complex file transfer security requirements.
[0051] In order to overcome the above technical problems, this application provides a file transfer traceability method to optimize the file traceability transfer method to efficiently meet the increasingly complex file transfer security requirements.
[0052] See Figure 1 As shown, the embodiments of the present invention disclose a file transfer traceability method, which is applied to a server and includes:
[0053] Step S11: Obtain the first entity attribute information from a preset asset database and a preset personnel database, and receive the log data sent by the target client; the log data is the data collected by the target client from a number of target security devices.
[0054] In this embodiment, first obtain the first entity attribute information from a preset asset database and a preset personnel database, and receive the log data sent by the target client. Among them, the first entity attribute information includes asset attribute information and personnel attribute information; the sources of the log data include security device logs, network traffic, etc., and the log data is the data collected by the target client from a number of target security devices.
[0055] It should be understood that due to the different manufacturers of the number of target security devices, the log formats of different security devices are different. Therefore, the log data is collected through the proxy of the target client, and the target client is docked with the number of security devices by using syslog (a standard for transmitting log messages on the Internet protocol network), jdbc (an application programming interface for standardizing how client programs access databases), etc., so that the target client can collect the log data from the number of security devices. In addition, after the log data is collected, in this embodiment, a detailed audit log can also be generated according to the log data, marking the detailed information of the file, the relevant operator, and recording the relevant file operation behavior for backup, so as to facilitate the management personnel to query and confirm according to the needs. As Figure 2 The figure shows the process timing diagram of a file transfer traceability method provided by this application.
[0056] It should be noted that before obtaining the first entity attribute information, this embodiment needs to collect asset data to construct the preset asset database. The processing flow is as follows: Use an asset automatic detection tool to scan the network environments of several target assets, and determine the first asset information of the target assets based on the obtained scan results; Receive the second asset information manually input through a preset interface, and perform unified processing on the formats of the first asset information and the second asset information to obtain processed asset information; Structurally store the processed asset information in the preset asset database according to the asset types of the processed asset information. Among them, the preset interface can be the front end of the server, and managers can manually input the asset data that needs to be supplemented and added through the preset interface. That is to say, first use an asset automatic detection tool to scan the network environments of several target assets to determine the first asset information of the target assets based on the obtained scan results. At the same time, receive the second asset information manually input by managers through the preset interface, perform unified processing on the formats of the first asset information and the second asset information, classify the processed asset information, and structurally store it in the preset asset database according to the asset types of the processed asset information for subsequent obtaining of corresponding asset attribute information from the preset asset database. In this way, this embodiment generates detailed audit logs based on traditional log auditing technology for subsequent traceability analysis, with a relatively low implementation complexity; Uses log data to identify detailed information of recorded files, relevant operators, operation behaviors, etc., and can provide detailed operation history data for auditing and tracing file operations; Collects and integrates log data of different security devices, details file operation behaviors, and improves the integrity of file association; Uses an asset automatic detection tool combined with manual input to collect asset data, improves the integrity of the data in the preset asset database, and further improves the accuracy of file transfer traceability.
[0057] Step S12: Preprocess the log data, and identify the obtained processed data to obtain second entity attribute information related to sensitive files.
[0058] In this embodiment, the received log data is preprocessed, and the obtained processed data is identified to obtain second entity attribute information related to sensitive files among them. That is to say, the second entity attribute information is sensitive file attribute information. Table 1 shows an entity association relationship provided by this application.
[0059] Table 1
[0060]
[0061] It should be noted that the file transfer traceability method of this application is a method for realizing file transfer traceability by constructing a knowledge graph model. The general representation of the knowledge graph is a "entity-relationship-entity" triple and key-value pairs of entity-related attributes, forming a networked knowledge structure. Among them, the entities in this embodiment are three types: personnel, assets, and sensitive files, which are used for correlation analysis; the assets refer to device information registered within the system and associated with electronic documents, such as host devices, network devices, etc.; the sensitive files are electronic documents containing confidential information, trade secrets, or other important information. As Figure 3 Shown is a schematic diagram of the mapping relationship between entities provided by this application. Among them, the asset can transmit electronic files as a storage medium or a processing device, and each asset has its own responsible person, that is, the personnel; the personnel can perform processing operations such as creating, modifying, copying, and printing electronic files, and transmit electronic files through tools such as email and instant messaging. Since there is an association between the front and back states of the same file, for example, the state connection before and after the same file is renamed, this embodiment abstracts it as a relationship to express the mapping relationship between and within each entity.
[0062] It can be understood that after receiving the log data, for the convenience of subsequent management of the log data, this embodiment can perform preprocessing operations on the log data, and its processing process is as follows: screening the log data based on a custom filtering rule to obtain target log data related to the sensitive file; parsing the target log data into a unified data format, and classifying the target log data to obtain classified data; complementing relevant information of the classified data based on the preset asset database and the preset personnel database. Among them, the types of the target log data include file operation type, email type, instant messaging system type, printing type, burning type, and cloud disk type. That is, first screen and filter the log data according to the custom filtering rule to screen out the target log data related to the sensitive file, reduce the amount of data to be processed, then parse the target log data into a unified intermediate data format, classify the target log data into the preset types of the target log data, and complement relevant information of the classified data based on the preset asset database and the preset personnel database to ensure the integrity of the log data.
[0063] It should be noted that after obtaining the processed data after preprocessing the log data, it is necessary to identify the processed data to obtain the second entity attribute information related to the sensitive file. The processing flow is as follows: Input the obtained processed data into a queue so that the processed data serves as the real-time data of the queue; Read the corresponding data to be identified from the real-time data of the queue, and identify the read data to be identified to obtain the corresponding identification result; Normalize the identification result according to the log type and eliminate the ambiguity of the same-name attributes in the identification result to obtain the second entity attribute information related to the sensitive file. That is, input the obtained processed data into a queue to use the processed data as the real-time data of the queue for batch processing of the processed data. Subsequently, read the corresponding data to be identified from the entity data of the queue and identify the read data to be identified to identify the knowledge elements in the data to be identified, such as sensitive file entities, attributes, association relationships, etc., and pay attention to the file name, classification level, file size, and file type to obtain the corresponding identification result. It can be understood that since the meanings of log fields are different under different file transmission methods, it is necessary to model them as independent relationships separately to avoid errors in the meanings of log fields. That is, normalize the identification result according to the log type. For example, abstract the file operation log as an independent relationship to identify the operation status of the entire life cycle of the file, and use attributes in the relationship to identify specific operation information; At the same time, due to the ambiguity of the same-name attributes between different logs, that is, the same field in different logs may represent different meanings, it is necessary to eliminate the ambiguity between the entity object attributes before performing knowledge fusion. That is, eliminate the ambiguity of the same-name attributes in the identification result to obtain the second entity attribute information related to the sensitive file. In this way, in this embodiment, the log data is preprocessed, the obtained heterogeneous log data is uniformly modeled, and the normalized log data is input into the queue for use as real-time data, improving the processing efficiency of the log data; During the preprocessing of the log data, the relevant information of the data is supplemented using the preset asset database and the preset personnel database to ensure the integrity of the log data.
[0064] Step S13: Perform entity mapping using the first entity attribute information and the second entity attribute information to obtain the mapped data.
[0065] In this embodiment, after obtaining the first entity attribute information and the second entity attribute information, entity mapping is performed using the first entity attribute information and the second entity attribute information to obtain the mapped data. That is, this embodiment needs to associate the entities in each dimension to form a complete knowledge graph model. Specifically, entities and entity relationships of a preset type are extracted according to the log type. Among them, as Figure 4 shown is a schematic diagram of an internal structure example of a "person - document" relationship provided by this application. Through fields that can identify identities in the log fields of the receiving and sending type, such as email addresses, employee numbers, etc., the senders and receivers of sensitive documents are mapped to associate person entities. Based on fields such as the operation name, file information, and employee number in the operation - type logs, the person behaviors associated with the files are mapped. By aligning the same entities in the first entity attribute information and the second entity attribute information, it is ensured that there is only one unique identifier for them in the knowledge graph, and the same relationships are merged, conflicting relationships are processed, and the priority of the relationships is determined to achieve knowledge fusion. In this way, the processing of entity mapping for the first entity attribute information and the second entity attribute information in this embodiment can enable the knowledge graph to form a complex and comprehensive knowledge system, which helps to discover new knowledge, insights, and patterns, thereby supporting more in - depth data analysis and decision - making.
[0066] Step S14: Write the mapped data into the graph database according to the preset knowledge graph model, continuously track the changes in the full - life - cycle state of the files in the graph database, and update the graph database using the tracked state - change information, so as to obtain query results related to the file to be traced from the currently latest graph database based on the received query instruction, and provide corresponding feedback on the query results.
[0067] In this embodiment, the mapped data is written into the graph database according to the preset knowledge graph model, and the changes in the full - life - cycle state of the files in the graph database are continuously tracked to update the graph database using the tracked state - change information, so as to obtain query results related to the file to be traced from the currently latest graph database based on the received query instruction, and provide corresponding feedback on the query results. Among them, the query instruction can be an instruction initiated by a management personnel through a preset interface, and the preset interface can be the front - end of the server.
[0068] It should be noted that after entity mapping, this embodiment needs to write the mapped data into the graph database, track the status changes of the files therein, and update the graph database according to the tracked status change information. The processing flow is as follows: Write the mapped data into the graph database in a multi-threaded concurrent manner according to a preset knowledge graph model, and assign a unique file identifier to the sensitive files in the mapped data; Continuously track the full life cycle status changes of the files in the graph database based on the file identifier, and when the full life cycle status of the files in the graph database changes, automatically record the corresponding status change information and update it to the graph database. That is, write the mapped data into the graph database in a multi-threaded concurrent manner according to a preset knowledge graph model, and assign a unique file identifier to the sensitive files in the mapped data to ensure accurate tracking and facilitate the management of sensitive files. Continuously track the full life cycle status changes of the files in the graph database based on the file identifier, and when the full life cycle status of the file changes, automatically record and update the operations of the file to the graph database. That is, automatically record the corresponding status change information and update it to the graph database to ensure accurate tracing of file operations and transfer records. It can be understood that the construction of the knowledge graph is a continuous iterative update process. To facilitate the update of the graph database, when new data arrives, this embodiment uses the incremental update method for update, that is, taking the currently newly added data as input, adding or updating entities to the current graph database, dynamically adding new nodes and relationships, and constructing a global transfer chain.
[0069] It should be further pointed out that after receiving the query instruction, the processing flow of this embodiment is as follows: determining the corresponding file to be traced based on the received query instruction and constructing query conditions; according to the query conditions, taking the target node of the current latest graph database as the traversal starting point to perform correlation analysis on the file to be traced, so as to obtain query results related to the file to be traced and feedback the query results accordingly. Among them, the target node can be a file node or a personnel node. That is, determining the relevant information of the corresponding file to be traced according to the received query instruction and constructing corresponding query conditions according to the relevant information, so as to perform correlation analysis on the file to be traced with the target node of the current latest graph database as the traversal starting point according to the query conditions, filtering and screening the query results based on the types, attributes of nodes and edges, and the in-out directions of edges, so as to obtain query results related to the file to be traced and feedback the query results accordingly, so as to ensure that the operator who initiates the query instruction can receive the feedback on the query instruction, understand the current life cycle stage of the file, help identify abnormal behaviors, and perform permission control and risk assessment as needed, reduce potential leakage risks, and ensure the security of the file during the transfer process. In addition, due to the existence of directed loops in the graph, this embodiment can adopt a graph traversal algorithm based on depth-first search and create an index for path queries to speed up the query speed and improve the efficiency of file transfer traceability. This embodiment can also support the time query function to expand the functions related to queries and expand the applicable scenarios of the file transfer traceability method.
[0070] It should be noted that since there may be operations such as type addition, editing, and deletion of data in the preset asset database and the preset personnel database, in order to ensure the accuracy of the file transfer traceability method, this embodiment can continuously monitor the change status of the database, and its processing flow is as follows: synchronously monitor the event change status of the preset asset database and the preset personnel database based on a preset interface; when the event change status is monitored, update the first entity attribute information based on the event change status so as to update the graph database by using the updated first entity attribute information. Among them, the preset interface can be a RESTFUL (an application programming interface designed based on the representational state transfer architectural style) interface. That is, synchronously monitor the event change status of the preset asset database and the preset personnel database based on a preset interface, and after the event change status is monitored, update the first entity attribute information based on the event change status so as to update the graph database by using the updated first entity attribute information. In addition, considering the security of the data stored in the graph database, this embodiment can manage the permissions of operators as needed to avoid data leakage. In this way, this embodiment stores data in a graphical structure, constructs a global view of the associated mapping of sensitive files, access personnel, and processing devices, and improves the integrity of file associations; combined with log data, complex relationships and multi-level association queries can be performed, query conditions can be dynamically constructed, global traceability analysis can be realized, different traceability scenarios can be applied, the upper limit of query flexibility can be improved, and the file information security management ability can be enhanced; it can intuitively display the state changes and transfer paths of files in the life cycle, trace file-related operations and personnel, and support efficient queries and analyses to ensure the efficient tracking and traceability of the file transfer process.
[0071] As can be seen from the above, the embodiments of the present application obtain the first entity attribute information from the preset asset database and the preset personnel database, and receive the log data sent by the target client to preprocess and identify the log data to obtain the second entity attribute information related to sensitive files. The first entity attribute information and the second entity attribute information are used for entity mapping, and the mapped data is written into the graph database according to the preset knowledge graph model. The full life cycle state change of the files in the graph database is continuously tracked, and the graph database is updated using the tracked state change information, so as to obtain the query results related to the files to be traced from the current latest graph database based on the received query instructions, and the query results are fed back accordingly. In this way, through the above process of the embodiments of the present application, on the one hand, data is stored in a graphical structure to construct a global view of the associated mapping of sensitive files, access personnel, and processing devices, improving the integrity of file associations; on the other hand, combined with log data, complex relationships and multi-level association queries can be performed, query conditions can be dynamically constructed, global traceability analysis can be realized, different traceability scenarios can be applied, the upper limit of query flexibility can be improved, and the file information security management ability can be enhanced; on the one hand, it can intuitively display the state changes and transfer paths of files during the life cycle, trace file-related operations and personnel, and support efficient query and analysis to ensure the efficient tracking and traceability of the file transfer process; on the one hand, based on traditional log audit technology, detailed audit logs are generated for subsequent traceability analysis, and the complexity is relatively low; on the one hand, an asset automatic detection tool is used in combination with manual input to collect asset data, improving the integrity of the data in the preset asset database, and thus improving the accuracy of file transfer traceability; on the one hand, the log data is preprocessed, the obtained heterogeneous log data is uniformly modeled, and the normalized log data is input into the queue for use as real-time data, improving the processing efficiency of the log data; on the other hand, during the preprocessing of the log data, the preset asset database and the preset personnel database are used to complement the relevant information of the data to ensure the integrity of the log data, and thus optimize the file traceability transfer method to efficiently meet the increasingly complex file transfer security requirements.
[0072] Correspondingly, referring to Figure 5 As shown, the embodiments of the present application further provide a file transfer traceability device applied to a server, including:
[0073] A data receiving module 11, configured to obtain first entity attribute information from a preset asset database and a preset personnel database, and receive log data sent by a target client; the log data is data collected by the client from a plurality of target security devices;
[0074] The data recognition module 12 is configured to preprocess the log data and recognize the processed data to obtain second entity attribute information related to sensitive files;
[0075] The entity mapping module 13 is configured to perform entity mapping using the first entity attribute information and the second entity attribute information to obtain mapped data;
[0076] The change tracking module 14 is configured to write the mapped data into the graph database according to a preset knowledge graph model, continuously track the change of the full life cycle state of the files in the graph database, and update the graph database using the tracked state change information, so as to obtain a query result related to the file to be traced from the current latest graph database based on the received query instruction, and feedback the query result accordingly.
[0077] As can be seen from the above, in the embodiment of the present application, the first entity attribute information is obtained from a preset asset database and a preset personnel database, and the log data sent by the target client is received to preprocess and recognize the log data to obtain second entity attribute information related to sensitive files. The first entity attribute information and the second entity attribute information are used for entity mapping, and the mapped data is written into the graph database according to a preset knowledge graph model. The change of the full life cycle state of the files in the graph database is continuously tracked, and the graph database is updated using the tracked state change information, so as to obtain a query result related to the file to be traced from the current latest graph database based on the received query instruction, and feedback the query result accordingly. In this way, through the above process of the embodiment of the present application, the graphic structure of the graph database is used to store data, the log data of different security devices can be collected and integrated, a globally associated view is constructed, the integrity of file association is improved, the query result can be fed back according to the received query instruction, the flexibility of the query is improved, and the graph database is generated according to the log data, the operation state of the file can be monitored, and then the file tracing and transfer method is optimized to efficiently meet the increasingly complex security requirements of file transfer.
[0078] In some specific embodiments, the file transfer and traceability device may further include:
[0079] The information determination unit is configured to scan the network environments of several target assets using an asset automatic detection tool, and determine the first asset information of the target assets based on the obtained scan results;
[0080] The format unification unit is configured to receive the second asset information manually input through a preset interface, and perform unified processing on the formats of the first asset information and the second asset information to obtain processed asset information;
[0081] An information storage unit for structurally storing the processed asset information into the preset asset database according to the asset type of the processed asset information.
[0082] In some specific embodiments, the data recognition module 12 may specifically include:
[0083] A data screening unit for screening the log data based on a custom filtering rule to obtain target log data related to the sensitive file;
[0084] A data classification unit for parsing the target log data into a unified data format and classifying the target log data to obtain classified data;
[0085] An information completion unit for completing relevant information of the classified data based on the preset asset database and the preset personnel database.
[0086] In some specific embodiments, the data recognition module 12 may specifically include:
[0087] A data input unit for inputting the obtained processed data into a queue so that the processed data is used as the real-time data of the queue;
[0088] A data recognition unit for reading corresponding data to be recognized from the real-time data of the queue and recognizing the read data to be recognized to obtain corresponding recognition results;
[0089] An ambiguity elimination unit for normalizing the recognition results according to the log type and eliminating the ambiguity of the same-name attributes in the recognition results to obtain second entity attribute information related to the sensitive file.
[0090] In some specific embodiments, the change tracking module 14 may specifically include:
[0091] An identification assignment unit for writing the mapped data into the graph database in a multi-threaded concurrent manner according to a preset knowledge graph model and assigning a unique file identifier to the sensitive file in the mapped data;
[0092] A database update unit for continuously tracking the change of the full life cycle state of the file in the graph database based on the file identifier, and automatically recording the corresponding state change information and updating it to the graph database when the full life cycle state of the file in the graph database changes.
[0093] In some specific embodiments, the change tracking module 14 may specifically include:
[0094] A condition construction unit, configured to determine a corresponding file to be traced based on the received query instruction and construct a query condition;
[0095] An association analysis unit, configured to perform an association analysis on the file to be traced starting from the target node of the current latest graph database according to the query condition, obtain a query result related to the file to be traced, and provide corresponding feedback on the query result.
[0096] In some specific embodiments, the file transfer and traceability device may further include:
[0097] A status monitoring unit, configured to synchronously monitor the event change status of the preset asset database and the preset personnel database based on a preset interface;
[0098] An information update unit, configured to update the first entity attribute information based on the event change status when the event change status is monitored, so as to update the graph database by using the updated first entity attribute information.
[0099] Furthermore, an embodiment of the present application also discloses an electronic device, Figure 6 It is a structural diagram of an electronic device 20 shown according to an exemplary embodiment, and the content in the figure should not be considered as any limitation on the scope of use of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the file transfer and traceability method disclosed in any of the foregoing embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.
[0100] In this embodiment, the power supply 23 is used to provide working voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows is any communication protocol applicable to the technical solution of the present application, and specific limitations are not imposed here; the input / output interface 25 is used to obtain external input data or output data to the outside, and the specific interface type can be selected according to specific application needs, and specific limitations are not imposed here.
[0101] In addition, as a carrier for resource storage, the memory 22 may be a read-only memory, a random access memory, a magnetic disk, or an optical disc, etc., and the resources stored thereon may include an operating system 221, a computer program 222, etc., and the storage method may be short-term storage or permanent storage.
[0102] Among them, the operating system 221 is used to manage and control each hardware device and computer program 222 on the electronic device 20, and it can be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program capable of implementing the file transfer traceability method executed by the electronic device 20 disclosed in any of the foregoing embodiments, the computer program 222 may further include computer programs capable of performing other specific tasks.
[0103] Furthermore, the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the file transfer traceability method disclosed above is implemented. For the specific steps of this method, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details will not be repeated here.
[0104] In this specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and reference can be made to the method part for the relevant parts.
[0105] Those skilled in the art can further realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed in this document can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0106] The steps of the method or algorithm described in combination with the embodiments disclosed in this document can be directly implemented by hardware, software modules executed by a processor, or a combination of the two. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.
[0107] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, article or device comprising said element.
[0108] The technical solutions provided in this application have been introduced in detail above. Specific examples are used in this text to elaborate on the principles and implementation manners of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to this application.
Claims
1. A file circulation tracing method, characterized in that: Applicable to servers, including: Acquire first entity attribute information from a preset asset database and a preset personnel database, and receive log data sent by a target client; the log data is data collected by the target client from a plurality of target security devices; Preprocessing the log data, and identifying the processed data to obtain second entity attribute information related to the sensitive file; Performing entity mapping using the first entity attribute information and the second entity attribute information to obtain mapped data; The mapped data is written into the graph database according to the preset knowledge graph model, and the status changes of the files in the graph database throughout their life cycle are continuously tracked, and the graph database is updated using the tracked status change information, so as to obtain query results related to the files to be traced from the current latest graph database based on the received query instructions, and provide corresponding feedback on the query results.
2. The file circulation tracing method according to claim 1, characterized in that: Before acquiring the first entity attribute information from the preset asset database and the preset personnel database, the method further includes: Scanning a network environment of a plurality of target assets using an automatic asset detection tool, and determining first asset information of the target assets based on the obtained scanning results; receiving manually input second asset information through a preset interface, and performing unified processing on the formats of the first asset information and the second asset information to obtain processed asset information; The processed asset information is structured and stored in the preset asset database according to the asset type of the processed asset information.
3. The file circulation tracing method according to claim 1 is characterized in that: The preprocessing of the log data includes: Filtering the log data based on a custom filtering rule to obtain target log data related to the sensitive file; Parsing the target log data into a unified data format, and classifying the target log data to obtain classified data; The relevant information of the classified data is supplemented based on the preset asset database and the preset personnel database.
4. The file circulation tracing method according to claim 1, characterized in that: The step of identifying the processed data to obtain second entity attribute information related to the sensitive file includes: Inputting the obtained processed data into a queue so as to use the processed data as real-time data of the queue; Reading corresponding data to be identified from the real-time data in the queue, and identifying the read data to be identified to obtain corresponding identification results; The identification result is normalized according to the log type, and ambiguity of attributes with the same name in the identification result is eliminated to obtain second entity attribute information related to the sensitive file.
5. The file circulation tracing method according to claim 1, characterized in that: The method of writing the mapped data into a graph database according to a preset knowledge graph model, continuously tracking the full life cycle status changes of files in the graph database, and updating the graph database using the tracked status change information includes: According to a preset knowledge graph model, the mapped data is written into a graph database in a multi-threaded concurrent form, and a unique file identifier is assigned to the sensitive file in the mapped data; Based on the file identifier, the full life cycle status changes of the file in the graph database are continuously tracked, and when the full life cycle status of the file in the graph database changes, the corresponding status change information is automatically recorded and updated in the graph database.
6. The file circulation tracing method according to claim 1, characterized in that: The method of obtaining query results related to the file to be traced from the latest graph database based on the received query instruction and providing corresponding feedback of the query results includes: Determine the corresponding file to be traced based on the received query instruction and construct the query condition; According to the query conditions, the association analysis of the file to be traced is performed with the current latest target node of the graph database as the traversal starting point to obtain query results related to the file to be traced, and the query results are fed back accordingly.
7. The file circulation tracing method according to any one of claims 1 to 6, characterized in that: Also includes: Synchronously monitoring the event change status of the preset asset database and the preset personnel database based on a preset interface; When the event change state is monitored, the first entity attribute information is updated based on the event change state, so as to update the graph database with the updated first entity attribute information.
8. A file circulation tracing device, characterized in that: Applicable to servers, including: A data receiving module, used to obtain the first entity attribute information from a preset asset database and a preset personnel database, and receive log data sent by a target client; the log data is data collected by the client from a number of target security devices; A data identification module, used for preprocessing the log data and identifying the processed data to obtain second entity attribute information related to the sensitive file; An entity mapping module, used to perform entity mapping using the first entity attribute information and the second entity attribute information to obtain mapped data; The change tracking module is used to write the mapped data into the graph database according to the preset knowledge graph model, and continuously track the status changes of the files in the graph database throughout their life cycle and update the graph database using the tracked status change information, so as to obtain the query results related to the files to be traced from the current latest graph database based on the received query instructions, and provide corresponding feedback on the query results.
9. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor is used to execute the computer program to implement the file flow tracing method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: Used to store computer programs; wherein, when the computer program is executed by a processor, the file flow tracing method as described in any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Data synchronization method, device and system of time sequence database and storage medium
CN121979952A
Data synchronization method, device and system of time series database and storage medium
CN121979952B