A method and system for sensitive data operation trajectory monitoring and compliance review
By extracting and embedding operation time, source identifier, and business association marker in a one-way network isolation environment, the problem of sensitive data operation trajectory losing context during cross-network transmission is solved. This enables integrity and reliability review of operation trajectory, ensures traceability of legal sequence and responsibility attribution, and improves the efficiency and accuracy of compliance review.
Patent Information
- Application Number
- CN202512036302.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-31
- Publication Date
- 2026-03-20
- Estimated Expiration
- 2045-12-31
AI Technical Summary
In a one-way network gateway isolation environment, sensitive data operation trajectories lose contextual information when transmitted across networks, making it difficult for reviewers to determine the legal order of operations and the attribution of responsibility. Existing technologies cannot ensure the integrity and reliability of trajectory records while maintaining one-way transmission isolation.
The operation time, source identifier, and business association marker are extracted from the sensitive data operation trajectory through a preset module in the online environment. These markers are then embedded into the core content of the trajectory record to form enhanced trajectory data. This data is then transmitted to offline storage through a one-way gateway. The business association marker is matched using an association algorithm, and combined with source identifier comparison and time series sorting, clustering algorithm grouping is performed. Finally, the data is cross-validated with historical trajectories in offline storage to reconstruct complete context information.
Under strict one-way transmission isolation conditions, the complete preservation and reliable review of sensitive data operation traces are achieved, enabling rapid confirmation of the legal sequence of operations and attribution of responsibility, significantly improving the efficiency and accuracy of compliance reviews.
Smart Images

Figure CN121458235B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of information technology, in particular to a sensitive data operation track monitoring and compliance review method and system. BACKGROUND
[0002] In the field of information system security and compliance review, complete recording and reliable review of sensitive data operation track are directly related to enterprise risk prevention and control and regulatory requirements fulfillment, and its importance is reflected in that once the track information is missing or incomplete, it will lead to difficulty in responsibility tracing and potential violation hidden trouble difficult to find.
[0003] Although some current methods have realized the collection and storage of operation track across systems, they often rely on simple physical isolation transmission method when data is transferred from online environment to offline environment, resulting in loss of necessary context information of track record after transmission, making it difficult for review personnel to independently judge the complete background of each record on the offline side.
[0004] The core difficulty of this transmission method is that one-way gateways can only support one-way ferry of data, but cannot add auxiliary information required for review, such as operation time, source identification and business association mark, in the transmission process.
[0005] These auxiliary information is originally scattered in the business system on the online side, once the data is separated from the original environment and enters the one-way transmission channel, it faces the risk of context stripping, and this stripping further amplifies the isolation of track records in offline storage, so that the originally interrelated operation steps become scattered fragments in review, and the review personnel need to manually supplement information to restore the complete process.
[0006] For example, in actual business, a user completes query and modification operations on sensitive data in an online system, and the multiple track records generated should be associated with each other through business serial numbers to form a clear operation chain, but after simple ferry through one-way gateways, the records received on the offline side only remain the core operation content, lacking timestamp, system source and association identification, resulting in that it is impossible to quickly confirm whether these records belong to the same business event, and it is also impossible to directly judge the legal order of operation and responsibility attribution.
[0007] Therefore, how to ensure that the sensitive operation track still retains complete review context when it is transmitted to offline storage across network boundaries, while strictly maintaining one-way transmission isolation, has become a key problem for efficient compliance review. SUMMARY
[0008] The present application provides a sensitive data operation track monitoring and compliance review method, mainly comprising:
[0009] The operation time, the source identifier and the business association mark are extracted from the sensitive data operation track through a preset module in an online environment, the mark is attached to the core content of the track record in an embedding mode, and enhanced track data containing complete context is obtained;
[0010] If the enhanced track data meets the transmission format requirement of the one-way gate, one-way ferry is performed through the gate, and it is judged that the data after ferrying retains the additional mark in offline storage.
[0011] According to the enhanced track data after ferrying, the business association mark is matched by using an association algorithm, and the matched record forms a continuous operation chain.
[0012] The source identifier in the operation chain is obtained, and the identifier is compared with a preset system source library to obtain a comparison result for subsequent review and judgment.
[0013] If the comparison result shows that the source identifiers are consistent, the chain is sorted by the operation time sequence, and it is judged that the sorted chain reflects the legal operation sequence.
[0014] According to the sorted operation chain, the related records are grouped by using a clustering algorithm, and the grouped chain provides a basis for responsibility attribution.
[0015] The grouped chain is cross-verified with historical tracks in offline storage to obtain complete context information after verification to support compliance review.
[0016] The application provides a sensitive data operation track monitoring and compliance review system, mainly comprising:
[0017] An enhanced track data generation module is used to extract operation time, source identifier and business association mark from sensitive data operation track through a preset module in an online environment, and the mark is attached to the core content of the track record in an embedding mode to obtain enhanced track data containing complete context.
[0018] A one-way ferry and mark retention judgment module is used to perform one-way ferry through the gate if the enhanced track data meets the transmission format requirement of the one-way gate, and it is judged that the data after ferrying retains the additional mark in offline storage.
[0019] A business association mark matching module is used to match the business association mark according to the enhanced track data after ferrying by using an association algorithm, and the matched record forms a continuous operation chain.
[0020] A source identifier comparison module is used to obtain the source identifier in the operation chain, compare the identifier with a preset system source library, and obtain a comparison result for subsequent review and judgment.
[0021] The operation chain sorting module is configured to sort the chains by operation time sequence if the comparison result shows that the source identifiers are consistent, and to determine whether the sorted chains reflect a legal operation sequence.
[0022] The clustering grouping module is configured to group the related records by using a clustering algorithm according to the sorted operation chains, and to determine that the grouped chains provide a basis for responsibility attribution.
[0023] The cross verification module is configured to cross verify the grouped chains with historical tracks stored offline to obtain complete context information after verification to support compliance review.
[0024] The technical scheme provided by the embodiment of the present application can include the following beneficial effects:
[0025] The present application discloses a method for realizing cross-network security compliance review of sensitive data operation track in a one-way gate isolation environment, which extracts operation time, source identifier and business association marker from the operation track in an online environment and forms enhanced track data by embedding and adding, and still retains the added markers after being transferred to an offline environment through a one-way gate, then matches and forms continuous operation chains based on the business association marker by using an association algorithm, verifies the legality by comparing the source identifier with a system source library, combines time sequence sorting to reconstruct the real operation sequence, and uses a clustering algorithm to group the chains to determine responsibility attribution, and finally cross verifies with offline historical tracks to form a complete context evidence chain, so as to realize traceability, verifiability and responsibility identification of operation behavior under strict physical isolation conditions, completely solve the integrity and credibility problem of cross-network sensitive data operation audit, and significantly improve the compliance review efficiency and accuracy. BRIEF DESCRIPTION OF DRAWINGS
[0026] Figure 1 The flowchart of the method for monitoring and compliance review of sensitive data operation track of the present application.
[0027] Figure 2 The schematic diagram of the enhanced track data generation process in the method for monitoring and compliance review of sensitive data operation track of the present application.
[0028] Figure 3 Another schematic diagram of the method for monitoring and compliance review of sensitive data operation track of the present application.
[0029] Figure 4 The structural schematic diagram of the system for monitoring and compliance review of sensitive data operation track of the present application.
[0030] Figure 5 The enhanced track data vector embedding processing schematic diagram of the present application.
[0031] Figure 6 The operation chain matching schematic diagram of the correlation algorithm of the application.
[0032] Figure 7 The hierarchical clustering algorithm grouping tree structure schematic diagram in S106 step of the application.
[0033] Figure 8 The source identification matching confidence comparison diagram of the application.
[0034] Figure 9 The support vector machine classification compliance review schematic diagram of the application.
[0035] Figure 10 The one-way gate data ferry physical connection schematic diagram of the application.
[0036] Figure 11 The enhanced trajectory data packet structure schematic diagram of the application.
[0037] Figure 12 The operation chain time sequence arrangement schematic diagram of the application. DETAILED DESCRIPTION
[0038] It should be noted that the embodiments and features in the embodiments in the present application can be combined with each other without conflict. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0039] As Figure 1As shown, the present application provides a method for sensitive data operation trajectory monitoring and compliance review, which realizes the whole-process monitoring and review of sensitive data operation through seven steps. First, in step S101, the operation time, source identifier and business association mark are extracted from the sensitive data operation trajectory by a preset module in the online environment, the mark is attached to the core content of the trajectory record in an embedded manner, and the enhanced trajectory data containing complete context is obtained. Next, in step S102, if the enhanced trajectory data meets the transmission format requirements of the one-way gate, the one-way transfer is performed through the gate, and it is judged whether the data after transfer retains the additional mark in the offline storage. Then, in step S103, according to the enhanced trajectory data after transfer, the business association mark is matched by using the correlation algorithm, and the matched record forms a continuous operation chain. In step S104, the source identifier in the operation chain is obtained, and the identifier is compared with the preset system source library to obtain the comparison result for subsequent review and judgment. In step S105, if the comparison result shows that the source identifiers are consistent, the chain is sorted by the operation time sequence, and it is judged whether the sorted chain reflects the legal operation sequence. In step S106, according to the sorted operation chain, the related records are grouped by using the clustering algorithm, and the grouped chain provides the basis for responsibility attribution. Finally, in step S107, the grouped chain is cross-verified with the historical trajectory in the offline storage, and the complete context information after verification is obtained to support the compliance review. The flowchart clearly shows the complete technical path of the method from data extraction, transmission, association, verification to final review.
[0040] As Figure 4As shown, the sensitive data operation track monitoring and compliance review system provided by the present application includes four functional layers and seven core modules. In the data collection layer, the enhanced track data generation module 401 is responsible for extracting operation time, source identification and business association markers from the online environment, and attaching the markers to the core content of the track record by embedding to form enhanced track data containing complete context. In the transmission and verification layer, the one-way transfer and marker retention judgment module 402 is represented by a hexagonal icon, embodying its gateway characteristics. This module is responsible for judging whether the enhanced track data meets the transmission format requirements of the one-way gateway, and verifying that the additional markers are retained in the offline storage after one-way transfer through the gateway. In the analysis and processing layer, the business association marker matching module 403 uses association algorithms to match business association markers, ensuring that the matched records form a continuous operation chain; the source identification comparison module 404 compares the source identification in the operation chain with the preset system source library to obtain the comparison result; the operation chain sorting module 405 sorts the chain by operation time sequence on the premise that the comparison result shows that the source identification is consistent, ensuring that the sorted chain reflects the legal operation sequence. In the verification and review layer, the clustering grouping module 406 uses clustering algorithms to group related records, so that the grouped chain can provide a basis for responsibility attribution; the cross-validation module 407 cross-validates the grouped chain with the historical track in the offline storage to obtain complete context information after verification to support compliance review. The offline storage is represented by a cylindrical icon, which provides historical track data to the cross-validation module 407 through a dashed arrow and receives storage data from the one-way transfer and marker retention judgment module 402 through a dotted arrow. The modules are connected by solid arrows, and the types of data flow are marked on the arrows, embodying the complete data processing flow from enhanced track data generation, one-way transfer, business association analysis, source comparison verification, chain sorting and arrangement, clustering grouping and attribution, to final cross-validation and review.
[0041] As Figure 11As shown, the enhanced trajectory data packet adopts a hierarchical structure design, containing, from top to bottom, a packet header (4 bytes), a version number (2 bytes), a core content area (256 bytes), an operation time (8 bytes), a source identifier (64 bytes), a service association marker area (128 bytes), and a checksum (4 bytes). Among them, the packet header records the length information of the entire data packet, and the version number identifies the data packet format version; the core content area, as the largest field area, stores the original sensitive data operation trajectory record, including operation type, data table name and involved field information; the operation time field records the specific timestamp in UTC standard format, such as "2023-10-15 14:30:00 UTC"; the source identifier field records the IP address and user identifier of the operation source, such as "IP: 192.168.1.1; User ID: UserA"; the service association marker area splits the multi-value type service marker into multiple independent sub-markers, each sub-marker is attached with a unique sequence number, such as "transaction type: query [sequence number: 001]", "association ID: TX12345 [sequence number: 002]", "department: finance department [sequence number: 003]", forming a standardized service marker set; the checksum field ensures the integrity of the data packet through CRC32 algorithm. This embedded structure design can organically integrate multi-dimensional meta-information into the trajectory record, maintaining the independence of the core content, and realizing the unified packaging of operation time, source identifier and service association marker, providing a standardized data basis for subsequent data retrieval, correlation analysis and traceability tracking.
[0042] The technical solutions in the embodiments of the present application will be described clearly and in detail below with reference to the drawings in the embodiments of the present application. The described embodiments are only some of the embodiments of the present application.
[0043] The method and system for sensitive data operation trajectory monitoring and compliance review in this embodiment can specifically include:
[0044] S101, extracting operation time, source identifier and service association marker from sensitive data operation trajectory through a preset module in an online environment, and attaching the markers to the core content of the trajectory record in an embedded manner to obtain an enhanced trajectory data containing complete context.
[0045] The original sensitive data operation trajectory record is obtained from the online environment log. The operation time field, the source identification field and the business association mark field in the original sensitive data operation trajectory record are parsed to obtain structured trajectory basic data. The business association mark in the structured trajectory basic data is standardized by a preset field mapping table. If the business association mark belongs to a multi-value type, it is split into multiple independent sub-marks and a unique serial number is attached to each of the sub-marks to obtain a standardized business mark set. The operation time, the source identification and the standardized business mark set in the structured trajectory basic data are converted into fixed-dimension numerical vectors by a vector embedding method, and are spliced to the core content field of the original sensitive data operation trajectory record to obtain enhanced trajectory vector data. Cosine similarity calculation is performed on the enhanced trajectory vector data to determine whether the similarity between adjacent records is higher than a preset threshold. If the similarity is higher than the preset threshold, the adjacent records are merged into the same business context chain. Otherwise, the adjacent records are kept as independent records to obtain enhanced trajectory data containing complete context.
[0046] As shown in Figure 2 The enhanced trajectory data generation process includes an input data area, a processing flow area and an output data area. First, the original sensitive data operation trajectory record is obtained from the online environment log as input data. In the processing flow area, the operation time field, the source identification field and the business association mark field are parsed by a field parsing module to obtain structured trajectory basic data. Then, the business association mark is standardized by a standardized processing module through a preset field mapping table. If the business association mark belongs to a multi-value type, it is split into multiple independent sub-marks and a unique serial number is attached to each of the sub-marks to obtain a standardized business mark set. Next, the operation time, the source identification and the standardized business mark set are converted into fixed-dimension numerical vectors by a vector embedding module, and are spliced to the core content field of the original sensitive data operation trajectory record to obtain enhanced trajectory vector data. Finally, cosine similarity calculation is performed on the enhanced trajectory vector data by a similarity calculation module to determine whether the similarity between adjacent records is higher than a preset threshold. If the similarity is higher than the preset threshold, the adjacent records are merged into the same business context chain. Otherwise, the adjacent records are kept as independent records. Finally, enhanced trajectory data containing complete context information is obtained in the output data area. The entire process clearly shows the conversion process from original data to enhanced trajectory data through data flow arrows.
[0047] As shown in Figure 5As shown, the vector embedding process includes the following steps: first, convert the operation time (e.g., 14:30:25) into a timestamp vector with 8 dimensions; second, convert the source identifier (e.g., IP: 192.168.1.100) into an 8-dimensional source vector through hash embedding; third, embed the service tag set (e.g., {login, query}) into an 8-dimensional tag vector; and finally, convert the core content (e.g., user query order) into a 128-dimensional content vector. These four vectors are sequentially concatenated in the order of time, source, tag, and content to form a 512-dimensional enhanced trajectory vector data, where the core content vector occupies the main dimension (128 dimensions), and the other metadata vectors each occupy 8 dimensions, collectively forming a complete vector representation. This fixed-dimension vectorization representation not only preserves the semantic information of the original record, but also facilitates subsequent vector similarity calculation and business context chain construction.
[0048] For example, when processing online environment logs, it is first necessary to extract those records that record the user's operation trajectory on sensitive data such as personal information or financial records from the log system. These trajectories usually include user login, data query and modification behaviors, etc., for subsequent audit and security analysis. Through this acquisition method, it can be ensured that the data comes from a real online interaction environment, thereby providing a reliable basis for the entire processing flow.
[0049] For example, suppose that in the logs of a bank online system, it is recorded that user A queried the account balance from IP address 192.168.1.1 at 2023-10-15 14:30:00, this record constitutes part of the original sensitive data operation trajectory, and its core content field may contain the specific query command and result summary.
[0050] Specifically, when parsing these original records, the operation time field such as "2023-10-15 14:30:00", the source identifier field such as "IP: 192.168.1.1; User ID: UserA", and the business association tag field such as "transaction type: query; association ID: TX12345" are extracted and structured. The operation time field indicates the exact time of the event, the source identifier field is used to track the initiator of the operation, and the business association tag field is linked to the specific business scenario, such as account query or transfer. Through parsing tools or scripts, these fields are separated from the original string of logs to form a structured data object, such as a JSON format dictionary, where key-value pairs clearly correspond to each field, thereby facilitating subsequent standardization and analysis. This structured trajectory base data helps to transform chaotic logs into an operable format, which can improve the efficiency and accuracy of data processing in business, and avoid errors caused by manual interpretation.
[0051] For example, for the standardization of business association tags, a preset field mapping table is used to unify the tag expressions in different systems. The mapping table is a predefined dictionary that maps, for example, "transaction type: query" to the standard code "OP_QUERY" to ensure consistency. If the tag belongs to a multi-value type, such as "association ID: TX12345, TX67890", it is split into "TX12345-001" and "TX67890-002" and a unique serial number is appended to form a standardized business tag set. This set can be a list, such as ["OP_QUERY-TX12345-001", "OP_QUERY-TX67890-002"], so that after processing, not only are the tags standardized, but the integrity of the multi-value relationship is also preserved, providing more detailed input for vector conversion.
[0052] In an embodiment, this standardization effectively solves the problem of inconsistent data across systems, such as in a multi-platform banking system, where the different branch tag expressions can be unified through a mapping table, thereby quickly identifying relevant operation chains during auditing.
[0053] Specifically, the vector embedding method converts the operation time, source identifier, and standardized business tag set into a fixed-dimensional numerical vector.
[0054] For example, using Word2Vec or BERT models to embed text fields, the operation time can be converted into a timestamp vector such as [0.5, 0.3,...] to represent the normalized hours and minutes, the source identifier is converted into [0.1, 0.7,...] based on IP and ID hash embedding, and the tag set is embedded into [0.2, 0.4,...]. These vectors are concatenated after the embedding vector of the core content field of the original record to form enhanced trajectory vector data, such as a 512-dimensional overall vector. This conversion principle is based on mapping discrete data to continuous vector space, which facilitates similarity calculation and can capture the semantic similarity of operations in business, such as making it easier to cluster analysis after vectorizing similar query operations.
[0055] For example, when performing cosine similarity calculation, the formula is applied to adjacent enhanced trajectory vector data to calculate the similarity score, and a preset threshold such as 0.8 is used. If it is higher, it is merged into the same business context chain.
[0056] For example, the vector similarity of two records is 0.85, indicating that they belong to consecutive steps of the same transfer business, such as querying immediately followed by transferring, then merging into the chain "query-transfer-TX12345", otherwise keeping them as independent records. This merging results in enhanced trajectory data that contains complete context, which can reveal potential abnormal patterns in security monitoring, such as abnormally frequent query chains, thereby improving the risk detection capability of the system.
[0057] Specifically, this complete process, from acquisition to merging, forms a logically progressive trajectory analysis chain. In extended solutions, machine learning models can be further integrated to classify the chain, such as identifying high-risk context chains for real-time alerts.
[0058] S102. If the enhanced trajectory data meets the transmission format requirements of the one-way gateway, then the data is transferred in one direction through the gateway, and it is determined that the transferred data retains additional markers in offline storage.
[0059] By processing the enhanced trajectory, data content conforming to the one-way gateway transmission specification is obtained, determining the pre-transmission preparation status. According to the one-way gateway's transmission specification, the format of the additional markers in the enhanced trajectory is validated to determine if they meet the requirements of the transmission process. If the additional markers conform to the transmission specification during transmission, data transfer is performed through the one-way gateway to obtain the transferred trajectory information. For the transferred trajectory information, data integrity is checked in the offline storage environment to determine if the additional markers are completely preserved. If the additional markers in offline storage are completely preserved, the preserved content is compared using an information verification tool to obtain the verification result. Based on the verification result, the trajectory information in the storage environment is categorized and archived to determine the final storage structure. Periodic scanning of the data integrity in the storage structure is performed to obtain the long-term preservation status of the markers, determining the continuous availability of the data.
[0060] like Figure 3 As shown, this invention achieves physical isolation between online and offline environments through a one-way gateway. The enhanced trajectory data in the online environment on the left contains key information such as business association markers, source identifiers, and operation time sequences. This data is unidirectionally transferred via the intermediate one-way gateway, ensuring that data can only flow from the online environment to the offline environment, thus guaranteeing system security. After the data is transmitted to the offline environment on the right, it enters a multi-stage processing flow: First, step S102 is executed to check the integrity of the received data and confirm whether the additional markers are completely retained; then, in step S103, an association algorithm is used to match the business association markers, forming a continuous operation chain; in step S104, the source identifier in the operation chain is extracted and compared with the system source database for verification; in step S105, the chain is sorted using the operation time sequence to establish a time-seriesd operation record; in step S106, a clustering algorithm is used to group the relevant records to determine the attribution of responsibility; finally, in step S107, the current operation chain is cross-validated with historical trajectory data in offline storage, and a compliance review result is output after comprehensive analysis. The entire process embodies a complete compliance review mechanism, from secure data transmission to multi-dimensional verification and analysis.
[0061] like Figure 10As shown, the physical connection relationship and data flow direction between the online environment, one-way gate and offline storage system are displayed. The left high-security area contains online environment servers, which store enhanced trajectory data; the middle is a one-way gate device that uses fiber optic one-way transmission technology, allowing data to flow only from the high-security area to the low-security area. The internal optical one-way transmission channel ensures that data cannot be transmitted back after flowing out; the right low-security area contains an offline storage system for receiving and storing data after being transferred by the one-way gate. The data flow is as follows: enhanced trajectory data in the online environment server enters the input port of the one-way gate through the input data interface, is processed through the optical one-way transmission channel, and then flows out from the output port, finally in the form of transferred data to the offline storage system. The entire physical connection architecture realizes the secure isolation and one-way transfer of sensitive data through the hardware-level one-way transmission mechanism, effectively preventing data from being stolen and leaked in the opposite direction.
[0062] Specifically, when processing enhanced trajectory data, first, the additional markers need to be standardized and adjusted to ensure that they meet the transmission requirements of the one-way gate.
[0063] For example, in the sensitive transaction trajectory of the financial industry, additional markers may include operation time such as "2025-12-1014:30:00", source identification such as "System A Terminal 001", and business-related markers such as "money transfer business - customer ID 12345". By checking the format of these markers, ensuring that the time is in UTC standard, the source identification is a uniform coded string, and the business marker does not exceed the specified length, it is determined that the data is ready for transmission, avoiding transmission interruption.
[0064] In one embodiment, format verification is performed on the additional markers in the enhanced trajectory.
[0065] It can be understood that the one-way gate is a hardware device that only allows data to flow from the high-security area to the low-security area, and its transmission specifications usually require that the data packet size be less than 1MB and that the marker field have no special characters.
[0066] For example, for trajectory data of medical records, the verification process involves verifying whether the operation time is in ISO 8601 format, whether the source identification matches a predefined whitelist such as "Hospital Server H001", and whether the business-related marker such as "patient diagnosis - case number P789" has been encrypted. If there are multiple values in the business such as "diagnosis + treatment", each sub-item needs to be split and verified to ensure that the whole meets the specifications, thereby determining the transmission feasibility.
[0067] For example, when the additional markers pass the verification, data transfer is performed through the one-way gate, and the transferred trajectory information is obtained.
[0068] Specifically, in the enterprise internal audit scenario, the enhanced trajectory data is transferred from the online environment to the offline environment, and the transfer process uses optical fiber one-way transmission, only allowing data to flow out but not back, ensuring that sensitive information such as employee access logs cannot be stolen in reverse, and the information after the transfer retains the original structure.
[0069] In an embodiment, for the trajectory information after the transfer, data integrity detection is performed in the offline storage environment.
[0070] It should be noted that the integrity detection compares the fingerprint values of the original and received data through a hash algorithm.
[0071] For example, in the order trajectory of an e-commerce platform, it is detected whether the additional markers such as “order time 2025-12-10 15:00:00”, source “APP user U456” and business “payment-order O987” are completely retained, and if the hash values match, it is confirmed that there is no loss.
[0072] Specifically, if the additional markers are completely retained, the content is compared using an information verification tool, and the result is obtained.
[0073] For example, the tool such as a custom script compares the marker fields, and in the bank transfer trajectory, the string consistency of the operation time, source and business markers is compared, and the result shows “complete” or “difference”, which is helpful for subsequent processing.
[0074] For example, according to the verification result, the trajectory information in the storage environment is classified and archived, and the storage structure is determined.
[0075] Specifically, in the insurance claim business, if the result is complete, the trajectory is archived to the “verified” folder, classified according to the business markers such as “claim-claim S654”, forming a tree structure, and ensuring fast retrieval.
[0076] In an embodiment, the retention status of the markers is obtained by periodically scanning the storage structure, and the continuous availability is determined.
[0077] For example, the storage files of e-commerce trajectories are scanned every week, and the integrity of the additional markers is checked using the verification and tool, and if the status is stable, the data is considered available, supporting long-term audit needs.
[0078] S103, according to the enhanced trajectory data after the transfer, the business association markers are matched using an association algorithm, and the matched records form a continuous operation chain.
[0079] According to the enhanced trajectory data, all trajectory data sequence and corresponding business association mark are extracted. The support and confidence of adjacent operation record identifiers in the trajectory data sequence are calculated by using the association rule algorithm, and the matching result of the mark is obtained. All record pairs with confidence higher than the pre-established association strength threshold are screened through the matching result of the mark, and a candidate matching record set is determined. According to the time sequence relationship of the operation record identifiers in the candidate matching record set, adjacent record pairs are connected, and an initial operation chain is generated. For each chain end record in the initial operation chain, other candidate records sharing the same business association mark are obtained, and if there are time-continuous and mark-consistent records, the chain end point is extended. The chain records after extension are used for chain continuity judgment, and if the time difference between adjacent records is consistent with the mark matching result, the chain forming record is determined. According to all chain forming records, the chains with overlapping operation record identifiers are merged, and the final continuous operation chain record is obtained.
[0080] As shown in Figure 6 , the complete process of forming the operation chain by the association rule algorithm to calculate the support and confidence in the S103 step is shown. The figure is divided into three levels: the top is the trajectory data sequence layer, which shows three trajectory data sequences extracted from the enhanced trajectory data (trajectory T1 contains operation records R1, R2, R3, R4, trajectory T2 contains R2, R3, R5, and trajectory T3 contains R1, R3, R4, R6), each operation record is represented by a circular node, and the time sequence relationship is connected by an arrow; the middle is the association rule algorithm calculation layer, which shows the support and confidence calculation process of adjacent operation record identifiers in the trajectory data sequence by the Apriori or FP-growth algorithm, wherein the support is the frequency of the co-occurrence of the two divided by the total number of records, and the confidence is the conditional probability of the occurrence of R2 when R1 occurs (for example, the support of R1->R2 is 0.55, and the confidence is 0.91, the support of R2->R3 is 0.67, and the confidence is 0.92, etc.), and the candidate matching record set {R1->R2, R2->R3, R3->R4} is obtained by threshold screening (confidence >= 0.85); the bottom is the formed operation chain layer, according to the time sequence relationship of the operation record identifiers, adjacent record pairs are connected to generate an initial operation chain, chain end points are extended, and chains with overlapping operation record identifiers are merged to obtain the final continuous operation chain (chain 1: R1->R2->R3->R4, chain 2: R2->R3->R5). The whole process is represented by a vertical dashed arrow indicating the data flow conversion relationship, forming a complete closed loop from trajectory extraction, algorithm calculation to chain generation, which clearly shows the application mechanism of the association rule algorithm in sensitive data operation trajectory monitoring.
[0081] For example, in one embodiment, when extracting all sequence of trajectory data from enhanced trajectory data, a specific business scenario can be considered, such as vehicle trajectory records in a logistics transportation management system. Assume that the enhanced trajectory data contains path point information of the vehicle from the starting point to the ending point, as well as corresponding business association labels, such as "cargo loading label" or "route optimization label". Through the data analysis tool, the sequence of trajectory points such as vehicle ID V001 is first extracted: timestamp T1 position A, T2 position B, T3 position C, and corresponding labels such as "loading complete". This process involves scanning the entire data set, filtering out the serialized operation records, and ensuring that each sequence corresponds to a business label one by one, thereby providing basic data for subsequent analysis.
[0082] For example, when calculating support and confidence using the association rule algorithm, it is necessary to first understand the basic principles of the association rule algorithm, which is a data mining method used to discover frequent patterns between item sets.
[0083] Specifically, in the sequence of trajectory data, for adjacent operation record identifiers such as record R1 (position A, label "loading") and R2 (position B, label "transportation"), the support is the frequency of their co-occurrence divided by the total number of records, and the confidence is the conditional probability of the occurrence of R2 when R1 occurs. Through the Apriori algorithm or FP-growth algorithm, for example, in a logistics system, all sequences are scanned to find patterns such as "loading" followed by "transportation", and if the support is 0.8 and the confidence is 0.9, the matching result is obtained, which helps to identify common business operation associations.
[0084] For example, when filtering all record pairs with a confidence higher than a threshold, the pre-established association strength threshold can be set to 0.85.
[0085] In one embodiment, for the matching results of the previous step, all record pairs are traversed, such as filtering out the "loading-transportation" pair with a confidence of 0.9 and the "transportation-unloading" pair with a confidence of 0.87, to form a candidate matching record set. This set is essentially a list of high-confidence association pairs, which can bring more accurate chain construction results in business, such as avoiding low-confidence noise data interference in transportation management.
[0086] For example, the process of connecting adjacent record pairs to generate an initial operation chain according to the time sequence relationship can be implemented through a sorting algorithm.
[0087] Specifically, the record pairs in the candidate set are sorted by timestamp, such as connecting R1 to R2 and then to R3, forming the chain "loading-transportation-unloading". In the logistics example, this corresponds to the continuous path of the vehicle from the warehouse to the destination, ensuring that the chain reflects the true business process.
[0088] For example, when extending a chain endpoint, other records sharing the same marker are retrieved. For instance, endpoint R3 (unloading marker) is searched for R4 (with the same marker "unloading confirmation"), which is consecutive in time. If the time difference is less than a threshold, such as 5 minutes, and the markers are consistent, the chain is extended. This extension mechanism enriches the chain's integrity, improving tracking accuracy in business applications, such as extending the transportation chain to include a confirmation step.
[0089] For example, when determining chain continuity, check whether the time difference between adjacent records is consistent with the matching result. If the time difference is within the expected 10 minutes, then the chain is confirmed to be formed.
[0090] In one embodiment, this is similar to verifying the continuity of a chain. For example, if all adjacent points in a logistics chain are time-series consecutive and their markers match, they are determined to be valid records, facilitating subsequent merging.
[0091] For example, when merging chains with overlapping identifiers to obtain the final record, all chains are scanned to form records. If chains 1 and 2 share the R2 identifier, they are merged into a single continuous chain. In logistics operations, this integrates fragmented trajectories into complete paths, enabling efficient route analysis and optimization.
[0092] S104. Obtain the source identifier in the operation chain, compare the identifier with the preset system source library, and obtain the comparison result for subsequent review and judgment.
[0093] Obtain the complete data packet of the operation chain and extract the source identifier from it. Perform standardization processing on the source identifier to generate a standardized source identifier. Load the set of all legal source records from the preset system source library. Perform precise string matching comparison on each item of the standardized source identifier and the set of legal source records to obtain a preliminary matching status. If the preliminary matching status is a miss, use the cosine similarity algorithm to calculate the similarity score between the standardized source identifier and each record in the set of legal source records, and obtain the highest similarity score and its corresponding record. If the highest similarity score is greater than a preset similarity threshold, the matching status is determined to be a match and the corresponding legal source record is bound; otherwise, the matching status is determined to be a miss. Generate the final comparison matching result based on the matching status and the bound legal source record, including the matching status field and the associated record identifier field, and output the final comparison matching result to the review process processing interface.
[0094] like Figure 8As shown, the present application uses the cosine similarity algorithm to match and compare the standardized source identifier with the legal source record set, and shows the matching similarity performance of different source identifier samples. The horizontal coordinate in the figure represents ten different source identifier samples (source 1 to source 10), and the vertical coordinate represents the cosine similarity score, which ranges from 0 to 1.0. Each column chart represents the similarity score of the source identifier and the best matching record in the legal source record set. The horizontal dashed line in the figure marks the preset similarity threshold T = 0.8. When the similarity score is greater than or equal to 0.8, the matching state is determined to be matched, and the corresponding column is filled with dark color; when the similarity score is less than 0.8, the matching state is determined to be not matched, and the corresponding column is filled with light color. As can be seen from the figure, the similarity scores of source 1, source 3, source 5, source 7, source 8 and source 10 are 0.95, 0.88, 0.91, 0.82, 0.97 and 0.85 respectively, all of which are higher than the threshold, and the matching is successful; while the similarity scores of source 2, source 4, source 6 and source 9 are 0.72, 0.63, 0.45 and 0.68 respectively, all of which are lower than the threshold, and the matching fails. This embodiment fully verifies the effectiveness of the cosine similarity algorithm in source identifier matching, which can accurately distinguish between legal sources and illegal sources, thereby effectively preventing content security risks.
[0095] In one possible implementation, when the system needs to verify the source reliability of the operation chain, it first obtains the operation chain complete data packet from the storage system. This data packet usually contains structured data such as user behavior sequence, timestamp and source information.
[0096] For example, in the user behavior analysis scenario of an e-commerce platform, the operation chain complete data packet may record the entire process from browsing goods to ordering and paying, where the source identifier is, for example, "APP" or "Web browser". These identifiers are extracted by parsing tools to ensure the accuracy of subsequent processing. This acquisition and extraction process helps to build a reliable data foundation and avoid verification failures due to data loss.
[0097] Specifically, the extracted source identifier is subjected to standardization processing, which can unify various formats of identifiers into a standard form, such as converting "app-end" to "APP_END", which involves removing extra spaces, converting case and replacing special characters.
[0098] In one possible implementation, for an operation chain of a bank transaction system, if the source identifier is "mobile app v1.2", it is standardized to "MOBILE_APP", which facilitates subsequent matching and improves processing efficiency and consistency. Then, a complete set of legal source records is loaded from a preset system source library, which is a pre-constructed database storing all authorized sources such as "OFFICIAL_APP" and "PARTNER_WEBSITE".
[0099] In one possible implementation, the system loads these records from the cloud database through SQL queries to ensure real-time and completeness, for example, in a supply chain management system, the set of legal source records may include "internal ERP system" and "external supplier API", and after loading, a list is formed for comparison.
[0100] In one possible implementation, an exact string matching comparison is performed for the standardized source identifier and the set of legal source records, and if the standardized identifier completely matches a record in the library, the preliminary matching state is "hit", otherwise it is "miss".
[0101] For example, in a medical record tracking system, if the standardized source is "HOSPITAL_PORTAL" and there is a same record in the library, the state is hit, and this step emphasizes accuracy to filter obvious matching items.
[0102] Specifically, if the preliminary matching state is miss, a cosine similarity algorithm is used to calculate the similarity score of the standardized source identifier and each legal record, which is based on the vector space model to convert the string into a vector and obtain the similarity by calculating the cosine value.
[0103] For example, the string is decomposed into a character vector, and the cosine similarity of "APP_END" and "APP_END_V2" may be 0.85, which helps to identify approximate matches and avoid the limitations of strict matching.
[0104] In one possible implementation, for an operation chain of a content distribution platform, if the highest similarity score is 0.92, corresponding to the record "CONTENT_APP", and greater than the threshold 0.8, it is determined to be matched and bound to the record, otherwise it is marked as not matched.
[0105] In one possible implementation, the final comparison matching result is generated according to the matching state and the bound legal source record, including a matching state field such as "matched" and an associated record identifier field such as "CONTENT_APP_ID", and then output to the review process processing interface. This output can be a JSON format data packet sent to the downstream audit system.
[0106] For example, in the insurance claim process, the final result helps to identify the operation chain of the legal source, thereby accelerating the review and reducing the risk of fraud, which improves the accuracy and automation level of source verification in business.
[0107] S105, if the comparison result shows that the source identifiers are consistent, then sort the chain by operation time sequence to determine whether the sorted chain reflects the legal operation order.
[0108] By obtaining the data of source identifiers and operation times from the data storage, an initial operation chain record is constructed. If the source identifiers in the obtained operation chain record are consistent, then a time sequence is constructed according to the operation times to generate sorted chain data. According to the sorted chain data, a preset legal order rule is used for comparison to determine whether the chain meets the requirements of legal order. If the comparison shows that the chain order does not meet the legal rule, then the time point of abnormal operation is determined through reverse verification of the time sequence. For the time point of abnormal operation, the associated operation chain segment is obtained, and the deviation of time sorting and legal order is analyzed. According to the deviation data obtained by analysis, a decision tree algorithm is used to classify the operation chain to obtain the classification label of the abnormal segment. By matching the classification label with the preset judgment basis, the overall legality of the operation chain is determined, and the final judgment conclusion is output.
[0109] As shown in FIG. 1, Figure 12 The operation chain arranged in time sequence shows the complete process from data acquisition to compliance judgment. The horizontal axis in the figure is the time axis, and five time points T1 to T5 are marked, corresponding to the time of operation (09:00:05, 09:00:12, 09:00:35, 09:00:58, 09:01:20). Above the time axis, five operation nodes are arranged in time sequence, including login 201, identity verification 202, balance inquiry 203, transfer confirmation 204, and operation completion 205, and the nodes are connected by arrows to form an operation chain. Each node is associated with the corresponding time point on the time axis by a dashed line, clearly showing the correspondence between operation and time. Two business events are identified by dashed boxes in the figure: business event 1 includes login and identity verification operations, and business event 2 includes balance inquiry, transfer confirmation, and operation completion operations, reflecting the business logic grouping of operations. The top of the figure marks the source identifier "User A-Terminal 1", indicating the source of the operation chain. The bottom of the figure shows the preset legal order rule "login→identity verification→balance inquiry→transfer confirmation", which is used to compare with the sorted chain to determine whether the operation order meets the expected legal process, achieving time sequence monitoring and compliance verification of the operation track.
[0110] For example, in a financial transaction auditing system, the operation source identification such as "bank APP" or "web page" and the corresponding operation timestamp such as "2023-10-15 14:30:00" of each transaction are extracted from the database. These data are integrated into the initial operation chain record, ensuring that each record contains source and time information. By comparing the source identification of all records, if they are all consistent, such as "bank APP", then according to the timestamp from early to late, a continuous time sequence chain data is formed. This helps subsequent sequential verification.
[0111] Specifically, the sorted chain data can be represented as a list, where each element is an operation event sorted in ascending order of time. Using a preset legal sequence rule, for example, in a transfer business, the rule requires "login" first, then "identity verification", and then "confirm transfer". The system will compare each item of chain data with the rule sequence. If "confirm transfer" appears before "login" in the chain, it is judged as not meeting the legal requirements.
[0112] In one embodiment, if the comparison result shows that the order is abnormal, the reverse verification of the time sequence is started, and each operation is checked from the end of the chain to the front, and the logical relationship between the time point of each operation and the previous operation is checked.
[0113] For example, in a supply chain management system, if the "shipment" operation is earlier than the "warehouse entry" operation, the reverse verification will locate the "shipment" time point as an anomaly.
[0114] For example, for this abnormal time point, the system obtains the operation chain segment associated before and after, such as three operation events before and after the abnormal point, forming a local sub-chain. Then analyze the deviation of time sorting and legal order, calculate the time difference between each operation, and compare the pre-set interval in the ideal order. If the deviation exceeds the threshold, such as the actual time difference is negative, it is marked as a serious deviation.
[0115] Specifically, based on these deviation data, a decision tree algorithm is used to classify the operation chain. Decision tree is a supervised learning model, starting from the root node, making branch decisions according to features such as deviation value and operation type, until the leaf node to get the classification label. First, select the deviation data as the input feature, train the decision tree model, and the node splitting criterion is based on information gain, for example, the root node checks if the deviation is greater than 0, if yes, it is divided into the "normal" branch, otherwise it further checks if the operation type matches the rule, and finally outputs the label such as "order inversion" or "time forgery". In e-commerce order processing, the decision tree may classify a chain as "abnormal segment" because "payment" is after "order" but the time interval is abnormally short.
[0116] In an embodiment, the system determines the legality of the entire operation chain by matching the classification label with a preset judgment basis, for example, the label "sequence inversion" corresponds to the basis "manual review is required". If the matching is successful, the output conclusion is "illegal, intervention is required", otherwise it is "legal". This method improves the audit efficiency and ensures the compliance of the business process.
[0117] S106, according to the sorted operation chain, the related records are grouped by using clustering algorithm, and the grouped chain provides the basis for responsibility attribution.
[0118] According to the sorted operation chain, the distance matrix between records is calculated by using hierarchical clustering algorithm to obtain initial distance data. The hierarchical clustering processing is performed on the records by using the distance data to generate a tree structure containing multiple hierarchical groupings. A cutting operation is performed on the tree structure to produce multiple mutually exclusive grouping sets to obtain grouping labels and corresponding chain segments. The operation personnel identifiers and operation time sequences in each group are obtained to determine the responsibility personnel set of each group. If different operation personnel identifiers appear in the same group, the group is marked as a cross-responsibility group to obtain a cross-responsibility list. The cross-responsibility list is compared with a preset single responsibility rule to determine whether the group meets the single responsibility requirement to obtain a legal grouping set. For the legal grouping set, the number of chain segments corresponding to each responsibility personnel is counted to determine the final responsibility attribution mapping relationship.
[0119] As shown in Figure 7 When clustering and grouping the sorted operation chain, a distance matrix is first calculated according to the similarity between each operation record, which is used as the input data of the hierarchical clustering algorithm. The hierarchical clustering algorithm starts from a single operation record, and each time the two closest clusters are merged to gradually build a tree-like hierarchical structure (dendrogram). The vertical axis of the tree structure represents the distance between clusters, reflecting the similarity between different operation records or clusters. The tree structure is cut at a preset cutting line (threshold T = 2.5) to divide it into multiple mutually exclusive grouping sets. In this embodiment, three groups are obtained by the cutting operation: group 1 contains records R1, R2, R6, corresponding to responsibility personnel U1; group 2 contains records R3, R4, R8, corresponding to responsibility personnel U2; group 3 contains records R5, R7, corresponding to responsibility personnel U3. The operation records in each group have high similarity and are executed by the same responsibility personnel, thereby realizing the identification of responsibility personnel based on operation behavior patterns.
[0120] In an embodiment, according to the sorted operation chain, the distance matrix between records is calculated by using hierarchical clustering algorithm to obtain initial distance data.
[0121] Specifically, hierarchical clustering is an unsupervised learning method that builds a clustering structure by progressively merging similar data points. First, a distance metric needs to be defined, such as using Euclidean distance to measure the similarity between operational records. In a supply chain management system, assume the operational chain includes records such as "purchase order submission," "goods warehousing," and "payment settlement." Each record contains a timestamp and an operation type feature. When calculating the distance matrix, these records are treated as multi-dimensional vectors, with the timestamp as one dimension and the operation type encoding as another dimension. Then, the distance between each pair of vectors is calculated to form a symmetric matrix, such as the distance between "purchase order submission" and "goods warehousing" may be small because they are consecutive in time and related in type. The resulting initial distance data can be represented as an n x n matrix, where n is the number of records, for subsequent clustering processing. In this way, the system can identify similar patterns in the operational chain, ensuring more accurate auditing.
[0122] For example, in a financial transaction auditing system, hierarchical clustering is performed on records using distance data to generate a tree structure containing multiple hierarchical groupings. Hierarchical clustering starts with individual records and progressively merges the closest clusters until all records form a root node, forming a tree diagram, also known as a dendrogram, where each branch represents a merging step. The specific process includes selecting a linkage method, such as single linkage, which takes the distance between the closest points in the two clusters when merging. Assume there are 10 records in the chain, and after processing, a tree structure is generated, with the bottom leaf nodes being individual records and the upper nodes being merged groupings, which helps visualize the hierarchical relationship of operations.
[0123] In one embodiment, a cutting operation is performed on the tree structure, resulting in multiple mutually exclusive grouping sets, obtaining grouping labels and corresponding chain segments. The cutting operation usually selects a height threshold on the tree diagram, such as cutting according to half of the tree height, thereby dividing the tree into several sub-trees, each corresponding to a grouping. For example, in e-commerce order processing, if the tree structure shows that "order placement," "payment," and "delivery" are closely merged in one branch, the cutting produces a grouping label such as "order execution group," corresponding to a chain segment that includes these consecutive operations, ensuring the exclusivity and integrity of the groupings.
[0124] For example, the operator identifiers and operation time sequences within each grouping are obtained to determine the set of responsible personnel for each grouping. In practical applications, personnel identifiers such as employee IDs are extracted from the database and associated with time sequences, such as a grouping with three records, identifiers "Employee A" and "Employee B," and a time sequence from "2023-10-16 09:00:00" to "2023-10-16 10:00:00," thereby determining the set as {Employee A, Employee B}. This method helps trace the chain of responsibility.
[0125] In one embodiment, if different operator identifiers appear in the same group, the group is marked as a cross-responsibility group, and a cross-responsibility list is obtained.
[0126] Specifically, the system iterates through the groups and checks the uniqueness of the identifiers. If they are not unique, such as multiple employees appearing in the example above, they are marked and added to a list. For example, the list items include the group ID and the personnel involved. This can quickly locate potential collaboration problems during auditing.
[0127] For example, by comparing a cross-responsibility list with a pre-defined single responsibility rule, it can be determined whether a group meets the single responsibility requirement, thus obtaining a set of valid groups. The single responsibility rule may stipulate that certain types of operations must be performed by a single person. During the comparison, each item is matched; if a group in the list meets the rule (e.g., all operations are the responsibility of one person), it is included in the valid set; otherwise, it is excluded. In the supply chain, this improves the clarity of responsibility and reduces disputes.
[0128] In one embodiment, for a set of legitimate groups, the number of chain segments corresponding to each responsible person is counted to determine the final responsibility attribution mapping relationship.
[0129] For example, count the number of segments each person in the set is responsible for. If employee A has 5 segments and employee B has 3 segments, forming a mapping such as {employee A: 5, employee B: 3}, this provides a quantitative basis for performance evaluation and ensures the fairness and traceability of business processes.
[0130] S107. Cross-validate the grouped chains with historical trajectories in offline storage to obtain complete contextual information after verification to support compliance review.
[0131] By associating grouped chains with historical trajectories, corresponding storage records are extracted from offline storage to obtain a preliminary trajectory data set. Based on this preliminary data set, a comparison and verification method is used to match key nodes in the grouped chains with storage records to determine consistency results. If the consistency results meet a preset threshold, the matched data is classified as valid context information; otherwise, it is marked as data to be processed, and supplementary trajectory analysis results are obtained. Using the trajectory analysis results, the data to be processed undergoes secondary verification, combining the continuity characteristics of historical trajectories to determine whether it can be included in context information. Based on the context information after secondary verification, a complete review support dataset is compiled according to compliance review requirements. After obtaining the review support dataset, a support vector machine algorithm is used to classify the dataset, distinguishing between parts that meet compliance standards and potentially abnormal parts, resulting in classified data groups. Using these classified data groups, structured compliance review evidence is generated, outputting the final review support content.
[0132] As Figure 9 shown, in the classification processing of the review support dataset, the support vector machine (SVM) algorithm is used to realize intelligent classification. The classification process is demonstrated in a two-dimensional feature space, with the horizontal axis representing feature 1 (trajectory deviation degree) and the vertical axis representing feature 2 (abnormal times). The sample points in the dataset are marked with different shapes: the circular mark represents the compliance data (normal trajectory), and the triangular mark represents the abnormal data (deviation trajectory). The support vector machine uses the radial basis function (RBF) kernel to process non-linear separable data, and effectively separates the two types of data by constructing a non-linear classification hyperplane (solid curve in the figure). The dashed line on both sides of the classification hyperplane represents the margin boundary, and the data points on the margin boundary are the support vectors, marked with double circles for highlighting. These support vectors are the key sample points that determine the position of the classification hyperplane. After SVM classification processing, the dataset is divided into two classification data groups: the compliance area and the abnormal area. The lower left gray area is the compliance area, and the upper right white area is the abnormal area, providing a clear data grouping basis for subsequent generation of structured compliance review basis. This classification result can effectively identify normal trajectory data that meets the compliance standard and abnormal trajectory data that has potential risks, realizing intelligent compliance review judgment.
[0133] For example, in the logistics transportation industry, by associating the grouping chain with the historical trajectory, the corresponding storage records can be extracted from the offline storage. This can be understood as first defining the grouping chain of the vehicle as a logical sequence formed according to the transportation route and time period, such as a continuous path chain of a group of trucks from warehouse A to destination B. The historical trajectory, such as the GPS positioning record in the past week, is associated with the historical trajectory, so that the matching storage records are pulled from the database. These records include vehicle ID, timestamp, position coordinates, and speed data, etc., and a preliminary trajectory data set is formed. This set may cover hundreds of records for subsequent verification to ensure the accuracy and integrity of data extraction.
[0134] Specifically, according to the preliminary trajectory data set, a comparison verification method is used to match the key nodes in the grouping chain with the storage records. The key nodes can be turning points in the chain, such as the location of a transfer station or a checkpoint, while the storage records are historical data in the offline database. The consistency result is determined by string matching or coordinate distance calculation, for example, by comparing the node coordinates (latitude and longitude) with the location in the record. If the deviation is less than the preset 50 meters, it is considered consistent, thereby quantifying the consistency score.
[0135] For example, if the consistency result meets the preset threshold, such as a score exceeding 80%, the matched data is classified as valid context information, which includes complete trajectory paragraphs for building a reliable review basis; if it does not meet the threshold, it is marked as pending data, and additional trajectory analysis results are obtained, which can be supplemented by external sensor data or satellite images to fill in the data gaps.
[0136] In one embodiment, the pending data is subjected to secondary verification through trajectory analysis results, combined with the continuity characteristics of historical trajectories, such as checking whether the trajectory is continuous in time without jumps or reasonable speed changes, to determine whether it can be included in the context information. For example, if the pending data shows that the vehicle suddenly jumps from position X to Y, but the analysis result shows that there is a reasonable explanation for the signal loss in between, it is included; otherwise, it is excluded, thereby improving the coherence of the context information, which can effectively reduce misjudgment and improve review efficiency in business.
[0137] For example, based on the context information after secondary verification, the information-complete review support dataset is sorted out according to the requirements of compliance review, which needs to include elements such as trajectory starting point, ending point, duration, and abnormality label, to ensure compliance with regulatory standards such as transportation time efficiency and route compliance requirements. Through this sorting, structured input can be provided for subsequent classification, avoiding review delays caused by scattered data.
[0138] Specifically, after obtaining the review support dataset, support vector machine algorithm is used for classification processing of the dataset. Support vector machine is a supervised learning model that separates data points by constructing a hyperplane. Its principle is to maximize the interval to separate different categories. Here, features such as trajectory deviation degree and speed anomaly frequency can be used as input vectors to train the model to distinguish between parts that meet compliance standards, such as normal transportation trajectories, and potential abnormal parts, such as trajectories that deviate from the route. The classified data groups are obtained, which involves feature extraction and kernel function selection, such as using radial basis function to process nonlinear data, thereby achieving efficient classification. In logistics compliance review, this can automatically identify potential violations, bringing higher automation level and accuracy.
[0139] For example, through the classified data groups, structured compliance review evidence is generated, such as sorting normal groups into report tables, including trajectory abstracts and compliance proofs, and labeling risk points for abnormal groups. Finally, the review support content is output, which can be a PDF report or a database entry for regulatory department review, ensuring logical closure from data extraction to output.
[0140] The present application also provides a system for sensitive data operation trajectory monitoring and compliance review, mainly including:
[0141] The enhanced trajectory data generation module is configured to extract operation time, source identifier and business association mark from the sensitive data operation trajectory by a preset module in an online environment, and to attach the marks to the core content of the trajectory record in an embedded manner to obtain enhanced trajectory data containing complete context.
[0142] The one-way transfer and mark retention judgment module is configured to perform one-way transfer through a one-way gate if the enhanced trajectory data meets the transmission format requirements of the one-way gate, and to judge whether the data after transfer retains the additional marks in offline storage.
[0143] The business association mark matching module is configured to match the business association mark by using an association algorithm according to the enhanced trajectory data after transfer, and to determine that the matched record forms a continuous operation chain.
[0144] The source identifier comparison module is configured to obtain the source identifier in the operation chain, compare the identifier with a preset system source library, and obtain a comparison result for subsequent review and judgment.
[0145] The operation chain sorting module is configured to sort the chain by operation time sequence if the comparison result shows that the source identifiers are consistent, and to judge whether the sorted chain reflects a legal operation sequence.
[0146] The clustering grouping module is configured to group related records by using a clustering algorithm according to the sorted operation chain, and to determine that the grouped chain provides a basis for responsibility attribution.
[0147] The cross-validation module is configured to cross-validate the grouped chain with historical trajectories in offline storage to obtain complete context information after validation to support compliance review.
[0148] The above is only a specific embodiment of the present specification, and those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the above-described system, module and unit can refer to the corresponding process in the foregoing method embodiments, which will not be described here. It should be understood that the protection scope of the present specification is not limited to this, and any person skilled in the art can easily think of various equivalent modifications or replacements within the technical range disclosed in the present specification, and these modifications or replacements should be covered within the protection scope of the present specification.
Claims
1. A method for monitoring and complying with sensitive data operation trajectories, characterized in that, The method includes: The operation time, source identifier, and business association marker are extracted from the sensitive data operation trajectory through a preset module in the online environment. The markers are then attached to the core content of the trajectory record by embedding them to obtain enhanced trajectory data containing complete context. If the enhanced trajectory data meets the transmission format requirements of the one-way gateway, then one-way transfer is performed through the gateway, and it is determined that the transferred data retains additional markers in offline storage. Based on the enhanced trajectory data after the transfer, an association algorithm is used to match business association markers to determine the matching records that form a continuous operation chain. This includes: extracting all trajectory data sequences and corresponding business association markers; using an association rule algorithm to calculate the support and confidence of adjacent record markers, and selecting high-confidence record pairs to form a candidate set; connecting candidate records in chronological order to generate an initial chain, and expanding endpoints that share the same marker and are time-continuous; verifying the consistency between the time difference and marker matching of adjacent records, and merging overlapping chains to obtain the final continuous operation chain. Obtain the source identifier in the operation chain, compare the identifier with the preset system source library, and obtain the comparison result for subsequent review and judgment; If the comparison results show that the source identifiers are consistent, the chain is sorted by the operation time series to determine whether the sorted chain reflects the legitimate operation sequence. Based on the sorted operation chain, a clustering algorithm is used to group the relevant records, and the grouped chain provides a basis for attribution of responsibility; By cross-validating the grouped chains with historical trajectories in offline storage, complete contextual information is obtained to support compliance reviews.
2. The method according to claim 1, characterized in that, The process involves extracting operation time, source identifier, and business association markers from sensitive data operation trajectories using a preset module in the online environment. These markers are then embedded into the core content of the trajectory record to obtain enhanced trajectory data containing complete context, including: The system acquires raw sensitive data operation trajectory records and parses the operation time, source identifier, and business association marker fields to obtain structured basic data. It standardizes business association markers using a preset field mapping table, splitting multi-value types into independent sub-markers and attaching unique sequence numbers. The system then uses vector embedding to convert the above data into fixed-dimensional vectors, which are concatenated to the core content of the original records to obtain enhanced trajectory vector data. Finally, through cosine similarity calculation, adjacent records with similarity exceeding a preset threshold are merged into the same business context chain; otherwise, independent records are retained.
3. The method according to claim 1, characterized in that, If the enhanced trajectory data meets the transmission format requirements of the one-way gateway, then a one-way transfer is performed through the gateway, and it is determined that the transferred data retains additional markers in offline storage, including: The enhanced trajectory data is made to conform to the one-way gateway transmission specification, and the format of the additional markers is verified. After the data is transferred through the gateway, the data integrity is checked in offline storage to confirm the integrity of the additional markers. The marker content is compared and classified for archiving, and the storage structure is periodically scanned to ensure the continuous availability of data.
4. The method according to claim 1, characterized in that, The process of acquiring the source identifier in the operation chain involves comparing the identifier with a preset system source database to obtain a comparison result for subsequent review and judgment, including: Extract the source identifier from the operation chain and standardize it; load the set of legal records from the preset system source library and perform precise matching item by item; if no match is found, calculate the cosine similarity. If it is higher than the preset threshold, it is determined to be a match and the corresponding record is bound; otherwise, it is determined to be a non-match; generate the comparison result containing the matching status and the associated record identifier and output it to the review interface.
5. The method according to claim 1, characterized in that, If the comparison results show that the source identifiers are consistent, the chain is sorted by the operation time series to determine whether the sorted chain reflects a valid operation sequence, including: Obtain source identifiers and operation time data to construct an initial operation chain; if the source identifiers are consistent, sort them by operation time and compare them against the preset legal order rules; if the order is abnormal, reverse the verification to locate the abnormal time point, analyze the deviation, use a decision tree to classify abnormal segments, and output the overall legality conclusion of the chain based on the matching judgment criteria.
6. The method according to claim 1, characterized in that, The process of grouping related records into groups based on the sorted operation chains, and determining the grouped chains, provides a basis for attribution of responsibility, including: Hierarchical clustering algorithm is used to calculate the distance matrix of records, generate a tree structure and cut it to obtain mutually exclusive groups; the operator identification and operation time in the group are extracted, and the cross-responsibility group is marked; the single responsibility rule is compared to obtain the set of legal groups, the number of chain segments corresponding to each person is counted, and the responsibility attribution mapping relationship is determined.
7. The method according to claim 1, characterized in that, The process involves cross-validating the grouped chains with historical trajectories in offline storage to obtain complete verified contextual information to support compliance reviews, including: Link offline storage historical trajectories, extract corresponding storage records to form a preliminary trajectory set; compare key nodes of the grouped chain with storage records, classify those that meet the threshold as valid context information, otherwise mark them as data to be processed and perform secondary verification; organize a complete review support dataset, use the support vector machine algorithm to classify compliant and potentially abnormal parts, generate structured review basis and output it.
8. A system for monitoring and complying with sensitive data operation trajectories, characterized in that, The system includes: The enhanced trajectory data generation module is used to extract operation time, source identifier and business association marker from sensitive data operation trajectory through preset modules in the online environment, and to attach the marker to the core content of trajectory record by embedding, so as to obtain enhanced trajectory data containing complete context; The one-way transfer and mark retention judgment module is used to determine whether the enhanced trajectory data meets the transmission format requirements of the one-way gateway, and to determine whether the data after transfer retains additional marks in offline storage. The business association tag matching module is used to match business association tags based on the enhanced trajectory data after the transfer, using an association algorithm to determine the matching records to form a continuous operation chain. This includes: extracting all trajectory data sequences and corresponding business association tags; using an association rule algorithm to calculate the support and confidence of adjacent record tags, and selecting high-confidence record pairs to form a candidate set; connecting candidate records sequentially to generate an initial chain, expanding endpoints that share the same tag and are temporally continuous; verifying the consistency between the time difference and tag matching of adjacent records, and merging overlapping chains to obtain the final continuous operation chain. The source identifier comparison module is used to obtain the source identifier in the operation chain, compare the identifier with the preset system source library, and obtain the comparison result for subsequent review and judgment. The operation chain sorting module is used to sort the chain by operation time series if the comparison results show that the source identifiers are consistent, and to determine whether the sorted chain reflects the legitimate operation sequence. The clustering and grouping module is used to group relevant records according to the sorted operation chain using a clustering algorithm, and the grouped chain provides a basis for responsibility attribution; The cross-validation module is used to cross-validate the grouped chains with historical trajectories in offline storage to obtain complete contextual information after validation to support compliance review.
Citation Information
Patent Citations
Government affair data sharing system based on data security law risk control mode
CN119989417A
National secret log auditing system
CN121151036A