Internet asset sensitive data management method and system based on digital model

By constructing a digital twin model and combining data type identification and spatiotemporal characteristics, dynamic monitoring and protection of sensitive Internet asset data throughout its entire lifecycle can be achieved. This solves the problems of insufficient monitoring of the dynamic data flow process and fragmented integrity verification in existing technologies, and improves the timeliness of data protection and the accuracy of anomaly identification.

CN120893078AActive Publication Date: 2025-11-04GUANGZHOU CANGHAI NETWORK TECHNOLOGY CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510973866.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-15
Publication Date
2025-11-04
Estimated Expiration
2045-07-15

AI Technical Summary

Technical Problem

Existing internet asset sensitive data management technologies lack the ability to proactively monitor the dynamic flow of data, making it difficult to block attacks in real time. Furthermore, data integrity verification and flow control are disconnected, making it difficult to locate the source and propagation path of data leaks. Traditional protection strategies have simple triggering conditions and are unable to cope with complex attack methods.

Method used

A digital model-based Internet asset sensitive data management system is constructed. By acquiring the type identification and feature encoding of target sensitive data to generate digital fingerprints, and combining spatiotemporal feature information to construct a digital twin model, the system can achieve real-time monitoring and protection of control behaviors, trigger preset protection strategies, and generate security audit logs.

Benefits of technology

It achieves proactive, multi-dimensional, real-time protection of sensitive data, improves the accuracy of abnormal behavior identification and the integrity of audit security, enhances the timeliness of data protection and the accuracy of feature extraction, and provides a cryptographic protection mechanism.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120893078A_ABST
    Figure CN120893078A_ABST
Patent Text Reader

Abstract

The invention relates to an internet asset sensitive data management method and system based on a digital model. The method comprises the steps of obtaining target sensitive data and performing type identification; performing feature coding on the target sensitive data based on the type identification result, extracting data features and generating digital fingerprints; obtaining spatio-temporal characteristic information associated with the target sensitive data, and constructing a digital twin model of the target sensitive data; the control behavior is monitored in real time based on the constructed digital twinborn model; when the abnormal control behavior is monitored, triggering a correspondingly set protection strategy; acquiring a data flow path based on the spatio-temporal feature information, and generating a security audit log; according to the method, the data-fingerprint-space-time three-in-one digital twinning model is constructed, so that active and multi-dimensional real-time protection of sensitive data is realized, and the effects of improving the timeliness of data protection, enhancing the accuracy of abnormal behavior recognition and improving the integrity of auditing safety are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of sensitive data management, and in particular to a method and system for managing sensitive data of Internet assets based on a digital model. Background Technology

[0002] With the rapid development of internet technology, the scale of internet assets of enterprises and institutions continues to expand, and the storage, transmission and processing of sensitive data (such as user privacy, trade secrets, financial information, etc.) are facing increasingly severe security challenges.

[0003] Currently, the management of sensitive data in internet assets mainly relies on traditional technologies such as static encryption, access control lists, and log auditing. While these methods provide basic security protection, they have significant limitations in practical applications. First, existing technologies typically only provide passive protection for data, lacking the ability to actively monitor the dynamic flow of data. For example, when data is abnormally copied or transmitted, existing technologies often can only detect the anomaly afterward through log analysis, making it difficult to block attacks in real time. Second, existing solutions mostly rely on a single dimension for data identification and control, judging solely based on data content or user permissions, while ignoring the spatiotemporal context of data operations. This makes it difficult for the system to distinguish between normal operations and malicious behavior, especially when facing internal threats or APT attacks, resulting in high false positive and false negative rates.

[0004] A more significant drawback is the disconnect between existing data integrity verification mechanisms and data flow control. While common hash verification or digital signature technologies can verify whether data has been tampered with, they cannot correlate when, where, or through what path the data was accessed or modified. This disconnect makes it difficult for security personnel to quickly pinpoint the source and propagation path of data breaches. Furthermore, the triggering conditions of traditional protection strategies are too simplistic, typically based on preset rules (such as IP blacklists) or fixed-time policies (such as work hour restrictions), making them ill-suited to increasingly complex attack methods, such as cross-timezone data penetration or covert theft using cloud platform APIs.

[0005] To address the aforementioned issues, existing technologies urgently need improvement. Summary of the Invention

[0006] To address the aforementioned shortcomings, this application provides a method and system for managing sensitive Internet asset data based on a digital model.

[0007] The above-mentioned objective of this application is achieved through the following technical solution:

[0008] A method for managing sensitive internet asset data based on a digital model, comprising the following steps:

[0009] Acquire target sensitive data and identify the type of the target sensitive data;

[0010] Based on the type recognition results, feature encoding is performed on the target sensitive data to extract data features and generate digital fingerprints;

[0011] Obtain the spatiotemporal feature information associated with the target sensitive data, and construct a digital twin model of the target sensitive data based on the spatiotemporal feature information and digital fingerprint;

[0012] The control behavior is monitored in real time based on the constructed digital twin model, and data integrity is verified based on digital fingerprints.

[0013] When abnormal control behavior is detected, the corresponding protection policy is triggered;

[0014] Data flow paths are obtained based on spatiotemporal feature information, and security audit logs are generated.

[0015] By adopting the above technical solutions, a digital twin model based on digital fingerprints and spatiotemporal features is constructed to achieve dynamic monitoring and protection of sensitive data throughout its entire lifecycle. First, by acquiring target sensitive data and identifying its type, a differentiated feature encoding method is used to extract the correlation features of structured data and the essential features of unstructured data, generating a uniquely identifiable digital fingerprint to solve the problem of single data identification dimensions in traditional technologies. Second, by collecting the timestamps, physical location identifiers of data storage nodes, and network device information in the data transmission path, a complete contextual environment containing spatiotemporal features is constructed to compensate for the lack of spatiotemporal information in existing technologies. Based on this, Digital fingerprints are linked and bound to spatiotemporal features in multiple dimensions to form the core verification basis of the digital twin model, achieving an organic unity of data integrity verification and flow control. The digital twin model monitors control behavior in real time; when anomalies in the digital fingerprint or mismatches in spatiotemporal features are detected, a preset protection strategy is immediately triggered, and a security audit log containing the data flow path is generated based on complete spatiotemporal feature information. This application, by constructing a data-fingerprint-spatiotemporal three-in-one digital twin model, achieves proactive, multi-dimensional, real-time protection of sensitive data, improving the timeliness of data protection, enhancing the accuracy of abnormal behavior identification, and improving the integrity of audit security.

[0016] In a preferred embodiment, this application can be further configured such that: the type recognition result includes structured data and unstructured data; the step of feature encoding the target sensitive data based on the type recognition result, extracting data features, and generating a digital fingerprint includes the following steps:

[0017] Extract structured feature vectors from structured data based on pre-set analysis and extraction strategies;

[0018] Extract unstructured feature vectors from unstructured data based on pre-set mining and extraction strategies;

[0019] The structured and unstructured feature vectors are standardized and fused to generate data features.

[0020] By adopting the above technical solution, differentiated feature extraction strategies are used to process different types of data, achieving accurate feature encoding of sensitive data. For structured data, the inherent inter-table relationships and constraint rules are obtained through analysis and extraction strategies, generating feature vectors that reflect the data structure characteristics. For unstructured data, mining and extraction strategies are used to deeply analyze content patterns and semantic relationships, extracting feature vectors that can characterize the essence of the data. This application standardizes, transforms, and fuses the two types of feature vectors, preserving the organizational characteristics of structured data while capturing the content characteristics of unstructured data, ultimately generating a digital fingerprint that comprehensively characterizes the data features. This achieves a unified feature representation for different types of data, improving the accuracy of feature extraction and enhancing the uniqueness of digital fingerprints, providing a feature foundation for the subsequent construction of digital twin models.

[0021] In a preferred embodiment, this application can be further configured as follows: the step of extracting structured feature vectors from structured data based on a pre-set analysis and extraction strategy includes the following steps:

[0022] Establish a connection graph with structured data as data nodes and analyze the reference relationships of structured data to quantify the topological importance index of data nodes;

[0023] Extract constraint rule features between structured data;

[0024] The connectivity graph, topological importance index, and constraint rule features are vectorized and encoded, and structured feature vectors are generated through feature dimensionality reduction.

[0025] By adopting the above technical solution, a multi-dimensional feature extraction system for structured data is established to achieve in-depth mining of the inherent characteristics of the data. First, by constructing a data node connection graph and analyzing the reference relationships between data, the topological importance of each data node in the overall structure is quantified. At the same time, constraint rule features between data are extracted to capture the correlation characteristics at the business logic level. Finally, the connection graph, topological importance index, and constraint rule features are vectorized and dimensionality reduced to generate structured feature vectors. This application, through the aforementioned hierarchical feature extraction method, achieves a comprehensive representation of data structure features, correlation features, and business features, which can comprehensively reflect data characteristics, improve feature discrimination, and enhance model interpretability.

[0026] In a preferred embodiment, this application can be further configured as follows: the step of extracting essential features of unstructured data based on a pre-set mining and extraction strategy includes the following steps:

[0027] Content slicing and pattern analysis are performed on unstructured data to identify recurring feature fragments and key content patterns in the unstructured data;

[0028] Construct a semantic association network among unstructured data;

[0029] Vectorize repetitive feature fragments, key content patterns, and semantic association networks;

[0030] Unstructured feature vectors are generated through feature selection.

[0031] By adopting the above technical solution, a multi-level feature mining system for unstructured data is constructed to extract the essential features of the data. First, content slicing technology is used to decompose unstructured data into analyzable units, and pattern recognition algorithms are used to capture recurring feature fragments and key content patterns. At the same time, a semantic association network between data is established to reveal deep-level connections at the content level. The extracted feature fragments, content patterns, and semantic association networks are transformed into a unified vectorized representation. Finally, feature selection algorithms are used to filter feature combinations and generate unstructured feature vectors. This application, through the aforementioned feature mining method, achieves multi-faceted extraction of content features and semantic features of unstructured data, which has the effect of improving feature representation capabilities and enhancing the depth of content understanding.

[0032] In a preferred embodiment, this application can be further configured as follows: the step of obtaining spatiotemporal feature information associated with the target sensitive data, and constructing a digital twin model of the target sensitive data based on the spatiotemporal feature information and the digital fingerprint, includes the following steps:

[0033] Identify the data storage nodes of the target sensitive data, and obtain the timestamps and physical location identifiers from the data storage nodes as basic spatiotemporal features;

[0034] Identify the data transmission path of the target sensitive data, and collect network device information in the data transmission path through the acquisition terminal as supplementary spatiotemporal features;

[0035] Spatiotemporal feature information is generated based on basic and supplementary spatiotemporal features, and the spatiotemporal feature information is then linked and bound to digital fingerprints in multiple dimensions.

[0036] A digital twin model containing data features, digital fingerprints, and spatiotemporal features is constructed based on the multi-dimensional association and binding results.

[0037] By adopting the above technical solutions, a multi-source spatiotemporal feature acquisition and fusion system is constructed to accurately depict the entire lifecycle trajectory of data. First, timestamps and physical location identifiers are obtained from data storage nodes as basic spatiotemporal features to ensure the authenticity and reliability of the data's static attributes. Simultaneously, network device information in the data transmission path is captured by the acquisition terminal as dynamic supplementary features to completely record the data's flow trajectory. The basic and supplementary features are then fused to generate unified spatiotemporal feature information. Finally, the spatiotemporal features are linked and bound to digital fingerprints in multiple dimensions to construct a three-dimensional digital twin model containing data features, digital fingerprints, and spatiotemporal features. This application, through the aforementioned multi-faceted and multi-layered spatiotemporal feature acquisition and fusion method, achieves complete recording and dynamic tracking of data spatiotemporal attributes, improving data tracing accuracy, enhancing anomaly detection capabilities, and supporting end-to-end auditing, providing reliable spatiotemporal dimension support for the secure management of sensitive data.

[0038] In a preferred embodiment, this application can be further configured as follows: the step of generating spatiotemporal feature information based on basic spatiotemporal features and supplementary spatiotemporal features, and binding the spatiotemporal feature information with digital fingerprints in a multi-dimensional manner, includes the following steps:

[0039] Establish temporal, spatial, and transmission associations between digital fingerprint features and timestamps, physical locations, and network paths, respectively;

[0040] Cross-validation of temporal correlation, spatial correlation, and transport correlation was performed.

[0041] Generate associated credentials containing digital fingerprints and spatiotemporal feature information using a pre-set encryption algorithm.

[0042] By adopting the above technical solution, a multi-dimensional data fingerprint and spatiotemporal feature association system is established to achieve the binding and trusted verification of data security attributes. First, the temporal association between digital fingerprints and timestamps, the spatial association with physical locations, and the transmission association with network paths are constructed respectively, forming a three-dimensional security feature network covering the entire data lifecycle. A cross-validation mechanism ensures the consistency of the associations across dimensions, eliminating potential spatiotemporal logical contradictions. Finally, a pre-set encryption algorithm solidifies the verified associations into tamper-proof electronic credentials, achieving cryptographic-level binding of data fingerprints and spatiotemporal features. This application, through multi-dimensional association and cross-validation, achieves comprehensive protection of data security attributes, enhancing the reliability of data traceability and improving the sensitivity of anomaly detection, providing key technical support for building a trusted data security protection system.

[0043] In a preferred embodiment, this application can be further configured as follows: the step of generating an associated credential containing digital fingerprint and spatiotemporal feature information through a pre-set encryption algorithm includes the following steps:

[0044] A layered encryption architecture is adopted to perform layered joint processing of digital fingerprints and spatiotemporal feature information;

[0045] The digital fingerprint is encrypted using a pre-set asymmetric encryption algorithm;

[0046] Spatiotemporal feature information is encrypted using a pre-set lightweight encryption algorithm;

[0047] The encrypted digital fingerprint and spatiotemporal feature information are combined and encapsulated to generate associated credentials.

[0048] By adopting the above technical solution, employing layered encryption and combined encapsulation techniques, secure binding and processing of digital fingerprints and spatiotemporal feature information are achieved. First, a layered encryption architecture is constructed, using an asymmetric encryption algorithm for the digital fingerprint to ensure its immutability, while a lightweight encryption algorithm is used for the spatiotemporal feature information to guarantee processing efficiency. The two types of encrypted secure data are then standardized and combined to generate an associated credential with integrity and confidentiality. This application, through the aforementioned differentiated layered encryption strategy, not only ensures the high security requirements of the core digital fingerprint but also meets the real-time processing needs of spatiotemporal feature information. It optimizes encryption processing efficiency, ensures the security of sensitive data, and enhances the verifiability of credentials, providing a cryptographic guarantee mechanism for the secure management of target privacy data.

[0049] The second objective of this invention is achieved through the following technical solution:

[0050] A digital model-based Internet asset sensitive data management system includes:

[0051] The type recognition module is used to acquire target sensitive data and perform type recognition on the target sensitive data;

[0052] The feature extraction module is used to encode the target sensitive data based on the type recognition results, extract data features and generate digital fingerprints;

[0053] The model building module is used to acquire spatiotemporal feature information associated with target sensitive data, and to build a digital twin model of the target sensitive data based on the spatiotemporal feature information and digital fingerprint.

[0054] The real-time monitoring module is used to monitor control behavior in real time based on the constructed digital twin model and to verify data integrity based on digital fingerprints.

[0055] The protection module is used to trigger the corresponding protection policy when abnormal control behavior is detected.

[0056] The audit log generation module is used to obtain the data flow path based on spatiotemporal feature information and generate security audit logs.

[0057] By adopting the above technical solution, the following modules are used: a type identification module for acquiring target sensitive data and identifying its type; a feature extraction module for encoding the target sensitive data based on the type identification results, extracting data features, and generating a digital fingerprint; a model building module for acquiring the spatiotemporal feature information associated with the target sensitive data and constructing a digital twin model of the target sensitive data based on this spatiotemporal feature information and the digital fingerprint; a real-time monitoring module for real-time monitoring of control behavior based on the constructed digital twin model and verifying data integrity based on the digital fingerprint; a protection module for triggering the corresponding protection strategy when abnormal control behavior is detected; and an audit log generation module for acquiring the data flow path based on the spatiotemporal feature information and generating a security audit log.

[0058] In summary, this application includes at least one of the following beneficial technical effects:

[0059] 1. This application achieves proactive, multi-dimensional, real-time protection of sensitive data by constructing a three-in-one digital twin model of data-fingerprint-spatiotemporal integration, which can improve the timeliness of data protection, enhance the accuracy of abnormal behavior identification, and improve the integrity of audit security.

[0060] 2. This application standardizes and fuses structured feature vectors with unstructured feature vectors, thus preserving the organizational features of structured data while capturing the content characteristics of unstructured data. This results in the generation of a digital fingerprint that comprehensively represents the data features, achieving a unified feature representation for different types of data. This improves the accuracy of feature extraction and enhances the uniqueness of digital fingerprints, providing a feature foundation for the subsequent construction of digital twin models.

[0061] 3. This application, through a differentiated layered encryption strategy, not only ensures the high security requirements of the core digital fingerprint, but also meets the real-time requirements of spatiotemporal feature information processing. It has the effects of optimizing encryption processing efficiency, ensuring the security of sensitive data, and enhancing the verifiability of credentials, providing a cryptographic guarantee mechanism for the secure management of target privacy data. Attached Figure Description

[0062] Figure 1 This is a flowchart of an embodiment of a digital model-based method for managing sensitive data of Internet assets according to this application;

[0063] Figure 2 This is a flowchart of step S20 in an embodiment of a digital model-based Internet asset sensitive data management method of this application;

[0064] Figure 3This is a flowchart of step S21 in an embodiment of a digital model-based Internet asset sensitive data management method of this application;

[0065] Figure 4 This is a flowchart of step S22 in an embodiment of a digital model-based Internet asset sensitive data management method of this application;

[0066] Figure 5 This is a flowchart of step S30 in an embodiment of a digital model-based Internet asset sensitive data management method of this application;

[0067] Figure 6 This is a flowchart of step S33 in an embodiment of a digital model-based Internet asset sensitive data management method of this application;

[0068] Figure 7 This is a flowchart of step S333 in an embodiment of a digital model-based Internet asset sensitive data management method of this application. Detailed Implementation

[0069] The following is in conjunction with the appendix Figure 1-7 This application will be described in further detail.

[0070] In one embodiment, such as Figure 1 As shown, this application discloses a method for managing sensitive data of Internet assets based on a digital model, which specifically includes the following steps:

[0071] S10: Acquire target sensitive data and identify the type of target sensitive data;

[0072] In this embodiment, the target sensitive data refers to various data assets that require special protection, including but not limited to personal privacy information, trade secrets, and financial transaction records, which have high value or high risk attributes. Type identification involves distinguishing between structured and unstructured data through predefined classification rules or machine learning algorithms. Structured data includes database tables and spreadsheets with clear patterns and fixed fields, while unstructured data includes text, images, audio, video, and other data types without fixed formats.

[0073] S20: Based on the type recognition results, perform feature encoding on the target sensitive data, extract data features, and generate a digital fingerprint;

[0074] In this embodiment, a digital fingerprint is a unique and irreversible data identifier generated by a feature extraction algorithm. Essentially, it is a mathematical abstraction of data features, enabling accurate identification and integrity verification of data identity without exposing the original data. The generation of a digital fingerprint involves topological relationship analysis of structured data and semantic feature extraction of unstructured data. For example, graph neural networks are used to process database relationships, and natural language processing techniques are used to parse document content.

[0075] S30: Obtain the spatiotemporal feature information associated with the target sensitive data, and construct a digital twin model of the target sensitive data based on the spatiotemporal feature information and the digital fingerprint;

[0076] In this embodiment, the spatiotemporal feature information may include the timestamp sequence of data operations, the geographical coordinates of storage nodes, and the network topology information of data transmission paths, which together constitute a spatiotemporal trajectory record of the entire data lifecycle. For example, precise time stamps can be obtained through distributed system clock synchronization, and the geographical location of the server can be recorded using a GPS module. The digital twin model is a virtual image system constructed by dynamically binding digital fingerprints with spatiotemporal features. It can reflect the changes in data status and security situation in the physical world in real time. Its construction requires multi-dimensional association between data fingerprints and spatiotemporal features, such as establishing time series indexes and spatial coordinate mapping relationships.

[0077] S40: Real-time monitoring of control behavior based on the constructed digital twin model, and data integrity verification based on digital fingerprints;

[0078] In this embodiment, the control behavior includes all operation requests such as access, modification, and transmission of sensitive data. Its monitoring scope includes both the normal operations of legitimate users and the abnormal behavior of potential attackers.

[0079] S50: When abnormal control behavior is detected, the corresponding protection policy is triggered;

[0080] In this embodiment, abnormal control behavior refers to the pre-defined identification of abnormal control behavior. The protection strategy is a preset gradient response mechanism that automatically triggers security measures of different intensities, from alarm prompts to data isolation, based on the risk level of the abnormal control behavior.

[0081] S60: Obtain the data flow path based on spatiotemporal feature information and generate security audit logs.

[0082] In this embodiment, the security audit log is a complete operation record that has been encrypted. It not only contains traditional event description information, but also integrates contextual data such as spatiotemporal features, forming a traceable chain of evidence.

[0083] Specifically, the system first acquires sensitive information such as transaction records and customer profiles through data interfaces, and then uses a pre-trained classification model to identify data types. For structured account balance tables, it can extract field constraints and primary / foreign key connection graphs, while for unstructured scanned contracts, it can perform text segmentation and keyword extraction. The extracted feature vectors are then standardized to generate unique digital fingerprints, and the location coordinates and recent access time of the data storage server are collected. In the digital twin model, each control behavior, such as the fingerprint of each transfer record, is bound to its occurrence time and the IP address of the operating terminal. When abnormal control behavior is detected, such as a batch data export request initiated from an unfamiliar device in the early morning, the model immediately compares the current operation with historical patterns, triggers an access blocking mechanism, and records the complete operation trajectory.

[0084] In one embodiment, the type identification result includes structured data and unstructured data, such as... Figure 2 As shown, step S20 includes the following steps:

[0085] S21: Extract structured feature vectors from structured data based on pre-set analysis and extraction strategies;

[0086] S22: Extract unstructured feature vectors from unstructured data based on pre-set mining and extraction strategies;

[0087] S23: Standardize and fuse structured and unstructured feature vectors to generate data features.

[0088] In this embodiment, structured data refers to data types with predefined formats and fixed organizational patterns, including tabular data, spreadsheet files, and XML / JSON in relational databases. Specifically, it can be implemented using a two-dimensional table structure in a relational database. Machine readability is achieved through predefined data patterns. A significant characteristic of structured data is its clearly defined fields, fixed data types, and predefined relational constraints. Unstructured data refers to data content without a fixed format or pattern, including text documents, images, videos, audio files, and social media content. Specifically, it can be implemented using natural language text or multimedia file formats. Its content features need to be analyzed using pattern recognition technology. The characteristic of unstructured data is its flexible and varied information organization, requiring the extraction of its inherent features through specific technical means. The structured feature vector is a mathematical representation generated by analyzing the inherent organizational characteristics of structured data, including quantitative indicators of table structure features, field relationships, and data integrity constraints, accurately reflecting the data's organizational structure and business logic. The unstructured feature vector is... Numerical features extracted after in-depth analysis of unstructured data content include semantic features of text, visual features of images, and spectral features of audio. Unstructured feature vectors can be extracted using machine learning algorithms, effectively capturing the essential attributes of unstructured data. Analysis and extraction strategies refer to the set of rules for extracting features from structured data. Specifically, graph theory algorithms can be used to construct a data node connection graph, and topological weight indicators can be calculated by quantifying the reference relationships between nodes. Mining and extraction strategies refer to the set of rules for extracting features from unstructured data. Specifically, natural language processing techniques can be used for semantic slicing, and pattern matching algorithms can be used to identify repetitive content fragments. Standardization is a mathematical process that converts different types of feature vectors into a unified dimension and distribution range, eliminating scale differences between different features and creating comparable conditions for subsequent feature fusion. Feature fusion is the process of organically integrating standardized structured and unstructured feature vectors. A specific fusion algorithm is used to establish a correlation mapping between the two types of features, forming a composite feature representation that comprehensively characterizes the data properties.

[0089] Specifically, differentiated feature extraction strategies are employed to process different types of data, achieving accurate feature encoding for sensitive data. For structured data, the inherent inter-table relationships and constraint rules are obtained through analysis and extraction strategies, generating feature vectors that reflect the data structure characteristics. For unstructured data, mining and extraction strategies are used to deeply analyze content patterns and semantic relationships, extracting feature vectors that can characterize the essence of the data. Among these, structured feature vectors can be generated by quantifying the strength of relationships and constraint rules between data tables, while unstructured feature vectors can be constructed by statistically analyzing the distribution density of key content patterns. After normalization, the two types of feature vectors are eliminated through orthogonal transformation, and finally fused to form data features with unique identifiers.

[0090] In one embodiment, such as Figure 3 As shown, step S21 includes the following steps:

[0091] S211: Establish a connection graph with structured data as data nodes and analyze the reference relationships of structured data to quantify the topological importance index of data nodes;

[0092] S212: Extract the constraint rule features between structured data;

[0093] S213: Vectorize the connectivity graph, topological importance index, and constraint rule features, and generate structured feature vectors through feature dimensionality reduction.

[0094] In this embodiment, the connection graph is a data relationship network model, which can be implemented using a graph database or graph computing framework. It visually reflects the references and dependencies between structured data. Each structured data table or data entity is abstracted as a graph node, and foreign key relationships or business associations between data tables are abstracted as edges, forming a visualized data topology. Reference relationships are the logical connections established between tables in structured data through primary and foreign keys, including one-to-one, one-to-many, and many-to-many association types. These reference relationships reflect the actual interaction patterns of business data. Topological importance indicators are quantitative parameters calculated using graph algorithms, including but not limited to graph theory indicators such as degree centrality, proximity centrality, and betweenness centrality of nodes, used to identify key data nodes and... Its impact on the overall data structure; constraint rules are predefined data integrity constraints in structured data, including primary key constraints, foreign key constraints, uniqueness constraints, check constraints, and other business rules. These constraints reflect the business logic and verification rules set during data modeling; vectorization encoding is the process of converting non-numerical graph structure information and constraint rules into numerical vectors. Commonly used techniques include graph embedding algorithms and feature hashing methods. The purpose is to convert complex structured information into a numerical representation that can be processed by machine learning models; feature dimensionality reduction is the compression and optimization of high-dimensional feature vectors through techniques such as principal component analysis or autoencoders, removing redundant information and retaining the most discriminative feature dimensions, ultimately generating low-dimensional dense structured feature vectors.

[0095] Specifically, firstly, a connection graph is constructed based on the storage structure and relationships of the structured data, with each data table or field as a node and foreign key references or logical dependencies as edges. By analyzing the connection density and path depth between nodes, in-degree, betweenness centrality, and other indicators of each node are calculated to quantify its importance in the data system. Simultaneously, constraint rules such as primary and foreign key constraints, uniqueness restrictions, and value ranges are extracted from the database schema definition or business rules to form a set of rule features. The adjacency matrix, node importance indicators, and rule features of the graph are normalized, and high-dimensional vectors are generated using graph convolutional networks or feature concatenation methods. Finally, noise and redundant dimensions are removed using dimensionality reduction algorithms to generate low-dimensional dense feature vectors, which serve as the representation of the structured data and input into subsequent processing flows.

[0096] In one embodiment, such as Figure 4 As shown, step S22 includes the following steps:

[0097] S221: Perform content slicing and pattern analysis on unstructured data to identify recurring feature fragments and key content patterns in unstructured data;

[0098] S222: Construct a semantic association network among unstructured data;

[0099] S223: Vectorize the repetitive feature fragments, key content patterns, and semantic association networks;

[0100] S224: Generate unstructured feature vectors through feature selection.

[0101] In this embodiment, content slicing involves dividing unstructured data into analyzable data units according to preset rules. Different segmentation strategies are used based on the data type: semantic paragraph segmentation or sliding window slicing is used for text data; region segmentation or grid partitioning is used for image data; and audio / video data is segmented according to time sequence to eliminate data redundancy and locate key information regions. Pattern analysis refers to identifying high-frequency fragments and latent semantic structures in the data through feature matching algorithms. Specifically, regular expression matching or deep learning models can be used to extract recognizable data patterns. Repeated feature fragments are high-frequency content units identified in unstructured data through pattern matching algorithms. In text, this manifests as the repetition of specific phrases or named entities; in images, it manifests as the repetition of similar visual patterns; and in audio / video, it manifests as the cyclical appearance of specific voiceprints or visual features. These repeated feature fragments carry key information about the data. Key content patterns are discriminative combinations of data features mined through machine learning algorithms, including topic distribution features in text, salient region features in images, and acoustic features in audio. These key content patterns can effectively distinguish different categories of unstructured data. Semantic association networks are graph structure models built by analyzing the contextual relationships between unstructured data. Nodes represent data entities or content elements, and edges represent semantic similarity or co-occurrence relationships. They can be implemented using graph neural networks or co-occurrence analysis algorithms to reveal implicit relationships between data. Vectorization is the process of converting unstructured content features into numerical vectors. It can be implemented using word embedding techniques or matrix factorization methods to achieve computer-processable feature representation. Feature selection is the process of selecting a discriminative subset of features from high-dimensional vectors. It can be implemented using information gain algorithms or principal component analysis methods to reduce computational complexity and improve feature quality.

[0102] Specifically, the content slicing operation for unstructured data involves: first, dividing text, image, or audio / video data into fixed-length blocks, such as using a sliding window mechanism to slice text data at fixed character intervals; then, detecting recurring feature fragments through a pattern analysis module, such as identifying high-frequency error code patterns in log files or discovering specific shape contours in image data; the semantic association network construction stage analyzes the contextual relationships between data slices, such as building a keyword co-occurrence map in a document collection or constructing a spatiotemporal association model in a video frame sequence; the vectorization representation process converts the above analysis results into multi-dimensional numerical vectors, such as using a bag-of-words model to encode text slices or using a convolutional neural network to extract image features; finally, removing redundant dimensions through feature selection algorithms, such as using a chi-square test to select feature terms strongly correlated with data sensitivity, forming unstructured feature vectors that can be used for subsequent processing.

[0103] In one embodiment, such as Figure 5 As shown, step S30 includes the following steps:

[0104] S31: Identify the data storage nodes of the target sensitive data, and obtain the timestamp and physical location identifier from the data storage nodes as basic spatiotemporal features;

[0105] S32: Identify the data transmission path of the target sensitive data, and collect network device information in the data transmission path as supplementary spatiotemporal features through the acquisition terminal;

[0106] S33: Generate spatiotemporal feature information based on basic spatiotemporal features and supplementary spatiotemporal features, and associate and bind the spatiotemporal feature information with digital fingerprints in multiple dimensions;

[0107] S34: Construct a digital twin model containing data features, digital fingerprints, and spatiotemporal feature information based on multi-dimensional association and binding results.

[0108] In this embodiment, data storage nodes are physical or logical storage units that carry target sensitive data, including but not limited to infrastructure such as database servers, cloud storage instances, and edge storage devices. These nodes are distinguished and managed by unique identifiers, forming the basic architecture of data storage. A timestamp is a precise time record of the data operation, which can be implemented using a clock source synchronized with a network time protocol, used to mark the time node of data access or modification. A physical location identifier is the spatial coordinate information of the data storage device, which can be implemented using a GPS positioning module or an IP geolocation database, used to determine the physical location of the data storage. The data transmission path is the communication trajectory composed of devices and links that the data passes through during transmission in the network, including interface information of network devices such as routers, switches, and gateways. The connectivity relationships reflect the actual network topology of data flow; network device information consists of device identifiers involved in the data transmission path, including network layer characteristics such as the device's MAC address, IP address, and VLAN identifier, as well as metadata such as the device's geographical location and management domain, used to supplement the description of the data transmission context; multi-dimensional association binding establishes a dynamic link between spatiotemporal characteristics of different dimensions and data fingerprints, which can be implemented using blockchain technology or distributed ledger technology to ensure the immutable association between spatiotemporal characteristics and data fingerprints; the digital twin model is a virtual image system constructed by integrating data characteristics, digital fingerprints, and spatiotemporal characteristics. The digital twin model not only statically reflects the current state of the data, but also dynamically tracks the data change process, realizing real-time synchronization and interaction between physical data and the virtual model.

[0109] Specifically, when sensitive target data is stored, the data storage node generates basic spatiotemporal features containing timestamps and physical location identifiers. For example, the timestamp can be recorded as a millisecond-level time value in UTC format, and the physical location identifier can be recorded as latitude and longitude coordinates. During the transmission of this data, the acquisition terminal deployed at the network boundary captures the network device information flowing through it in real time, such as router IP addresses or switch port numbers, forming supplementary spatiotemporal features. After the basic spatiotemporal features and supplementary spatiotemporal features are normalized, spatiotemporal feature information is generated through a hash algorithm. This spatiotemporal feature information is bound to the previously generated data fingerprint through an encryption algorithm, such as using asymmetric encryption to encapsulate the data fingerprint and spatiotemporal feature information into a digital envelope. Finally, a digital twin model is constructed based on the binding result. This model maintains the mapping relationship between data features, digital fingerprints, and spatiotemporal features in memory, such as using a graph database to store multidimensional relationships.

[0110] In one embodiment, such as Figure 6 As shown, step S33 includes the following steps:

[0111] S331: Establish temporal, spatial, and transmission associations between digital fingerprint features and timestamps, physical locations, and network paths, respectively;

[0112] S332: Perform cross-validation of temporal correlation, spatial correlation, and transport correlation;

[0113] S333: Generate associated credentials containing digital fingerprints and spatiotemporal feature information through a pre-set encryption algorithm.

[0114] In this embodiment, temporal association refers to the dynamic association established between digital fingerprints and operation timestamps through specific time series analysis methods. Temporal association not only records the correspondence at a single point in time but also captures the temporal patterns and abnormal temporal characteristics of data operations through time series pattern analysis. Spatial association refers to the geospatial binding relationship between digital fingerprints and physical location identifiers. Geocoding technology transforms location information into spatial feature vectors and establishes a multi-dimensional spatial mapping with the digital fingerprint, achieving geofencing and regional access control for data operations. Transmission association refers to the network topology relationship between digital fingerprints and network path features, achieved by analyzing device sequences in the path... Network characteristics such as transmission delay and hop count are used to construct a network topology fingerprint of data flow, providing a network layer verification basis for data transmission security. Cross-validation is a technical process for verifying the consistency of three types of associations. By verifying the logical rationality of time sequence and spatial movement, the degree of matching between network path and geographical location, and the correspondence between operation sequence and transmission delay, the integrity and credibility of the association system are ensured. The association credential is a tamper-proof data packet generated by cryptographic methods. It adopts a layered encryption architecture to protect the association, including the encrypted digest of the digital fingerprint, the signature verification of spatiotemporal characteristics, and the integrity check code of the association, forming a verifiable security token.

[0115] Specifically, a multi-dimensional data fingerprint and spatiotemporal feature association system is established to achieve data security attribute binding and trusted verification. First, the temporal association between digital fingerprint and timestamp, the spatial association with physical location, and the transmission association with network path are constructed to form a three-dimensional security feature network covering the entire data lifecycle. The consistency of the association relationships in each dimension is ensured through a cross-validation mechanism to eliminate potential spatiotemporal logical contradictions. Finally, a pre-set encryption algorithm is used to solidify the verified association relationships into tamper-proof electronic credentials, achieving cryptographic-level binding of data fingerprint and spatiotemporal features.

[0116] In one embodiment, such as Figure 7 As shown, step S333 includes the following steps:

[0117] S3331: Employs a layered encryption architecture to perform layered joint processing of digital fingerprints and spatiotemporal feature information;

[0118] S3332: Encrypt the digital fingerprint using a pre-set asymmetric encryption algorithm;

[0119] S3333: Encrypt spatiotemporal feature information using a pre-set lightweight encryption algorithm;

[0120] S3334: Combine and encapsulate the encrypted digital fingerprint and spatiotemporal feature information to generate associated credentials.

[0121] In this embodiment, the layered encryption architecture employs differentiated encryption methods for data with varying sensitivities. Specifically, it can be implemented using an encryption framework that separates the core data layer from the auxiliary information layer. This separation ensures strong protection for critical data while maintaining efficient processing of auxiliary information. The asymmetric encryption algorithm uses a public-key and private-key pairing for data encryption, specifically employing RSA or ECC algorithms. This asymmetric characteristic ensures the immutability and verifiability of the digital fingerprint. The lightweight encryption algorithm is suitable for high-frequency, low-latency scenarios, specifically employing AES-GCM or ChaCha20 algorithms. This symmetric encryption mechanism reduces computational resource consumption while ensuring the security of spatiotemporal feature information. The combined encapsulation is the operation of integrating different encryption results into a single data packet, specifically employing TLV encoding or ASN.1 structures. This standardized encapsulation format ensures the integrity and resolvability of the associated credential. The associated credential is the final generated data credential entity with complete security attributes. The associated credential not only contains the encrypted original data but also integrates verification information such as digital signatures, timestamps, and version identifiers, forming a self-contained, verifiable, and secure data unit capable of secure transmission and verification in a distributed environment.

[0122] Specifically, in the process of generating associated credentials, the digital fingerprint is first used as the core data layer and encrypted using an asymmetric encryption algorithm, such as RSA-2048, to form an irreversible encrypted data block. Simultaneously, spatiotemporal feature information is used as an auxiliary information layer and encrypted quickly using a lightweight encryption algorithm, such as AES-128-GCM, to encrypt timestamps, physical locations, and network path information. Then, a layered encryption architecture is used to jointly process the two encryption layers, establishing a logical relationship between the encrypted data. Finally, the encrypted digital fingerprint and spatiotemporal feature information are combined and encapsulated according to a predefined data structure, such as using TLV encoding to encapsulate the encrypted data block, checksum, and metadata into a binary data packet, generating an associated credential with integrity and verifiability.

[0123] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0124] In one embodiment, a digital model-based internet asset sensitive data management system is provided, which corresponds one-to-one with the digital model-based internet asset sensitive data management method described in the previous embodiment. The digital model-based internet asset sensitive data management system includes:

[0125] The type recognition module is used to acquire target sensitive data and perform type recognition on the target sensitive data;

[0126] The feature extraction module is used to encode the target sensitive data based on the type recognition results, extract data features and generate digital fingerprints;

[0127] The model building module is used to acquire spatiotemporal feature information associated with target sensitive data, and to build a digital twin model of the target sensitive data based on the spatiotemporal feature information and digital fingerprint.

[0128] The real-time monitoring module is used to monitor control behavior in real time based on the constructed digital twin model and to verify data integrity based on digital fingerprints.

[0129] The protection module is used to trigger the corresponding protection policy when abnormal control behavior is detected.

[0130] The audit log generation module is used to obtain the data flow path based on spatiotemporal feature information and generate security audit logs.

[0131] For specific limitations regarding the internet asset sensitive data management system based on a digital model, please refer to the limitations of the internet asset sensitive data management method based on a digital model mentioned above, which will not be repeated here. Each module in the aforementioned internet asset sensitive data management system based on a digital model can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the corresponding operations of each module.

[0132] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A method for managing sensitive Internet asset data based on a digital model, characterized in that: Including the following steps: Acquire target sensitive data and identify the type of the target sensitive data; Based on the type recognition results, feature encoding is performed on the target sensitive data to extract data features and generate digital fingerprints; Obtain the spatiotemporal feature information associated with the target sensitive data, and construct a digital twin model of the target sensitive data based on the spatiotemporal feature information and digital fingerprint; The control behavior is monitored in real time based on the constructed digital twin model, and data integrity is verified based on digital fingerprints. When abnormal control behavior is detected, the corresponding protection policy is triggered; Data flow paths are obtained based on spatiotemporal feature information, and security audit logs are generated.

2. The method for managing sensitive Internet asset data based on a digital model according to claim 1, characterized in that: The type identification result includes structured data and unstructured data. The step of feature encoding the target sensitive data based on the type identification result, extracting data features, and generating a digital fingerprint includes the following steps: Extract structured feature vectors from structured data based on pre-set analysis and extraction strategies; Extract unstructured feature vectors from unstructured data based on pre-set mining and extraction strategies; The structured and unstructured feature vectors are standardized and fused to generate data features.

3. The method for managing sensitive Internet asset data based on a digital model according to claim 2, characterized in that: The step of extracting structured feature vectors from structured data based on a pre-set analysis and extraction strategy includes the following steps: Establish a connection graph with structured data as data nodes and analyze the reference relationships of structured data to quantify the topological importance index of data nodes; Extract constraint rule features between structured data; The connectivity graph, topological importance index, and constraint rule features are vectorized and encoded, and structured feature vectors are generated through feature dimensionality reduction.

4. The method for managing sensitive Internet asset data based on a digital model according to claim 2, characterized in that: The step of extracting essential features of unstructured data based on a pre-set mining and extraction strategy includes the following steps: Content slicing and pattern analysis are performed on unstructured data to identify recurring feature fragments and key content patterns in the unstructured data; Construct a semantic association network among unstructured data; Vectorize repetitive feature fragments, key content patterns, and semantic association networks; Unstructured feature vectors are generated through feature selection.

5. The method for managing sensitive Internet asset data based on a digital model according to claim 1, characterized in that: The step of acquiring the spatiotemporal feature information associated with the target sensitive data, and constructing a digital twin model of the target sensitive data based on the spatiotemporal feature information and the digital fingerprint, includes the following steps: Identify the data storage nodes of the target sensitive data, and obtain the timestamps and physical location identifiers from the data storage nodes as basic spatiotemporal features; Identify the data transmission path of the target sensitive data, and collect network device information in the data transmission path through the acquisition terminal as supplementary spatiotemporal features; Spatiotemporal feature information is generated based on basic and supplementary spatiotemporal features, and the spatiotemporal feature information is then linked and bound to digital fingerprints in multiple dimensions. A digital twin model containing data features, digital fingerprints, and spatiotemporal features is constructed based on the multi-dimensional association and binding results.

6. The method for managing sensitive Internet asset data based on a digital model according to claim 5, characterized in that: The step of generating spatiotemporal feature information based on basic spatiotemporal features and supplementary spatiotemporal features, and then associating and binding the spatiotemporal feature information with digital fingerprints in multiple dimensions, includes the following steps: Establish temporal, spatial, and transmission associations between digital fingerprint features and timestamps, physical locations, and network paths, respectively; Cross-validation of temporal correlation, spatial correlation, and transport correlation was performed. Generate associated credentials containing digital fingerprints and spatiotemporal feature information using a pre-set encryption algorithm.

7. The method for managing sensitive Internet asset data based on a digital model according to claim 6, characterized in that: The step of generating an associated credential containing digital fingerprints and spatiotemporal feature information using a pre-set encryption algorithm includes the following steps: A layered encryption architecture is adopted to perform layered joint processing of digital fingerprints and spatiotemporal feature information; The digital fingerprint is encrypted using a pre-set asymmetric encryption algorithm; Spatiotemporal feature information is encrypted using a pre-set lightweight encryption algorithm; The encrypted digital fingerprint and spatiotemporal feature information are combined and encapsulated to generate associated credentials.

8. A digital model-based Internet asset sensitive data management system, characterized in that: include: The type recognition module is used to acquire target sensitive data and perform type recognition on the target sensitive data; The feature extraction module is used to encode the target sensitive data based on the type recognition results, extract data features and generate digital fingerprints; The model building module is used to acquire spatiotemporal feature information associated with target sensitive data, and to build a digital twin model of the target sensitive data based on the spatiotemporal feature information and digital fingerprint. The real-time monitoring module is used to monitor control behavior in real time based on the constructed digital twin model and to verify data integrity based on digital fingerprints. The protection module is used to trigger the corresponding protection policy when abnormal control behavior is detected. The audit log generation module is used to obtain the data flow path based on spatiotemporal feature information and generate security audit logs.

Citation Information

Patent Citations

  • Data security management method and system and storage medium

    CN117574458A

  • Enterprise sensitive data security access management method and system

    CN118656870A

  • Multi-service scheduling optimization algorithm for infrastructure safety management

    CN119620717A

  • Safety management method and system for data asset transaction

    CN120106841A