Data asset automatic classification method and system based on data analysis
Patent Information
- Application Number
- CN202611230300.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-14
- Publication Date
- 2026-09-25
AI Technical Summary
本发明解决了现有技术存在的静态特征易被伪装绕过、多模态特征融合效果差以及模型闭环自进化能力弱的问题
[0014]通过安全探针不仅提取静态语义特征,还针对具有伪装或动态执行特征的文件利用轻量级沙箱提取动态衍生特征,并计算动态行为异常度评估值,有效应对了传统静态规则容易被绕过的问题,提升了对恶意文件及敏感动态数据的识别精度;提出基于门控交叉注意力机制的多模态特征自适应融合模型,通过交叉注意力提取动态特征与静态文本语义高度关联的部分,并在动态衍生特征为零矩阵时自动退化为纯静态特征提取,既保证了多模态信息的深度融合,又避免了无效特征对模型的干扰,提高了特征表达的鲁棒性;设计了双流分类预测与加权聚类模型,监督流负责输出具体标签,非监督流通过引入动态行为异常度作为惩罚权重的自适应距离加权聚类机制发现未知数据资产,实现了已知分类与未知数据资产发现的协同处理;引入了带时间衰减弹性遗忘因子的对比学习主动闭环微调机制,通过动态调节间隔并结合对比损失、分类损失及参数正则化,使得模型在持续学习新样本的同时有效缓解灾难性遗忘,显著提升了模型的自进化能力和长期的分类准确率。
Smart Images

Figure CN122818031A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data security management technology, and in particular to a method and system for automatic classification of data assets based on data analysis. Background Technology
[0002] With the advent of the big data era, enterprise data assets are characterized by massive volume, heterogeneity, and dynamic changes. Accurate and efficient classification and grading of data assets is fundamental to data security governance. Existing data asset classification methods largely rely on static rule matching or single machine learning models. However, this approach has several limitations:
[0003] 1) Static features are easily disguised and bypassed: Relying solely on static feature extraction cannot effectively deal with malicious files and sensitive data streams that are disguised or have dynamic execution capabilities, which can easily lead to false negatives or false positives.
[0004] 2) Poor multimodal feature fusion: When fusing data features from different modalities such as text and images, simple splicing is often used, ignoring the semantic relationships between modalities, which leads to limited classification accuracy;
[0005] 3) Weak self-evolution ability of closed-loop model: Most existing classification techniques are open-loop techniques. Once deployed, the model parameters are fixed. Faced with constantly evolving unknown data assets and misclassified samples, they lack the ability to actively learn and fine-tune in a closed loop, leading to model aging and a decline in classification accuracy.
[0006] Therefore, there is an urgent need for an automatic data asset classification method and system that can adaptively fuse dynamic and static features, possess dual-stream classification prediction and unknown asset identification capabilities, and support closed-loop fine-tuning. Summary of the Invention
[0007] This invention provides a method and system for automatic classification of data assets based on data analysis. This invention addresses the problems of existing technologies, such as the susceptibility of static features to spoofing, poor multimodal feature fusion performance, and weak self-evolutionary ability of closed-loop models.
[0008] In a first aspect, embodiments of the present invention provide an automatic data asset classification method based on data analysis, the method comprising:
[0009] Security probes are used to collect data streams of data assets, extract static semantic feature matrices of static files and dynamic derived feature matrices of disguised or dynamic files from the data streams, and calculate dynamic behavior anomaly assessment values.
[0010] Based on the static semantic feature matrix and the dynamic derived feature matrix, a multimodal feature adaptive fusion model based on a gated cross-attention mechanism is used to generate a dimensionality-reduced fusion feature vector.
[0011] Based on the dynamic behavior anomaly assessment value and the dimensionality reduction fusion feature vector, a dual-stream classification prediction and weighted clustering model is used to generate classification prediction results and unknown asset determination results, and to collect a set of error correction samples that are misjudged and / or need to be reclassified.
[0012] Based on the error correction sample set and the historical sample set, active closed-loop fine-tuning of the multimodal feature adaptive fusion model and the two-stream classification prediction and weighted clustering model is carried out based on contrastive learning with forgetting factor.
[0013] The technical solution provided in this application has at least the following beneficial effects:
[0014] This approach not only extracts static semantic features using security probes, but also utilizes a lightweight sandbox to extract dynamic derived features from files with camouflaged or dynamically executed characteristics. It also calculates anomaly assessment values for dynamic behavior, effectively addressing the issue of traditional static rules being easily bypassed and improving the accuracy of identifying malicious files and sensitive dynamic data. Furthermore, it proposes a multimodal feature adaptive fusion model based on a gated cross-attention mechanism. This model extracts highly semantically related parts of dynamic features from static text through cross-attention, and automatically degenerates to pure static feature extraction when the dynamic derived feature matrix is zero. This ensures deep fusion of multimodal information while avoiding the impact of invalid features on the model. To mitigate interference and improve the robustness of feature representation, a dual-stream classification prediction and weighted clustering model was designed. The supervised stream is responsible for outputting specific labels, while the unsupervised stream discovers unknown data assets through an adaptive distance-weighted clustering mechanism that introduces dynamic behavioral anomaly as a penalty weight, achieving collaborative processing of known classifications and the discovery of unknown data assets. A contrastive learning active closed-loop fine-tuning mechanism with a time-decaying elastic forgetting factor was introduced. By dynamically adjusting the interval and combining contrastive loss, classification loss, and parameter regularization, the model can effectively mitigate catastrophic forgetting while continuously learning new samples, significantly improving the model's self-evolution ability and long-term classification accuracy.
[0015] In one alternative implementation, a security probe is used to collect the data stream of the data asset, extract the static semantic feature matrix of static files and the dynamic derived feature matrix of disguised or dynamic files from the data stream, and calculate the dynamic behavior anomaly assessment value, including:
[0016] Security probes are used to collect data streams from data assets, and stream reassembly and application-layer parsing are performed on the data streams to extract and restore independent file objects.
[0017] Static parsing is performed on static files in the file object. If the file object is a text and / or office document, the text in the file object is extracted and input into the pre-trained language model to obtain the corresponding static semantic feature matrix. If the file object is an image file, the text in the file object is extracted by OCR and then converted into a static semantic feature matrix.
[0018] Extract file metadata features from file objects and input them into a lightweight decision tree to determine whether they have spoofing or dynamic execution characteristics. If they do, the file objects are automatically isolated into a lightweight sandbox for controlled pre-execution.
[0019] During the lightweight sandbox pre-execution, capture the dynamic behavior sequence generated by the dynamic execution of disguised or dynamic files in the file object, and vectorize the dynamic behavior sequence to obtain the dynamic derived feature matrix;
[0020] If the file metadata features do not have spoofing or dynamic execution features, the output dynamic derived feature matrix is a zero matrix;
[0021] The dynamic behavior anomaly assessment value is calculated based on the degree to which the dynamic derived feature matrix deviates from the normal baseline; wherein, the normal baseline is constructed based on the statistical analysis of the historical dynamic behavior sequence vectors of historical normal files in the lightweight sandbox.
[0022] In one alternative implementation, a dynamic behavior anomaly assessment value is calculated based on the degree to which the dynamically derived feature matrix deviates from the normal baseline, including:
[0023] Determine whether the dynamic derived feature matrix of the file object is a zero matrix. If it is a zero matrix, output the dynamic behavior anomaly evaluation value as 0 directly. If it is not a zero matrix, proceed to the next step.
[0024] The set of historical dynamic behavior feature vectors of historical normal files in the lightweight sandbox is used to form a normal baseline, and its mean vector and covariance matrix are statistically analyzed.
[0025] In the case that the file metadata features have spoofing or dynamic execution features, calculate the sequence mean vector of the dynamic derived feature matrix of the spoofed or dynamic file in the current file object;
[0026] Based on the mean vector and covariance matrix of the normal baseline, and the sequence mean vector of the dynamically derived feature matrix, the Mahalanobis distance between the dynamically derived feature matrix and the normal baseline is calculated, which is used as the degree to which the dynamically derived feature matrix deviates from the normal baseline.
[0027] The Mahalanobis distance is normalized using the Sigmoid function to obtain the dynamic behavior anomaly assessment value;
[0028] If the file object does not have masquerading or dynamic execution characteristics, the output dynamic behavior anomaly evaluation value is 0.
[0029] In one optional implementation, the multimodal feature adaptive fusion model based on the gated cross-attention mechanism includes a cross-attention calculation module, a dynamic gating coefficient calculation module, and a fusion feature generation module.
[0030] In one alternative implementation, based on the static semantic feature matrix and the dynamically derived feature matrix, a multimodal feature adaptive fusion model constructed using a gated cross-attention mechanism is used to generate a dimensionality-reduced fusion feature vector, including:
[0031] The static semantic feature matrix is mapped to the query matrix, and the dynamic derived feature matrix is mapped to the key matrix and the value matrix, respectively.
[0032] Based on the query matrix, key matrix, and value matrix, the cross-attention calculation module of the multimodal feature adaptive fusion model built on a gated cross-attention mechanism is used to extract the part of the dynamic derived feature matrix that is highly related to the semantics of the static text and generate the attention feature matrix.
[0033] Determine whether the dynamic derived feature matrix is a zero matrix. If it is a zero matrix, set the dynamic gate coefficient to 0 so that the multimodal feature adaptive fusion model automatically degenerates into pure static feature extraction. If it is not a zero matrix, proceed to the next step.
[0034] The static semantic feature matrix and the dynamic derived feature matrix are input into the fully connected layer of the dynamic gating coefficient calculation module of the multimodal feature adaptive fusion model to generate dynamic gating coefficients.
[0035] Based on dynamic gating coefficients, the attention feature matrix and static semantic feature matrix are input into the fusion feature generation module of the multimodal feature adaptive fusion model to generate a dimensionality-reduced fusion feature vector.
[0036] In one optional implementation, the dual-stream classification prediction and weighted clustering model includes a supervised stream classification prediction branch and an unsupervised stream weighted clustering branch;
[0037] The supervised flow classification prediction branch is equipped with a DNN classifier;
[0038] The unsupervised flow weighted clustering branch is equipped with an adaptive distance weighted clustering program, and the unsupervised flow weighted clustering branch is connected to a known data asset library, which contains several known data asset clusters.
[0039] In one optional implementation, based on the dynamic behavior anomaly assessment value and the dimensionality-reduced fused feature vector, a two-stream classification prediction and weighted clustering model is used to generate classification prediction results and unknown asset determination results, and to collect a set of error-corrected samples that are misclassified and / or require reclassification, including:
[0040] The dimensionality-reduced fused feature vector is input into the two-stream classification prediction and weighted clustering model. The DNN classifier of the supervised stream classification prediction branch is used to calculate the probability distribution of each category of the dimensionality-reduced fused feature vector, output the classification label and sensitivity level of the current data asset, and generate the classification prediction result.
[0041] Based on the dynamic behavior anomaly evaluation value and the dimensionality reduction fusion feature vector, the unsupervised flow weighted clustering branch is used to call the adaptive distance weighted clustering program to calculate the comprehensive weighted distance between the dimensionality reduction fusion feature vector and the cluster centers of several known data asset clusters in the known data asset library, and obtain the minimum comprehensive weighted distance among all comprehensive weighted distances;
[0042] If the minimum weighted average distance is greater than the distance threshold, the dimensionality reduction and fusion feature vector of the current file object is divided into unknown asset clusters, and all unknown asset clusters are aggregated into an active learning request sample set.
[0043] The current dimensionality reduction and fusion feature vector with classification prediction results and the active learning request sample set are used to generate a visual data asset map on the front end. The error correction sample set obtained by the operation and maintenance personnel in judging the data asset map and / or needing to be reclassified is also collected.
[0044] In one alternative implementation, based on the error correction sample set and the historical sample set, the multimodal feature adaptive fusion model and the two-stream classification prediction and weighted clustering model are actively fine-tuned using contrastive learning with a forgetting factor, including:
[0045] The error correction samples in the error correction sample set are used as anchor samples. Several historical samples of the same category as the anchor samples are retrieved from the historical sample set in the known data asset library as positive samples. Several historical samples whose true category is different from the anchor samples but whose features are similar are retrieved as negative samples.
[0046] Calculate the time decay elastic forgetting factor, and dynamically adjust the interval of contrastive learning based on the time decay elastic forgetting factor to obtain the adaptive interval for the current batch.
[0047] Using a multimodal feature adaptive fusion model, sample dimensionality reduction fusion feature vectors are generated for each anchor sample, each positive sample, and each negative sample. The sample dimensionality reduction fusion feature vectors and the adaptive interval are then substituted into the joint contrastive learning loss function with parameter regularization to obtain the comprehensive loss value.
[0048] The gradient vector is obtained by taking the partial derivative of the comprehensive loss value with respect to the model parameters. Then, the gradient descent algorithm is used to perform gradient descent with a set learning rate to complete the synchronous iterative update of the parameters and obtain the updated model parameters.
[0049] Based on the updated model parameters, the multimodal feature adaptive fusion model and the two-stream classification prediction and weighted clustering model are fine-tuned to obtain the updated multimodal feature adaptive fusion model and the updated two-stream classification prediction and weighted clustering model, which are used for the automatic classification of data assets in the next batch.
[0050] In one alternative implementation, the model parameters include a first model weight matrix and a first model bias vector for the fully connected layer, and a second model weight matrix and a second model bias vector for the DNN classifier.
[0051] Secondly, embodiments of the present invention provide an automatic data asset classification system based on data analysis, used to implement an automatic data asset classification method, the system comprising:
[0052] The data acquisition and feature extraction unit is used to collect data streams of data assets using security probes, extract static semantic feature matrices of static files and dynamic derived feature matrices of disguised or dynamic files in the data stream, and calculate dynamic behavior anomaly evaluation values.
[0053] The multimodal feature fusion unit is used to generate a dimensionality-reduced fusion feature vector based on a multimodal feature adaptive fusion model constructed using a gated cross-attention mechanism, according to the static semantic feature matrix and the dynamic derived feature matrix.
[0054] The dual-stream classification and clustering unit is used to generate classification prediction results and unknown asset determination results based on dynamic behavior anomaly assessment values and dimensionality-reduced fusion feature vectors using a dual-stream classification prediction and weighted clustering model, and to collect error correction sample sets of misjudged and / or reclassified cases.
[0055] The closed-loop fine-tuning unit is used to perform active closed-loop fine-tuning of the multimodal feature adaptive fusion model and the two-stream classification prediction and weighted clustering model based on contrastive learning with a forgetting factor, according to the error correction sample set and the historical sample set.
[0056] A third aspect of this invention provides an electronic device, which includes:
[0057] At least one processor; and a memory communicatively connected to the at least one processor; wherein,
[0058] The memory stores instructions that can be executed by at least one processor, such that the at least one processor can perform the method proposed in the first aspect of the present invention.
[0059] A fourth aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method as described in the first aspect of the present invention. Attached Figure Description
[0060] Figure 1 This is a schematic diagram of the electronic device structure of the hardware operating environment involved in the embodiments of the present invention;
[0061] Figure 2 This is a flowchart illustrating the steps of an automatic data asset classification method based on data analysis provided in an embodiment of the present invention.
[0062] Figure 3 This is a schematic diagram of the functional units of an automatic data asset classification system based on data analysis provided in an embodiment of the present invention. Detailed Implementation
[0063] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0064] The present invention will be further described below with reference to the accompanying drawings.
[0065] Reference Figure 1 , Figure 1 This is a schematic diagram of the electronic device structure of the hardware operating environment involved in the embodiments of the present invention.
[0066] like Figure 1 As shown, the electronic device may include: a processor 1001, such as a central processing unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. The communication bus 1002 is used to enable communication between these components. The user interface 1003 may include a display screen or an input unit such as a keyboard; optionally, the user interface 1003 may also include a standard wired interface or a wireless interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface). The memory 1005 may be a high-speed random access memory (RAM) or a stable non-volatile memory (NVM), such as a disk drive. The memory 1005 may also optionally be a storage device independent of the aforementioned processor 1001.
[0067] Those skilled in the art will understand that Figure 1The structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0068] like Figure 1 As shown, the memory 1005, which serves as a storage medium, may include an operating system, a data storage module, a network communication module, a user interface module, and an electronic program for an automatic data asset classification system based on data analysis.
[0069] exist Figure 1 In the electronic device shown, the network interface 1004 is mainly used for data communication with the network server; the user interface 1003 is mainly used for data interaction with the user; the processor 1001 and the memory 1005 in the electronic device of the present invention can be set in the electronic device. The electronic device calls the electronic program of the data asset automatic classification system based on data analysis stored in the memory 1005 through the processor 1001, and executes the data asset automatic classification method based on data analysis provided in the embodiment of the present invention.
[0070] Reference Figure 2 The present invention provides an automatic data asset classification method based on data analysis, the method comprising:
[0071] S201: Use security probes to collect data streams of data assets, extract static semantic feature matrices of static files and dynamic derived feature matrices of disguised or dynamic files in the data stream, and calculate dynamic behavior anomaly assessment values.
[0072] S202: Based on the static semantic feature matrix and the dynamic derived feature matrix, a multimodal feature adaptive fusion model based on a gated cross-attention mechanism is used to generate a dimension-reduced fusion feature vector;
[0073] S203: Based on the dynamic behavior anomaly assessment value and the dimensionality reduction fusion feature vector, use a dual-stream classification prediction and weighted clustering model to generate classification prediction results and unknown asset determination results, and collect a set of error correction samples that are misjudged and / or need to be reclassified.
[0074] S204: Based on the error correction sample set and the historical sample set, perform active closed-loop fine-tuning of the multimodal feature adaptive fusion model and the dual-stream classification prediction and weighted clustering model based on contrastive learning with forgetting factor.
[0075] The technical solution provided in this application has at least the following beneficial effects:
[0076] This approach not only extracts static semantic features using security probes, but also utilizes a lightweight sandbox to extract dynamic derived features from files with camouflaged or dynamically executed characteristics. It also calculates anomaly assessment values for dynamic behavior, effectively addressing the issue of traditional static rules being easily bypassed and improving the accuracy of identifying malicious files and sensitive dynamic data. Furthermore, it proposes a multimodal feature adaptive fusion model based on a gated cross-attention mechanism. This model extracts highly semantically related parts of dynamic features from static text through cross-attention, and automatically degenerates to pure static feature extraction when the dynamic derived feature matrix is zero. This ensures deep fusion of multimodal information while avoiding the impact of invalid features on the model. To mitigate interference and improve the robustness of feature representation, a dual-stream classification prediction and weighted clustering model was designed. The supervised stream is responsible for outputting specific labels, while the unsupervised stream discovers unknown data assets through an adaptive distance-weighted clustering mechanism that introduces dynamic behavioral anomaly as a penalty weight, achieving collaborative processing of known classifications and the discovery of unknown data assets. A contrastive learning active closed-loop fine-tuning mechanism with a time-decaying elastic forgetting factor was introduced. By dynamically adjusting the interval and combining contrastive loss, classification loss, and parameter regularization, the model can effectively mitigate catastrophic forgetting while continuously learning new samples, significantly improving the model's self-evolution ability and long-term classification accuracy.
[0077] In one alternative implementation, a security probe is used to collect the data stream of the data asset, extract the static semantic feature matrix of static files and the dynamic derived feature matrix of disguised or dynamic files from the data stream, and calculate the dynamic behavior anomaly assessment value, including:
[0078] S2011: Use security probes to collect data streams of data assets, and perform stream reassembly and application-layer parsing on the data streams to extract and restore independent file objects;
[0079] In this embodiment, the security probe can be deployed at the enterprise's core switch or data gateway in a bypass mirroring or serial access manner to collect data streams of data assets;
[0080] Stream reconstruction and application layer parsing: After the probe captures network packets at the data link layer, it performs state tracking and reconstruction of the Transmission Control Protocol Stream (TCP) based on the five-tuple (source IP, destination IP, source port, destination port, protocol) to restore the complete application layer session such as Hypertext Transfer Protocol (HTTP), File Transfer Protocol (FTP), or Server Message Block (SMB). When stripping file objects, it not only extracts the file content but also retains its context metadata (such as Uniform Resource Locator (URL) path, request headers, transmission time, etc.).
[0081] S2012: Static parsing is performed on static files in the file object. If the file object is a text and / or office document, the text in the file object is extracted and input into the pre-trained language model to obtain the corresponding static semantic feature matrix. If the file object is an image file, the text in the file object is extracted by Optical Character Recognition (OCR) and then converted into a static semantic feature matrix.
[0082] In this embodiment, static semantic feature extraction is performed as follows: For text or office documents, encoding conversion and noise removal are first performed, and then the data is input into a lightweight pre-trained language model (such as lightweight Transformer-based bidirectional encoder representations from Transformers (ALBERT)). The hidden layer output of the CLS token is extracted as a static semantic feature matrix, and the dimension of this matrix can be set to [1, 768]. For image files, the same process is used after extracting the text using Tesseract or PaddleOCR. If there is no valid text in the image, the visual feature vector extracted by the lightweight residual network (ResNet) is mapped to the same dimensional space as the text features.
[0083] S2013: Extract the file metadata features of the file object and input them into the lightweight decision tree to determine whether it has spoofing or dynamic execution features. If it does, the file object is automatically isolated into the lightweight sandbox for controlled pre-execution.
[0084] In this embodiment, a lightweight sandbox and dynamic behavior sequence vectorization are implemented: a lightweight decision tree performs spoofing judgment based on file entropy, file header features, and extension consistency; the sandbox can be built using a lightweight container (such as Docker + seccomp) to limit CPU and memory usage, and deploys an Application Programming Interface Hook (API Hook) tool within it to capture file operation sequences to the registry, file system, network, and processes; when vectorizing dynamic behavior sequences, a pre-trained behavior embedding model can be used to transform discrete operation sequences into a continuous dynamic derived feature matrix, with the matrix dimension set to [l, 128], where l is the number of captured behaviors;
[0085] S2014: During the pre-execution of the lightweight sandbox, capture the dynamic behavior sequence generated by the dynamic execution of the disguised or dynamic file in the file object, and vectorize the dynamic behavior sequence to obtain the dynamic derived feature matrix.
[0086] S2015: If the file metadata features do not have masquerading or dynamic execution features, then the output dynamic derived feature matrix is a zero matrix;
[0087] S2016: Calculate the dynamic behavior anomaly assessment value based on the degree to which the dynamic derived feature matrix deviates from the normal baseline; wherein, the normal baseline is constructed based on the statistical analysis of historical dynamic behavior sequence vectors of historical normal files in the lightweight sandbox.
[0088] In one alternative implementation, a dynamic behavior anomaly assessment value is calculated based on the degree to which the dynamically derived feature matrix deviates from the normal baseline, including:
[0089] S20161: Determine whether the dynamic derived feature matrix of the file object is a zero matrix. If it is a zero matrix, output the dynamic behavior anomaly evaluation value as 0 directly. If it is not a zero matrix, proceed to the next step.
[0090] S20162: Construct a normal baseline by taking the set of historical dynamic behavior feature vectors of historical normal files in a lightweight sandbox and calculating their mean vector and covariance matrix.
[0091] S20163: When the file metadata features have camouflage or dynamic execution features, calculate the sequence mean vector of the dynamic derived feature matrix of the camouflage or dynamic file in the current file object, using the following formula:
[0092]
[0093] In the formula, It is the vector of the sequence mean; For dynamically derived feature matrices In the sequence, i is the sequence index; m is the total number of sequences; if the file object does not have masquerading or dynamic execution characteristics, then output ;
[0094] S20164: Based on the mean vector and covariance matrix of the normal baseline, and the sequence mean vector of the dynamically derived feature matrix, calculate the Mahalanobis distance between the dynamically derived feature matrix and the normal baseline. This distance serves as the degree to which the dynamically derived feature matrix deviates from the normal baseline. The formula is as follows:
[0095]
[0096] In the formula, The Mahalanobis distance between the dynamically derived feature matrix and the normal baseline; It is the vector of the sequence mean; This is the mean vector of the normal baseline; The covariance matrix of the normal baseline; The symbol for the inverse matrix; It is the transpose symbol;
[0097] S20165: Normalize the Mahalanobis distance using the Sigmoid function to obtain the dynamic behavior anomaly assessment value, the formula is as follows:
[0098]
[0099] In the formula, This is a dynamic behavioral anomaly assessment value; The Mahalanobis distance between the dynamically derived feature matrix and the normal baseline; It is a natural constant; It is a bias constant used to control the translation of the Sigmoid curve on the X-axis (i.e., to determine the threshold benchmark for anomaly determination).
[0100] S20167: If the file object does not have the characteristics of masquerading or dynamic execution, the output dynamic behavior anomaly evaluation value is 0.
[0101] In one optional implementation, the multimodal feature adaptive fusion model based on the gated cross-attention mechanism includes a cross-attention calculation module, a dynamic gating coefficient calculation module, and a fusion feature generation module.
[0102] In one alternative implementation, based on the static semantic feature matrix and the dynamically derived feature matrix, a multimodal feature adaptive fusion model constructed using a gated cross-attention mechanism is used to generate a dimensionality-reduced fusion feature vector, including:
[0103] S2021: Map the static semantic feature matrix to the query matrix, and map the dynamic derived feature matrix to the key matrix and the value matrix respectively;
[0104] S2022: Based on the query matrix, key matrix, and value matrix, the cross-attention calculation module of the multimodal feature adaptive fusion model built on a gated cross-attention mechanism is used to extract the part of the dynamic derived feature matrix that is highly correlated with the semantics of the static text, generating an attention feature matrix. The formula is as follows:
[0105]
[0106] In the formula, This is the attention feature matrix; It is a normalized exponential function; These are the query matrix, the key matrix, and the value matrix, respectively. It is the transpose symbol; This is the scaling factor; For feature dimensions;
[0107] S2023: Determine whether the dynamic derived feature matrix is a zero matrix. If it is a zero matrix, set the dynamic gate coefficient to 0 so that the multimodal feature adaptive fusion model automatically degenerates into pure static feature extraction. If it is not a zero matrix, proceed to the next step.
[0108] In this embodiment, when it is determined that there is no disguise or dynamic execution feature, a zero matrix with dimensions [1, 128] is generated as a dynamic derived feature matrix and input to the linear mapping layer, thereby ensuring that the tensor dimension does not cause tensor shape mismatch anomaly during the calculation process;
[0109] S2024: Input the static semantic feature matrix and the dynamic derived feature matrix into the fully connected layer of the dynamic gating coefficient calculation module of the multimodal feature adaptive fusion model to generate dynamic gating coefficients, the formula of which is:
[0110]
[0111] In the formula, For dynamic gating coefficients; It is a static semantic feature matrix; For dynamically derived feature matrices; Symbols for feature splicing operations; This is the weight matrix of the first model; This is the bias vector for the first model; The activation function is Sigmoid; if the dynamically derived feature matrix is a zero matrix, When the value approaches 0, it automatically degenerates into pure static feature extraction;
[0112] S2025: Based on dynamic gating coefficients, the attention feature matrix and static semantic feature matrix are input into the fusion feature generation module of the multimodal feature adaptive fusion model to generate a dimensionality-reduced fusion feature vector, as shown in the formula:
[0113]
[0114] In the formula, To reduce the dimensionality and fuse the feature vectors; This is the attention feature matrix; For dynamic gating coefficients; It is a static semantic feature matrix; The symbol for element-wise multiplication.
[0115] In one optional implementation, the dual-stream classification prediction and weighted clustering model includes a supervised stream classification prediction branch and an unsupervised stream weighted clustering branch;
[0116] The supervised flow classification prediction branch is equipped with a deep neural network (DNN) classifier;
[0117] The unsupervised flow weighted clustering branch is equipped with an adaptive distance weighted clustering program, and the unsupervised flow weighted clustering branch is connected to a known data asset library, which contains several known data asset clusters.
[0118] In one optional implementation, based on the dynamic behavior anomaly assessment value and the dimensionality-reduced fused feature vector, a two-stream classification prediction and weighted clustering model is used to generate classification prediction results and unknown asset determination results, and to collect a set of error-corrected samples that are misclassified and / or require reclassification, including:
[0119] S2031: Input the dimensionality-reduced fused feature vector into the two-stream classification prediction and weighted clustering model. Use the DNN classifier of the supervised stream classification prediction branch to calculate the probability distribution of each category of the dimensionality-reduced fused feature vector, output the classification label and sensitivity level of the current data asset, and generate the classification prediction result. The formula is:
[0120]
[0121] In the formula, For classification prediction results; It is a normalized exponential function; This is the weight matrix for the second model; This is the bias vector for the second model; To reduce the dimension and fuse the feature vectors;
[0122] S2032: Based on the dynamic behavior anomaly assessment value and the dimensionality-reduced fusion feature vector, using the unsupervised flow-weighted clustering branch, the adaptive distance-weighted clustering program is called to calculate the comprehensive weighted distance between the dimensionality-reduced fusion feature vector and the cluster centers of several known data asset clusters in the known data asset database. The minimum comprehensive weighted distance among all comprehensive weighted distances is obtained, and the formula is:
[0123]
[0124] In the formula, Let be the weighted distance from the j-th file object to the c-th cluster center; j is the file object index; c is the cluster center index; The dimension-reduced fusion feature vector of the j-th file object; For the j-th file object; Let c be the eigenvector of the c-th cluster center; Spatial distance weights; The weight for dynamic anomaly penalty; It is an L2 norm; The dynamic behavior anomaly evaluation value of the dynamic derived feature matrix of the j-th file object;
[0125] S2033: If the minimum comprehensive weighted distance is greater than the distance threshold, the dimensionality reduction fusion feature vector of the current file object is divided into unknown asset clusters, and all unknown asset clusters are aggregated into an active learning request sample set;
[0126] In this embodiment, the distance threshold is not a fixed constant, but is dynamically calculated based on the historical distribution of the known data asset database. Specifically, the strategy is to periodically (e.g., every 24 hours) statistically analyze the comprehensive weighted distance distribution from samples within all asset clusters in the known data asset database to their cluster centers; calculate the 95th percentile of this distribution as the basic threshold; and in actual determination, the formula for the distance threshold is:
[0127]
[0128] In the formula, Distance threshold; Basic threshold; This is the tolerance factor (which can be set between 1.1 and 1.3).
[0129] When the minimum comprehensive weighted distance is greater than When the cluster is identified as an unknown asset cluster, this dynamic threshold setting can effectively prevent the cluster boundaries from becoming rigid as the known asset library expands, and improve the sensitivity to new forms of data assets.
[0130] S2034: Generate a visualized data asset map on the front end using the current dimensionality reduction fusion feature vector with classification prediction results and the active learning request sample set, and collect the error correction sample set obtained by the operation and maintenance personnel in judging the data asset map and / or needing to be reclassified.
[0131] In this embodiment, the front-end visualized data asset map can be rendered using a force-directed graph algorithm. The node size represents the number of assets, the node color represents the category label, and the unknown asset clusters and manually corrected sample points are highlighted to guide maintenance personnel to focus on them.
[0132] In one alternative implementation, based on the error correction sample set and the historical sample set, the multimodal feature adaptive fusion model and the two-stream classification prediction and weighted clustering model are actively fine-tuned using contrastive learning with a forgetting factor, including:
[0133] S2041: Use the error correction samples in the error correction sample set as anchor samples, retrieve several historical samples of the same category as the anchor samples from the historical sample set in the known data asset library as positive samples, and retrieve several historical samples whose true category is different from the anchor samples but whose features are similar and are prone to misjudgment as negative samples.
[0134] In this embodiment, a cache queue and a cold start mechanism are introduced during the closed-loop fine-tuning process:
[0135] When operations and maintenance personnel mark misjudged and / or reclassified assets on the front-end map, the generated error correction sample set will be pushed into the fine-tuning cache queue; set trigger conditions: when the accumulation of new error correction samples in the cache queue reaches the set batch size (e.g., Batch Size = 32), or when the time since the last fine-tuning exceeds the set period (e.g., 72 hours), active closed-loop fine-tuning will be triggered.
[0136] When retrieving positive and negative samples, the K-nearest neighbor algorithm based on the Euclidean distance of the dimensionality-reduced fusion feature vector is adopted. For example, positive samples are selected from the closest historical samples of the same class, and negative samples are selected from the closest historical samples of different classes, so as to form high-quality contrastive learning triples.
[0137] S2042: Calculate the time-decay elastic forgetting factor, and dynamically adjust the contrastive learning interval based on the time-decay elastic forgetting factor to obtain the adaptive interval for the current batch. The formula is:
[0138]
[0139] In the formula, The time-decaying elastic forgetting factor; This is the proportionality coefficient; It is a natural constant; The decay rate over time; The time interval between the current incremental adjustment batch and the previous batch; The basic interval; For adaptive interval; This is the interval adjustment coefficient; Base retention rate (can be set to 0.8);
[0140] S2043: Using a multimodal feature adaptive fusion model, a sample dimensionality reduction and fusion feature vector is generated for each anchor sample, each positive sample, and each negative sample. The sample dimensionality reduction and fusion feature vector and the adaptive margin are then substituted into a joint contrastive learning loss function with parameter regularization to obtain the comprehensive loss value, as shown in the formula:
[0141]
[0142] In the formula, This is the overall loss value; For the maximum value function, when the calculation result within the parentheses is less than or equal to 0, it means that the distance between positive and negative samples has met the set interval requirement. At this time, the loss is recorded as 0 and no gradient is generated. When the result is greater than 0, it means that the distance does not meet the requirement, and this value is taken as the penalty loss for backpropagation. To compare the loss values; This is the classification loss value; The classification weight hyperparameter; This is the regularization weight hyperparameter; , Used to balance the proportions of "contrast learning term", "classification term" and "parameter penalty term" in the total loss; The time-decaying elastic forgetting factor; The model parameters for the current batch of multimodal feature adaptive fusion model and two-stream classification prediction and weighted clustering model; These are the model parameters for the previous batch of multimodal feature adaptive fusion models and two-stream classification prediction and weighted clustering models;
[0143]
[0144] In the formula, This is a distance metric function, usually referring to cosine distance or Euclidean distance, used to measure how close two feature vectors are in a multidimensional space; For dimension reduction and fusion feature mapping function; For anchor point samples, positive samples, and negative samples; For adaptive interval;
[0145]
[0146] In the formula, For anchor point samples In the True labels in each category are typically encoded using one-hot encoding; Indexed by category; Model parameters for the current batch Below, the DNN classifier predicts anchor samples. In the The probability of each category is usually output by the Softmax layer and takes values between (0,1);
[0147] S2044: Take the partial derivative of the comprehensive loss value with respect to the model parameters to obtain the gradient vector, and then use the stochastic gradient descent algorithm with the set learning rate to perform gradient descent, completing the synchronous iterative update of the parameters and obtaining the updated model parameters. The formula is as follows:
[0148]
[0149] In the formula, For updated model parameters; These are the model parameters for the current batch; The learning rate; The gradient vector;
[0150] S2045: Based on the updated model parameters, fine-tune the multimodal feature adaptive fusion model and the two-stream classification prediction and weighted clustering model to obtain the updated multimodal feature adaptive fusion model and the updated two-stream classification prediction and weighted clustering model, which will be used for the automatic classification of data assets in the next batch.
[0151] In one alternative implementation, the model parameters include a first model weight matrix and a first model bias vector for the fully connected layer, and a second model weight matrix and a second model bias vector for the DNN classifier.
[0152] This invention also provides an automatic data asset classification system 300 based on data analysis, referring to... Figure 3 The system may include the following units:
[0153] The data acquisition and feature extraction unit 301 is used to acquire the data stream of data assets using security probes, extract the static semantic feature matrix of static files and the dynamic derived feature matrix of disguised or dynamic files in the data stream, and calculate the dynamic behavior anomaly evaluation value.
[0154] The multimodal feature fusion unit 302 is used to generate a dimensionality-reduced fusion feature vector based on the static semantic feature matrix and the dynamic derived feature matrix, using a multimodal feature adaptive fusion model constructed based on a gated cross-attention mechanism.
[0155] The dual-stream classification and clustering unit 303 is used to generate classification prediction results and unknown asset determination results based on the dynamic behavior anomaly degree evaluation value and the dimensionality reduction fusion feature vector, using the dual-stream classification prediction and weighted clustering model, and to collect the error correction sample set of misjudgment and / or reclassification.
[0156] The closed-loop fine-tuning unit 304 is used to perform active closed-loop fine-tuning of the multimodal feature adaptive fusion model and the dual-stream classification prediction and weighted clustering model based on contrastive learning with a forgetting factor, according to the error correction sample set and the historical sample set.
[0157] Based on the same inventive concept, another embodiment of the present invention provides an electronic device, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus.
[0158] Memory, used to store computer programs;
[0159] The processor, when executing a program stored in memory, implements the data asset automatic classification method based on data analysis of the present invention.
[0160] The communication bus mentioned above can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used in the diagram, but this does not indicate that there is only one bus or one type of bus. The communication interface is used for communication between the aforementioned terminal and other devices. The memory can include Random Access Memory (RAM) or Non-Volatile Memory (NVM), such as at least one disk storage device. Optionally, the memory can also be at least one storage device located remotely from the aforementioned processor.
[0161] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0162] Furthermore, to achieve the above objectives, embodiments of the present invention also propose a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the data asset automatic classification method based on data analysis according to embodiments of the present invention.
[0163] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, apparatus, or computer program products. Therefore, embodiments of the present invention can take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware aspects. Furthermore, embodiments of the present invention can take the form of computer program products implemented on one or more computer-usable hardware devices (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0164] The embodiments of the present invention are described with reference to flowchart illustrations and / or block diagrams of methods, terminal devices (apparatus), and computer program products according to embodiments of the invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0165] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0166] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal equipment, causing a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable terminal equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0167] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. "And / or" indicates that either one or both can be chosen. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes the element.
[0168] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for automatically classifying data assets based on data analysis, characterized in that, The method includes: Security probes are used to collect data streams of data assets, extract static semantic feature matrices of static files and dynamic derived feature matrices of disguised or dynamic files from the data streams, and calculate dynamic behavior anomaly assessment values. Based on the static semantic feature matrix and the dynamic derived feature matrix, a multimodal feature adaptive fusion model based on a gated cross-attention mechanism is used to generate a dimensionality-reduced fusion feature vector. Based on the dynamic behavior anomaly assessment value and the dimensionality reduction fusion feature vector, a dual-stream classification prediction and weighted clustering model is used to generate classification prediction results and unknown asset determination results, and to collect a set of error correction samples that are misjudged and / or need to be reclassified. Based on the error correction sample set and the historical sample set, active closed-loop fine-tuning of the multimodal feature adaptive fusion model and the two-stream classification prediction and weighted clustering model is carried out based on contrastive learning with forgetting factor.
2. The automatic data asset classification method based on data analysis according to claim 1, characterized in that, Security probes are used to collect data streams from data assets, extracting static semantic feature matrices from static files and dynamic derived feature matrices from disguised or dynamic files, and calculating dynamic behavior anomaly assessment values, including: Security probes are used to collect data streams from data assets, and stream reassembly and application-layer parsing are performed on the data streams to extract and restore independent file objects. Static parsing is performed on static files in the file object. If the file object is a text and / or office document, the text in the file object is extracted and input into the pre-trained language model to obtain the corresponding static semantic feature matrix. If the file object is an image file, the text in the file object is extracted by OCR and then converted into a static semantic feature matrix. Extract file metadata features from file objects and input them into a lightweight decision tree to determine whether they have spoofing or dynamic execution characteristics. If they do, the file objects are automatically isolated into a lightweight sandbox for controlled pre-execution. During the lightweight sandbox pre-execution, capture the dynamic behavior sequence generated by the dynamic execution of disguised or dynamic files in the file object, and vectorize the dynamic behavior sequence to obtain the dynamic derived feature matrix; If the file metadata features do not have spoofing or dynamic execution features, the output dynamic derived feature matrix is a zero matrix; The dynamic behavior anomaly assessment value is calculated based on the degree to which the dynamic derived feature matrix deviates from the normal baseline; wherein, the normal baseline is constructed based on the statistical analysis of the historical dynamic behavior sequence vectors of historical normal files in the lightweight sandbox.
3. The automatic data asset classification method based on data analysis according to claim 2, characterized in that, Based on the degree to which the dynamic derived feature matrix deviates from the normal baseline, a dynamic behavioral anomaly assessment value is calculated, including: Determine whether the dynamic derived feature matrix of the file object is a zero matrix. If it is a zero matrix, output the dynamic behavior anomaly evaluation value as 0 directly. If it is not a zero matrix, proceed to the next step. The set of historical dynamic behavior feature vectors of historical normal files in the lightweight sandbox is used to form a normal baseline, and its mean vector and covariance matrix are statistically analyzed. In the case that the file metadata features have spoofing or dynamic execution features, calculate the sequence mean vector of the dynamic derived feature matrix of the spoofed or dynamic file in the current file object; Based on the mean vector and covariance matrix of the normal baseline, and the sequence mean vector of the dynamically derived feature matrix, the Mahalanobis distance between the dynamically derived feature matrix and the normal baseline is calculated, which is used as the degree to which the dynamically derived feature matrix deviates from the normal baseline. The Mahalanobis distance is normalized using the Sigmoid function to obtain the dynamic behavior anomaly assessment value; If the file object does not have masquerading or dynamic execution characteristics, the output dynamic behavior anomaly evaluation value is 0.
4. The automatic data asset classification method based on data analysis according to claim 3, characterized in that, The multimodal feature adaptive fusion model based on the gated cross-attention mechanism includes a cross-attention calculation module, a dynamic gating coefficient calculation module, and a fusion feature generation module.
5. The automatic data asset classification method based on data analysis according to claim 4, characterized in that, Based on the static semantic feature matrix and the dynamic derived feature matrix, a multimodal feature adaptive fusion model based on a gated cross-attention mechanism is used to generate a dimensionality-reduced fusion feature vector, including: The static semantic feature matrix is mapped to the query matrix, and the dynamic derived feature matrix is mapped to the key matrix and the value matrix, respectively. Based on the query matrix, key matrix, and value matrix, the cross-attention calculation module of the multimodal feature adaptive fusion model built on a gated cross-attention mechanism is used to extract the part of the dynamic derived feature matrix that is highly related to the semantics of the static text and generate the attention feature matrix. Determine whether the dynamic derived feature matrix is a zero matrix. If it is a zero matrix, set the dynamic gate coefficient to 0 so that the multimodal feature adaptive fusion model automatically degenerates into pure static feature extraction. If it is not a zero matrix, proceed to the next step. The static semantic feature matrix and the dynamic derived feature matrix are input into the fully connected layer of the dynamic gating coefficient calculation module of the multimodal feature adaptive fusion model to generate dynamic gating coefficients. Based on dynamic gating coefficients, the attention feature matrix and static semantic feature matrix are input into the fusion feature generation module of the multimodal feature adaptive fusion model to generate a dimensionality-reduced fusion feature vector.
6. The automatic data asset classification method based on data analysis according to claim 5, characterized in that, The dual-stream classification prediction and weighted clustering model includes a supervised stream classification prediction branch and an unsupervised stream weighted clustering branch; The supervised flow classification prediction branch is equipped with a DNN classifier; The unsupervised flow weighted clustering branch is equipped with an adaptive distance weighted clustering program, and the unsupervised flow weighted clustering branch is connected to a known data asset library, which contains several known data asset clusters.
7. The automatic data asset classification method based on data analysis according to claim 6, characterized in that, Based on the dynamic behavior anomaly assessment value and the dimensionality-reduced fusion feature vector, a two-stream classification prediction and weighted clustering model is used to generate classification prediction results and unknown asset determination results. A set of error-corrected samples that are misclassified or require reclassification is also collected, including: The dimensionality-reduced fused feature vector is input into the two-stream classification prediction and weighted clustering model. The DNN classifier of the supervised stream classification prediction branch is used to calculate the probability distribution of each category of the dimensionality-reduced fused feature vector, output the classification label and sensitivity level of the current data asset, and generate the classification prediction result. Based on the dynamic behavior anomaly evaluation value and the dimensionality reduction fusion feature vector, the unsupervised flow weighted clustering branch is used to call the adaptive distance weighted clustering program to calculate the comprehensive weighted distance between the dimensionality reduction fusion feature vector and the cluster centers of several known data asset clusters in the known data asset library, and obtain the minimum comprehensive weighted distance among all comprehensive weighted distances; If the minimum weighted average distance is greater than the distance threshold, the dimensionality reduction and fusion feature vector of the current file object is divided into unknown asset clusters, and all unknown asset clusters are aggregated into an active learning request sample set. The current dimensionality reduction and fusion feature vector with classification prediction results and the active learning request sample set are used to generate a visual data asset map on the front end. The error correction sample set obtained by the operation and maintenance personnel in judging the data asset map and / or needing to be reclassified is also collected.
8. The automatic data asset classification method based on data analysis according to claim 7, characterized in that, Based on the error correction sample set and the historical sample set, active closed-loop fine-tuning of the multimodal feature adaptive fusion model and the two-stream classification prediction and weighted clustering model is performed based on contrastive learning with a forgetting factor, including: The error correction samples in the error correction sample set are used as anchor samples. Several historical samples of the same category as the anchor samples are retrieved from the historical sample set in the known data asset library as positive samples. Several historical samples whose true category is different from the anchor samples but whose features are similar are retrieved as negative samples. Calculate the time decay elastic forgetting factor, and dynamically adjust the interval of contrastive learning based on the time decay elastic forgetting factor to obtain the adaptive interval for the current batch. Using a multimodal feature adaptive fusion model, sample dimensionality reduction fusion feature vectors are generated for each anchor sample, each positive sample, and each negative sample. The sample dimensionality reduction fusion feature vectors and the adaptive interval are then substituted into the joint contrastive learning loss function with parameter regularization to obtain the comprehensive loss value. The gradient vector is obtained by taking the partial derivative of the comprehensive loss value with respect to the model parameters. Then, the gradient descent algorithm is used to perform gradient descent with a set learning rate to complete the synchronous iterative update of the parameters and obtain the updated model parameters. Based on the updated model parameters, the multimodal feature adaptive fusion model and the two-stream classification prediction and weighted clustering model are fine-tuned to obtain the updated multimodal feature adaptive fusion model and the updated two-stream classification prediction and weighted clustering model, which are used for the automatic classification of data assets in the next batch.
9. The automatic data asset classification method based on data analysis according to claim 8, characterized in that, The model parameters include the first model weight matrix and the first model bias vector of the fully connected layer, and the second model weight matrix and the second model bias vector of the DNN classifier.
10. A data asset automatic classification system based on data analysis, used to implement the data asset automatic classification method as described in any one of claims 1-9, characterized in that, The system includes: The data acquisition and feature extraction unit is used to collect data streams of data assets using security probes, extract static semantic feature matrices of static files and dynamic derived feature matrices of disguised or dynamic files in the data stream, and calculate dynamic behavior anomaly evaluation values. The multimodal feature fusion unit is used to generate a dimensionality-reduced fusion feature vector based on a multimodal feature adaptive fusion model constructed using a gated cross-attention mechanism, according to the static semantic feature matrix and the dynamic derived feature matrix. The dual-stream classification and clustering unit is used to generate classification prediction results and unknown asset determination results based on dynamic behavior anomaly assessment values and dimensionality-reduced fusion feature vectors using a dual-stream classification prediction and weighted clustering model, and to collect error correction sample sets of misjudged and / or reclassified cases. The closed-loop fine-tuning unit is used to perform active closed-loop fine-tuning of the multimodal feature adaptive fusion model and the two-stream classification prediction and weighted clustering model based on contrastive learning with a forgetting factor, according to the error correction sample set and the historical sample set.