Heterogeneous data utility dynamic measurement method and device, electronic equipment and storage medium

CN122778293APending Publication Date: 2026-09-18北京国信中健人工智能科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610918989.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-24
Publication Date
2026-09-18

AI Technical Summary

Technical Problem

[0006]本发明的目的在于提供异构数据效用动态度量方法、装置、电子设备及存储介质,能够解决现有异构数据特征提取失真以及度量指标僵化的问题

Benefits of technology

[0017]The beneficial effects of this invention are as follows: First, this invention parses the security classification labels of heterogeneous source data in real time and performs controlled computing environment routing, thus ensuring absolute physical and logical isolation security in the data measurement process. Second, because this invention performs dual-track feature decoupling extraction on the benchmark sample set, it abandons the traditional flat splicing method, thus preserving the internal high-dimensional topological logic of heterogeneous data and greatly improving the representation accuracy of feature vectors. Furthermore, this invention can introduce security classification labels as prior control signals through a dynamic attention network, enabling adaptive adjustment of the feature capture perspective according to the compliance constraints of the data. In addition, for highly sensitive data that is deeply masked, this invention can automatically enhance the weight of topological features, completely solving the problem of blind spots in dense data evaluation caused by static weights. Finally, this invention introduces time decay and quality penalty in the measurement calculation, filtering out redundant noise from high-frequency invalid updates from a mathematical logic perspective, thereby ensuring that computing resources are accurately allocated to active data with real business utility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122778293A_ABST
    Figure CN122778293A_ABST
Patent Text Reader

Abstract

This invention belongs to the field of data processing and information representation technology, and discloses a method, device, electronic device, and storage medium for dynamic measurement of heterogeneous data utility. The method mainly includes: receiving heterogeneous source data in real time and identifying security classification labels; routing the heterogeneous source data to a controlled computing environment and constructing a consistent read-only view of the corresponding source data and generating a benchmark sample set; performing dual-track feature decoupling extraction on the benchmark sample set to obtain topological structure feature vectors and deep content feature vectors respectively; in a dynamic attention network, adaptively weighting and fusing the topological structure feature vectors and deep content feature vectors according to dynamic weight coefficients to generate a comprehensive data feature vector; obtaining the time decay factor and quality penalty factor of the benchmark sample set, and calculating the final utility index based on the time decay factor, quality penalty factor, and comprehensive data feature vector; when the final utility index meets the admission threshold, triggering data object admission and encapsulation flow.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing and information representation technology, and in particular to a method, apparatus, electronic device and storage medium for dynamic measurement of the utility of heterogeneous data. Background Technology

[0002] In distributed heterogeneous data processing systems, objectively and accurately measuring the "effective information utility" of massive amounts of raw data is a core prerequisite for achieving standardized data encapsulation and on-demand scheduling of system resources. Currently, when extracting and measuring features from heterogeneous data, some common approaches involve extracting data features using general-purpose large language models (such as BERT) and performing similarity comparisons.

[0003] Due to the diversity of heterogeneous data sources, existing conventional measurement methods have significant drawbacks: First, existing feature extraction methods typically forcibly concatenate the table structure (Schema) and content of relational databases into a one-dimensional long text input model. This flattening and dimensionality reduction destroys the original high-dimensional topological structure features of the data, such as primary and foreign key dependencies and field type distribution, resulting in distorted feature representation.

[0004] Second, existing algorithms typically assign static weights to content and structural features. In controlled computing scenarios, for highly sensitive data, the plaintext content is often severely anonymized or masked. Algorithms with static weights cannot dynamically adjust their attention based on the data security level, leading to the failure of feature extraction from dense data.

[0005] Third, existing systems typically calculate data activity by simply multiplying data volume by update frequency. This results in redundant log data containing a large number of null values ​​or meaningless automatic system refreshes receiving extremely high scores, failing to accurately reflect the lifecycle and true timeliness characteristics of valid data. Summary of the Invention

[0006] The purpose of this invention is to provide a method, apparatus, electronic device and storage medium for dynamic measurement of heterogeneous data utility, which can solve the problems of distortion in existing heterogeneous data feature extraction and rigidity in measurement indicators.

[0007] The technical solution adopted by this invention to solve its technical problem is as follows: In a first aspect, the present invention provides a method for dynamically measuring the utility of heterogeneous data, comprising the following steps: It receives heterogeneous source data in real time and identifies the security classification labels of the heterogeneous source data. Based on the security classification label, heterogeneous source data is routed to the controlled computing environment corresponding to the security classification label; Construct a consistent read-only view of the corresponding source data in memory within a controlled computing environment and generate a benchmark sample set; Dual-track feature decoupling extraction is performed on the benchmark sample set to obtain the topological feature vector of the structure flow and the deep content feature vector of the content flow, respectively. The security classification label is input into the dynamic attention network to obtain dynamic weight coefficients. Based on the dynamic weight coefficients, the topological structure feature vector and the deep content feature vector are adaptively weighted and fused to generate a comprehensive data feature vector. Obtain the time decay factor and quality penalty factor of the benchmark sample set, and calculate the final utility index based on the time decay factor, quality penalty factor and comprehensive data feature vector. When the final utility index meets the admission threshold, the data object admission and encapsulation flow are triggered.

[0008] In some embodiments, after receiving heterogeneous source data, the security classification label for identifying heterogeneous source data refers to: It receives the connection request from the current data source, intercepts the data source's message header and data definition language metadata through the tag parsing layer, calls the regular expression rule engine to extract the table name and field name set, and inputs the table name and field name set into a preset compliance dictionary for matching to obtain the security classification tag represented by four-dimensional one-hot encoding.

[0009] In some embodiments, routing heterogeneous source data to a controlled computing environment corresponding to a security classification label based on the security classification label means: Security classification labels are captured through the hardware routing and scheduling layer, and conditional branch instructions are executed based on the security classification labels. The conditional branch instruction based on the security classification label is: When the last bit of the four-dimensional one-hot encoding of the current security level label is detected to be 1, it indicates that the source data corresponding to the current security level label is high-sensitivity data. At this time, a CPU hardware interrupt is triggered, and the subsequent processing flow of the source data stream corresponding to the current security level label is routed to the encrypted memory page of the TEE using direct memory access technology and page table isolation mechanism. Otherwise, it indicates that the source data corresponding to the current security level label is low-sensitivity data. At this time, the low-sensitivity data is mounted to the cgroup resource pool of the standard Docker data sandbox.

[0010] In some embodiments, constructing a consistent read-only view of the corresponding source data in the memory of a controlled computing environment and generating a benchmark sample set refers to: In a controlled computing environment, the database's multi-version concurrency control protocol is invoked to send read-only transaction instructions, or a memory-mapped file stream is created at the file system layer to directly map local slices of the source data into the volatile memory of the TEE in the form of a binary stream, and generate a basic sample set.

[0011] In some embodiments, the dual-track feature decoupling extraction of the benchmark sample set to obtain the topological feature vector of the structure flow and the deep content feature vector of the content flow respectively includes the following steps: The baseline sample set is scanned using a data stream separation algorithm, and the samples are physically divided into two independent data streams based on ASCII control characters and data table structure identifiers. The two independent data streams include a structure stream and a content stream. The topology parsing engine is used to parse the structured flow into graph structured data. Each field of the structured flow is mapped to a set of graph nodes, the foreign keys and logical dependencies between fields are mapped to an adjacency matrix, and the data types are mapped to a node feature matrix through embedding. The entity extraction engine is used to clean up garbled characters and format control characters in the content stream, generating a clean serialized text sequence, and then tokenizing and encoding it. The adjacency matrix and node feature matrix are input into the graph convolutional neural network, and the feature aggregation operator of two layers of graph convolution is calculated, and global average pooling is performed. The serialized text sequence after tokenization and encoding is input into the quantized language model loaded in the video memory, and the contextual semantics are captured through the internal multi-head self-attention mechanism, and the hidden state of the last layer [CLS] flag bit is extracted. The result of global average pooling is solidified and output as a specified-dimensional topological structure feature vector representing the intrinsic topological dimension of the data, and the [CLS] flag is solidified and output as a specified-dimensional deep content feature vector representing the extrinsic semantic dimension of the data.

[0012] In some embodiments, the adaptive weighted fusion of the topological structure feature vector and the deep content feature vector based on dynamic weight coefficients to generate a comprehensive data feature vector is calculated using the following formula: ; in, To synthesize the feature vector of the data, For safety classification labels, For deep content feature vectors, It is a topological feature vector. and For dynamic attention networks, these are the dynamic weight coefficients output in real time based on safety classification labels, and Norm is the normalization function.

[0013] In some embodiments, the formula for calculating the final utility index based on the time decay factor, the quality penalty factor, and the comprehensive data feature vector is as follows: ; in, For the final utility index, For the i-th reference vector in the pre-set reference feature library, For similarity weights, The time decay factor, This is the time difference since the last valid update. As a quality penalty factor, For the data sample size, This is a weight based on activity level.

[0014] Secondly, the present invention also provides a device for dynamically measuring the utility of heterogeneous data, comprising: The data receiving and identification module is used to receive heterogeneous source data in real time and identify the security classification labels of the heterogeneous source data. The environment routing module is used to route heterogeneous source data to the controlled computing environment corresponding to the security classification label based on the security classification label. The sample set construction module is used to build a consistent read-only view of the corresponding source data in the memory of a controlled computing environment and generate a benchmark sample set. The feature extraction module is used to perform dual-track feature decoupling extraction on the benchmark sample set, and to obtain the topological feature vector of the structure flow and the deep content feature vector of the content flow respectively. The dynamic fusion module is used to input the security classification label into the dynamic attention network, obtain the dynamic weight coefficients, and perform adaptive weighted fusion of the topological structure feature vector and the deep content feature vector according to the dynamic weight coefficients to generate a comprehensive data feature vector. The utility measurement module is used to obtain the time decay factor and quality penalty factor of the benchmark sample set, and calculate the final utility index based on the time decay factor, quality penalty factor and comprehensive data feature vector. When the final utility index meets the admission threshold, the data object admission and encapsulation flow are triggered.

[0015] Thirdly, the present invention also provides an electronic device, comprising: a processor, a memory, and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor communicates with the memory via the bus, and the processor executes the machine-readable instructions to perform the steps of the heterogeneous data utility dynamic measurement method.

[0016] Fourthly, the present invention also provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, performs the steps of the heterogeneous data utility dynamic measurement method.

[0017] The beneficial effects of this invention are as follows: First, this invention parses the security classification labels of heterogeneous source data in real time and performs controlled computing environment routing, thus ensuring absolute physical and logical isolation security in the data measurement process. Second, because this invention performs dual-track feature decoupling extraction on the benchmark sample set, it abandons the traditional flat splicing method, thus preserving the internal high-dimensional topological logic of heterogeneous data and greatly improving the representation accuracy of feature vectors. Furthermore, this invention can introduce security classification labels as prior control signals through a dynamic attention network, enabling adaptive adjustment of the feature capture perspective according to the compliance constraints of the data. In addition, for highly sensitive data that is deeply masked, this invention can automatically enhance the weight of topological features, completely solving the problem of blind spots in dense data evaluation caused by static weights. Finally, this invention introduces time decay and quality penalty in the measurement calculation, filtering out redundant noise from high-frequency invalid updates from a mathematical logic perspective, thereby ensuring that computing resources are accurately allocated to active data with real business utility. Attached Figure Description

[0018] Figure 1 This is a flowchart of the dynamic measurement method for heterogeneous data utility in Embodiment 1 of the present invention; Figure 2 This is a flowchart of the dual-track feature decoupling extraction of the benchmark sample set in Embodiment 1 of the present invention; Figure 3 This is a flowchart illustrating the adaptive weighted fusion of topological structure feature vectors and deep content feature vectors based on dynamic weight coefficients in Embodiment 1 of the present invention. Figure 4 This is a flowchart illustrating the calculation of the final utility value using a spatial geometric matching and time decay model in Embodiment 1 of the present invention. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0020] Example 1:

[0021] In this embodiment, the system is deployed on an edge server node. The hardware environment includes a central processing unit supporting the SGX instruction set (for allocating memory pages for the TEE trusted execution environment) and a tensor acceleration computing card with dedicated video memory (for supporting low-precision quantization model inference). The underlying computer execution of the entire heterogeneous data utility dynamic measurement method is based on the above deployment configuration. This embodiment provides a heterogeneous data utility dynamic measurement method, the flowchart of which can be found in [link to flowchart]. Figure 1 The method may include the following steps: S1. Receive heterogeneous source data in real time and identify the security classification labels of heterogeneous source data; S2. Based on the security classification label, route the heterogeneous source data to the controlled computing environment corresponding to the security classification label; S3. Construct a consistent read-only view of the corresponding source data in the memory of the controlled computing environment and generate a benchmark sample set; S4. Perform dual-track feature decoupling extraction on the benchmark sample set to obtain the topological feature vector of the structure flow and the deep content feature vector of the content flow, respectively. S5. Input the security classification label into the dynamic attention network, obtain the dynamic weight coefficients, and perform adaptive weighted fusion of the topological structure feature vector and the deep content feature vector according to the dynamic weight coefficients to generate a comprehensive data feature vector. S6. Obtain the time decay factor and quality penalty factor of the benchmark sample set, and calculate the final utility index based on the time decay factor, quality penalty factor and comprehensive data feature vector. When the final utility index meets the admission threshold, trigger the data object admission and encapsulation flow.

[0022] It should be noted that the above method in this embodiment needs to solve the problem of secure access and physical isolation reading of heterogeneous data sources in the initial execution stage (corresponding to steps S1 to S3 of the above method in this embodiment). Therefore, this embodiment needs to limit the above steps S1 to S3 respectively.

[0023] First, regarding the security classification labels for identifying heterogeneous source data mentioned in step S1 above, which is the security classification label parsing based on metadata feature matching, when a data source connection request is received, the label parsing layer intercepts the data source's message header and DDL (Data Definition Language) metadata, and calls the internal regular expression rule engine to extract the table name and field name set. This set is then input into a pre-set compliance dictionary for matching, and the resulting security classification label is represented by a four-dimensional one-hot encoding. For example, if a match is found in fields such as "default" or "case", then output... It represents the highest security level, L4.

[0024] Secondly, after identifying the security classification labels of heterogeneous source data, the data can be routed to the controlled computing environment corresponding to the security classification label, i.e., hardware routing scheduling based on label interrupts is executed. In this embodiment, the data can be captured at the hardware routing scheduling layer. Then, the conditional branch instruction is executed: when the last bit of the four-dimensional one-hot encoding of the current security level label is detected to be 1 (i.e., L4 level), it indicates that the source data corresponding to the current security level label is high-sensitivity data. At this time, a CPU hardware interrupt is triggered, and the subsequent processing flow of the data stream is hard-routed to the encrypted memory page (Enclave) of the TEE using direct memory access (DMA) technology and page table isolation mechanism; when the last bit of the four-dimensional one-hot encoding of the current security level label is detected to be not 1, it indicates that the source data corresponding to the current security level label is low-sensitivity data (such as [1,0,0,0], L1 level). At this time, the low-sensitivity data is mounted to the cgroup resource pool of the standard Docker data sandbox.

[0025] Then, a consistent read-only view of the corresponding source data can be built in the memory of the controlled computing environment and a benchmark sample set can be generated. In this embodiment, within the determined controlled computing environment, the database's multi-version concurrency control (MVCC) protocol can be called to send a read-only transaction instruction (Start Transaction Read Only), or a memory-mapped file stream (MMF) can be created at the file system layer to directly map local slices of the source data into the volatile memory of the TEE in the form of a binary stream, thereby generating a benchmark sample set, and any persistent disk I / O operations are strictly prohibited.

[0026] It should be noted that after the benchmark sample set is generated as described above, the benchmark sample set can be decoupled and extracted using a dual-track feature extraction method to obtain the topological feature vector of the structure flow and the deep content feature vector of the content flow, respectively.

[0027] In this embodiment, a graph encoder can be invoked to traverse the field relationships and data type distributions in the benchmark sample set and output the topological feature vector of the structure stream; at the same time, a quantized language model encoder can be invoked to perform semantic representation on the data content of the benchmark sample set and output the deep content feature vector of the content stream.

[0028] Here, a graph encoder is used to process the structure flow, which can accurately capture the primary and foreign key dependencies and logical topology of relational databases; a quantized language model is used to process the content flow, which can effectively reduce the computational overhead in the controlled environment and ensure deep semantic extraction. The two processes are decoupled and do not interfere with each other, laying a high-precision feature foundation for subsequent dynamic fusion.

[0029] Specifically, in this embodiment, a dual-track feature decoupling extraction is performed on the benchmark sample set to obtain the topological structure feature vector of the structure flow and the deep content feature vector of the content flow, respectively. This involves decoupling the benchmark sample set in controlled memory and converting it into a high-dimensional feature tensor. (See [link to relevant documentation]). Figure 2 In practical applications, it can be achieved through the following steps: S410. Sample Stream Splitting and Decomposition. This embodiment can use a data stream separation algorithm to scan the benchmark sample set and, based on ASCII control characters and data table structure identifiers, physically split the samples into two independent data streams: a "Schema Stream" containing field names, foreign keys, and constraints, and a "Content Stream" containing specific numerical values ​​and strings.

[0030] S420, concurrent parsing of topology and entity engines.

[0031] S420A (Topology Parsing) is an engine that parses structured streams into graph-structured data. It maps each field to a graph node set V, maps foreign keys and logical dependencies between fields to an adjacency matrix A, and maps data types (such as VARCHAR and INT) to a node feature matrix X via embeddings.

[0032] The S420B (Entity Extraction) engine cleans up garbled characters and formatting control characters in the content stream, generates a clean serialized text sequence, and performs tokenization and word encoding.

[0033] S430, feature extraction calculation of the underlying encoder.

[0034] S430A (graph encoder processing) inputs the adjacency matrix A and the feature matrix X into a graph convolutional neural network (GCN) to perform feature aggregation operator computation over two layers of graph convolution: ; in, Let A represent the identity matrix, with the same dimensions as the adjacency matrix A. It is used to add self-connections between nodes in the original adjacency matrix, i.e., to construct an adjacency matrix with self-loops. , for The degree matrix, This represents a non-linear activation function used to perform non-linear transformations on the features after graph convolution aggregation, thereby enhancing the model's ability to express complex structural features such as field relationships and data type distributions. In practical implementation, Activation functions such as ReLU, Sigmoid, and Tanh can be used, with ReLU being the preferred choice. Indicates the first The node feature matrix output by the layer graph convolutional network, that is, after the th layer... The new layer of node representation is obtained after layer graph convolution aggregation, weight transformation, and nonlinear activation. Indicates the first The node feature matrix input to the layer graph convolutional network, when When =0, That is, the initial node feature matrix is ​​the feature matrix X, when When >0, This represents the feature representation of the intermediate nodes output by the previous layer's graph convolutional network and passed to the current layer. Indicates the first The trainable weight matrix of a layer graph convolutional network is used to train the weights of the first layer graph convolutional network. The layer node features are linearly transformed to learn the feature representation of different field nodes and their relationships in the topology.

[0035] Global average pooling is then performed. The S430B (quantized SLM processing) inputs the segmented text sequence into a vertical small language model loaded in GPU memory with extremely low precision INT4 quantization. Contextual semantics are captured through the multi-head attention mechanism within the larger model, and the [CLS] flag of the last hidden states is extracted.

[0036] S440, Dual-track Feature Tensor Solidification and Output.

[0037] In this step, the system allocates specific memory addresses to store the extracted high-dimensional features: S440A (Topological Feature Output Action): Receives the result of global pooling performed by the S430A graph encoder, solidifies it, and outputs it as a 768-dimensional topological structure feature vector representing the inherent topological dimension of the data. .

[0038] S440B (Content Feature Output Action) receives the hidden state of the [CLS] flag bit extracted by the S430B quantized language model, solidifies it, and outputs it as a 768-dimensional deep content feature vector representing the external semantic dimension of the data. .

[0039] At this point, the benchmark sample set has been completely transformed into two parallel, purely mathematical high-dimensional tensors at the physical level, asynchronously awaiting invocation by the downstream dynamic attention network.

[0040] In this embodiment, the topological structure feature vector and the deep content feature vector are adaptively weighted and fused according to dynamic weight coefficients to generate a comprehensive data feature vector. The calculation formula is as follows: ; in, To synthesize the feature vector of the data, For safety classification labels, For deep content feature vectors, It is a topological feature vector. and For dynamic attention networks, these are the dynamic weight coefficients output in real time based on safety classification labels, and Norm is the normalization function.

[0041] It should be noted that for step S5 of this embodiment, see [link to relevant documentation]. Figure 3 This can be achieved through the following steps: S510: Loading computational context and control signals. This embodiment can synchronously allocate space in the video memory of the tensor accelerator card to load 768-dimensional... 768-dimensional and four-dimensional control vector .

[0042] S520, Miniature Multilayer Perceptron (MLP) gated computation. Four-dimensional vector... The input is fed into an MLP network containing two fully connected hidden layers, and nonlinear dimension mapping calculations are performed.

[0043] S530, dynamic attention weight calculation: After receiving the input signal, the miniature multilayer perceptron (MLP) performs forward propagation calculation. Due to the input from the preceding steps... For discrete security compliance labels (such as L1 to L4), the encoding operator is first invoked to map them into column vectors in the real number field. : ; in, The weight matrix is ​​the learnable weight matrix in the hidden layer. The system's preset security classification dimensions are divided into 4 security levels in this embodiment. =4.

[0044] Then, the column vector Input the hidden layer of the MLP network and compute the feature representation h: ; in, The hidden layer is a learnable weight matrix. For the hidden layer neurons, Here, ReLU is the hidden layer bias vector, and ReLU is the non-linear activation function. Subsequently, the output layer's logits vector is calculated. : ; in, This is the output layer weight matrix. The output layer bias vector is forced to map the dimension to a two-dimensional tensor. Finally, the logistic vector z is input into the Softmax activation function, which maps it to a probability distribution that sums to 1, to obtain the final dynamic scalar coefficients: ; in, This represents the dynamic attention weight allocation coefficients for deep content feature vectors. This represents the dynamic attention weight allocation coefficients for the topological feature vectors, and mathematically it always satisfies... Thus, a precise mapping from discrete security compliance signals to continuous tensor calculation coefficients was achieved.

[0045] S540, Feature Adaptive Tilt Intervention Execution: If the input is an L4 level high-sensitivity vector, the MLP weight network automatically outputs an extremely skewed coefficient allocation after calculation (e.g., =0.05 corresponds to the content, =0.95 corresponding structure), at the mathematical level, when the data is in a state of high-frequency desensitization and content is invisible, this embodiment can automatically reduce the dependence on the distorted content vector and forcibly shift the computational attention to the undamaged topological structure vector.

[0046] S550-S560, tensor weighting and normalization are combined, the tensor dot multiplication operator is called to perform feature fusion, and L2 normalization is used to smooth the dimensions: ; The underlying variables and their physical meanings in the formula are defined as follows: This represents the deep content feature vector output by step S430B (output dimension d=768 in this embodiment). This represents the topological feature vector output by step S430A (with dimensions d=768, satisfying the alignment requirements of high-dimensional matrix addition). These are the scalar coefficients calculated by the pre-MLP gating network, which serve as dynamic weighting factors to control the attention ratio of the corresponding modality in the fused representation; The L2 norm operator (i.e., the square root of the sum of the squares of the vector elements) of a tensor mathematically smooths the convergence of a weighted high-dimensional tensor to the unit hypersphere. The final system output is a comprehensive data feature vector, which fully preserves the structure and content information after dynamic tilting and is standardized and mapped to a unified feature space for cosine distance calculation in subsequent steps.

[0047] Therefore, this embodiment ultimately outputs a unified 768-dimensional comprehensive data feature vector. .

[0048] It should be noted that after completing the comprehensive data feature vector as described above, this embodiment can calculate the final utility index based on the time decay factor, quality penalty factor, and comprehensive data feature vector. The calculation formula is as follows: ; in, For the final utility index, For the i-th reference vector in the pre-set reference feature library, For similarity weights, The time decay factor, This is the time difference since the last valid update. As a quality penalty factor, For the data sample size, This is a weight based on activity level.

[0049] Specifically, this embodiment can abandon static rule scoring and calculate the final utility value through spatial geometric matching and time decay model, see [link to relevant documentation]. Figure 4 The specific implementation process can be achieved through the following steps: S610, High-dimensional semantic tensor space matching. This embodiment pre-sets a superior data feature matrix. (Each column represents a standard high-value feature vector) ), and call the matrix multiplication operator to calculate The cosine similarity with each column in the matrix is ​​used, and the maximum scalar value is extracted as the semantic matching score.

[0050] S620, extraction of time-domain and physical quality parameters. Here, the operating system's underlying API can be called to obtain the inode metadata of the benchmark sample set and extract the time difference since the last valid write operation. (Unit: hours), and the physical storage capacity occupied. (Unit: MB), and simultaneously iterate through the sample set to count the proportion of empty cells, thus obtaining the penalty factor. ( ).

[0051] S630-S640, nonlinear time-domain decay calculation and final utility output. This embodiment uses the above discrete parameter inputs based on the Sigmoid activation function and the natural constant. Floating-point operations are performed in the constructed utility mathematical model: ; in, The preset time decay factor, and These are the preset similarity weight and activity weight, respectively. Under the action of this operator, when... When the value is extremely large (data is outdated for a long time), the exponential decay term Approaching 0 severely compresses the activity score brought by large files, eliminating the possibility of redundant large files being used to generate scores from the algorithm's underlying layer.

[0052] S650, state machine memory interrupt and closed-loop triggering. In this embodiment, the calculated... The data is stored in the floating-point comparator register and compared with a preset threshold (e.g., 0.85) for one clock cycle. Once the condition is met, the state machine triggers a "pass successful" interrupt signal, immediately releasing the reference sample set temporarily stored in the TEE to clear the controlled memory, and... and The pointer address is transferred to the downstream sandbox scheduling and code encapsulation bus through inter-process communication (IPC) to complete the fully automated dynamic measurement process of heterogeneous data utility.

[0053] Example 2: Based on Embodiment 1, this embodiment provides a dynamic measurement device for heterogeneous data utility, wherein the device includes: The data receiving and identification module is used to receive heterogeneous source data in real time and identify the security classification labels of the heterogeneous source data. The environment routing module is used to route heterogeneous source data to the controlled computing environment corresponding to the security classification label based on the security classification label. The sample set construction module is used to build a consistent read-only view of the corresponding source data in the memory of a controlled computing environment and generate a benchmark sample set. The feature extraction module is used to perform dual-track feature decoupling extraction on the benchmark sample set, and to obtain the topological feature vector of the structure flow and the deep content feature vector of the content flow respectively. The dynamic fusion module is used to input the security classification label into the dynamic attention network, obtain the dynamic weight coefficients, and perform adaptive weighted fusion of the topological structure feature vector and the deep content feature vector according to the dynamic weight coefficients to generate a comprehensive data feature vector. The utility measurement module is used to obtain the time decay factor and quality penalty factor of the benchmark sample set, and calculate the final utility index based on the time decay factor, quality penalty factor and comprehensive data feature vector. When the final utility index meets the admission threshold, the data object admission and encapsulation flow are triggered.

[0054] As can be seen from the description of Embodiment 1, the application scenario and working principle of this embodiment are the same as those of Embodiment 1, so they will not be repeated here.

[0055] Example 3: This embodiment provides an electronic device, including a processor, a memory, and a bus. The memory stores machine-readable instructions that the processor can execute. When the electronic device is running, the processor communicates with the memory via the bus, and the processor executes the machine-readable instructions to perform the steps of the heterogeneous data utility dynamic measurement method as described in Embodiment 1.

[0056] Example 4:

[0057] This embodiment provides a computer-readable storage medium on which a computer program is stored. When the computer program is run by a processor, it executes the steps of the heterogeneous data utility dynamic measurement method as described in Embodiment 1.

[0058] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for dynamically measuring isomorphic data utility, the method comprising: Includes the following steps: It receives heterogeneous source data in real time and identifies the security classification labels of the heterogeneous source data. Based on the security classification label, heterogeneous source data is routed to the controlled computing environment corresponding to the security classification label; Construct a consistent read-only view of the corresponding source data in memory within a controlled computing environment and generate a benchmark sample set; Dual-track feature decoupling extraction is performed on the benchmark sample set to obtain the topological feature vector of the structure flow and the deep content feature vector of the content flow, respectively. The security classification label is input into the dynamic attention network to obtain dynamic weight coefficients. Based on the dynamic weight coefficients, the topological structure feature vector and the deep content feature vector are adaptively weighted and fused to generate a comprehensive data feature vector. Obtain the time decay factor and quality penalty factor of the benchmark sample set, and calculate the final utility index based on the time decay factor, quality penalty factor and comprehensive data feature vector. When the final utility index meets the admission threshold, the data object admission and encapsulation flow are triggered.

2. The method for dynamically measuring the utility of heterogeneous data according to claim 1, characterized in that, Upon receiving heterogeneous source data, the security classification label for identifying heterogeneous source data refers to: It receives the connection request from the current data source, intercepts the data source's message header and data definition language metadata through the tag parsing layer, calls the regular expression rule engine to extract the table name and field name set, and inputs the table name and field name set into a preset compliance dictionary for matching to obtain the security classification tag represented by four-dimensional one-hot encoding.

3. The method for dynamically measuring the utility of heterogeneous data according to claim 2, characterized in that, The step of routing heterogeneous source data to a controlled computing environment corresponding to a security classification label based on the security classification label means: Security classification labels are captured through the hardware routing and scheduling layer, and conditional branch instructions are executed based on the security classification labels. The conditional branch instruction based on the security classification label is: When the last bit of the four-dimensional one-hot encoding of the current security level label is detected to be 1, it indicates that the source data corresponding to the current security level label is high-sensitivity data. At this time, a CPU hardware interrupt is triggered, and the subsequent processing stream of the source data stream corresponding to the current security level label is routed to the encrypted memory page of the TEE using direct memory access technology and page table isolation mechanism. Otherwise, it indicates that the source data corresponding to the current security classification label is low-sensitivity data. In this case, the low-sensitivity data will be mounted to the cgroup resource pool of the standard Docker data sandbox.

4. The method for dynamically measuring the utility of heterogeneous data according to claim 1, characterized in that, The process of constructing a consistent read-only view of the corresponding source data in the memory of a controlled computing environment and generating a benchmark sample set refers to: In a controlled computing environment, the database's multi-version concurrency control protocol is invoked to send read-only transaction instructions, or a memory-mapped file stream is created at the file system layer to directly map local slices of the source data into the volatile memory of the TEE in the form of a binary stream, and generate a basic sample set.

5. The method for dynamically measuring the utility of heterogeneous data according to claim 1, characterized in that, The step of performing dual-track feature decoupling extraction on the benchmark sample set to obtain the topological structure feature vector of the structure flow and the deep content feature vector of the content flow includes the following steps: The baseline sample set is scanned using a data stream separation algorithm, and the samples are physically divided into two independent data streams based on ASCII control characters and data table structure identifiers. The two independent data streams include a structure stream and a content stream. The topology parsing engine is used to parse the structured flow into graph structured data. Each field of the structured flow is mapped to a set of graph nodes, the foreign keys and logical dependencies between fields are mapped to an adjacency matrix, and the data types are mapped to a node feature matrix through embedding. The entity extraction engine is used to clean up garbled characters and format control characters in the content stream, generating a clean serialized text sequence, and then tokenizing and encoding it. The adjacency matrix and node feature matrix are input into the graph convolutional neural network, and the feature aggregation operator of two layers of graph convolution is calculated, and global average pooling is performed. The serialized text sequence after tokenization and encoding is input into the quantized language model loaded in the video memory, and the contextual semantics are captured through the internal multi-head self-attention mechanism, and the hidden state of the last layer [CLS] flag bit is extracted. The result of global average pooling is solidified and output as a specified-dimensional topological structure feature vector representing the intrinsic topological dimension of the data, and the [CLS] flag is solidified and output as a specified-dimensional deep content feature vector representing the extrinsic semantic dimension of the data.

6. The method for dynamically measuring the utility of heterogeneous data according to claim 1, characterized in that, The adaptive weighted fusion of the topological structure feature vector and the deep content feature vector based on dynamic weight coefficients to generate a comprehensive data feature vector is calculated using the following formula: ; in, To synthesize the feature vector of the data, For safety classification labels, For deep content feature vectors, It is a topological feature vector. and For dynamic attention networks, these are the dynamic weight coefficients output in real time based on safety classification labels, and Norm is the normalization function.

7. The method for dynamically measuring the utility of heterogeneous data according to claim 6, characterized in that, The final utility index is calculated based on the time decay factor, quality penalty factor, and comprehensive data feature vector. The calculation formula is as follows: ; in, For the final utility index, For the i-th reference vector in the pre-set reference feature library, For similarity weights, The time decay factor, This is the time difference since the last valid update. As a quality penalty factor, For the data sample size, This is a weight based on activity level.

8. A dynamic measurement device for the utility of heterogeneous data, characterized in that, include: The data receiving and identification module is used to receive heterogeneous source data in real time and identify the security classification labels of the heterogeneous source data. The environment routing module is used to route heterogeneous source data to the controlled computing environment corresponding to the security classification label based on the security classification label. The sample set construction module is used to build a consistent read-only view of the corresponding source data in the memory of a controlled computing environment and generate a benchmark sample set. The feature extraction module is used to perform dual-track feature decoupling extraction on the benchmark sample set, and to obtain the topological feature vector of the structure flow and the deep content feature vector of the content flow respectively. The dynamic fusion module is used to input the security classification label into the dynamic attention network, obtain the dynamic weight coefficients, and perform adaptive weighted fusion of the topological structure feature vector and the deep content feature vector according to the dynamic weight coefficients to generate a comprehensive data feature vector. The utility measurement module is used to obtain the time decay factor and quality penalty factor of the benchmark sample set, and calculate the final utility index based on the time decay factor, quality penalty factor and comprehensive data feature vector. When the final utility index meets the admission threshold, the data object admission and encapsulation flow are triggered.

9. Electronic devices, including: The device includes a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the memory via the bus, and the processor executes the machine-readable instructions to perform the steps of the heterogeneous data utility dynamic measurement method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, performs the steps of the dynamic measurement method for heterogeneous data utility as described in any one of claims 1 to 7.