Blood relationship path tracking system and method based on quantum coding and GPU acceleration
The lineage path tracing system, which uses quantum coding and GPU acceleration, solves the problem of inefficiency in lineage path tracing in cross-platform data interaction, and achieves efficient lineage relationship storage and fast query. It is suitable for scenarios with high requirements for system availability, such as finance and government affairs.
Patent Information
- Application Number
- CN202511111439.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-08
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2045-08-08
AI Technical Summary
Existing lineage path tracing methods are inefficient when processing large-scale data and are unable to cope with the real-time matching requirements of field changes and massive path data. Especially in cross-platform data interaction scenarios, traditional parsing algorithms are unable to meet the needs of real-time auditing and troubleshooting.
A lineage path tracing system based on quantum coding and GPU acceleration is adopted. Field mapping relationships are established through probabilistic superposition state representation, Huffman coding compression is performed, and a memory pooling strategy is used to manage data blocks. Combined with GPU resource allocation and parallel computing, lineage path feature matching and graph assembly are achieved.
It improves the storage and query efficiency of lineage path data, ensures the complete traceability of field-level lineage relationships across distributed systems, meets high-frequency and real-time lineage query requirements, and supports rapid positioning and access to target path data.
Smart Images

Figure CN120610944A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and more specifically, to a bloodline path tracing system and method based on quantum coding and GPU acceleration. Background Art
[0002] In the field of data governance and traceability analysis, lineage path tracking technology provides key support for data quality control, compliance auditing, and anomaly detection by recording the complete flow of data from generation to final application. With the surge in distributed systems and multi-source heterogeneous data, the complexity of lineage paths has grown exponentially. Traditional tracking solutions based on relational databases or graph databases face problems such as low processing efficiency and excessive storage costs. In particular, in scenarios involving cross-platform data interaction, the dynamic changes in field mapping relationships between different systems and the need for real-time matching of massive path data pose greater challenges to the performance and scalability of lineage tracking.
[0003] Existing lineage path tracing methods primarily rely on parsing algorithms, such as syntax tree analysis, which present significant bottlenecks when processing large amounts of data. On the one hand, traditional rule-based parsing algorithms struggle to effectively handle frequent field changes, resulting in low recognition accuracy for field mappings. On the other hand, they are inefficient when processing large or frequent lineage queries, making it difficult to meet the rapid response requirements of scenarios such as real-time auditing and immediate troubleshooting.
[0004] To address the above-mentioned problems, no effective solutions have been proposed so far. Summary of the Invention
[0005] The embodiments of the present application provide a lineage path tracing system and method based on quantum coding and GPU acceleration to solve the above technical problems.
[0006] This application provides a lineage path tracing system based on quantum coding and GPU acceleration, including: The log parsing module is used to parse the original execution plan log and abstract the cross-platform operations into a probabilistic superposition state representation including mapping, filtering, and connection operators; A lineage path set forming module is used to establish a field mapping relationship between different platforms based on the probabilistic superposition state representation to form a quantum state lineage path set; A compressed data block generation module is used to perform Huffman coding compression on the quantum state lineage path set to generate a compressed data block composed of binary codes corresponding to the operation sequence; A memory pool address generating module, configured to manage the compressed data block using a memory pooling strategy and output a memory pool address corresponding to the compressed data block; a target memory pool address locating module, configured to respond to a lineage query request received in real time and locate a target memory pool address based on a target path included in the lineage query request; A GPU resource allocation weight calculation module is used to calculate the GPU resource allocation weight according to the query priority, the size of the target compressed data block corresponding to the target memory pool address, and the video memory remaining rate; A path feature matching module is used to schedule the GPU to perform lineage path feature matching operations in parallel, including: comparing the target path with the encoding path in the target compressed data block, and screening a subset of encoding paths that meet the path features; The bloodline graph assembly module is used to assemble a bloodline graph in JSON format based on the filtered encoding path subset.
[0007] This application provides a lineage path tracing method based on quantum coding and GPU acceleration, including: Parse the original execution plan log and abstract the cross-platform operations into a probabilistic superposition representation including mapping, filtering, and joining operators; Based on the probabilistic superposition state representation, a field mapping relationship between different platforms is established to form a quantum state lineage path set; Performing Huffman coding compression on the quantum state lineage path set to generate a compressed data block consisting of binary codes corresponding to the operation sequence; Adopting a memory pooling strategy to manage the compressed data block, and outputting a memory pool address corresponding to the compressed data block; In response to a lineage query request received in real time, locating a target memory pool address based on a target path included in the lineage query request; Calculate the GPU resource allocation weight according to the query priority, the size of the target compressed data block corresponding to the target memory pool address, and the remaining rate of the video memory; Scheduling the GPU to perform lineage path feature matching operations in parallel, including: comparing the target path with the encoding path in the target compressed data block, and screening a subset of the encoding paths that meet the path features; The bloodline graph is assembled into JSON format based on the filtered encoding path subset.
[0008] Furthermore, after generating the JSON-formatted bloodline graph, an asynchronous checkpoint operation is performed: Backing up a snapshot of the quantum state encoding library of the quantum state lineage path set to the Iceberg column storage system every 5 minutes; When a failure occurs, the quantum state encoding library is restored based on the amount of data lost and the solid-state hard drive reading rate; Convert the recovered quantum state encoding library into a lineage matrix in Apache Arrow column format; In response to the received full blood relationship query request, the blood relationship matrix is output through the zero-copy interface.
[0009] Furthermore, ; in, is the probability superposition state representation; Represents the base state of the mapping operation, is the base state of the filtering operation, It is the base state of connection operation; is the probability weight of the mapping operation, ranging from 0.6 to 0.8; is the probability weight of the filtering operation, ranging from 0.1 to 0.3; is the connection operation probability weight, ranging from 0.05 to 0.15; The establishment of the field mapping relationship includes: based on 、 、 , calculates the mapping probability from Kafka source fields to Flink target fields.
[0010] Furthermore, the compression rate of the Huffman coding compression Calculated by the following formula:
[0011] in, is the total number of bloodline paths in the quantum state bloodline path set; is the index of the bloodline path; For the The encoding length of the lineage path in the compressed data block, in bits; For the The original length of the lineage path in the quantum state lineage path set, in bytes; For the The frequency of occurrence of the lineage path in the original execution plan log; The GPU resource allocation weight is determined based on the following formula: weight value = (query priority × target compressed data block size) / video memory remaining rate; wherein, the video memory remaining rate is obtained in real time through the GPU driver interface; the bloodline path feature matching operation adopts a video memory back pressure mechanism, and when the video memory remaining rate is lower than the set remaining rate threshold, new task scheduling is suspended.
[0012] Furthermore, the fault recovery time Calculated by the following formula:
[0013] in, The total amount of quantum state encoded data that was not persisted when the failure occurred, in GB; is the solid-state drive read rate, in GB / s; This is the inherent overhead time for system recovery, set to 10ms.
[0014] Furthermore, the bloodline path feature matching operation includes: Parsing the target path in the blood relationship query request to generate a target feature vector; the target feature vector includes a mapping operation probability weight, a filtering operation probability weight, and a connection operation probability weight of the target path; Divide the target compressed data block into multiple sub-blocks of equal length according to the number of GPU stream processors, and bind each GPU stream processor to one sub-block; Each GPU stream processor sequentially reads the Huffman coded bit stream of the bound sub-blocks and initializes the weight accumulation value of the candidate path when the bit pattern 1111 is recognized; When bit pattern 1100 is detected, the mapping operator is recorded and the accumulated mapping weight of the current candidate path is increased. When bit pattern 1101 is detected, the filtering operator is recorded and the accumulated filtering weight of the current candidate path is increased. When bit pattern 1110 is detected, the connection operator is recorded and the accumulated connection weight of the current candidate path is increased. When the bit pattern 1111 is recognized again, the absolute error sum between the cumulative weight vector of the candidate path and the target feature vector is calculated; if the absolute error sum is less than the set tolerance threshold, the storage offset of the candidate path in the target compressed data block is written into the global matching list; After all GPU stream processors have completed processing, the corresponding encoding paths are extracted according to the storage offsets in the global matching list and aggregated into the encoding path subset.
[0015] Furthermore, the method also includes an incremental path processing mechanism, including: Parse the newly added execution plan log into incremental lineage paths; Merging the incremental lineage path with the quantum state lineage path set to form an updated lineage path set; Recalculate the occurrence frequency of each path based on the updated bloodline path set; Reconstruct the Huffman coding tree according to the recalculated occurrence frequencies of each path; Using the reconstructed Huffman coding tree, a compression operation is performed on the updated lineage path set to generate an updated compressed data block; Release the storage space of the original compressed data block, store the updated compressed data block, and update the corresponding memory pool address mapping table.
[0016] Furthermore, the memory pooling strategy is implemented through data heat level storage, including: For each compressed data block, sort it in descending order according to the total number of occurrences of its lineage path in the original execution plan log; define the top 20% as high-heat blocks, the bottom 30% as low-heat blocks, and the middle 50% as medium-heat blocks; Establish a three-level storage architecture, including: a high-heat storage layer: resident in the GPU video memory, storing the high-heat blocks; a medium-heat storage layer: resident in the host memory, storing the medium-heat blocks; a low-heat storage layer: resident in the solid-state drive, storing the low-heat blocks; When receiving the lineage query request, if the target compressed data block corresponding to the target memory pool address is not in the high-heat storage layer, the current query thread is paused, the target compressed data block is copied from the current storage layer to the GPU memory, the memory pool address mapping table is updated and the query thread is restarted; Monitor GPU memory usage in real time; when GPU memory usage exceeds a set usage threshold, select N compressed data blocks with the earliest access time and the fewest total occurrences from the high-heat storage layer as migration objects, where N is dynamically determined based on the memory overlimit ratio; perform a hierarchical downgrade operation on each migration object; and synchronously update the memory pool address mapping table.
[0017] Furthermore, the element values of the blood relationship matrix Indicates the Kafka source field With Flink target fields The dependence strength between them is calculated as:
[0018] in, For the connection field and The set of all bloodline paths; is the total number of paths in the lineage path set; is an element in the bloodline path set; For path The sequence of operation types included; For sequence length; for The type of operation in For operation The composite weight function is defined as:
[0019] is the weight factor of the mapping operation type in the path; The weight factor for the filtering operation type in the path; is the weight factor of the connection operation type in the path; is the probability weight of the mapping operation, ranging from 0.6 to 0.8; is the probability weight of the filtering operation, ranging from 0.1 to 0.3; is the connection operation probability weight, ranging from 0.05 to 0.15.
[0020] Based on the embodiments provided in this application, the lineage path tracing method that combines quantum coding with GPU acceleration has significant advantages over traditional solutions. By abstracting cross-platform operations into probabilistic superposition state representations and establishing field mapping relationships based on quantum states, the lineage chain break problem caused by differences in operational semantics between different systems is effectively solved, ensuring that field-level lineage relationships across distributed systems are complete and traceable. At the same time, Huffman coding compression is performed on the quantum state lineage path set, and the compressed data blocks are managed in combination with a memory pooling strategy, which significantly reduces the storage overhead of the lineage path data, improves memory utilization efficiency, and supports rapid positioning and access to target path data.
[0021] Furthermore, GPU resources are dynamically allocated based on the available graphics memory, allowing for parallel execution of lineage path feature matching operations. This fully leverages the GPU's parallel computing capabilities, effectively improving the efficiency of matching massive lineage paths and meeting the needs of high-frequency, real-time lineage queries. Ultimately, the selected lineage path subset is assembled into a lineage graph in JSON format, achieving a standardized representation of lineage relationships. This facilitates seamless integration with downstream analysis tools and visualization systems, improving the overall effectiveness of data governance. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The drawings described herein are used to provide a further understanding of the embodiments of the present invention and constitute a part of this application. The illustrative embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation of this application. In the drawings: Figure 1 1 is a structural diagram of an optional lineage path tracing system based on quantum coding and GPU acceleration according to an embodiment of the present application; Figure 2 This is a flowchart of an optional lineage path tracing method based on quantum coding and GPU acceleration according to an embodiment of the present application; Figure 3 This is a flowchart of another optional bloodline path tracing method based on quantum coding and GPU acceleration according to an embodiment of the present application.
[0023] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION
[0024] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0025] Alternatively, as Figure 1 As shown, the present application provides a lineage path tracing system based on quantum coding and GPU acceleration, including: The log parsing module 101 is used to parse the original execution plan log and abstract the cross-platform operation into a probabilistic superposition state representation including mapping, filtering and connection operators; The lineage path set forming module 102 is used to establish a field mapping relationship between different platforms based on the probabilistic superposition state representation to form a quantum state lineage path set; The compressed data block generation module 103 is used to perform Huffman coding compression on the quantum state lineage path set to generate a compressed data block composed of binary codes corresponding to the operation sequence; A memory pool address generation module 104 is configured to manage compressed data blocks using a memory pooling strategy and output a memory pool address corresponding to the compressed data block; The target memory pool address locating module 105 is configured to respond to a lineage query request received in real time and locate the target memory pool address based on a target path included in the lineage query request; The GPU resource allocation weight calculation module 106 is used to calculate the GPU resource allocation weight according to the query priority, the size of the target compressed data block corresponding to the target memory pool address, and the remaining rate of the video memory; The path feature matching module 107 is used to schedule the GPU to perform lineage path feature matching operations in parallel, including: comparing the target path with the encoding path in the target compressed data block, and screening the subset of encoding paths that meet the path features; The bloodline map assembly module 108 is used to assemble a bloodline map in JSON format based on the filtered encoding path subset.
[0026] Alternatively, as Figure 2 As shown, the present application provides a lineage path tracing method based on quantum coding and GPU acceleration, including: S201, parsing the original execution plan log, abstracting the cross-platform operation into a probabilistic superposition state representation including mapping, filtering and connection operators; The original execution plan log, also known as the ETL execution plan log, stands for "Extract-Transform-Load Execution Plan" and is a core concept in the data processing field. This includes Spark, Flink, and Iceberg logs.
[0027] Specifically, Extract: extract the required data from the original data source (such as databases, files, sensors, logs, etc.); Transform: process the extracted data (such as cleaning error values, format conversion, calculation, filtering, merging, etc.) to make the data meet the requirements of the target system; Load: store the final transformed data into the target system (such as data warehouse, database, big data platform, etc.).
[0028] The ETL execution plan log is the specific arrangement and detailed planning of these three steps.
[0029] Among them, the mapping operator is used to perform field conversion (such as desensitizing the user ID to a hash value); the filtering operator is used to filter data according to conditions (such as retaining records with a status of "valid"); and the join operator is used to associate multi-source data (such as joining the order table and the user table by ID).
[0030] In S201, field conversions in cross-platform operations (e.g., Kafka to Flink) are decomposed into three base states: map, filter, and join. Their probability is quantified using probability weights (α, β, γ). Traditional rule-based matching cannot adapt to dynamically changing field mappings (e.g., a user ID may be mapped to a hash value or associated with an order on different platforms). Probabilistic superposition representation transforms discrete operations into a continuous probability space, resolving the problem of disconnected relationships caused by semantic differences across platforms. By modeling operational uncertainty using the concept of "superposition" in quantum mechanics, this allows for a probabilistic description of relationship relationships, providing a mathematical foundation for subsequent path matching.
[0031] S202, based on the probabilistic superposition state representation, establish a field mapping relationship between different platforms to form a quantum state lineage path set; In one embodiment, the process of establishing the field mapping relationship includes: Parsing operation logs: Read cross-platform execution plan logs and identify key operation records. For example, in the conversion log from Kafka to Flink, it was found that a user ID field was processed 1,000 times; 720 times it was converted to a user hash value (mapping operation); 210 times it was used for state filtering (filtering operation); and 70 times it was associated with an order ID (join operation). Calculate the probability weights: Statistical distribution ratio of the three types of operations: mapping operations account for 72%; filtering operations account for 21%; and joining operations account for 7%. The operation probability superposition state representation is formed: mapping operation weight is 0.72, filtering operation weight is 0.21, and joining operation weight is 0.07.
[0032] Constructing a quantum mapping: Establishing a probabilistic association between source and target fields: When the source system's user ID field changes, the target system's user hash field has a 72% probability of being affected, the status field has a 21% probability of being affected, and the order ID field has a 7% probability of being affected. This association reflects the characteristics of quantum entanglement: a single point of change triggers a probabilistic response in multiple targets.
[0033] The formation of a bloodline path set includes: Combine consecutive operations into complete data links, for example: Path 1: user ID → user hash → user dimension table; Path 2: user ID → status filter → active user table; Path 3: user ID → order association → order table Each path is labeled with a probability weight: Path 1 includes two operations: user ID conversion (weight 0.72) and dimension storage (weight 0.33), and its overall reliability is the product of the two operations, approximately 24%. Probabilistic entanglement is formed between paths: when path 1 fails, path 2 has a 21% probability of failing simultaneously. When a new field is converted (such as user geographic location), a new mapping relationship is automatically generated (user ID → geographic location, weight 0.05); the original mapping weight is dynamically adjusted (user hash weight is reduced from 0.72 to 0.68); a new lineage path is added: user ID → geographic location → regional analysis table.
[0034] In S202, field-level dependencies are established based on probability weights, and sequential operations are combined into complete paths (e.g., user ID → mapping → filtering → target table). A single operation cannot reflect the end-to-end lineage chain, but probabilistic entanglement between paths (e.g., the chain reaction of path failures) can improve the fault tolerance of lineage analysis. By incorporating quantum entanglement, the association probabilities of multiple target fields are automatically calculated when source fields change, enabling dynamic probabilistic deduction of lineage across systems.
[0035] S203, performing Huffman coding compression on the quantum state lineage path set to generate a compressed data block consisting of binary codes corresponding to the operation sequence; Among them, the operation sequence is the atomic component unit of the lineage path (such as a single Map operation); the lineage path is the complete data flow composed of the operation sequence (such as [source system→Map→Join→target system]).
[0036] The compression process includes: assigning short binary codes to high-frequency lineage paths (such as [Map→Filter] that appears 1000 times); assigning long codes to low-frequency paths (such as [Join→Map→Join] that appears once).
[0037] In S203, variable-length codes are assigned based on the frequency of path occurrence (short codes such as 1100 for high-frequency paths and long codes for low-frequency paths), converting the operation sequence into a binary bit stream. The original set of lineage paths is large (e.g., millions of paths), and direct storage would require too much memory. Huffman coding exploits the repetitive patterns of operators (e.g., frequent filtering operations) to compress the data volume. The coding table is designed based on the probability distribution characteristics of quantum state paths. High-frequency operators (e.g., mapping) are fixed to a short 4-bit code (1100), reducing parsing complexity and enabling the GPU to identify the operation type in a single cycle.
[0038] S204, adopting a memory pooling strategy to manage the compressed data block, and outputting a memory pool address corresponding to the compressed data block; In S204, a memory pool is used to centrally manage compressed blocks, outputting the memory pool address for fast addressing. Distributed, stored lineage data is difficult to locate efficiently. The memory pool strategy implements centralized address mapping, accelerating access to target data blocks. The address abstraction layer decouples physical storage locations, laying the foundation for subsequent rapid target path location (S205).
[0039] S205, in response to the lineage query request received in real time, locating the target memory pool address based on the target path included in the lineage query request; In S205, the memory pool address of the corresponding compressed block is located according to the target path in the lineage query request. Combining the memory pool address with the path characteristics, accurate target positioning is achieved, avoiding the inefficiency of full scanning.
[0040] S206, calculating the GPU resource allocation weight according to the query priority, the size of the target compressed data block corresponding to the target memory pool address, and the remaining rate of the video memory; In S206, GPU resources are dynamically allocated based on query priority, target compressed data block size, and available GPU memory. Highly concurrent queries must avoid GPU resource contention. The available GPU memory ratio is used to control task scheduling in real time. A memory-aware dynamic weighting mechanism prioritizes critical queries, ensuring immediate response for scenarios requiring high real-time performance, such as troubleshooting.
[0041] S207, scheduling the GPU to perform lineage path feature matching operations in parallel, including: comparing the target path with the encoding path in the target compressed data block, and selecting a subset of encoding paths that meet the path features; In S207, the GPU is scheduled to compare the target path in parallel with the encoded paths in the compressed block, selecting a subset that meets the characteristics (for example, identifying the bit pattern 1111 as a path boundary). Since serial matching of massive paths by the CPU is inefficient, the GPU stream processor parses the bitstream in parallel (for example, 1100 triggers the accumulation of mapping weights). Hard-coded logic is used to implement single-cycle operation recognition (for example, 4-bit pattern detection), and tolerance thresholds are combined to balance matching accuracy and flexibility.
[0042] S208, assembling a kinship map in JSON format based on the filtered encoding path subset.
[0043] In S208, the subset of matched encoding paths is converted to a standard JSON format. The JSON format supports seamless integration with downstream tools, such as visualization systems. This standardizes the output of lineage results and improves cross-platform compatibility.
[0044] The embodiments provided in this application address the broken-link problem inherent in traditional methods by abstracting and encoding complex cross-platform lineage relationships into quantum states. Simultaneously, the parallel computing capabilities of GPUs are leveraged to accelerate the matching and querying of lineage paths, breaking through CPU performance bottlenecks. This combination aims to significantly improve the integrity, speed, and efficiency of cross-system lineage tracing.
[0045] Furthermore, if Figure 3 As shown, after generating the JSON-formatted bloodline graph, an asynchronous checkpoint operation is performed: S301, backing up a snapshot of the quantum state encoding library of the quantum state lineage path set to the Iceberg column storage system every 5 minutes; Among them, Iceberg columnar is a tabular storage strategy for data lakes. It focuses on using columnar storage to manage large-scale offline data and solve the consistency and manageability issues of data lakes. In S301, Iceberg's transactional and versioning capabilities ensure the consistency and traceability of snapshot data, preventing data corruption during the backup process. Columnar storage adapts to the batch read characteristics of lineage paths, providing an efficient data source for fault recovery (S302).
[0046] S302, when a failure occurs, restoring the quantum state encoding library based on the amount of data loss and the solid-state drive reading rate; In S302, the code base is restored based on the amount of data lost and the SSD read rate. Traditional database recovery takes a long time. This step utilizes the high-speed I / O of the SSD to accelerate the recovery process.
[0047] S303, converting the recovered quantum state encoding library into a lineage relationship matrix in Apache Arrow column format; Among them, the Apache Arrow columnar format is a memory-level data exchange standard, focusing on unifying the memory data formats of different tools and reducing the loss of cross-system data transmission.
[0048] In S303, the recovered codebase is converted into a matrix in Apache Arrow columnar format. The Apache Arrow memory format optimizes large-scale matrix query efficiency. The columnar memory layout accelerates full-database lineage analysis (such as field dependency strength calculation). In S304, in response to a received full-database lineage query request, the lineage matrix is output via a zero-copy interface.
[0049] In S304, traditional serialization consumes CPU resources. In this step, memory data is shared across systems, serialization overhead is eliminated, and matrix data is directly output to the query end through zero-copy technology. This avoids duplicate copies of data in memory and compresses the full query response time to sub-seconds, meeting the efficiency requirements of scenarios such as batch auditing and historical path tracing. The standardized interface design supports direct docking with big data analysis frameworks (such as Spark) and expands the analysis dimensions of lineage data (such as path clustering and abnormal pattern mining). Furthermore, ; in, is a probabilistic superposition state representation; Represents the base state of the mapping operation, is the base state of the filtering operation, It is the base state of connection operation; is the probability weight of the mapping operation, ranging from 0.6 to 0.8; is the probability weight of the filtering operation, ranging from 0.1 to 0.3; is the connection operation probability weight, ranging from 0.05 to 0.15; The establishment of field mapping relationship includes: 、 、 , calculates the mapping probability from Kafka source fields to Flink target fields.
[0050] Based on the embodiments provided in this application, cross-platform operations are abstracted into probabilistic representations using quantum superposition formulas, and the probability of field mapping from Kafka to Flink is calculated based on α, β, and γ. Traditional solutions use deterministic rules to match field relationships, which makes it difficult to cope with the dynamic changes in cross-platform operation semantics (for example, the same field may be mapped, filtered, or connected in different systems). This solution, however, quantifies the operation possibilities through probability weights (α / β / γ), upgrading the field mapping relationship from a discrete judgment to a dynamic description in a continuous probability space. This design effectively captures the uncertainty of cross-system operations (for example, a Kafka source field may be mapped to a Flink target field with a 70% probability and participate in filtering with a 30% probability), avoids lineage breaks caused by the failure of a single rule, and ensures complete traceability of field-level lineage relationships in distributed scenarios. Furthermore, the probability weights provide a mathematical foundation for subsequent path matching, making the matching process compatible with subtle differences in operation sequences.
[0051] Furthermore, the compression rate of Huffman coding is Calculated by the following formula:
[0052] in, is the total number of bloodline paths in the quantum state bloodline path set; is the index of the bloodline path; For the The encoding length of the lineage path in the compressed data block, in bits; For the The original length of the lineage path in the quantum state lineage path set, in bytes; For the The frequency of occurrence of the lineage path in the original execution plan log; The GPU resource allocation weight is determined based on the following formula: Weight = (query priority × target compressed data block size) / video memory remaining ratio. The video memory remaining ratio is obtained in real time through the GPU driver interface. The lineage path feature matching operation uses a video memory backpressure mechanism. When the video memory remaining ratio falls below the set remaining ratio threshold, new task scheduling is suspended.
[0053] The remaining ratio threshold is a pre-set threshold for the remaining GPU memory ratio to prevent insufficient GPU memory from impacting lineage path feature matching operations. When the remaining GPU memory ratio (the ratio of remaining GPU memory to total GPU memory) obtained in real time through the GPU driver interface falls below this threshold, the memory backpressure mechanism is triggered, pausing the scheduling of new tasks to ensure that currently executing lineage path matching tasks (such as matching the target path with the encoded path in the compressed data block) are not affected by insufficient GPU memory.
[0054] For example, in a real-time audit scenario in a financial system, it is necessary to respond to concurrent lineage query requests at a high frequency (e.g., 100+ times per second), and the compressed data blocks involved in each query (containing molecular encoding paths) need to be quickly loaded into the GPU memory for matching. In this case, the "Set Remaining Rate Threshold" can be set to 30%: when the GPU memory remaining rate drops below 30%, the scheduling of new query tasks will be suspended. This is because under high concurrency, the continuous loading of data blocks by new tasks will quickly consume the memory. The 30% threshold can reserve enough memory for the current task to complete the encoding path comparison (such as screening the subset of encoding paths that meet the characteristics), avoiding task interruption or matching errors due to memory overflow.
[0055] For example, during nighttime batch processing of historical lineage paths, query concurrency is low, but the compressed data blocks involved in a single task are large (e.g., containing millions of quantum state path encodings). In this case, the "Set Remaining Rate Threshold" setting can be set to 15%. Due to the low concurrency and low frequency of new task scheduling, a 15% threshold ensures sufficient memory for the current batch task (e.g., Huffman coded bitstream parsing of historical paths) to complete the matching process, while also making better use of GPU memory and reducing resource idleness caused by prematurely pausing new tasks.
[0056] Based on the embodiments provided in this application, we clarify the calculation method for Huffman coding compression ratio (taking into account path length and frequency), as well as the formula for determining GPU resource allocation weights (combining query priority, data block size, and memory availability), and introduce a memory backpressure mechanism. Traditional Huffman coding compression ratio calculations fail to consider the impact of path frequency, which can lead to insufficient compression efficiency for high-frequency paths. This solution uses freq(i) weighting to achieve a better compression ratio for high-frequency paths (such as recurring mapping-filtering sequences), significantly reducing storage overhead. Regarding GPU resource allocation, traditional static allocation methods struggle to adapt to dynamic query loads. This solution dynamically adjusts weights based on real-time memory availability, ensuring that high-priority tasks (such as real-time auditing) and large data blocks receive priority computing resources, improving resource utilization. Furthermore, the memory backpressure mechanism suspends new task scheduling when memory is insufficient, preventing matching task crashes caused by memory overflow. This balances parallel computing efficiency and system stability, making it particularly suitable for high-frequency, high-traffic lineage query scenarios.
[0057] Furthermore, the fault recovery time Calculated by the following formula:
[0058] in, The total amount of quantum state encoded data that was not persisted when the failure occurred, in GB; is the solid-state drive (SSD) read rate, in GB / s; This is the inherent overhead time for system recovery, set to 10ms.
[0059] Based on the embodiments provided in this application, traditional fault recovery lacks a clear time quantification model, making it difficult to ensure availability in critical scenarios. This solution breaks down the recovery time into data reading time and system-inherent overhead, and combines the high-speed reading characteristics of SSDs (which significantly shorten data reading time compared to mechanical hard drives) and fixed latency (10ms) to make the recovery process predictable and controllable. In scenarios such as finance and government affairs that have extremely high requirements for system availability, this quantitative design can help operations and maintenance personnel assess the scope of fault impact in advance (for example, when 10GB of data is lost, the recovery time can be estimated based on the SSD rate), and proactively reduce recovery time by optimizing SSD performance or reducing the amount of non-persistent data (such as shortening the snapshot interval), ensuring the rapid recovery of lineage tracking services and reducing business interruption losses.
[0060] Furthermore, the bloodline path feature matching operation includes: Parse the target path in the lineage query request and generate a target feature vector; the target feature vector includes the mapping operation probability weight, the filtering operation probability weight, and the connection operation probability weight of the target path; Divide the target compressed data block into multiple sub-blocks of equal length according to the number of GPU stream processors, and bind one sub-block to each GPU stream processor; Each GPU stream processor sequentially reads the Huffman coded bit stream of the bound sub-blocks and initializes the weight accumulation value of the candidate path when the bit pattern 1111 is recognized; When bit pattern 1100 is detected, the mapping operator is recorded and the accumulated mapping weight of the current candidate path is increased. When bit pattern 1101 is detected, the filtering operator is recorded and the accumulated filtering weight of the current candidate path is increased. When bit pattern 1110 is detected, the connection operator is recorded and the accumulated connection weight of the current candidate path is increased. When the bit pattern 1111 is recognized again, the absolute error sum between the cumulative weight vector of the candidate path and the target feature vector is calculated; if the absolute error sum is less than the set tolerance threshold, the storage offset of the candidate path in the target compressed data block is written into the global matching list; Wherein, cumulative weight vector = [mapping weight cumulative value, filtering weight cumulative value, connection weight cumulative value]; The tolerance threshold refers to the maximum absolute error and critical value allowed between the pre-set target feature vector and the cumulative weight vector of the candidate path during the lineage path feature matching process. When the sum of the absolute errors is less than the threshold, the candidate path's operational characteristics (mapping, filtering, and probability weight distribution of connections) are considered consistent with the target path and are included in the encoded path subset for subsequent assembly of the JSON lineage graph. Its core function is to balance path matching accuracy and flexibility, adapting to the stringency requirements of lineage path feature matching in different scenarios.
[0061] For example, in a bank data compliance audit, it is necessary to accurately track the lineage path of a transaction data (such as the operation sequence of "Kafka source field → Flink filtering → Snowflake connection"). The weights of the target feature vector are α = 0.7 (mapping), β = 0.2 (filtering), and γ = 0.1 (connection). In this case, the "Set tolerance threshold" can be set to 0.1: A match is considered when the sum of the absolute errors between the cumulative weight vector of a candidate path (e.g., α = 0.68, β = 0.21, γ = 0.11) and the target vector (|0.68 - 0.7| + |0.21 - 0.2| + |0.11 - 0.1| = 0.01 + 0.01 + 0.01 = 0.03) is less than 0.1. This threshold ensures that only candidate paths with an operation weight distribution that closely approximates the target path are selected, meeting the path accuracy requirements for compliance audits.
[0062] For example, in e-commerce platform data exploration, it is necessary to quickly locate various lineage paths related to "user behavior data → order data" (which may include mapping and filtering operations of varying proportions). The weights of the target feature vector are α=0.65, β=0.25, and γ=0.1. In this case, the "Set Tolerance Threshold" can be set to 0.3: A match is considered when the sum of the absolute errors between the cumulative weight vector of a candidate path (e.g., α = 0.5, β = 0.35, γ = 0.15) and the target vector (|0.5 - 0.65| + |0.35 - 0.25| + |0.15 - 0.1| = 0.15 + 0.1 + 0.05 = 0.3) equals 0.3. This threshold allows for greater fluctuation in the weights, screening out more relevant lineage paths and meeting the need for rapid discovery of potential connections in data exploration.
[0063] In this embodiment, the Huffman coding table design rules include: Operator code assignment rules: The mapping operator is assigned a 4-bit fixed code 1100, and the code value corresponds to the mapping operation base state probability weight defined by weight 3; the filtering operator is assigned a 4-bit fixed code 1101, and the code value corresponds to the filtering operation base state probability weight defined by weight 3; the connection operator is assigned a 4-bit fixed code 1110, and the code value corresponds to the connection operation base state probability weight defined by weight 3.
[0064] Path boundary control code design: The path start flag is assigned a 4-bit fixed code 1111, and its physical meaning is to initialize a new lineage path record; the path end flag reuses the 4-bit fixed code 1111, and its function is distinguished by the context position: when it is located before the operator sequence, it indicates the beginning of the path; when it is located after the operator sequence, it indicates the end of the path.
[0065] Coding space isolation mechanism: The Hamming distance between all operator codes (1100 / 1101 / 1110) and the control code (1111) is not less than 2; operator codes are prohibited from containing two consecutive 3-bit patterns starting with 1 to prevent confusion with control codes.
[0066] Weight value embedding method: When an operator encoding is detected, a fixed weight value is directly associated: 1100 triggers the accumulation of the mapping operation weight α; 1101 triggers the accumulation of the filtering operation weight β; 1110 triggers the accumulation of the connection operation weight γ. The weight value is preset in the stream processor register in the form of a constant and does not participate in the encoding transmission; Ensuring codec consistency: The compression phase generates codes according to the same rules; the matching phase uses hard-coded logic gates to achieve real-time bit pattern detection. In some embodiments of the present application, the binary code generated by Huffman coding compression includes two types of coding units: Operator encoding unit: fixed 4-bit length, where the mapping operator is encoded as 1100, the filtering operator is encoded as 1101, and the connection operator is encoded as 1110; Path boundary unit: fixed 4-bit length, coded as 1111. When this code appears at the beginning of a path, it indicates the initialization of a new path. When it appears at the end of a path, it triggers a path matching decision. The operator encoding unit and the path boundary unit are arranged in the order of the operation sequence of the lineage path to form a complete compressed data block.
[0067] The above encoding rules can reduce the false detection rate. Specifically, the minimum Hamming distance between the operator code (1100 / 1101 / 1110) and the path marker (1111) is 2, ensuring that the GPU stream processor can recognize it in a single cycle. Decoding efficiency is improved. Specifically, a fixed 4-bit length eliminates the bit operation overhead of traditional Huffman variable-length codes. Weight consistency is maintained. Specifically, the operator code is directly bound to a fixed weight constant to avoid numerical deviations introduced during the encoding and decoding process. After all GPU stream processors have finished processing, the corresponding encoding paths are extracted according to the storage offset in the global matching list and aggregated into encoding path subsets.
[0068] Based on the embodiments provided in this application, traditional CPU serial matching struggles to meet the real-time matching requirements for massive paths. This solution overcomes this bottleneck through the following design considerations: First, the target path is converted into a feature vector containing operation weights, upgrading the matching from string comparison to vector space computation, improving semantic matching accuracy; second, sharding processing based on the number of GPU stream processors enables parallel parsing, fully leveraging the GPU's parallel computing capabilities; third, operators are quickly identified through fixed bit patterns (e.g., 1100 corresponds to a mapping operation), reducing decoding time; and fourth, by comparing absolute error with a tolerance threshold, this approach ensures matching accuracy while accommodating subtle path differences (e.g., weight fluctuations). These design considerations enable GPUs to efficiently handle parallel matching of millions of paths, making them particularly suitable for scenarios requiring high real-time performance (e.g., immediate troubleshooting) while maintaining both flexibility and accuracy.
[0069] Furthermore, the method also includes an incremental path processing mechanism, including: Parse the newly added execution plan log into incremental lineage paths; Merge the incremental lineage path with the quantum lineage path set to form an updated lineage path set; Recalculate the occurrence frequency of each path based on the updated bloodline path set; Reconstruct the Huffman coding tree according to the recalculated occurrence frequencies of each path; Using the reconstructed Huffman coding tree, a compression operation is performed on the updated lineage path set to generate an updated compressed data block; Release the storage space of the original compressed data block, store the updated compressed data block, and update the corresponding memory pool address mapping table.
[0070] Based on the embodiments provided in this application, the traditional full reprocessing mechanism needs to repeatedly encode all data when adding a new path, which is inefficient. This solution uses an incremental update strategy to only merge the newly added paths, recalculate the frequency of occurrence and partially rebuild the Huffman tree, avoiding the waste of resources for full encoding. In real-time data flow scenarios (such as e-commerce platforms that generate a large number of new paths per second), this design can ensure that the system quickly adapts to path changes and maintains the compression efficiency of Huffman coding (high-frequency new paths automatically obtain short codes). At the same time, by releasing the storage space of the original data block and updating the address mapping, storage redundancy is avoided, and the dynamic scalability of the system is improved. It is suitable for business scenarios where new bloodline paths are continuously generated.
[0071] Furthermore, the memory pooling strategy is implemented through hierarchical storage of data heat, including: For each compressed data block, sort it in descending order based on the total number of occurrences of its lineage paths in the original execution plan log. Define the top 20% as high-heat blocks, the bottom 30% as low-heat blocks, and the middle 50% as medium-heat blocks. Establish a three-level storage architecture, including: high-heat storage layer: resides in GPU video memory and stores high-heat blocks; medium-heat storage layer: resides in host memory and stores medium-heat blocks; low-heat storage layer: resides in solid-state drives and stores low-heat blocks; When a lineage query request is received, if the target compressed data block corresponding to the target memory pool address is not in the high-heat storage layer, the current query thread is paused, the target compressed data block is copied from the current storage layer to the GPU memory, the memory pool address mapping table is updated, and the query thread is restarted; Monitor GPU memory usage in real time; when GPU memory usage exceeds a set usage threshold, select N compressed data blocks with the earliest access time and the fewest total occurrences from the high-hot storage tier as migration targets, where N is dynamically determined based on the memory overage ratio; perform a graded downgrade operation on each migration target; and synchronously update the memory pool address mapping table; N is a positive integer. In some embodiments, the graded downgrade operation includes: migrating data blocks ranked in the top 5 percent by total occurrence to a medium-hot storage tier; and migrating all other data blocks to a low-hot storage tier.
[0072] Based on the embodiment provided in this application, a three-level storage architecture is implemented based on data heat (high heat resides in GPU memory, medium heat resides in host memory, and low heat resides in SSD), and dynamic migration is supported. Traditional memory management does not distinguish between data heat, resulting in high-frequency access data and low-frequency data competing for limited memory resources, reducing access efficiency. This solution divides the heat level according to the number of path appearances, and stores high-frequency access high-heat blocks directly in GPU memory (with the lowest access latency), medium-frequency data in host memory, and low-frequency data in SSD, thereby achieving hierarchical optimization of storage resources. When the query involves low-frequency data, it is dynamically migrated to the memory and the address mapping is updated to ensure fast access to high-frequency data; when the memory is insufficient, the earliest accessed and least frequently accessed data blocks are downgraded to release resources. This dynamic grading mechanism balances storage cost and access efficiency, and is particularly suitable for scenarios where the access frequency of bloodline paths varies greatly (such as some core paths are frequently queried, while historical paths are rarely accessed), improving the resource utilization of the overall system.
[0073] Furthermore, the element values of the blood relationship matrix Indicates the Kafka source field With Flink target fields The dependence strength between them is calculated as:
[0074] in, For the connection field and The set of all bloodline paths; is the total number of paths in the lineage path set; is an element in the bloodline path set; For path The sequence of operation types included; For sequence The length of , i.e. the number of operation steps; for The type of operation in For operation The composite weight function is defined as:
[0075] is the weight factor of the mapping operation type in the path; The weight factor for the filtering operation type in the path; is the weight factor of the connection operation type in the path; is the probability weight of the mapping operation, ranging from 0.6 to 0.8; is the probability weight of the filtering operation, ranging from 0.1 to 0.3; is the connection operation probability weight, ranging from 0.05 to 0.15.
[0076] Based on the embodiments provided in this application, traditional blood relationships are mostly represented by linked lists or graph structures, which makes it difficult to quantify the degree of dependence between fields. The calculation quantifies the dependency between the Kafka source field and the Flink target field into a specific value, which is a composite weight of all operations in the path ( Combine the operation probability weights α / β / γ with the type weight factor) and take the average of multiple paths to ensure the statistical significance of the results. This quantization matrix can not only intuitively reflect the dependency strength between fields (such as =0.8 indicates strong dependency), and also supports rapid output to downstream analysis tools (such as Apache Arrow format-compatible data visualization platforms) through a zero-copy interface, providing a quantitative basis for data governance (such as identifying strongly dependent fields to focus on monitoring data quality), and improving the depth and operability of lineage analysis.
[0077] It should be noted that in this application, the embodiments implemented on the lineage path tracing system side based on quantum coding and GPU acceleration can be referenced with the embodiments implemented on the lineage path tracing method side based on quantum coding and GPU acceleration, and this application will not go into details one by one.
[0078] The above are only preferred embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention description and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A lineage path tracing system based on quantum coding and GPU acceleration, characterized by: include: The log parsing module is used to parse the original execution plan log and abstract the cross-platform operations into a probabilistic superposition state representation including mapping, filtering, and connection operators; A lineage path set forming module is used to establish a field mapping relationship between different platforms based on the probabilistic superposition state representation to form a quantum state lineage path set; A compressed data block generation module is used to perform Huffman coding compression on the quantum state lineage path set to generate a compressed data block composed of binary codes corresponding to the operation sequence; A memory pool address generating module, configured to manage the compressed data block using a memory pooling strategy and output a memory pool address corresponding to the compressed data block; a target memory pool address locating module, configured to respond to a lineage query request received in real time and locate a target memory pool address based on a target path included in the lineage query request; A GPU resource allocation weight calculation module is used to calculate the GPU resource allocation weight according to the query priority, the size of the target compressed data block corresponding to the target memory pool address, and the video memory remaining rate; A path feature matching module is used to schedule the GPU to perform lineage path feature matching operations in parallel, including: comparing the target path with the encoding path in the target compressed data block, and screening a subset of encoding paths that meet the path features; The bloodline graph assembly module is used to assemble a bloodline graph in JSON format based on the filtered encoding path subset.
2. A lineage path tracing method based on quantum coding and GPU acceleration, characterized in that: include: Parse the original execution plan log and abstract the cross-platform operations into a probabilistic superposition representation including mapping, filtering, and joining operators; Based on the probabilistic superposition state representation, a field mapping relationship between different platforms is established to form a quantum state lineage path set; Performing Huffman coding compression on the quantum state lineage path set to generate a compressed data block consisting of binary codes corresponding to the operation sequence; Adopting a memory pooling strategy to manage the compressed data block, and outputting a memory pool address corresponding to the compressed data block; In response to a lineage query request received in real time, locating a target memory pool address based on a target path included in the lineage query request; Calculate the GPU resource allocation weight according to the query priority, the size of the target compressed data block corresponding to the target memory pool address, and the remaining rate of the video memory; Scheduling the GPU to perform lineage path feature matching operations in parallel, including: comparing the target path with the encoding path in the target compressed data block, and screening a subset of the encoding paths that meet the path features; The bloodline graph is assembled into JSON format based on the filtered encoding path subset.
3. The lineage path tracing method based on quantum coding and GPU acceleration according to claim 2 is characterized in that: After generating the JSON-formatted lineage graph, perform an asynchronous checkpoint operation: Backing up a snapshot of the quantum state encoding library of the quantum state lineage path set to the Iceberg column storage system every 5 minutes; When a failure occurs, the quantum state encoding library is restored based on the amount of data lost and the solid-state hard drive reading rate; Convert the recovered quantum state encoding library into a lineage matrix in Apache Arrow column format; In response to the received full blood relationship query request, the blood relationship matrix is output through the zero-copy interface.
4. The lineage path tracing method based on quantum coding and GPU acceleration according to claim 2 is characterized in that: ; in, is the probability superposition state representation; Represents the base state of the mapping operation, is the base state of the filtering operation, It is the base state of connection operation; is the probability weight of the mapping operation, ranging from 0.6 to 0.8; is the probability weight of the filtering operation, ranging from 0.1 to 0.3; is the connection operation probability weight, ranging from 0.05 to 0.15; The establishment of the field mapping relationship includes: based on 、 、 , calculates the mapping probability from Kafka source fields to Flink target fields.
5. The lineage path tracing method based on quantum coding and GPU acceleration according to claim 2 is characterized in that: The compression rate of the Huffman coding compression Calculated by the following formula: in, is the total number of bloodline paths in the quantum state bloodline path set; is the index of the bloodline path; For the The encoding length of the lineage path in the compressed data block, in bits; For the The original length of the lineage path in the quantum state lineage path set, in bytes; For the The frequency of occurrence of the lineage path in the original execution plan log; The GPU resource allocation weight is determined based on the following formula: weight value = (query priority × target compressed data block size) / video memory remaining rate; wherein, the video memory remaining rate is obtained in real time through the GPU driver interface; the bloodline path feature matching operation adopts a video memory back pressure mechanism, and when the video memory remaining rate is lower than the set remaining rate threshold, new task scheduling is suspended.
6. The lineage path tracing method based on quantum coding and GPU acceleration according to claim 3 is characterized in that: Fault recovery time Calculated by the following formula: in, The total amount of quantum state encoded data that was not persisted when the failure occurred, in GB; is the solid-state drive read rate, in GB / s; This is the inherent overhead time for system recovery, set to 10ms.
7. The lineage path tracing method based on quantum coding and GPU acceleration according to claim 2 is characterized in that: The bloodline path feature matching operation includes: Parsing the target path in the blood relationship query request to generate a target feature vector; the target feature vector includes a mapping operation probability weight, a filtering operation probability weight, and a connection operation probability weight of the target path; Divide the target compressed data block into multiple sub-blocks of equal length according to the number of GPU stream processors, and bind each GPU stream processor to one sub-block; Each GPU stream processor sequentially reads the Huffman coded bit stream of the bound sub-blocks and initializes the weight accumulation value of the candidate path when the bit pattern 1111 is recognized; When bit pattern 1100 is detected, the mapping operator is recorded and the accumulated mapping weight of the current candidate path is increased. When bit pattern 1101 is detected, the filtering operator is recorded and the accumulated filtering weight of the current candidate path is increased. When bit pattern 1110 is detected, the connection operator is recorded and the accumulated connection weight of the current candidate path is increased. When the bit pattern 1111 is recognized again, the absolute error sum between the cumulative weight vector of the candidate path and the target feature vector is calculated; if the absolute error sum is less than the set tolerance threshold, the storage offset of the candidate path in the target compressed data block is written into the global matching list; After all GPU stream processors have completed processing, the corresponding encoding paths are extracted according to the storage offsets in the global matching list and aggregated into the encoding path subset.
8. The lineage path tracing method based on quantum coding and GPU acceleration according to claim 2 is characterized in that: The method also includes an incremental path processing mechanism, including: Parse the newly added execution plan log into incremental lineage paths; Merging the incremental lineage path with the quantum state lineage path set to form an updated lineage path set; Recalculate the occurrence frequency of each path based on the updated bloodline path set; Reconstruct the Huffman coding tree according to the recalculated occurrence frequencies of each path; Using the reconstructed Huffman coding tree, a compression operation is performed on the updated lineage path set to generate an updated compressed data block; Release the storage space of the original compressed data block, store the updated compressed data block, and update the corresponding memory pool address mapping table.
9. The lineage path tracing method based on quantum coding and GPU acceleration according to claim 2, characterized in that: The memory pooling strategy is implemented through data heat level storage, including: For each compressed data block, sort it in descending order according to the total number of occurrences of its lineage path in the original execution plan log; define the top 20% as high-heat blocks, the bottom 30% as low-heat blocks, and the middle 50% as medium-heat blocks; Establish a three-level storage architecture, including: a high-heat storage layer: resident in the GPU video memory, storing the high-heat blocks; a medium-heat storage layer: resident in the host memory, storing the medium-heat blocks; a low-heat storage layer: resident in the solid-state drive, storing the low-heat blocks; When receiving the lineage query request, if the target compressed data block corresponding to the target memory pool address is not in the high-heat storage layer, the current query thread is paused, the target compressed data block is copied from the current storage layer to the GPU memory, the memory pool address mapping table is updated and the query thread is restarted; Real-time monitoring of GPU memory usage; when the GPU memory usage exceeds a set usage threshold, selecting N compressed data blocks with the earliest access time and the least total number of occurrences from the high-heat storage layer as migration objects, where N is dynamically determined based on the memory overlimit ratio; performing a hierarchical downgrade operation on each migration object; and synchronously updating the memory pool address mapping table.
10. The lineage path tracing method based on quantum coding and GPU acceleration according to claim 3, characterized in that: The element value of the blood relationship matrix Indicates the Kafka source field With Flink target fields The dependence strength between them is calculated as: in, For the connection field and The set of all bloodline paths; is the total number of paths in the lineage path set; is an element in the bloodline path set; For path The sequence of operation types included; For sequence length; for The type of operation in For operation The composite weight function is defined as: is the weight factor of the mapping operation type in the path; The weight factor for the filtering operation type in the path; is the weight factor of the connection operation type in the path; is the probability weight of the mapping operation, ranging from 0.6 to 0.8; is the probability weight of the filtering operation, ranging from 0.1 to 0.3; is the connection operation probability weight, ranging from 0.05 to 0.15.
Citation Information
Patent Citations
Data quality monitoring system based on traceability analysis technology
CN111723082A
Multi-link service data traceability technology
CN114491045A
Parameter optimization method and system based on genetic breeding method
CN114492165A
Blood relationship analysis method and system for new energy data
CN115204163A
Combination optimization problem solving method and related device
CN118211668A
Cited By
Large model training resource optimization method and system based on Hadoop ecology
CN121070632A