Single-cell data format conversion method and system based on intermediate representation

CN122842718APending Publication Date: 2026-09-29NANJING STOMATOLOGICAL HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611339655.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-09-01
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0005]本发明的一个目的在于提出基于中间表示的单细胞数据格式转换方法及系统,针对现有技术中异构单细胞对象的逻辑轴、稀疏布局、数据层角色、元数据类型和对象结构难以一致映射,超大对象转换容易发生语义错位、静默丢失、索引溢出、峰值内存过高及故障后全量重做的问题,提出了以语义约束清单和稳定逻辑令牌建立格式无关中间表示,依据稀疏结构与内存上限可变分块,为源块生成布局无关语义摘要,受语义约束选择转置同构三数组复用或内存受限重排,写入目标临时块后回读生成同规则摘要,验证通过方可原子登记账本,并按摘要差异和模式依赖关系确定待重放集合的技术方案,本发明具备在目标对象发布前形成逐块语义验证闭环并减少非零元素重排、峰值内存占用和故障恢复量的技术效果

Benefits of technology

[0038]1、通过语义约束清单、稳定逻辑令牌和布局无关语义摘要,将物理轴方向及稀疏存储差异与细胞、基因、数据层和对象引用的逻辑语义分离,使目标临时块能够按与源块相同的规则回读验证,从而在发布前识别轴、注释、类型、数据层角色和引用关系错误。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122842718A_ABST
    Figure CN122842718A_ABST
Patent Text Reader

Abstract

The application discloses a single-cell data format conversion method and system based on an intermediate representation, and belongs to the field of biological information data processing. In order to solve the problems of mispositioning of axes and annotations, loss of semantic fields, index overflow, excessively high memory peak value and full rework after interruption in the conversion of heterogeneous single-cell objects, the application constructs a semantic constraint list and a stable logic token, and divides blocks according to variable upper limit of memory, so as to control transpose isomorphism reuse or memory limited rearrangement by layout-independent semantic abstract, and to realize atomic accounting after read-back verification and replay according to differences, thereby achieving the technical effects of detecting semantic errors before the target object is published and reducing rearrangement, memory and recovery overhead.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of bioinformatics data processing, and in particular to a method and system for converting single-cell data formats based on intermediate representations. Background Technology

[0002] Single-cell sequencing and spatial omics data typically consist of expression matrices, cell annotations, gene annotations, data layers, dimensionality reduction results, graph structures, spatial data, and object references. Different software ecosystems employ different cell and gene axis orientations, compressed sparse row or column layouts, data layer or analysis object roles, metadata types, and RDS and S4 object slot structures. Existing conversion tools often determine conversion success based on file openability, matrix shape, or offline round-trip results after conversion, making it difficult to verify that the logical coordinates, annotations, types, data layer roles, and reference relationships actually written to the target object remain consistent.

[0003] When a single-cell object contains a large number of cells, genes, and non-zero elements, rearranging the sparse matrix all at once can cause a surge in peak memory usage, and differences in numerical values ​​or index widths may also cause overflows. If the transformation process is interrupted, rules change, or some target blocks are corrupted, the lack of block-level verification commit status will require a complete re-execution of the transformation. Furthermore, when fields, slots, or references are silently lost in the format mapping, structural checks alone often cannot locate the difference paths and affected objects before the target data becomes publicly visible.

[0004] Therefore, there is a need for a single-cell data format conversion method and system based on intermediate representation that can overcome the shortcomings of the existing technology. Summary of the Invention

[0005] One objective of this invention is to propose a method and system for single-cell data format conversion based on intermediate representation. Addressing the challenges of consistent mapping between the logical axes, sparse layouts, data layer roles, metadata types, and object structures of heterogeneous single-cell objects in existing technologies, and the potential for semantic misalignment, silent data loss, index overflow, excessive peak memory usage, and full rework after failures in the conversion of extremely large objects, this invention proposes a format-independent intermediate representation established using a semantic constraint list and stable logical tokens. Based on the sparse structure and variable memory limit, it generates layout-independent semantic summaries for the source blocks. Subject to semantic constraints, it selects between transposed isomorphic three-array reuse or memory-constrained rearrangement. After writing to the target temporary block, it reads back to generate a summary with the same rules. Only after successful verification can the ledger be atomically registered. Finally, it determines the set to be replayed based on summary differences and pattern dependencies. This invention achieves the technical effect of forming a block-by-block semantic verification closed loop before the target object is published, and reduces non-zero element rearrangement, peak memory usage, and failure recovery.

[0006] This invention provides a method for converting single-cell data formats based on intermediate representation, comprising:

[0007] S1. Parse the format version, logical axis, sparse layout, data type, data layer and object reference relationship of the source single cell object, generate a list of semantic constraints, and generate stable logical tokens for cells, genes, data layers and referenced objects.

[0008] S2. Determine the variable block boundaries based on the semantic constraint list, sparse matrix pointer array, numerical width, index width, buffer usage, and memory limit, and generate a block conversion plan;

[0009] S3. Read the source block according to the block conversion plan, generate a source block semantic summary that is independent of sparse layout based on the stable logical token and type normalization value, select transpose isomorphic three-array reuse or memory-constrained rearrangement according to the logical axis and physical storage direction of the source object and the target object, and write the conversion result into the target temporary block.

[0010] S4. Read back the target temporary block and generate the target block semantic summary according to the rules for generating the source block semantic summary. Compare the semantic summaries and associated semantic states of the source block and the target block according to the semantic constraint list. When the verification is successful, register the corresponding block atomically to the committed block ledger.

[0011] S5. Verify the submitted block ledger, determine the set to be replayed based on unregistered blocks, digest differences and schema dependencies, perform transformation and verification on the set to be replayed, until all blocks are registered to the submitted block ledger and then submit the target root object.

[0012] Optionally, S1 includes: taking the source single-cell object as input, and extracting the format version and object topology from the file header, object class identifier, slot description and reference entry of the source single-cell object;

[0013] The logical axes corresponding to each physical dimension are determined based on cell identifier sequences and gene identifier sequences, and the data layer names are mapped to data layer roles.

[0014] The semantic constraint list records logical axis direction, sparse layout, data layer roles, metadata types, format version compatibility relationships, and object reference closures.

[0015] The stable logical tokens for cells and genes are generated by the object namespace, logical axis role, and in-axis identifier. The stable logical tokens for the data layer are generated by the object namespace, data layer role, and data layer identifier. The stable logical tokens for referenced objects are generated by the object namespace, object category, object identifier, and reference path.

[0016] Furthermore, the semantic constraint list also records the field name, data type, missing value representation, category value set and correspondence with stable logical token of the axis metadata field, and records the object class, slot name, slot type and slot reference target of RDS object and S4 object;

[0017] When converting axis metadata and object slots, records are aligned according to stable logical tokens. Fields, slots, or reference targets that are not mapped in the semantic constraint list are marked as validation failures and their corresponding blocks are not registered.

[0018] Optionally, S2 includes: accumulating the number of non-zero elements in adjacent logical axis intervals along the sparse matrix pointer array, calculating the storage amount of non-zero elements according to the numerical width and index width, and superimposing the buffer occupancy required for pointer array, sorting, rearrangement, digest calculation and target writing, and forming a block when the estimated occupancy does not exceed the memory limit;

[0019] When the estimated occupancy of a logical axis interval that cannot be further divided along the logical axis boundary exceeds the memory limit, it is further divided into sub-blocks along the non-zero element sequence. When the target format does not allow the sub-blocks to be written independently, the corresponding logical axis interval is marked as a conversion failure.

[0020] When the upper bound of the global index value exceeds the upper limit of the target index type, if the target format allows intra-block indexing, the global index value is mapped to an intra-block index value and the block offset is recorded. If the target format does not allow intra-block indexing, an index type whose representation range covers the upper bound of the global index value is selected. If there is no index type that satisfies the representation range, it is marked as a verification failure.

[0021] Optionally, S3 includes: generating semantic digests includes: generating leaf digests using stable logical tokens corresponding to logical coordinates, data layer roles, type identifiers, and type normalization values; merging leaf digests into block digests according to a predetermined logical token order; and forming a hierarchical semantic digest tree level by level.

[0022] The hierarchical semantic summary tree is set up with sub-summary nodes for the data layer, axis data, dimensionality reduction data, graph data, spatial data and object references. The parent node is generated by its child node identifier and child node summary in a predetermined node order.

[0023] Furthermore, the selection of transpose isomorphic three-array reuse includes: when the physical storage directions of the source sparse matrix and the target sparse matrix are transpose correspondences, verifying the structural correspondence of the pointer array, index array and value array, and verifying the logical axis fingerprint, the merging rules of repeated logical coordinates, the explicit zero retention rules, the value width, the index width and the data layer role.

[0024] When all verification results satisfy the semantic constraint list, the three arrays are transposed and reused isomorphically; otherwise, the non-zero elements are rearranged in segments within the memory limit according to the stable logic token.

[0025] Both conversion results are then read back for verification in step S4.

[0026] Optionally, S4 includes: reading back the logical axis identifier from the target temporary block, the logical coordinates of the non-zero element, the data layer role, the metadata type, the format version, and the object reference identifier;

[0027] Perform the same type normalization rule as the source block on the readback data, and arrange the normalization results according to the stable logical token and logical coordinate order to generate the root digest and sub-digest of the target block;

[0028] Based on the semantic constraint list and the object categories contained in the current block, select the corresponding sub-item summary from the hierarchical semantic summary tree as the necessary sub-item summary;

[0029] When the root digest, the necessary sub-item digests, the logical axis fingerprint, the data layer role, the metadata type, the format version, and the object reference closure are all consistent, a ledger record containing the belonging object identifier, block identifier, source block digest, target block digest, digest rule version, schema dependency fingerprint, and commit status is generated, and the committed block ledger is atomically updated by replacing the formal record with a temporary record.

[0030] Optionally, S5 includes: verifying the readability of the target block, the target block digest, and the schema dependency fingerprint of the ledger records in the committed block ledger;

[0031] Blocks with missing ledger records, blocks with ledger records but unreadable target blocks, abnormal blocks with inconsistent target block digests, and blocks whose pattern-dependent fingerprint correspondence rules have changed are added to the abnormal recovery replay set.

[0032] The field-level, data-level, or block-level replay granularity is determined based on the lowest-level node that produces the difference in the hierarchical semantic summary tree, and the target objects that depend on the difference node are added to the anomaly recovery replay set along the object reference relationship.

[0033] Furthermore, the target root object commit includes: establishing a mapping between object identifiers and corresponding block identifier sets based on the object identifiers in the ledger records; when the ledger records of each target temporary block are in a committed state and all block identifiers corresponding to each referenced object in the object reference closure have committed ledger records, a target root object list is generated. The target root object list records the target object identifier, format version, logical axis fingerprint, data layer role set, object root summary, and referenced object identifier. The object root summary is generated by the target block summaries in all ledger records of the corresponding object in the order of block identifiers.

[0034] The target root object list and the target temporary block set are published as target objects through atomic substitution. If the publication fails, the committed block ledger and the target temporary block are retained for verification in step S5.

[0035] On the other hand, the present invention also provides a single-cell data format conversion system based on intermediate representation, comprising:

[0036] The parsing and identification module parses the source single-cell object and generates a list of semantic constraints and stable logical tokens; the block planning module determines the variable block boundaries and generates a block transformation plan; the block transformation module generates a semantic digest of the source block, performs transpose isomorphic three-array reuse or memory-constrained rearrangement, and writes it to the target temporary block; the readback verification module generates a semantic digest of the target block, performs semantic comparison, and updates the committed block ledger; and the recovery and commit module determines the set to be replayed and commits the target root object after all blocks have been verified.

[0037] The beneficial effects of this invention are:

[0038] 1. By using a semantic constraint list, stable logical tokens, and layout-independent semantic digests, the physical axis direction and sparse storage differences are separated from the logical semantics of cells, genes, data layers, and object references. This allows the target temporary block to be read back and verified according to the same rules as the source block, thereby identifying errors in axes, annotations, types, data layer roles, and reference relationships before release.

[0039] 2. Variable blocks are determined by combining the sparse matrix pointer array, numerical width, index width, buffer usage, and memory limit. When the semantic constraints are met, the transposed isomorphic three arrays are reused; otherwise, memory-constrained rearrangement is performed. This controls memory usage during the transformation and reduces the rearrangement of non-zero elements that meet the conditions.

[0040] 3. By atomically registering the ledger only for the verified target blocks, and determining the set of fields, data layers, or blocks to be replayed based on the differences in hierarchical semantic digests, schema dependencies, and object references, recovery after interruption, unreadable target blocks, inconsistent digests, or rule changes does not require re-performing all transformations, and the target root object is only published after all related blocks have been verified. Attached Figure Description

[0041] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0042] Figure 1 This is a flowchart of the single-cell data format conversion method based on intermediate representation of the present invention.

[0043] Figure 2 This is a flowchart of generating a layout-independent semantic summary for S3 of the present invention and selecting transpose isomorphic reuse or memory-constrained rearrangement. Detailed Implementation

[0044] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0045] refer to Figures 1-2 Single-cell data format conversion methods based on intermediate representation include:

[0046] S1. Parse the format version, logical axis, sparse layout, data type, data layer and object reference relationship of the source single cell object, generate a list of semantic constraints, and generate stable logical tokens for cells, genes, data layers and referenced objects.

[0047] S2. Determine the variable block boundaries based on the semantic constraint list, sparse matrix pointer array, numerical width, index width, buffer usage, and memory limit, and generate a block conversion plan;

[0048] S3. Read the source block according to the block conversion plan, generate a source block semantic summary that is independent of sparse layout based on the stable logical token and type normalization value, select transpose isomorphic three-array reuse or memory-constrained rearrangement according to the logical axis and physical storage direction of the source object and the target object, and write the conversion result into the target temporary block.

[0049] S4. Read back the target temporary block and generate the target block semantic summary according to the rules for generating the source block semantic summary. Compare the semantic summaries and associated semantic states of the source block and the target block according to the semantic constraint list. When the verification is successful, register the corresponding block atomically to the committed block ledger.

[0050] S5. Verify the submitted block ledger, determine the set to be replayed based on unregistered blocks, digest differences and schema dependencies, perform transformation and verification on the set to be replayed, until all blocks are registered to the submitted block ledger and then submit the target root object.

[0051] In this specific embodiment, S1 includes:

[0052] In this specific embodiment, the data conversion server uses the RDS serialization file of a Seurat object as the source single-cell object and the AnnData object carried by the hierarchical data file as the target object. When the conversion session is started, a session identifier, source object content fingerprint and target format version number are generated, and the object namespace is set as a normalized combination of source file content fingerprint, object class identifier and object internal path. All subsequent steps use this object namespace to identify the same object.

[0053] Before parsing object semantics, the source carrier adapter first reads the serialization version, character encoding, byte order, and compression flags of the RDS header. It decompresses the compressed stream sequentially with an 8MB fixed input buffer and parses the object class, attributes, vector length, and reference number according to the R serialization token. When encountering a dgCMatrix object, the complete matrix is ​​not kept resident in memory. Instead, the p, i, and x slot vectors are written sequentially to three disk-supported flat array files under ir / source-arrays in the session directory. At the same time, the element type, number of elements, 64-bit file start and end offsets, segment check value, and source object content fingerprint are recorded for each array. Cell and gene names, axis data, object slots, and reference number tables are written to key-value record files in the same directory. The reference number is parsed to the canonical object identity. After sequential parsing is completed, the files and directories are refreshed and a source carrier index is generated.

[0054] The source carrier adapter rereads the first segment, last segment, and segment boundaries of each written array by index, requiring the number of elements, file byte length, object path, and content fingerprint to be consistent with the parsed record. If the decompression or serialization token is not supported, the object reference sequence number is dangling, the temporary space is lower than the expected materialized bytes, the post-write verification fails, or the source RDS content fingerprint changes during adaptation, the incomplete carrier index is deleted and the session is marked as source adaptation failure, and the S2 block generation plan is prohibited. When executed again, the complete carrier is reused only if the source content fingerprint and the adapter version are the same; otherwise, it is rematerialized from the RDS sequential stream.

[0055] The parsing and identification module loads the RDS file header and serialized object directory in read-only mode, reads the Seurat object class, object version, analysis dataset, cell metadata, dimensionality reduction results, adjacency graph, spatial image, and current analysis data slot description, and traverses the referenced objects in breadth-first order from the slot reference entry. The traversal queue records the parent object identifier, slot name, reference path, and target object class. The visited set uses the canonical object identity, which does not change with the traversal path, as the deduplication key. This identity consists of the object namespace and the object's internal stable identifier. The reference path is only stored as a parent-child edge attribute. When the same object is visited again, a new parent-child edge is recorded without re-queuing. If the object's internal stable identifier is missing or the same canonical object identity corresponds to different object content fingerprints, a parsing failure status is written and the expansion stops. This ensures that the circular reference object graph terminates within a finite set of objects and forms an object topology record with parent-child edges.

[0056] For each expression matrix, the parsing and identification module reads the matrix dimensions, row name sequence, column name sequence, sparse object class, and pointer array length. It then compares the cell identifier sequence with the cell metadata row names one by one, and the gene identifier sequence with the analysis data feature names one by one. When the column names of the source compressed sparse column matrix match the cell identifier sequence and the row names match the gene identifier sequence, the first physical dimension is registered as the gene logical axis and the second physical dimension is registered as the cell logical axis. If the axis name is duplicated, empty, or the length is inconsistent with the matrix dimensions, a parsing failure status is written.

[0057] Logical axis judgment adopts the evidence priority rule. The axis role explicitly declared in the file or object pattern is the first evidence. The one-to-one correspondence between the axis identifier and the primary key of the axis metadata is the second evidence. The matrix dimension and the number of identifiers are equal to the third evidence. The next evidence is used only when the preceding evidence does not exist. When two valid pieces of evidence give opposite axis roles, automatic transposition is not performed. Instead, the evidence source, matrix path and conflicting axis are written to the parsing error record.

[0058] The parsing and labeling module registers the counting slots as the original counting data layer, the normalized expression slots as the normalized expression data layer, and the scaled expression slots as the scaled expression data layer. It also registers the roles of dimensionality reduction data, graph data, spatial data, and object reference for dimensionality reduction coordinates, adjacency graphs, spatial coordinates, and image references, respectively. The data layer role mapping record simultaneously saves the source slot path, target AnnData path, logical axis order, source sparse layout, and target sparse layout.

[0059] The parsing and identification module also completes the materialization of source carriers for non-expression sub-items according to object categories: for dimensionality reduction coordinates and spatial coordinate arrays, it records the shape, cell or feature logical axis, element type, flat array file offset, segment check value, and content fingerprint; for adjacency graphs, it records the row and column logical axes, sparse layout, and the offset, quantity, and check value of the three disk support arrays p, i, and x; for spatial images or external payloads, it records the media type, byte length, content fingerprint, controlled source file path, and target UNS landing point, without embedding image bytes into object reference tokens; for object references, it records the canonical node identity, parent-child edge tokens, source path, and unique target landing point respectively; the above records are written to the source carrier index and use the same refresh, readback, and content fingerprint checks as the expression matrix carrier; if any array, payload, or reference edge materialization is incomplete, the corresponding object is added to the parsing failure set;

[0060] The semantic constraint list is recorded using versioned JSON. The list header contains the list version, source format version, target format version, summary rule version, and character normalization version. Each object entry contains the object namespace, object category, logical axis direction, sparse layout, data layer role, metadata type, allowed numerical width and index width, format version compatibility relationship, and object reference closure. The compatibility relationship is given by the format capability table that is released with the conversion program and verified by the version signature. The source and target version combinations that do not match are registered as incompatible.

[0061] The format capability table uses the source object class, source format version, target object class, target format version, and data role as the joint lookup key. The value fields include allowed logical axis order, sparse layout, numeric type, index type, partial write capability, missing value encoding, and reference landing point. The capability table is generated by running the structure and performing round-trip verification on the corresponding RDS and AnnData schema samples when the conversion program is released. The table header stores the program build identifier, schema file fingerprint, and effective time. When any schema fingerprint changes, a new version is generated without overwriting the old version.

[0062] For axis metadata, the semantic constraint list records the field name, source data type, target data type, missing value representation, category value set, stable logical token correspondence, and target path for each field. For RDS objects and S4 objects, the list records the object class, slot name, slot type, slot cardinality, and slot reference target for each slot. One implementation record registers the cell type factor field in the cell metadata as an ordered categorical type, the missing value representation as NA, and the target path as the cell type field in the observation metadata. It also registers the RNA analysis data slot as the Assay class and the reference target as the RNA data layer object.

[0063] Field mapping rules are first searched by normalized field name and data layer role. If no match is found, they are then searched by the alias set explicitly registered in the semantic constraint list. When multiple source fields match the same target field, they are only accepted if the merging rule is explicitly registered in the list and the source types are compatible. Category fields are reconstructed according to stable logical tokens and the original labels are saved. Time, integer and floating-point type conversions all record the source type, target type, allowed range and overflow handling method.

[0064] Stable logical tokens are generated according to deterministic byte rules. The parsing and identification module first performs UTF-8 encoding and UnicodeNFC normalization on the object namespace, logical axis role or data layer role, intra-axis identifier or data layer identifier, object category and reference path respectively. Then, it concatenates them in the order of field sequence number, four-byte field length and field byte value, and calculates the SHA-256 value of the concatenation result. Cell and gene tokens use object namespace, logical axis role and intra-axis identifier, while data layer tokens use object namespace, data layer role and data layer identifier. The canonical node token of the referenced object uses the canonical object identity without the reference path, and the reference edge token uses the parent object canonical identity, slot name, referenced object canonical identity and reference path, so that the same referenced object has only one transformation entity and each reference edge still has a path-related stable token.

[0065] The parsing and identification module builds a token dictionary in ascending order of token byte values. The dictionary entries store the token, the pre-normalization identifier, the post-normalization identifier, the logical role, and the source path. When two different pre-normalization identifiers obtain the same post-normalization identifier or the same stable logical token, the corresponding axis or object is marked as a verification failure and the generation of blocks for it is stopped. Entries that fail to establish target fields, slots, or reference target mappings in the semantic constraint list are written into the unmapped item set and carried to the readback verification stage.

[0066] The object reference closure uses canonical node tokens as nodes and reference edge tokens as directed edges, and separately stores the merged mapping from the canonical object identity to all path-related reference edge tokens. When the same canonical object is reached via different paths, the unique node, the unique block plan object identifier, and the unique ledger object identifier are reused, and only reference edge entries are added. After parsing, a reachability traversal is performed once for each root object. The closure record simultaneously stores the canonical object identity, all incoming edge tokens, the number of incoming edges, the number of outgoing edges, the reference slot type, and the target format landing point. A closure record in the semantic constraint list consists of the root object canonical node token, the RNA analysis data canonical node token, the reference edge token, and the target unstructured data path. If the target landing point does not exist, the reference edge cannot be resolved to the unique canonical object identity, or the same reference edge token points to two different canonical nodes, the root object is marked as verification failed.

[0067] Before writing the intermediate representation, the module rereads each object entry and looks up its axis, data layer, metadata field and reference closure according to the object namespace. It requires that each physical matrix dimension maps exactly to a logical axis, each data layer maps exactly to a data layer role, and each reference entry maps exactly to a referenced object token. The verification results form the object-level preparation state. Only objects with a successful preparation state enter the variable block computation of S2.

[0068] After parsing is complete, the parsing and identification module writes the semantic constraint list to ir / semantic-manifest-v1.json in the session directory, writes the token dictionary to ir / token-dictionary.bin, writes the object topology and reference closure to ir / object-closure.json, and calculates the content fingerprint for the three files. The intermediate representation index containing the file path, byte length and content fingerprint is used as the input for S2 to determine the block boundaries.

[0069] In this specific embodiment, S2 includes:

[0070] The block planning module reads the intermediate representation index, source carrier index, and semantic constraint list output by S1. It opens the pointer array, index array, and value array supported by the disk according to the source carrier index, and first checks the number of array elements, file offset, segment check value, and source object content fingerprint. In this specific embodiment, the server provides 16 GiB of working memory, i.e., 17179869184 bytes. 20% of this is reserved as a safe reserve for runtime and hierarchical data file writing. The upper limit of block availability is 13743895347 bytes, approximately 12.8 GiB, obtained by multiplying 17179869184 by 0.8 and rounding down. The module also obtains the value width of 8 bytes, the source index width of 4 bytes, the target index width of 8 bytes, and the maximum block length of the target dataset from the list.

[0071] The block planning module establishes a plan record for the dimensionality reduction, graph, spatial, and object reference carriers within the complete transformation range, along with the representation matrix: the dimensionality reduction and spatial numerical arrays are block-based according to the continuous intervals of the first logical axis and the estimated number of bytes; the sparse graph is block-based according to the difference of its p array pointers and uses the same peak memory constraint as the representation matrix; the image or external payload is registered as a whole copy task of the controlled file, and the single task buffer is required not to exceed the block's available limit; object reference nodes and edges are segmented according to the parent object's token prefix; each plan saves the source offset, estimated number of records, estimated number of bytes, target temporary group path, necessary sub-item identifiers, and reference predecessors; the whole payload exceeding the limit is converted to a fixed 8MiB streaming copy buffer; payloads that cannot be streamed are entered into the block plan failure set;

[0072] The block planning module scans the pointer array along the cell logical axis and calculates the candidate axis intervals. ,in Indicates the position of the axis To axis position The number of non-zero elements and dimensionless; This represents the total number of dimensionless positions along the logical axes of the current object's cell, and is a positive integer. Indicates the starting axis position of the candidate interval, which is a dimensionless integer index read sequentially from the source cell axis; Indicates the end axis position of the candidate interval, which is a dimensionless integer index read sequentially from the source cell axis; and satisfy ; Indicates the axis position read from the source sparse matrix pointer array. The cumulative count starts at a point and is a dimensionless non-negative integer. Indicates the axis position read from the same array of pointers. The next digit is the cumulative count, which is a dimensionless non-negative integer; both satisfy the following conditions: ; This represents the total number of dimensionless non-zero elements in the current source sparse matrix;

[0073] Calculate for each candidate axis interval ,in This represents the peak memory estimate for the candidate interval, in bytes. This indicates the axis position obtained according to the counting formula in the previous section. To axis position The number of dimensionless non-zero elements; This indicates the starting axis position of the candidate interval as described in the previous complete explanation; This indicates the position of the end axis of the candidate interval as described in the previous complete explanation; Indicates the width in bytes of a single source numeric element read from the source numeric element type field of the semantic constraint list, taken as a positive integer; Indicates the width of a single source index byte read from the source index element type field of the source carrier index, and is a positive integer; This represents the total byte width of a single target index and the rearranged position record, and is a positive integer. The target index width is read from the semantic constraint list, and the rearranged position record uses an unsigned 64-bit integer and is fixed at 8 bytes. Therefore, in this embodiment, the target index width is 8 bytes. Take 16 bytes; This indicates the byte width of a single pointer value read from the element type field of the source pointer array, and it must be a positive integer. In this embodiment, the source pointer array uses 64-bit integer elements. Take 8 bytes; range to pointer slices by to common It consists of pointer values; This represents the number of bytes in the sorting work area obtained from the preheating measurement in this section, and is a non-negative integer. This indicates the number of bytes in the working area calculated from the summary obtained from the preheating measurements in this section, and it is a non-negative integer; This represents the number of bytes written to the target buffer obtained from the preheating measurement in this segment, and is a non-negative integer. This represents the baseline byte count of the conversion thread, object description, and fixed cache obtained from this preheating measurement, and is a non-negative integer; each buffer item is obtained by performing a preheating measurement on 10,000 non-zero elements with the same version of the conversion kernel at startup and rounding up to the 1MiB boundary;

[0074] The preheating measurement is executed in three rounds. The first round is used to establish a file page cache. The next two rounds record the process resident memory increments during the sorting, digest, and target writing stages, respectively. The buffer item takes the maximum value of the corresponding increments from the two rounds and adds the thread stack limit. The measurement record saves the conversion kernel build identifier, the number of concurrent threads, the target compression parameters, and the timestamp. When any parameter changes, the old measurement record is discarded and remeasured before the block plan is generated.

[0075] The upper limit of available blocks is denoted as ,in Indicates the number of bytes of allocatable memory after deducting runtime safety reserves, subscript This indicates a runtime retention factor deducted from the scenario and is only a mnemonic index; this embodiment It is 13743895347 bytes; This represents the estimated peak memory of the candidate interval obtained according to the aforementioned peak memory formula and its complete definition, in bytes. This indicates the starting axis position of the candidate interval in the peak memory formula; This indicates the end axis position of the candidate interval in the peak memory formula; the module expands backward from the current starting axis position to the end position and checks it. Retain the maximum value that satisfies the constraints. As the end position of the current block, if two end positions obtain the same estimated value, the position that covers more logical axis markers is retained, and then the scanning continues from the next axis position until the entire logical axis range of the current object is covered;

[0076] When a single axis interval that cannot be further divided along the logical axis boundary still exceeds the block's available limit, the module retains a parent block identifier and extracts the longest continuous interval that satisfies the memory constraints according to the sequence of non-zero elements to form an execution segment within the parent block. Each segment records the parent axis token, segment number, start and end positions of non-zero elements, number of non-zero elements in the segment, total number of non-zero elements in the parent axis, and a logical axis incomplete flag. Within the same parent axis, the parent axis token, segment number, and start and end positions of non-zero elements are used as the joint overwrite key. The segment number must be continuous, the non-zero intervals must be contiguous, and their union must equal the complete non-zero interval of the parent axis. The hierarchical dataset of the target AnnData supports writing continuous superslices to the same parent block temporary group by offset and allows the generation of a parent block pointer array after all segments are completed. The segment is not registered as an independent final block. If the target capability table marks the corresponding path as prohibiting the above-mentioned delayed assembly writing, the parent axis interval is registered as a conversion failure and no write task is generated.

[0077] The module compares the maximum gene index of the source index array within the current block with the upper limit of the target index type. In this embodiment, the blocks are divided along the cell logical axis, and the index array of the source dgCMatrix stores the global gene axis index. Therefore, the indices of the target CSR directly retain the dimensionless global gene index without deducting the cell block start position. The cell axis start position in the block plan is only used for target row offset and indptr writing. In this specific embodiment, the upper limit of the 32-bit signed index is 2147483647. When the maximum gene index exceeds this upper limit, a 64-bit signed target index is used. If there is no target index type in the format capability table that covers the upper limit of the global gene index, the corresponding block is written to the failure list and the generation of write tasks is prohibited. Local conversion is only performed when another target path is explicitly divided along the gene axis, the gene slice start point is stored in the root object list, and the capability table allows local indexing. This condition does not apply to the AnnDataCSR path in this embodiment.

[0078] Each block transformation plan records the canonical object identity, data layer role, block identifier, parent block identifier, logical axis start and end tokens, source axis start and end positions, target axis start and end sequence, source pointer start and end values, non-zero element start and end values, source and target numerical widths, source and target index widths, block offset, estimated peak memory, write path, allowed transformation modes, and dependent object canonical identities. The module uses the source cell identifier sequence and source gene identifier sequence registered in the S1 semantic constraint list as the unique axis order, establishing a one-to-one mapping from cell tokens to target row sequences and from gene tokens to target column sequences, with sequences numbered consecutively from zero. The axis sequence mapping table stores each token, source axis position, and... The target axis sequence and mapping table content fingerprint are recorded. When there are execution segments within the parent block, the segment sequence number, segment non-zero interval, total number of non-zero elements in the parent axis, and delayed pointer assembly flag are also recorded. The module assigns parent block identifiers in ascending order according to the object identity, data layer role, and minimum target cell sequence. After the plan is generated, each parent block is required to cover continuous target cell sequences. The target sequence intervals of adjacent parent blocks have no overlap and the union is equal to zero to the total number of cells minus one. The source axis position and token correspondence are checked according to the axis sequence mapping table. Within the same parent axis, the non-zero intervals are checked to be continuous and non-overlapping according to the segment joint coverage key. It is forbidden to misjudge execution segments that share the parent axis token as two logical axis blocks.

[0079] The block plan allows the number of non-zero elements in the data layer to be zero, but only records that meet the conditions for a valid empty block are registered as empty blocks: the logical axis token interval is not empty, the first cumulative count and the last cumulative count of the source p slice are equal, the length of the target indptr in the plan is equal to the number of logical axes plus one, and the expected lengths of the target indices and data are both zero; the plan also registers the empty node identifier of the expression item, the empty node summary rule version, and the null value allowance flag. If any of the conditions are not met, the failure list cannot be circumvented with an empty block;

[0080] In this specific implementation, a block plan record covers 4096 consecutive cell logical tokens, with 12,653,472 non-zero elements. The value width is 8 bytes, the source index width is 4 bytes, the combined width of the target index and rearranged position record is 16 bytes, and the pointer width is 8 bytes. , , and The total is 805,306,368 bytes, or 768 MiB. Substituting into the peak memory formula, we get 12,653,472 multiplied by 28, plus 4,097 multiplied by 8, plus 805,306,368, which equals 115,963,636 bytes, approximately 1.08 GiB, which is lower than the block availability limit of 137,438,953,47 bytes. This record registers the allowed transformation mode as transpose isomorphic reuse or memory-constrained rearrangement, and writes the target data layer path and parent object token into the dependency field.

[0081] After the plan is completed, the module calculates the union of the source pointer intervals and the union of the logical axis token intervals of all parent blocks according to the identity of the specified object. It requires that the first parent block starts from the zero pointer, the last parent block ends at the last value of the pointer array, the boundaries of adjacent parent blocks are continuous, and the sequence of non-zero elements is not omitted. For parent blocks containing execution segments, the module calculates the union of non-zero element intervals according to the segment number, requiring that the first segment starts at the first non-zero position of the parent axis, the last segment ends at the last non-zero position of the parent axis, and the beginning and end of adjacent segments are continuous. Then, the module simulates the target write offset of each parent block and segment, allowing the same parent block segment to write to pre-allocated and non-overlapping continuous intervals, prohibiting different parent blocks from covering the same target interval, and any unauthorized overwriting, hole, or segment estimated peak exceeding the block available limit is moved from the execution plan to the failure list.

[0082] The number of concurrent blocks allowed during execution is determined by the lower bound of the integer after dividing the upper limit of available blocks by the estimated peak value of the current candidate blocks, and is further restricted by the single writer constraint of the target file. In this specific implementation, only one write task is allowed in the same target data layer, and other tasks remain in the ready queue without pre-allocation of sorting and write buffers, so that the total estimated occupancy after concurrent scheduling still does not exceed the upper limit of available blocks.

[0083] The block planning module writes the checked plans to ir / block-plan-v1.jsonl, and the failed blocks and their reasons to ir / block-plan-errors.jsonl. It also merges the S1 object-level preparation status, the set of parsing failures, the set of unmapped blocking blocks, and the S2 failure records into a complete transformation range list. This list retains the status and failure reason for each expected object, logical axis interval, and expected block identity, ensuring that failed items do not disappear from the transformation range due to not being included in the execution plan. It also calculates fingerprints for the list content. The plan header stores the plan content fingerprint, the complete transformation range list fingerprint, the semantic constraint list fingerprint, and the memory measurement version. The block transformation plan, the complete transformation range list, and the read-only token dictionary together serve as inputs for S3 to read source blocks, generate semantic digests, select transformation paths, and for S5 to execute the release gate.

[0084] In this specific embodiment, S3 includes:

[0085] The block conversion module receives the write task according to the parent block identifier and fragment sequence number of S2. After verifying the semantic constraint list fingerprint, source carrier index fingerprint, token dictionary fingerprint, and axis sequence mapping table content fingerprint in the block conversion plan, it reads the p, i, x array range corresponding to the current parent block or execution fragment in memory mapping based on the file offset given by the source carrier index. At the same time, it reads the cell token, gene token, source axis position, target axis sequence, data layer token, axis metadata record, and object reference record covered by the block. The module requires the cell token to be mapped to the target row sequence registered in the plan and the gene token to be mapped to the target column sequence registered in the plan. If any source carrier file offset, element length, segment check value, source content fingerprint, or axis sequence correspondence is inconsistent with the plan, the current parent block is terminated and the corresponding index is rebuilt. No target temporary block is created.

[0086] For non-expressive sub-tasks, the block transformation module reads the dimensionality reduction or spatial numerical slices according to the planned offset and writes them into independent target temporary groups named with object tokens and logical axis intervals; the adjacency graph reads its p, i, x intervals according to the dual-axis tokens, and writes them into the corresponding temporary groups of obsp or varp after checking the index range and symmetry requirements; the image or external payload is copied to the corresponding candidate payload file of uns in a fixed buffered streaming manner and the content fingerprint is calculated at the same time; the object reference writes the canonical node identity and parent-child edge tokens into the independent node table and edge table; each temporary group records the shape, element type, logical axis, actual number of records, target standard path and write completion flag, and can only generate sub-specification records after closing and refreshing. If any payload fingerprint, reference landing point or record number is inconsistent, the temporary group is deleted and its parent block is blocked;

[0087] The module normalizes values ​​according to the digest rule version. 64-bit integers use fixed-width big-endian two's complement bytes. Floating-point values ​​first normalize negative zero to positive zero, normalize all NaN payloads to the same silent NaN bit type and use IEEE754 big-endian bytes. Boolean values ​​use single-byte 0 or 1. Strings are normalized to UnicodeNFC and then use UTF-8 bytes with a four-byte length prefix. Missing values ​​are encoded according to the missing marker of the corresponding metadata type in the semantic constraint list. Category values ​​use category-stable logical tokens instead of the integer encoding inside the source format.

[0088] For non-zero elements in the expression matrix, the module performs deterministic semantic preprocessing on the source records before generating the summary, based on cell stable logical tokens, gene stable logical tokens, and data layer roles: the original counts of the same logical coordinates are accumulated as unsigned 64-bit integers, the normalized expression values ​​given by the source object are summed deterministically with compensation based on the source sequence number, and the overflow, explicit zero, and type byte specification rules adopted by the target write are executed simultaneously; the type byte specification is based on the source value, only unifying the integer width, byte order, negative zero, and NaN bit type, without performing numerical scaling, division, or interval pruning; then, the merged cell tokens, gene tokens, data layer roles, type identifiers, and type specification values ​​are used to form a specification record, which is arranged in ascending order of cell token byte value, gene token byte value, and data layer role code. Axis data, dimensionality reduction data, graph data, spatial data, and object reference records are formed into specification records with the same structure according to their respective object tokens, field roles, and logical coordinates, so that the order of specification records does not depend on the physical storage direction of CSC or CSR;

[0089] No. Summary of each standard record Calculation, where Indicates the first The 32-byte leaf digest of the specification record, This represents the total number of canonical records in the current source block and is a positive integer. Indicates satisfaction The standard record number, This refers to the SHA-256 hash function. The stable logical token byte representing the object or axis to which the record belongs. Represents the data layer or field role code byte. Represents the logical coordinate token combination bytes, This represents the normalized type identifier byte. Represents a type-normalized value byte. This indicates that byte strings are concatenated by a prefix indicating the field length.

[0090] The root digest of the current block is... The calculations are performed using parameters derived from the output records of the S2 block plan and the previous leaf summary formula; where... This represents the 32-byte source block digest of the current block; This indicates the current dimensionless block number determined by the block identifier in the S2 block plan; This indicates the total number of canonical records in the current source block as fully explained in the previous leaf summary formula; This indicates the SHA-256 hash function that uses the previous leaf summary formula and its complete explanation; This indicates the first leaf summary obtained according to the previous leaf summary formula. A 32-byte leaf digest of the specification record; This indicates the dimensionless record number in the leaf abstract connection order and satisfies... ; to The stable order of records is obtained according to the previously defined specifications. This indicates that the byte strings are concatenated by field length prefixes, following the complete explanation in the previous paragraph; the module simultaneously establishes sub-nodes according to six categories: expressive data layer, axis data, reduced-dimensional data, graph data, spatial data, and object references. The sub-nodes are connected in ascending order according to their child node identifiers, and the parent summary is calculated, writing a hierarchical semantic summary tree level by level;

[0091] Before selecting a transformation path, the module obtains the logical axis order and physical storage direction from the source matrix description and the target AnnData dataset description. In this specific implementation, the source dgCMatrix is ​​stored in the CSC method of gene multiplication by cell, and the target X dataset is interpreted in the CSR method of cell multiplication by gene. When the two are transpose correspondences, the transpose isomorphism check is entered; otherwise, the memory-constrained rearrangement is directly entered.

[0092] The transpose isomorphic check verifies that the length of the source p array is equal to the number of rows in the target CSR plus one, the p array is monotonically non-decreasing and its last value is equal to the number of non-zero elements, the i array value is within the target column range, and the index order within each pointer interval meets the target requirements. The logical axis fingerprint, repeated logical coordinate merging rules, explicit zero retention rules, numerical width, index width, and data layer role are compared with the semantic constraint list. Only when all conditions are met is the transformation mode registered as transpose-isomorphic-reuse.

[0093] When using transposed isomorphic three-array reuse and the parent block has no internal execution fragments, the module first checks line by line to ensure that the increasing order of the source cell axis positions is completely consistent with the increasing order of the target row positions in the axis position mapping table. It then checks that the target column position obtained from the global gene index in the source i array, obtained through gene token lookup, is equal to the index value. Only when both identity conditions are met does the module read the p, i, and x arrays in the order of the start and end positions of the source arrays in the block plan. It then subtracts the current block's first pointer value from the p array and writes it to the target indptr. Finally, it selects the i array according to S2. Write the target index type of a sufficiently wide target index to the target indices without loss and do not deduct the cell block starting point. Then, perform a lossless transformation on the x array according to the target numerical width required by the list and write it to the target data. The relative order of the elements of the three arrays remains unchanged. If any identity condition is not met, reuse is prohibited. Instead, the memory-constrained path is rearranged. The cell token and gene token of each record are converted into the target row order and target column order respectively according to the axis order mapping table. After being stably sorted by the target row order and target column order, it is written. The original physical row order is not allowed to be used.

[0094] When a parent block contains an internal execution segment of a single axis, the complete p slice of that parent axis is not directly reused. The module writes the i and x of each segment consecutively into the pre-allocated and non-overlapping indices and data ranges in the same temporary group of the parent block according to the segment number. Each segment only saves the local pointer boundary from zero to the end of the number of non-zero elements in the segment, the non-zero offset within the parent axis, and the write completion flag. After all segments are closed and refreshed, the module accumulates the number of non-zero elements in each segment according to the segment number. The accumulated value is required to be equal to the total number of non-zero elements in the parent axis. Then, the parent block indptr boundary corresponding to the parent axis is generated once again. The segment local pointer is not published as the official indptr of the target AnnData.

[0095] When any one of the transpose isomorphism conditions is not met, the module employs a two-stage external sorting process: In the first stage, segmented buckets are established based on the high-order prefix of the target cell token. Source non-zero elements are read one by one, and the gene token, cell token, type normalization value, and source sequence number are written to the corresponding bucket. When a bucket reaches its memory limit, it is stably sorted by the target row token, target column token, and source sequence number, resulting in a disk-ordered run segment with the run segment number, minimum and maximum row tokens, record count, and content fingerprint. No run segment is directly published as a CSR. In the second stage, bounded memory multi-way merging is performed bucket by bucket in ascending prefix order. Using the target row token, target column token, and source sequence number as comparison keys, the indptr of a row is determined once all run segments involved in that row are exhausted. Indices and data are output according to a predetermined repeating coordinate rule. If a run segment is missing, the fingerprint is incorrect, temporary space is insufficient, or the number of records after merging is not conserved, the output temporary group is deleted, and the conversion is marked as failed.

[0096] For duplicate logical coordinates, both source digest preprocessing and target writing use the same rules: the original counting layer accumulates unsigned 64-bit integers and marks verification failure when the value exceeds the target value limit; the normalization layer performs deterministic compensation summation according to the source sequence number; the explicit zero retention rule is to write zero-value records when retaining them, and to delete them synchronously before digest generation and target writing when deleting them; the module records the number of original records before merging, the number of semantic records after merging, and the number of target written records respectively. The former number is only used for transmission integrity auditing, while the latter two must be equal and used for source-target digest comparison; records that fail to resolve the target row and column positions by stable logical tokens are written to the unmapped item set and the current block commit is blocked;

[0097] For a valid empty expression block registered by S2, the module still creates a target temporary group corresponding to the non-empty logical axis range as planned, writes an indptr with a length equal to the number of logical axes plus one and all items being zero, does not write indices and data, and uses the pre-registered empty node rules to generate source expression sub-item summaries and target empty node identifiers to be read back respectively; before closing the refresh, it is required again that the cumulative counts of the first and last slices of the source p slice are equal, the target indptr is all zero and the lengths of the three arrays meet the plan respectively, if any condition is not met, the temporary group is deleted and the conversion failure is registered;

[0098] Before writing the array, the ordinary parent block recalculates the first value, last value, and length of the target pointer array according to the block plan, requiring the first value to be zero, the last value to be equal to the length of the target value array, and the sum of the differences between adjacent pointers to be equal to the number of non-zero elements in the block; the parent block containing execution fragments first requires the local pointers to have a first value of zero and a last value equal to the number of values ​​in that fragment, and then requires the last value of the parent block's indptr to be equal to the sum of the number of values ​​in each fragment and equal to the total number of non-zero elements in the parent axis after all fragments are assembled; both types of paths require that each value in the target index array falls within the range of the target logical column, and if the verification fails, delete the temporary dataset of the parent block that has not yet been closed and set the status of the parent block and all its fragments to conversion failure;

[0099] When writing axis data and object slots with the expression matrix block, a stable logical token is used as the connection key. The module sorts the source records according to the token byte value and generates the target row number mapping. The field value is converted according to the target type registered in S1 and missing values ​​are written as a target type-specific missing mark. When the target field does not exist, the category label is not in the list set, or the object reference does not exist, the current block writing is stopped and the unmapped items are saved.

[0100] Before the temporary group is closed, the module calculates the byte length and fast transmission check value for the target pointer array, index array and value array respectively, writes the three check values, the number of standard records, the number of six types of sub-item summary nodes and the number of bytes written into the completion record, and then forces a refresh of the target file and its directory. If any field is missing in the completion record, the temporary group will not be published to the readback verification queue.

[0101] When all verification results satisfy the semantic constraint list, the three arrays are transposed and reused isomorphically. Otherwise, the non-zero elements are rearranged in segments within the memory limit according to the stable logical token. The reuse result and the rearrangement result both save the same logical axis token range, data layer role and digest rule version. Both transformation results enter the readback verification in step S4.

[0102] After the write task is completed, the block conversion module will check the actual length of the pointer array, index array, value array, number of target axis tokens, and number of sub-summary nodes against the block plan item by item, and write the check results, along with the source read offset, target write offset, and target temporary group content fingerprint, into the handover record. If any quantity or offset is inconsistent, the handover record will not be published to S4.

[0103] Both transformation paths write the results to an independent hierarchical data temporary group under the temporary block directory of the target session directory. The temporary group records the block identifier, object identifier, data layer role, logical axis token range, transformation mode, source block digest, digest rule version, semantic constraint list fingerprint, and write completion flag. After closing and refreshing the target file handle, only the target temporary group path and the source block hierarchical semantic digest tree are handed over to S4 for readback verification.

[0104] In this specific embodiment, S4 includes:

[0105] The readback verification module does not reuse the S3 memory array. Instead, it reopens the hierarchical data file after the target parent block temporary group is closed. It reads back the logical axis identifier, parent block pointer array, index array, value array, data layer role, axis metadata type, target format version, and object reference identifier according to the block plan. It also verifies the block identifier, canonical object identity, logical axis token range, and write completion flag of the target temporary group. When the parent block contains execution fragments, it also reads back the sequence number, non-zero offset within the parent axis, local quantity, and completion flag of each fragment. It requires that the fragment ranges be continuous and non-overlapping, the sum of the fragment quantities equals the total number of non-zero elements in the parent axis, and the final value of the assembled parent block indptr equals the length of the parent block value array. If the read fails or the record header is inconsistent, the entire parent block is registered as a verification failure.

[0106] For non-expressive sub-items, the readback verification module reopens the dimensionality reduction and spatial numerical temporary groups according to the semantic constraint list and verifies the shape, logical axis, element type, number of records, and numerical content. It also reopens the graph data temporary group and checks the lengths of p, i, and x, pointer monotonicity, index range, and dual-axis tokens. The module rereads the image or external payload according to the fixed buffer and verifies the media type, byte length, content fingerprint, and target uns landing point. It reads the object reference node table and edge table and verifies that both ends of each edge exist and the target path is unique. Subsequently, it reconstructs the dimensionality reduction, graph, spatial, and object reference specification records and necessary sub-item summaries from the above target physical carriers. It is prohibited to use S3 source records or summary caches to replace target readback.

[0107] The readback task distinguishes the input source from the transformation mode field in S3. All verification results that satisfy the semantic constraint list and have an identity mapping from the source axis position to the target axis position are entered into the multiplexing result check for the block that performs transpose isomorphism three-array multiplexing. The remaining blocks that perform segmented rearrangement of non-zero elements within the memory limit according to the axis position mapping table are entered into the rearrangement result check. Both transformation results enter the readback verification in step S4 and use the same pass condition. Both must prove that the stable logical token obtained by reverse lookup of each readback row position and column position is consistent with the cell token and gene token in the source canonical record.

[0108] The module uses the same digest rule version and character normalization version as S3 to process the readback value. First, it uses the S2 axis sequence mapping table to reverse look up the physical row sequence of the target compressed sparse row to find a unique cell token, and reverse look up the target column sequence in indices to find a unique gene token. Then, it restores the logical coordinates composed of cell tokens and gene tokens. If the target sequence is missing, one item corresponds to multiple tokens, the block target row interval is inconsistent with the plan registration interval, or the fingerprint of the mapping table content is inconsistent, the verification is directly judged as failing. After the reverse lookup, the same type normalization rule as the source block is executed on the readback data. Then, the target specification record is generated according to the stable logical token, logical coordinates and data layer role order. Reading the source digest cache to replace the target data for recalculation is prohibited.

[0109] No. Summary of the target specification record by leaf Calculation, where Indicates the first The 32-byte leaf summary of the target specification record. This indicates the total number of canonical records in the current target block, and it must be a positive integer. Indicates satisfaction Target specification record number, This refers to the SHA-256 hash function. This indicates the stable logical token byte that represents the object or axis to which the readback record belongs. This indicates reading back the data layer or field role code bytes. This indicates that the logical coordinate token combination bytes are read back. This indicates the type identifier byte after the readback value has been normalized. Indicates a readback of a normalized value byte. This indicates that byte strings are concatenated by a prefix indicating the field length.

[0110] target block root summary by Calculation, where This represents the 32-byte root digest of the current target block; This indicates the dimensionless block number corresponding to the current block identifier; This represents the total number of records in the current target block specification and is a positive integer. This indicates the SHA-256 hash function that uses the target leaf summary formula and its complete explanation from the previous section. This indicates the first segment obtained according to the target leaf summary formula of the previous segment. A 32-byte leaf summary of the target specification record; Indicates the dimensionless record number in the target leaf summary connection order and satisfies ; to Obtained in ascending order of stable logical token, logical coordinate, and role code; This indicates that the byte strings are concatenated by field length prefix, following the complete definition of the previous paragraph; the leaf summary is concatenated in the above order, and empty entries with zero record count are allowed to use the empty node bytes pre-registered in the summary rules to calculate the summary;

[0111] The target specification records calculate the leaf summary according to the leaf summary connection rules defined in S3, and generate the sub-summaries of the expression data layer, axis data, dimensionality reduction data, graph data, spatial data and object references in ascending order of child node identifiers. Finally, the target block root summary is generated. The summary calculation task records the actual number of records read, the number of records of each sub-item, the summary rule version and the target block root summary.

[0112] The back-read verification module determines the necessary sub-items based on the object category of the current block in the semantic constraint list. For the expression matrix block, at least the expression data layer, axis data, and object reference sub-items are selected. For the dimensionality reduction object, the dimensionality reduction data, axis data, and object reference sub-items are selected. For the adjacency graph object, the graph data, axis data, and object reference sub-items are selected. For the spatial object, the spatial data, axis data, and object reference sub-items are selected. The version of the necessary sub-item set and the list fingerprint are written together into the verification record.

[0113] The necessary item selection rules are derived from the mapping of object categories to item sets in the semantic constraint list. The mapping record contains the object category, required item identifier, nullable flag, minimum number of records, and rule version. Ordinary expression matrix blocks require the existence of three items: expression data layer, axis data, and object reference, and the number of expression data layer records must be greater than zero. Valid empty expression blocks registered by S2 require that the logical axis token interval is not empty, the cumulative counts of the first and last slices of the source p are equal, the length of the readback indptr is equal to the number of logical axes plus one and all items are zero, and the lengths of indices and data are both zero. Furthermore, the source and target expression items must calculate their summaries according to the same pre-registered empty node rule. When these conditions are met, the number of expression records is allowed to be zero, and full conjunctive verification continues. Verification fails directly when the mapping key is not hit, the empty block conditions are incomplete, or the actual number of records is lower than the corresponding minimum number.

[0114] The module first searches the format capability table using the source object class, source format version, target object class, target format version, and data role recorded in the semantic constraint list. If a match is found, a common compatibility relationship identifier containing the capability table version and the composite key is obtained. If a match is not found, signature verification fails, or entry tags are incompatible, the current block is directly judged as verification failed. Then, the logical axis fingerprint, axis sequence mapping table content fingerprint, data layer role, metadata type set, common compatibility relationship identifier, and object reference closure of the target block are normalized in ascending order of field name to calculate the target associated semantic state fingerprint. The corresponding field in the S3 source record is compared with the same axis sequence mapping table content fingerprint and common compatibility relationship identifier to calculate the source associated semantic state fingerprint according to the same rules. Both fingerprints are 32 bytes. The original values ​​of the source format version and the target format version are not directly used as fingerprint fields that are required to be equal. If a field is missing, a missing field marker with the field name is used in the calculation to avoid the same result as an empty string due to missing fields.

[0115] Verification results by Calculation, where Representation block The dimensionless binary verification result is given, with 1 indicating success and 0 indicating failure. This indicates the dimensionless block number corresponding to the current block identifier; An indicator function that returns 1 if the condition is true and 0 otherwise; This represents the 32-byte root digest of the current target block obtained by using the aforementioned target block root digest formula and its complete interpretation; This represents the 32-byte source block root digest generated by S3; Represents a 32-byte target-associated semantic state fingerprint; This represents a 32-byte source-associated semantic state fingerprint. Representation block The necessary set of item identifiers; Indicates necessary sub-item identifiers; This represents a 32-byte digest of the necessary sub-items corresponding to the target block; This represents a 32-byte digest of the necessary sub-items corresponding to the source block; This indicates that the product of all indicator values ​​in the necessary item set is taken;

[0116] The verification process employs a full conjunctive rule, and a pass is only achieved when the target block root digest equals the source block root digest, each necessary sub-item digest is equal, the logical axis fingerprint is equal, the axis ordinal mapping table content fingerprint is equal, the data layer roles are equal, the metadata type sets are equal, the source and target version combinations have obtained the same common compatibility relationship identifier, and the target object reference closure is equal to the source object reference closure. If any item is inconsistent, the first difference node is searched in the order of root digest to sub-item digest to leaf digest, and the path of that node, the source and target digests, the logical coordinates, and all inconsistent fields are recorded, while keeping the target temporary group in an uncommitted state. When the format capability table version or the axis ordinal mapping table content fingerprint changes, the schema dependency fingerprint changes accordingly, and S5 adds the relevant blocks to the replay set.

[0117] The reuse result check reconfirms the length of the readback pointer array, monotonicity, last value, index range, source axis position and target axis ordinal identity condition, duplicate logical coordinate merging result, number of explicit zeros, numerical width, and index width; the rearrangement result check reconfirms that there are no records outside the bucket boundary, the target row ordinal and column ordinals come from the same axis ordinal mapping table, the target row intra-index order is stable, and all source logical coordinates appear exactly once; only after the path-specific check passes can the same summary and associated semantic state concatenation determination be performed;

[0118] For the set of unmapped items carried by S1, the module checks each axis metadata field, RDS object slot and S4 object slot for the existence of target fields, target slots or reference targets according to the stable logic token. If the mapping is missing, the target type is incompatible, the category value set is missing or the reference target is not in the target closure, the corresponding block is marked as verification failure. Even if the root digest is the same by chance due to the lack of this item, the block is not registered.

[0119] For the reuse results of the transposed isomorphic three arrays of S3 and the memory-constrained rearrangement results, the back-read verification reconstructs the same logical specification record from the target physical data. The conversion mode is only used as an audit field and does not enter the semantic equality condition. The reuse path still needs to be checked by pointer boundary, index range, summary and associated semantic state. The rearrangement path still needs to be checked by the duplicate coordinate merging rule and the explicit zero retention rule.

[0120] When the summaries are consistent but the associated semantic states are inconsistent, the module first locates the differences in logical axis fingerprints, data layer roles, metadata types, format versions, or reference closures by field name, and then uses stable logical tokens to look up the source object path and target object path. The difference record saves the block identifier, object category, difference field, source value summary, target value summary, and affected sub-items, so that S5 can determine the field-level, data layer-level, or block-level replay granularity.

[0121] Before writing a ledger record, it checks whether a commit record already exists for the same identifier. During normal writes, if the source block digest, target block digest, pattern dependency fingerprint, and target temporary group path of an existing record are all the same, it returns idempotent success. If any field is different, the new record is written to the conflict file and the replacement of the official record is rejected. Only records that pass verification with the S5 authorized replay flag, the old record generation, and the old record content fingerprint can enter the conditional atomic replacement branch. This branch requires that the current official record content fingerprint is equal to the expected old fingerprint saved by the task, the plan version, and the rule version match, and the new target temporary group has not been referenced by the official record. Otherwise, the replacement is rejected, thus preserving the conflict isolation of normal concurrent repeated writes.

[0122] After successful verification, the module generates a final ledger record containing the object identifier, block identifier, record generation, source block digest, target block digest, necessary sub-item digests, digest rule version, schema dependency fingerprint, target temporary group path, verification time, and commit status. The commit status is set to "committed" during generation. Before replacing an existing official record, the module reads and verifies the old record, writes it to a read-only versioned backup in the same directory named after the block identifier and record generation, saves the content fingerprint during backup, refreshes the backup file and directory, and retains at least the most recently verified valid generation. Then, the new record is written to the temporary record file. After file and directory refresh, the official record is atomically replaced by comparison with the file system and swapped. This atomic replacement is the only persistent commit point. After replacement, the official record is read again and its content fingerprint and "committed" status are verified. Only after successful verification is the old temporary group, no longer referenced by any official record, allowed to be reclaimed.

[0123] If the process is interrupted after backup refresh but before atomic replacement of a new temporary record, the recovery program selects to discard the temporary record or re-execute the conditional atomic replacement based on the temporary record content fingerprint, the expected old fingerprint, and the generation of the formal record. If the process is interrupted after atomic replacement but before readback audit, the recovery program reopens the formal record. If the formal record content fingerprint is correct, the status is committed, and the target temporary group is readable, the existing commit is directly recognized. If any condition is not met, the program searches for a read-only backup with the correct content fingerprint by block identifier and generation decrement. After recovery, the program re-verifies its target temporary group and mode dependency fingerprint. If there is no valid backup, the corrupted record is isolated and the block is handed over to S5 for replay without rewriting through memory status or inferring the committed status.

[0124] Blocks that fail to be verified retain the target temporary group and difference record but are not written to the commit status ledger record. The ledger record that passes verification is indexed uniquely by block identifier and duplicate writing with different digests is rejected. The updated committed block ledger, difference node path and object reference closure index are jointly handed over to S5 to determine the set to be replayed and control the commit of the target root object.

[0125] In this specific embodiment, S5 includes:

[0126] When the recovery and commit module starts, it loads the committed block ledger generated by S4, the complete block transformation plan and complete transformation range list generated by S2 in read-only mode, and loads the S5 cyclic replay group transaction log and group commit pointer. The module first verifies the fingerprint of the complete transformation range list and requires that the S1 parsing failure set, the unmapped blocking set, and the S2 block plan failure set are all empty. If any set is not empty, the session is set to a waiting-for-input-repair or-failure state and publication is prohibited. After passing through the pre-gate, it verifies the fingerprint of each ledger record, the uniqueness of the block identifier, the readability of the target temporary group, the target block digest, the digest rule version, and the schema dependency fingerprint. The verification time, actual summary, and verification conclusion are written to the session-level recovery status file. When the ledger file is corrupted, the previous valid version is selected from the read-only versioned backup retained in S4 by block identifier, record generation decrement, and content fingerprint. After recovery, the target temporary group and mode dependency fingerprint are re-verified. If there is no valid backup, the corrupted record is isolated and the block is added to the replay set. For circular replay groups with a status of prepared in the transaction log but whose group commit pointer does not point to the corresponding new version, the recovery program keeps the old group version visible and selects to delete the complete prepared version or continue to generate missing members according to the group record. It is forbidden to interpret a single prepared member as committed.

[0127] The module constructs an anomaly recovery replay set based on the complete transformation range list and the complete block plan. Objects, axis intervals and expected blocks in the complete range that have no execution plan or retain a failure state are registered as pre-blocking items and are not considered as completed. Blocks that lack a commit status ledger record in the plan are added to the set. Blocks that have ledger records but whose target temporary groups cannot be opened or whose dataset length does not meet the plan are added to the set. Blocks whose recalculated target block summary is inconsistent with the ledger target block summary are added to the set. Blocks whose pattern dependency fingerprint calculated by the current summary rule version, format capability table version or field mapping version is inconsistent with the ledger record are added to the set.

[0128] The block identifier and object identifier in the set to be replayed are both dimensionless discrete identifiers. The target block digest and pattern dependency fingerprint are both 32-byte values. The replay granularity adopts three finite enumeration values: field level, data level, and block level. The difference depth adopts non-negative integer records starting from the root node being zero. The module only compares the digests of the same byte length and the granularity within the same enumeration field, and does not add the digest bytes, record number, or node depth to form a mixed quantity.

[0129] The number of abnormal blocks, the number of dependent blocks, the number of task attempts, and the number of recovery rounds are all count-type results, with the unit being 1 and the value range being from zero to the total number of block transformation plan records or the maximum number of attempts configured. The unit of the depth of the difference node is 1, with a lower limit of zero and an upper limit of the number of layers of the hierarchical semantic summary tree minus one. Before writing the recovery status, the module judges out-of-bounds values ​​according to the above range. When out of bounds, the corresponding task is marked as corrupted and regenerated from the ledger and summary tree.

[0130] For blocks with inconsistent summaries, the module compares from the root node down along the hierarchical semantic summary tree of the source block and the target block. If the difference falls only on a single leaf node or a single axis data element field node, field-level replay is selected. If the difference falls on multiple leaf nodes of the same data layer item, data-level replay is selected. If there are more than two item differences under the root node, the block is unreadable, or the rule version change affects the block structure, block-level replay is selected. The selected granularity, the identifier of the node with the lowest difference, and the affected target path are written into the replay task.

[0131] The module builds a dependency index using the reverse edge of the S1 object reference closure. Starting from the object to which each difference node belongs, it traverses the parent objects that reference the object in breadth-first order. The target blocks that depend on the corresponding fields, data layers or block summaries in the parent objects are added to the set to be replayed. The accessed object and block association key prevents circular expansion until the queue is empty. This results in a closure set containing direct exception blocks and all referenced dependency blocks.

[0132] Before generating replay tasks, the module performs strong connected component identification on the dependency graph of the objects to be replayed. Objects or blocks referencing each other within the same strong connected component are merged into a cyclic replay group. Predecessor relationships are established only between different cyclic replay groups, and scheduling is performed according to the directed acyclic graph obtained after the strong connected components are shrunk. The module generates a stable group identifier and a persistent group transaction record for each cyclic replay group, recording all member block identifiers, the expected old record generation and content fingerprint for each member, the new record content fingerprint, the old group version pointer, the new group version path, and the preparing, prepared, committed, or aborted status. Each task within the group performs write-time replication using the old committed group as a read-only baseline. After completing S3 transformation and S4 readback verification, only the new ledger record is written to the pre-slot of the new group version and marked as pre. The prepared slot does not replace the current official ledger record and is not visible to other cyclic replay groups, integrity checks, and root object commits. It also prohibits the reclamation of old temporary groups. The group-level gate checks that all members are prepared, all reference targets, item summaries, object closures, schema dependency fingerprints, and that the old records are expected to be consistent with the currently visible old group version. After the check is passed, the new group version file and directory are refreshed, and then the group commit pointer is atomically replaced with the same file system, switching it from the old group version to the new group version in one go. This group commit pointer replacement is the only visible commit point in the cyclic replay group. If any task in the group fails or is interrupted before the pointer switch, the transaction status is set to aborted or kept prepared pending recovery. The old group version and all old official records remain visible, and members that have passed the check cannot be considered committed by other groups.

[0133] Each replay task records the task identifier, session identifier, block identifier, object identifier, replay granularity, lowest difference node, set of dependent predecessor tasks, number of attempts, plan content fingerprint, and task status. The task status is in order of pending, running, pending verification, submitted, or failed. The scheduler only retrieves records of all predecessor tasks that have been submitted and whose task status is pending, and prevents the same task from being retrieved by two executors at the same time through conditional updates with version numbers.

[0134] Replay tasks are scheduled in the order of referenced object before referenced object, field level before data level, and data level before block level. All granularities use copy-on-write without in-situ overwriting of committed temporary groups: field-level tasks are copied to the new version temporary group using the old commit group as a read-only baseline, only rewriting the corresponding fields and recalculating the ancestor summary node; data-level tasks re-execute the S3 transformations involved in the data layer in the new version temporary group; block-level tasks create a new temporary group according to the original block plan and re-execute S3; tasks not belonging to a multi-member cyclic replay group, after independently passing S4 readback verification on the new temporary group, enter the S4 authorized replay condition atomic replacement with the saved old record generation and content fingerprint; those belonging to... The task of the circular replay group is only written to the prepared slots specified in the corresponding group transaction record. The visibility of the entire group is controlled by the aforementioned group commit pointer, and no member-level formal record replacement is performed. When the recovery program reads the group commit pointer, if it still points to the old group version, it deletes the aborted prepared slot with a complete content fingerprint or continues to fill in the prepared members. If it has atomically pointed to the new group version, it checks the new group list and all member fingerprints, then adds the transaction status to committed and reclaims the old group version. If the pointer value or the new group list is corrupted, the old group version is restored from the group pointer versioned backup. Therefore, in the event of any interruption, other groups will only read the complete old version or the complete new version, and will not read mixed generations.

[0135] After each replay round, the module rechecks the fingerprint of the complete transformation range list, requiring that the parsing failure set, the unmapped blocking set, and the block plan failure set are all empty. It then performs the difference between the complete block plan and the committed block ledger selected by the current group commit pointer, as well as the target block readability, target digest, and schema dependency fingerprint verification. Any cyclic replay group in the preparing, prepared, or aborted state cannot be considered a completed group, nor can it meet the predecessor conditions of other groups or the root object commit gate. The next round continues when the waiting replay set is not empty. In this specific implementation, if the same block fails three times consecutively due to the same input error, automatic replay stops and the diagnostic state is retained. After the input is repaired, it resumes along the original session identifier. Only when all three predecessor failure sets are empty, the waiting replay set is empty, all cyclic replay groups have committed group commit pointers with correct content fingerprints, and each expected item in the complete transformation range has a corresponding block in the plan and a committed ledger record, does it proceed to root object commit.

[0136] To determine whether the recovery status has progressed between two rounds, the module connects the current block identifier to be replayed, the task status, and the difference node identifier in ascending order of block identifier and calculates a 32-byte recovery status summary. If the summaries of two consecutive rounds are the same and no task has changed from running or pending verification to submitted, the session is set to the waiting input repair state without idling. If the summary changes, the dimensionless round count is incremented and verification continues. The round count is only used for diagnosis and does not participate in semantic summary.

[0137] Before the root object is committed, the module re-verifies the complete transformation range list fingerprint and confirms that the parsing failure set, the unmapped blocking set, and the block plan failure set are all empty. Then, it parses each path-related reference edge token in the object reference closure to a unique canonical object identity. It aggregates ledger records according to the canonical object identity and connects the target block digests in ascending order of the minimum target axis ordinal and parent block identifier byte values ​​to calculate the object root digest. The generated target root object list records the target canonical object identity, target format version, logical axis fingerprint, axis ordinal mapping table content fingerprint, data layer role set, object root digest, all reference edge tokens and their referenced object canonical identities, parent block identifier sets of each object, digest rule version, and list content fingerprint. It also checks that each canonical object in the complete transformation range has a complete committed parent block set, the parent block target axis ordinal interval continuously covers the complete axis, the internal fragments of each parent block have been assembled, and each reference edge has a target landing point.

[0138] After integrity checks, the module creates a candidate AnnData object outside the release directory. It calculates the sum of the number of non-zero elements preceding each parent block in ascending order of the target cell starting sequence registered in the block plan, using this as a global non-zero prefix. If any parent block's target cell sequence interval overlaps, breaks, or conflicts with the parent block's identifier order, assembly is blocked. The module pre-allocates standard datasets for data, indices, and indptr under the target X path. The data and indices of each parent block are written according to the block plan offset corresponding to the target cell sequence. Each target column sequence in indices must be back-looked up to find the unique gene token in the axis sequence mapping table. The first parent block writes all local indptrs. Subsequent parent blocks skip the first item of the local indptr and add the global non-zero prefix of the parent block to each of the remaining items before writing, thus forming a formal indptr with a length equal to the number of cells plus one, a starting value of zero, and a last value equal to the total number of global non-zero elements. Parent blocks containing internal execution fragments only participate in the above assembly after the fragment has been closed into a parent block indptr; fragment local pointers are not written to the candidate object.

[0139] The obs records of candidate AnnData objects are written in ascending order of the target cell order in the axis order mapping table, and the var records are written in ascending order of the target gene order in the same mapping table to the unique standard path registered in the semantic constraint list. Axis data, reduced-dimensional data, graph data, spatial data, and object reference records are all first queried in the same axis order mapping table with stable logical tokens, and then written to the obsm, varm, layers, obsp, varp, uns, or other unique landing points registered in the list according to the obtained target row order, target column order, and the identity and field role of the corresponding canonical object. It is forbidden to directly replace the target axis order with the token byte order in ascending order. Duplicate records with the same stable token are deduplicated only when the type normalization value digest is the same. If the digests are different, assembly is blocked. Reference edges are simultaneously written to the mapping from the canonical object identity to the target path to ensure that each reference target exists and has only one readable landing point.

[0140] After assembly, the module is closed and the candidate AnnData files and their directories are refreshed. The candidate objects are reopened using the AnnData reader and pattern validator that are compatible with the format capability table version. The global shape, the length of the three arrays under the X path, the monotonicity and final value of indptr, the range of indices, the uniqueness of the obs and var primary keys, the fingerprint of the axis ordinal mapping table, the target cell ordinal of obs corresponding to each row of X, the target gene ordinal of var corresponding to each column of X, the ordinal consistency of each axis-related sub-item, the standard path of each sub-item, the object reference landing point, the object root summary and reference closure. If any check fails, no candidate root pointer is generated. The relevant parent block and its dependent blocks are added to the block-level replay set. Only when all checks pass is the candidate AnnData object recognized as a publishable object.

[0141] After the candidate object passes the reader and pattern verification, the module closes all target temporary group handles, moves the target root object list and candidate AnnData objects to the same release directory; before release, it reads the full byte value, version number and content fingerprint of the previous official target root pointer and writes them to the versioned backup file in the same directory, refreshes the backup file and directory, and records the absence of a previous version flag if there is no previous official root pointer; then it writes the candidate root pointer with the session identifier and refreshes the file and directory, and then uses atomic substitution to release the candidate root pointer as the official target root pointer. The official root pointer points to the target AnnData object and its root list that has been assembled according to the standard path. After release, the target object is reopened and the object root digest and reference closure are checked;

[0142] Before the official release, a bidirectional difference set will be performed between the block identifier set in the block plan and the block identifier set in the target root object list, requiring both difference sets to be empty. The number of blocks, the total number of non-zero elements, the number of axis data records, and the number of referenced objects for each object will be compared with the semantic constraint list item by item. The quantity will only be compared with the same type of dimensionless count. If any quantity is inconsistent, the release will be blocked and all blocks of the corresponding object will be added to the block-level replay set.

[0143] When an atomic release fails, the previous official target root pointer is not changed. The committed block ledger, target temporary block, target root object list, and failure reason are retained for the next S5 verification. When an atomic replacement succeeds but verification fails after release, if a previous version exists, the versioned backup is written to the recovery temporary pointer. After refreshing the files and directories, the previous official target root pointer is restored by atomic replacement with the same file system, and its version number and content fingerprint are verified again. If a previous version does not exist, the official entry atomic is switched to an unreleased state and external reading is prohibited. After the recovery is completed, the new candidate corresponding block is added back to the replay set. Only when both release and verification are successful is the session state set to completed and the target single-cell object that has passed block-by-block semantic verification is output.

[0144] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

[0145] This invention uses a semantic constraint list and stable logical tokens as intermediate representations between the source and target formats, unifying the physical layout of sparse matrices, logical axes, data layer roles, metadata types, and object reference relationships into a verifiable logical semantic space. Furthermore, it ensures that the source block digest and the target readback digest adopt the same logical token order and type normalization rules, thus enabling the generation of locatable verification results for axis misalignment, annotation mismatch, and field or slot loss caused by format differences.

[0146] This invention changes the traditional processing structure of unified inspection after transformation, embedding a hierarchical semantic summary tree, a transpose isomorphic reuse path controlled by semantic constraints, and a verification submission ledger into the block-by-block transformation process; the summary difference is used to select the replay granularity, and the ledger status is used to control the recovery and root object publication, so that the low-reorder path is still subject to semantic verification constraints, and the processing scope after the transformation is interrupted or the rules change is limited to the affected fields, data layers, blocks and their referenced dependent objects.

Claims

1. A single-cell data format conversion method based on intermediate representation, characterized in that, include: S1. Parse the format version, logical axis, sparse layout, data type, data layer and object reference relationship of the source single cell object, generate a list of semantic constraints, and generate stable logical tokens for cells, genes, data layers and referenced objects. S2. Determine the variable block boundaries based on the semantic constraint list, sparse matrix pointer array, numerical width, index width, buffer usage, and memory limit, and generate a block conversion plan; S3. Read the source block according to the block transformation plan, generate a source block semantic summary that is independent of the sparse layout based on the stable logical token and type normalization value, select transpose isomorphic three array reuse or memory-constrained rearrangement according to the logical axis and physical storage direction of the source object and the target object, and write the transformation result into the target temporary block. S4. Read back the target temporary block and generate the target block semantic digest according to the rules for generating the source block semantic digest. Compare the semantic digests and associated semantic states of the source block and the target block according to the semantic constraint list. When the verification is successful, atomically register the corresponding block to the committed block ledger. S5. Verify the committed block ledger, determine the set to be replayed based on unregistered blocks, digest differences and schema dependencies, perform transformation and verification on the set to be replayed, and submit the target root object after all blocks are registered to the committed block ledger.

2. The single-cell data format conversion method based on intermediate representation according to claim 1, characterized in that, S1 includes: Taking the source single-cell object as input, the format version and object topology are extracted from the file header, object class identifier, slot description, and reference entry of the source single-cell object; the logical axis corresponding to each physical dimension is determined according to the cell identifier sequence and gene identifier sequence, and the data layer name is mapped to the data layer role; the semantic constraint list records the logical axis direction, sparse layout, data layer role, metadata type, format version compatibility relationship, and object reference closure; the stable logical tokens of cells and genes are jointly generated by the object namespace, logical axis role, and intra-axis identifier; the stable logical tokens of the data layer are jointly generated by the object namespace, data layer role, and data layer identifier; and the stable logical tokens of referenced objects are jointly generated by the object namespace, object category, object identifier, and reference path.

3. The single-cell data format conversion method based on intermediate representation according to claim 1, characterized in that, S2 includes: The number of non-zero elements in adjacent logical axis intervals is accumulated along the sparse matrix pointer array. The storage amount of non-zero elements is calculated according to the numerical width and index width. The buffer usage required for pointer array, sorting, rearrangement, digest calculation and target writing is superimposed. When the estimated usage does not exceed the memory limit, a block is formed. When the estimated usage of a logical axis interval that cannot be further divided along the logical axis boundary exceeds the memory limit, it is further divided into sub-blocks along the non-zero element sequence. If the target format does not allow the sub-block to be written independently, the corresponding logical axis interval is marked as conversion failure. When the upper bound of the global index value exceeds the representation limit of the target index type, if the target format allows intra-block indexing, the global index value is mapped to an intra-block index value and the block offset is recorded. If the target format does not allow intra-block indexing, an index type whose representation range covers the upper bound of the global index value is selected. If there is no index type that meets the representation range, it is marked as verification failure.

4. The single-cell data format conversion method based on intermediate representation according to claim 1, characterized in that, Step S3, generating a semantic summary, includes: generating a leaf summary using the stable logical tokens corresponding to the logical coordinates, the data layer role, the type identifier, and the type normalization value; merging the leaf summaries into a block summary according to a predetermined logical token order; and forming a hierarchical semantic summary tree level by level. The hierarchical semantic summary tree is set with sub-summary nodes representing the data layer, axis data, dimensionality reduction data, graph data, spatial data, and object references. The parent node is generated by its child node identifier and child node summary according to a predetermined node order.

5. The single-cell data format conversion method based on intermediate representation according to claim 4, characterized in that, Step S3 involves selecting transpose isomorphic three-array reuse, which includes: when the physical storage directions of the source sparse matrix and the target sparse matrix are transpose correspondences, verifying the structural correspondence of the pointer array, index array, and value array, and verifying the logical axis fingerprint, the merging rules of repeated logical coordinates, the explicit zero retention rules, the value width, the index width, and the data layer role; when all verification results satisfy the semantic constraint list, transpose isomorphic reuse is performed on the three arrays; otherwise, the non-zero elements are rearranged in segments within the memory limit according to the stable logical token; both transformation results are entered into the readback verification in step S4.

6. The single-cell data format conversion method based on intermediate representation according to claim 4, characterized in that, S4 includes: reading back the logical axis identifier, non-zero element logical coordinates, data layer role, metadata type, format version, and object reference identifier from the target temporary block; performing the same type normalization rule as the source block on the read-back data, and arranging the normalization results according to the stable logical token and logical coordinate order to generate the root digest and sub-item digests of the target block; selecting the corresponding sub-item digests as necessary sub-item digests from the hierarchical semantic digest tree according to the semantic constraint list and the object categories contained in the current block; when the root digest, the necessary sub-item digests, logical axis fingerprints, data layer roles, metadata types, format versions, and object reference closures are all consistent, generating a ledger record containing the belonging object identifier, block identifier, source block digest, target block digest, digest rule version, schema dependency fingerprint, and commit status, and atomically updating the committed block ledger by replacing the formal record with a temporary record.

7. The single-cell data format conversion method based on intermediate representation according to claim 6, characterized in that, S5 includes: Verify the readability of the target block, the target block digest, and the schema dependency fingerprint of the ledger records in the submitted block ledger; Blocks with missing ledger records, blocks with ledger records but unreadable target blocks, abnormal blocks with inconsistent target block digests, and blocks whose pattern-dependent fingerprint correspondence rules have changed are added to the abnormal recovery replay set. The field-level, data-level, or block-level replay granularity is determined based on the lowest-level node that produces the difference in the hierarchical semantic summary tree, and the target objects that depend on the difference node are added to the anomaly recovery replay set along the object reference relationship.

8. The single-cell data format conversion method based on intermediate representation according to claim 2, characterized in that, The semantic constraint list also records the field name, data type, missing value representation, category value set, and correspondence with stable logic tokens of axis metadata fields, and records the object class, slot name, slot type, and slot reference target of RDS objects and S4 objects; when converting axis metadata and object slots, the records are aligned according to the stable logic tokens, and fields, slots, or reference targets that are not mapped in the semantic constraint list are marked as verification failures and are not registered in the corresponding blocks.

9. The single-cell data format conversion method based on intermediate representation according to claim 6, characterized in that, The target root object commit includes: establishing a mapping between object identifiers and corresponding block identifier sets based on the object identifiers in the ledger records; generating a target root object list when all ledger records for each target temporary block are in a committed state and all block identifiers corresponding to each referenced object in the object reference closure have committed ledger records; the target root object list records the target object identifier, format version, logical axis fingerprint, data layer role set, object root digest, and referenced object identifier; the object root digest is generated by the target block digests in all ledger records of the corresponding object in block identifier order; and publishing the target root object list and the target temporary block set as the target object through atomic replacement. In the event of a publishing failure, the committed block ledger and the target temporary block are retained for verification in step S5.

10. A single-cell data format conversion system based on intermediate representation, used to execute the single-cell data format conversion method based on intermediate representation as described in any one of claims 1 to 9, characterized in that, include: The parsing and identification module is used to parse the source single-cell object and generate a list of semantic constraints and stable logic tokens. The block planning module is used to determine variable block boundaries and generate block transition plans. The block transformation module is used to generate a semantic digest of the source block, perform transpose isomorphic three-array reuse or memory-constrained rearrangement, and write it to the target temporary block; The readback verification module is used to generate a semantic summary of the target block, perform semantic comparisons, and update the ledger of committed blocks; The recovery and submission module is used to determine the set to be replayed and submit the target root object after all blocks have been verified.