Formal modeling and verification method of a TLA+ based vector database Milvus

CN122195961BActive Publication Date: 2026-09-15NANJING UNIV OF INFORMATION SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610640315.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-05-11
Publication Date
2026-09-15
Estimated Expiration
2046-05-11

AI Technical Summary

Technical Problem

参考Raft协议的形式化验证经验可知,即使开发者使用TLA+等工具对协议设计进行了验证,实际系统在实现过程中仍可能违背协议设计并引入复杂缺陷

Benefits of technology

[0012]Beneficial Effects: Compared with existing technologies, this invention has the following significant advantages: This invention compensates for the shortcomings of traditional testing methods in vector database verification through formal verification, enabling efficient and comprehensive testing of the Milvus system, which is of great value and significance for improving Milvus reliability. Within a single TLA+ framework, it formalizes four core semantics: strong consistency, finite expiration consistency, session consistency, and eventual consistency, providing precise and unambiguous formal definitions for different consistency levels. By simulating a sharded migration state machine (idle, replicating, ready switching) and imposing maximum concurrency constraints and state transition conditions on the migration process, it provides formal guarantees for data migration during load balancing, preventing data loss or inconsistency due to migration logic defects. By setting key invariants such as StateBasedStrongConsistency and using an exhaustive verification method with a TLC model checker, it covers all possible execution interleaved paths, including message out-of-order delivery, node latency, and migration interference, far exceeding the scope of traditional testing, proving the correctness of the timestamp-based asynchronous replication mechanism. Taking into account the interaction between access control and data consistency, and verifying that all successful read operations meet role-based access control requirements through invariants such as SecurityInvariant, a formal model for integrated security verification is established.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122195961B_ABST
    Figure CN122195961B_ABST
Patent Text Reader

Abstract

The application discloses a formal modeling and verification method of a vector database Milvus based on TLA+, and comprises the following steps: TLA+ abstract modeling of the vector database Milvus, definition simulation of basic behaviors of the Milvus, and then imitation of a conventional workflow. Firstly, a variable set is defined according to the operation principle of the Milvus system; then, an authorization process is modeled to ensure the legality of all operations; then, a client write operation behavior is simulated to write data; finally, a state machine model containing four consistency semantics, i.e., strong consistency, limited expiration consistency, session consistency and eventual consistency, is constructed, and a migration state machine and a maximum concurrent migration constraint are introduced; safety properties are defined through formal specification, including authorized access, node load balancing and migration data integrity, and the model is run to detect whether the database meets the safety properties under all execution paths; and the application improves the credibility and safety.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of database modeling and verification technology, and in particular to a formal modeling and verification method for the vector database Milvus based on TLA+. Background Technology

[0002] Milvus, an open-source project, is one of the most popular vector databases. It employs a sharding mechanism to horizontally scale high-dimensional vector data and uses asynchronous replication to synchronize data between nodes. Milvus supports four adjustable consistency levels: strong consistency, limited expiration consistency, session consistency, and eventual consistency. These consistency models are implemented using a timestamp mechanism; query nodes wait for all data before the corresponding consistency timestamp to become available before executing a query, depending on the required consistency level. However, Milvus's distributed architecture involves complex mechanisms such as data sharding, segment management, multi-replica synchronization, and fault recovery, and the correctness of its consistency guarantee implementation lacks formal verification. Although Milvus has clearly defined consistency levels, there is a significant gap between high-level design and code implementation. Referring to the formal verification experience of the Raft protocol, even if developers use tools such as TLA+ to verify the protocol design, the actual system may still violate the protocol design and introduce complex defects during implementation. Furthermore, exhaustive model testing for large-scale distributed systems like Milvus faces serious scalability bottlenecks. Summary of the Invention

[0003] Purpose of the invention: The purpose of this invention is to provide a formal modeling and verification method for the vector database Milvus based on TLA+, to detect whether Milvus' node load balancing, authorized access, and shard migration processes are secure, and whether data consistency is guaranteed.

[0004] Technical solution: The formal modeling and verification method based on the TLA+ vector database Milvus, as described in this invention, includes: Step 1: Analyze the code and design of the vector database Milvus, abstract and construct its state machine model for data migration and consistent read functions respectively, and define the set of variables used to maintain the database's running state; Step 2: Set the maximum number of concurrent migrations for the shard migration process, and introduce idle state, copying state and ready to switch state to describe the life cycle of a single shard migration, and limit the number of concurrent migration tasks in the system through state transitions; Step 3: Provide four semantics for database read operations: strong consistency, limited expiration consistency, session consistency, and eventual consistency, and set corresponding comparison constraints between node service timestamps and global write timestamps for each semantic. Step 4: Use the model checker to perform formal verification of the constructed TLA+ model. By setting invariants to check node load balancing, authorized access security, and migration data integrity, verify that the model satisfies the invariants in all execution paths.

[0005] Furthermore, in step 1, the variable set includes: global timestamp, message queue, shard location mapping, node status, vector main storage, local data replicas of each node, user roles, role permissions, client session timestamp, migration task status and target, and replication watermark.

[0006] Furthermore, in step 2, before changing the shard migration status from idle to replication, the following steps are also included: confirming that the current shard is in an idle state and that the shard data is consistent in the source node and the vector main storage; confirming that the target node and the source node are different nodes and that the load of the source node is greater than the load of the target node.

[0007] Furthermore, in step 2, when the shard migration is in the ready-to-switch state, the shard location mapping is updated, the shard status is reset to the idle state, and the load records of the source node and the target node, as well as the local data copy of the target node, are updated.

[0008] Furthermore, in step 3, the comparison constraint for strong consistency semantics is specifically: the service timestamp corresponding to the node providing the read service is not less than the global last write timestamp.

[0009] Furthermore, in step 3, the comparison constraint of the finite expiration consistency semantics is specifically as follows: the lower limit of reading is calculated according to the preset latency tolerance window, and the service timestamp corresponding to the node providing the read service is not less than the lower limit of reading.

[0010] Furthermore, in step 3, the comparison constraint of session consistency semantics is as follows: before the read operation is executed, the service timestamp corresponding to the node providing the read service is not less than the client session timestamp; and after the read operation is executed, the client session timestamp is updated to the maximum value between the client's current session timestamp and the global last written timestamp.

[0011] Furthermore, in step 4, the invariants set include: a first invariant used to verify the legality of all variable values; a second invariant used to verify that any successful read operation has the corresponding permissions; a third invariant used to verify that the data of the node that has been processed up to the latest written data is consistent with the latest global value; a fourth invariant used to verify that at least one node holds the latest data during the migration process; and a fifth invariant used to verify that the load difference between any two nodes does not exceed a preset threshold.

[0012] Beneficial Effects: Compared with existing technologies, this invention has the following significant advantages: This invention compensates for the shortcomings of traditional testing methods in vector database verification through formal verification, enabling efficient and comprehensive testing of the Milvus system, which is of great value and significance for improving Milvus reliability. Within a single TLA+ framework, it formalizes four core semantics: strong consistency, finite expiration consistency, session consistency, and eventual consistency, providing precise and unambiguous formal definitions for different consistency levels. By simulating a sharded migration state machine (idle, replicating, ready switching) and imposing maximum concurrency constraints and state transition conditions on the migration process, it provides formal guarantees for data migration during load balancing, preventing data loss or inconsistency due to migration logic defects. By setting key invariants such as StateBasedStrongConsistency and using an exhaustive verification method with a TLC model checker, it covers all possible execution interleaved paths, including message out-of-order delivery, node latency, and migration interference, far exceeding the scope of traditional testing, proving the correctness of the timestamp-based asynchronous replication mechanism. Taking into account the interaction between access control and data consistency, and verifying that all successful read operations meet role-based access control requirements through invariants such as SecurityInvariant, a formal model for integrated security verification is established. Attached Figure Description

[0013] Figure 1 This is a flowchart illustrating the formal modeling and verification process of the TLA+-based vector database Milvus according to the present invention. Figure 2 This is the fragmentation migration state transition diagram of the present invention; Figure 3 This is a flowchart illustrating the logic for determining the four types of consistent reads in this invention. Detailed Implementation

[0014] The technical solution of the present invention will be further described below with reference to the accompanying drawings.

[0015] like Figure 1 As shown, embodiments of the present invention provide a formal modeling and verification method based on the TLA+ vector database Milvus, including: (1) Analyze the code and design of the vector database Milvus, and abstract and construct its state machine model for data migration and consistent read functions respectively, and define the set of variables used to maintain the database running state; (2) Set the maximum number of concurrent migrations for the shard migration process, and introduce idle state, copying state and ready to switch state to describe the life cycle of a single shard migration, and limit the number of concurrent migration tasks in the system through state transitions; (3) Provide four semantics for database read operations: strong consistency, limited expiration consistency, session consistency and eventual consistency, and set corresponding comparison constraints between node service timestamp and global write timestamp for each semantic; (4) Use the model checker to perform formal verification of the constructed TLA+ model. By setting invariants to check node load balancing, authorized access security and migration data integrity, the model is verified to satisfy the invariants in all execution paths.

[0016] The specific process is as follows: Step 1: Define a set of variables based on the principles of Milvus, including: global timestamp, message queue, last write timestamp, user role, role permissions, audit log, client session timestamp, vector main storage, node local data replica, shard location mapping, node status, migration status, migration target, and replication watermark. Use this set of variables to describe the state of the database. Finally, these variables are initialized. The global timestamp is an integer with an initial value of 1; the message queue is a sequence with an initial value of an empty sequence; the last written timestamp is a function, representing the last write timestamp for each set, initially all sets have timestamp values ​​of 0; the user role is a function, initially each user role in the role set is empty (i.e., no role); the role permission is a function, initially each role permission in the role set is empty (i.e., no permission); the audit log is a set, initially an empty set; the client session timestamp is a function, initially the set storing each client session timestamp is empty (i.e., each client's current logical clock or session timestamp), initially 0; the vector main storage is a nested function, initially each set is mapped to an initial vector timestamp mapping, this inner mapping displays the vector's timestamp, which is 0; the node local data replica is a three-level nested function, initially each node is mapped to a set mapping, this set mapping further maps each set to an initial vector timestamp mapping, this inner mapping displays the vector's timestamp, which is 0. The initial value is 0; the type of the shard location mapping is a function. Initially, there exists a mapping function from all shards to all nodes as the initial location allocation, making the shard location variable equal to this initial allocation. This is a non-deterministic initialization, meaning that the shard location can be any mapping function from the shard set to the node set. That is, when the system starts, shards can be distributed on any node, as long as they conform to the type definition. The type of the node state is a function. Initially, each node is mapped to a state record, which contains a service timestamp initialized to 0, and a status record based on the initial shard location allocation and the current status. The load value calculated by the node indicates that the node's state includes the service timestamp and load status; the migration state type is a function, and initially, each shard is mapped to its migration state "idle", meaning that the initial migration state of each shard is "idle"; the migration target type is a function, and initially, each shard is mapped to an empty set, which is the set of target nodes for the shard migration. Initially, it is an empty set, indicating that there are no target nodes for migration; the replication watermark type is a function, and initially, each shard is mapped to a replication watermark with a value of 0, indicating that the replication watermark of the initial shard is 0.

[0017] Step 2: Write the authorization access functions, including the following atomic operations: an "Assign Roles" function to assign one or more roles to each user; and an "Assign Permissions" function to grant one or more <set, permission> pairs to each role. Also, set up an auxiliary function: a "Permission" function, to query whether a user's roles possess the corresponding operation permissions for a specific set.

[0018] Step 3: Write the write operation for the client response, including the following atomic operations: Before writing, check if the user has write permissions for the set. After confirming permissions, update the global timestamp, add the written message to the message queue, update the last write timestamp of the corresponding set, update the record in the audit log and the information in the vector main storage, that is, generate a version identifier based on the global timestamp and write the version identifier to the vector main storage, then append the message containing the set identifier and the version identifier to the message queue, and update the last write timestamp and the audit log.

[0019] Step 4: The node synchronization action includes the following atomic operations: Retrieve the first message from the message queue; compare the node's current service timestamp with the message's timestamp; if the timestamp is less, update the node's current service timestamp; update the message queue to include all remaining elements except the first element (i.e., all message records except those already synchronized); finally, synchronize the data to the node's local data copy. One new message is synchronized at a time until the message queue is empty.

[0020] Step 5: Fragment migration section as follows Figure 2 As shown, there are three shard migration states: idle, replicating, and ready. The three migration states change under different conditions, and the specific steps include the following operations: To limit the number of migration tasks in the non-idle state of the system, you can first set a maximum number of concurrent migrations, and then set a function to count the number of non-idle shards in the current system; compare the current number of non-idle shards with the maximum number of concurrent migrations.

[0021] Before starting the migration operation, it is necessary to confirm that: the current shard is in an idle state, and the shard data in the source node is consistent with the vector main storage data; the target node and the source node are not the same node, and the load of the source node is greater than the load of the target node, with a load difference of more than 2. Then, the shard is converted to a replication state.

[0022] The data replication operation is a repeatable incremental synchronization action used to copy data from a specified shard 's' from the source node to the target node, continuously catching up on new writes generated during the migration. The operation first obtains the source node, target node, shard set, global latest version number, and the currently synchronized replication watermark. Then, it enters a branch decision: if the replication watermark is less than the global latest version number, it indicates incremental data exists. In this case, the data in the corresponding set on the target node is directly updated to the global latest version, and the replication watermark is synchronized to the global latest version number, while the migration status remains "replicating" to allow for continued catching up. If the replication watermark is greater than or equal to the global latest version number, it indicates that the data has been caught up. In this case, the migration status is switched to "ready for switch," and the node data remains unchanged. Throughout the execution process, the data replication operation does not modify variables such as the global timestamp, message queue, last write timestamp, user role, role permissions, audit logs, client session timestamps, vector main storage, shard location mapping, node status, and migration target, thus ensuring a single responsibility for the operation. Through this cyclic incremental synchronization mechanism, the data replication operation can ensure that new writes generated during the migration are not lost, and the switchover preparation state is only allowed after the data is fully synchronized, providing a correctness prerequisite for the subsequent completion of the migration operation.

[0023] When the migration operation is completed, it is necessary to first confirm that the current shard is in a ready-to-switch state and whether the latest data on the target node is consistent with the latest data in the vector main storage. If they are inconsistent, the subsequent operations are stopped until the conditions are met after data synchronization and data copying operations are performed again, in order to prevent new writes during the migration and maintain data consistency. Then, the shard location mapping is updated, the load records in the source node and target node status are updated, the shard migration status is changed back to idle, allowing future migrations, and the migration target is cleaned up.

[0024] Step 6: The four types of consistent reads include the following atomic operations: A limited version lag number K is set for finite expiration consistency; before all read operations, a permission function must be called to check if the user has read permissions for the node, and the target node position to be read is obtained through shard position mapping; after each read, the audit log must be updated. The specific logic for detecting and judging the four types of consistency is as follows: Figure 3As shown, the requirements include the following: Strong consistency requires the service node's timestamp to be greater than or equal to the last write timestamp; finite expiration consistency reads require setting a finite expiration lower bound, i.e., Max(0, last write timestamp - latency tolerance window), and then comparing the service node's timestamp to ensure it is greater than or equal to the finite expiration lower bound; session consistency requires maintaining the client session timestamp, requiring that before a read operation: the node requesting service is compared after a session write operation, i.e., the node's service timestamp is greater than or equal to the client session timestamp; after a read operation: the client session timestamp is updated to Max(the client's current session timestamp, last write timestamp). Eventual consistency has no special constraints.

[0025] Step 7: Write the Next code, which chains together all possible operations, meaning the system can only atomically execute one of these operations at any given time. This constructs a concurrent state space, allowing these operations to be executed in any interleaved order, thereby simulating complex concurrent scenarios in a real distributed environment.

[0026] Step 8: Set invariants: TypeOK (type correctness invariant) checks all variables; SecurityInvariant (security invariant) checks that the user has permission for any successful read operation; StateBasedStrongConsistency (state-based strong consistency invariant) checks that if a node has processed the latest write time, its data must be equal to the globally latest value; MigrationSafety (migration safety invariant) checks that at least one node holds the latest data during migration; BoundedImbalance (load balancing invariant) checks that the load difference between nodes does not exceed 2. Then, use the TLC model checker to formally verify the written TLA+ model, inputting the above invariants to check for violations and ensure that all execution paths can satisfy the security properties of node load balancing, authorized access, and migration data integrity.

[0027] Step 9: Finally, constrain the model rows in the TLC detector, requiring global timestamp <= 4; message queue length <= 2; number of audit logs <= 3 to prevent state explosion; and configure the model to include: maximum concurrent migrations = 1; only Client1 as the client; K = 2; a total of S1 and S2 as shards; only C1 as the set; both shards are in set C1; the permission set includes read permissions and write permissions; roles include administrator and reader; users include u1 and u2.

Claims

1. A formal modeling and verification method based on the TLA+ vector database Milvus, characterized in that, include: Step 1: Analyze the code and design of the vector database Milvus, abstract and construct its state machine model for data migration and consistent read functions respectively, and define the set of variables used to maintain the database's running state; Step 2: Set the maximum number of concurrent migrations for the shard migration process, and introduce idle, replicating, and ready-to-switch states to describe the lifecycle of a single shard migration. Limit the number of concurrent migration tasks in the system through state transitions. Before the shard migration state is changed from idle to replicating, the following steps are also taken: confirm that the current shard is in an idle state and that the shard data is consistent in the source node and the vector main storage; confirm that the target node and the source node are different nodes and that the load of the source node is greater than the load of the target node; when the shard migration is in the ready-to-switch state, update the shard location mapping, reset the shard state to idle, and update the load records of the source node and the target node, as well as the local data copy of the target node. Step 3: Provide four semantics for database read operations: strong consistency, limited expiration consistency, session consistency, and eventual consistency. For each semantic, set corresponding comparison constraints between the node service timestamp and the global write timestamp. Specifically, the comparison constraint for strong consistency is: the service timestamp of the node providing read services is not less than the global last write timestamp. The comparison constraint for limited expiration consistency is: based on a preset latency tolerance window, the read lower limit is calculated, and the service timestamp of the node providing read services is not less than the read lower limit. The comparison constraint for session consistency is: before a read operation is executed, the service timestamp of the node providing read services is not less than the client session timestamp; and after the read operation is executed, update the client session timestamp to the maximum value between the current client session timestamp and the global last write timestamp. Step 4: Perform formal verification of the constructed TLA+ model using the model checker. By setting invariants to verify node load balancing, authorized access security, and migration data integrity, verify that the model satisfies the invariants in all execution paths. The set invariants include: a first invariant to verify the legality of all variable values; a second invariant to verify that any successful read operation has the corresponding permissions; a third invariant to verify that the data of nodes that have processed the latest written data is consistent with the latest global value; a fourth invariant to verify that at least one node holds the latest data during the migration process; and a fifth invariant to verify that the load difference between any two nodes does not exceed a preset threshold.

2. The formal modeling and verification method based on the TLA+ vector database Milvus according to claim 1, characterized in that, In step 1, the set of variables includes: global timestamp, message queue, shard location mapping, node status, vector main storage, local data replicas of each node, user roles, role permissions, client session timestamp, migration task status and target, and replication watermark.

Citation Information

Patent Citations

  • Formalized modeling and verification method and system for distributed database

    CN116257590A

  • Methods of Measuring Consistability of a Distributed Storage System

    US20100192018A1