Enterprise attribute zero-knowledge proof method and system based on dynamic positioning anchor point

CN122824409APending Publication Date: 2026-09-25SOUTHWEST JIAOTONG UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202611034899.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-13
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

然而,这类方案在处理企业身份认证等私密数据时存在天然缺陷:为了对数据真实性达成共识,所有预言机节点通常需要读取明文数据,这直接导致了企业核心隐私的泄露

Benefits of technology

[0046]1、本发明提供的基于动态定位锚点的企业属性零知识证明方法,通过动态定位锚点机制,将JSON格式的权威响应明文数据中的目标属性提取转化为独热编码约束、模式匹配约束和内积提取约束的组合,利用布尔指示向量的独热性质将属性定位问题转化为线性扫描与内积选择。具体而言,独热编码约束强制整个布尔指示向量中恰好有一个分量为1,从而唯一确定锚点字符串在明文中的起始索引;模式匹配约束验证该起始索引处的字节序列与锚点字符串逐字节吻合,确保定位的准确性;内积提取约束利用布尔指示向量的选择特性,通过加权求和将锚点后紧跟的字节序列抬升为电路中间变量,结合布尔掩码向量截取真实有效长度的目标字节序列。上述机制避免了传统正则表达式匹配或有限状态机解析所需的复杂电路结构,使得电路约束总量与明文长度呈线性关系,而非随状态空间指数级膨胀,显著降低了零知识证明的生成开销,解决了大规模非结构化数据上链时电路规模爆炸的问题,实现了高效的企业属性隐私提取与验证。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122824409A_ABST
    Figure CN122824409A_ABST
Patent Text Reader

Abstract

The application discloses an enterprise attribute zero-knowledge proof method and system based on dynamic positioning anchor points, relates to the technical field of blockchains, and has the technical scheme as follows: obtaining authoritative response plaintext data and a digital signature thereof; performing arithmetic preprocessing on the plaintext data to generate an arithmetic vector; extracting a target attribute value based on dynamic positioning anchor points, locking an anchor point string starting index through one-hot encoding constraint, verifying an anchor point byte sequence through pattern matching constraint, and extracting an attribute value through inner product constraint combined with a Boolean mask vector; applying validity check constraint to the attribute value; constructing a rank-1 constraint system circuit based on the constraint relationship that passes the check, and generating a zero-knowledge proof. The application degrades field extraction into linear scanning and inner product selection, and the total amount of circuit constraints and the length of plaintext are in a linear relationship, thereby avoiding the state explosion problem of traditional schemes and realizing efficient enterprise attribute privacy extraction and on-chain verification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of blockchain technology, and more specifically, to a method and system for zero-knowledge proof of enterprise attributes based on dynamic positioning anchors. Background Technology

[0002] With the widespread application of blockchain technology in supply chain finance, digital copyright protection, and electronic evidence preservation, smart contracts, as the execution carriers of blockchain code as law, rely on deterministic logic within a closed network for operation. However, the blockchain network is essentially a closed environment isolated from the outside world; smart contracts cannot actively perceive or obtain data from the real world outside the chain, such as business registration information and financial statements. To overcome this barrier, blockchain oracles have emerged, responsible for securely and reliably inputting external data onto the chain. In enterprise collaboration scenarios, verifiers often need to verify the business qualifications of partner companies, such as whether the registered capital meets the entry threshold and whether there are any administrative penalty records. Existing oracle solutions, such as Chainlink, mostly adopt decentralized consensus mechanisms, with multiple independent nodes jointly providing data and reaching consensus. However, these solutions have inherent flaws when handling private data such as enterprise identity authentication: to reach consensus on the authenticity of data, all oracle nodes usually need to read plaintext data, which directly leads to the leakage of core enterprise privacy. On the other hand, while centralized oracles are simple to deploy, they introduce a serious single point of failure risk. Once the central node goes down or is attacked, on-chain applications that rely on its data will face huge risks.

[0003] To address privacy concerns, academia has proposed web data proof technologies based on zero-knowledge proofs, such as the DECO protocol, which allows users to prove they have accessed a specific website and extracted specific fields without disclosing sensitive information. However, existing zero-knowledge proof-based schemes face severe performance bottlenecks in practical applications. When processing large-scale unstructured data, such as JSON-formatted business files containing hundreds of fields, traditional schemes require constructing complex finite state machines or general regular expression matching logic within the zero-knowledge proof circuit to locate and extract target fields. These schemes require evaluating the transition function on the state set at each character position, causing the circuit constraint size to grow superlinearly with the data length, and the number of states to expand exponentially with pattern complexity. When processing 10KB-level enterprise qualification data, the total number of circuit constraints often reaches millions, resulting in excessively long proof generation times, and even preventing computation on conventional hardware due to memory overflow. Furthermore, existing schemes require the introduction of pushdown stacks to handle nested structures in JSON format when dealing with nested structures, further exacerbating the expansion of circuit size and severely restricting the practical application of zero-knowledge proof technology in enterprise collaboration scenarios.

[0004] Therefore, researching and designing a zero-knowledge proof method and system for enterprise attributes based on dynamic positioning anchors that can overcome the above-mentioned defects is an urgent problem to be solved. Summary of the Invention

[0005] To address the shortcomings of existing technologies, the present invention aims to provide a method and system for zero-knowledge proof of enterprise attributes based on dynamic positioning anchor points. This method can efficiently extract target attributes from large-scale unstructured JSON data and generate zero-knowledge proofs while protecting enterprise privacy. It also ensures that the total number of circuit constraints is linearly related to the plaintext length rather than growing exponentially.

[0006] The above-mentioned technical objective of the present invention is achieved through the following technical solution:

[0007] Firstly, a zero-knowledge proof method for enterprise attributes based on dynamically located anchor points is provided, including the following steps:

[0008] Obtain authoritative plaintext response data and its digital signature for enterprise attributes;

[0009] The authoritative response plaintext data is preprocessed arithmetically to generate an arithmetic vector;

[0010] Calculate the hash of the arithmetic vector and apply consistency constraints to the digital fingerprint signed by the notary node.

[0011] Extracting target attribute values ​​from the arithmetic vector based on dynamically located anchor points, the extraction includes: presetting an anchor string for each target attribute, introducing a Boolean indicator vector and applying one-hot encoding constraints to lock the starting index of the anchor string in the plaintext, verifying the byte sequence at the starting index matches the anchor string byte by byte through pattern matching constraints, extracting the byte sequence immediately following the anchor string using the selection characteristics of the Boolean indicator vector through inner product extraction constraints, and combining the Boolean mask vector to truncate the target byte sequence of the actual effective length and reconstruct it into the target attribute value;

[0012] Apply validity checks to the reconstructed target attribute values, including at least ASCII range delimitation constraints, terminal symbol boundary constraints, range proof constraints, and Boolean satisfaction constraints.

[0013] Based on the verified constraint relationships, a rank-1 constraint system circuit is constructed, generating a zero-knowledge proof.

[0014] Furthermore, the arithmetic preprocessing includes:

[0015] The authoritative response plaintext data is mapped byte by byte to a domain element expanded vector, and the domain element expanded vector is aggregated into a compact domain element vector with a preset step size.

[0016] Furthermore, the hash calculation for the arithmetic vector specifically involves:

[0017] Calculate a hash for the compact field element vector, reducing the hash input size from the plaintext length to one-inverse of the preset step size.

[0018] Furthermore, the one-hot encoding constraint includes:

[0019] Apply a global uniqueness constraint and a Boolean constraint to the Boolean indicator vector. The global uniqueness constraint forces that exactly one component in the Boolean indicator vector is 1, and the Boolean constraint forces that each component can only take the values ​​0 or 1.

[0020] Furthermore, the pattern matching constraint specifically includes:

[0021] For any candidate starting position and offset within the anchor point, apply a constraint that the product of the component of the Boolean indicator vector and the byte difference is equal to 0, such that when the component is equal to 1, the plaintext byte and the anchor byte are equal byte by byte.

[0022] Furthermore, the inner product extraction constraint is specifically as follows:

[0023] The target attribute is preset with a maximum byte length. A weighted summation constraint of the Boolean indicator vector and the plaintext byte is applied to each byte offset. The one-hot property of the Boolean indicator vector is used to collapse the summation result into the byte value of the corresponding offset after the anchor point.

[0024] The first few bits of the Boolean mask vector are 1s and the rest are 0s, which are used to extract the target byte sequence of the actual effective length within the maximum byte length.

[0025] Furthermore, the terminal symbol boundary constraint is implemented using a polynomial root-finding structure:

[0026] Suppose that the set of JSON syntax terminators includes the ASCII codes of commas, double quotes, right curly braces, and newlines. Based on the locked start index in the Boolean indicator vector, apply a constraint that the product of the component corresponding to the start index and the terminator polynomial is equal to 0, wherein the terminator polynomial is the product of the difference between the byte immediately following the end of the attribute value and the codes of each terminator.

[0027] The constraint is satisfied when the product of the consecutive byte belongs to the set of terminal symbols; otherwise, the constraint is not satisfied to prevent truncation attacks.

[0028] Furthermore, the target attribute value includes a set of public attributes and a set of private attributes, which do not overlap;

[0029] After extracting and reconstructing the attributes in the public attribute set, the hash is calculated to obtain the public attribute hash commitment, which is then written into the on-chain identity credential.

[0030] The verifier obtains the plaintext of the public attributes off-chain, calculates the hash locally, and compares it with the on-chain hash commitment to complete the authenticity verification;

[0031] For numerical attributes in the privacy attribute set, the attribute value is proven to be no less than the threshold value declared by the verifier without exposing the specific value through range proof constraints. For administrative credit attributes in the privacy attribute set, the attribute value is proven to be equal to the compliance code without exposing the specific content through Boolean satisfaction constraints.

[0032] Furthermore, the method also includes:

[0033] Submit the zero-knowledge proof, the digital fingerprint, the digital signature, and the enterprise pledge deposit to the on-chain identity registration center to complete the registration;

[0034] The on-chain identity registration center pre-sets the validity period of the certificate and the minimum staking threshold. Enterprises need to re-execute the extraction and proof process within the validity period of the certificate to refresh the certificate. If the enterprise's qualifications deteriorate and it is unable to regenerate a valid certificate, the old certificate will be determined to be expired in real time when the verifier queries after the validity period of the certificate expires, thus realizing implicit revocation.

[0035] Furthermore, the total number of constraints in the rank-1 constraint system circuit is linearly related to the plaintext length. The single-attribute constraint scale is a constant multiplied by the plaintext length plus another constant. The constant is determined by the anchor length, the maximum byte length, and the size of the terminal symbol set, but does not include the number of finite automaton states.

[0036] Furthermore, the authoritative response plaintext data is obtained through MPC-TLS collaborative acquisition:

[0037] The proving enterprise and the collaborative notary node share the ECDHE temporary private key of the TLS handshake additively, so that the aggregated temporary public key of both parties is a standard elliptic curve point in algebraic terms to pass the verification of the authoritative data source. Both parties work together to complete AES-GCM decryption and hash calculation through Yao's obfuscation circuit and Beaver triple, and output a digital fingerprint with the signature of the collaborative notary node. The authoritative data source does not need to be modified.

[0038] Secondly, it provides a zero-knowledge proof system for enterprise attributes based on dynamic positioning anchor points, including:

[0039] The data acquisition module is used to acquire authoritative response plaintext data of enterprise attributes and their digital signatures;

[0040] The data processing module is used to perform arithmetic preprocessing on the authoritative response plaintext data to generate an arithmetic vector;

[0041] The hash verification module is used to calculate the hash of the arithmetic vector and perform consistency constraints with the digital fingerprint signed by the notary node.

[0042] An anchor point extraction module is used to extract target attribute values ​​from the arithmetic vector based on dynamically located anchor points. The extraction includes: presetting an anchor string for each target attribute, introducing a Boolean indicator vector and applying one-hot encoding constraints to lock the starting index of the anchor string in the plaintext, verifying the byte sequence at the starting index matches the anchor string byte by byte through pattern matching constraints, extracting the byte sequence immediately following the anchor string by using the selection characteristics of the Boolean indicator vector through inner product extraction constraints, and combining the Boolean mask vector to truncate the target byte sequence of the actual effective length and reconstruct it into the target attribute value.

[0043] The validation constraint module is used to apply validity validation constraints to the reconstructed target attribute values, including at least ASCII range delimitation constraints, terminal symbol boundary constraints, range proof constraints, and Boolean satisfaction constraints.

[0044] The proof generation module is used to construct a rank-1 constraint system circuit based on verified constraint relationships and generate zero-knowledge proofs.

[0045] Compared with the prior art, the present invention has the following beneficial effects:

[0046] 1. The zero-knowledge proof method for enterprise attributes based on dynamic anchor points provided by this invention transforms the extraction of target attributes from authoritative response plaintext data in JSON format into a combination of one-hot encoding constraints, pattern matching constraints, and inner product extraction constraints through a dynamic anchor point mechanism. It utilizes the one-hot property of Boolean indicator vectors to transform the attribute localization problem into linear scanning and inner product selection. Specifically, the one-hot encoding constraint forces that exactly one component in the entire Boolean indicator vector is 1, thus uniquely determining the starting index of the anchor string in the plaintext; the pattern matching constraint verifies that the byte sequence at this starting index matches the anchor string byte by byte, ensuring the accuracy of localization; the inner product extraction constraint utilizes the selection characteristics of Boolean indicator vectors, using weighted summation to elevate the byte sequence immediately following the anchor point into an intermediate variable of the circuit, and combines this with a Boolean mask vector to extract the target byte sequence of the true effective length. The above mechanism avoids the complex circuit structure required by traditional regular expression matching or finite state machine parsing, making the total number of circuit constraints linearly related to the plaintext length, rather than expanding exponentially with the state space. This significantly reduces the generation overhead of zero-knowledge proofs, solves the problem of circuit size explosion when large-scale unstructured data is uploaded to the blockchain, and achieves efficient extraction and verification of enterprise attribute privacy.

[0047] 2. This invention reduces the dimensionality of the Poseidon hash input from the plaintext length to one-inverse of the preset step size by aggregating the expanded vectors of field elements into a compact vector of field elements with a preset step size. Specifically, with a step size of 31 bytes, every 31 bytes is merged into a compact field element through polynomial reconstruction constraints, ensuring that the value of the compact field element does not exceed the order of the BN254 elliptic curve scalar field, thereby guaranteeing that the subsequent Poseidon hash input is within a safe range. This aggregation operation significantly reduces the number of hash constraints while ensuring data integrity, reducing the hash input size from N to N / 31, further reducing the total number of circuit constraints, improving proof generation efficiency, and breaking through the performance bottleneck of large-scale unstructured data on-chain.

[0048] 3. This invention achieves accurate extraction of variable-length fields through the coordinated use of one-hot encoding constraints, pattern matching constraints, and inner product extraction constraints. The global uniqueness and Boolean constraints in the one-hot encoding constraints strictly shrink the value space of the indicator vector to a standard one-hot vector; the pattern matching constraints utilize the selective activation property of the multiplication gate to precisely concentrate the verification effect of the constraints on the selected starting position; the inner product extraction constraints utilize the selector property of the one-hot vector to elevate the entire byte sequence immediately following the anchor point into an intermediate variable of the circuit; the Boolean mask vector truncates the target byte sequence of the actual effective length within the static maximum boundary, filtering out trailing characters after the field value. These mechanisms ensure the accuracy and completeness of the extraction results while maintaining the static nature of the circuit structure, i.e., the number of constraints is fixed at compile time, maintaining the flexible extraction capability for variable-length fields.

[0049] 4. This invention enforces a polynomial root-finding structure with terminal symbol boundary constraints, ensuring that the last byte of an attribute value belongs to the JSON syntax terminal symbol set. Specifically, assuming the JSON syntax terminal symbol set includes the ASCII codes for commas, double quotes, right curly braces, and newlines, the circuit applies a constraint that the sum of the product of a Boolean indicator vector and a terminal symbol polynomial equals 0. The terminal symbol polynomial is the product of the differences between the last byte of the attribute value and the codes of each terminal symbol. When the last byte belongs to the terminal symbol set, at least one term in the product is zero, making the entire product zero and automatically satisfying the constraint. When the last byte does not belong to the terminal symbol set, the product is not zero, the constraint is not valid, and generation fails. This constraint, algebraically, forces attribute values ​​to have correct boundaries in the JSON structure, achieving zero-knowledge protection against numerical truncation and tampering with constant-level constraint overhead. This fundamentally blocks truncation attack paths and enhances the security of the solution.

[0050] 5. This invention achieves fine-grained privacy protection by dividing the target attributes into a public attribute set and a private attribute set, and employing different verification strategies for each. After extracting and reconstructing attributes from the public attribute set, a Poseidon hash is calculated to obtain the public attribute hash commitment, which is then written into the on-chain identity credential. The verifier obtains the plaintext of the public attribute off-chain, calculates the hash locally, and compares it with the on-chain hash commitment to complete the authenticity verification. For numerical attributes in the private attribute set, a range proof constraint is used to prove that the attribute value is not less than the threshold value declared by the verifier without revealing the specific value. For administrative credit attributes in the private attribute set, a Boolean satisfaction constraint is used to prove that the attribute value is equal to the compliance code without revealing the specific content. This mechanism protects sensitive enterprise data while meeting the business needs of the verifier, resolving the contradiction between verifying identity and protecting privacy in enterprise collaboration.

[0051] 6. This invention achieves automated maintenance of credential status through an on-chain identity credential lifecycle management mechanism. Specifically, zero-knowledge proofs, digital fingerprints, digital signatures, and enterprise collateral are submitted to the on-chain identity registration center for registration. A credential validity period and minimum collateral threshold are pre-set. Enterprises must re-execute the extraction and proof process within the credential validity period to refresh the credential. If an enterprise's qualifications deteriorate to the point that it cannot regenerate a valid proof, the old credential is immediately deemed expired upon verification by the verifier after its validity period expires, achieving implicit revocation. This mechanism, leveraging the cryptographic fact that authoritative certification cannot be re-obtained, completes the natural elimination of ineffective enterprises with a bounded time delay. It achieves automatic elimination of ineffective credentials without requiring active transaction cancellation, solving the synchronization problem between on-chain credential status and dynamic changes in real-world enterprise qualifications, and reducing on-chain management costs. Attached Figure Description

[0052] The accompanying drawings, which are included to provide a further understanding of embodiments of the invention and form part of this application, do not constitute a limitation thereof. In the drawings:

[0053] Figure 1 This is a flowchart from Embodiment 1 of the present invention;

[0054] Figure 2 This is the logic diagram in Embodiment 1 of the present invention;

[0055] Figure 3 This is a schematic diagram showing the results of circuit constraint scale and generation efficiency based on R1CS in Embodiment 1 of the present invention;

[0056] Figure 4 This is a schematic diagram illustrating the results of proving the efficiency of volume and verification overhead in Embodiment 1 of the present invention;

[0057] Figure 5 This is a system block diagram in Embodiment 1 of the present invention. Detailed Implementation

[0058] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the embodiments and accompanying drawings. The illustrative embodiments and descriptions of the present invention are only used to explain the present invention and are not intended to limit the present invention.

[0059] Example 1: A zero-knowledge proof method for enterprise attributes based on dynamic positioning anchor points, such as... Figure 1 and Figure 2 As shown, it includes the following steps:

[0060] S1: Obtain the authoritative plaintext response data of the enterprise attributes and its digital signature;

[0061] S2: Perform arithmetic preprocessing on the authoritative response plaintext data to generate arithmetic vectors;

[0062] S3: Calculate the hash of the arithmetic vector and perform consistency constraints with the digital fingerprint signed by a notary node;

[0063] S4: Extract target attribute values ​​from arithmetic vectors based on dynamic anchor points. The extraction includes: presetting anchor strings for each target attribute, introducing Boolean indicator vectors and applying one-hot encoding constraints to lock the starting index of the anchor string in the plaintext, verifying the byte sequence at the starting index matches the anchor string byte by byte through pattern matching constraints, and extracting the byte sequence immediately following the anchor string by using the selection characteristics of Boolean indicator vectors through inner product extraction constraints, and combining Boolean mask vectors to extract the target byte sequence of the actual effective length and reconstruct it into the target attribute value.

[0064] S5: Apply validity checks to the reconstructed target attribute values, including at least ASCII range delimitation constraints, terminal symbol boundary constraints, range proof constraints, and Boolean satisfaction constraints.

[0065] S6: Construct a rank-1 constraint system circuit based on the verified constraint relationships and generate a zero-knowledge proof.

[0066] In step S1, the proving company ε wants to prove to the verifying party a specific attribute in its business registration information, such as the registered capital amount not being lower than a certain threshold. To this end, the proving company ε first needs to obtain the original response data containing this attribute from the authoritative data source S and ensure the authenticity and integrity of this data. This embodiment adopts the MPC-TLS collaborative acquisition mechanism, where the proving company ε and the collaborative notary node N jointly complete the TLS handshake and data decryption of the authoritative data source S, thereby outputting a digital fingerprint signed by the collaborative notary node N without exposing the complete plaintext to either party. The specific process is as follows.

[0067] First, the proving entity ε establishes a connection with the collaborative notary node N and jointly initiates an HTTPS request to the authoritative data source S. During the TLS handshake phase, both parties need to share an additive secret of the ECDHE temporary private key. Additive secret sharing is a cryptographic technique that randomly splits the secret value into multiple shares. Each share does not reveal any information about the secret individually; only by adding all the shares can the original secret be recovered. Specifically, the proving entity ε randomly selects a share of the temporary private key. The collaborative notary node N randomly selects another temporary private key share. Both are scalars on the secp256r1 elliptic curve. According to the principle of additive secret sharing, the complete temporary private key... satisfy ,in, Let q be the order of the elliptic curve group. The modulo operation mod q here strictly restricts the computation result to the elliptic curve scalar field [0, q-1], ensuring that the private key always remains within a secure finite field. This is the core mechanism in elliptic curve cryptography to prevent numerical overflow and guarantee the difficulty of the discrete logarithm problem. The share held by either party is uniformly distributed within the field, guaranteeing the privacy of the private key from an information theory perspective. Both parties exchange their respective temporary public key shares through a secure channel. and ,in This is the base point of the elliptic curve. Subsequently, both parties each compute the aggregate temporary public key. Due to the additive homomorphic property, the aggregated temporary public key Q is algebraically equivalent to a standard secp256r1 elliptic curve point, which can be verified by the TLS certificate of the authoritative data source S. Based on this, the authoritative data source S completes ECDHE key negotiation and generates a session key. This allows for completely non-intrusive access to Web2 authoritative data sources without the parties being aware that the private key is shared. The authoritative data source S requires no modification and can simply provide standard HTTPS services.

[0068] Next, the authoritative data source S returns an HTTPS response. The HTTP body of this response is plaintext data D in JSON format containing the company's business information, with a length of N bytes. This response is transmitted encrypted using AES-GCM via the TLS record layer, with the encryption key being the session key. The proving entity ε and the collaborating notary node N need to collaboratively decrypt the ciphertext. This is because both parties each hold a share of the ECDHE private key. and This requires jointly computing the symmetric key needed for AES-GCM decryption using a secure multi-party computation protocol. Specifically, both parties first utilize the additive secret sharing property to share the session key. Split into two shares and The plaintext data D is held by the proving enterprise ε and the collaborating notary node N, respectively. Then, using Yao's obfuscation circuit and Beaver triplet technology, both parties jointly perform AES-GCM decryption to obtain the individual bytes of the plaintext data D without revealing their respective private key shares. Yao's obfuscation circuit is a cryptographic protocol for secure two-party computation, transforming each logic gate in a Boolean circuit into an encrypted obfuscation table. The evaluator obtains the input tags through an inadvertent transmission protocol, thus calculating the circuit output without exposing the input. Beaver triplet is a pre-computation technique used to efficiently handle multiplication operations over finite fields. In the online phase, only one lightweight communication interaction is needed to transform the multiplication into a local linear reconstruction. During decryption, Yao's obfuscation circuit handles nonlinear operations such as AES S-box substitution, while Beaver triplet efficiently handles multiplication operations over finite fields. The entire decryption process ensures that neither party can obtain the complete plaintext D independently, thus protecting the privacy of sensitive enterprise information.

[0069] After obtaining the plaintext data D, the proving entity ε and the collaborating notary node N need to jointly calculate the digital fingerprint of the plaintext for subsequent integrity constraints in zero-knowledge proofs. The digital fingerprint is calculated using the Poseidon hash function. Both parties perform a Poseidon hash operation on the plaintext data D off-chain to obtain the digital fingerprint Cm_D. This digital fingerprint Cm_D uniquely identifies the original authoritative response plaintext data D and is collision-resistant. After verifying the correctness of the decryption process locally, the collaborating notary node N uses its own private key. Perform ECDSA signing on Cm_D to generate a signature. Specifically, notary node N calculates the hash digest of the message to be signed. ,in To prove the long-term identity public key of enterprise ε, The hash value is used to bind the digital fingerprint to the identity of the certifier, preventing the credential from being stolen. Then, the notary node N selects a random number. ,in, Let represent the multiplicative group of non-zero integers modulo q, i.e., the set of all integers between 1 and q-1. This restriction ensures that the random number k is non-zero and has a multiplicative inverse, which is a necessary prerequisite for the security of the ECDSA signature algorithm. Next, the notary node N calculates the elliptic curve points. ,in The default base point for the secp256k1 elliptic curve is a fixed reference point used to generate public keys and signature points. ,in Let R be the x-coordinate of the point on the elliptic curve. Modulo q operations are performed to strictly constrain it within a valid scalar domain, thereby extracting a fixed-length signature component r. Then, the signature is calculated. Where s is the signature proof value in the ECDSA signature algorithm, which is obtained by combining the message hash digest c, the x-coordinate of the elliptic curve point r, and the private key of the notary node. The signature is calculated by performing a linear combination and multiplying by the modular inverse of a random number k. This not only proves that the notary node N holds a legitimate private key, but also ensures the uniqueness of each signature result by introducing a nonlinear transformation of the random number k, effectively preventing replay attacks and data tampering. The final output signature is... ,in To recover the ID, a point R on the elliptic curve is used to uniquely determine the signature during verification. The proving enterprise ε ultimately obtains the authoritative response plaintext data D, the digital fingerprint Cm_D, and the signature of the collaborative notary node N. Throughout the process, the authoritative data source S requires no modification and only needs to provide standard HTTPS services, thus achieving completely non-intrusive access to the Web2 authoritative data source.

[0070] In step S2, within the zero-knowledge proof circuit, the authoritative response plaintext data D obtained in step S1 needs to be converted into an arithmetic representation that the circuit can perform algebraic operations on, and then consistently bound to the publicly input digital fingerprint Cm_D. This is because the Groth16 zero-knowledge proof system used in this invention is built on the prime number domain. The authoritative response plaintext data D is a binary byte sequence, therefore a mapping relationship must be established between the text byte stream and the cryptographic arithmetic object. This step uses two sub-operations, byte-by-byte mapping and aggregation compression, to losslessly compress the massive JSON byte stream into a high-density array of field elements, and uses integrity hash constraints to ensure that the private data processed inside the circuit is consistent with the data source of the off-chain notarized endorsement.

[0071] First, the circuit maps the authoritative response plaintext data D byte by byte into an expanded vector of field elements. Where N is the total number of bytes in the plaintext data D. Each field element The value of the i-th byte corresponding to the plaintext data D ranges from 0 to 255, belonging to the prime number field. The mapping operation converts binary byte data into an algebraic representation over a finite field, allowing subsequent arithmetic constraints to be performed on the field. The expanded vector d of the field elements preserves the complete order and numerical information of the plaintext data, serving as the fundamental input for subsequent dynamic positioning anchor extraction and validity verification. Simultaneously, it acts as a private witness input circuit, never exposed externally.

[0072] Next, to address the state explosion problem caused by the massive byte stream directly participating in Poseidon hashing, while also ensuring the flexibility of subsequent sliding window matching, the circuit introduces algebraic aggregation constraints. With a preset step size K = 31 bytes, the expanded vector d of the field elements is dynamically reconstructed into a compact vector of field elements. The number of groups For the t-th group, its compact field elements The calculation formula is:

[0073] ;

[0074] Where t ranges from 1 to m, representing the current group number being processed; k is the byte offset within the group, from 0 to K−1; This represents the byte value at the corresponding position in the expanded vector of the field element; The weighting factor is used to encode the byte sequence into an integer in big-endian order. This aggregation operation combines the 31 bytes within each group into a compact field element. Its value does not exceed The order p is less than that of the BN254 elliptic curve scalar field, thus ensuring that the input to the subsequent Poseidon hash is within a safe range, avoiding the risk of overflow or hash collisions due to excessively large field elements. Through this polynomial constraint, the huge single-byte sequence is losslessly compressed into a high-density array of field elements with a length reduced to one-thirty-first of the original.

[0075] In step S3, after completing the arithmetic preprocessing in step S2, a compact field element vector has been generated inside the circuit. The goal of this step is to compute the Poseidon hash of the compact field element vector within the zero-knowledge proof circuit, and to perform a mandatory consistency check between the computation result and the digital fingerprint Cm_D endorsed by the collaborative notary node N in step S1, thereby establishing a complete trust transfer link from the off-chain authoritative data source to the on-chain circuit verification at the cryptographic level.

[0076] Specifically, the circuit imposes the following integrity constraint on the compact domain element vector E:

[0077] ;

[0078] In this constraint, This represents the hash value calculated by sequentially inputting all m elements of the compact field element vector E into the Poseidon hash function; Cm_D is the digital fingerprint in the public input, which has been used by the collaborative notary node N with its private key in step S1. ECDSA signing is performed, and the legality of the signature is verified by the identity registry R through the ecrecover pre-compiled contract during the on-chain registration phase. The combination of a minus sign and zero indicates that the circuit requires the two to be strictly equal; if they are not equal, it proves that the generation has failed.

[0079] This constraint has critical business security implications. Since Cm_D is published after the collaborative notary node N signs the Poseidon hash result of the authoritative response plaintext data D following the completion of the off-chain MPC-TLS protocol execution, and the compact field element vector E within the circuit is calculated by the private witness d through the aggregation constraint in step S2, therefore… The constraint is algebraically equivalent to the following fact: the private byte vector d sent to the circuit by the proving enterprise ε is strictly consistent with the original HTTPS response data D returned by the authoritative data source S at the hash level, and this consistency has been recognized on-chain by the signature of the notary node N.

[0080] By reducing the input size of the nonlinear hash from the original plaintext length N to the number of blocks This reduces the size to one-thirty-first of the original, significantly decreasing the number of expensive Poseidon hash constraints in the circuit. This is because without aggregation and compression, Poseidon hashing requires processing N field elements directly, with the number of constraints increasing linearly with N and having a large coefficient; however, after aggregation, only m compact field elements need to be processed, significantly reducing the proportion of hash constraints and thus overcoming the performance bottleneck of large-scale unstructured data on-chain. This constraint is related to the subsequent public attribute hash commitment constraint. Together, they form a complete trust transfer chain: This ensures that the underlying byte vector d originates from a genuine response from an authoritative data source. This further ensures that the public attribute plaintext extracted from d strictly corresponds to the on-chain commitment.

[0081] In step S4, after completing the arithmetic preprocessing in step S2 and the integrity constraints in step S3, the circuit internally possesses the expanded vectors of the domain elements. The consistency between the fingerprint and the publicly available digital fingerprint Cm_D has been verified. The goal of this step is to accurately extract the byte sequence of each target attribute from the expanded vector d of the domain elements using a dynamic anchoring mechanism, and reconstruct it into an intermediate variable that the circuit can process, providing input for subsequent validity verification.

[0082] The core idea of ​​the dynamic anchor extraction mechanism is to leverage the deterministic position of attribute keys (anchor strings) in the plaintext of a JSON message. This position is then locked using a Boolean indicator vector and one-hot encoding constraints. Pattern matching constraints are then used to verify the correctness of the anchor string. Finally, the attribute value byte sequence is extracted using inner product extraction constraints and a Boolean mask vector. This mechanism avoids the constraint explosion caused by nested structures in traditional general-purpose JSON parsers, keeping the extraction complexity within a linear range.

[0083] Specifically, for each target attribute, the circuit pre-sets an anchor string. This string represents the key name of the attribute in the JSON message. For example, for the registered capital attribute, the anchor string would be `regCapital:`. The length of the anchor string is denoted as... The circuit introduces a Boolean indicator vector. Each of its components The value can be 0 or 1, used to indicate the anchor string. The starting index in the plaintext. The anchor string is used if and only if it begins at the i-th byte of the plaintext. All other components are 0.

[0084] To ensure Boolean indicator vector To satisfy the above semantics, the circuit applies one-hot encoding constraints. These constraints consist of two parts: global uniqueness constraints and Boolean property constraints. The global uniqueness constraint ensures that exactly one component in the entire Boolean indicator vector is 1, i.e., it satisfies... This constraint ensures that the anchor string has exactly one starting position in the plaintext. The Boolean constraint forces each component to take only 0 or 1, i.e., satisfies... This constraint ensures that each component is a valid binary value. Through the one-hot encoding constraint, the Boolean indicator vector... The starting index of the anchor string is uniquely determined.

[0085] After locking the starting index, the circuit needs to verify that the byte sequence at that index does indeed match the preset anchor string. Byte-by-byte matching. The circuit applies a pattern matching constraint, defined as: for any candidate start position i (value range 0 to...) ) and the offset k within the anchor point (range 0 to Apply the following equality constraints:

[0086] ;

[0087] in, Let i be the i-th component of the Boolean indicator vector; Expand the vector of the field elements. A value of one byte; Anchor string The ASCII code value of the k-th byte. The mathematical meaning of this constraint is: when... At that time, it must meet the following conditions. That is, the continuous sequence starting at index i Each byte is equal to the anchor string byte by byte; when When the product is zero, the condition is automatically satisfied, without generating any constraints. Through this pattern matching constraint, the circuit verifies the existence and accuracy of the anchor string at the arithmetic level.

[0088] After confirming the position of the anchor string, the circuit needs to extract the byte sequence of attribute values ​​immediately following the anchor string. The circuit applies an inner product extraction constraint, which utilizes the one-hot property of Boolean indicator vectors to extract the byte values ​​at a specified offset after the anchor string through a weighted summation. Specifically, a maximum byte length is preset for the target attribute. This length should be sufficient to accommodate the actual value of the attribute. For each byte, offset k (ranging from 0 to...) Circuit calculations:

[0089] ;

[0090] in, This refers to the extracted k-th byte value; Let i be the i-th component of the Boolean indicator vector; This refers to the k-th byte following the end of the anchor string in the expanded vector of the domain elements. Because... It satisfies the one-hot encoding constraint, i.e., it has only one component. Since all others are 0, the summation result directly collapses to the value of the k-th byte after the anchor string, without needing to traverse all possible starting positions. This inner product extraction constraint extracts the attribute value byte sequence losslessly from the expanded vector of domain elements, forming a length of... byte sequence .

[0091] However, the actual length of the attribute value Typically smaller than the preset maximum byte length byte sequence The latter part may contain trailing characters that do not belong to that attribute. In order to... To extract the target byte sequence of real effective length within the static boundary, the circuit introduces a Boolean mask vector. , before One bit is 1, and the rest are 0. Boolean mask vector Its value is determined by additional constraints: circuit requirements The former The first component is 1, the second... The first component is 0, and the second component is 0. The plaintext position corresponding to each byte belongs to the JSON syntax terminator set, such as commas, double quotes, and closing curly braces. This is achieved through a boolean mask vector. With byte sequence The element-wise product, the circuit only retains the first element. After filtering out trailing characters from each valid byte, the true byte sequence of the target attribute value is finally reconstructed. .

[0092] Following the dynamic positioning anchor point extraction process described above, the circuit outputs the byte sequence for each target attribute. The sequence will be subject to validity checks in subsequent steps, including ASCII range delimitation constraints, terminal symbol boundary constraints, range proof constraints, and Boolean satisfaction constraints, to ensure that the extracted attribute values ​​are syntactically and semantically consistent with expectations.

[0093] In step S5, after completing the dynamic positioning anchor point extraction in step S4, the circuit has obtained the byte sequence for each target attribute. The goal of this step is to impose a series of validity constraints on these extracted attribute values ​​to ensure that they conform to expectations both syntactically and semantically, thereby preventing the prover from deceiving the verifier by forging or tampering with attribute values. Validity constraints include ASCII range delimitation constraints, terminal symbol boundary constraints, range proof constraints, and Boolean satisfaction constraints. Furthermore, this step distinguishes between a public attribute set and a private attribute set, employing different validation and verification strategies for each.

[0094] First, the circuit applies an ASCII range constraint to the extracted attribute values. This constraint, applied to numeric attributes, checks byte-by-byte whether the ASCII code falls within the numeric character range. Specifically, for attribute values... Each byte in The circuit requires its values ​​to meet the following conditions. The constraint is implemented by using a double bit decomposition method, where 48 corresponds to the ASCII code for the character 0 and 57 corresponds to the ASCII code for the character 9. Each byte's value is decomposed into bits and then compared with both the lower and upper bounds (48 and 57) to ensure that all bytes are decimal numeric characters, thus preventing the inclusion of illegal characters.

[0095] Next, the circuit applies terminal symbol boundary constraints to the extracted attribute values ​​to prevent truncation attacks. A truncation attack refers to the prover truncating a long attribute value (such as registered capital of 10 million yuan) into a shorter value (such as 1 million yuan), thereby bypassing the lower bound check of the range proof. Terminal symbol boundary constraints are implemented using a polynomial root-finding structure. Let the set of JSON syntax terminal symbols be... This includes the ASCII codes for commas, double quotes, closing curly braces, and newlines, where the ASCII code for a comma is 44, for a double quote it is 34, for a closing curly brace it is 125, and for a newline it is 10. The circuit is based on the Boolean indicator vector in step S4. The locked starting index is subject to the following constraints:

[0096] ;

[0097] in, The i-th component of the Boolean indicator vector is used to lock the starting position of the anchor string; This is the value of the byte immediately following the end of the attribute value in the expanded vector of the domain element. The length of the anchor string. The actual valid length of the attribute value; The ASCII encoding of the m-th terminator in the JSON syntax terminator set; Let `next` be a terminal polynomial, representing the product of the differences between the last byte of the attribute value and the codes of each terminal symbol. When the last byte belongs to the terminal symbol set, at least one term in the product is zero, making the entire product zero and automatically satisfying the constraint. When the last byte does not belong to the terminal symbol set, the product is not zero, the constraint is not valid, and generation fails. This constraint enforces the correct boundaries of the attribute value in the JSON structure at an algebraic level, effectively preventing truncation attacks.

[0098] After completing the syntactic validation, the circuit further imposes semantic constraints on the attribute values. This invention categorizes the target attributes into a public attribute set. With privacy attribute set The two do not intersect. Public attributes refer to attributes that a company is willing to disclose publicly without hiding specific values, such as the company's unified social credit code and the name of its legal representative; privacy attributes refer to attributes that a company wants to prove it meets certain conditions without disclosing specific values, such as registered capital and tax credit rating.

[0099] For public attribute set The circuit extracts and reconstructs the attributes, calculates their Poseidon hashes, and obtains the public attribute hash commitments. This hash commitment is then written to the on-chain identity credential. Specifically, for each attribute value in the public attribute set... Circuit calculation And combine the hash commitments of all public attributes into Where P is the number of public attributes. This hash commitment is written into the public field of the enterprise identity credential during the on-chain registration phase. When the verifier needs to verify a public attribute, it obtains the plaintext of the attribute through off-chain channels, calculates the Poseidon hash locally, and compares it with the hash commitment stored on-chain. A comparison is performed. If the two match, it proves that the plaintext is consistent with the attribute value at the time of registration, thus completing the authenticity verification. This method ensures the verifiability of public attributes while avoiding the storage of large amounts of plaintext data on the blockchain.

[0100] For privacy attribute set For attributes within the privacy attribute set, the circuit employs range proof constraints and Boolean satisfaction constraints to prove that the attribute value satisfies the conditions declared by the verifier without revealing the specific numerical value. For numerical attributes in the privacy attribute set, such as registered capital, the circuit uses range proof constraints to prove that the attribute value is not less than the threshold value declared by the verifier. However, the specific numerical value of the attribute value is not exposed. Range proof constraints are implemented in zero-knowledge proof circuits through bit decomposition and comparison circuits: first, the attribute value... Decompose it into bit representations, then construct a comparison circuit to output a Boolean value. Whether the condition is met or not, the Boolean value is finally constrained to 1. The entire process is completed inside the circuit, and the verifier can only know that the attribute value meets the threshold condition, but cannot know the specific value of the attribute value.

[0101] For administrative credit attributes in the privacy attribute set, such as a company's tax credit rating, the circuit proves that the attribute value equals the compliance code through Boolean satisfaction constraints. This does not expose the specific content of the attribute value. Boolean satisfaction constraints are implemented through byte-by-byte comparisons: for attribute values... Each byte and compliance encoding The corresponding bytes are XORed, and the sum of all XOR results is constrained to zero, thus proving that the two are completely equal. This constraint is also performed internally within the circuit. The verifier can only know that the attribute value is equal to the compliance code, but cannot know the specific content of the compliance code, thereby protecting the company's privacy information.

[0102] Through the above validity verification constraints, the circuit ensures that the extracted target attribute values ​​conform to the JSON specification in terms of syntax and meet the conditions set by the verifier in terms of semantics. At the same time, it distinguishes between different processing strategies for public and private attributes, providing complete security for zero-knowledge proofs of enterprise attributes.

[0103] In step S6, after completing all the arithmetic preprocessing and validity verification constraints in steps S2 to S5, a complete set of constraint relationships has been constructed within the circuit, covering data integrity constraints, attribute location and extraction constraints, and attribute validity and satisfaction verification constraints. The goal of this step is to compile these constraint relationships into a rank-1 constraint system circuit and call the Groth16 zero-knowledge proof algorithm to generate the corresponding zero-knowledge proof, enabling the verifier to be certain that the proving entity ε does indeed hold a legitimate witness that satisfies all constraints without accessing private data.

[0104] Specifically, the zero-knowledge proof relation constructed in this invention It can be formally described as the conjunction of the following set of constraints:

[0105] ;

[0106] in, The representative public input includes the prover's identity public key hash. Trusted digital fingerprint Cm_D, public attribute hash commitment Identity extension anchor set A and corresponding type set and the set of threshold values ​​for each privacy attribute. ; Representing a private witness, it includes an arithmetic vector d of the complete plaintext, a high-density field element vector E, and a set of privacy attribute extracted values. The set of sliding window indicator vectors for each attribute With the actual effective length Boolean mask vector set Range proof auxiliary difference And low-level algebraic intermediate variables such as bit decomposition auxiliary vectors. The constraint set is divided into three main categories: data preprocessing and integrity assurance categories, including aggregation and compression constraints. Authoritative data source traceability constraints and public attribute commitment constraints Attribute localization and extraction classes include one-hot encoding constraints. Pattern matching constraints Inner product extraction constraints and polymorphic reconstruction constraints Attribute validity and satisfaction validation classes include ASCII range delimitation constraints. Terminal symbol boundary constraints Scope proof constraints Boolean satisfaction constraints .

[0107] During the circuit construction phase, all the above constraints are compiled into a rank-1 constraint system circuit. A rank-1 constraint system is a system that represents the computational process as a series of rules of the form... The arithmetic form of the quadratic constraints is given, where A, B, and C are linear combinations of field element vectors. Each constraint corresponds to a multiplication gate, while addition operations can be implemented free of charge through linear combinations. After compilation, the circuit calls the Groth16 proof generation algorithm, inputting the proof key pp and the public input... and private witness Output zero-knowledge proof ,in and Elliptic curve group The point on, Elliptic curve group The proof has a constant size of 128 bytes, regardless of the length of the input plaintext.

[0108] The total number of circuit constraints in this invention increases strictly linearly with the length N of the input plaintext. The fundamental reason is that the circuit degenerates field extraction from a traditional syntax parsing problem into a linear scan plus one-hot inner product selection, without any state machine iteration, backtracking, or state space expansion due to pattern complexity. Specifically, let the anchor length be L and the maximum byte length of a single field be... The number of extracted attributes is U, which is a design constant independent of N. Aggregate compression constraint. by Bytes are packaged in increments, with a constraint number of 100. The size is O(N); integrity hash The Poseidon input has been reduced to dimensionality. Its constraint number increases linearly with the input length, and remains at 1. In the attribute localization and extraction stage, one-hot constraints are used. It contains one summation constraint and N Boolean constraints, for Pattern matching constraints Traversal There are 1 candidate window, and each window compares L constant bytes. Inner product extraction constraints right Perform an inner product of length N on each of the output bytes, resulting in... The remaining polymorphic reconstructions, ASCII ranges, terminal symbols, and Boolean satisfaction constraints are all bounded by constant boundaries. Framing, for or at most Adding all the terms together, the constraint size of a single attribute is... , where the coefficient , Only L, Size of the set of terminal symbols Determined by constants; the total of U attributes is Therefore, when the number of attributes U is fixed, the total number of circuit constraints is a linear function of the plaintext length N, i.e. .

[0109] To verify the aforementioned linear growth characteristics, this embodiment underwent experimental evaluation. The experimental environment and tools are shown in Table 1. The host machine was equipped with an Intel i5-9300H processor and 16GB of memory, running the Ubuntu 20.04 LTS operating system. The zero-knowledge proof arithmetic circuit was written and compiled using the Circom 2.0 framework, and the proof generation and verification were performed using the SnarkJS library.

[0110] Table 1 Experimental Environment and Tools

[0111] CPU Intel® Core™ i5-9300H CPU @ 2.40 GHz, 4 cores and 8 threads RAM 16 GB operating system Ubuntu 20.04 LTS Operating environment and language Node.js 18.20, Rust 1.77.1, Solidity 0.8.19 Zero-knowledge arithmetic compiler Circom 2.0 Proof generation and verification library SnarkJS (a Groth16 library based on JavaScript / WASM) Smart Contract Testnet Ganache (Ethereum EVM local emulation environment)

[0112] like Figure 3 As shown, the experiment fixed the number of extracted attributes U=3, only varying the input plaintext length N. The total number of R1CS constraints showed a linear growth relationship with the plaintext size N, increasing from approximately 43,000 for 1KB to approximately 407,000 for 10KB, with the growth factor matching the plaintext length factor. This directly verifies the correctness of the sliding window dynamic anchor point design scheme of this invention. Even when processing 10KB unstructured JSON messages, the total number of constraints remained at a controllable level of around 400,000. For 10KB of qualification data, the peak time for generating Groth16 proofs in a typical office computer environment was approximately 20.42 seconds. Considering that identity authentication in enterprise collaboration is a low-frequency asynchronous initialization business, this latency fully meets the performance expectations of real-world scenarios.

[0113] like Figure 4 As shown, the zero-knowledge proof of this invention generates a proof size of 128 bytes regardless of changes in the input load, with each of the two G1 points being 32 bytes and the G2 point being 64 bytes. The on-chain verification time remains stable at approximately 8.5 milliseconds regardless of changes in the plaintext input load, with fluctuations not exceeding 1.3 milliseconds. This constant characteristic stems from the algebraic nature of Groth16 verification overhead, which depends only on the number of public inputs and is independent of the circuit constraint size, laying the foundation for subsequent low-cost, constant-gas on-chain verification.

[0114] Furthermore, the plaintext size N is fixed at 5KB while the number of extracted attributes U is varied. As shown in Table 2, each group adds a range proof and a Boolean proof sequentially, increasing from U=1 to U=5. The total number of R1CS constraints increases by approximately 43,000, an increase of about 23%. This indicates that the circuit size of this invention is mainly dominated by the plaintext length N, and the linear impact of the number of attributes U is much smaller than the impact of the plaintext length. The system can support parallel verification of multi-dimensional enterprise qualifications without significantly increasing computational overhead, demonstrating a certain degree of scalability.

[0115] Table 2. Impact of the number of extracted attributes on circuit size

[0116] U=1 0 180200 9.35 s U=3 2 204500 10.50 s U=5 4 223200 11.23 s

[0117] Compared to the traditional approach, which requires compiling the matching pattern into a finite automaton, the circuit must evaluate the transition function on the state set Q at each character position, with a constraint size of... However, the number of states |Q| in a finite automaton after determinism expands exponentially with the pattern complexity, reaching a maximum in the worst case. General JSON parsing requires the introduction of a pushdown stack to handle arbitrary nesting, further increasing overhead beyond linear limits. This invention does not parse JSON syntax; instead, it leverages the business characteristic of closely spaced key-value pairs in business registration documents to degenerate matching into a one-hot indicator vector of length N. The choice of inner product eliminates the |Q| term entirely in the constraint coefficients, thus preventing state explosion at its source. This is the fundamental reason why the constraints grow linearly rather than exponentially with N, and it is also the key reason why this invention has an order-of-magnitude efficiency advantage over general-purpose circuits on large-scale unstructured data.

[0118] After the above steps, the proving company ε generated a zero-knowledge proof locally. This proof, along with the publicly available input The digital fingerprint Cm_D and the signature of the collaborative notary node N. It will be submitted to the on-chain identity registration center. Registration and verification are performed. The entire proof generation process, conducted on a typical office computer environment, takes approximately 20.42 seconds to generate proofs for 10KB of plaintext enterprise attribute information, with a total constraint of approximately 407,000. Proof verification takes a stable 8.4 milliseconds, and the proof size remains constant at 128 bytes, fully meeting the balance between efficiency and privacy protection requirements in enterprise collaboration scenarios.

[0119] In some optional examples, after completing the zero-knowledge proof generation in step S6, the proving firm ε holds a valid proof containing the target identity attributes locally. Public input The digital fingerprint Cm_D and the signature of the collaborative notary node N. This invention submits these cryptographic materials, along with corporate collateral, to an on-chain identity registration center. Complete registration and achieve automated lifecycle management of enterprise identity credentials through credential validity period design and implicit revocation mechanism.

[0120] First, the proving party enterprise ε constructs a registration transaction. ,in The enterprise is provided with collateral and submitted to the on-chain identity registration center. Identity Registration Center The long-term public key of the collaborative notary node N has been pre-configured during the deployment phase. And a Groth16 verification key VK that matches the off-chain arithmetic circuit. Upon receiving the registration transaction, the identity registry... The atomic multidimensional verification is triggered, which includes the following three core gating logics.

[0121] The first layer of control consists of an authoritative root of trust and tamper-proof verification. Identity Registration Center Extract public input The digital fingerprint Cm_D and the public key hash of the current initiator's identity. Calculate the complete message digest to be verified. Subsequently, the contract extracts the endorsement signature of the notary node. It directly calls the Ethereum low-gas-consumption ecrecover pre-compiled interface, whose underlying elliptic curve algebra logic is: using the recovery identifier from The random point R of the signing process is reconstructed, and the issuer's public key is reconstructed on the secp256k1 curve. ,Right now The contract strictly asserts the reconstructed It must match the pre-defined public key in the whitelist of notary nodes. Consistency, that is The validity of this algebraic equation mathematically proves that the underlying data fed into the zero-knowledge proof circuit for computation absolutely originates from authentic and authoritative data sources and remains unaltered.

[0122] The second layer of control consists of identity binding and replay protection verification. Identity Registration Center From public input Extract the public key hash embedded during the TLS handshake phase. This derives its Ethereum address and verifies the consistency of the current transaction initiator address msg.sender. Meanwhile, the identity registration center Calculate the unique fingerprint for this proof. And check the list of used items. This is to prevent the same certificate from being registered repeatedly.

[0123] The third gate is the verification of the zero-knowledge proof algorithm. The contract supports zero-knowledge proofs. Perform Groth16 bilinear pairwise verification. The contract first computes the linear state combination points using the public input vector x{\text{pub}}. ,in, To verify the pre-computed points in the key VK, the contract then calls the underlying blockchain's BN254 curve pre-compiled contract to perform the core bilinear pairing equality verification: ,in These are all constants in the verification key VK. In engineering implementation, due to... and These are all constants in the verification key VK, and the contract can pre-calculate their values ​​in the extended domain during the initialization and deployment phase. Pairing results It is also cached, reducing the real-time verification overhead on the chain from four pairings to three pairings, which significantly reduces the gas cost for enterprise registration.

[0124] If and only if all the above multi-dimensional verifications pass, and the company's attached pledge deposit... Minimum staking threshold required to meet system requirements ,Right now Only then does the contract confirm that the proving entity ε did indeed honestly execute the attribute extraction circuit off-chain and update the global mapping. Store enterprise identity credentials as a seven-tuple state. This 7-tuple contains the following fields: the hash of the authoritative response plaintext Cm_D, and the public attribute hash digest. Privacy Attribute Threshold Statement Collection Validity timestamp of the certificate Voucher expiration timestamp Corporate pledge deposit and Boolean state bits At this point, the proving company ε has officially completed the registration and activation of its trusted on-chain identity.

[0125] Regarding the validity period of the voucher, this mechanism adopts a short-cycle dynamic refresh strategy. Effective timestamp. The `block.timestamp` of the Ethereum block at the time the registration transaction is uploaded to the blockchain is retrieved, which is objectively determined by the on-chain environment. The contract calculates the expiration timestamp based on the effective time and the system's preset valid window parameter `T`. Where T is a global system parameter, which is set to 90 days in this paper. If an enterprise wishes to maintain a valid on-chain identity, it must proactively re-execute the complete enterprise identity authentication process before the certificate expires, including all operations from steps S1 to S6, generate the corresponding zero-knowledge proof based on the latest enterprise information, and submit an update transaction to overwrite the old certificate status.

[0126] This invention employs an implicit revocation mechanism, enabling the automatic elimination of expired credentials without requiring active transaction cancellation. Specifically, when a company re-authenticates its credentials within their validity period, if its actual business status has changed—for example, insufficient registered capital, expiration of business term, or inclusion in the list of companies with abnormal operations—then the corresponding range proof constraint in the zero-knowledge proof circuit... or Boolean satisfaction constraint This will fail to meet the requirements, resulting in the inability to generate valid proof and the rejection of the registration transaction. At this time, the status of the existing on-chain credentials remains unchanged. Once the validity window T expires, the credentials will be immediately determined to be expired when queried by the verifier. It was determined to be expired during the query. This determination is made by the identity registry center. The calculation is performed in real time when responding to a verifier's query, without needing to write to the on-chain state, thus avoiding additional gas overhead. This invention leverages the cryptographic fact that authoritative certification cannot be regained, using a bounded time delay T to achieve the natural elimination of failing enterprises.

[0127] For the verification company The verification process begins with the verifier calling the identity registration center. read-only interface Obtain the seven-tuple state of the target company. The verifier sequentially performs status and deposit verification to determine... Is it active? And assess the pledged funds linked to the enterprise. Does it meet the requirements of subsequent business operations? Perform timeliness checks to determine whether the current block timestamp meets the requirements. The verifier proceeds to the attribute content comparison and verification phase only if the credential is valid and has not been cancelled. For public attribute sets... The proving entity (ε) presents the plaintext attributes to the verifying entity through an off-chain encrypted channel. The verifying entity calculates the Poseidon hash locally and compares it with the data stored on-chain. If the comparison is successful and the data matches, the authenticity of the public attribute is confirmed. For the set of privacy attributes... The verifier does not need to obtain any sensitive plaintext; they can directly query the on-chain data. Collection of privacy threshold statements By logically comparing the results with the company's own business risk control requirements, it can be confirmed that the company meets specific qualifications. The entire verification process requires only one on-chain read-only query and one local hash calculation, with an end-to-end latency of approximately 52.5 milliseconds in a wide area network environment, which fully meets the needs of enterprises for efficient verification in collaborative scenarios.

[0128] Example 2: A zero-knowledge proof system for enterprise attributes based on dynamic positioning anchors. This system is used to implement the zero-knowledge proof method for enterprise attributes based on dynamic positioning anchors described in Example 1, such as... Figure 5 As shown, it includes a data acquisition module, a data processing module, a hash verification module, an anchor point extraction module, a verification constraint module, and a proof generation module.

[0129] The system comprises the following modules: a data acquisition module for acquiring authoritative response plaintext data of enterprise attributes and its digital signature; a data processing module for performing arithmetic preprocessing on the authoritative response plaintext data to generate an arithmetic vector; a hash verification module for calculating the hash of the arithmetic vector and constraining consistency with the digital fingerprint signed by a notary node; an anchor extraction module for extracting target attribute values ​​from the arithmetic vector based on dynamically located anchor points, including: pre-setting anchor strings for each target attribute, introducing Boolean indicator vectors and applying one-hot encoding constraints to lock the starting index of the anchor string in the plaintext, verifying the byte sequence at the starting index matches the anchor string byte by byte through pattern matching constraints, and extracting the byte sequence immediately following the anchor string using the selection characteristics of the Boolean indicator vector through inner product extraction constraints, and combining the Boolean mask vector to truncate the target byte sequence of the actual effective length and reconstruct it into the target attribute value; a verification constraint module for applying validity verification constraints to the reconstructed target attribute value, including at least ASCII range delimitation constraints, terminal symbol boundary constraints, range proof constraints, and Boolean satisfaction constraints; and a proof generation module for constructing a rank-1 constraint system circuit based on the verified constraint relationships to generate zero-knowledge proofs.

[0130] Working Principle: This invention utilizes a dynamic anchor point positioning mechanism to transform the extraction of target attributes from authoritative response plaintext data in JSON format into a combination of one-hot encoding constraints, pattern matching constraints, and inner product extraction constraints. Leveraging the one-hot property of Boolean indicator vectors, the attribute positioning problem is transformed into linear scanning and inner product selection. Specifically, the one-hot encoding constraint forces that exactly one component in the entire Boolean indicator vector is 1, thus uniquely determining the starting index of the anchor string in the plaintext; the pattern matching constraint verifies that the byte sequence at this starting index matches the anchor string byte-by-byte, ensuring accurate positioning; the inner product extraction constraint utilizes the selection characteristics of Boolean indicator vectors, using weighted summation to elevate the byte sequence immediately following the anchor point into an intermediate variable of the circuit, and combining this with a Boolean mask vector to extract the target byte sequence of the truly effective length. This mechanism avoids the complex circuit structures required by traditional regular expression matching or finite state machine parsing, ensuring that the total number of circuit constraints is linearly related to the plaintext length, rather than expanding exponentially with the state space. This significantly reduces the generation overhead of zero-knowledge proofs, solves the problem of circuit size explosion when large-scale unstructured data is uploaded to the blockchain, and achieves efficient enterprise attribute privacy extraction and verification.

[0131] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0132] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0133] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0134] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0135] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A zero-knowledge proof method for enterprise attributes based on dynamic positioning anchor points, characterized in that, Includes the following steps: Obtain authoritative plaintext response data and its digital signature for enterprise attributes; The authoritative response plaintext data is preprocessed arithmetically to generate an arithmetic vector; Calculate the hash of the arithmetic vector and apply consistency constraints to the digital fingerprint signed by the notary node. Extracting target attribute values ​​from the arithmetic vector based on dynamically located anchor points, the extraction includes: presetting an anchor string for each target attribute, introducing a Boolean indicator vector and applying one-hot encoding constraints to lock the starting index of the anchor string in the plaintext, verifying the byte sequence at the starting index matches the anchor string byte by byte through pattern matching constraints, extracting the byte sequence immediately following the anchor string using the selection characteristics of the Boolean indicator vector through inner product extraction constraints, and combining the Boolean mask vector to truncate the target byte sequence of the actual effective length and reconstruct it into the target attribute value; Apply validity checks to the reconstructed target attribute values, including at least ASCII range delimitation constraints, terminal symbol boundary constraints, range proof constraints, and Boolean satisfaction constraints. Based on the verified constraint relationships, a rank-1 constraint system circuit is constructed, generating a zero-knowledge proof.

2. The zero-knowledge proof method for enterprise attributes based on dynamic positioning anchor points according to claim 1, characterized in that, The arithmetic preprocessing includes: The authoritative response plaintext data is mapped byte by byte to a domain element expanded vector, and the domain element expanded vector is aggregated into a compact domain element vector with a preset step size. Furthermore, the hash calculation for the arithmetic vector specifically involves: Calculate a hash for the compact field element vector, reducing the hash input size from the plaintext length to one-inverse of the preset step size.

3. The zero-knowledge proof method for enterprise attributes based on dynamic positioning anchor points according to claim 1, characterized in that, The one-hot encoding constraints include: Apply a global uniqueness constraint and a Boolean constraint to the Boolean indicator vector. The global uniqueness constraint forces that exactly one component in the Boolean indicator vector is 1, and the Boolean constraint forces that each component can only take the values ​​0 or 1. Furthermore, the pattern matching constraint specifically includes: For any candidate starting position and offset within the anchor point, apply a constraint that the product of the component of the Boolean indicator vector and the byte difference is equal to 0, such that when the component is equal to 1, the plaintext byte and the anchor byte are equal byte by byte.

4. The zero-knowledge proof method for enterprise attributes based on dynamic positioning anchor points according to claim 1, characterized in that, The inner product extraction constraint is specifically as follows: The target attribute is preset with a maximum byte length. A weighted summation constraint of the Boolean indicator vector and the plaintext byte is applied to each byte offset. The one-hot property of the Boolean indicator vector is used to collapse the summation result into the byte value of the corresponding offset after the anchor point. The first few bits of the Boolean mask vector are 1s and the rest are 0s, which are used to extract the target byte sequence of the actual effective length within the maximum byte length.

5. The zero-knowledge proof method for enterprise attributes based on dynamic positioning anchor points according to claim 1, characterized in that, The terminal symbol boundary constraints are implemented using a polynomial root-finding structure: Suppose that the set of JSON syntax terminators includes the ASCII codes of commas, double quotes, right curly braces, and newlines. Based on the locked start index in the Boolean indicator vector, apply a constraint that the product of the component corresponding to the start index and the terminator polynomial is equal to 0, wherein the terminator polynomial is the product of the difference between the byte immediately following the end of the attribute value and the codes of each terminator. The constraint is satisfied when the product of the consecutive byte belongs to the set of terminal symbols; otherwise, the constraint is not satisfied to prevent truncation attacks.

6. The zero-knowledge proof method for enterprise attributes based on dynamic positioning anchor points according to claim 1, characterized in that, The target attribute values ​​include a set of public attributes and a set of private attributes, which do not overlap; After extracting and reconstructing the attributes in the public attribute set, the hash is calculated to obtain the public attribute hash commitment, which is then written into the on-chain identity credential. The verifier obtains the plaintext of the public attributes off-chain, calculates the hash locally, and compares it with the on-chain hash commitment to complete the authenticity verification; For numerical attributes in the privacy attribute set, the attribute value is proven to be no less than the threshold value declared by the verifier without exposing the specific value through range proof constraints. For administrative credit attributes in the privacy attribute set, the attribute value is proven to be equal to the compliance code without exposing the specific content through Boolean satisfaction constraints.

7. The zero-knowledge proof method for enterprise attributes based on dynamic positioning anchor points according to claim 1, characterized in that, The method also includes: Submit the zero-knowledge proof, the digital fingerprint, the digital signature, and the enterprise pledge deposit to the on-chain identity registration center to complete the registration; The on-chain identity registration center pre-sets the validity period of the certificate and the minimum staking threshold. Enterprises need to re-execute the extraction and proof process within the validity period of the certificate to refresh the certificate. If the enterprise's qualifications deteriorate and it is unable to regenerate a valid certificate, the old certificate will be determined to be expired in real time when the verifier queries after the validity period of the certificate expires, thus realizing implicit revocation.

8. The zero-knowledge proof method for enterprise attributes based on dynamic positioning anchor points according to claim 1, characterized in that, The total number of constraints in the rank-1 constraint system circuit is linearly related to the plaintext length. The size of a single-attribute constraint is a constant multiplied by the plaintext length plus another constant. The constant is determined by the anchor length, the maximum byte length, and the size of the terminal symbol set, but does not include the number of finite automaton states.

9. The zero-knowledge proof method for enterprise attributes based on dynamic positioning anchor points according to claim 1, characterized in that, The authoritative response plaintext data is obtained through MPC-TLS collaboration: The proving enterprise and the collaborative notary node share the ECDHE temporary private key of the TLS handshake additively, so that the aggregated temporary public key of both parties is a standard elliptic curve point in algebraic terms to pass the verification of the authoritative data source. Both parties work together to complete AES-GCM decryption and hash calculation through Yao's obfuscation circuit and Beaver triple, and output a digital fingerprint with the signature of the collaborative notary node. The authoritative data source does not need to be modified.

10. A zero-knowledge proof system for enterprise attributes based on dynamic positioning anchor points, characterized in that: include: The data acquisition module is used to acquire authoritative response plaintext data of enterprise attributes and their digital signatures; The data processing module is used to perform arithmetic preprocessing on the authoritative response plaintext data to generate an arithmetic vector; The hash verification module is used to calculate the hash of the arithmetic vector and perform consistency constraints with the digital fingerprint signed by the notary node. An anchor point extraction module is used to extract target attribute values ​​from the arithmetic vector based on dynamically located anchor points. The extraction includes: presetting an anchor string for each target attribute, introducing a Boolean indicator vector and applying one-hot encoding constraints to lock the starting index of the anchor string in the plaintext, verifying the byte sequence at the starting index matches the anchor string byte by byte through pattern matching constraints, extracting the byte sequence immediately following the anchor string by using the selection characteristics of the Boolean indicator vector through inner product extraction constraints, and combining the Boolean mask vector to truncate the target byte sequence of the actual effective length and reconstruct it into the target attribute value. The validation constraint module is used to apply validity validation constraints to the reconstructed target attribute values, including at least ASCII range delimitation constraints, terminal symbol boundary constraints, range proof constraints, and Boolean satisfaction constraints. The proof generation module is used to construct a rank-1 constraint system circuit based on verified constraint relationships and generate zero-knowledge proofs.

Citation Information

Patent Citations

  • System and method for intermediating electronic commerce

    CN1386231A