Strain copyright protection system and method based on DNA watermark technology

By embedding encrypted copyright information in the strain genome and storing it on the chain, the problem that the existing technology cannot effectively protect the strain intellectual property rights is solved, and efficient and safe strain copyright protection and management are achieved.

CN120197159APending Publication Date: 2025-06-24JIYIN CHUANGWU (SHANGHAI) TECHNOLOGY CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510385194.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The prior art cannot effectively protect the intellectual property rights of bacterial strains, prevent misappropriation and illegal spread, and traditional biosecurity measures cannot achieve comprehensive protection.

Method used

A strain copyright protection system based on DNA watermarking technology is adopted to realize copyright protection and management by embedding encrypted copyright information in the strain genome and storing it on the chain.

Benefits of technology

It has achieved efficient, safe and reliable intellectual property protection for biological bacteria species, prevented misappropriation and illegal dissemination, and ensured the permanence and authenticity of copyright records.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120197159A_ABST
    Figure CN120197159A_ABST
Patent Text Reader

Abstract

The invention discloses a strain copyright protection system based on a DNA watermark technology. The system comprises a DNA watermark embedding module, a block chain storage and verification module, a strain authentication and tracking module and a security and privacy protection module. The invention further discloses a strain copyright protection method based on the DNA watermarking technology. The method comprises the steps of strain copyright information generation and coding, DNA watermark site selection and editing, strain screening and verification, copyright information uplink and strain use and tracking. According to the method, the copyright information of the strain is coded into the DNA watermark, the DNA watermark is embedded into the genome non-functional region of the strain, and the copyright information is stored and verified in combination with a block chain technology, so that intellectual property protection with high concealment and high safety of the biological strain is realized; illegal copying and tampering of strains are prevented, reliable copyright authentication and tracking functions are provided, and the method has wide application prospects and remarkable market value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of biological resource protection technology, and in particular to a bacterial strain copyright protection system and method based on DNA watermark technology, which is suitable for copyright protection, traceability and security management of biological strains. Background Art

[0002] In recent years, synthetic biology has been widely used in many fields such as industry, medicine and agriculture. In particular, the use of microbial strains has achieved remarkable results in biomanufacturing, drug production and environmental protection. However, with the widespread application of high-value strains, the theft and unauthorized spread of strains have become urgent problems to be solved. As the core production factor in synthetic biology, strains have become targets of theft and imitation due to their unique biological characteristics and commercial value. This not only damages the intellectual property rights of the original developers, but may also bring safety and ecological risks.

[0003] At present, the protection technology against the misappropriation of bacterial strains mainly focuses on physical isolation, legal measures and limited biosafety means. However, these measures have many shortcomings in actual operation. First, physical isolation and legal means cannot effectively prevent the illegal replication and distribution of bacterial strains, especially in the absence of effective tracking and identification methods. Secondly, traditional biosafety measures, such as biological isolation and environmental-dependent bacterial strain growth restrictions, are often only effective under specific conditions and cannot achieve comprehensive protection of bacterial strains. In addition, traditional protection methods rely on legal means, which have disadvantages such as high execution costs and difficulty in tracing.

[0004] As an emerging information hiding technology, DNA watermarking technology draws on the application ideas of traditional digital watermarking in the multimedia field. By embedding specific watermark information in the gene sequence, a unique identification method can be provided for the source and ownership of the gene product. DNA watermarking technology has become an emerging technology for protecting biological strains due to its high concealment and uniqueness. The decentralized and tamper-proof characteristics of blockchain technology also provide new solutions for the management of biological strains. However, the application of these technologies alone still cannot fully meet the needs of strain protection.

[0005] Therefore, there is an urgent need for an innovative solution that combines DNA watermarking technology, blockchain technology and strain protection to achieve efficient, safe and reliable protection and management of biological strains. Summary of the invention

[0006] To solve the technical problems existing in the prior art, the present invention provides a strain copyright protection system and method based on DNA watermarking technology. By combining the high concealment of DNA watermarking, the immutability of blockchain, and the professional requirements of strain protection, efficient, safe, and reliable intellectual property protection and management of biological strains are achieved.

[0007] The present invention provides a strain copyright protection system based on DNA watermarking technology, including: A DNA watermark embedding module, which is used to encode strain copyright information into a DNA watermark sequence and embed it into a specific region of the strain's genome. Specifically, it includes: A strain copyright information generation unit: used to generate a unique identifier for a specific strain and encrypt the unique identifier of the strain using an encryption algorithm.

[0008] A DNA sequence encoding unit: used to convert the encrypted unique identifier of the strain into a DNA sequence suitable for genome embedding to generate a unique DNA watermark sequence.

[0009] A DNA watermark embedding site selection and analysis unit: screening a specific region suitable for inserting the DNA watermark through a bacterial genome annotation database.

[0010] A DNA watermark embedding unit: used to embed the DNA watermark sequence into a specific region of the strain's genome.

[0011] A strain screening and verification unit: screening and functionally evaluating the strains into which the DNA watermark has been successfully inserted to ensure that the strain performance is not affected.

[0012] A blockchain storage and verification module, which is used to store the strain copyright information and its hash value on the blockchain and automatically execute copyright verification, authorization, and transactions through a smart contract. Specifically, it includes: A blockchain node unit: used to maintain the nodes of the blockchain network, including full nodes and light nodes, to ensure the decentralization and high availability of the network.

[0013] A smart contract execution unit: used to write, deploy, and manage smart contracts, and define the automated rules for strain copyright verification, authorization, and transactions.

[0014] A data storage and indexing unit: used to permanently store the copyright information and its hash value on the blockchain to ensure the immutability and permanent traceability of the data.

[0015] A transaction generation and submission unit: used to submit the constructed transaction to the blockchain network through a blockchain client or API to participate in the consensus mechanism verification.

[0016] Consensus Mechanism Management Unit: It is used to implement and manage the consensus algorithm adopted by the blockchain network to ensure the consistency and security of the network.

[0017] User Interface and API Unit: It is used to provide a user-friendly interface and API, allowing users to upload strain copyright information, submit transactions, query and verify copyright status, etc.

[0018] Monitoring and Log Management Unit: It is used to monitor the running status of the blockchain storage and verification module in real time, including node status, transaction processing, smart contract execution, etc.

[0019] Strain Authentication and Traceability Module: It is used to compare the DNA watermark extracted and decoded from the strain with the blockchain records to achieve strain copyright confirmation and traceability. Specifically, it includes: DNA Watermark Extraction Unit: It is used to extract the embedded DNA watermark sequence from a specific region of the genome of the strain to be verified.

[0020] Information Decoding Unit: It is used to restore the extracted DNA watermark sequence to the original strain copyright information according to the predetermined coding rules.

[0021] Strain Copyright Comparison and Verification Unit: It is used to compare the decoded copyright information with the copyright information stored on the blockchain to verify the copyright ownership and integrity.

[0022] Tracking and Traceability Unit: It is used to record the transfer of strain ownership, authorized use, and transaction history.

[0023] Database Access Unit: It is used to retrieve relevant copyright information and transaction records from the blockchain storage and verification module.

[0024] Reporting and Visualization Unit: It is used to generate detailed copyright verification reports and traceability analysis reports based on the authentication and traceability results.

[0025] User Notification Unit: It is used to send notifications and alerts to users in a timely manner when strain copyright verification results and traceability events occur.

[0026] Security and Privacy Protection Module: It is used to implement encryption, protection, auditing, and compliance management for various types of data and operations in the system. Specifically, it includes: Data Encryption Unit: It uses asymmetric encryption algorithms to encrypt key data to prevent leakage and tampering.

[0027] Privacy Protection Unit: It protects user privacy through anonymization and zero-knowledge proof technologies.

[0028] Access Control Unit: It implements multi-factor authentication (MFA) and role-based access control (RBAC).

[0029] Security Audit and Log Management Unit: Real-time record system operation logs and use automated analysis tools to detect anomalies.

[0030] Key Management Unit: Securely generate, distribute, and store encryption keys, and rotate them regularly.

[0031] Security Policy and Compliance Management Unit: Ensure that the system complies with laws and regulations such as GDPR and ISO 27001.

[0032] The present invention also provides a method for protecting the copyright of bacterial strains based on DNA watermarking technology, including the following steps: Step 1, Generation and Encoding of Bacterial Strain Copyright Information: Convert the bacterial strain copyright information into a DNA watermark sequence and add redundancy and verification. Step 2, Selection and Editing of DNA Watermark Sites: Insert the DNA watermark sequence into the non-functional region of the bacterial strain genome. Step 3, Screening and Verification of Strains: Verify the correctness of the watermark sequence and the functionality of the strain through selective markers and sequencing. Step 4, Uploading of Bacterial Strain Copyright Information to the Chain: Package the bacterial strain copyright information and its hash value into a transaction and submit it to the blockchain network, and generate a copyright ID. Step 5, Use and Tracking of Bacterial Strains: When the strain is actually used, confirm the copyright ownership by extracting the DNA watermark and comparing it with the blockchain, and record the usage path.

[0033] Step 6, Report Generation and Result Display: Provide users with visual information on the copyright status, usage records, and traceability paths of bacterial strains.

[0034] In the above steps, data and operations are encrypted, access-controlled, audited, and abnormal responses are performed through security and privacy protection methods.

[0035] In a preferred embodiment of the method for protecting the copyright of bacterial strains based on DNA watermarking technology provided by the present invention, the CRISPR-Cas9 technology is used for genome editing in Step 2 to ensure high precision and low off-target rate of watermark insertion.

[0036] In a preferred embodiment of the method for protecting the copyright of bacterial strains based on DNA watermarking technology provided by the present invention, when uploading copyright information to the chain in Step 4, multi-chain collaboration or cross-chain technology is supported to achieve compatibility and expansion of different blockchain networks.

[0037] In a preferred embodiment of the method for protecting the copyright of bacterial strains based on DNA watermarking technology provided by the present invention, the visual report generated in Step 6 includes the usage location, usage time, user information, and usage authorization scope of the strain, and supports export to a third-party system or sharing with partners.

[0038] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0039] The present invention deeply embeds copyright information into the genome of the strain, and uses non-functional regions to ensure the concealment of the watermark without affecting the normal growth and metabolic functions of the strain. Compared with traditional digital watermarks, DNA watermarks are more difficult to detect, remove or tamper with, providing strong protection at the physical level. In addition, error checking and redundancy design ensure the high stability of the watermark in the genome, and even in the case of gene mutation or recombination, the copyright information can still be accurately extracted and verified.

[0040] The present invention uploads the copyright information to the blockchain, and uses the decentralized and immutable characteristics of blockchain technology to ensure the permanence and authenticity of the copyright record. Any change to the copyright information needs to be verified by the consensus mechanism of the blockchain network, greatly improving the credibility of the copyright data. At the same time, the application of smart contracts realizes the automated management of copyright verification and authorization, reduces the risks and errors of manual operations, and improves the overall credibility and transparency of the system.

[0041] The present invention realizes the automated management of copyright verification and authorization through smart contracts, reduces manual intervention, and improves management efficiency and accuracy. The design of the user interface is simple and intuitive, reducing the operation complexity and improving the user experience. The reporting and visualization functions enable users to intuitively understand the copyright status and the dissemination path of the strain, supporting quick decision-making and management operations. This intelligent and automated management method significantly improves the operation efficiency and management level of the system, meeting the needs of modern biotechnology enterprises for efficient management.

[0042] The present invention comprehensively guarantees the confidentiality, integrity and availability of data through multi-level security measures. Data encryption ensures the high security of sensitive information during transmission and storage, preventing data leakage and unauthorized access. Access control and multi-factor authentication mechanisms strictly limit users' access rights to system resources, avoiding unauthorized operations. Security auditing and log management provide detailed operation records for post-event review and problem tracking. The privacy protection unit ensures that users' privacy is not leaked when using the system through anonymization and zero-knowledge proof technologies, improving the compliance and user trust of the system.

[0043] The present invention realizes the automated management of copyright verification and authorization through smart contracts and automated processes, significantly reducing the need for manual operations and lowering management costs. At the same time, high-throughput sequencing technology and automated screening methods are adopted to improve the efficiency of DNA watermark embedding and strain verification, reducing experimental costs. The efficient operation and cost optimization of the overall system not only improve the economic benefits of the system, but also promote the sustainable development of biological resource management, meeting the needs of enterprises for efficient and low-cost solutions. Brief Description of the Drawings

[0044] Figure 1 It is the architecture diagram of the strain copyright protection system based on DNA watermark technology provided by the present invention.

[0045] Figure 2 It is the structural diagram of the DNA watermark embedding module of the strain copyright protection system based on DNA watermark technology provided by the present invention.

[0046] Figure 3 It is the structural diagram of the blockchain storage and verification module of the strain copyright protection system based on DNA watermark technology provided by the present invention.

[0047] Figure 4 It is the structural diagram of the strain authentication and tracking module of the strain copyright protection system based on DNA watermark technology provided by the present invention.

[0048] Figure 5 It is the structural diagram of the security and privacy protection module of the strain copyright protection system based on DNA watermark technology provided by the present invention.

[0049] Figure 6 It is the flowchart of the strain copyright protection method based on DNA watermark technology provided by the present invention. Detailed Embodiments

[0050] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0051] Refer to Figures 1-5 , which are respectively the architecture diagram of the strain copyright protection system based on DNA watermark technology, the structural diagram of the DNA watermark embedding module, the structural diagram of the blockchain storage and verification module, the structural diagram of the strain authentication and tracking module, and the structural diagram of the security and privacy protection module provided by the present invention.

[0052] 1.1 Strain Copyright Information Generation Unit: Used to generate a unique identifier for a specific strain and encrypt the unique identifier of the strain using an encryption algorithm.

[0053] First: Assign a unique identifier (ID) to each strain protected by this technology. This ID can be generated in various ways, such as the patent number of the patented strain, a randomly generated globally unique identifier (GUID), or identification information generated based on the biological characteristics of the strain itself. This ID is the basis for strain identity authentication and is used to uniquely identify each specific strain.

[0054] For example, for a strain with a patent number of "CN20240101", by combining the embedded date "20241129" and the registration ID "12345", a fixed-length unique ID: "3a7d67f..." (the actual generated content depends on the algorithm implementation) can be generated through string concatenation and then using a hashing algorithm (such as SHA256). This identifier will serve as the "ID card" of the strain and will be recorded in the strain copyright database, which is a blockchain database in this case.

[0055] Second: To ensure that the identifier of the strain cannot be maliciously read or tampered with, this invention introduces an encryption mechanism. The specific steps are as follows: Use an asymmetric encryption algorithm (such as RSA) to encrypt the unique identifier of the strain. Asymmetric encryption algorithms are particularly advantageous in some high-security scenarios because public-key encryption and private-key decryption can be used, which ensures that only legitimate users who possess the decryption private key can interpret the ID.

[0056] Therefore, we will publicly disclose the public key, further transform the encrypted ciphertext into a DNA sequence, and prepare to embed it into the genome of the strain. For each user, we will prepare a corresponding private key, which can be used for decryption when the strain is found to be misappropriated. This private key will be kept by the protected person.

[0057] For example, if the original identifier is "3a7d67f...", after encryption with the public key, it may generate "eb67a93...". This ciphertext can only be decrypted by legitimate users who possess the private key, ensuring the security and uniqueness of the identifier.

[0058] 1.2 DNA sequence encoding unit: Used to transform the encrypted unique identifier of the strain into a DNA sequence suitable for genome embedding, generating a unique DNA watermark sequence.

[0059] Specifically: Transform the encrypted strain ID into a DNA sequence suitable for genome embedding. The specific encoding method can adopt the mapping relationship between binary and four nucleotide bases (A, T, C, G). We can use the following rules: A represents binary "00" T represents binary "01" C represents binary "10" G represents binary "11" According to this rule, the encrypted identifier is converted bit by bit into a DNA sequence to generate a unique DNA watermark sequence. The length of this sequence can be flexibly adjusted according to the strain ID and the complexity of the encrypted data, ensuring that it is short enough to avoid affecting the normal functions of the strain, has sufficient encryption strength, and has a low enough probability of randomly appearing in the genome to ensure sufficient persuasiveness.

[0060] For example, a mapping rule from binary to bases (A, T, C, G) is adopted: A corresponds to "00", T corresponds to "01", C corresponds to "10", and G corresponds to "11". Taking the binary encoding of the encrypted ciphertext "eb67a93..." as an example, after conversion, a DNA sequence is generated, such as "ACTGTGCC...". The length of this DNA sequence will be flexibly adjusted according to the complexity of the encrypted content, being short enough to avoid affecting the normal functions of the strain and having sufficient randomness and encryption strength.

[0061] 1.3 Embedding site selection and analysis unit: Screen non-functional regions suitable for inserting DNA watermarks through a bacterial genome annotation database.

[0062] Specifically: 1) Genome retrieval Log in to the NCBI official website and enter the Genome or Assembly database; Enter the scientific name of the target bacterial species (such as "Escherichia coli") in the search box, and screen specific strains as needed (such as "K-12 MG1655" or other sequenced strain versions); Select the gene assembly result with higher quality (high coverage, accuracy) in the RefSeq (reference sequence) or GenBank record, and click to enter this entry.

[0063] 2) Annotation file acquisition Under the corresponding genome entry, the complete genome sequence (in FASTA format) and annotation files (in GFF, GBK, etc. formats) can be downloaded; The annotation file usually contains information such as the start position (start), end position (end), function description, CDS (protein coding region), or rRNA / tRNA of the gene, providing basic data for subsequent screening of embedding sites.

[0064] 3) Annotation file parsing Use bioinformatics tools (such as BioPython, BEDTools, Prokka, or other custom scripts) to read the GFF / GBK file and extract information such as the coordinates, gene function annotations, and positive and negative strands of each gene (CDS or RNA gene); Storing the annotation information in a structured database or table (such as MySQL, CSV) facilitates subsequent screening and analysis.

[0065] 4) Intergenic region extraction Based on the start and end positions of all genes, calculate the intergenic regions between adjacent genes; Record the start and end coordinates, length, strand (positive or negative strand) of the intergenic region, as well as the distance information from the upstream and downstream genes; Eliminate overly short or overlapping intergenic regions to ensure a feasible embedding length and safety margin for subsequent processing.

[0066] 5) Excluding coding regions and key regulatory regions Filter out all regions annotated as protein-coding regions (CDS), rRNA, tRNA, and other regions that may contain important regulatory elements (such as promoters, operons, enhancers, etc.); If the annotation file contains information such as "repetitive elements" or "transposon / IS element", they can also be excluded as appropriate to avoid unstable integration.

[0067] 6) Setting length and position thresholds Set a minimum length threshold (such as 100 bp), and only retain intergenic regions that meet certain length requirements to ensure sufficient space for inserting the DNA watermark sequence and necessary redundancy and check codes; At the same time, an upper length limit can be set according to actual needs to avoid potential impacts on unknown functions after inserting into overly long intergenic regions.

[0068] 7) Assessing the safety distance of important genes For the genes on both sides of the intergenic region, if they belong to essential genes or genes in key metabolic pathways, a safety buffer distance can be increased; Refer to the list of essential genes or key metabolic pathway annotations in databases such as EcoCyc, KEGG, UniProt, etc., and carefully evaluate whether inserting the watermark too close will interfere with their expression regulation.

[0069] 8) Cross-validation with literature and databases Compare the information of the candidate intergenic regions obtained through preliminary screening with the literature data, or search in databases such as EcoCyc, UniProt, KEGG, etc. to check if there are reports that this region contains potential regulatory elements or unknown RNA genes; If important cis-acting elements (such as promoters, terminators, enhancers, etc.) are found, then exclude or treat this site with caution.

[0070] 9) Candidate site ranking Score and rank all candidate spacer regions according to the above analysis (metrics such as length, functional impact, database annotation, conservation, etc.); Determine the top 1 - 3 alternative sites with the highest priority to provide replacement options in case of unexpected failures or poor results during genome editing.

[0071] 10) Embedding site determination Finally, confirm an optimal site for the subsequent design of the CRISPR - Cas9 vector and the insertion of the watermark sequence; Record key metadata such as the coordinates of this site, upstream and downstream gene information, acceptable insertion sequence length, and the expected impact on strain function in the database or system.

[0072] 1.4 DNA watermark embedding unit: Use gene editing technology to precisely embed the DNA watermark sequence into the designated site.

[0073] Specifically: Embed the encrypted DNA watermark into a specific region of the strain genome through gene editing tools (such as CRISPR - Cas9). This region can be selected in the non - coding region of the strain genome or other parts that do not affect its functional expression to ensure the stability of the watermark and not interfere with the normal physiological functions of the strain.

[0074] For example, use the CRISPR - Cas9 tool to precisely locate the target region by designing guide RNA (gRNA) and use homologous recombination technology to achieve site - directed insertion. Taking the target region "ATG...TAA" as an example, after the watermark sequence "ACTGTGCC..." is embedded, the result is "ATGACTGTGCCTAA", which neither affects gene function expression nor ensures the stability of the watermark.

[0075] In addition, to prevent the DNA watermark from being tampered with or deleted, the present invention designs a redundant embedding mechanism. Specifically, segment the encrypted DNA watermark. For example, divide a 100 - bp long watermark into three segments, each with a length of 33 bp, and insert these segments at multiple sites in the genome. For example, the watermark "ACTGTGCC..." is divided into "ACTG...", "TGCC...", and "CCGG...", and are respectively embedded in three different regions of the genome. This multi - site embedding method can extract complete information through the remaining sites even if some sites are damaged, thus greatly improving the anti - tampering ability.

[0076] 1.5 Strain screening and verification unit: Screen and evaluate the functions of the strains into which the watermark has been successfully inserted to ensure that the strain performance is not affected.

[0077] Specifically: When it is necessary to verify the legitimacy of the strain, the following steps are used for dual identity authentication: Watermark extraction: Through PCR amplification and DNA sequencing technology, the embedded encrypted DNA watermark sequence is extracted from a specific region of the strain genome.

[0078] Decryption verification: The extracted DNA sequence is decoded into binary data after sequencing. The decoded data is decrypted with a specific key to obtain the original strain identifier. This identifier is compared with the strain ID in the record to verify its legitimacy. If the decryption is successful and the ID matches, it proves that the strain is a protected strain. In case of infringement, this phenomenon will serve as an important basis for safeguarding intellectual property rights.

[0079] For example, when actually verifying the legitimacy of a strain, first, the DNA watermark sequence is extracted from a specific region of the strain genome through PCR amplification and DNA sequencing technology. The extracted sequence is then decoded back into binary form and decrypted using a legitimate private key. If the decryption is successful and matches the original identifier in the database, for example, the identifier "3a7d67f..." is consistent with the database record, it proves that the strain is a legally protected strain; otherwise, it may indicate illegal replication or dissemination of the strain, and these data can serve as important evidence for legal rights protection.

[0080] In addition, the DNA watermark embedding module may further include the following units: Genomic data acquisition unit: It can perform whole-genome sequencing on the strain to obtain high-accuracy genomic sequence data.

[0081] System integration and coordination unit of the embedding module: It is responsible for the data flow management of each unit within this module and the interface connection with other modules.

[0082] 2.1 Blockchain node unit: It is used to maintain the nodes of the blockchain network, including full nodes and light nodes, to ensure the decentralization and high availability of the network.

[0083] 2.2 Smart contract execution unit: It is used to write, deploy, and manage smart contracts, and define the automated rules for strain copyright verification, authorization, and transactions.

[0084] 2.3 Data storage and indexing unit: It is used to permanently store the copyright information and its hash value on the blockchain to ensure the immutability and permanent traceability of the data.

[0085] 2.4 Transaction generation and submission unit: It is used to submit the constructed transaction to the blockchain network through the blockchain client or API and participate in the consensus mechanism verification.

[0086] 2.5 Consensus Mechanism Management Unit: It is used to implement and manage the consensus algorithm adopted by the blockchain network to ensure the consistency and security of the network.

[0087] 2.6 User Interface and API Unit: It is used to provide a user-friendly interface and API, allowing users to upload strain copyright information, submit transactions, query and verify the copyright status, etc.

[0088] 2.7 Monitoring and Log Management Unit: It is used to monitor the running status of the blockchain storage and verification module in real time, including node status, transaction processing, smart contract execution, etc.

[0089] 3.1 DNA Watermark Extraction Unit: It extracts the embedded DNA watermark sequence from a specific region of the genome of the strain to be verified.

[0090] 3.2 Information Decoding Unit: It decodes the extracted DNA watermark sequence according to the encoding rules and restores it to the original strain copyright information.

[0091] 3.3 Copyright Comparison and Verification Unit: It compares the decoded copyright information with the copyright information stored on the blockchain to verify the copyright ownership and integrity.

[0092] 3.4 Tracking and Tracing Unit: It records the transfer of strains between different locations and users to form a visual traceability path.

[0093] 3.5 Database Access Unit: It retrieves relevant copyright information and transaction records from the blockchain storage and verification module. It retrieves information from the blockchain and the local database uniformly to ensure data consistency.

[0094] 3.6 Report and Visualization Unit: It generates detailed copyright verification reports and traceability analysis reports based on the authentication and tracking results.

[0095] 3.7 User Notification Unit: It sends notifications and alerts to users in a timely manner when the strain copyright verification results and tracking events occur.

[0096] In addition, the strain authentication and tracking module also has an independent system integration and coordination unit: It is responsible for the data flow management of each unit within this module and the interface connection with other modules.

[0097] 4.1 Data Encryption Unit: It encrypts key data using an asymmetric encryption algorithm (such as RSA) to prevent leakage and tampering.

[0098] 4.2 Privacy Protection Unit: It protects user privacy through anonymization and zero-knowledge proof technologies.

[0099] 4.3 Access Control Unit: It implements multi-factor authentication (MFA) and role-based access control (RBAC).

[0100] 4.4 Security Audit and Log Management Unit: Real-time records system operation logs and uses automated analysis tools to detect anomalies.

[0101] 4.5 Key Management Unit: Securely generates, distributes, and stores encryption keys, and rotates them regularly.

[0102] 4.6 Security Policy and Compliance Management Unit: Ensures that the system complies with laws and regulations such as GDPR and ISO 27001.

[0103] In addition, the security and privacy protection module may also include the following units: Firewall and Intrusion Detection Unit: Monitors network traffic, blocks attacks, and implements automated responses.

[0104] Data Backup and Recovery Unit: Backs up critical data in multiple locations to provide disaster recovery capabilities.

[0105] Security Incident Response and Policy Management Unit: Responds quickly to sudden security incidents, conducts root cause analysis of incidents, dynamically adjusts security policies, and uniformly publishes and maintains them.

[0106] Reference Figure 6 , which is the flow step diagram of the strain copyright protection method provided by the present invention based on DNA watermark technology.

[0107] The system collects and collates copyright information (including owner information, patent number / registration number, protection period, etc.), uses the DNA sequence encoding unit to convert the copyright information into a DNA sequence, and adds redundancy and verification to improve stability.

[0108] According to the genome annotation results, select non-functional regions suitable for watermark embedding or spacer regions that do not affect the growth and metabolism of the strain; Adopt CRISPR-Cas9 or similar gene editing technologies to accurately embed the DNA watermark into the selected genomic loci.

[0109] Use selective markers or PCR screening to identify strains with successfully embedded watermarks; Ensure the correctness and stability of watermark embedding through sequencing and functional evaluation without affecting the normal functions of the strain.

[0110] The system packages the copyright information and its hash value into a transaction and submits it to the blockchain network; After verification by the consensus mechanism, the smart contract generates the corresponding copyright ID and completes the on-chain storage.

[0111] When the strain needs to be used or traded, the user can extract the DNA watermark from the strain through sequencing or PCR; The system decodes the DNA watermark and compares it with the records on the blockchain to confirm the copyright ownership and authorization scope; Meanwhile, update the usage and transfer records of the strain in the blockchain to form a traceable transfer chain.

[0112] Users can view the strain copyright information, usage records, and traceability paths through the system interface; The system automatically generates visual reports or authorization certificates to improve efficiency and transparency.

[0113] In addition, in the above steps, data and operations are encrypted, access-controlled, audited, and abnormal response operations are performed through security and privacy protection methods.

[0114] Background: A biotechnology company has been engaged in microbial transformation research for a long time and has successfully cultivated an Escherichia coli strain with unique metabolic capabilities. This strain can efficiently synthesize an important industrial compound (such as a new biocatalyst or special metabolite) under specific conditions, and both its economic value and innovation significance are quite remarkable.

[0115] To prevent the strain from being stolen or illegally replicated, the company decides to protect the copyright and manage the traceability of this strain through the "Strain Copyright Protection System and Method Based on DNA Watermark Technology" of the present invention.

[0116] Sample acquisition: The company's scientific research personnel take samples from the successfully cultivated target Escherichia coli strain to ensure the purity and viability of the strain.

[0117] Genome sequencing: Use high-throughput sequencing technology (such as Illumina) to perform whole-genome sequencing on the strain to obtain complete and accurate genome data.

[0118] Genome annotation: Use bioinformatics tools (such as Prokka, BLAST, etc.) to assemble and annotate the sequencing results to identify functional gene regions and non-functional regions.

[0119] Copyright information collection: The company confirms the key copyright information of the strain, including the owner (such as "XYZ Biotech Co., Ltd."), description of unique metabolic functions, registration number "XYZ-12345", protection period (such as 20 years), and applicable scope (such as industrial catalytic enzyme production).

[0120] Structuring of copyright information: Organize the above information into JSON or XML format for subsequent automated processing.

[0121] DNA encoding: Convert the copyright information into a DNA base sequence through the DNA encoding algorithm unit. To improve stability and anti-interference ability, redundant bits and error check codes are added during encoding to ensure that the watermark can still be accurately decoded after genome evolution or mutation.

[0122] Non-functional region screening: Based on the genome annotation results, target sites suitable for embedding DNA watermarks (such as certain spacer regions or pseudogene regions) are screened without affecting the unique metabolic functions of the strain.

[0123] Safety assessment: Through experiments or literature research, confirm that the embedded watermark will not affect the normal growth characteristics and unique metabolic pathways of the strain, while avoiding the insertion site being too close to the key regulatory region.

[0124] Functional verification: If there are multiple optional insertion sites, preliminary experimental insertions can be performed to test whether the strain undergoes obvious phenotypic changes or reduced product yields after insertion at different sites.

[0125] Vector construction: A guide RNA (gRNA) sequence is inserted into the CRISPR-Cas9 gene editing vector to enable it to cut at the selected site and carry a repair template containing a DNA watermark sequence.

[0126] Transformation and editing: The constructed editing vector is introduced into the target strain, the genome is cut at the predetermined position by the CRISPR-Cas9 system, and the DNA watermark sequence is inserted into the target site by homologous recombination.

[0127] Improved editing efficiency: Key reagents (such as proteins or small molecules that enhance Cas9 activity) can be added to the editing reaction system to improve editing efficiency and reduce off-target events.

[0128] Selective screening: Antibiotic resistance markers or fluorescent protein genes can be added to the watermark insertion sequence to preliminarily screen out positive strains that have completed the insertion.

[0129] PCR and sequencing verification: PCR amplification is performed on the screened positive strains, and specific primers for the insertion site are designed to verify whether the watermark sequence is successfully inserted. Sequencing (Sanger sequencing or high-throughput sequencing) is then performed to confirm the correctness and integrity of the inserted sequence and the accuracy of the insertion site.

[0130] Functional testing: The strains with successfully inserted watermarks are subjected to growth curve assays, product detection or enzyme activity analysis to ensure that their unique metabolic capabilities are not negatively affected.

[0131] Blockchain transaction construction: copyright information (such as owner, registration number, strain characteristic description, etc.) and corresponding hash values ​​are packaged into a transaction data packet.

[0132] On-chain submission: The transaction is submitted to the blockchain network through the blockchain storage and verification module and verified by the consensus mechanism (such as PoS).

[0133] Smart contract execution: Deploy a smart contract to automatically record copyright information and generate a unique copyright ID. Once the transaction is confirmed, the copyright information is permanently written to the blockchain, making it difficult to tamper with.

[0134] Strain delivery and use: When the company provides strains carrying DNA watermarks to partners or downstream enterprises, it can require the other party to conduct copyright verification before actual use.

[0135] DNA watermark extraction: Use PCR amplification to extract the watermark sequence from the strain; Information decoding: Decode the watermark sequence into the original copyright information according to the preset DNA coding rules; Blockchain comparison: The system retrieves the corresponding copyright records from the blockchain to confirm the copyright ownership and authorization scope. If the information matches, it proves that the use of the strain is compliant; if the information does not match or no corresponding record is retrieved, it indicates potential infringement or illegal use.

[0136] Traceability tracking: The system records every authorization, transfer, and usage path of the strain in the blockchain, enabling traceability in subsequent audits or rights protection.

[0137] Data encryption: During the process of uploading strain copyright information, submitting transactions, and user access, use the RSA encryption algorithm to prevent data theft; Multi-factor authentication: Adopt multi-factor authentication methods such as password + SMS verification or password + biometric identification for enterprise management personnel and authorized users to prevent account theft; Privacy protection: Anonymize user and transaction information or use zero-knowledge proof technology to prevent the leakage of trade secrets or personal privacy; Security audit: The system automatically records each operation log and conducts real-time security analysis, issuing alerts for any abnormal access or transactions.

[0138] Regular follow-up and testing: Conduct periodic tracking tests on strains successfully embedded with watermarks to confirm that the DNA watermark has not been accidentally mutated or deleted, while maintaining the expected metabolic performance of the strain; Report generation: The system can automatically generate PDF or online visual reports according to enterprise needs, covering information such as strain DNA watermark verification records, blockchain authorization history, and traceability paths, facilitating review by regulatory agencies or inspection by partners.

[0139] Technical effect: Through the above complete implementation steps, the target Escherichia coli strain has successfully embedded a highly secure DNA watermark that is difficult to remove or tamper with, and the copyright information is stored on the blockchain in an immutable manner. The company can verify the copyright ownership of the strain at any time and quickly trace the responsible party in case of intellectual property disputes. At the same time, the security and privacy protection modules ensure the legal compliance of the trade secrets and usage scenarios of all parties.

[0140] This embodiment effectively illustrates the actual application process and significant technical advantages of the present invention in the copyright protection of specific Escherichia coli strains, providing a practical reference for the copyright management of more biological strains.

[0141] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A bacterial strain copyright protection system based on DNA watermark technology, characterized in that: include: A DNA watermark embedding module, which is used to encode the copyright information of the bacterial strain into a DNA watermark sequence and embed it into a specific region of the genome of the bacterial strain; The blockchain storage and verification module is used to store the copyright information of bacterial strains and their hash values ​​on the blockchain, and automatically perform copyright verification, authorization and transactions through smart contracts; The strain authentication and tracking module is used to extract and decode the DNA watermark in the strain and compare it with the blockchain record to achieve strain copyright confirmation and traceability; The security and privacy protection module is used to implement encryption, protection, auditing and compliance management of various data and operations in the system.

2. The bacterial strain copyright protection system based on DNA watermarking technology according to claim 1 is characterized in that: The DNA watermark embedding module includes: a strain copyright information generating unit, a DNA sequence encoding unit, a DNA watermark embedding site selecting and analyzing unit, a strain screening and verifying unit and a DNA watermark embedding unit.

3. The bacterial strain copyright protection system based on DNA watermarking technology according to claim 2 is characterized in that: The blockchain storage and verification module includes: a blockchain node unit, a smart contract execution unit, a data storage and indexing unit, a transaction generation and submission unit, a consensus mechanism management unit, a user interface and API unit, and a monitoring and log management unit.

4. The bacterial strain copyright protection system based on DNA watermarking technology according to claim 3 is characterized in that: The strain authentication and tracking module includes: a DNA watermark extraction unit, an information decoding unit, a strain copyright comparison and verification unit, a tracking and tracing unit, a database access unit, a report and visualization unit and a user notification unit.

5. The bacterial strain copyright protection system based on DNA watermarking technology according to claim 4 is characterized in that: The security and privacy protection module includes: a data encryption unit, a privacy protection unit, an access control unit, a security audit and log management unit, a key management unit and a security policy and compliance management unit.

6. A method for protecting bacterial strain copyright based on DNA watermark technology, characterized in that: The following steps are involved: Step 1: Generate and encode the copyright information of bacterial strains: convert the copyright information of bacterial strains into a DNA watermark sequence and add redundancy and verification; Step 2: DNA watermark site selection and editing: inserting DNA watermark sequences into non-functional regions of bacterial genomes; Step 3: Strain screening and verification: Verify the correctness of the watermark sequence and the functionality of the strain through selective labeling and sequencing; Step 4: Upload the copyright information of the strain to the blockchain: Package the copyright information and hash value of the strain into a transaction and submit it to the blockchain network, and generate a copyright ID; Step 5: Use and tracking of strains: When the strain is actually used, the copyright ownership is confirmed and the usage path is recorded by extracting the DNA watermark and comparing it with the blockchain; Step 6. Report generation and result display: Provide users with visualized strain copyright status, usage records and traceability paths; In the above steps, data and operations are encrypted, access controlled, audited, and responded to exceptions through security and privacy protection.

7. The method for protecting bacterial strain copyright based on DNA watermarking technology according to claim 6, characterized in that: In step 2, CRISPR-Cas9 technology is used for genome editing to ensure high accuracy and low off-target rate of watermark insertion.

8. The method for protecting bacterial strain copyright based on DNA watermarking technology according to claim 6, characterized in that: In step 4, when uploading copyright information to the chain, multi-chain collaboration or cross-chain technology is supported to achieve compatibility and expansion of different blockchain networks.

9. The method for protecting bacterial strain copyright based on DNA watermarking technology according to claim 6, characterized in that: The visualization report generated in step 6 includes the location and time of use of the strain, user information, and scope of authorization for use, and can be exported to a third-party system or shared with partners.

Citation Information

Cited By

  • Information encryption method and system based on strain genome coding and identification

    CN121603197A

  • An information encryption method and system based on strain genome coding and identification

    CN121603197B

  • Copyright authentication method and device for double-key model and storage medium

    CN121841866A