A rare and endangered species DNA barcode encryption storage inference method and system
By constructing an anti-parsing database framework and a functionally restricted DNA barcode set, the problems of data security and identification accuracy in rare and endangered species monitoring have been solved, and the simultaneous output of rare species identity authentication, habitat assessment and endangered level determination has been achieved, thereby improving the decision-making efficiency of biodiversity conservation.
Patent Information
- Application Number
- CN202510963101.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-14
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-07-14
AI Technical Summary
Existing technologies for monitoring rare and endangered species have the following problems: morphological identification is highly subjective, and DNA analysis lacks a standardized encrypted storage mechanism, which makes genetic data vulnerable to illegal use, makes it difficult to achieve collaborative reasoning of morphological characteristics, geographical distribution, and genetic verification, and cannot support refined decision-making for biodiversity conservation.
The encrypted storage and inference method of DNA barcodes of rare and endangered species is adopted. By constructing an anti-parsing database framework, a barcode set of population-specific mutant bases is implanted to generate an encrypted barcode set. A single valid sequence is generated by randomly masking the primer region to achieve collaborative reasoning verification and comprehensive judgment, and the identification results are output in combination with morphological characteristics and geographical distribution.
It has achieved active protection of genetic data of rare species, improved the accuracy of non-destructive sample identification and data usage security, supported decision-making support for the trinity of morphology, space and population, and improved the decision-making efficiency of biodiversity conservation.
Smart Images

Figure CN120452560B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of encryption reasoning technology, and in particular to a method and system for encryption storage reasoning of DNA barcodes of rare and endangered species. Background Art
[0002] Against the backdrop of increasingly urgent global biodiversity conservation, the conservation of endemic species in the mountains of southwestern China, one of the world's 34 biodiversity hotspots, holds significant strategic value for maintaining ecosystem balance. Current monitoring of rare and endangered species relies primarily on traditional morphological identification and fragmented genetic data analysis. Existing technologies have significant limitations when dealing with non-invasive samples such as feces and hair obtained in the wild: Morphological methods are subject to the subjectivity of empirical judgment and the integrity of specimens. While conventional DNA analysis can improve accuracy, the lack of a standardized barcode encryption storage mechanism makes genetic data vulnerable to illegal exploitation. Furthermore, it is difficult to integrate multi-source ecological information, such as spatiotemporal distribution and population dynamics. More critically, existing solutions are unable to achieve collaborative reasoning between morphological characteristics, geographic distribution, and genetic verification while ensuring data security. This leaves key conservation efforts, such as habitat delineation and poaching tracing, at the fragmented decision-making level, severely restricting the refinement of biodiversity conservation.
[0003] Based on the above shortcomings of the existing technology, there is an urgent need for a rare and endangered species DNA barcode encryption storage inference method and system. Summary of the Invention
[0004] The present invention aims to provide a method for storing and reasoning DNA barcodes of rare and endangered species to improve the above-mentioned problem. To achieve the above-mentioned object, the present invention adopts the following technical solutions:
[0005] In a first aspect, the present application provides a method for encrypted storage and reasoning of DNA barcodes of rare and endangered species, comprising:
[0006] Obtain rare and endangered species monitoring datasets and corresponding biological sample sets;
[0007] An encrypted database is constructed based on the monitoring data set, and a composite coding structure is generated by fusing monitoring time and geographic coordinates to obtain an anti-parsing database framework;
[0008] Performing DNA barcode encryption on the biological sample set, generating a barcode set with restricted functionality by implanting population-specific mutant bases in the standard primer region, obtaining an encrypted barcode set and storing it in the anti-parsing database framework;
[0009] Obtaining a DNA sample to be identified, generating a verification sequence based on the DNA sample to be identified, and generating a single valid sequence by randomly masking three consecutive base positions in the primer region;
[0010] Performing collaborative reasoning and verification based on the single valid sequence and the encrypted barcode set to generate a species identification code;
[0011] A comprehensive judgment is made based on the species identification code and the monitoring data set, and an identification result is output by correlating changes in morphological characteristics with geographical distribution density. The identification result includes the species name, degree of morphological consistency, distribution probability range and endangered status level.
[0012] In a second aspect, the present application also provides a rare and endangered species DNA barcode encryption storage and inference system, comprising:
[0013] Acquisition module, used to obtain rare and endangered species monitoring data sets and corresponding biological sample sets;
[0014] A construction module is used to construct an encrypted database based on the monitoring data set, generate a composite coding structure by fusing monitoring time and geographic coordinates, and obtain an anti-parsing database framework;
[0015] an encryption module for encrypting DNA barcodes based on the biological sample set, generating a barcode set with restricted functionality by implanting population-specific mutant bases in the standard primer region, obtaining an encrypted barcode set, and storing the encrypted barcode set in the anti-parsing database framework;
[0016] A generation module is used to obtain a DNA sample to be identified, generate a verification sequence based on the DNA sample to be identified, and generate a single valid sequence by randomly masking three consecutive base positions in the primer region;
[0017] an inference module, configured to perform collaborative inference verification based on the single valid sequence and the encrypted barcode set to generate a species identification code;
[0018] The identification module is used to make a comprehensive judgment based on the species identification code and the monitoring data set, and output the identification results by correlating the changes in morphological characteristics with the geographical distribution density. The identification results include the species name, the degree of morphological consistency, the distribution probability range and the endangered status level.
[0019] The beneficial effects of the present invention are:
[0020] The present invention realizes the active protection of genetic data of rare species by constructing an anti-parsing encrypted database framework and a functionally restricted DNA barcode set, thus solving the key hidden danger of sensitive biological information being used for poaching tracing; by introducing species-specific mutation implantation and dynamic masking mechanisms, it ensures the security of data use while improving the accuracy of non-destructive sample identification; through multi-level collaborative reasoning and verification, the morphological manifold characteristics, geographic distribution parameters and genetic identifiers are deeply coupled, and the simultaneous output of species identity verification, habitat assessment and endangered level determination is achieved within a single technical framework, greatly improving the efficiency of conservation decision-making in biodiversity hotspots. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.
[0022] Figure 1 A schematic diagram of a rare and endangered species DNA barcode encryption storage reasoning method according to an embodiment of the present invention;
[0023] Figure 2 This is a schematic diagram of the structure of a rare and endangered species DNA barcode encryption storage and inference system according to an embodiment of the present invention;
[0024] Figure 3 This is a schematic diagram of the structure of a rare and endangered species DNA barcode encryption storage and inference device described in an embodiment of the present invention.
[0025] Markings in the figure: 800, a rare and endangered species DNA barcode encryption storage and inference device; 801, processor; 802, memory; 803, multimedia component; 804, I / O interface; 805, communication component; 901, acquisition module; 902, construction module; 903, encryption module; 904, generation module; 905, inference module; 906, identification module. DETAILED DESCRIPTION
[0026] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. The components of the embodiments of the present invention generally described and shown in the drawings herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0027] It should be noted that similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings. At the same time, in the description of the present invention, the terms "first", "second", etc. are used only to distinguish the description and should not be understood as indicating or implying relative importance.
[0028] Example 1:
[0029] This embodiment provides a method for inferring the encrypted storage of DNA barcodes of rare and endangered species.
[0030] See also Figure 1 , the figure shows that the method includes steps S100 to S600.
[0031] Step S100: Acquire a rare and endangered species monitoring data set and a corresponding biological sample set;
[0032] Understandably, monitoring datasets encompass multidimensional spatiotemporal information, including species movement patterns, population changes, and habitat parameters; biological sample collections specifically refer to non-invasive genetic material collected in the wild, such as feces and hair. This step aims to address the fragmented information found in traditional conservation efforts by integrating ecological monitoring records and physical samples of species endemic to the study area. Its core value lies in systematically linking dispersed field survey data with laboratory-analyzable samples, providing a unified input source for subsequent encrypted storage and collaborative analysis. This directly addresses the current situation in which research institutions struggle to effectively integrate massive amounts of data.
[0033] Step S200: construct an encrypted database based on the monitoring data set, generate a composite coding structure by fusing the monitoring time and geographic coordinates, and obtain an anti-parsing database framework;
[0034] It's important to note that this step generates a non-reversible composite encoding structure through the nonlinear fusion of monitoring time (e.g., breeding season dates) and geographic coordinates (e.g., habitat latitude and longitude). This structure leverages the chaotic nature of temporal parameters and the topological relationships of spatial coordinates to create a database index framework resistant to reverse engineering. Its purpose is to transform sensitive spatiotemporal information into mathematical abstractions, preserving data relevance while preventing location inference. This prevents poachers from exploiting monitoring data to locate rare species, achieving an underlying architectural innovation that achieves "data available, invisible."
[0035] Step S300: encrypting DNA barcodes based on the biological sample set, generating a barcode set with restricted functionality by implanting population-specific mutant bases in the standard primer region, obtaining an encrypted barcode set and storing it in an anti-parsing database framework;
[0036] As can be understood, population-specific mutant bases are targeted into the standard primer region (the starting point for polymerase chain reaction amplification): frequent mutation sites are screened based on the species' genetic characteristics, and base substitutions are performed at homologous positions in the primer sequence. The resulting barcode set retains the core recognition region functionality but loses primer binding activity, rendering the stored data unusable for secondary amplification. This technology addresses the risk of genetic data theft by achieving "storage-as-encryption" through biometric binding and modification, maintaining identification accuracy while preventing illegal sample duplication, addressing a core pain point in biobank security management.
[0037] Step S400: Obtain a DNA sample to be identified, generate a verification sequence based on the DNA sample to be identified, and generate a single valid sequence by randomly masking three consecutive base positions in the primer region;
[0038] It should be noted that this step dynamically masks the primer region of the DNA sample to be identified. Specifically, three consecutive base positions are randomly selected and replaced with invalid characters to generate a one-time verification sequence. This operation, based on the principle of polymerase chain reaction amplification, destroys primer function by masking key sites. Its core value lies in establishing a "self-destruct mechanism" for sample use, ensuring that the sequence is invalid after a single verification, preventing malicious reuse of field sample data during circulation.
[0039] Step S500: Perform collaborative reasoning and verification based on the single valid sequence and the encrypted barcode set to generate a species identification code;
[0040] As you can understand, this step generates species identification through dual coupled verification: first, the masked sequence is parsed to restore the complete barcode, and then the core identification region of the sample under test is compared with the encrypted barcode set. This process incorporates population genetic benchmarks from the monitoring database. Unlike traditional full-sequence comparison, only key region matching is required to determine species. This lightweight "partial verification-global confirmation" verification mechanism ensures identification accuracy while increasing verification efficiency for common species by over 50%.
[0041] Step S600: Perform a comprehensive assessment based on the species identification code and monitoring data set, and output an identification result by correlating morphological feature changes with geographic distribution density. The identification result includes the species name, morphological consistency, distribution probability range, and endangered status level.
[0042] It is important to note that this step maps morphological parameters to manifold spatial curvature characteristics, calculates their geodesic distance to the historical baseline, and quantifies the degree of morphological consistency. The species diffusion equation is then solved using habitat parameters to generate a presence probability distribution map and extract the core distribution range. Finally, genetic isolation and population size parameters are integrated to construct a stochastic survival model to output the endangered species level. This step, through the intersection of manifold geometry and biogeography, transcends the limitations of single-species identification and achieves a trinity of "morphological, spatial, and population" decision support.
[0043] Corresponding to the above steps and methods, the present invention realizes a complete technical closed loop from field sample collection to standardized database construction: first, through a standardized process, non-destructive samples of rare species are obtained and matched with spatiotemporal monitoring data, and a strict mapping between sample codes and ecological parameters is established; then, based on the standard operating procedure of DNA barcoding, the biological samples are encrypted using primer region directed mutagenesis technology (to generate a set of functionally restricted barcodes carrying the genetic fingerprint of the population); then, an anti-parsing database framework integrating spatiotemporal encryption indexes is constructed, and the standardized barcodes of 100,000 biological samples are deeply bound to habitat characteristics; finally, through a multi-source data integration mechanism, a bioinformatics database platform covering 56 rare species in Sichuan's biological hotspots is formed, realizing standardized management of the entire process from "sample collection-barcode generation-data storage."
[0044] Furthermore, step S200 includes steps S210 to S230.
[0045] Step S210: Perform phenological cycle encryption processing based on the monitoring data set, and obtain a phenological encryption vector by converting the start and end dates of the breeding season into the initial phase of the chaotic oscillator;
[0046] Step S220: performing habitat spatial binding processing based on the phenological encryption vector, and obtaining a habitat fusion factor by performing niche convolution operation on the core habitat coordinates and the phenological vector;
[0047] Step S230: embedding the population genetic structure according to the habitat fusion factor, generating a biological constraint coding structure by associating the effective population size with the fractal dimension, and obtaining an anti-parsing database framework.
[0048] Furthermore, step S300 includes steps S310 to S330.
[0049] Step S310: performing species-specific mutation hotspot location processing based on the biological sample set, generating a mutation fingerprint by identifying highly variable sites in the mitochondrial control region, and obtaining a set of targeted mutation coordinates;
[0050] This step analyzes the genetic sequences of the mitochondrial control region of the biological sample collection, screening for high-frequency mutation sites within the population (mutation frequency >55%) and constructing a mutation fingerprint map containing site coordinates and variation characteristics. This map identifies a set of target coordinates representative of the population, such as site 15,794 in the giant panda control region, providing a standardized biological target foundation for subsequent encryption and addressing the lack of objective basis for mutation site selection in traditional barcoding technology.
[0051] Step S320: performing biological mirroring of the primer region according to the target mutation coordinate set, generating a variant primer set by mapping the high-frequency mutation sites to homologous positions in the standard primer region, and obtaining biometric binding primers;
[0052] This step maps high-frequency sites in the mutation fingerprint to homologous positions in a standard primer sequence (e.g., base 7 in the COI gene Folmer primer), generating biomarker-binding primers through base-directed substitution. This process preserves the primer framework but incorporates species-specific genetic markers (e.g., A→G transitions). This allows the primers to amplify while preventing unauthorized population tracing, thus overcoming the universal limitations of conventional primer design.
[0053] Step S330: Perform restricted amplification encryption processing based on the biometric binding primers, implant targeted mutations through annealing temperature offset control to generate function-restricted barcodes, obtain an encrypted barcode set and store it in the anti-parsing database framework.
[0054] This step utilizes biomarker-binding primers for polymerase chain reaction (PCR). Precisely controlling the annealing temperature offset (e.g., a ±2°C gradient) allows for efficient targeted mutagenesis. The generated barcodes fail under standard amplification conditions (failure rate >99%), and core region identification functionality can only be restored through temperature adaptation. Ultimately, the encrypted barcode set is integrated into a robust database framework, creating a standardized storage solution that combines species identification with data protection.
[0055] This process realizes the standardized operation system of DNA barcode "generation-encryption-storage" through three levels of innovation: biological target positioning-homologous encryption modification-temperature restricted amplification, providing a replicable technical template for the construction of biological genetic resource libraries.
[0056] Furthermore, step S400 includes steps S410 to S430.
[0057] Step S410: Predicting the species type based on the DNA sample to be identified, determining the potential species classification by screening the base composition of the mitochondrial conserved region, and obtaining a predicted species identifier;
[0058] This step rapidly screens the base composition patterns of conserved mitochondrial regions (such as the cytochrome b gene) within the DNA sample to be identified, comparing them to a library of known species signatures (such as the 146-site deletion mutation unique to the giant panda) to determine the sample's potential species classification. This process, based on the stability of conserved region sequences and interspecies variability, outputs a predicted species identifier (such as "PANDA") within 10 minutes, addressing the inefficiency of early classification in traditional identification and providing a rapid classification basis for subsequent targeted conservation efforts.
[0059] Step S420: Masking the primer region risk based on the predicted species identifier, dynamically adjusting the masking position in conjunction with the poaching heat map to generate a directional masking sequence;
[0060] This step dynamically selects masked primer regions based on data related to predicted illegal activity hotspots associated with species identification: core functional sites (positions 1-3) are masked for high-risk species, while auxiliary regions (positions 9-11) are masked for medium- and low-risk species. By replacing specific segments of a standard primer sequence (e.g., LCO1490) with "NNN," a targeted masked sequence is generated that is resistant to secondary amplification. This mechanism surpasses conventional full-sequence encryption, achieving "risk-adaptive precision masking."
[0061] Step S430: Perform time-sensitive encryption processing according to the directional masking sequence, and generate a single valid sequence including a self-destruction countdown by binding the sample collection time.
[0062] This step aligns sample collection time with species phenological characteristics (e.g., breeding season / hibernation period), sets differentiated self-destruction countdowns (breeding season samples are valid for 15 days, hibernation samples are valid for 60 days), and embeds a timestamp and countdown code (e.g., "EF15|20250801") in the sequence header. When the validity period expires, the sequence automatically obfuscates key bases, rendering the data invalid. This design integrates biological principles with data security requirements to establish a time-driven self-destruction verification mechanism.
[0063] This process achieves full life cycle management of the DNA sample identification process through three levels of protection: classification prediction, risk response, and time self-destruction. This ensures that each field sample enters an expiration countdown from the start of identification, completely eliminating the risk of secondary abuse of genetic data.
[0064] Furthermore, step S500 includes steps S510 to S530.
[0065] Step S510: performing primer masking recovery processing based on the single valid sequence and the encrypted barcode set, restoring the complete sequence fragment by matching the variant primer framework to obtain the unmasked DNA sequence;
[0066] This step, based on a pre-set variant primer framework (preferably, a Folmer primer sequence containing targeted mutagenesis), analyzes the masked regions of a single valid sequence using the principle of base complementarity, restoring the randomly masked invalid characters to their original functional bases to generate a complete, unmasked DNA sequence. This process utilizes the primer mutation signatures stored in the encrypted barcode as a decoding key, overcoming the limitation of traditional verification requiring full sequence decryption. It allows for the precise recovery of masked sites while ensuring that the original sample data is not exposed.
[0067] Step S520: performing genetic similarity calculation based on the unmasked DNA sequence, and generating a genetic consistency index by comparing the base matching rate of the sequence to be tested with the core region of the encrypted barcode set;
[0068] This step compares the base matches between the unmasked sequence and the encrypted barcode set within the core recognition region (e.g., bases 200-600 of the mitochondrial cytochrome oxidase I gene), calculates the proportion of matching bases among all sites, and generates a genetic match index (GCI) ranging from 0 to 1. This process focuses on the core region (20% of the total length) rather than the entire sequence, addressing the compatibility issues of traditional methods with degraded samples. The method allows for up to 15% base mismatches while still effectively determining species, significantly improving the identification success rate for low-quality samples such as field feces.
[0069] Step S530: Perform threshold determination processing based on the genetic matching index, and generate a species identification code by comparing the genetic matching index with the species minimum matching threshold.
[0070] This step dynamically adjusts the minimum matching threshold based on the species' endangerment level. When the genetic match index exceeds the threshold, the species identifier (e.g., "PANDA-092") is output; otherwise, it is marked as unknown. This determination integrates the genetic baseline of the population in the monitoring database, balancing conservation needs with identification accuracy through a floating threshold mechanism: leniency is relaxed for endangered species to save marginal populations, while strict control is maintained for stable species to prevent misidentification.
[0071] Furthermore, step S600 includes steps S610 to S630.
[0072] Step S610: performing morphological feature manifold alignment processing according to the species identification code, mapping the anatomical features to the Riemannian manifold space and calculating the geodesic distance between the Gaussian curvature field and the reference manifold to obtain the degree of morphological fit;
[0073] This step maps anatomical features (such as skull curvature and limb proportions) into a Gaussian curvature field in a Riemannian manifold, and calculates the geodesic distance deviation between the measured morphology and the standard reference manifold to generate a quantitative degree of morphological fit. This technology breaks through the dimensional limitations of traditional caliper measurements, and unifies two-dimensional morphological parameters (such as pattern direction) and three-dimensional structural features (such as bone angles) in a continuous curvature space for analysis. Its core lies in establishing a topological invariant expression of morphological features - for example, converting the angular length ratio into the ratio of the principal curvature of the surface, and mapping the distribution of fur color patches into a gradient change pattern of the curvature field. This manifold space expression is not only compatible with incomplete samples (local curvature can still be calculated even if some features are missing), but also realizes the continuous quantification of morphological differences, providing a millimeter-level precision scale for monitoring the degradation of endangered populations. Specifically, the morphological fit calculation model is:
[0074]
[0075] in, M match Indicates the degree of morphological fit; K ref is the Gaussian curvature field of the standard reference form, representing the historical population baseline; K sample is the Gaussian curvature field of the sample to be tested, representing the measured anatomical parameter mapping; d g represents the geodesic distance between curvature fields, which is a geometric invariant used to quantify morphological differences; λ represents the sensitivity adjustment factor; is a Riemannian manifold space, representing the topological space of anatomical feature mapping.
[0076] Step S620: Deducing the geographic distribution probability based on the degree of morphological consistency, solving the time-varying diffusion equation including the morphological gradient term and performing spatial discretization calculation using the adaptive finite element method, to obtain the distribution probability range;
[0077] This step is based on a diffusion model driven by the degree of morphological matching, integrating species migration capacity parameters (maximum diffusion radius) and environmental barrier data (rivers, roads) to construct an adaptive hierarchical grid system. A hundred-meter grid is used in the core habitat to capture microhabitat differences, and a kilometer-level grid is used in the edge area to improve computational efficiency. By iteratively solving the spatial propagation equation, a thermal distribution map of the probability of species existence is output, and three types of key areas are extracted: core areas (probability > 0.8), diffusion corridors (probability 0.5-0.8) and isolation zones (probability < 0.3). The coupled modeling of morphological adaptability (such as the matching degree of the paw morphology and the terrain) and geographic diffusion is achieved to solve the defect of traditional models that ignore biological functional traits. Specifically, the time-varying diffusion equation is:
[0078]
[0079] in, P represents the probability density of species existence; D represents the diffusion coefficient, which is calibrated by the migration rate of monitoring data; α represents the population growth rate; K represents the environmental carrying capacity; β represents the morphological gradient coupling coefficient, the default value is 0.3; Indicates the spatial variation rate of the degree of morphological fit; represents the Laplace operator; t represents the time index; δ represents the partial derivative.
[0080] Step S630: Analyze the population viability based on the degree of morphological consistency and the distribution probability range, construct a random population model by integrating genetic isolation, effective population size, and habitat connectivity parameters, and generate an endangered status level;
[0081] This step, strictly adhering to the international endangered species classification framework, integrates morphological consistency, genetic isolation, and habitat connectivity parameters for localized dynamic classification. A "genetic degradation" sub-level is automatically triggered when morphological consistency falls below 0.6; an "isolated population" designation is added when the effective population size is less than 50 and the distribution is fragmented (connectivity index <0.3); and a "habitat crisis" sub-level is upgraded when the core distribution area shrinks by more than 30% compared to the historical baseline. This process precisely matches the priority needs of actual conservation scenarios by embedding regional response rules in the algorithm (such as a compensation mechanism for altitude isolation of mountain species).
[0082] Step S640: Generate an identification result based on the degree of morphological matching, the distribution probability range, and the generated endangered status level.
[0083] This step constructs a three-in-one decision matrix of "biological characteristics-geographic distribution-population risk" by integrating the degree of morphological fit (quantized value of manifold space curvature), distribution probability range (core area boundary deduced by adaptive grid) and dynamically refined endangered level (including sub-level coding); and finally outputs a structured report.
[0084] Example 2:
[0085] like Figure 2 As shown, this embodiment provides a rare and endangered species DNA barcode encryption storage and inference system, the system comprising:
[0086] Acquisition module 901, used to obtain rare and endangered species monitoring data sets and corresponding biological sample sets;
[0087] Construction module 902 is used to construct an encrypted database based on the monitoring data set, generate a composite coding structure by fusing the monitoring time and geographic coordinates, and obtain an anti-parsing database framework;
[0088] Encryption module 903 is used to encrypt DNA barcodes based on the biological sample set by implanting population-specific mutant bases in the standard primer region to generate a barcode set with restricted functionality, obtain the encrypted barcode set, and store it in the anti-parsing database framework;
[0089] A generation module 904 is used to obtain a DNA sample to be identified, generate a verification sequence based on the DNA sample to be identified, and generate a single valid sequence by randomly masking three consecutive base positions in the primer region;
[0090] Inference module 905, for performing collaborative inference verification based on a single valid sequence and an encrypted barcode set to generate a species identification code;
[0091] Identification module 906 is used to make a comprehensive judgment based on the species identification code and monitoring data set, and output the identification results by correlating morphological feature changes with geographical distribution density. The identification results include species name, morphological consistency, distribution probability range and endangered status level.
[0092] In a specific embodiment disclosed in this application, the building block 902 includes:
[0093] The first construction unit is used to perform phenological cycle encryption processing based on the monitoring data set, and obtain the phenological encryption vector by converting the start and end dates of the breeding season into the initial phase of the chaotic oscillator;
[0094] The second construction unit is used to perform habitat spatial binding processing based on the phenological encryption vector, and obtain the habitat integration factor by performing niche convolution operation on the core habitat coordinates and the phenological vector;
[0095] The third construction unit is used to embed the population genetic structure according to the habitat fusion factor, generate the biological constraint coding structure by associating the effective population size with the fractal dimension, and obtain the anti-parsing database framework.
[0096] In a specific embodiment disclosed in this application, the encryption module 903 includes:
[0097] The first encryption unit is used to perform species-specific mutation hotspot location processing based on the biological sample set, generate a mutation fingerprint by identifying the highly variable sites in the mitochondrial control region, and obtain a set of targeted mutation coordinates;
[0098] The second encryption unit is used to perform biological mirroring of the primer region according to the target mutation coordinate set, and generate a variant primer set by mapping the high-frequency mutation sites to the homologous positions in the standard primer region to obtain the biological feature binding primers;
[0099] The third encryption unit is used to perform restricted amplification encryption processing based on the biometric binding primers, implant directed mutations through annealing temperature offset control to generate function-restricted barcodes, obtain an encrypted barcode set and store it in the anti-parsing database framework.
[0100] In a specific embodiment disclosed in this application, the generation module 904 includes:
[0101] The first generation unit is used to pre-judge the species type based on the DNA sample to be identified, determine the potential species classification by screening the base composition of the mitochondrial conserved region, and obtain the predicted species identification;
[0102] The second generation unit is used to perform risk masking of the primer region based on the predicted species identification, and to generate a directional masking sequence by dynamically adjusting the masking position in combination with the poaching heat map;
[0103] The third generation unit is used to perform time-sensitive encryption processing according to the directional masking sequence, and generate a single valid sequence including a self-destruction countdown by binding the sample collection time.
[0104] In a specific embodiment disclosed in this application, the reasoning module 905 includes:
[0105] The first inference unit is used to perform primer masking recovery processing based on the single valid sequence and the encrypted barcode set, restore the complete sequence fragment by matching the variant primer framework, and obtain the unmasked DNA sequence;
[0106] The second inference unit is used to perform genetic similarity calculation based on the unmasked DNA sequence, and generate a genetic consistency index by comparing the base matching rate of the sequence to be tested with the core region of the encrypted barcode set;
[0107] The third reasoning unit is used to perform threshold determination processing according to the genetic matching index, and generate a species identification code by comparing the relationship between the genetic matching index and the species minimum matching threshold.
[0108] In a specific embodiment disclosed in this application, the identification module 906 includes:
[0109] The first identification unit is used to perform morphological feature manifold alignment processing based on the species identification code, map the anatomical features to the Riemannian manifold space and calculate the geodesic distance between the Gaussian curvature field and the reference manifold to obtain the degree of morphological consistency;
[0110] The second identification unit is used to deduce the geographical distribution probability based on the degree of morphological consistency. The distribution probability range is obtained by solving the time-varying diffusion equation containing the morphological gradient term and using the adaptive finite element method for spatial discretization calculation;
[0111] The third identification unit is used to analyze population viability based on the degree of morphological consistency and distribution probability range, and to construct a random population model by integrating genetic isolation, effective population size and habitat connectivity parameters to generate an endangered status level;
[0112] The fourth identification unit is used to generate an identification result based on the degree of morphological matching, the distribution probability range and the generated endangered status level.
[0113] Example 3:
[0114] Corresponding to the above method embodiment, this embodiment also provides a rare and endangered species DNA barcode encryption storage and reasoning device. The rare and endangered species DNA barcode encryption storage and reasoning device described below and the rare and endangered species DNA barcode encryption storage and reasoning method described above can be referenced to each other.
[0115] Figure 3 FIG. 8 is a block diagram of a rare and endangered species DNA barcode encryption storage inference device 800 according to an exemplary embodiment. Figure 3 As shown, the rare and endangered species DNA barcode encryption storage inference device 800 may include: a processor 801, a memory 802. The rare and endangered species DNA barcode encryption storage inference device 800 may also include one or more of a multimedia component 803, an I / O interface 804, and a communication component 805.
[0116] The processor 801 is used to control the overall operation of the device 800 for encrypted storage and inference of DNA barcodes of rare and endangered species to complete all or part of the steps in the aforementioned method for encrypted storage and inference of DNA barcodes of rare and endangered species. The memory 802 is used to store various types of data to support the operation of the device 800. This data may include, for example, instructions for any application or method operating on the device 800, as well as application-related data such as contact information, sent and received messages, images, audio, video, and the like. The memory 802 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk. The multimedia component 803 may include a screen and an audio component. The screen may be, for example, a touch screen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone for receiving external audio signals. The received audio signal may be further stored in the memory 802 or transmitted via the communication component 805. The audio component also includes at least one speaker for outputting audio signals. The I / O interface 804 provides an interface between the processor 801 and other interface modules, and the above-mentioned other interface modules can be a keyboard, a mouse, buttons, etc. These buttons can be virtual buttons or physical buttons. The communication component 805 is used for wired or wireless communication between the rare and endangered species DNA barcode encryption storage inference device 800 and other devices. Wireless communication, such as Wi-Fi, Bluetooth, near field communication (NFC), 2G, 3G or 4G, or a combination of one or more of them, so the corresponding communication component 805 can include: Wi-Fi module, Bluetooth module, NFC module.
[0117] In an exemplary embodiment, a rare and endangered species DNA barcode encryption storage inference device 800 can be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components to execute the above-mentioned rare and endangered species DNA barcode encryption storage inference method.
[0118] In another exemplary embodiment, a computer-readable storage medium including program instructions is also provided. When executed by a processor, the program instructions implement the steps of the aforementioned method for inferring DNA barcodes for encrypted storage of rare and endangered species. For example, the computer-readable storage medium may be the aforementioned memory 802 including the program instructions. The program instructions may be executed by the processor 801 of the device 800 for inferring DNA barcodes for encrypted storage of rare and endangered species to implement the aforementioned method for inferring DNA barcodes for encrypted storage of rare and endangered species.
[0119] The above description is only a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed by the present invention, which should be covered by the scope of protection of the present invention.
Claims
1. A rare and endangered species DNA barcode encryption storage inference method, characterized by: include: Obtain rare and endangered species monitoring datasets and corresponding biological sample sets; An encrypted database is constructed based on the monitoring data set, and a composite coding structure is generated by fusing monitoring time and geographic coordinates to obtain an anti-parsing database framework; Performing DNA barcode encryption on the biological sample set, generating a barcode set with restricted functionality by implanting population-specific mutant bases in the standard primer region, obtaining an encrypted barcode set and storing it in the anti-parsing database framework; Obtaining a DNA sample to be identified, generating a verification sequence based on the DNA sample to be identified, and generating a single valid sequence by randomly masking three consecutive base positions in the primer region; Performing collaborative reasoning and verification based on the single valid sequence and the encrypted barcode set to generate a species identification code; A comprehensive judgment is made based on the species identification code and the monitoring data set, and an identification result is output by correlating changes in morphological characteristics with geographical distribution density. The identification result includes the species name, degree of morphological consistency, distribution probability range and endangered status level.
2. The method for encrypted storage and reasoning of DNA barcodes of rare and endangered species according to claim 1, characterized in that: An encrypted database is constructed based on the monitoring data set. A composite coding structure is generated by fusing monitoring time and geographic coordinates to obtain an anti-parsing database framework, including: Performing phenological cycle encryption processing on the monitoring data set, and obtaining a phenological encryption vector by converting the start and end dates of the breeding season into the initial phase of a chaotic oscillator; Performing habitat spatial binding processing according to the phenological encryption vector, and obtaining a habitat fusion factor by performing niche convolution operation on the core habitat coordinates and the phenological vector; The population genetic structure is embedded according to the habitat fusion factor, and a biological constraint coding structure is generated by associating the effective population size with the fractal dimension to obtain an anti-parsing database framework.
3. The method for encrypted storage and reasoning of DNA barcodes of rare and endangered species according to claim 1, characterized in that: DNA barcode encryption is performed according to the biological sample set, and a barcode set with limited functionality is generated by implanting population-specific mutant bases in the standard primer region, to obtain an encrypted barcode set and store it in the anti-parsing database framework, including: Performing species-specific mutation hotspot location processing based on the biological sample set, generating a mutation fingerprint by identifying highly variable sites in the mitochondrial control region, and obtaining a targeted mutation coordinate set; Performing biological mirroring of the primer region according to the target mutation coordinate set, generating a variant primer set by mapping the high-frequency mutation sites to homologous positions in the standard primer region, and obtaining biological feature binding primers; Restricted amplification encryption processing is performed according to the biometric binding primers, and targeted mutations are implanted through annealing temperature offset control to generate function-restricted barcodes, and an encrypted barcode set is obtained and stored in the anti-parsing database framework.
4. The method for encrypted storage and reasoning of DNA barcodes of rare and endangered species according to claim 1, characterized in that: Obtain a DNA sample to be identified, generate a verification sequence based on the DNA sample to be identified, and generate a single valid sequence by randomly masking three consecutive base positions in the primer region, including: Predict species types based on the DNA sample to be identified, determine potential species classification by screening the base composition of the mitochondrial conserved region, and obtain the predicted species identification; Performing risk masking of the primer region based on the predicted species identification, and dynamically adjusting the masking position in combination with the poaching heat map to generate a directional masking sequence; A time-sensitive encryption process is performed according to the directional masking sequence, and a single valid sequence including a self-destruction countdown is generated by binding the sample collection time.
5. The method for encrypted storage and reasoning of DNA barcodes of rare and endangered species according to claim 1, characterized in that: Performing collaborative reasoning and verification based on the single valid sequence and the encrypted barcode set to generate a species identification code, including: Performing primer masking recovery processing based on the single valid sequence and the encrypted barcode set, restoring the complete sequence fragment by matching the variant primer framework, and obtaining an unmasked DNA sequence; Performing genetic similarity calculation based on the unmasked DNA sequence, and generating a genetic match index by comparing the base matching rate of the sequence to be tested with the core region of the encrypted barcode set; A threshold determination process is performed according to the genetic matching index, and a species identification code is generated by comparing the genetic matching index with a species minimum matching threshold.
6. A rare and endangered species DNA barcode encryption storage and inference system, characterized by: include: Acquisition module, used to obtain rare and endangered species monitoring data sets and corresponding biological sample sets; A construction module is used to construct an encrypted database based on the monitoring data set, generate a composite coding structure by fusing monitoring time and geographic coordinates, and obtain an anti-parsing database framework; an encryption module for encrypting DNA barcodes based on the biological sample set, generating a barcode set with restricted functionality by implanting population-specific mutant bases in the standard primer region, obtaining an encrypted barcode set, and storing the encrypted barcode set in the anti-parsing database framework; A generation module is used to obtain a DNA sample to be identified, generate a verification sequence based on the DNA sample to be identified, and generate a single valid sequence by randomly masking three consecutive base positions in the primer region; an inference module, configured to perform collaborative inference verification based on the single valid sequence and the encrypted barcode set to generate a species identification code; The identification module is used to make a comprehensive judgment based on the species identification code and the monitoring data set, and output the identification results by correlating the changes in morphological characteristics with the geographical distribution density. The identification results include the species name, the degree of morphological consistency, the distribution probability range and the endangered status level.
7. The rare and endangered species DNA barcode encryption storage and inference system according to claim 6, characterized in that: The building blocks include: The first construction unit is configured to perform phenological cycle encryption processing based on the monitoring data set, and obtain a phenological encryption vector by converting the start and end dates of the breeding season into the initial phase of a chaotic oscillator; The second construction unit is used to perform habitat space binding processing according to the phenological encryption vector, and obtain a habitat fusion factor by performing niche convolution operation on the core habitat coordinates and the phenological vector; The third construction unit is used to embed the population genetic structure according to the habitat fusion factor, generate a biological constraint coding structure by associating the effective population size with the fractal dimension, and obtain an anti-parsing database framework.
8. The rare and endangered species DNA barcode encryption storage and inference system according to claim 6, characterized in that: The encryption module includes: A first encryption unit is used to perform species-specific mutation hotspot location processing based on the biological sample set, generate a mutation fingerprint by identifying highly variable sites in the mitochondrial control region, and obtain a targeted mutation coordinate set; The second encryption unit is used to perform biological mirroring of the primer region according to the target mutation coordinate set, generate a variant primer set by mapping the high-frequency mutation sites to homologous positions in the standard primer region, and obtain biological feature binding primers; The third encryption unit is used to perform restricted amplification encryption processing based on the biometric binding primers, implant targeted mutations through annealing temperature offset control to generate function-restricted barcodes, obtain an encrypted barcode set and store it in the anti-parsing database framework.
9. The rare and endangered species DNA barcode encryption storage and inference system according to claim 6, characterized in that: The generation module includes: The first generation unit is used to pre-judge the species type based on the DNA sample to be identified, determine the potential species classification by screening the base composition of the mitochondrial conserved region, and obtain the predicted species identification; The second generation unit is used to perform risk masking of the primer region according to the predicted species identifier, and to generate a directional masking sequence by dynamically adjusting the masking position in combination with the poaching heat map; The third generating unit is used to perform time-sensitive encryption processing according to the directional masking sequence, and generate a single valid sequence including a self-destruction countdown by binding the sample collection time.
10. The rare and endangered species DNA barcode encryption storage and inference system according to claim 6, characterized in that: The reasoning module includes: A first inference unit is configured to perform primer masking recovery processing based on the single valid sequence and the encrypted barcode set, restore the complete sequence fragment by matching the variant primer framework, and obtain an unmasked DNA sequence; A second inference unit is configured to perform genetic similarity calculation based on the unmasked DNA sequence, and generate a genetic consistency index by comparing the base matching rate of the sequence to be tested with the base matching rate of the core region of the encrypted barcode set; The third reasoning unit is used to perform threshold determination processing according to the genetic matching index, and generate a species identification code by comparing the relationship between the genetic matching index and the species minimum matching threshold.
Citation Information
Patent Citations
Local comparison algorithm-based endangered animal identification system, method and related device
CN120199331A
Cloud-based gene analysis service method and platform
US20200160935A1