New-generation artificial intelligence bioinformatics individual geographical traceability method

By constructing a new generation of artificial intelligence DNA tracing model based on spatial attention mechanism and geographic distance loss function, the multiple adaptation obstacles of DNA tracing in existing technologies have been solved, and accurate prediction and efficient tracing of individual geographic locations have been achieved.

CN120808897AActive Publication Date: 2025-10-17BEIJING FORESTRY UNIVERSITY
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510909816.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-02
Publication Date
2025-10-17
Estimated Expiration
2045-07-02

AI Technical Summary

Technical Problem

In the field of wildlife protection, existing DNA traceability technology faces conflicts between the development of universal cross-species models and the specificity of genetic markers between species, the difficulty in balancing the high-precision requirements of geographic traceability with the applicability of low-quality DNA samples, and the insufficient learning ability of the model for small samples. These problems have led to unmet needs for rapid traceability.

Method used

A new generation of artificial intelligence bioinformatics individual tracing method based on spatial attention mechanism and new geographic distance loss function is adopted to construct an artificial intelligence DNA individual tracing model, and convolutional neural network and multi-layer perceptron are combined with CBAM spatial attention module to predict individual geographic location.

Benefits of technology

It achieves accurate prediction of individual geographic locations, simplifies the dataset integration process, improves the repeatability and accuracy of the model, reduces dependence on the amount of sample data, and adapts to the computational efficiency and model generalization capabilities of different traceability scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120808897A_ABST
    Figure CN120808897A_ABST
Patent Text Reader

Abstract

The invention discloses a new-generation artificial intelligence bioinformatics individual geographical traceability method, and belongs to the technical field of wild animal DNA geographical traceability, and the method comprises the following steps: S1, obtaining DNA of a biological sample; s2, on the basis of a convolutional neural network module, constructing a novel artificial intelligence DNA individual traceability model in combination with a CBAM space attention module and a multi-layer perceptron; s3, taking the genotype data of the sample and the background sampling geographic data of the sample as the input of the model, training the created artificial intelligence DNA individual traceability model, adopting cross validation to adjust parameters, and optimizing model parameters; and S4, performing geographic coordinate prediction on the individual DNA of the unknown sample by using the artificial intelligence DNA individual traceability model. The method has the capability of integrating data of different species and batches, the integration process of the data set is greatly simplified, the complexity is remarkably reduced, the repeatability is improved, and the prediction of the accurate geographic position of an individual can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of wild animal DNA geographic tracing, and in particular to a new generation of artificial intelligence bioinformatics individual geographic tracing method. BACKGROUND

[0002] With the rapid development of machine learning and high-throughput sequencing technology, DNA molecular provenance technology has made new progress. DNA molecular provenance technology has undergone several technical iterations, as follows:

[0003] First generation DNA tracing technology (1980s-2000s): haplotype network analysis based on mitochondrial single gene markers (COI / Cytb / D-loop) successfully realized species identification (accuracy > 90%) and large-scale geographic tracing (depending on known population background). China has realized the breakthrough of the judicial identification system from scratch through the introduction of digestion, but is subject to access restrictions of international genetic databases. This technology supports the law enforcement needs of the initial stage of global wild animal judicial identification, but cannot meet the fine tracing challenges brought by current network smuggling.

[0004] Second generation DNA tracing technology (2000s-2010s): population genetic structure analysis represented by SNP chip and STRUCTURE model improves the tracing resolution to the subspecies population level. This technology system supports the application of conservation genetics in tracing, but has three major limitations: ① dependent on prior population division; ② computational complexity increases exponentially with the number of sites; ③ unable to analyze the genetic geographic gradient of continuously distributed species.

[0005] Third generation DNA tracing technology (2020s-): deep learning-based whole genome spatial decoding technology breaks through the linear dimensionality reduction limitation and achieves <100km positioning accuracy in human population tracing. However, it faces three adaptation barriers in the field of wildlife conservation: ① 90% of the samples involved are low-quality degraded DNA (genotype missing rate > 35%); ② 85% of CITES appendix species lack reference genomes; ③ cross-border data sharing barriers result in insufficient training samples (species with n<20 account for 67%). These technical bottlenecks seriously restrict the rapid tracing needs of the involved items.

[0006] The existing DNA tracing technology has the following "three difficult balances" problems, namely: (1) there is a conflict between the development of cross-species general models and the specificity requirements of genetic markers between species; (2) the high accuracy requirement of geographic tracing (individual positioning) and the applicability of low-quality DNA samples are difficult to balance; (3) the insufficient learning ability of small samples of the model restricts the application of endangered species tracing. Therefore, it is urgent to develop a new generation of DNA tracing method to solve the above problems. SUMMARY

[0007] The purpose of the present application is to provide a new generation of artificial intelligence bioinformatics individual geographic tracing method, based on a spatial attention mechanism and a new geographic distance loss function, to construct a new generation of artificial intelligence DNA individual tracing model, thereby realizing the prediction of the precise geographic location of individuals.

[0008] To achieve the above-mentioned purpose, the present application provides a new generation of artificial intelligence bioinformatics individual geographic tracing method, comprising the following steps:

[0009] S1, obtaining the DNA of a biological sample;

[0010] S2, based on a convolutional neural network module, combining a CBAM spatial attention module and a multi-layer perceptron to construct a new generation of artificial intelligence DNA individual tracing model;

[0011] S3, using the genotype data of the sample and the background sampling geographic data of the sample as the input of the model, training the created artificial intelligence DNA individual tracing model, and optimizing the model parameters by cross-validation parameter tuning;

[0012] S4, using the artificial intelligence DNA individual tracing model to predict the geographic coordinates of the individual DNA of an unknown sample.

[0013] Preferably, in S1, the microsatellite of the sample is obtained by species-specific genetic markers and PCR simple and rapid amplification technology, and if the species is not clear, the whole genome resequencing or simplified genome data of the sample is quickly obtained by high-throughput sequencing.

[0014] Preferably, in S2, the input of the CBAM spatial attention module in the artificial intelligence DNA individual tracing model is the re-encoded genotype feature matrix reshaped as Specifically:

[0015] Spatial descriptor generation:

[0016] Max-pooling and average-pooling along the channel dimension:

[0017] F max = MaxPool(X);

[0018] F avg = AvgPool(X);

[0019] Wherein,

[0020] Concatenate the pooling results:

[0021]

[0022] Spatial attention weight calculation:

[0023] Extract spatial dependencies by convolutional layer:

[0024]

[0025] where, Conv 3×3 represents a 3x3 convolution kernel, and σ is a Sigmoid activation function;

[0026] Feature enhancement:

[0027] Apply attention weight to the original feature matrix:

[0028] X attn = X ⊙ A spatial ;

[0029] where, ⊙ represents element-wise multiplication, and enhances spatial correlation features;

[0030] The input of the convolutional neural network module in the artificial intelligence DNA individual tracing model is the spatial attention enhanced feature matrix

[0031] The network structure of the convolutional neural network module is as follows:

[0032] Convolutional layer 1:

[0033]

[0034] Use 64 3x3 convolution kernels, step = 1, and the activation function is ReLU;

[0035] Pooling layer 1:

[0036]

[0037] Convolutional layer 2:

[0038]

[0039] Pooling layer 2:

[0040]

[0041] Feature flattening:

[0042]

[0043] The input of the multi-layer perceptron in the artificial intelligence DNA individual tracing model is the flattened feature vector

[0044] The network structure of the multi-layer perceptron is as follows:

[0045] Fully connected layer 1:

[0046]

[0047] wherein,

[0048] Fully connected layer 2:

[0049]

[0050] wherein, The output is the predicted latitude and longitude coordinates

[0051] Preferably, in the artificial intelligence DNA individual tracing model training process, if the sample background information contains the geographic coordinates of the individual's source place, the artificial intelligence DNA individual tracing model integrates the geographic information as a reference; if the sample does not have the geographic information of the sampling place, the artificial intelligence DNA individual tracing model predicts the geographic coordinates of the sample by analyzing the population genetic characteristic value and converting the relationship between the genetic characteristic value and the geographic information.

[0052] Preferably, in S3, the following steps are specifically included:

[0053] S31, data augmentation based on random mask upsampling;

[0054] S32, re-encoding the genotype data;

[0055] S33, adaptive adjustment of the CBAM spatial attention module, retaining its spatial attention mechanism and omitting the channel attention part;

[0056] S34, convolution regression processing;

[0057] S35, geographic spatial dimension mapping in the artificial intelligence DNA individual tracing model.

[0058] Preferably, in S31, a given original data set D:

[0059]

[0060] wherein, x i represents the genotype data of sample i, y i represents the geographic coordinates, and N is the number of original samples;

[0061] A new sample is generated by the following steps:

[0062] First, a sample (x i , y i ) is randomly selected from D to generate a new coordinate yi', which is within a 50-kilometer range of the original coordinate y i , ensuring that:

[0063] Haversine(y i ,y′ i )≤50km;

[0064] where Haversine denotes the spherical distance function:

[0065]

[0066] Δλ=λ2-λ1;

[0067] Δφ=φ2-φ1;

[0068] where r is the earth radius, φ1, φ2 are the latitudes of the original point and the perturbed point, and λ1 and λ2 are the longitudes of the original point and the perturbed point, respectively;

[0069] At the same time, a random mask is generated for the SNP data, and the SNP genotype x i is generated after random masking x i , where each SNP site is independently masked with a probability of 20%:

[0070]

[0071] where x ij represents the jth SNP site of sample i;

[0072] The final augmented dataset is defined as:

[0073]

[0074] where M is the number of newly generated samples.

[0075] Preferably, in S32, different machine learning algorithms, including multilayer perceptron, variational autoencoder and generative adversarial network, are used to compress or expand the SNP data into a fixed 1024-dimensional representation; given an input SNP matrix where N represents the number of samples, d represents the original feature dimension, and a conversion function f θ is defined by the neural network parameters Each sample is mapped to a standardized 1024-dimensional vector:

[0076] z i =f θ (x i );

[0077]

[0078] To improve the balance between computational efficiency and model generalization ability in different traceability scenarios, the model contains a two-layer fully connected coding layer neural network architecture, which is represented as follows:

[0079] z i = σ(Wx i + b);

[0080] wherein, W is a weight matrix, b is a bias term, and σ(·) represents a nonlinear activation function.

[0081] Preferably, in S33, based on the given reconstructed 32x32 feature matrix max-pooling and average-pooling are performed along the channel dimension to obtain two spatial descriptors:

[0082] X max = MaxPool(X);

[0083] X avg = AvgPool(X);

[0084] The pooled feature maps are concatenated in the channel dimension to obtain:

[0085] X cat = [X max , X avg ];

[0086] Convolutional layer processing is performed to extract spatial dependencies to obtain:

[0087] X conv = σ(Conv(X cat ));

[0088] wherein σ(·) represents a sigmoid activation function, ensuring that the spatial attention weights are normalized in the range [0, 1];

[0089] Finally, the learned spatial attention weights are applied to the original feature maps through element-level multiplication:

[0090] X' = X ⊙ X conv .

[0091] Preferably, in S34, a distance-based loss function is defined The calculation formula is:

[0092]

[0093] wherein, d represents the distance between the true coordinates and the predicted coordinates;

[0094] Given a dataset containing m samples, let and are the true and predicted latitude and longitude coordinates of the i-th sample, respectively, and the Euclidean distance between them is:

[0095]

[0096] The distance-weighted MSE loss is defined as:

[0097]

[0098] where w i is the weight assigned to each sample according to the Euclidean distance d i , which is defined as:

[0099]

[0100] Preferably, in S35, the goodness of fit of the regression model is measured by the coefficient of determination R 2 :

[0101]

[0102] where d i and are the actual and predicted geographic coordinates, is the mean of the actual coordinates;

[0103] The average distance error is calculated as:

[0104]

[0105] where Haversine refers to the spherical distance between the actual and predicted coordinates;

[0106] A spatial confidence-based criterion is introduced, where a prediction is considered accurate if the predicted coordinate falls within a confidence radius r' of the true location. For a given confidence radius r', the accuracy is defined as:

[0107]

[0108] where is an indicator function that is 1 if the predicted coordinate falls within the confidence radius r' and 0 otherwise.

[0109] Therefore, the present application adopts the above-mentioned new generation of artificial intelligence bioinformatics individual geographic tracing method, based on the spatial attention mechanism and the new geographic distance loss function, a new generation of artificial intelligence DNA individual tracing model is constructed, so that the prediction of the accurate geographic location of the individual can be realized.

[0110] The technical solutions of the present application will be further described in detail through the accompanying drawings and examples. BRIEF DESCRIPTION OF THE DRAWINGS

[0111] Figure 1 This is a flow chart of an embodiment of the new generation artificial intelligence bioinformatics individual geographic tracing method of the present invention;

[0112] Figure 2 This is a schematic diagram of the artificial intelligence DNA individual provenance model framework and training process of an embodiment of the new generation artificial intelligence bioinformatics individual geographic provenance method of the present invention;

[0113] Figure 3 2 is a schematic diagram of the test structure of an embodiment of the present invention; wherein, a is the accuracy of predicting the destination within 200 km of the actual location; b is the accuracy of predicting the destination within 50 km of the actual location. DETAILED DESCRIPTION

[0114] The technical solution of the present invention is further described below with reference to the accompanying drawings and embodiments.

[0115] Example 1

[0116] like Figure 1 As shown, the present invention provides a new generation of artificial intelligence bioinformatics individual geographic tracing method, comprising the following steps:

[0117] S1. Obtain DNA from a biological sample.

[0118] Microsatellites of samples are obtained through species-specific genetic markers and simple and rapid PCR amplification technology. If the species is unclear, whole-genome resequencing or simplified genomic data of the sample can be quickly obtained through high-throughput sequencing.

[0119] S2. Based on the convolutional neural network module, combined with the CBAM spatial attention module and multi-layer perceptron, a new generation of artificial intelligence DNA individual tracing model is constructed.

[0120] The core of the artificial intelligence DNA individual tracing model consists of the CBAM spatial attention module, the convolutional neural network module (2 convolution 2 pooling layers), and the multi-layer perceptron ( Figure 2 a). Figure 2 As shown, Figure 2 a in it is the model framework; Figure 2 The b in is the spatial attention mechanism; Figure 2 Where ce is the workflow of the model, c represents the integration of heterogeneous datasets of different species and geographical locations, d represents the data being standardized separately before training and restored after training, and e represents the pre-training loss rate.

[0121] The input of the CBAM spatial attention module in the artificial intelligence DNA individual tracing model is the re-encoded genotype feature matrix Reshaping to Specifically:

[0122] Spatial descriptor generation:

[0123] Max-pooling and average-pooling along the channel dimension:

[0124] F max = MaxPool(X)

[0125] F avg = AvgPool(X)

[0126] where,

[0127] Concatenate the pooling results:

[0128]

[0129] Spatial attention weight calculation:

[0130] Extract spatial dependencies through convolutional layers:

[0131]

[0132] where, Conv 3×3 represents a 3x3 convolutional kernel, and σ is the Sigmoid activation function.

[0133] Feature enhancement:

[0134] Apply attention weights to the original feature matrix:

[0135] X attn = X ⊙ A spatial

[0136] where, ⊙ represents element-wise multiplication, enhancing spatially related features.

[0137] The input of the convolutional neural network module in the artificial intelligence DNA individual tracing model is the spatial attention-enhanced feature matrix

[0138] The network structure of the convolutional neural network module is as follows:

[0139] Convolutional layer 1:

[0140]

[0141] Use 64 3x3 convolutional kernels with a step size of 1 and a ReLU activation function.

[0142] Pooling layer 1:

[0143]

[0144] Convolutional layer 2:

[0145]

[0146] Pooling layer 2:

[0147]

[0148] Feature flattening:

[0149]

[0150] The input of the multi-layer perceptron in the artificial intelligence DNA individual tracing model is the flattened feature vector

[0151] The network structure of the multi-layer perceptron is as follows:

[0152] Fully connected layer 1:

[0153]

[0154] Wherein,

[0155] Fully connected layer 2:

[0156]

[0157] Wherein, The output is the predicted latitude and longitude coordinates

[0158] The artificial intelligence DNA individual tracing model dynamically weights the spatial signal contribution of multi-modal data through the attention mechanism, enabling the model to integrate data of different species and batches, greatly simplifying the integration process of the data set, and realizing the geographical information fusion of multi-modal genetic data. Significantly reduce the complexity and improve the repeatability, at the same time through the data enhancement, reduce the sample data amount to the influence of model performance.

[0159] S3, training the constructed artificial intelligence DNA individual tracing model.

[0160] During the training of the artificial intelligence DNA individual tracing model, if the sample background information contains the geographical coordinates of the individual source place, the artificial intelligence DNA individual tracing model integrates the geographical information as a reference to optimize and improve the model performance. If the sample does not have the geographical information of the sampling place, the artificial intelligence DNA individual tracing model analyzes the population genetic characteristic value, and converts the relationship between the genetic characteristic value and the geographical information to predict the geographical coordinates of the sample.

[0161] The training process of the AI ​​DNA individual traceability model specifically includes the following steps:

[0162] S31. Data enhancement based on random mask upsampling.

[0163] Samples of endangered wild animals are extremely limited in quantity and quality. In order to improve the accuracy and reliability of traceability, this embodiment adopts an upsampling strategy based on random masking to intelligently increase the size of the dataset while maintaining the integrity of its genetic data and background information.

[0164] Given an original dataset D:

[0165]

[0166] Among them, x i represents the genotype data of sample i (taking SNP as an example), y i represents the geographical coordinates (latitude and longitude), and N is the number of original samples.

[0167] Generate a new sample by following these steps:

[0168] First, a sample (x i ,y i ), generate a new coordinate yi', which is located at the original coordinate y i Within 50 km of the

[0169] Haversine(y i ,y′ i )≤50km

[0170] Among them, Haversine represents the spherical distance function.

[0171]

[0172] Δλ=λ2-λ1

[0173] Δφ=φ2-φ1

[0174] Where r is the radius of the Earth, φ1 and φ2 are the latitudes of the original and perturbation points (in radians), and λ1 and λ2 are the longitudes of the original and perturbation points, respectively.

[0175] At the same time, a random mask is generated for the SNP data, SNP genotype x i Generate x through random mask i ', where each SNP site is masked independently with a probability of 20%:

[0176]

[0177] where x ij denotes the jth SNP locus of sample i.

[0178] The final augmented dataset is defined as:

[0179]

[0180] where M is the number of newly generated samples.

[0181] To ensure the validity of the geographical information of samples, the perturbation follows a uniform distribution within a specified 50-kilometer range, avoiding areas where there are no known real samples. The random masking process preserves the integrity of the genetic structure while simulating the pattern of missing data in the real world.

[0182] S32, re-encode the genotype data.

[0183] Due to the limitations of sequencing methods, datasets obtained from different batches of the same species often exhibit different dimensions. Even if the same sequencing technology is used, differences in genome size between different species can lead to inconsistencies in dataset dimensions. These differences hinder the analysis of geographic genetic information in DNA markers across species. To address this issue, a re-encoding strategy is applied in this embodiment to standardize genotype data to a unified dimensional space.

[0184] A number of different machine learning algorithms are used, including multi-layer perceptron (MLP), variational autoencoder (VAE), and generative adversarial network (GAN), to compress or expand SNP data to a fixed 1024-dimensional representation.

[0185] Given an input SNP matrix where N represents the number of samples and d represents the original feature dimension (which varies in different datasets), a conversion function f θ is defined, which is parameterized by a neural network maps each sample to a standardized 1024-dimensional vector:

[0186] z i = f θ (x i )

[0187]

[0188] To improve the balance between computational efficiency and model generalization ability in different traceability scenarios, the model includes a two-layer fully connected encoding layer neural network architecture, represented as follows:

[0189] z i = σ(Wx i +b)

[0190] where, is the weight matrix, is the bias term, and σ(·) denotes the nonlinear activation function. This re-encoding framework can effectively achieve dimension standardization while preserving genetic variations essential for traceability analysis.

[0191] S33, adaptively adjust the CBAM spatial attention module, retain its spatial attention mechanism, and omit the channel attention part.

[0192] To enhance the genetic feature representation of re-encoded genotype data, the CBAM spatial attention module is adaptively adjusted, retaining its spatial attention mechanism and omitting the channel attention part. The re-encoded 1×1024 vector is reshaped into a 32×32 matrix, and spatial dependency is extracted through convolution operations.

[0193] To enhance spatial feature representation, a spatial attention mechanism is introduced during feature extraction. The spatial attention module generates an attention weight matrix through max-pooling, average-pooling, and convolution operations, thereby optimizing the feature map.

[0194] Based on the given reconstructed 32×32 feature matrix Max-pooling and average-pooling are performed along the channel dimension to obtain two spatial descriptors:

[0195] X max = MaxPool(X)

[0196] X avg = AvgPool(X)

[0197] The pooled feature maps are concatenated along the channel dimension to obtain:

[0198] X cat = [X max , X avg ]

[0199] Convolution layer processing to extract spatial dependencies results in:

[0200] X conv = σ(Conv(X cat ))

[0201] where σ(·) denotes the sigmoid activation function, ensuring that the spatial attention weights are normalized within the range [0, 1].

[0202] Finally, the learned spatial attention weights are applied to the original feature map through element-level multiplication:

[0203] X' = X ⊙ X conv

[0204] This spatial attention mechanism dynamically enhances spatially relevant features while suppressing less informative regions, improving the feature extraction capability of downstream tasks.

[0205] S34, a convolutional regression process is performed.

[0206] The input of the convolutional regression is a 32x32 feature matrix with spatial attention weights, which is then processed through a series of convolutional layers and fully connected layers.

[0207] In regression tasks, mean squared error (MSE) is one of the most widely used loss functions. Although MSE provides a global loss measure, it ignores local variations. This often leads the model to predict clusters of points around a few central locations to minimize the overall loss, but reduces the accuracy of individual sample predictions.

[0208] To address this issue, a distance-based loss function is defined in this embodiment This loss function does not treat all samples equally, but introduces weights based on the Euclidean distance between predicted and actual positions. This adjustment encourages the model to minimize errors for distant points, thereby improving the accuracy of geographic predictions.

[0209] Distance-based loss function The calculation formula is:

[0210]

[0211] where d represents the distance between the true and predicted coordinates. The parameters of this model are optimized using the Adam optimizer through backpropagation.

[0212] Given a dataset containing m samples, let and be the true and predicted latitude and longitude coordinates of the i-th sample, respectively, and their Euclidean distance is:

[0213]

[0214] Define the distance-weighted MSE loss as:

[0215]

[0216] where w i is the weight assigned to each sample based on the Euclidean distance d i , and the weight is defined as:

[0217]

[0218] This weighting scheme ensures that samples with larger errors contribute more to the total loss, allowing the model to improve its performance on special samples.

[0219] S35, geographical space dimension mapping is performed in the artificial intelligence DNA individual traceability model.

[0220] The numerical values of latitude and longitude exhibit relative sparsity, which may hinder model convergence if directly used as regression targets. Previous studies directly trained neural networks on raw geographic coordinates, which performed poorly in practical applications. In contrast, the model introduces an optimization strategy that normalizes coordinates during feature encoding while preserving the basic statistical properties of each dataset. By normalizing geographic data during training and applying an inverse transformation after inference, the model can significantly accelerate model convergence and improve regression accuracy.

[0221] To comprehensively evaluate the performance of different traceability models, in addition to the standard regression R 2 score, a new evaluation metric is introduced in this embodiment. Previous studies mainly relied on mean distance error (MDE) for evaluation. In this embodiment, a prediction accuracy metric based on spatial proximity is further incorporated.

[0222] The coefficient of determination R 2 measures the goodness of fit of a regression model:

[0223]

[0224] where d i and are the actual and predicted geographic coordinates, is the mean of the actual coordinates;

[0225] The mean distance error is calculated as:

[0226]

[0227] where Haversine refers to the spherical distance between the actual and predicted coordinates;

[0228] A spatial confidence-based criterion is introduced, where a prediction is considered accurate if the predicted coordinate falls within the confidence radius r' of the true location. For a given confidence radius r', accuracy is defined as:

[0229]

[0230] where is an indicator function that is 1 if the predicted coordinate falls within the confidence radius r' and 0 otherwise.

[0231] In this example, three confidence radii are set to evaluate the model's prediction accuracy: 100 km, 200 km, and 500 km. By incorporating this additional metric, the model's geographic provenance performance becomes more interpretable and practical.

[0232] S4. Use the artificial intelligence DNA individual tracing model to predict the geographic coordinates of individual DNA of unknown samples.

[0233] When the sample's genotype data includes geographic coordinates, these coordinates are also used as input for the AI ​​DNA individual traceability model. Genotype data is derived from resequencing of biological samples, simplified genomes, and screening for genetic variations such as microsatellites, while geographic coordinates correspond to the latitude and longitude at the time of sample collection.

[0234] In practical applications, the model can be fine-tuned based on existing sample data to adapt it to different application scenarios. The model seamlessly combines non-parametric (predicting coordinates based solely on genotype data) and parametric (using genotype data and some geographic coordinates for fine-tuning before prediction), allowing users to choose the appropriate mode based on their needs.

[0235] In this embodiment, tests were conducted on Chinese pangolins, Asian locusts, and musk deer. The test data are shown in Table 1:

[0236] Table 1 Test dataset

[0237]

[0238] The final test results are as follows Figure 3 As shown, it can be seen that the prediction result of this embodiment is relatively accurate and reliable, and has certain reference value.

[0239] Therefore, the present invention adopts the above-mentioned new generation artificial intelligence bioinformatics individual geographic tracing method, and constructs a new generation artificial intelligence DNA individual tracing model based on the spatial attention mechanism and the new geographic distance loss function, so as to realize the prediction of the individual's precise geographic location.

[0240] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit the same. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that they can still modify or replace the technical solutions of the present invention with equivalents, and these modifications or equivalent replacements cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A new generation of artificial intelligence bioinformatics individual geographic tracing method, characterized by: The following steps are involved: S1. Obtain DNA from biological samples; S2. Based on the convolutional neural network module, combined with the CBAM spatial attention module and multi-layer perceptron, a new generation of artificial intelligence DNA individual tracing model is constructed; S3. Using the sample's genotype data and the sample's background sampling geographic data as the model input, the created artificial intelligence DNA individual traceability model is trained, and cross-validation is used to adjust the parameters and optimize the model parameters; S4. Use the artificial intelligence DNA individual tracing model to predict the geographic coordinates of individual DNA of unknown samples.

2. The next-generation artificial intelligence bioinformatics individual geographic tracing method according to claim 1 is characterized by: In S1, the microsatellites of the sample are obtained through species-specific genetic markers and simple and rapid PCR amplification technology. If the species is unclear, the whole genome resequencing or simplified genome data of the sample can be quickly obtained through high-throughput sequencing.

3. The next-generation artificial intelligence bioinformatics individual geographic tracing method according to claim 1 is characterized by: In S2, the input of the CBAM spatial attention module in the artificial intelligence DNA individual tracing model is the re-encoded genotype feature matrix Reshape into Specifically: Spatial descriptor generation: Perform maximum pooling and average pooling along the channel dimension: F max =MaxPool(X); F avg =AvgPool(X); Among them, F max , Splicing pooling results: Spatial attention weight calculation: Extract spatial dependencies through convolutional layers: Among them, Conv 3×3 represents a 3×3 convolution kernel, and σ is the Sigmoid activation function; Feature Enhancement: Apply attention weights to the original feature matrix: X attn =X⊙A spatial ; Among them, ⊙ represents element-by-element multiplication to enhance spatial correlation features; The input of the convolutional neural network module in the artificial intelligence DNA individual tracing model is the feature matrix after spatial attention enhancement. The network structure of the convolutional neural network module is as follows: Convolutional layer 1: Use 64 3×3 convolution kernels, stride = 1, and ReLU activation function; Pooling layer 1: Convolutional layer 2: Pooling layer 2: Feature Flattening: The input of the multilayer perceptron in the artificial intelligence DNA individual tracing model is the flattened feature vector The network structure of the multilayer perceptron is as follows: Fully connected layer 1: in, Fully connected layer 2: in, The output is the predicted latitude and longitude coordinates 4. The next-generation artificial intelligence bioinformatics individual geographic tracing method according to claim 1 is characterized by: In S3, during the training process of the artificial intelligence DNA individual traceability model, if the sample background information contains the geographical coordinates of the individual's place of origin, the artificial intelligence DNA individual traceability model integrates the geographical information as a reference; if the sample does not have the geographical information of the sampling location, the artificial intelligence DNA individual traceability model analyzes the genetic characteristic values ​​of the population and predicts the geographical coordinates of the sample by converting and rotating the relationship between the genetic characteristic values ​​and the geographical information.

5. The next-generation artificial intelligence bioinformatics individual geographic tracing method according to claim 1 is characterized by: In S3, the following steps are specifically included: S31, data enhancement based on random mask upsampling; S32, recode genotype data; S33, adaptively adjust the CBAM spatial attention module, retain its spatial attention mechanism, and omit the channel attention part; S34, performing convolution regression processing; S35. Map geographic spatial dimensions in the artificial intelligence DNA individual tracing model.

6. The next-generation artificial intelligence bioinformatics individual geographic tracing method according to claim 5 is characterized by: In S31, given an original data set D: Among them, x i represents the genotype data of sample i, y i represents the geographical coordinates, N is the number of original samples; Generate a new sample by following these steps: First, a sample (x i ,y i ), generate a new coordinate y i ', the coordinate is located at the original coordinate y i Within 50 km of the Haversine(y i ,the' i )≤50 km: Among them, Haversine represents the spherical distance function: Δλ=λ2-λ1; Δφ=φ2-φ1; Where r is the radius of the earth, φ1 and φ2 are the latitudes of the original point and the disturbance point, respectively, and λ1 and λ2 are the longitudes of the original point and the disturbance point, respectively; At the same time, a random mask is generated for the SNP data, SNP genotype x i Generate x through random mask i ', where each SNP site is masked independently with a probability of 20%: Among them, x ij represents the j-th SNP site of sample i; The final augmented dataset is defined as: Where M is the number of newly generated samples.

7. The next-generation artificial intelligence bioinformatics individual geographic tracing method according to claim 5 is characterized by: In S32, different machine learning algorithms, including multilayer perceptron, variational autoencoder and generative adversarial network, are used to organize SNP data into a fixed 1024-dimensional representation. Given an input SNP matrix Where N represents the number of samples, d represents the original feature dimension, and a transformation function f parameterized by a neural network is defined. θ , each sample Mapped to a normalized 1024-dimensional vector: z i =f θ (x i ); To balance computational efficiency and model generalization in different traceability scenarios, the model includes a two-layer fully connected encoding layer neural network architecture, as shown below: With i =σ(Wx i +b); in, is the weight matrix, is the bias term, and σ(·) represents the nonlinear activation function.

8. The next-generation artificial intelligence bioinformatics individual geographic tracing method according to claim 5 is characterized by: In S33, based on the given reconstructed 32×32 feature matrix Perform maximum pooling and average pooling along the channel dimension to obtain two spatial descriptors: X max =MaxPool(X); X avg =AvgPool(X); The pooled feature maps are spliced ​​in the channel dimension to obtain: X cat =[X max ,X avg ]; Convolutional layers are used to extract spatial dependencies and obtain: X conv =σ(Conv(X cat )); Where σ(·) represents the sigmoid activation function, which ensures that the spatial attention weights are normalized in the range of [0, 1]; Finally, the learned spatial attention weights are applied to the original feature map via element-wise multiplication: X'=X⊙X conv 。 9. The next-generation artificial intelligence bioinformatics individual geographic tracing method according to claim 5 is characterized by: In S34, customize a distance-based loss function The calculation formula is: in, Represents the distance between the true coordinates and the predicted coordinates; Given a dataset containing m samples, let and are the true and predicted latitude and longitude coordinates of the i-th sample, respectively. The Euclidean distance between them is: Define the distance weighted MSE loss as: Among them, w i It is based on the Euclidean distance d i The weight assigned to each sample is defined as:

10. The next-generation artificial intelligence bioinformatics individual geographic tracing method according to claim 5 is characterized by: In S35, the coefficient of determination R 2 Measure the goodness of fit of the regression model: Among them, d i and are the actual and predicted geographic coordinates, is the mean of the actual coordinates; The average distance error is calculated as: Among them, Haversine refers to the spherical distance between the actual coordinates and the predicted coordinates; A spatial confidence-based criterion is introduced. When the predicted coordinates fall within the confidence radius r' of the true position, the prediction is considered accurate. For a given confidence radius r', the accuracy is defined as: in, is an indicator function that equals 1 if the predicted coordinate falls within the confidence radius r' and 0 otherwise.

Citation Information

Patent Citations

  • Single nucleotide polymorphism (SNP) marker combination for geographical traceability identification of elaphurus davidianus in Beijing and Jiangsu and identification method of SNP marker combination

    CN112458180A

  • Pedigree tracing method based on whole genome re-sequencing SNP big data and deep learning

    CN118248210A

  • Geographic tracing method for endangered wild animals and products thereof based on machine learning and DNA imprinting and application thereof

    CN119170102A

  • Artificial intelligence and high-throughput sequencing-based accurate geographical traceability method for pangolin scales and nail tips thereof and application of accurate geographical traceability method

    CN119229968A

  • Water pollutant tracing method based on environmental DNA and machine learning

    CN119763660A