Application of local ancestor inference and multi-gene risk scoring for predicting complex disease risk of mixed lineage individuals
By calculating single ancestor and effect size-weighted multigene risk scores, the problem of decreased performance of the multigene risk scoring model in non-European and newly mixed ancestry individuals was solved, achieving more accurate disease risk prediction.
Patent Information
- Application Number
- CN202380075617.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-13
- Filing Date
- 2023-10-12
- Publication Date
- 2025-06-06
AI Technical Summary
The performance of polygenic risk scoring models in non-European and newly mixed ancestry individuals was reduced mainly due to the underrepresentation of these individuals in the publicly available training cohort.
A method was adopted to improve performance in mixed-born individuals by calculating single ancestor and effect size-weighted polygenic risk scores (PRSs). This method utilizes multiple PRS scores that demonstrated optimal performance for a given ancestry, their effect size in non-mixed ancestry ancestry individuals, and local ancestral decomposition.
Through this method, the performance of the multigene risk scoring model in mixed-course individuals can be improved, and individuals with increased risk of disease can be identified, and more accurate disease risk prediction can be provided.
Smart Images

Figure CN120113002A_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims the benefit of U.S. Provisional Application No. 63 / 379,395, filed on October 13, 2022, which is incorporated herein by reference in its entirety. Technical Field
[0003] The present disclosure relates generally to determining disease risk, and more particularly to methods for determining the risk of disease occurrence in individuals of mixed ancestry Background Art
[0004] The present disclosure relates generally to determining disease risk, and more particularly to methods for determining the risk of disease occurrence in individuals of mixed ancestry Summary of the invention
[0005] Polygenic risk scores (PRS) have been used to successfully predict complex phenotypes such as coronary artery disease (CAD) or breast cancer (BC). However, the main limitation of polygenic risk scores is the reduced performance in non-European and recently admixed individuals, which stems from the underrepresentation of non-European individuals in publicly available training cohorts.
[0006] The proposed method / workflow aims to improve the performance of PRS models in individuals of recent admixture.
[0007] The method uses multiple PRS scores that demonstrate the best performance for a given ancestry, their effect sizes in individuals of non-admixed ancestry, and local ancestry decomposition to calculate a single ancestry and effect size weighted PRS score. The obtained composite PRS score can be used as a feature / predictor for a downstream classification model that is used to identify individuals with elevated disease risk.
[0008] enter:
[0009] -Query sample (phased) VCF file
[0010] - Known ancestor reference (phased) VCF file
[0011] -PRS model weights used to query sample scores
[0012] - Effect sizes of PRS models estimated in individuals of non-admixed ancestry
[0013] Output:
[0014] - A composite PRS score calculated using one of the following methods:
[0015] - Sum of partial PRS model scores weighted by the global ancestry score
[0016] - Sum of partial PRS model scores weighted by the global ancestry score and the PRS effect size estimated in individuals of non-admixed ancestry
[0017] - Sum of partial PRS model scores weighted by the global ancestry score and the estimated partial PRS effect size in individuals of non-admixed ancestry
[0018] Compared to existing methods that use local ancestry deconvolution for PRS, our approach includes additional weighting of partial model scores by effect sizes of full or partial PRS models estimated in an independent non-admixed ancestry training cohort, whereas existing methods weight partial scores only by the estimated ancestry fraction and an additional scaling factor from other previously used methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Having generally described certain example embodiments above, reference will now be made to the accompanying drawings, which are not necessarily drawn to scale. Some embodiments may include fewer or more components than shown in the figures.
[0020] Figure 1 A schematic block diagram of an example method for calculating partial ancestry specific PRS scores and their coefficients using 2-source admixture ancestry as an example according to some example embodiments described herein is shown.
[0021] Figure 2 The performance of the method according to some example embodiments described herein on a cohort of mixed-race individuals of Latino or Hispanic descent is shown.
[0022] Figure 3 Schematic block diagrams showing example circuitry implementing apparatus that may perform various operations according to some example embodiments described herein. DETAILED DESCRIPTION
[0023] Some example embodiments will now be described more fully hereinafter with reference to the accompanying drawings, in which some, but not necessarily all, embodiments are shown. Because the invention described herein can be embodied in many different forms, the invention should not be limited to the embodiments set forth herein; rather, these embodiments are provided so that this disclosure will satisfy applicable legal requirements.
[0024] Definition of Certain Terms
[0025] Unless otherwise defined, technical and scientific terms used herein have the meaning commonly understood by one of ordinary skill in the art to which the invention belongs. Unless otherwise stated, the materials cited in the following description and examples are all available from commercial sources.
[0026] The terms "computer-readable medium" and "memory" refer to non-transitory storage hardware, non-transitory storage devices, or non-transitory computer system memory that can store computer-executable instructions or software programs that can be accessed by a controller, a microcontroller, a computing system, or a module of a computing system. The non-transitory computer-readable medium can be accessed by a computing system or a module of a computing system to retrieve and / or execute the computer-executable instructions or software programs stored on the medium. Exemplary non-transitory computer-readable media can include, but are not limited to, one or more types of hardware memory, non-transitory tangible media (e.g., one or more magnetic storage disks, one or more optical disks, one or more USB flash drives), computer system memory, or random access memory (such as DRAM, SRAM, EDO RAM), etc.
[0027] The term "computing device" may refer to any computer implemented in hardware, software, firmware, and / or any combination thereof. Non-limiting examples of computing devices include personal computers, servers, laptops, mobile devices, smartphones, fixed terminals, personal digital assistants ("PDAs"), kiosks, custom hardware devices, wearable devices, smart home devices, Internet of Things ("IoT") enabled devices, and network-linked computing devices.
[0028] Example Implementation Device
[0029] Figure 3 A device 300 is shown that may include an example system that may implement the example embodiments described herein. The device may include a processor 302, a memory 304, a communication circuit system 306, and an input-output circuit system 308, each of which will be described in more detail below, and Figure 3 any number of additional hardware components not explicitly represented in the Figure 3 300 is only shown as being connected to the processor 302, but it will be understood that the device 300 may further include a bus ( Figure 3 (not explicitly indicated in the figure) for transmitting information between any combination of various components of the device 300. The device 300 may be configured to perform the various operations described above, as well as the following combined Figure 3 The operations described.
[0030] The processor 302 (and / or a co-processor or any other processor that assists or is otherwise associated with the processor) can communicate with the memory 304 via a bus for transferring information between components of the device. The processor 302 can be implemented in a variety of different ways, and can include, for example, one or more processing devices configured to execute independently. In addition, the processor can include one or more processors configured in series via a bus to enable independent execution of software instructions, pipelines, and / or multithreading. The use of the term "processor" can be understood to include a single-core processor, a multi-core processor, multiple processors of the device 300, a remote or "cloud" processor, or any combination thereof.
[0031] The processor 302 may be configured to execute software instructions stored in the memory 304 or otherwise accessible to the processor (e.g., software instructions stored on a separate storage device). In some cases, the processor may be configured to perform hard-coded functions. Thus, whether configured by hardware or software methods, or by a combination of hardware and software, the processor 302 represents an entity (e.g., physically implemented in a circuit system) that is capable of performing operations according to various embodiments of the present invention when configured accordingly. Alternatively, as another example, when the processor 302 is implemented as an executor of software instructions, the software instructions may specifically configure the processor 302 to perform the algorithms and / or operations described herein when executing the software instructions.
[0032] The memory 304 is non-transitory and may include, for example, one or more volatile and / or non-volatile memories. In other words, for example, the memory 304 may be an electronic storage device (e.g., a computer-readable storage medium). The memory 304 may be configured to store information, data, content, applications, software instructions, etc. to enable the device to perform various functions according to the example embodiments contemplated herein.
[0033] The communication circuit system 306 can be any component, such as a device or circuit system implemented in hardware or a combination of hardware and software, which is configured to receive and / or send data from / to a network and / or any other device, circuit system, or module in communication with the device 300. In this regard, the communication circuit system 306 may include, for example, a network interface for implementing communications with a wired or wireless communication network. For example, the communication circuit system 306 may include one or more network interface cards, antennas, buses, switches, routers, modems, and supporting hardware and / or software, or any other device suitable for implementing communications via a network. In addition, the communication circuit system 306 may include processing circuit systems for sending such signals to the network or for processing the reception of signals received from the network.
[0034] Equipment 300 may include input-output circuit system 308, which is configured to provide output to the user and receive the indication of user input in some embodiments. It will be noted that some embodiments will not include input-output circuit system 308, in which case, user input can be received via a separate device. Input-output circuit system 308 may include a user interface, such as a display, and may further include a component for controlling the use of the user interface, such as a web browser, mobile application, dedicated client device, etc. In some embodiments, input-output circuit system 308 may include a keyboard, a mouse, a touch screen, a touch area, a soft key, a microphone, a loudspeaker and / or other input / output mechanisms. Input-output circuit system 308 may utilize processor 302 to control one or more functions of one or more of these user interface elements by software instructions (e.g., application software and / or system software, such as firmware) stored in a memory (e.g., memory 304) accessible to processor 302.
[0035] In some embodiments, the various components of the device 300 can be remotely hosted (e.g., by one or more cloud servers), and therefore not all components must reside in one physical location. In addition, some of the functions described herein can be provided by third-party circuit systems. For example, the device 300 can access one or more third-party circuit systems via any type of networked connection, which facilitates the transmission of data and electronic information between the device 300 and the third-party circuit systems. In turn, the device 300 can communicate remotely with one or more components that constitute the device 300 described above.
[0036] As will be appreciated based on this disclosure, some example embodiments may take the form of a computer program product that includes software instructions stored on at least one non-transitory computer-readable storage medium (e.g., memory 304). Any suitable non-transitory computer-readable storage medium may be used in such embodiments, some examples of which are non-transitory hard disks, CD-ROMs, flash memory, optical storage devices, and magnetic storage devices. It should be understood that with respect to the computer program product, such as the one provided by the computer program product, the computer program product may be stored on at least one non-transitory computer-readable storage medium. Figure 3 In some embodiments, the software instructions are loaded onto a computing device or apparatus to produce a special-purpose machine that includes components for implementing the various functions described herein.
[0037] Having described the specific components of device 300, example embodiments are described below.
[0038] Example Operation
[0039] Figure 1Depicted is an example method for calculating partial ancestry specific PRS scores and their coefficients using 2-source admixture ancestry as an example. As described above, Figure 1 The steps shown in may be performed by a computing device such as the apparatus 300 described above.
[0040] Step 0. Evaluate the performance of candidate PRS models for each continental ancestry using a non-admixed ancestry training cohort (e.g., UKBB or other cohort with available genotype and phenotype labels), and identify the best performing model for each continental ancestry.
[0041] Step 1. Collect the patient's DNA sample and perform whole genome sequencing WGS, genotyping and phasing. This analysis can be done using long-read sequencing technology (i.e., at least about 5 kb or longer, including about 20 kb or longer read length, and about 100 kb or longer ultra-long read sequencing read length), these services can be obtained through existing suppliers such as Pacific Biosciences, Oxford Nanopore Technologies and Illumina.
[0042] Step 2. Estimate the local ancestry of the patient sample using a reference cohort of samples of known ancestry such as the 1000 Genomes Project and one of the methods described previously.
[0043] After ancestry inference, each marker of the patient sample is labeled with its inferred ancestor, and the haplotypes are divided into regions corresponding to each inferred ancestor.
[0044] Step 3. Score the subject's ancestry-specific regions using the best performing PRS model for a given ancestry (as identified in step 0) to obtain a raw partial PRS score. Simultaneously, score the same segments in a non-admixed ancestry reference cohort (such as the 1000 Genomes Project samples).
[0045] Additionally, in a variation of this approach, the same regions are scored in individuals of non-admixed ancestry from a training cohort for which phenotypic information (e.g., UKBB or other biobank data) is available.
[0046] Step 4. Calculate the mean and standard deviation of the partial PRS scores in the reference cohort and use this mean and standard deviation to center and scale each partial PRS score for the patient. Similarly, center and scale the partial scores for the training cohort using the same mean and standard deviation.
[0047] Step 5. In embodiments of the method utilizing a non-admixed ancestry ancestral training cohort, an additional step is performed to estimate the effect size of the ancestry-specific partial PRS score on the phenotype of interest (partial_β in Equation 3) i ). This was achieved by fitting a linear / logistic regression model for each ancestry with the corresponding partial PRS score as a predictor.
[0048] Alternative methods for estimating effect sizes of ancestry-specific partial PRS scores ( Figure 1 not depicted) is the effect size using the corresponding full PRS score (β in Equation 2 i , calculated using the complete genomes of the training cohort samples). This is also done by fitting a linear / logistic regression.
[0049] Step 6. Calculate the admixed ancestry PRS score for the admixed ancestry sample as a weighted sum of the partial PRS scores using one of the following three equations:
[0050] Equation 1: Composite PRS score with partial scores weighted by the global ancestry score:
[0051]
[0052] Equation 2: The partial score is a composite PRS score weighted by the global ancestry score and the full PRS model effect size estimated in an independent non-admixed ancestry (training) cohort.
[0053]
[0054] Equation 3: The partial score is a composite PRS score weighted by the global ancestry score and the estimated partial PRS model effect size in an independent non-admixed ancestry (training) cohort.
[0055]
[0056] where i indexes the fractionalized ancestral component, partial_score is the centered and scaled partial score calculated as described in step 4, hap1 and hap2 index the query sample haplotypes, and anc_fraction is the global estimate of the given fractionalized ancestry (the fraction of the genome length assigned to this ancestor).
[0057] Figure 2The performance of this method in a cohort of individuals of mixed ancestry of Latino / Hispanic descent is shown. PGS000008 is a single PRS model that does not utilize ancestral inference and is included as a performance baseline. score_gw and score_bw are composite scores calculated according to equations 1 and 2, respectively. The values on the x-axis are the odds ratios (expressed in standard deviation units of control samples) for a logistic regression model using breast cancer as the outcome. The error bars correspond to the standard deviation of 10 replicates of 10-fold cross validation.
[0058] in conclusion
[0059] With the benefit of the foregoing description and the associated drawings, those skilled in the art to which these inventions belong will think of many modifications and other embodiments of these inventions set forth herein. Therefore, it should be understood that these inventions are not limited to the specific embodiments disclosed, and modifications and other embodiments are intended to be included within the scope of the appended claims. In addition, although the foregoing description and the associated drawings describe example embodiments in the context of certain example combinations of elements and / or functions, it should be understood that different combinations of elements and / or functions may be provided by alternative embodiments without departing from the scope of the appended claims. In this regard, for example, combinations of elements and / or functions different from those explicitly described above may also be envisioned, as may be set forth in some of the appended claims. Although specific terms are employed herein, these terms are used only in a general and descriptive sense and not for purposes of limitation.
Claims
1. A method for determining a mixed ancestry polygenic risk score (PRS) for a mixed ancestry subject, the method comprising: include: assigning ancestral labels to one or more phased subject genotype segments; generating one or more ancestry-specific groupings, wherein each ancestry-specific grouping includes the one or more phased subject genotype segments corresponding to a specific ancestry tag; as well as For each ancestor specific grouping: applying a polygenic risk model corresponding to the ancestry signature for the ancestry-specific grouping to each phased subject genotype segment of the ancestry-specific grouping to generate one or more ancestry-specific raw partial PRSs, applying the polygenic risk model to corresponding non-admixed ancestry genotype segments of a non-admixed ancestry reference cohort corresponding to the same ancestry label as the ancestry-specific grouping to generate one or more non-admixed ancestry raw partial PRSs, determining a mean PRS and a standard deviation PRS for the non-admixed ancestry reference cohort based on the one or more non-admixed ancestry raw partial PRSs, normalizing the one or more ancestor-specific raw partial PRSs based on the mean PRS and the standard deviation PRS to generate a normalized partial PRS, and The mixed-ancestry PRS for the mixed-ancestry subject is generated based on a weighted sum of the standardized partial PRSs for each ancestry-specific grouping.
2. The method according to claim 1, in: The phased subject genotype segments are markers or haplotypes, and Assigning the ancestry labels to the one or more phased subject genotype segments further comprises at least one of: assigning the ancestry label to each marker of the phased subject genotype based on a reference cohort of samples of known ancestry; as well as The ancestral label is assigned to each haplotype of the phased subject genotype.
3. The method according to claim 1, in, Normalizing the one or more ancestor-specific origin portions PRS further comprises: Centering the one or more ancestor-specific raw PRSs based on the mean PRS and the standard deviation PRS; and The one or more ancestor-specific raw PRSs are scaled based on the mean PRS and the standard deviation PRS.
4. The method of claim 1, further comprising: include: obtaining a mixed-ancestry genotype from the mixed-ancestry subject; as well as The subject genotype is phased to generate the one or more phased subject genotype segments.
5. The method according to claim 4, in, The phasing of the admixed ancestry genotypes is performed using one or more of a population-based approach or a molecular-based approach.
6. The method of claim 4, further comprising: include: Whole genome sequencing is performed on a biological sample obtained from the mixed-ancestry subject to determine the mixed-ancestry genotype.
7. The method according to claim 1, in, Generating the mixed-ancestry PRS further comprises: determining an ancestry-specific sum for each ancestry-specific grouping based on the corresponding normalized partial PRS and the global ancestry score; and The mixed-ancestry PRS is determined based on each ancestor-specific sum for the one or more ancestor-specific groupings.
8. The method according to claim 1, in, Generating the mixed-ancestry PRS further comprises: determining an ancestry-specific sum for each ancestry-specific grouping based on the corresponding standardized partial PRS, global ancestry score, and full PRS model effect size parameter; and The mixed-ancestry PRS is determined based on each ancestor-specific sum for the one or more ancestor-specific groupings.
9. The method according to claim 1, in, Generating the mixed-ancestry PRS further comprises: determining an ancestry-specific sum for each ancestry-specific grouping based on the corresponding standardized partial PRS, the global ancestry score, and the partial PRS model effect size parameter; and The mixed-ancestry PRS is determined based on each ancestor-specific sum for the one or more ancestor-specific groupings.
10. The method of claim 1, further comprising: include: identifying one or more non-admixed ancestry training sets corresponding to each non-admixed ancestry cohort, wherein each non-admixed ancestry training set comprises one or more non-admixed ancestry training genotype segments; and For each non-admixed ancestry training set: applying a polygenic risk model corresponding to the ancestral labels of the non-admixed ancestry training set to each non-admixed ancestry training genotype segment of the non-admixed ancestry training set to generate one or more non-admixed ancestry training partial PRSs, and The one or more non-admixed ancestry training portion PRSs are standardized based on the mean PRS and the standard deviation PRS to generate a standardized non-admixed ancestry training portion PRS.
11. The method of claim 10, further comprising: include: A partial PRS was trained based on each non-admixed ancestry, and regression models were used to determine the partial PRS model effect size parameters.
12. The method of claim 1, further comprising: include: identifying one or more non-admixed ancestry training sets corresponding to each non-admixed ancestry cohort, wherein each non-admixed ancestry training set comprises one or more non-admixed ancestry training genotype segments, and each non-admixed ancestry training genotype segment corresponds to a complete genotype of a corresponding non-admixed ancestry individual; and For each non-admixed ancestry training set: applying a polygenic risk model corresponding to the ancestral signature of the non-admixed ancestry training set to each non-admixed ancestry training genotype segment of the non-admixed ancestry training set to generate one or more non-admixed ancestry training full PRSs, and The one or more non-admixed ancestry training full PRSs are standardized based on the mean PRS and the standard deviation PRS to generate a standardized non-admixed ancestry training full PRS.
13. The method of claim 12, further comprising: include: A full PRS was trained based on each non-admixed ancestry, and regression models were used to determine full PRS model effect size parameters.
14. An apparatus for generating a mixed-ancestry PRS for a mixed-ancestry subject, the apparatus comprising a processor and a memory storing software instructions which, when executed by the processor, cause the apparatus to perform the steps as recited in any one of claims 1 to 13.
15. A computer program product that generates a mixed-ancestry PRS for a mixed-ancestry subject, the computer program product comprising at least one non-transitory computer-readable storage medium storing software instructions that, when executed by a device, cause the device to perform the steps described in any one of claims 1 to 13.