Biomarker for lung disease predisposion

By analyzing genetic markers in the MUC5AC gene and protein composition, the method predicts lung disease predisposition, facilitating early intervention and personalized healthcare for COPD and IPF.

WO2026005676A1PCT designated stage Publication Date: 2026-01-02HANSSON GUNNAR C +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/SE2025/050549
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-26
Filing Date
2025-06-11
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

There is a need for predicting lung disease predisposition in humans, particularly identifying biomarkers for conditions like COPD and IPF to enable early intervention.

Method used

A method involving determining the genotype of SNP rs878913005 in the MUC5AC gene and the presence of arginine or tryptophan at amino acid position 1201 in the MUC5AC protein to predict lung disease predisposition, using nucleic acid and protein analysis techniques.

Benefits of technology

Enables the prediction of lung disease predisposition, allowing for targeted surveillance and treatment strategies based on genetic markers associated with increased risk of COPD and IPF.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SE2025050549_02012026_PF_FP_ABST
    Figure SE2025050549_02012026_PF_FP_ABST
Patent Text Reader

Abstract

Lung disease predisposition of a human subject is predicted by determining, in a sample comprising nucleic acid molecules from the human subject, genotype of a SNP rs878913005 located at the MUC5AC gene at chromosome 11 and predicting lung disease predisposition of the human subject based on the genotype of the SNP rs878913005.5.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] BIOMARKER FOR LUNG DISEASE PREDISPOSION

[0002] TECHNICAL FIELD

[0003] The present invention generally relates to a biomarker for lung diseases in human, and in particular to such a biomarker that can be used to predict lung disease predisposition and uses thereof.

[0004] BACKGROUND

[0005] Today there are many people suffering from different types of disorders related to dysfunctions of the respiratory tract. Common examples are bronchitis, chronic bronchitis, chronic obstructive pulmonary disease (COPD), asthma, emphysema, cystic fibrosis, common colds and especially idiopathic pulmonary fibrosis (IPF). Some of these diseases are chronic conditions and these have large negative impact on the life of the person.

[0006] These diseases can be caused by different mechanisms and generally give inflammation. However, common to all these disorders are an increased production and accumulation of mucus in the lungs. In chronic lung diseases, the accumulated mucus often becomes colonized by bacteria, further worsening the disease problems.

[0007] Mucus is a mixture of molecules where the large polymer forming mucins are a major constituent. The mucins present in the lungs are MUC5B and MUC5AC. These molecules are stored in goblet cells and undergo a >1 , 000-fold expansion upon secretion, a process requiring sufficient amount of liquid and bicarbonate to raise the pH. In diseases, the mucus remains attached to the epithelium.

[0008] In chronic lung diseases, the amount of lung mucins and especially the MUC5AC mucin are significantly increased. The MUC5AC mucin can form net-like polymeric structures and is therefore a target for binding bacteria. The MUC5AC mucin is also important for attaching the mucus to the epithelial cells.

[0009] Chronic obstructive pulmonary disease (COPD) is a progressive lung disease characterized by long-term respiratory problems and respiratoryfailure and eventually death. COPD is now the third most common cause of death, COPD is a heterogeneous lung disease characterized by chronic respiratory symptoms (dyspnea, cough, sputum production and / or exacerbations) due to mucus accumulation and abnormalities of the airways (bronchitis) and / or later emphysema.

[0010] The main causes of the development of COPD are the long-term exposure to harmful particles or gases, including tobacco smoke, that irritate the lung causing inflammation and mucus accumulation that together with predisposing host factors trigger the disease. The most well-known genetic risk factor today is alpha-1 antitrypsin deficiency (AATD).

[0011] Idiopathic pulmonary fibrosis (IPF) is a severe lung disease affecting the peripheral airways. As its name state, the cause of IPF is not understood today. However, a single nucleotide polymorphism (SNP) in the MLIC5B promotor (rs35705950) is strongly associated with increased risk of developing IPF suggesting a link to mucus accumulation. Later stages of IPF is characterized by the thickening and stiffening of lung tissue due to fibrosis. The development of fibrosis is progressive, sometimes fast, and the irreversible decline in lung function.

[0012] There is a need for predicting lung disease predisposition in humans, and in particular to identify and use biomarkers that can be used to identify human subjects having a predisposition for COPD or IPF allowing for early intervention.

[0013] SUMMARY

[0014] It is a general objective to provide a biomarker for predicting lung disease predisposition in humans.

[0015] This and other objectives are met by embodiments disclosed herein.

[0016] The present invention is defined in the independent claims. Further embodiments of the invention are defined in the dependent claims.

[0017] An aspect of the invention relates to method for predicting lung disease predisposition of a human subject. The method comprises determining, in a sample comprising nucleic acid molecules from the human subject, genotype of a SNP rs878913005 located in the MUC5AC gene on chromosome 11 and predicting lung disease predisposition of the human subject based on the genotype of the SNP rs878913005. Another aspect of the invention relates to a method for predicting lung disease predisposition of a human subject. The method comprises extracting proteins from a sample from a human subject. The method also comprises determining presence of arginine or tryptophan at amino acid position 1201 in the MLIC5AC protein. The method further comprises predicting lung disease predisposition of the human subject based on the determined presence or absence of arginine or tryptophan at amino acid position 1201 in the MLIC5AC protein.

[0018] The present invention involves the use a SNP indicative of lung disease predisposition in human subject. The SNP could thereby be used to predict lung disease predisposition of human subjects and identify individuals that are predisposed to develop lung diseases, such as COPD or IPF.

[0019] BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The embodiments, together with further objects and advantages thereof, may best be understood by making reference to the following description taken together with the accompanying drawings, in which:

[0021] Figure 1. CryoEM structure of the MUC5AC-D3 assembly.

[0022] (1A) MUC5AC-D3 sequence (SEQ ID NO: 8). VWD3 sequence is shown, C8-3, TIL3 and E3. The residues affected by the SNPs rs36189285 (R996) and rs878913005 (R1201) are marked by arrows. (1 B) Schematic sketch of the domains of MLIC5AC mucin with the N-terminal region (VWD1, VWD2, VWD’ and VWD3), nine CysD domains surrounded by PTS sequences densely O-glycosylated to form mucin domains and the C- terminal region (VWD4, VWCs and CK). The fragments analyzed are extracted and marked 1 , 2, and 3. (10) SDS-PAGE analysis of reduced and non-reduced D3 (1), D3-CysD (2) and D’-D3-CysD (3) reveals the formation of reducible dimers in all three fragments. (1D) MLIC5AC D3 assembly cryoEM density map and cartoon representation showing the disulfide bonds. The map and model of the two monomers are shown. The Ca2+ions are shown as spheres. The top left figure represents the top view of the molecule. It is rotated anticlockwise by 90°around the x-axis in the top right figure, and rotated clockwise by 45°around the y-axis and enlarged by 50% in the bottom figure showing the details of the front view. Putative intermolecular disulfide bonds are marked (Cys1132-Cys1132’ bond seems to be reduced). N-terminal (N) and C-terminal (C) of each monomer are marked. (1 E) Detail of the MUC5AC-D3 covalent dimerization interface zoomed in from (1 D) showing the interaction between C8-3 domains. (1 F) Detail of the MUC5AC-D3 covalent dimerization interface zoomed in from (1 D) showing the TIL3-TIL3’ interaction.

[0023] Figure 2. MUC5AC-D3 domains.

[0024] (2A) Front view of MUC5AC-D3 assembly dimer from cryoEM density map and model. The two VWD3 domains are shown light (left) and dark (right), C8-3 domains in dark and light grey, and TIL3 domains are marked. (2B) Structural alignment of MUC5AC-D3 (dark) and MUC2-D3 (light. PDB code: 6rbf). Left image is presented in the same orientation as (2A). Right image is rotated clockwise 45°. Regions with high variability are marked by black arrows. The / V-glycosylated Asp1154 in lateral chain of MUC2, not present in MUC5AC, is marked with a grey star. (2C) Detail of TIL3 structurally aligned of MUC5AC-D3 (grey), MUC2-D3 (dark grey. PDB code: 6rbf) and VWF-D3 (light grey. PDB code: 6n29). The TIL3 (31 -(32 loop is highlighted in brighter grey. To the left, superposition of all three structures showing the larger distance between the TIL3 and VWD3 domains in MUC5AC. The distinct disulfide bonds are marked by stars, the one connecting the TIL3 domain with C8-3 domain and the internal TIL3 (31 -(32 loop disulfide bond. All three structures are shown separately showing the TIL3 (31 -(32 loop and interfacing residues of lateral chains. The cysteines involved in the distinctive disulfide pattern are labeled. Hydrogen bonds between VWD3 and TIL3 are showed with dashed black lines and the residues involved are annotated. In MUC5AC, the residues affected by SNP variation at the amino acids R996 and R1201 are marked. (2D) Amino acid sequence alignment of MUC5AC (SEQ ID NO: 9), MUC5B (SEQ ID NO: 10), MUC2 (SEQ ID NO: 11) and VWF (SEQ ID NO: 12) TIL3 domains. Disulfide bonds are marked. Stars mark distinct disulfide bonds as in (2C). Cysteines are light grey and arginines in dark.

[0025] Figure 3. MUC5AC-D3 assembly open conformation.

[0026] (3A) CryoEM 2D classes, box size 220 A. The top figure shows the closed conformation 2D classes from the high-resolution structure shown in Figures 1 and 2. The discarded particles from an initial 3D classification were further 2D classified. These 2D classes are shown in the bottom panel, open conformation. (3B) CryoEM low-resolution map generated using the particles from (3A) bottom panel. The top figure shows the fitting on the closed conformation model and the bottom the proposed model for the open conformation. The VWD3 domain of one monomer is shown in light grey, C8-3 and TIL3 domains in dark grey and the connecting loop in dark grey. The VWD3 domains are marked VWD and VWD’, C8-3 and C8-3’, and TIL3 and TIL3’ are marked. The black arrows show the densities not covered by the models. The non-occupied densities in the closed form are explained by the movement of VWD3 as shown by the arrows. The C-terminal of TIL3 points toward the marked densities in the open conformation as they could represent E3 and / or CysD. (3C) Surface representation of the closed (left) and open (right) conformation marked by molecular lipophilicity potential (MPL) from light grey (most hydrophilic) via white to dark grey (most lipophilic). The newly exposed hydrophobic pocket residues in the open conformation are labeled. (3D) Detail of proposed model for MUC5AC-D3 open conformation. VWD3, C8-3 and TIL3 are marked. Putative salt bridges between K962 and E1148, and E981 and R1201, and the hydrogen bond between K979 and V1149 are represented by dashed lines. These residues and the interfacing histidines His977 and His 1109 are labeled. (3E) MUC5AC-D3 closed and open conformation and FCGBP D10 alignment. In the left figure MUC5AC-D3 closed (C8-3 and TIL3 marked) and open (C8-3 and TIL3 marked) conformation and FCGBP D10 (C8-10 marked) were aligned to the VWD domain (dark). The N-terminal (N) of VWD and C-terminal (C) of the different C8 domains are marked. In the right figure MUC5AC-D3 closed and open (dark) conformation and FCGBP D10 (top) were aligned by the C8 and TIL domains.

[0027] Figure 4. MUC5AC-D3 tetramerization.

[0028] (4A) Detail of TIL3-VWD3 interface mutants aligned with the MUC5AC-D3 wild-type (WT) (dark grey). The left figure shows the Arg996Gln mutant in light grey, the middle panel highlights the Arg1201 Trp mutant, and the right shows the double mutant Arg996Gln-Arg1201 Trp in dark grey. The mutations are pointed by arrows. (4B) CryoEM density map and cartoon representation of the MUC5AC-D3 dimeric assemblies WT, Arg996Gln, Arg1201 Trp and Arg996Gln-Arg 1201 Trp at two 45° angles. (4C) CryoEM 2D classes of higher order oligomers. Box sizes are specified for every group of classes. The groups of closed conformation oligomers are marked with “C” and the open conformation with “O”. (4D) MUC5AC-D3 Arg996Gln tetrameric assembly cryoEM density map and cartoon representation. One covalent dimer is shown in light grey (chain A) and dark grey (chain B) and the other in middle grey (chain C) and light grey (chain D). The His-tag density is shown in white. The top image shows a lateral view of the tetramer. It is rotated clockwise by 90° around y-axis and reduced 1.5 times in the middle-left figure, rotated clockwise by 90° around y-axis again in the middle-right figure, and rotated anticlockwise by 90°around x-axis and rescaled to the original size in the bottom figure. (4E) Detail of the MUC5AC-D3 non-covalent tetramerization interface zoomed from (4D). The predicted salt bridges and hydrogens bonds are shown as dashed lines. (4F) Structural alignment of the tetramer chain B and chain D against the R996Q dimer chain A.

[0029] Figure 5. MUC5AC-D3 physiological properties, SNPs and correlation with disease (5A) MUC5AC-D3 Arg996Gln tetrameric assembly cryoEM density map and cartoon representation. The PTS domains are schematically represented as polyalanine straightened chains protruding from the TIL3 C- termini. The PTS from each covalent dimer extend in opposite directions forming an angle of about 40° with the PTS from the other dimer. (5B) Ideal schematic representation of the MLIC5AC network generated by repetitions of (5A) linked by covalent dimerization at the end of the PTS domains (cystine-knot domain). (5C) Carnoy fixed human stomach biopsy paraffin section stained with a monoclonal anti-human MLIC5AC antibody (45M1) and Hoechst (nuclei; round white). Theen larged white square shows stratified surface mucus positive for MLIC5AC. (5D) Scanning electron micrograph of a piglet airway showing a MLIC5B bundled strand, MLIC5AC mucus attached to the bundle, and cilia. (5E) Frequency of SNPs increases in COPD (left) and IPF (right). ACMstands for mutant MUC5AC Arg1201 Trp (rs878913005), BMfor mutant MUC5B promotor (rs35705950), ACwtfor wild type in MLIC5AC Arg1201 position and Bwtfor wild type MLIC5B promotor in the position for rs35705950. Significance with the Fisher exact test is shown by three stars (p<0.001 ), one star (p<0.05) or a triangle (p<0.1 ). The bottom table shows the raw values used for the frequency calculations. (5F) Genomic organization of the MUC5AC and MUC5B gene locus and chromosome 11 . (5G) Linkage disequilibrium between MUC5AC Arg1201 Trp (rs878913005) and MUC5B promotor (rs35705950) in (5E) control, COPD and IPF groups. The graphic shows the frequency increase of both mutations appearing in the same subject in relation to the expected frequency if both mutations were independent. (5H) Formalin fixed paraffin section from an IPF lung explanted at lung transplantation, stained with a polyclonal anti-human MUC5B antibody (middle), a monoclonal anti-human MUC5AC antibody (45M1 ; right) and Hoechst (nuclei; white).

[0030] Figure 6. PCR-assay to detect rs36189285 and rs378913005 SNPs in MUC5AC

[0031] (6A) PCR amplification curves showing specific amplification of reference and alternate for the rs36189285 locus using control plasmid DNA. Also shown are the amplification curves of heterozygous control reactions indicating co-amplification of both reference and alternate alleles. Scatter plot indicates the distinct grouping of homozygous reference, homozygous alternate and heterozygous alleles allowing for identification of the rs36189285 genotype. (6B) PCR amplification curves showing specific amplification of reference and alternate for the rs878913005 locus using control plasmid DNA. Also shown are the amplification curves of heterozygous control reactions indicating co-amplification of both reference and alternate alleles. Scatter plot indicates the distinct grouping of homozygous reference, homozygous alternate and heterozygous alleles allowing for identification of the rs878913005 genotype. (6C) Scatter plots showing the identification of homozygous reference and heterozygous individuals tested for the rs36189285 SNP. (6D) Scatter plots showing the identification of homozygous reference, homozygous alternate and heterozygous individuals tested for the rs878913005 SNP.

[0032] Figure 7. MUC5AC and MUC5B SNP in IPF patients

[0033] (7A) Genomic organization of the MLIC5AC and MLIC5B gene locus in chromosome 11. (7B) Frequency of SNPs in control and IPF groups. ACMstands for mutant MLIC5AC Arg1201 Trp (rs878913005), BMfor mutant MLIC5B promotor (rs35705950), ACwtfor wild-type in MLIC5AC Arg1201 position and Bwtfor wild-type MLIC5B promotor in the position for rs35705950. Significance with the Fisher’s exact test is shown by three stars (p<0.001) or two stars (p<0.01 ). The Table shows the raw values used for the frequency and relative risk calculations. (7C) Relative Risk, ratio of the probability of developing IPF in participants carrying the SNPs to the probability of developing IPF in WT participants. For ACMand BMindependent analysis the relative risk of presenting one allele (Het) or two alleles (Hom) is showed separate. (7D) Linkage disequilibrium between MUC5AC Arg1201Trp (rs878913005) and MUC5B promotor (rs35705950) in (7B, 7C) control, and IPF groups. The graphic shows the frequency increase of both mutations appearing in the same subject in relation to the expected frequency if both mutations were independent.

[0034] Figure 8. MUC5AC SNP in COPD patients

[0035] (8A) Frequency of SNPs increases in COPD. ACMstands for mutant MUC5AC Arg 1201 Trp (rs878913005), BMfor mutant MUC5B promotor (rs35705950), ACwtfor wild-type in MUC5AC Arg1201 position and Bwtfor wild-type MUC5B promotor in the position for rs35705950. Significance with the Fisher's exact test is shown by one star (p<0.05) or a triangle (p<0.1 ). Table shows the raw values used for the frequency calculations. (8B) Linkage disequilibrium between MUC5AC Arg1201Trp (rs878913005) and MUC5B promotor (rs35705950) in control and COPD groups. The graphic shows the frequency increase of both mutations appearing in the same subject in relation to the expected frequency if both mutations were independent.

[0036] DETAILED DESCRIPTION

[0037] The present invention generally relates to a biomarker for lung diseases in human, and in particular to such a biomarker that can be used to predict lung disease predisposition and uses thereof. Mucins are a group of highly glycosylated molecules that cover and protect all mucosal surfaces of the body. The two gel-forming mucins MLIC5AC and MLIC5B constitute the main structural components of the mucus protecting the underlying epithelia in the respiratory system. Both mucins share the same domain organization and are large proteins, 5,654 or 5,762 amino acids and a mass of 586 or 596 kDa without the characteristic O-glycosylation. They are related to the other gel-forming mucins MLIC2 and MLIC6 and the von Willebrand factor (VWF). All these mucins have an N-terminal part built by 3.5 von Willebrand D assemblies, VWD1, VWD2, VWD’ and VWD3. The VWD assemblies are formed by the VWD domain, C8, trypsin inhibitor like (TIL) and E domains, except the VWD’ domain that only has the TIL and E domains. The gel-forming mucins have their N-termini followed by one or several PTS domains rich in proline (P), threonine (T) and serine (S), often in a repetitive fashion. The hydroxy amino acids in the PTS domains become heavily O-glycosylated to form the extended rod-like mucin domains. The dense sugar coating is responsible for the mucin hygroscopicity necessary for binding water and to mucus formation. The PTS domains are typically interrupted by a variable number of CysD domains, nine in MUC5AC and seven in MUC5B, involved in homotypic interactions as shown for CysD2 of MUC2. The MUC5AC, MUC5B and MUC2 mucin C-terminal region is formed by a VWD4 assembly, 3.5 VWC domains and a C-terminal cysteine-knot.

[0038] Secreted mucins interact specifically with each other and other molecules giving mucus specific properties. These mucins are shown to be orderly packed in the goblet cell granule due to low pH and high Ca2+. Upon secretion, they are unpacked into large disulfide-based polymers. The process requires an increase in pH and maybe calcium removal, both of which can be achieved by bicarbonate transported via the cystic fibrosis transmembrane conductance regulator (CFTR). The inter-molecular disulfide bonds formed in the endoplasmic reticulum (ER) and trans-Golgi network in the C-terminus and the N-terminal regions pave the ground for the expansion into covalent linear oligomers. However, additional non-covalent interactions are required for mucus organization.

[0039] The respiratory system is constantly exposed to inhaled particles, bacteria and viruses. Humans and pigs have submucosal glands down to the 10thairway branch. They secrete a chloride and bicarbonate rich fluid that pulls out the MUC5B mucin into long polymers exiting the glands as >20 m thick bundled strands. These contain >1,000 parallel MUC5B molecules that interact laterally. These are patchily coated by the MUC5AC from the surface goblet cells, which controls the bundle movement by attach ment / detachment events that together with MUC5AC threads clean the larger airways. The MLIC5B mucin is required for the normal lung homeostasis in mice whereas the MLIC5AC is not. Interestingly, the MLIC5AC mucin is induced and increased in amount at sites of metaplasia and disease in the lung. This affects the surface goblet cells during metaplasia, leading to co-expression of both MLIC5B and MLIC5AC in the same cell and even in the same granule, leading to the formation of an attached stratified mucus layer in the airways.

[0040] The question as to why we have two different mucins in the lung and in what way they differ has been a longstanding puzzle. The MLIC5B mucin is clearly forming linear molecules, which has been confirmed by electron microscopy. The MLIC5AC mucin on the other hand forms net-like structures. Experimental data as presented herein show that the disulfide bonded linear MLIC5AC mucin also interacts non-covalently through its VWD3 assembly and that there are genetic variants in this region that are linked to increased risk of lung disease.

[0041] Reference to positions within the human genome as used herein is to the Genome Reference Consortium Human Build 38 patch release 14 (GRCH38.p14) deposited on 3 February 2022, having GenBank assembly accession: GCA_000001405.29 and RefSeq assembly accession: GCF_000001405. This genome version is referred to as GRCh38 herein.

[0042] An aspect of the invention relates to a method for predicting lung disease predisposition of a human subject. The method comprises determining, in a sample comprising nucleic acid molecules from the human subject, genotype of a single nucleotide polymorphism (SNP) rs878913005 located in the MUC5AC gene on chromosome 11. The method also comprises predicting lung disease predisposition of the human subject based on the genotype of the SNP rs878913005.

[0043] The sample is a biological sample that comprises nucleic acid molecules and is obtained from the human subject. The nucleic acid molecules could be present as free nucleic acid molecules in the sample. Alternatively, the sample could contain nucleated cells from the human subject and where these nucleated cells contain a nucleus comprising the nucleic acid molecules, and in particular a copy of the genome of the human subject. The sample could be a body fluid sample comprising nucleic acid molecules, such as a body fluid sample comprising nucleated cells. Illustrative, but non-limiting, examples of such body fluid samples include a blood sample, a plasma sample, a serum sample, a saliva sample, a mammary gland milk sample, a vaginal cell smear or lubrication sample, and a semen sample. Alternatively, the sample could be a body tissue sample including, but not limited to, a hair sample, a hair root sample or a biopsy sample.

[0044] The nucleic acid molecules could be deoxyribonucleic acid (DNA) molecules or ribonucleic acid (RNA) molecules, including complementary DNA (cDNA) molecules. In a particular embodiment, the nucleic acid molecules are DNA molecules and preferably genomic DNA.

[0045] Nucleic acid molecules can extracted, isolated and optionally purified from the sample according to well- known nucleic acid extraction, isolation and purification methods. For instance, standard protocol for the isolation of genomic DNA could be used, such as are inter alia referred to in Sambrook et al., 2001 , and Sharma et al., 1993.

[0046] Single nucleotide polymorphism or SNP as used herein refers to a single nucleotide polymorphism at a particular position in the human genome that varies among a population of individuals. A SNP can be identified by its location on chromosome 11 , i.e., nucleotide (nt) position on chromosome, or by its name.

[0047] The SNPs referred to herein are listed in Table 1 below.

[0048] Table 1 - SNPs

[0049] In an embodiment, determining genotype comprises determining, in the sample comprising nucleic acid molecules from the human subject, whether the human subject has the C allele or the T allele at position 1180138 on chromosome 11.

[0050] Experimental data as presented herein shows that human subjects having the allele T have a comparatively high predisposition for lung diseases whereas a human subject having the allele C has a comparatively lower predisposition for lung diseases. Thus, a human subject having the allele T at position 1180138 on chromosome 11 (11 : 1180138) is more likely to suffer from a lung disease as compared to a human subject having the allele C at the position 11 : 1180138.

[0051] In an embodiment, predicting lung predisposition comprises predicting a high lung disease predisposition if the human subject has the T allele at position 1180138 on chromosome 11 and predicting a low lung disease predisposition if the human subject has the C allele at position 1180138 on chromosome 11 .

[0052] There are several methods known by those skilled in the art for determining whether a particular nucleotide sequence is present in a nucleic acid molecule and for identifying the nucleotide in a given position in a nucleic acid sequence. These include the amplification of a nucleic acid segment encompassing the genetic marker by means of polymerase chain reaction (PCR) or any other amplification method, interrogate the genetic marker by means of allele specific hybridization, 3'-exonuclease assay (Taqman assay), fluorescent dye and quenching agent-based PCR assay, the use of allele-specific restriction enzymes (RFLP-based techniques), direct sequencing, oligonucleotide ligation assay (OLA), pyrosequencing, invader assay, mini-sequencing, denaturing high pressure liquid chromatography (DHPLC) based techniques, single strand conformational polymorphism (SSCP), allele-specific PCR, denaturating gradient gel electrophoresis (DGGE), temperature gradient gel electrophoresis (TGGE), chemical mismatch cleavage (CMC), heteroduplex analysis based system, techniques based on mass spectroscopy (MS), invasive cleavage assay, polymorphism ratio sequencing (PRS), microarrays, rolling circle extension assay, high pressure liquid chromatography (HPLC) based techniques, extension based assays, amplification refractory mutation system (ARMS), amplification refractory mutation linear extension (ALEX), single base chain extension (SBCE), molecular beacon assays, invader (Third wave technologies), ligase chain reaction assays, 5'-nuclease assay-based techniques, hybridization capillary array electrophoresis (CAE), protein truncation assays (PTT), immunoassays, and solid phase hybridization (dot blot, reverse dot blot, chips). This list of methods is not meant to be exclusive, but just to illustrate the diversity of available methods. Some of these methods can be performed in microarray format (microchips) or on beads.

[0053] The SNP rs878913005 is in the coding region of the MUC5AC gene and changes the amino acid sequence of the MUC5AC protein. Thus, the SNP rs878913005 is a missense SNP, in which a single nucleotide change C^T results in a codon coding for another amino acid, in the present case tryptophan (Trp, W) instead of arginine (Arg, R).

[0054] The SNP rs36189285 is in the coding region of the MUC5AC gene and changes the amino acid sequence of the MLIC5AC protein. Thus, the SNP rs36189285 is a missense SNP, in which a single nucleotide change G^A results in a codon coding for another amino acid, in the present case glutamine (Gin, Q) instead of arginine (Arg, R).

[0055] Hence, in an embodiment, determining genotype comprises determining, in the sample comprising nucleic acid molecules from the human subject, whether a codon coding for amino acid number 1201 in the MLIC5AC protein codes for arginine (R) or tryptophan (W).

[0056] In an embodiment, predicting lung predisposition comprises predicting a high lung disease predisposition if the codon coding for amino acid number 1201 in the MLIC5AC protein codes for tryptophan and predicting a low lung disease predisposition if the codon coding for amino acid number 1201 in the MLIC5AC protein codes for arginine. rs35705950 is a well-known SNP in the promoter of the MUC5B gene that has been said to be associated with the lung disease IPF. Experimental data as presented herein, however, show that the association of this SNP rs35705950 with IPF was only significant when combined with the SNP rs878913005.

[0057] Thus, in an embodiment, the method comprises determining the genotype of both the SNP rs878913005 and the SNP rs35705950. In such an embodiment, the method further comprises determining, in a sample comprising nucleic acid molecules from the human subject, genotype of a SNP rs35705950 located at the promoter of the MUC5B gene at chromosome 11 . In this embodiment, predicting lung disease predisposition of the human subject comprises predicting lung disease predisposition of the human subject based on the genotype of the SNP rs878913005 and the SNP rs35705950.

[0058] The sample comprising nucleic acid molecules used to determine the genotype of the SNP rs35705950 could be a same sample or a different sample from the human subject used to determine the genotype of the SNP rs878913005. In the former case, genotypes of the two SNPs rs878913005 and rs35705950 are determined from the same sample comprising nucleic acid molecules from the human subject. In the latter case, the genotype of the SNP rs878913005 is determined in a first sample comprising nucleic acid molecules from the human subject, whereas the genotype of the SNP rs35705950 is determined in a second sample comprising nucleic acid molecules from the human subject.

[0059] In an embodiment, determining genotype comprises determining, in the sample comprising nucleic acid molecules from the human subject, whether the human subject has the T or A allele or the G allele at position 1219991 on chromosome 11. Thus, in this embodiment, the genotyping of the SNP rs35705950 comprises determining whether the human subject has the T or A allele or the G allele at position 1219991 on chromosome 11 (11 : 1219991).

[0060] In an embodiment, predicting lung predisposition comprises predicting a high lung disease predisposition if the human subject has the T or A allele at position 1219991 on chromosome 11 and i) the codon coding for amino acid number 1201 in the MLIC5AC protein codes for tryptophan and / or ii) if the human subject has the T allele at position 1180138 on chromosome 11. In this embodiment, a low lung disease predisposition is predicted if the human subject has the G allele at position 1219991 on chromosome 11 and i) the codon coding for amino acid number 1201 in the MLIC5AC protein codes for arginine (R) and / or ii) if the human subject has the C allele at position 1180138 on chromosome 11 .

[0061] Another aspect of the invention relates to a method for predicting lung disease predisposition of a human subject. The method comprises extracting proteins from a sample from a human subject. The method also comprises determining presence of arginine or tryptophan at amino acid position 1201 in the MLIC5AC protein and predicting lung disease predisposition of the human subject based on the determined presence or absence of arginine or tryptophan at amino acid position 1201 in the MLIC5AC protein.

[0062] The amino acid sequence of the MLIC5AC protein is presented in SEQ ID NO: 1.

[0063] The sample is a biological sample that comprises proteins and is obtained from the human subject. The sample could be a body fluid sample comprising proteins, such as a body fluid sample comprising nucleated cells. Illustrative, but non-limiting, examples of such body fluid samples include a blood sample, a plasma sample, a serum sample, a saliva sample, a mammary gland milk sample, a vaginal cell, smear or lubrication sample, and a semen sample. Alternatively, the sample could be a body tissue sample including, but not limited to, a biopsy sample.

[0064] Proteins can be extracted from the sample according to techniques well known in the art including, but not limited to, centrifugation, electrophoresis, chromatography, and precipitation.

[0065] Presence or absence of arginine or tryptophan at amino acid position 1201 in the MLIC5AC protein can be determined according to various techniques for single amino acid substitution (SAAS) identification. An example of such techniques is mass spectrometry (MS) analysis, such as tandem mass spectrometry (MS / MS) or liquid chromatography MS / MS (LC-MS / MS). Also immunoassays using antibodies specifically binding to an epitope on MLIC5AC encompassing the amino acid position 1201 and having different binding characteristics to the MLIC5AC protein depending on whether the amino acid position 1201 is arginine or tryptophan could be used. Further examples include various proteomics technologies and analyses.

[0066] In an embodiment, predicting lung predisposition comprises predicting a high lung disease predisposition if the amino acid position 1201 in the MLIC5AC is tryptophan and predicting a low lung disease predisposition if the amino acid number 1201 in the MLIC5AC protein is arginine.

[0067] As mentioned in the foregoing, the prediction of lung disease predisposition could also be based on the SNP rs35705950, in addition to the amino acid at position 1201. In such an embodiment, the method further comprises determining, in a sample comprising nucleic acid molecules from the human subject, genotype of a SNP rs35705950 located at the promoter of the MUC5B gene at chromosome 11. In this embodiment, predicting lung disease predisposition comprises predicting lung disease predisposition of the human subject based on the determined presence or absence of arginine (R) or tryptophan (W) at amino acid position 1201 in the MLIC5AC protein and the SNP rs35705950.

[0068] In an embodiment, determining genotype comprises determining, in the sample comprising nucleic acid molecules from the human subject, whether the human subject has the T or A allele or the G allele at position 1219991 of chromosome 11 . In an embodiment, predicting lung predisposition comprises predicting a high lung disease predisposition if the human subject has the T or A allele at position 1219991 on chromosome 11 and the amino acid position 1201 in the MLIC5AC is tryptophan. Further, a low lung disease predisposition is predicted if the human subject has the G allele at position 1219991 on chromosome 11 and the amino acid number 1201 in the MLIC5AC protein is arginine.

[0069] The lung disease is preferably a disease or disorder causing or associated with dysfunctions of the respiratory tract. The lung disease is preferably characterized by an increased production and accumulation of mucus in the lungs. Examples of such lung diseases include bronchitis, chronic bronchitis, chronic obstructive pulmonary disease (COPD), asthma, emphysema, cystic fibrosis, common colds and especially idiopathic pulmonary fibrosis (IPF). In an embodiment, the lung disease is selected from the group consisting of COPD and IPF. In an embodiment, the lung disease is COPD. In another embodiment, the lung disease is IPF.

[0070] The method for predicting lung disease predisposition of a human according to any of the embodiments above may also comprise selecting a surveillance schedule for the human based on the predicted lung disease predisposition. Thus, an optimal or at least suitable surveillance schedule or scheme is selected for the human subject based on the whether the human is predicted to be predisposed to develop or suffer from a lung disease or not. This means that human subjects predicted to be lung disease predisposed could be selected for a more frequent surveillance and follow-up (first surveillance schedule) as compared to human subjects predicted not to be lung disposed predisposed, which instead can follow a less frequent surveillance and follow-up (second surveillance schedule) if any.

[0071] The method for predicting lung disease predisposition of a human subject according to any of the embodiments above may further also comprise selecting a lung disease treatment for the human subject based on the predicted lung disease predisposition. Thus, an optimal or at least suitable treatment is selected for the human subject based on the whether the human subject is predicted to be predisposed to develop or suffer from a lung disease or not.

[0072] EXAMPLE

[0073] Gel-forming mucins MUC5AC and MUC5B constitute the main structural component of the mucus in the respiratory system. Secreted mucins interact specifically with each other and other molecules giving mucus specific properties. The cryoEM structures of the wild type MUC5AC-D3 assembly and the structural SNP variants R996Q and R1201W were determined. The structures explain the basis of MLIC5AC N-terminal non- covalent oligomerization upon secretion. The MUC5AC-D3 assembly forms covalent dimers in two alternative conformations, open and closed. The closed conformation dimers interact through an arginine rich loop in the TIL3 domain forming tetramers. Moreover, a positive disease correlation was found between the SNP (R1201W, rs878913005), Chronic Obstructive Pulmonary Disease (COPD), and Idiopathic Pulmonary Fibrosis (IPF). The well-known MLIC5B promotor SNP (rs35705950) associated with IPF was most significant when combined with the MLIC5AC SNP. The Example provides a model to explain the formation of MLIC5AC net-like structures and how both SNPs will affect mucus organization and increase risk of lung disease.

[0074] RESULTS

[0075] MUC5AC-N covalent dimer

[0076] Expression of the secreted complete N-terminal part of MUC5AC (MUC5AC-N) has proved to be challenging and has only been achieved at low yields in polarized airway cell lines. Efforts have previously been made to produce the complete secreted MUC5AC in Chinese hamster ovary (CHO) and human embryonic kidney (HEK) cells, but without any success. The VWD1-VWD2 (D1-D2) assemblies of VWF and MUC2 are cleaved off after polymerization suggesting that these domains are required for packing, but less important for the mature protein. MUC2 D1-D2 assemblies are compactly packed at low pH in the goblet cell granule at the same time as they are involved in intracellular polymerization. MUC2-N expressed in CHO cells and analyzed by cryoEM at pH 7.4 showed that only the compact D3 dimer assembly could be visualized, the D1, D2 and D’ were too flexible. This suggests less importance of D1-D2 for the mature secreted mucins at physiological, neutral pH. Accordingly, the present Example focused on the MUC5AC-D3 (Figures 1A and 1 B) assembly and its interactions.

[0077] Expression plasmids including the N-terminal D’ assembly, the D3 assembly and the CysD1 domain with and without an N-terminal 6xHis-tag were designed (Figure 1 B). These plasmids were expressed in Lee 3.2.8.1 CHO cells and the recombinant proteins purified and analyzed by gel electrophoresis showing the expected band sizes after reduction (Figure 1 C). Without reduction, the bands migrate approximately at the double size, suggesting that all three recombinant proteins are disulfide bonded dimers. The structure of the MUCAC-D3 assembly was analyzed by cryoEM at pH 7.4. The cryoEM reconstruction of the D’-D3-CysD1 recombinant protein reached higher resolution (3.2 A) than the isolated D3, even if only the D3 assembly is visible due to flexibility of the D’ and CysD1 domains (Figure 1 D). The overall structure of the MUC5AC-D3 assembly dimer is similar to the previously reported D3 assemblies. The covalent MUC5AC-D3 assembly dimer is formed via two intermolecular disulfide bonds, Cys1132-Cys1132 in the C8-3 domain and Cys1174-Cys1174 in the TIL3 domain (Figure 1 D-1 F). However, the Cys1132 inter-molecular disulfide bond is partially reduced. The interaction surface is highly conserved between MUC5AC-D3 and MUC2-D3 with only three relevant substitutions. In the C8-3 domain interaction region, MUC5AC-D3 presents a phenylalanine (Phe1086, Figure 1 E) instead of a histidine (MUC2-D3 His1042), potentially affecting the effect of pH in the intracellular packing. In the TIL3 domain interface, Ser1154 (Figure 1 F) substitutes a phenylalanine present in an equivalent position in MLIC2 (Phe1110). This serine is highly conserved in MLIC5AC and MLIC5B between species. The hydrophobic pocket formed by Phe1164’, Tyr1167’ and Tyr1168’ interacts with Pro1158, while Tyr1178, Leu1163 and the aliphatic chain of Arg1156 creates another pocket that accommodates Phe1164’ (Figure 1 F). In MLIC2 the second hydrophobic pocket is also formed by Phe1110 stabilizing the interaction with Phe1164’. In the same area, MLIC2 interfacing residues Arg973’- Asp1115 are substituted in MLIC5AC by Seri 159-Lys101 T resulting in a loss of a salt bridge. No hydrogen bonds between Seri 159-Lys1017’ were observed.

[0078] The structure of the MUC2-E3 domain could be solved by crystallography due to the crystal packaging contacts showing high B-factors. However, in solution MLIC5AC is too flexible to be modeled. After 3D refinement only some noisy densities appeared in the area where it should be located.

[0079] MUC5AC-D3 TIL3 structure

[0080] The overall structure of MLIC5AC (Figure 2A) is similar to MLIC2 and VWF. However, an overlay of MLIC5AC and MLIC2 showed differences including the loops |34-[35 and |38-[39 in VWD3 domain and the loop [31 -[32 in the TIL3 domain, all located on the same side of the molecule (Figure 2B). The TIL3 domain is the region of the molecule presenting more differences when compared with MUC2-D3 and VWF-D3 assemblies (Figure 2C). The MUC5AC-TIL3 [31 -[32 loop, unlike the other mucins and VWF, is remarkably rich in arginines (20%) (Figure 2D). The disulfide bond organization in this domain is unique but still shows characteristics in common with the structures previously described. The disulfide bonds between MUC5AC-TIL3 cysteines 1165-1206, 1189-1228 and 1210-1224 are conserved at similar positions in all three proteins, in addition to the Cys1174 forming the intermolecular dimer disulfide bond. The Cys1181 interacts with the C8-3 region in MLIC5AC and MLIC2 (Cys1137) stabilizing the domain organization while the equivalent cysteine in VWF (Cys1149) is involved in an extra interdomain disulfide bond. On the other hand, MLIC5AC and VWF have an identical disulfide bond in the center of the loop (31-|32 that is absent in MLIC2, Cys1185-Cys1196 in MLIC5AC and Cys1153-Cys1165 in VWF. MLIC2 contains an / V-glycosylation site in this area, Asn1154, not present in the other mucins or VWF. The MUC5B-TIL3 domain shares 74.2% identity with MLIC5AC and almost certainly displays the same disulfide bond arrangement.

[0081] There were also remarkable differences in the relative position of the whole TIL3 domain in relation to VWD3. In MLIC2 and VWF, the center of the TIL3 [31 -[32 loop is located only 6.2 A and 8.8 A away from VWD3 [39, while in MLIC5AC they are separated by 14.4 A. This close contact is explained by the presence of multiple hydrogen bonds between these domains in MLIC2 (Gln956-Glu1161 , Glu961-Ser1153 and Lys979-Tyr1159) and VWF (G lu 954-His 1159, Ser958-Ala1152 and Ser958-H is 1174) that are completely absent in MLIC5AC.

[0082] MUC5AC-D3 open conformation

[0083] During the 3D-classification processes, a second MUC5AC-D3 assembly conformation was observed. The 3D volume refined with these particles showed low resolution with a severe preferred orientation problem. The particles were 2D classified showing all classes having the same top view orientation (Figure 3A). The maximum diameter of the new open conformation classes is 30% higher than in the standard or closed conformation. The 2D top view of both conformations display enough details to recognize the VWD3 domain [3-sandwich and the C8-3 domain, and together with the low-resolution 3D map allowed us to model this conformation into the density maps (Figure 3B). The VWD-C8 interaction dissociates and exposes a hydrophobic surface in VWD3 to the solvent (Figure 3C). The domains establish a novel interaction between the loops [37-[38 and [311 -[312 in VWD3 and the C-terminal side of o4 in C8-3, and between the VWD3 [38 and [31-[32 TIL3 loop (Figure 3D). The long connecting loop between VWD3 and C8-3 domains makes this conformational change possible (Figure 3B). This conformational change could explain the previous observation that MLIC5AC binds significantly more to hydrophobic surfaces compared to MLIC5B.

[0084] Recently an open structure from the D10 assembly in FCGBP has been published, supporting the idea that the opening of the VWD assemblies could have a physiological function. However, there are multiple differences between these two structures (Figure 3E). The FCGBP structure was solved by crystallography and therefore the crystal packaging could promote or alter the structure of the open conformation. The relative position of the domains is completely different. In the MUC5AC-D3 closed conformation the N-terminal of VWD and the C-terminal of C8 points in the same direction and are only 12 A apart. In FCGBP they are 55 A apart and rotated 90°, while in MLIC5AC open conformation they are at 34 A and rotated -45°. Furthermore, in FCGBP the domains interact only through the connecting loop as there are no other intramolecular interactions in contrast to MLIC5AC (Figure 3D). The FCGBP-D10 structure lacks the TIL domain known to interact with the VWD domain in other VWD assemblies to maintain the compact structure where the conformational changes were explained by the GDPH autocatalytic cleavage not present in MUC5AC-D3.

[0085] MUC5AC-D3 variants

[0086] Interestingly, when genomic databases of MUC5AC were analyzed, the parts of MUC5AC most different from the other mucins as discussed above (Figure 2) contained two variants where a single amino acid had been replaced (Figure 4A). In both of these, one of the arginines typical for MUC5AC was replaced by a glutamine (Arg996Gln, rs36189285) or by tryptophan (Arg1201Trp, rs878913005). We produced recombinant proteins with each variant separate as well as with both arginines replaced. The structure of each of these were solved by cryoEM and compared. The general appearance of the dimers is essentially identical (Figure 4B). However, careful comparison of each variant to the WT elucidates the effects of the mutations (Figure 4A). The Arg996Gln mutation, located in VWD3 p9, is placed directly in the interface between the VWD3 and the TIL3 domain. The substitution results in a slight reduction of the distance between the two domains (Figure 4A). No interactions involving this residue are observed in the WT or in Arg996Gln, but considering the resolution limitations a hydrogen bond between Arg996 or Gln996 and a TIL3 [31 -[32 loop residue cannot be completely ruled out. The observed alteration can also be attributed to arginine steric repulsion.

[0087] In contrast, the Arg1201 Trp mutation, placed in the [31 -[32 TIL3 loop, only affects the adjacent residues. The solvent exposed arginine is substituted by a tryptophan pointing toward the VWD-C8 / TIL interface. The tryptophan accommodates into the VWD3 hydrophobic interfacing region in close proximity to Leu1013. In the double mutant, the effect of the Arg996Gln mutation is more pronounced, bringing the VWD3 and the TIL3 domain closest together. Even though no major rearrangements were found, both SNPs can affect the equilibrium between the dimer open and closed conformation. Open conformation dimers were constantly found during cryoEM 2D and 3D classifications in the WT assembly. Only when 2D templates from WT open conformation were used for particle picking, these could be detected in the Arg996Gln sample. MUC5AC-D3 tetramers

[0088] The grids of the MUC5AC-D’D3-CysD1 assembly showed the covalent dimers (Figure 2), but also numerous images showed tetrameric particles (Figure 4C). In the same way, all MLIC5AC variants showed some tetramers in cryoEM. The interaction is shown to be flexible as most of the 2D classes are noisy and 3D reconstructions only show elongated blobs, but the presence of the His-tag stabilizes it. The WT 6xHis- MUC5AC-D3 assembly shows 2D classes with structural features in just one orientation and the 3D reconstructions thus have low quality. Similarly, the structure of the tetrameric 6xHis-MUC5AC-D3 Arg996Gln assembly was solved at low resolution (6-7 A), not enough to reveal the molecular mechanism of the non- covalent oligomerization. Surprisingly, the 6xHis-MUC5AC-D3 Arg996Gln adopts another conformation not observed in the other variants as the protein forms ring shaped octamers (Figure 4C). The ring formation gives stability to the interactions. However, after refinement only one of the tetramers was well defined, the others were poorly defined due to remaining flexibilities and difficulties in isolating octamers from tetramers in some orientations. Particle subtraction followed by local refinement let us solve the structure of a tetramer forming part of the octamer at 3.69 A resolution (Figures 4D and 4E)

[0089] The MUC5AC-D3 non-covalent tetramer interaction is driven by the TIL3 domain and involves an interface area of 385.4 A2(Figure 4D). The interaction between the MLIC5AC TIL3 domains occurs in the arginine-rich loop [31 -[32 stabilized by a disulfide bond between Cys1185-Cys1196 (Figures 2C and 2D). The interaction is mainly hydrophilic (Figure 4E). Intermolecular salt bridges and hydrogen bonds are formed between lateral chains of Asp1199-chainB (Asp1199B) and Arg1198-chainD (Arg1198D), and between Arg1198B lateral chain and the main chain of Arg1193D and Asp1195D, Arg1193D and Cys1185B, and between Arg1198D and Leu1197B and Arg1198B. The model also shows the presence of the lateral chain of Arg1187D and Arg1193D in the proximity of the Cys1185B-Cys1196B disulfide bond and Arg1198B pointing to the same disulfide bond in chain D.

[0090] The comparison between the tetrameric MUC5AC-D3 Arg996Gln assembly and the free closed conformation dimer shows significant differences only in the TIL3 domain region (Figure 4F). While the tetramer chain D shows only minor deviations, the TIL3 loop [31 -[32 in chain B changes its relative position with C8-3 and VWD3 domains allowing the interaction with chain D. The solved structure confirms the role of the His-tag in the interaction. No interactions between the His-tag of one monomer and the core of the other one were observed, just His-tag— His-tag interactions, supporting that it is an unspecific interaction that stabilizes the TIL3-TIL3 interface (Figure 4E).

[0091] Moreover, 2D classifications suggest that the MUC5AC-D3 assembly can also form an additional kind of tetramers involving the open conformation (Figure 4C). This interaction is linear and leads to the formation of high-order oligomers that look longest in Arg996Gln. However, the open conformation oligomers were only found in assemblies lacking D’ or N-terminal 6xHis-tag. The 2D classes together with the open conformation low-resolution map (Figure 3B) suggested a major steric impediment at the N-terminal region. The symmetric interaction of two VWD3 domains through the external side of the [3-sheet 1 seems to bury both N-terminal residues. Assemblies containing D’ or N-terminal His-tag were only found to form tetramers based in closed conformation covalent dimers, making it likely that the open form tetramers are an artefact.

[0092] MUC5AC-D3 tetramerization and physiological properties

[0093] The tetrameric MLIC5AC will cross-link linear molecules with an angle of about 40° as illustrated in Figure 5A. The O-glycosylated PTS domains are extended rods and, thus, expected to drive the MLIC5AC N-termini apart. This will generate a net-like structure as suggested in its ideal form in Figure 5B. This model is well in line with what we know of MLIC5AC organization from staining of tissues. When tissue sections of the surface mucus of the stomach was stained for MLIC5AC, a stratified and laminated organization appears (Figure 5C). This is similar to the inner mucus layer of the colon made up by the MLIC2 mucin. Electron microscopy of the tracheal surface shows the different organization of the two lung mucins, linear bundled strands of MLIC5B and more net-like appearances of the MLIC5AC (Figure 5D). The different organization of these two mucins is as shown by Carpenter et al. 2021 when pure MLIC5B and MLIC5AC mucins were studies by EM.

[0094] MUC5AC-D3 SNPs and disease

[0095] The SNP sequence variant where Arg1201 is replaced by Trp (rs878913005) has not normally been included in GWAS studies as it is located close to the variable region of the MLIC5AC mucin. To analyze if there is any coupling of this SNP to disease, the complete genome of individuals of the UK-biobank were analyzed for this SNP and compared to available medical information. The SNP (Arg 1201 Trp) was found to be overrepresented in patients with COPD (n=1 ,139). The frequency was 0.3740 in the COPD cohort while it was 0.3397 in the control group (n=5, 196), representing a statistically significant increase at p<0.05 (Fisher exact test statistic value 0.0278) of 0.0343 (10%) (ACMin Figure 5E). This SNP was also analyzed for patients with idiopathic pulmonary fibrosis (IPF). Although this cohort was small (n=49), an increased frequency rise of 0.1297 (38%) was observed, although this was only at p<0.1 due to the small number of individuals (Figure 5E).

[0096] In order to discard the influence of other known SNPs in IPF, the well documented rs35705950 SNP affecting the MLIC5B mucin promotor region was analyzed in the UK-biobank individuals. A significant (p<0.001) increase in frequency of this SNP in IPF was found for rs35705950, 0.2136 (110%) (BMin Figure 5E). Interestingly, when both SNPs were analyzed in relation to each other, we observed that the association between SNPs and disease is only significant when both SNPs appear at the same time with an increased frequency of 0.178 p<0.001 as compared to only 0.036 with WT MLIC5AC and the rs35705950 in MLIC5B promotor (Figure 5E). In contrast, the effect of the MLIC5AC SNP rs878913005 on the frequency of COPD was not related to the rs35705950 MLIC5B.

[0097] The MLIC5AC and MLIC5B mucins are located after each other on chromosome 11 with 39,853 nucleotides in between (Figure 5F). A genetic link between the two is, thus, possible. Analysis for linkage disequilibrium revealed a value of 0.062 and 0.064 between rs878913005 (MLIC5AC) and rs35705950 (MLIC5B) in the control and COPD groups, respectively. However, the linkage disequilibrium between the two SNPs increased to almost double, 0.115, in the IPF cohort (Figure 5G). Consequently, the frequency of having both SNPs increases from the expected 0.066 to 0.128 (0.066-H).062) in the control group while it increases from 0.191 to 0.306 (0.191 -+0.115) in IPF, meaning that the probability of presenting both SNPs in IPF is 139% higher than in the control group. Together the results suggest that the MLIC5AC variant (Arg1201 Trp) show increased risk of COPD, but the risk for IPF is especially increased.

[0098] The analysis of the other MLIC5AC SNP (rs36189285, Arg996Gln) studied here, was inconclusive due to its low frequency in the European population and therefore in the UK-biobank. This SNP is more frequent in the Asian population.

[0099] IPF is characterized by fibrosis in the smaller airways. When peripheral human lung tissue from patients with IPF were stained, most surface secretory cells expressed both the MUC5B and MUC5AC mucins (Figure 5H) further suggesting that the MUC5AC could have a role in IPF. MUC5AC-D3 SNP and IPF

[0100] The rs878913005 SNP, resulting in an Arg to Trp substitution at position 1201 has not normally been included in GWAS studies. This is due to its absence in SNP arrays and that it is normally not included at genomic exon sequencing. The reason for this is its location close to the large very variable region of the MLIC5AC mucin (Figure 7A). Therefore, whole random genomic sequencing databases was utilized for its identification. To analyze any coupling of this SNP to disease, the complete genome of individuals included in the UK- Biobank were analyzed for this SNP and compared to the available medical information (Figure 7B). The SNP rs8789913005 (Arg1201 Trp) was analyzed for patients classified to have IPF. 107 patients were identified, four of these patients that had post-mortem diagnosis and were not classified as IPF, but as three were WT and the fourth heterozygous, these only lowered the significance and were not removed from the study. Although this IPF cohort was relatively small (n= 107), 53 of these patients had at least one allele with the rs878913005 SNP. This is compared to the 33.9% (1 ,765) of the control group (n=5, 196). This means that the IPF patients showed a significant 1 .46 fold increase in frequency (p<0.01, Fisher exact test statistic value 0.0013, Figure 27) and a 1.9 fold risk increase for any allele (1.5 for one allele and 3.4 for two alleles, Figure 7C).

[0101] In order to analyze the influence of other known SNPs on IPF, the well documented rs35705950 SNP affecting the MLIC5B mucin promotor region was analyzed in the same UK-Biobank cohort. As expected, a significant 2.64 fold (p<0.001 ) increase in frequency of this SNP in IPF was found (BMin Figure 7B) and a 4.2 fold risk increase for any allele (3.1 for one and 5.8 for both alleles, Figure 7C). Interestingly, when both SNPs were analyzed in relation to each other, we observed that the association between SNPs and disease was stronger when both SNPs appeared at the same time. The increased frequency of presenting the two SNPs at the same time in IPF patients is higher than when only the rs35705950 in MLIC5B promotor is combined with WT for MUC5AC, 2.70 fold increase (p<0.001, ACM / BM) and 2.11 (p<0.01, ACwt / BM), respectively (Figure 7B). When only the MLIC5B SNP was present together with WT MLIC5AC the risk increase frequency of IPF was reduced by almost half from 4.2 fold increase to 2.2 suggesting that the well-studied MLIC5B SNP variant in isolation might be of lower importance in the absence of the MLIC5AC SNP (Figure 7C). However, the MLIC5AC SNP is even more dependent on the MLIC5B SNP as the relative risk of MLIC5AC in isolation is 0.7 (Figure 7C). The MLIC5AC and MLIC5B mucins are located close together on chromosome 11 with 39,853 nucleotides between the two IPF associated SNPs (Figure 7A). A genetic link between the two is therefore possible. Analysis for linkage disequilibrium revealed a value of 0.062 and 0.064 between rs878913005 (MLIC5AC) and rs35705950 (MLIC5B) in the control group and a group of Chronic Obstructive Pulmonary Disease (COPD) patients, respectively (Figure 7D, Figure 8B). However, the linkage disequilibrium between the two SNPs increases up to 0.092 (48%) in the IPF cohort (Figure 7D). Consequently, the frequency of having both SNPs increases from the expected 0.066 to 0.128 (0.066-H).062) in the control group while it increases from 0.254 to 0.346 (O.254-H).O92) in IPF, meaning that the probability of presenting both SNPs is almost 3 times higher in IPF patients (Figure 7). Together the results suggest that the MLIC5AC variant (Arg1201 Trp) together with the MLIC5B promoter SNP are associated with an increased risk of IPF.

[0102] The rs878913005 MLIC5AC SNP (Arg1201 Trp) was also found to be overrepresented in patients with COPD (n=1,139). The frequency is 0.37 in the COPD cohort while it is 0.34 in the control group (n=5,196), representing a statistically significant increase (p<0.05, Fisher’s exact test statistic value 0.0278) of 0.03 (10%) (ACMin Figure 8), whereas no association the MLIC5B SNP rs35705950 was observed.

[0103] Mucins and IPF

[0104] The MLIC5B mucin is normally absent from the smallest human airways, terminal and respiratory bronchioles, but found in surface cells of the slightly larger airways without submucosal glands. The MLIC5AC mucin is also normally absent from peripheral human airways. To study the localization of MLIC5B and MLIC5AC in peripheral airways, tissues were obtained from the removed lung following transplant of IPF patients and from tissues resected for lung adenocarcinoma. The presence of MLIC5AC and MLIC5B SNP genotypes of these patients were determined. When IPF tissue were stained for MLIC5B and MLIC5AC, most surface cells in the few remaining small airways stained positive for MLIC5B. Some of these cells also stained strongly for MLIC5AC, showing that both mucins could be produced in the same cell. The absence of the two mucins in ‘normal’ small airways is shown by staining with the same antibodies in control patients.

[0105] Most of the tissue in a removed lung from a transplanted IPF patient show severe destruction and fibrosis as revealed by H&E and Masson-Goldner trichrome staining, such tissue destruction is absent in control samples. In IPF patients the peripheral architecture and alveoli are lost, and the tissue is exchanged for fibroblasts and fibrotic tissue and microscopic honeycomb cysts characteristic for IPF. These cysts are often filled with large amounts of mucus where the MLIC5B mucins predominates, but the MLIC5AC mucin is also observed in parts of this mucus accumulation.

[0106] More detailed studies of IPF lungs show the MLIC5AC mucin often is localized close to the apical surface of the cells lining of the cysts and often in direct contact with the cyst epithelium. The MLIC5AC also appears mixed with MLIC5B out in the accumulated mucus. Interestingly, the MLIC5AC often appears as sheets that might suggest that they have been formed along the epithelial lining and then moved out into the mucus. These MLIC5AC sheets are formed and likely also attached to the cyst cell surface where the MLIC5B mucus has shrunk due to fixation and separated from the MLIC5AC which remain attached to the surface of the cyst epithelium.

[0107] The cells ling the typical IPF microscopic honeycombs often appear as flat cells and sometimes as more cuboidal cells. Some areas contain more typical goblet cells and when present more or less every cell is a goblet cell. These cells are the likely origin of the mucus and would explain the large amounts of mucus found in the cysts. The goblet cells here contain both the MLIC5B and MLIC5AC mucins and show mucus secreting from the cells that then mixes with cyst luminal mucus.

[0108] The D3 mediated covalent dimerization in the MLIC5AC mucin described in this Example as predicted by homology with MLIC2, together with the C-terminal dimerization observed in all gel-forming mucins, lead to the formation of long linear polymers. Different assemblies and structural rearrangements occur along the secretory pathway to achieve an efficient unpacking upon release and thus the N- and the C-terminal dimerization must occur in an ordered way. First, the C-terminal inter disulfide bond is formed in the ER (pH 7.4) and then in the trans-Golgi network (pH 6.2) the N-termini are coupled together. To prevent the N-terminal disulfide bond formation in the ER, the interaction of the covalent oligomerization interface should only happen at lower pH. The dominance of charged residues close to the interface could have a role regulating the formation of the covalent link along the secretory pathway, especially His1177 and Asp1166, conserved in all gel-forming mucins. Different structures of the MUC2-D3 assembly have been solved by crystallography and cryoEM, all of them at low pH. We present the first mucin D3 assembly solved at neutral pH, but interestingly no major structural differences were observed. It indicates that once the covalent dimerization is stabilized the interface is locked, even if the Cys1132-Cys1132 dimer bond is partially reduced. This disulfide bond reduction was first observed for the equivalent Cys observed in the MLIC2 crystallographic structure. This was atributed to radiation damage, an event that could also occur in cryoEM and can provide information about “weak links” that can be of structural significance. This suggests that the formation of the Cys1132- Cys1132 disulfide bond could be tightly regulated by its redox potential, requiring the higher oxidizing environment found in the Golgi compared to the ER. Its higher radiation damage susceptibility could also be explained by the formation of a stabilizing S---0 interaction with the carbonyl 0-atom from the same cysteine located just at 3.3 A from the S-atom.

[0109] Nevertheless, observation of MLIC5AC under physiological conditions has shown that it is different from MLIC5B in that it does not appear as long linear polymers. Instead, it forms more complex and likely net-like structures as shown here and previously, requiring the existence of other intermolecular interactions. The here observed capability of MLIC5AC to form non-covalent dimers through the TIL3 domain and thus MUC5AC-D3 tetramers explains the molecular mechanism for the net-like polymer formation. The interaction is mainly hydrophilic as it is based on multiple hydrogen bonds and salt bridges through a densely positively charged region. Unlike the covalent MUC5AC-D3 dimerization that locks the assembly, the non-covalent formation of tetramers is flexible and can be regulated by pH and ionic strength upon secretion. The unique MUC5AC arginine rich TIL3 composition suggests that this interaction only occurs in MUC5AC.

[0110] MUC5AC can appear both in a closed and an open conformation. Although the open model is derived from low resolution cryoEM, it shows that the VWD assemblies can open in solution supporting the physiological relevance of the previously reported FCGBP VWD10 structure. Considering other known VWD3 assemblies, we cannot exclude the existence of an open VWF or MUC2-D3 open conformation, but the closer interactions between the TIL domain and VWD make them less energetically favorable. The open conformation implies the exposure of a highly hydrophobic surface that is not evident in FCGBP-D10. The likely artefactual open conformation tetramers and high order oligomers found in MUC5AC-D3 suggest that the external site of the VWD3 [3-sheet 1 becomes a new interaction surface that could establish contacts with other mucins, mucus associated proteins, or other domains from the same molecule. This could explain the hydrophobic properties atributed to the mucins but do not explain any of the actual structures.

[0111] In addition to the observed non-covalent interactions within the VWD3 assembly explaining the net-like appearance of MUC5AC, there are additional possible interactions within mucins. The CysD2 domain of MUC2 was recently shown to form weak homotypic dimers that was further stabilized by transglutamination catalyzed by TGM3. Human MLIC5AC contains nine CysD domains, thus, similar mechanisms could take place also in MLIC5AC to produce a highly crosslinked mucin. However, although all the CysD domains have a common structure, there is a large variation in especially surface amino acids opening for variation in the interactions.

[0112] SNPs affecting the respiratory gel-forming mucins have been previously associated with disease (Seibold et al., 2011 ; Shrine et al., 2019; Sabo et al., 2023). However, mucin sequencing has proved to be challenging due to their highly repetitive sequences in the PTS domain. Mucin sequences and SNPs have traditionally been poorly annotated, especially in these domains and flanking regions, including the VWD3 assembly. Here we observed that two frequent missense SNPs in the MUC5AC-D3 assembly, rs36189285 and rs878913005, are located in the interaction surfaces involved in the MLIC5AC tetramerization. The rs36189285 SNP corresponds to Arg 996G In , whi le rs878913005 corresponds to Arg 1201 T rp, both substituting the for MLIC5AC typical arginines. Arg996Gln is present in 32.9% of East Asian population and in 0.3% of European population in the Genome Aggregation Database (gnomAD), whereas Arg1201 Trp is present in 18.6% of the European population and only 0.6% in the Asian population. A high prevalence of double mutants is not observed except in Finnish population, where both mutants are present independently in 10.9% and 23.9% of the population.

[0113] The two SNPs (Arg996Gln and Arg 1201 Trp) can be involved in the closed-open conformation equilibrium and the described non-covalent oligomerization. However, no wide-ranging conformation differences were observed in the dimeric D3 mutants (Figure 4B), although minor alterations could be overlooked due to the low resolution in the flexible open conformation. Anyhow, the two mutations are likely to affect the molecular dynamics rather than the structure.

[0114] The mutation Arg996Gln, frequent in Asian population, affects the distance and the dynamics between the TIL3 domain and the VWD3, directly interfering in the conformational change and oligomerization. The open conformation was not detected in the D’D3CysD1 Arg99Gln cryoEM preparation during standard particle picking. Only template picking based on the WT open structure revealed a few particles. This observation, together with the described reduction of the distance between VWD3 and TIL3 domains, suggest that this mutation stabilizes the closed conformation and reduces the steric repulsion between Arg996 and TIL3. However, the VWD3 Arg996Gln mutant showed a higher occurrence of high order oligomers with an open conformation (Figure 4C), supporting that the opening conformation can still occur. When it forms, the absence of the positive charge residue near the newly exposed hydrophobic interaction face can modulate its properties. On the other hand, the Arg996Gln mutation produces a dramatic conformational change in the standard closed conformation tetramer. The His-tagged Arg996Gln mutant showed a high proportion of octameric particles in cryoEM that were essential for solving the high-resolution structure of the non-covalent interaction. The TIL3-TIL3 interaction is flexible in all cases, but the Arg99Gln mutation could add more freedom to allow the octameric arrangement. This mutation probably also increases the interaction affinity and therefore stabilizes the TIL interface, consequently also the second TIL3 domain in each disulfide-bonded dimer (Figure 4D) appears involved in non-covalent interactions making the formation of octamers possible. We could not find any significant disease association of the Arg996Gln mutation as the mutation has a low frequency in the used database (UK-Biobank).

[0115] The Arg1201 Trp SNP, frequent in European population, shows a double role in the conformation alteration of the VWD3 assembly. In the closed conformation it establishes a novel interaction between the TIL3 domain and VWD3, and in the open conformation removes the possibility of the predicted salt bridge between Arg1201 and Glu981. By that we can speculate that this mutation should favor the closed conformation, even though open conformation molecules were still observed in cryoEM. The role of Arg1201 Trp in MUC5AC VWD3 tetramerization is not fully understood. Arg 1201 is located in the interaction site but not directly involved in any contact. However, it is in close proximity to Asp1195 and the limitations in resolution and flexibility of the interaction make it not possible to disregard the importance of the 1201 amino acid in the formation of the interface. On the other hand, the presence of the tryptophan interacting with VWD3 could affect the TIL3 flexibility that is important for the interaction. The double mutant shows that once Arg996 is mutated, Trp1201 stabilizes the interface bringing TIL3 closer to VWD3. In theory Arg1201 Trp should limit the TIL3 flexibility and have the opposite effect of Arg996Gln. Nevertheless, we found a significant association between the Arg1201 Trp SNP with IPF and COPD supporting that the mutation could favor tetramerization and thus net- like formation. This mutation has also recently been found to be associated with dysregulated inflammatory responses across keratoconic cone where MUC5AC mucin plays a protective role (Jaskiewicz et al., 2023). Arg996Gln and Arg1201Trp have, thus, the same effects in the closed-open conformation equilibrium, stabilizing the closed form. In the same way, the effect of both mutations in tetramerization seems to be equivalent even if the outcome has not been completely elucidated. That the Arg 1201 Trp SNP promotes tetramerization suggests that it should increase mucus crosslinking. The amount of MLIC5AC is known to be increased in lung disease and especially COPD and its role in trapping bacteria could be increased. MLIC5AC is known to anchor the MLIC5B bundled strands to the surface goblet cells and a more cross-linked MLIC5AC might increase the attachment and retard clearance. This should be especially deleterious in the tiny airways affected by IPF. That there is an association of the Arg 1201 Trp SNP with COPD suggests a close association between the MLIC5AC mucin properties and the risk of COPD.

[0116] The strong association of the MLIC5B promotor SNP (rs35705950) leading to increased levels of MLIC5B protein and IPF has been known for some time (Seibold et al., 2011). We observed a strong association between the MLIC5AC Arg 1201 Trp SNP and IPF. The fact that this has been overlooked might at first be a surprise, but this is explained by poor sequence coverage close to the repetitive and polymorphic parts of the MUC5AC gene. However, interestingly it is when both SNPs appear at the same time that the MLIC5B SNP shows a strong association with IPF. It is, thus, likely that the strong association of the MLIC5B SNP (rs35705950) with IPF is caused by combination with the MLIC5AC SNP (rs878913005).

[0117] In the normal lung, the small airways without submucosal glands are characterized by surface goblet cells producing only MLIC5B mucin. However, recent observations suggest that the MLIC5AC mucin is increased in peripheral airways in IPF. The reason for the association of the MLIC5B promotor SNP with IPF has been suggested to be due to increased levels of the MLIC5B mucin. Although this is a likely link to IPF, an association of IPF susceptibility with a more cross-linked MLIC5AC is easier to understand as MLIC5AC is known to attach mucus bundles in the upper airways A more cross-linked MLIC5AC might be more attached and, thus, difficult to remove from the thin peripheral airways. This might lead to mucus retardation and poor bacterial clearance. Previous observations that show the importance of MLIC5B for IPF has been puzzling as the MLIC5B promoter polymorphism was not associated with IPF in Asian population. However, this might be easier to explain by the very low abundance of the MLIC5AC Arg1201 Trp SNP in this population. This further supports the conclusion that it is the combination of the MLIC5AC and MLIC5B SNPs that is the required driver for IPF susceptibility.

[0118] METHODS

[0119] Production and purification of MUC5AC plasmids The recombinant MUC5AC-D’-D3-CysD (WT and Arg996Gln), MUC5AC-D3-CysD (WT) and MUC5AC-D3 (WT and Arg996Gln) (GenBank accession number NM_001304359, residues 800-1481, 901-1481 , 901- 1366) plasmids were expressed with an N-terminal Hisx6 tag and a C-terminal Myc tag using the mammalian episomal expression vector pCEP-His. MUC5AC-D3 (WT, Arg996Gln, Arg1201Trp and Arg996Gln Arg 1201 Trp) was also expressed without any tag using the same expression vector.

[0120] CHO-Lec 3.2.8.1 -S cells (Nilsson et al., 2014) were grown in 300 ml Freestyle™ CHO with 8 mM L-glutamine in an Erlenmeyer flask in 5% CO2. Transfection with NovaCHOice transfection kit (Merck, Nottingham, UK) was performed according to the manufacturer’s instructions. Four hours after transfection the temperature was decreased to 31 °C. The supernatant was harvested after 48h by centrifugation for 10 min at 200 x g at room temperature (20-23.5°C) and then immediately dialyzed against phosphate-buffered saline (PBS) 10 mM imidazole (His-tagged proteins) or 20 mM Tris pH 8 (His-tag free proteins) at 4°C. Several MUC5AC-D3 batches of 300 ml each were made in CHO-Lec 3.2.8.1-S.

[0121] MUC5AC was filtered (Durapore® Membrane Filter, 0.22 pm GVWP, Millipore) and further purified using an AKTA purifier (GE Healthcare). The His-tag containing proteins were loaded onto a HiTrap chelating HP nickel affinity 1-ml column (GE Healthcare). The bound components were eluted with a gradient of 10-300 mM imidazole in 20 mM Tris pH 7.4 and 150 mM NaCI. The His-tag free MUC5AC-D3 variants were loaded onto a HiPrep Q HP 16 / 10 anion exchange chromatography column (GE Healthcare) and eluted in a 0-500 mM NaCI gradient in 20 mM Tris pH 8.

[0122] The protein containing fractions were dialyzed against 20 mM Tris (pH 7.4) and 50 mM NaCI, loaded onto a Mono Q™ HR 10 / 10 anion exchange column (GE Healthcare) and eluted in a linear gradient from 50 to 500 mM NaCI. It was followed by size exclusion fractionation on a Superose 6 10 / 300 column (GE Healthcare) eluted in 20 mM Tris pH 7.4, 50 mM NaCI and 10 mM CaCh and collected in fractions of 0.5 ml.

[0123] Protein purity was checked by SDS-PAGE using 4-15% Mini-PROTEAN® TGX™ precast protein gels (BioRad). Proteins were diluted in 2x SDS-PAGE loading buffer reaching a final concentration of 100 mM DTT or in DTT-free buffer, heated 5 minutes at 95 °C and loaded into the gel together with the Precision Plus Protein Unstained Standard (Bio-Rad). Gels were stained with Coomassie brilliant blue G-250. Single particle cryoEM

[0124] In order to optimize the sample homogeneity only the central fraction of the size exclusion was used for all recombinant proteins sample preparation (MUC5AC-D’-D3-CysD1 (WT and Arg996Gln), MUC5AC-6xHis-D3 (WT and Arg996Gln) and MUC5AC-D3 (Arg1201 Trp and Arg996Gln-Arg1201 Trp). The protein concentration was adjusted to 0.6 M in 20 mM Tris pH 7.4, 150 mM NaCI and 10 mM CaCh.

[0125] The samples were loaded onto UltrAuFoil R1.2 / 1.3 300# (SPT Labtech) holey gold grids (D3 Arg1201 Trp, D3 Arg996Gln-Arg1201 Trp, 6xHis-D3 Arg996Gln and 6xHis-D3 WT) or Quantifoil Cu 1.2 / 1.3 300# (SPT Labtech) copper grids (D’-D3-CysD1 WT and D’-D3-CysD1 Arg996Gln) that were previously glow discharged at 15 mA for 40 s with a negative charge. The grids were plunge frozen using a Vitrobot Mark IV (ThermoFisher) set at 100% humidity and 4°C. 6xHis-D3 Arg996Gln data collections were performed in a Titan Krios microscope (Thermo Fisher) at 0.83 A / pix and 300 kV acceleration voltage using a K2 Summit 4k x 4k detector (Gatan). 40 frames per movie were collected with an average electron dose per image of 1.15 e / A2and nominal defocus between -0.5 and -3.5 m in 0.3 steps. The data from five grids collected using the same conditions were merged for reconstruction. The datasets from the other five recombinant proteins were collected using EPU software (Thermo Fisher Scientific) in Aberration-free image shift (AFIS) mode at 0.86 A / pix and 300 kV acceleration voltage and a K3 Summit 6k x 4k detector (Gatan). 40 frames per movie were collected with an average electron dose per image of 1.25 (6xHis-D3 WT), 1.26 (D3 Arg996Gln Arg1201Trp), 1.29 (D’-D3- CysD1 WT and D’-D3-CysD1 Arg996Gln) or 1.49 e / A2(D3 Arg1201 Trp) and defocus between -0.5 and -2.5 pm (D’-D3-CysD1 WT, 6xHis-D3 WT and D3 Arg996Gln Arg1201 Trp) or -0.5 and -3.0 pm (D’-D3-CysD1 Arg996Gln and D3 Arg 1201 Trp). The micrographs were imported in cryoSPARC v.4.2.1 (Punjani etal., 2017) and patch motion corrected. CTF correction was performed in cryoSPARC using Patch CTF. The micrographs were manually curated and outliers were removed.

[0126] D’-D3-CysD1 WT. A first round of manual picking was performed. The particles were 2D classified and the bests classes were used for template-based automatic particle picking. The particles were extracted using a 256 pixel box size and 2D classified. The junk particles were removed and the rest were used for ab-initio reconstruction (2 classes) and heterogeneous refinement. A fraction of the particles from the best locking class was used for Topaz particle picking training (Bepler ef al., 2019). The Topaz picking model was applied to the curated dataset and the particles were extracted using a 256 pixel box size and 2D classified for junk particle removal. The particles were used for ab-initio reconstruction (4 classes) and heterogeneous refinement. A fraction of the particles from the best locking class was used for a new Topaz particle picking training. A new particle extraction was performed and the previous steps were repeated generating new 4 ab- initio classes that were used again in heterogeneous refinement. The heterogeneous refinement classes were used as templates for non-uniform refinement using C1 and C2 symmetries. The particles from the highest resolution class in the non-uniform refinement with C2 symmetry were further 3D classified by a new ab-initio reconstruction round in 2 classes and heterogeneous refinement. The best quality class was non-uniform refined imposing C2 symmetry that was used as an input for a final local refinement also imposing C2 symmetry.

[0127] D’-D3-CysD1 Arq996Gln. The same protocol as for D’-D3-CysD1 WT Dimers was applied until the first Topaz particle picking and extraction. Then the particles were directly classified in 2 ab-initio classes. The particles from the best class were 2D classified and only the high-resolution classes were selected. These particles were used fora new Topaz training and the process was repeated. The particles were used in a new ab-initio reconstruction round in 2 classes and heterogeneous refinement. The best quality class was non-uniformly refined imposing C2 symmetry that was used as an input for a global CTF refinement and a final local refinement also imposing C2 symmetry. To look for the presence of open dimers in the sample, D’-D3-CysD1 WT open dimer templates were used for template-based automatic particle picking. Particles were extracted using a 256 pixel box size and 2D classified for junk particle removal. The remaining particles were 3D classified by ab-initio reconstruction in 2 classes and heterogeneous refinement, and further 2D classified.

[0128] D3 Arq1201 Trp. An initial blob picking was performed using 1,000 micrographs setting a minimum particle diameter of 70 A and maximum of 140 A. The particles were 2D classified and the bests classes were used for template-based automatic particle picking. The particles were extracted using a 256 pixel box size and 2D classified. The junk particles were removed and the rest were used tor ab-initio reconstruction (2 classes) and heterogeneous refinement. A fraction of the particles from the best locking class was used for Topaz particle picking training. The Topaz picking model was applied to the curated dataset and the particles were extracted using 256 pixels box size and 2D classified for junk particle removal. Particles were divided in two ab-initio classes. A fraction of the particles from the best locking class was used for a new Topaz particle picking training. A new particle extraction was performed and the previous steps were repeated generating new 2 ab- initio classes that were used in heterogeneous refinement. The particles from the highest quality class were 3D classified again by ab-initio reconstruction in two classes and heterogeneous refinement. The best class was non-uniform refined imposing C2 symmetry that was used as an input for a global CTF refinement and a final local refinement also imposing C2 symmetry.

[0129] D3 Arq996Gln Arq1201 Trp. The same protocol than in D’-D3-CysD1 WT Dimers was applied changing the ab-initio and heterogeneous refinement with 4 classes for 2 classes and skipping the further repetition.

[0130] 6xHis-D3 Arq996Gln. The same protocol as for D’-D3-CysD1 R996 Dimers, but using a 450 pixel box size, was applied until the ab-initio reconstruction following the first Topaz particle picking and extraction step. Both ab-initio classes, one representing the tetrameric conformation and the other the octameric, were further 2D classified to remove junk particles and used independently in a second Topaz training. Both models were used for particle picking and two group of particles were newly extracted. The particles based in the tetrameric conformation training were 3D ab-initio reconstructed in 2 classes and further heterogeneously refined. The particles from the heterogeneous refinement class corresponding to a tetramer were 3D ab-initio reconstructed in 2 classes and further heterogeneously refined again. The best quality class was non- uniformly refined. Finally, the non-uniform refinement was used as an input for a final local refinement. In parallel, the particles based in the octameric conformation training were also 3D ab-initio reconstructed in 2 classes and further heterogeneously refined. The particles from the heterogeneous refinement class corresponding to an octamer were 3D ab-initio reconstructed in 2 classes and further heterogeneously refined again. The best quality class was non-uniformly refined. The 3D volume was used to create two masks in Chimera (Pettersen et al., 2004), one including the best defined tetramer in the octameric assembly (A) and the inverted mask (B). The masks ware newly imported in cryoSPARC and the mask A was dilated 6 A and soft padded 8 A. The particles from the non-uniform refinement were 3D classified in 4 classes using the non- uniformly refinement mask as a solvent mask and the mask A as focused mask. All four classes were further non-uniform refined. The particles and volume from the best resolution class were used for particle subtraction applying mask B. The newly generated particles together with the non-uniform refinement volume and the mask A were used as inputs in a last local refinement step.

[0131] D’-D3-CysD1 WT, D’-D3-CysD1 Arq996Gln, D3 Arq1201 Trp, D3 Arq996Gln and 6xHis-D3 WT Tetramers: Manual and template picking was performed as for most of the dimers but using a bigger box size, 320, 450 or 512 pixels. Particles were 2D classified and the best classes were used for Topaz particle picking. Particles were extracted using the same box size and further 2D classified. Model building

[0132] An initial model of MUC5AC-D3 Arg996Gln dimer was built in AlphaFold2 (Jumper et al., 2021). It was fitted into the density map with Molrep (Vagin and Teplyakov 1997) and manually built using Coot 0.9.8.1. The model was refined along different iterations using the real-space refinement tool in Phenix 1.20.1-4487 (Liebschner et al., 2019) and manual refinement in Coot. The final structure refinement validation was performed by Molprobity (Wiliams et al., 2018). The other models were built following the same protocol but using the final MUC5AC-D3 Arg996Gln dimer model (WT and Arg996Gln Arg 1201 Trp dimers, and Arg996Gln tetramer) or the WT dimer model (Arg 1201 Trp) as initial models. PyMol (Schrodinger, LLC (2015) The PyMOL Molecular Graphics System, version 2.5) and UCSF ChimeraX (Meng et al., 2023) were used for structure analysis and figures generation.

[0133] The MUC5AC-D3 open conformation model was generated manually fitting the MUC5AC-D3 WT closed conformation C8-3 and TIL3 dimer and two independent VWD3 domains into the low-resolution density assisted by the fit in map tool. The model building was finished using ISOLDE (Croll, 2018).

[0134] Human samples

[0135] Transbronchial cryo-biopsies were acquired during routine bronchoscopy for diagnosis of parenchymal / interstitial lung disease at the Sahlgrenska University Hospital, Gothenburg, Sweden. Briefly, a flexible cryoprobe was inserted through the working channel of the bronchoscope into the peripheral lung, the probe tip rapidly cooled by compressed gas, causing the surrounding tissue to freeze and adhere to the probe. The resulting cryo-biopsy will generally be larger than a traditional forceps biopsy and lack crush artifacts. The biopsy was extracted by removing bronchoscope and cryoprobe from the airway. The frozen biopsy still attached to the probe tip was submerged in saline to release the biopsy from the cryo probe and then fixed in 10% neutral buffered formalin, transferred to 70% ethanol and embedded in paraffin according to standard pathology procedures (ethical permission 543-11).

[0136] Stomach biopsies were acquired during routine endoscopies at the Sahlgrenska University Hospital (ethical permission 085-06). All subjects gave written informed consent. Samples were fixed directly in Carnoy's solution (composition 60% methanol, 30% chloroform and 10% glacial acetic acid). Lung samples from IPF patients were acquired during resection of diseased lungs at transplantation (ethical permission 2020-03693). Control lung samples were acquired from patients undergoing resection for lung adenocarcinoma (ethical permission 543-11). Pieces of lung tissue were fixed in 10% neutral buffered formalin according to routine procedures at the Sahlgrenska University Hospital. Fixed human tissue was transferred to 70% ethanol and stored until paraffin embedding, sectioning into 4 m sections and mounting on superfrost plus slides (Cat# 631-9483, VWR, Avantor, Radnor Township, PA). Researchers had access to medical records.

[0137] Histological staining

[0138] For standard histological evaluation, sections were stained with Harris hematoxylin (Cat# 01800, Histolab, Askim, Sweden) and Eosin Y (Cat# 01650, Histolab, Askim, Sweden) according to standard protocols. For visualization of connective tissue, a Masson-Goldner staining kit was used (Cat# 1.00485, Sigma- Aldrich / Merck) according to the manufacturer's instructions. Nuclei were stained with Weigert's iron hematoxylin (Cat# 1,15973, Sigma- Aldrich / Merck, Darmstadt, Germany). Sections were mounted with borosilicate cover glass (Cat# 831-0137, VWR international, Radnor, PA) using DPX mounting medium (Cat# 06522, Sigma-Aldrich / Merck).

[0139] Immunofluorescent staining of histological sections

[0140] Sections were baked on the slides at 60°C for 2 hours, dewaxed using xylene and hydrated. Antigen heat- induced epitope retrieval was performed with 10 mM citrate buffer pH 6.0 at 100°C for 20 min and then room temperature for another 20 min. Sections were washed in PBS, a barrier was drawn with an ImmEdge® Hydrophobic Barrier PAP Pen (H-4000, Vector laboratories, Newark, CA). Unspecific epitopes were blocked with 3% donkey serum in Tris-buffered saline and sections were permeabilized with 0.1 % Triton X-100. Staining was performed with sequential incubation with custom made polyclonal rabbit anti-human MUC5B antibodies (1 :200) in block solution (Fakih et al., 2020) overnight at 4° C and monoclonal mouse anti-human MUC5AC (Lidell et al., 2008) (1 :200) in block solution over night at 4°C (Cat# ab3649, Abeam, Cambridge, UK, RRID:AB_2146844) and purified goat anti-human EpCAM / TROP-1 (1 :200) in block solution over night at 4°C (Cat# AF960, R and D Systems / Biotechne, Minneapolis, MN, RRID:AB_355745). Secondary antibody was polyclonal donkey anti-rabbit Alexa Fluor 488 (Cal# A-21206, Thermo Fisher Scientific, Waltham, MA, RRID:AB_2535792), polyclonal donkey anti-rabbit Alexa Fluor 555 (Cat# A-31572, Thermo Fisher Scientific, RRID:AB_162543), and polyclonal donkey anti-mouse Alexa Fluor 647 (Cat# A-31571 , Thermo Fisher Scientific, RRID:AB_162542). Nuclei were stained with Hoechts 34580 (Cat# 565877, BD Biosciences, Franklin Lakes, NJ, RRID:AB_2869723). Sections were mounted using Prolong™ Gold Antifade Mountant (Cat# P36930, Thermo Fischer Scientific) and borosilicate glass coverslips (Cat# 831-0137, VWR international, Radnor, PA).

[0141] Imaging

[0142] Sections stained with histological stains (H&E, Masson-Goldner) were scanned using a NanoZoomer-SQ digital slide scanner (Cat# C13140-01, Hamamatsu Photonics, Shizuoka, Japan, RRID:SCR_023763) and NDP.view2 Image viewing software (Hamamatsu Photonics, Shizuoka, Japan, RRID:SCR_025177). Epi- fluorescent images were acquired using a Nikon Eclipse Ci-E microscope (Nikon instruments, Tokyo, Japan), a Nikon microscope camera DS-Fi3 (RRID:SCR_018857) and NIS-Elements D software version 5.02.03 (Nikon instruments, RRID:SCR_014329). High resolution images were acquired using ZEISS ZEN Blue Microscopy Software (Carl Zeiss, Oberkochen, Germany, RRID:SCR_013672) on the ZEISS LSM900 with Airyscan 2 (Carl Zeiss, Oberkochen, Germany, RRID:SCR_022263) and processed with Imaris version 9.5 (Oxford Instruments, Abingdon, UK, RRID:SCR_007370).

[0143] Piglet airway tissue

[0144] Ethical permission for experiments involving newborn piglets (Sus scrota domesticus) were obtained from Regierungen von Oberbayern, Munich, Germany (AZ55.2-1 -54-2531 -78-07) and Jordbruksverket, Jonkoping, Sweden (Dnr 6.7.18-12708 / 2019). Tracheas were acquired from wild type piglets (Sus scrota domesticus). To induce birth, intramuscular administration of 0.175 mg Cloprostenol (Estrumate®, Intervet GmbH, Unterschleissheim, Germany), on gestation day 112-114. Within 24 h of birth, piglets were anesthetized by Ketamine (Ursotamin®, Serumwerk Bernburg, Germany) and Azaperone (Stresnil®, Elanco Animal Health, Bad Homburg, Germany) and killed by intracardial injection of Tanax® T61 euthanasia solution (Intervet GmbH, Unterschleissheim, Germany). Tracheas from the larynx and lungs were excised and the lung parenchyma removed under Perfadex® solution, pH 7.2 (XVIVO Perfusion, Gothenburg, Sweden) before the prepared airways including larynx, trachea and bronchi were transferred to a 50 ml tube with Perfadex® solution, pH 7.2 and shipped at 4°C overnight to Gothenburg.

[0145] Electron microscopy

[0146] Distal tracheal tissue (two to three cartilage rings in length) from newborn piglets were fixed in modified Karnovsky’s fixative (2% paraformaldehyde, 2.5% glutaraldehyde in 0.05 M sodium cacodylate buffer, pH 7.2) for 24 h at 4°C. Postfixation was performed in 1% OsO4 at 4°C three times with intervening 1 % thiocarbohydrazide steps. The samples were dehydrated with increasing concentrations of ethanol followed by hexamethyldisilazane that was allowed to evaporate. Samples were mounted on aluminum specimen pin stubs (Cat# AGG301, Agar Scientific, Stansted, Essex, UK) with carbon tabs (Cat# AGG3347N, Agar Scientific, Stansted, Essex, UK) and conductive silver paint (Cat# 16040-30, Ted Pella, Redding, CA). To decrease charging, samples were sputter-coated with palladium before imaging at 3 kV in a field emission scanning electron microscope (Zeiss DSM 982 Gemini, Carl Zeiss, Oberkochen, Germany).

[0147] UK-biobank

[0148] Two cohorts were defined using the UK-biobank Research Analysis Platform (DNAnexus), patients diagnosed with COPD (n= 1 , 139) and patients diagnosed with IPF (n=107). A randomized control group (n=5,196) was established with the rest of patients not included in these cohorts. The genetic information regarding chromosome 11 contained in the available variant call files (VCFs) was extracted (COPD n= 1,139, IPF n= 107 and control n=5,196). Annotations in GRCh38 position 1177533 (rs36189285, MUC5AC Arg996Gln, nucleotide G>A), 1180138 (rs878913005, MUC5AC Arg1201 Trp, nucleotide OT) and 1219991 (rs35705950, MUC5B promotor, nucleotide G>T, A) were searched and participants were classified as SNP carriers (heterozygous and homozygous) or WT (wild-type or not annotated). The database used was from August 30, 2023. The diseases association statistical analysis was performed using Fisher’s exact test.

[0149] PCR detection of rs36189285, rs878913005 and rs 35705950 SNPs

[0150] Patient blood (5 ml) was collected, and the buffy coat was separated by centrifugation at 2500 x g for 10 minutes at room temperate. The buffy coat was transferred to a DNA free microcentrifuge tube and stored at -80°C. DNA was isolated from the buffy coat using QiAamp DNA mini kit according to manufacturer’s instructions. Following isolation DNA was quantified using a Nanodrop 2000 spectrophotometer (Thermo Fisher Scientific) and diluted to 10 ng / l in TE buffer. DNA was stored at -20°C until use.

[0151] Specific assays were designed to identify reference and alternate alleles for the rs36189285, rs878913005 and rs35705950 SNPs using the following primers: rs36189285 allele primer 1 : / rhAmp-FAM / GTGCCATACACCATCCGrGCAGA / GT1 / (SEQ ID NO: 2); rs36189285 allele primer 2: / rhAmp-Yakima Yellow / GTGCCATACACCATCCArGCAGA / GT1 / (SEQ ID NO: 3); rs36189285 locus primer: GCCTGAGGTTGATGAAGATGCTrGGTCT / GT1 / (SEQ ID NO: 4); rs878913005 allele primer 1 : / rhAmp-FAM / ACCTTCCAGGCCCCGrGACGT / GT3 / (SEQ ID NO: 5); rs878913005 allele primer 2: / rhAmp-Yakima Yellow / ACCTTCCAGGCCCCArGACGT / GT3 / (SEQ ID NO: 6); rs878913005 locus primer: GCAGCTCTGTTCTGCGACTACTArCAACC / GT3 / (SEQ ID NO: 7); rs35705950 allele primer 1 : / rhAmp-FAM / TTCCTTTATCTTCTGTTTTCAGCGrCCTTC / GT4 / (SEQ ID NO: 13); rs3570595 allele primer 2: / rhAmp-Yakima Yellow / CTCCTTTATCTTCTGTTTTCAGCGrCCTTC / GT4 / (SEQ ID NO: 14); rs3570595 locus primer: GCGTTTGCTCAGCGTCTTTCAArGAGTT / GT3 / (SEQ ID NO: 15).

[0152] Control plasmids were synthesized containing the DNA sequence -200 to +200 base pairs either side of the reference or alternate alleles for the rs36189285, rs878913005 and rs35705950 SNPs. Primers were synthesized by Integrated DNA Technologies and control plasmids were synthesized by Thermo Fisher Scientific.

[0153] PCR detection of rs36189285, rs878913005 and rs35705950 SNPs was performed using rhAmp™ SNP Genotyping (Integrated DNA Technologies) according to manufacturer’s protocol using a CFX 96 real-time PCR system (Bio-rad). The following cycling conditions were used; enzyme activation for 10 minutes at 95°C followed 40 cycles of 95°C for 10 seconds, 60°C for 30 seconds and 68°C for 20 seconds. Following each cycle emission of fluorescent signal was detected in the FAM and VIC / Yakima Yellow channels. Reactions contained 10 ng of template DNA or 500 copies of control plasmid. Amplification data was exported to Excel (Microsoft) and processed using Graphpad Prism (v 10.2.3).

[0154] MUC5B antibody

[0155] The MUC5B-D3 and MUC5AC-D3 assemblies were heterologous expressed in CHO-Lec 3.2.8.1-S cells and purified as described in (Fakih et al., 2020). The purified proteins were immobilized in agarose resin to create affinity purification chromatography columns using the AminoLink™ Immobilization Kit (Thermo Fisher Scientific) according to the manufacturer's instructions.

[0156] MLIC5B polyclonal antibodies were produced in rabbit as detailed in (Fakih et al., 2020). The antiserum was loaded onto a HiTrap Protein A HP 5ml column (GE Healthcare) equilibrated with 20 mM sodium phosphate pH 8 using a peristaltic pump. The antibodies were eluted in 0.1 M citric acid pH 3 and the pH was neutralized adding 1 M Tris-HCI pH 9. The Protein A purified serum was loaded onto the AminoLink-MUC5B column equilibrated with PBS pH 7.2 and incubated for 1 hour at room temperature. The column was washed with 6 CV of PBS and the antibodies bound to MLIC5B D3 were eluted in 4 CV of 0.1 M citric acid pH 3, neutralized with 1 M Tris-HCI pH 9, and loaded onto the equilibrated AminoLink-MUC5AC column. After 1 hour at room temperature the unbound material was washed with 6 CV of PBS pH 7.2.

[0157] Data availability

[0158] The structural data from cryoEM is deposited at PDB and EMDB with accession 8QTV and 18654 (D’-D3- CysD1 WT), 8QTB and 18648 (D’-D3-CysD1 Arg996Gln), 8R1 U and 18828 (D3 Arg1201 Trp), 8R1Z and 18829 (D3 Arg996Gln Arg1201 Trp) and 8QSP and 18638 (6xHis-D3 Arg996Gln).

[0159] The embodiments described above are to be understood as a few illustrative examples of the present invention. It will be understood by those skilled in the art that various modifications, combinations and changes may be made to the embodiments without departing from the scope of the present invention. In particular, different part solutions in the different embodiments can be combined in other configurations, where technically possible.

[0160] REFERENCES

[0161] Bepler.T., Morin, A., Rapp,M., Brasch.J., Shapiro, L., Noble, A.J., and Berger, B. (2019). Positive-unlabeled convolutional neural networks for particle picking in cryo-electron micrographs. Nat. Meth. 16, 1153-1160.

[0162] Carpenter, J., Wang,Y., Gupta, R., Li,Y., Haridass.P., Subramani.D.B., Reidel.B., Morton, L., Ridley, C., ONeal.W.K.et al. (2021). Assembly and organization of the N-terminal region of mucin MLIC5AC: Indications for structural and functional distinction from MLIC5B. PNAS 118, e2104490118.

[0163] Croll.T. (2018). ISOLDE: a physically realistic environment for model building into low-resolution electrondensity maps. Acta Cryst. D 74, 519-530. Fakih, D., Rodriguez Pineiro, A.M., Trillo-Muyo,S., Evans, C.M., Ermund.A., and Hansson.G.C. (2020). Normal murine respiratory tract has its mucus concentrated in clouds based on the Muc5b mucin. Am. J. Physiol. Lung Cell Mol. Physiol. 318, L1270-L1279.

[0164] Jaskiewicz.K., Maleszka-Kurpiel.M., Kabza.M., Karolak.JA, and Gajecka.M. (2023). Sequence variants contributing to dysregulated inflammatory responses across keratoconic cone surface in adolescent patients with keratoconus. Front. Immunol. 14, 1197054.

[0165] Jumper, J., Evans, R., Pritzel.A., Green, T., Figurnov.M., Ronneberger.O., Tunyasuvunakool.K., Bates, R., Zidek.A., otapenko.A.et al. (2021). Highly accurate protein structure prediction with AlphaFold. Nature 596, 583-589.

[0166] Lidell.M.E., Bara, J., and Hansson.G.C. (2008). Mapping of the 45M1 epitope to the C-terminal cysteine-rich part of the human MUC5AC mucin. FEBS J. 275, 481-489.

[0167] Liebschner.D., Afonine.P.V., Baker, M.L., Bunkoczi.G., Chen.V.B., Croll.T.I., Hintze, B., Hung.L.W., Jain,S., McCoy, A. J. et al. (2019). Macromolecular structure determination using X-rays, neutrons and electrons: recent developments in Phenix. Acta Cryst. D 75, 861-877.

[0168] Meng.E.C., Goddard, T.D., Pettersen.E.F., Couch, G.S., Pearson, Z.J., Morris, J. H., and Ferrin, T.E. (2023). UCSF ChimeraX: Tools for structure building and analysis. Protein Science 32, e4792.

[0169] Nilsson, H.E., Ambort,D., Backstrom.M., Thomsson.E., Koeck.P.J., Hansson.G.C., and Hebert, H. (2014). Intestinal MUC2 mucin supramolecular topology by packing and release resting on D3 domain assembly. J. Mol. Biol. 426, 2567-2579.

[0170] Pettersen.E.F., Goddard, T.D., Huang, C.C., Couch, G.S., Greenblatt,D.M., Meng.E.C., and Ferrin, T.E. (2004). UCSF Chimera-A visualization system for exploratory research and analysis. J. Comput. Chem. 25, 1605- 1612. Punjani.A., Rubinstein, J. L., Fleet, D. J., and Brubaker, M.A. (2017). cryoSPARC: algorithms for rapid unsupervised cryo-EM structure determination. Nat. Meth. 14, 290-296.

[0171] Sambrook, J., Russell. D. W.. Molecular Cloning: A Laboratory Manual, the third edition, Cold Spring Harbor Laboratory Press, Cold Spring Harbor. New York, 1.31-1.38, 2001.

[0172] Seibold, M.A., Wise,A.L., Speer, M.C., Steele, M.P., Brown, K.K., Loyd, J. E., FingerlinJ.E., Zhang, W., Gudmundsson, G., Groshong.S.D.et al. (2011). A Common MUC5B Promoter Polymorphism and Pulmonary Fibrosis. New England Journal of Medicine 364, 1503-1512.

[0173] Sharma. R.C., et al. "A rapid procedure for isolation of RNA-free genomic DNA from mammalian cells", BioTechniques, 14, 176-178 (1993).

[0174] Vagin, A. and Teplyakov.A. (1997). MOLREP: an Automated Program for Molecular Replacement. J. Appl. Cryst. 30, 1022-1025.

Claims

CLAIMS1 . A method for predicting lung disease predisposition of a human subject, the method comprising: determining, in a sample comprising nucleic acid molecules from the human subject, genotype of a single nucleotide polymorphism (SNP) rs878913005 located in the MUC5AC gene on chromosome 11 ; and predicting lung disease predisposition of the human subject based on the genotype of the SNP rs878913005.

2. The method according to claim 1, wherein determining genotype comprises determining, in the sample comprising nucleic acid molecules from the human subject, whether the human subject has the C allele or the T allele at position 1180138 on chromosome 11.

3. The method according to claim 2, wherein predicting lung predisposition comprises: predicting a high lung disease predisposition if the human subject has the T allele at position 1180138 on chromosome 11; and predicting a low lung disease predisposition if the human subject has the C allele at position 1180138 on chromosome 11.

4. The method according to any one of claims 1 to 3, wherein determining genotype comprises determining, in the sample comprising nucleic acid molecules from the human subject, whether a codon coding for amino acid number 1201 in the MUC5AC protein codes for arginine (R) or tryptophan (W).

5. The method according to claim 4, wherein predicting lung predisposition comprises: predicting a high lung disease predisposition if the codon coding for amino acid number 1201 in the MUC5AC protein codes for tryptophan (W); and predicting a low lung disease predisposition if the codon coding for amino acid number 1201 in the MUC5AC protein codes for arginine (R).

6. The method according to any one of claims 1 to 5, further comprising determining, in a sample comprising nucleic acid molecules from the human subject, genotype of a SNP rs35705950 located at the promoter of the MUC5B gene at chromosome 11, wherein predicting lung disease predisposition comprisespredicting lung disease predisposition of the human subject based on the genotype of the SNP rs878913005 and the SNP rs35705950.

7. The method according to claim 6, wherein determining genotype comprises determining, in the sample comprising nucleic acid molecules from the human subject, whether the human subject has the T or A allele or the G allele at position 1219991 on chromosome 11.

8. The method according to claim 7, wherein predicting lung predisposition comprises: predicting a high lung disease predisposition if the human subject has the T or A allele at position 1219991 on chromosome 11 and i) the codon coding for amino acid number 1201 in the MLIC5AC protein codes for tryptophan (W) and / or ii) if the human subject has the T allele at position 1180138 on chromosome 11; and predicting a low lung disease predisposition if the human subject has the G allele at position 1219991 on chromosome 11 and i) the codon coding for amino acid number 1201 in the MLIC5AC protein codes for arginine (R) and / or ii) if the human subject has the C allele at position 1180138 on chromosome 11.

9. A method for predicting lung disease predisposition of a human subject, the method comprising: extracting proteins from a sample from a human subject; determining presence of arginine (R) or tryptophan (W) at amino acid position 1201 in the MLIC5AC protein; and predicting lung disease predisposition of the human subject based on the determined presence or absence of arginine (R) or tryptophan (W) at amino acid position 1201 in the MLIC5AC protein.

10. The method according to claim 9, wherein predicting lung predisposition comprises: predicting a high lung disease predisposition if the amino acid position 1201 in the MLIC5AC is tryptophan (W); and predicting a low lung disease predisposition if the amino acid number 1201 in the MLIC5AC protein is arginine (R).11 . The method according to claim 9 or 10, further comprising determining, in a sample comprising nucleic acid molecules from the human subject, genotype of a single nucleotide polymorphism (SNP) rs35705950located at the promoter of the MUC5B gene at chromosome 11, wherein predicting lung disease predisposition comprises predicting lung disease predisposition of the human subject based on the determined presence or absence of arginine (R) or tryptophan (W) at amino acid position 1201 in the MLIC5AC protein and the SNP rs35705950.

12. The method according to claim 11, wherein determining genotype comprises determining, in the sample comprising nucleic acid molecules from the human subject, whether the human subject has the T or A allele or the G allele at position 1219991 of chromosome 11 .

13. The method according to claim 12, wherein predicting lung predisposition comprises: predicting a high lung disease predisposition if the human subject has the T or A allele at position 1219991 on chromosome 11 and the amino acid position 1201 in the MLIC5AC is tryptophan (W); and predicting a low lung disease predisposition if the human subject has the G allele at position 1219991 on chromosome 11 and the amino acid number 1201 in the MLIC5AC protein is arginine (R).

14. The method according to any one of claims 1 to 13, wherein the lung disease is selected from the group consisting of chronic obstructive pulmonary disease (COPD) and idiopathic pulmonary fibrosis (IPF).