Humanized nova1 splicing factor
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- THE ROCKEFELLER UNIV
- Filing Date
- 2026-02-02
- Publication Date
- 2026-08-06
Smart Images

Figure IMGF000001_0001 
Figure IMGF000002_0001 
Figure IMGF000002_0002
Abstract
Description
[0001] HUMANIZED NOVAI SPLICING FACTOR
[0002] STATEMENT OF GOVERNMENT RIGHTS
[0003]
[0001] This invention was made with government support under R35NS097404, R01 DCOI 8691, and R35GM127070, awarded by the National Institutes of Health. The government has certain rights in the invention.
[0004] INCORPORATION OF SEQUENCE LISTING
[0005]
[0002] A Sequence Listing conforming to the rules of WIPO Standard ST.26 is hereby incorporated by reference. Said Sequence Listing has been filed as an electronic document encoded as XML in UTF-8 text. The electronic document, created on February 1, 2026, is entitled “1119-84_PCT_ST26.xml”, and is 29,922 bytes in size.
[0006] FIELD OF THE INVENTION
[0007]
[0003] The present invention relates to a genetically modified non-human animal comprising in its genome a nucleic acid encoding a human neuro-oncological ventral antigenl (NOVAI) protein, or a portion thereof, particularly comprising the human NOVAI 1197 V variant, wherein amino acid 197 is a valine. Such genetically modified non-human animals that express human NOVAI protein, or a portion thereof, particularly comprising the human NOVAI 1197V variant, may be used as models for evaluating NOVAI in neural and motor functions, assessing NOVAl-based therapeutics, and in methods to alter, modulate, or evaluate vocalization.
[0008] BACKGROUND OF THE INVENTION
[0009]
[0004] Humans differ significantly from their closest living relatives, the great apes, particularly in their ability to communicate through complex learned vocal communication, a necessary' component of spoken language13. This complexity is driven by some anatomical adaptions of the vocal tract and intricate neural networks linking various brain regions3^ However, the genetic basis underlying these specialized human traits remains to be fully identified.
[0010]
[0005] The transcription factor forkhead box P2 (FOXP2) is a potential driver of human language function, as it harbors two amino acid substitutions present in human but not in chimpanzee and many other mammal genomes. Families with FOXP2 mutations exhibit severe speech defects ^46w]Tj ]eFOXP2 disruption in mice leads to vocalization abnormalities
[0011]
[0012] ’ suggesting a role in spoken language function. Studies on mice with the two amino acids substituted to the human version have reported vocal changes both in the neonatal and adult stages^^l While Hammerschmidt et al. observed minimalvocal changes, von Merten et al. reported qualitative changes under a more natural vocalization paradigm ’, suggesting the involvement of these two amino acids in vocalization. However, these substitutions arc also present in archaic humans, and comprehensive analyses using diverse human genome datasets have found no evidence of recent selection. This suggests that the FOXP2 substitutions occurred earlier than initially thought ’
[0013]
[0014] .
[0015]
[0006] Similarly, the TKTL1 gene contains a human-specific amino acid thought to influence greater neurogenesis in human than Neanderthal frontal cortex, though this finding is based on European ancestry genome datasets
[0016]
[0017] . Broader analyses of modem human genomes reveal that 0.03-0.2% of individuals possess the ‘putative Neanderthal variant’, indicating its presence in a significant portion of the population’ These findings underscore the importance of incorporating diverse human samples to identify and validate the genetic background of modern human traits through genomic comparisons.
[0018]
[0007] Genomic comparisons between archaic humans, ape genomes, and the broader human population have identified 61 human-specific nonsynonymous coding variants that are fixed or nearly fixed in modem humans13. One of the genes includes an isoleucine to valine substitution at position 197 (I197V) in the RNA binding protein neuro-oncological ventral antigenl (NOVAI). NOVAI is highly expressed in neurons of the central nervous system (CNS) in both mice and humans'^ and its expression has also been observed in cultured human and rat cells
[0019]
[0020] . NOVAI was first identified as an autoantigen targeted in the paraneoplastic neurologic disorder (PND) opsoclonus-myoclonus ataxia (POMA)^^. PNDs develop when tumor cells ectopically express proteins normally restricted to the nervous system, triggering an antitumor immune response that breaches the blood-brain barrier, leading to autoimmune neurologic disease ’. In POMA, a robust immune response is mounted against NOVAI and its paralog, N0VA2 70 The autoimmune disorder is characterized by motor dysfunction due to failure of inhibition of midbrain neurons, which results in the hyperactivity associated with opsoclonus-myoclonus ataxia 74 Noval is required for the survival of neurons in the spinal cord and brainstem after birth. In mice, homozygous deletion of the Noval gene results in an early postnatal lethal phenotype due to abnormal motor function^ ’. Therefore, NOVAI plays a crucial role in neural development and neuromuscular control in mammals.
[0021] 7 *'7
[0022]
[0008] NOVA proteins directly bind RNA in the mouse brain3--33to regulate pre-mRNA processing^ 1,34,35, translation-^^ and neurophysiology^. Genetic studies mapping NOVA target RNAs TO
[0023] in mice and humans have also linked it to autism. A human patient with a heterozygous deletion of NO VAI presented with delay of language development, learning disabilities, motor hyperactivity andbehavioral dysregulation3^. A number of neurological disorders feature or present with motor vocalization disability — including development disorders (non-verbal autism and others) and degenerative disorders (c.g. frontal dementias).
[0024]
[0009] Studies have explored the NOVAI I197V variant by reverting the ancestral isoleucine 197 variant back into human iPSC-derived organoids, revealing morphological and electrophysiological changes In vitro^. However, these effects were not observed in another study that reintroduced the same substitution q
[0025] in different iPSCs, therefore effects and changes are unclear and unverified.
[0026]
[0010] There is a need for understanding of human-specific gene and protein alterations that drive vocalization and speech. Further, animal models carrying human-specific variants directly linked to alterations in vocalization and speech or other relevant neural and motor functions have application in understanding and impacting vocal, neural and motor function.
[0027] [OU] The citation of references herein shall not be construed as an admission that such is prior art to the present invention.
[0028] SUMMARY OF THE INVENTION
[0029]
[0012] NOVAI is a neuronal RNA-binding protein expressed in the central nervous system and is essential for survival in mice and normal development in humans. Modern humans specifically share and carry a single amino acid change (I197V) in NOVACs second RNA binding domain. In an aspect of the invention, mice were generated carrying the human-specific II 97V variant (denoted Noval^l,^ni) and the molecular and behavioral consequences were assessed and evaluated. The 1197 V substitution had minimal impact on NOVAl’s RNA binding capacity. The I197V substitution led to specific effects on alternative splicing, and multiple binding peaks in mouse brain transcripts involved in vocalization were revealed. The 1197V substitution was associated with behavioral differences in vocalization patterns in Noval^u^umice as pups and adults. Thus, this human-specific NOVAI substitution may have been part of an ancient evolutionary selective sweep in a common ancestral population of Homo sapiens, linked to the development of spoken language via differential RNA regulation during brain development.
[0030]
[0013] In accordance with aspects of the invention, gene-editing is implemented to substitute the NOVAI isoleucine (I) isoform present in most mammals and archaic hominids (Neanderthals and Denisovans) with the human-specific valine (V) variant at position 197 in a non-human mammal (such as mice). Comparison of humanized NOVAI mice (Noval^w^11) with wild-type mice carrying the ancestral Noval gene ( ' v<27w? / w?) reveals specific transcriptomic and behavior differences related to vocalization. The humanspecific NOVAI amino acid 197 variant confers vocalization changes in humanized mice.
[0014] In a general aspect, the present invention provides genetically altered non-human cells and animals carrying the human-specific I197V variant (denoted Noval^u^u. In aspects, the non-human animals or cells arc non-human mammals, or arc derived from non-human mammals, altered in Noval.
[0031]
[0015] Provided herein are genetically modified non-human animals encoding a human neuro-oncological ventral antigen 1 (NOVA 1) protein, or a portion thereof, particularly comprising the human NOVAI 1197V variant, wherein amino acid 197 is a valine. Also provided herein are compositions and methods for generating and using such modified non-human animals. Provided herein are cells derived from genetically modified non-human animals encoding a human neuro-oncological ventral antigenl (NOVAI) protein, or a portion thereof, particularly comprising the human NOVAI II 97V variant, wherein amino acid 197 is a valine. Also provided herein are compositions and methods for generating and using such cells from modified non-human animals.
[0032]
[0016] Described herein are genetically engineered non-human animal genomes, engineered cells, and non-human animals comprising a human Noval gene, or a portion thereof, particularly comprising the human NOVA! I197V variant, wherein amino acid 197 is a valine. In some embodiments, genetically engineered animals express a human NOVAI protein from a desired locus (e.g., from an endogenous NOVAI segment). The non-human animal may be a mammal, such as a rodent (e.g., a mouse or a rat). The non-human animal cell can be a mammalian cell, such as a rodent cell (e.g., a mouse cell or a rat cell). The non-human animal genome can be a mammalian nucleic acid, such as a rodent nucleic acid ( e.g., a mouse nucleic acid or a rat nucleic acid).
[0033]
[0017] In some embodiments, a non-human animal, a non-human animal cell, or non-human animal genome comprises a nucleic acid sequence encoding a human Noval gene, or a portion thereof, particularly comprising the human NOVAI 1197V variant, wherein amino acid 197 is a valine.
[0034]
[0018] In some embodiments, the non-human animal, a non-human animal cell, or non-human animal genome further comprises a nucleic acid sequence encoding a human or humanized Nova2 gene. In some embodiments, the non-human animal, a non-human animal cell, or non-human animal genome further comprises a nucleic acid sequence encoding a human or humanized FoxP2 gene.
[0035]
[0019] In some embodiments, the non-human animal expresses only human NOVAI, or only NOVA 1 having valine at amino acid 197. The non-human animal may particularly be homozygous for human NOVAI. In some embodiments, the non-human animal is heterozygous and expresses human NOVAI, or only NOVA 1 having valine at amino acid 197, and also expresses a non-human NOVAL
[0036]
[0020] An exemplary mouse NOVAI protein is a follows:
[0037] Mouse NQVA1 protein 19 / Isoleucine (I) in bold (507 amino acids ) 1 MMAAAPIQQNGTHTGVPIDLDPPDSRKRPLEAPPEAGSTKRTNTGEDGQYFLKVLIPSYA 61 AGSIIGKGGQTIVQLQKETGATIKLSKSKDFYPGTTERVCLIQGTIEALNAVHGFIAEKI121 REMFQNVAKTE PVSILQPQTTVNPDRI KQTLPSS PTTTKSS PSDPMTTSRANQVKI I VPN 181 STAGLIIGKGGATVKAIMEQSGAWVQLSQKPDGINLQERVVTVSGEPEQNRKAVELI IQK 241 IQEDPQSGSCLNI SYANVTGPVANSNPTGSPYANTAEVLPTAAAAAGLLGHANLAGVAAF 301 PAVLSGFTGNDLVAI TSALNTLASYGYNLNTLGLGLSQAAATGALAAAAAS N PAAAAAN 361 LLATYASEASASGSTAGGTAGTFALGSLAAATAATNGYFGAASPLAAS ILGTEKSTDGS 421 KDWEIAVPENLVGAILGKGGKTLVEYQELTGARIQISKKGEFVPGTRNRKVTITGTPAA 481 TQAAQYLITQRITYEQGVRAA PQKVG ( SEQ ID NO: 1)
[0038]
[0021] An exemplary human NOVAI protein is a follows:
[0039] Human NOVAI protein - human specific 197 (V) Valine in bold ( 507 amino acids )
[0040] 1 MMAAAPIQQNGTHTGVPIDLDPPDSRKRPLEAPPEAGSTKRTNTGEDGQYFLKVLIPSYA 61 AGSIIGKGGQTIVQLQKETGATIKLSKSKDFYPGTTERVCLIQGTVEALNAVHGFI EKI 121 REMFQNVAKTE PVS I LQPQTTVNPDRI KQTLPS S PTTTKSS PS DPMTTSRANQVKI I VPN 181 STAGLIIGKGGATVKAVMEQSGAWVQLSQKPDGINLQERVVTVSGEPEQNRKAVELIIQK 241 IQEDPQSGSCLNI SYANVTGPVANSNPTGSPYANTAEVLPTAAAAAGLLGHANLAGVAAF 301 PAVLSGFTGNDLVAI TSALNTLASYGYNLNTLGLGLSQAAATGALAAAAAS PAAAAAN 361 LLATYASEASASGSTAGGTAGTFALGSLAAATAATNGYFGAASPLAASAILGTEKSTDGS 421 KDWEIAVPENLVGAILGKGGKTLVEYQELTGARIQISKKGEFVPGTRNRKVTITGTPAA 481 TQAAQYLITQRITYEQGVRAANPQKVG ( SEQ ID NO: 2 )
[0041]
[0022] Human NOVAI or NOVAI comprising the human NOVAI I197V variant can be expressed in the non-human animal via modes and approaches known and available to one skilled in the art. Human NOVAI or NOVAI comprising the human NOVAI I197V variant can be expressed in the non-human animal via direct alteration of its genome. In one such embodiment, gene editing, such as CRISPR / Cas9-based gene editing, is utilized to introduce nucleotide changes that result in human-type NOVAI animals (such as knock-in mice). Alternative approaches may utilize homologous recombination or gene replacement techniques to replace the native animal genome nucleic acid encoding NOVI, or relevant portion thereof, with human or humanized NOVAI.
[0042]
[0023] In some embodiments, the nucleic acid sequence encoding a human NOVAI protein or portion thereof comprising the human NOVAI 1197V variant is incorporated into endogenous NOVAI locus (of the genome, cell, or non-human animal). In some embodiments, the nucleic acid sequence encoding a human NOVAI protein or portion thereof comprising the human NOVAI 1197V variant replaces (at an endogenous locus of the non-human animal genome, non-human animal cell, or non-human animal) an orthologous endogenous nucleic acid sequence encoding an endogenous non-human NOVAI protein or a portion thereof. In some embodiments, the endogenous NOVAI locus comprises a heterozygous or homozygous replacement of an endogenous nucleic acid sequence encoding an endogenous NOVAI protein or a portion thereof with the nucleic acid sequence encoding a human NOVAI protein or portion thereof comprising the human NOVAI II 97V variant.
[0024] In some embodiments, the human NOVAI protein or portion thereof comprising the human NOVAI 1197V variant comprises an amino acid sequence of a human NOVAI protein or portion thereof. In some embodiments, the human NOVAI protein or portion thereof comprising the human NOVAI I197V variant comprises the human NOVAI sequence as follows, or a portion thereof including valine 197:
[0043]
[0025] Human NOVAI protein - human specific 197 (V) Valine in bold (507 amino acids)
[0044] i MMAAAPIQQNGTHTGVPIDLDPPDSRKRPLEAPPEAGSTKRTNTGEDGQYFLB; VLIPSYA 61 AGSIIGKGGQTIVQLQKETGATIKLSKSKDFYPGTTERVCLIQGTVEALNAVHGFIAEKI 121 REMPQNVAKTE PVS I LQPQTTVNPDRIKQTLPS S PTTTKSS PS DPMTTSRANQVKI I VPN 181 STAGLI IGKGGATVKAVMEQSGAWVQLSQKPDGINLQERWTVSGEPEQNRKAVELI IQK 241 IQEDPQSGSCLNISYANVTGPVANSNPTGSPYANTAEVLPTAAAAAGLLGHANLAGVAAF 301 PAVLSGFTGNDLVAITSALNTLASYGYNLNTLGLGLSQAAATGALAAAAASANPAAAAAN 361 LL TYASEAS SGSTAGGTAGTFALGSLAAATAATNGYFGAASPLAASAILGTEKSTDGS 421 KDWEIAVPENLVGAILGKGGKTLVEYQELTGARIQISKKGEFVPGTRNRKVTITGTPAA 481 TQAAQYLITQRITYEQGVRAANPQKVG (SEQ ID NO: 2 )
[0045]
[0026] Also described herein is a chimeric nucleic acid molecule that encodes a functional NOVAI protein comprising a nucleic acid sequence of a modified non-human animal NOVAI gene that encodes a nonhuman NOVAI protein or portion thereof, wherein the modified non-human animal NOVAI gene comprises a replacement of a nucleic sequence encoding a portion of the non-human animal NOVAI protein with a nucleic acid sequence encoding a human NOVAI protein or portion thereof comprising the humanNOVAl I197V variant. In some embodiments, a chimeric nucleic acid molecule as described herein comprises a nucleic acid sequence of a non-human animal NOVAI gene that (a) encodes a NOVAI protein and (b) is modified to comprise a replacement of a sequence encoding the NOVA 1 protein or portion thereof w’ith a sequence encoding a human NOVAI protein or portion thereof comprising the human NOVAI 1197V variant, wherein the chimeric nucleic acid molecule encodes a functional NOVAI protein, and optionally, wherein the chimeric nucleic acid sequence further comprises promoter andor regulator ' sequences of the non-human animal NOVAI gene. In some embodiments the modified NOVAI gene further comprises a drug selection cassette. In some embodiments, a chimeric nucleic acid molecule described herein further comprises (i) a 5' homology arm upstream of the modified non-human animal NOVAI gene and (ii) a 3' homology aim downstream of the modified non-human animal NOVAI gene. In some cases, the 5' homology arm and 3' homology arm can undergo homologous recombination with a non-human animal NOV Al locus of interest, and wherein following homologous recombination with the non-human animal NOVA! locus of interest, the modified NOV Al gene can replace the non-human animal NOVAI gene at the non-human animal NOV Al locus of interest and is operably linked to an endogenous promoter that drives expression of the non-human animal NOVAI gene at the non-human animal NOVAI locus of interest.
[0027] Also described are methods of making a non-human animal, the non-human animal cell, or the non-human animal genome described herein by inserting the nucleic acid sequence encoding the human NOVAI protein or portion thereof comprising the human NOVAI I197V variant into the genome of the non-human animal, the genome of the non-human animal cell, or the non-human animal genome. In some embodiments, the non-human animal cell is a non-human animal embryonic stem (ES) cell, and wherein the inserting comprises inserting the nucleic acid sequence encoding the a human NOVA 1 protein or portion thereof comprising the human NOVAI I197V variant into the genome of the non-human animal ES cell to form a modified non-human animal ES cell comprising, in its genome, the nucleic acid sequence encoding the human NOVAI protein or portion thereof comprising the human NOVAI I197V variant. In some embodiments the method comprises introducing the modified non-human animal ES cell into host embryo cells in vitro. In some embodiments the method comprises gestating, in a suitable non-human surrogate mother animal, the host embryo cells comprising the modified non-human animal ES cell, and allowing the non-human surrogate mother animal to birth non-human animal progeny comprising a germ cell comprising the nucleic acid sequence encoding the human NOVAI protein or portion thereof comprising the human NOVAI 1197V variant. In some embodiments, the nucleic acid sequence encoding the human NOVAI protein or portion thereof comprising the human NOVAI I197V variant is inserted into an endogenous NOVAI locus. In such embodiments, the step of inserting comprises replacing an endogenous nucleic sequence encoding an endogenous NOVAI protein or portion thereof with the nucleic acid sequence encoding the human NOVAI protein or portion thereof comprising the human NOVAI 1197V variant, wherein the endogenous nucleic sequence encoding an endogenous NOVAI protein or portion thereof and the nucleic acid sequence encoding the human NOVAI protein or portion thereof comprising the human NOVAI I197V variant are orthologous.
[0046]
[0028] In some embodiments, the inserting of the nucleic acid comprises contacting the genome of the non-human animal, the genome of the non-human animal cell, or the non-human animal genome with any chimeric nucleic acid molecule of the disclosure.
[0047]
[0029] A non-human animal, non-human animal cell, or non-human animal genome can be made according to any method of the disclosure. In some embodiments, the non-human animal is a rodent, dog, horse, cattle, bird, a non-human primate such as a macaque or monkey or chimpanzee. In some embodiments, the non-human animal is a rodent. In some embodiments, the non-human animal is a rodent and is a mouse.
[0048]
[0030] In some embodiments, the non-human animal, non-human animal cell, or non-human animal genome com-prises a knockout mutation of an endogenous NOV Al gene. In some cases, the knockout mutation comprises a deletion of the NOV Al gene or a portion thereof. In specific embodiments, the knockout mutation comprises a deletion of the entire coding sequence of the NOVAI gene.
[0031] In some embodiments the disclosure provides a targeting vector comprising: (i) a 5’ homology arm and (ii) a 3’ homology arm, wherein the 5’ homology arm and 3' homology arm undergo homologous recombination with a non-human animal NOVAI locus of interest, and wherein following homologous recombination with the non-human animal NOVAI locus of interest, the targeting vector inserts a knockout mutation in the non-human animal NOVAI gene at the non-human animal NOVAI locus of interest.
[0049]
[0032] Methods are provided for altering or modifying vocalization or vocal communication in a non-human animal comprising expressing in a non-human animal the human NOVAI protein, or a portion thereof, particularly comprising the human NOVAI 1197V variant, wherein amino acid 197 is a valine. In some aspects of the method, the non-human animal is genetically altered to express the human NOVAI protein, or a portion thereof, particularly comprising the human NOVAI I197V variant, wherein amino acid 1 7 is a valine. In aspects, the non-human animal only expresses the human NOVA! protein, or a portion thereof, particularly comprising the human NOVAI II 97V variant, wherein amino acid 197 is a valine. In aspects the non-human animal is heterozygous and heterozygously expresses its native NOVAI protein and also the human NOVAI protein, or a portion thereof, particularly comprising the human NOVAI I197V variant, wherein amino acid 197 is a valine.
[0050]
[0033] Other objects and advantages will become apparent to those skilled in the art from a review of the ensuing detailed description, which proceeds with reference to the following illustrative drawings, and the attendant claims.
[0051] BRIEF DESCRIPTION OF THE DRAWINGS
[0052]
[0034] Figure 1A-1E: Evolutionarily conserved NOVAI harbors a modern human-specific amino acid at position 197. (a) Conservation analysis of NOVAI and NOVA2 protein conservation across species using high-quality genome assemblies from NCBI. The analysis was performed using NCBI’s Constraintbased Multiple Alignment Tool (Cobalt). “Conserved Region” is defined by the relative entropy threshold of the amino acid residue. Green bars indicate highly conserved positions. Gray-red bars indicate the Column Quality score: scores for amino acid residues based on agreement within the column / position. Rare residues are highlighted in darker red, while positions with any mismatch are anchored, (b) Sequence comparison of NOVA! CDS in 4 ancient humans (3 Neanderthals and 1 Denisovan; blue bar) and 8 modem humans: red bar. The upper panel shows the 50 bases around the 197^ amino acid and the lower shows the entire NOVAI CDS. (c) SNP frequency analysis of NOVAI gene in modem humans. Mean allele frequencies (MAFs) on the NOVAI gene detected in the genome analysis of 121,412 modern humans are indicated by gray dots; the upper bound of 95% confidence interval for SNPs in the NOVAI gene is indicated by red dotted lines, and for comparison, those for NOVA2 or all genes on chromosome 14 (w'here NOVAI gene is located) are indicated by blue- or black- dotted lines, respectively. The red star indicates the V197Ivariant: rs762662114, chrl4:26448894 / GRCh38.pl4. (d) Comparison of normalized Tajima’s D values. The first gene set includes NOVAI and NOVA2, and N0K4 / -neighboring genes (F0XG1 andSTXBP6) on chrl4. The second gene set includes all genes on chrl4 where X l7A / is located, (e) A model of the evolutionary timing of the 197^ amino acid change that occurred in the NOVAI gene, noting the mice generated in this study. Noval^u / ^umice have the modern human specific amino acid in NOVAI protein. The bottom panel shows the corresponding location within the KH2 domain of the NOVA 1 protein. Amino acids structurally proximal (<5 A) to the 197^ amino acid predicted by AlphaFold2 are boxed in. The KH2 domain sequence of Novalw, / wtANQVKI I VPNSTAGL 11 GKGGATVKAIMEQSGAWQLSQKP DG INLQERWTVSGEPEQNRKAVEL I IQK IQE (SEQ ID NO:3) and of Novalhu h‘ ANQVKIIVPNSTAGLIIGKGGATVKAVMEQSGAWQLSQKPDGINLQERWTVSGEPEQNRKAVELIIQK IQE (SEQ ID NO:4) are depicted.
[0053]
[0035] Figure 2A and 2B: (a) The SNP report from dbSNP database. The single nucleotide variation th
[0054] (rs762662114) responsible formodem human specific amino acid in NOVAI (197LnVai). Alternate allele frequencies in the total number of samples for each genome project are shown, (b) MAF analysis in the NOVAI gene across human ancestries. SNPs detected in the genome analysis of 121,410 modern humans on the NOVAI CDS, along with the NOVAI protein structure. Each SNP is color- coded by ethnicity. SNP corresponding to the 197^ amino acid (ancient human-type variant) is indicated by a star.
[0055]
[0036] Figure 3: Comparison of Normalized Tajima’s D values. The closer the normalized Tajima’s D value approaches - 1, the greater the likelihood that the gene has undergone strong purifying selection The first gene set includes NOVAI anANOVA2, and closest protein-coding M? T 47-neighboring genes (FOXGI and STXBP6) on chr!4. Each gray dot indicates a gene in each gene list. The normalized Tajima’s D value for NOVAI (red) is -0.9993. The second gene set includes all genes on chr.14 where NOVAI is located. These gene sets are also shown in Figure Id. The third and fourth gene sets are based on Meyer's report-, with genes related to the 260 human-specific single- nucleotide changes (SNCs) that cause fixed amino acid substitutions in well-defined human coding sequence, or to the subset of these genes (eight among the 260 SNCs) whose function is associated with brain function or nervous system development. The fifth gene set is based on Trujillo’s report, with genes associated with 61 autosomal fixed derived mutations in all humans compared to Neanderthal genomes and Denisovan genome. The sixth and seventh gene sets include human RNA binding proteins (RBPs) with KH domain or all annotated RBPs.
[0056]
[0037] Figure 4: Recent burst of coalescence of a sweep forNOVAl 197V. The left tree shows the NOVAI SNP, with the shortened branches and burst of relatively recent coalescence events leading to the modem humans (reflecting a rise in frequency of the derived allele). The right tree shows FRMD8 for contrast, withlonger branches and delayed coalescence: more typical of what would be expected in the absence of a selective sweep. The sampled ARGs included two Yoruba (HGDP00927, SS6004475), two Mbuti (SS6004471, HGDP00456), and two San (HGDP01029, SS6004473) individuals, as well as the Altai Neanderthal and Denisovan sequences and a chimpanzee outgroup (panTro4). The numbers following the samples represent the separate haplotypes from each individual. Red lines indicate derived allele.
[0057]
[0038] Figure 5A-5E: Generation of humanized Noval mice Novam,^ni). (a) Overview of the strategy to generate mice with a modern human-specific amino acid substitution in NOVAI protein. Using CRISPR / Cas9, nucleotide substitutions (that lead to a single amino acid change from Isoleucine to Valine) arc introduced. Two silent mutations were introduced for genotyping, (b) DNA sequencing of Noval allele in wild-type mice and in mice in which correct knock-in was introduced (
[0058]
[0059] Noval^11). (c) Genotyping of Noval^u^ll<mice using restriction enzymes. The introduction of silent mutations creates a Btsla recognition site. Noval^uallele is distinguished from wild-type Noval allele by restriction enzyme treatment after PCR. (d) The gRNA sequence TGCTACTGTGAAGGCTATAATGG (SEQ ID NO:5) and predicted off-target site information predicted by CRISPR direct (crispr.dbcls.jp). There are 10 potential off target loci with mismatches (chr2 TGAAGGCTATAACGG (SEQ ID NO:7); chr5 TG AAGGC TATAATGG (SEQ ID NO:8); chr6 TGAAGGCTATAAAGG (SEQ ID NOV); chr7 TGAAGGCTATAAAGG (SEQ ID NO: 10); chr8 TGAAGGCTATAAAGG (SEQ ID NO: II); chr9 TGAAGGCTATAAGGG (SEQ ID NO: 12); chr 17 TGAAGGCTATAAGGG (SEQ ID NO; 13); chrl CCTTTATAGCCTTCA (SEQ ID NO:14); chr3 CCTTTATAGCCTTCA (SEQ ID NO:15); chr!8 CCTTTATAGCCTTCA (SEQ ID NO: 16) outside of the PAM+12mer core sequences chr!2 CCATTATAGCCTTCA (SEQ ID NO: 6). (e) The genomic sequencing of the potential off target (POT) loci. zAlignment of each genotype and reference genome for the genomic sequence of 100 bases around the POTs are shown. Asterisks indicate identical nucleotides. All POT sites were identical between genotypes and the reference genome, with the target site (responsible for II 97V substitution) being the only detectable edits. The mouse reference and wt target site sequence NOVAI sequence corresponds to AGTTCCCAAC AGCACAGCAGGTCTGATAATAGGGAAGGGAGGTGCTACTGTGAAGGCTATAATGGAGCAGT CAGGGGCTTGGGTGCAGCTTTCCCAGAAACCCGATGGGATCAAC (SEQ ID NO:21) (mouse sequence regions around the Noval He amino acid 197 sequence underlined). The human sequence at the target site is AGTTCCCAACAGCACAGCAGGTCTGATAATAGGGAAGGGAGGTGCTACTGTGAAGGCAGTG ATGGAGCAGTCAGGGGCTTGGGTGCAGCTTTCCCAGAAACCCGATGGGATCAAC (SEQ ID NO:22) (mouse sequence regions around the Noval amino acid 197 sequence underlined and the valine VI 97 is shown in bold).[039| Figure 6A-6J: (a) Body weight comparison of wild type (Novalwt / wt, N=19), heterozygous (Noval^u,wt, N=12), and homozygous Noval^u^u, N=14) mice, measured from 2-12 weeks postnatal, (b) Brain-to-body weight ratio comparison in 3-week-old mice (N=8 per group), (c) Gene expression correlations between Novalwt,'wtand Noval^u'^uin midbrain at E18.5 and P21. Scatterplots show gene expression in average TPMs (log2 scale). The red dot marks one differentially expressed gene (Gkn3 p<0.()5, FDR<0.1), and the yellow dot indicates Naval. Pearson correlations are noted above the plots. El 8.5 midbrain: Novalwt,'wtN=6, Noval^u'^uN=6; P21 midbrain: Novalwt / w^ N=4, Noval^U / ^niN=4.
[0060] (d) Autoradiography images from N0VA1- CLIP of 3-week-old midbrain in Novalw / wtand Novalu'^umice. The yellow line marks the N0VA1 protein size, and the red outline highlights RNA extraction by sectioning, (e) Distribution of NOVA 1 CLIP peaks on the mouse genome, (f) The most enriched binding sequence from N0VA1 CLIP peaks (top) and frequency of that sequence (UCAU) present around the binding site (lower part), (g) Scatterplot of CLIP tag number per peak between Navalwt / wtand Noval^u / ^umice at P21 midbrain. The axes are shown in log2 scales. Sample sizes are Abvt? / te / "" N=3, Nova lwt'm=3.
[0061] (h) Gene annotation analysis52of N0VA1 bound transcripts. Transcripts with the top 1% peak height (read count) for each genotype were analyzed. Genes expressed in the midbrain of P21 mice were used as background for analysis. The term ‘behavior’ is indicated with a black arrowhead, (i) Representative data from gel shift assay using each purified N0VA1 protein and ■’■^P-labeled UCAU RNA oligo probe. Leftmost lane is probe only, no purified protein. The bottom band represents the free probe, and the top shifted band represents the purified protein bound to the RNA probe, (j) The bands were quantified from gel images of the gel shift assay and the amount of binding per protein concentration were plotted. The dissociation constants (Kd values) for each purified N0VA1 protein are shown in the graph. N=5 for each point.
[0062]
[0040] Figure 7A-7F: Comprehensive gene expression analysis in the brain, (a) Global correlation matrix of gene expression levels between brain samples: midbrain at E18.5, cortex, midbrain, cerebellum at P21 in Novam h" and Nova ' ' mice. Heatmap showing correlation coefficients for log2 (TPM+1), color intensity and the size of the circle are proportional to the correlation coefficients, (b) Principal component analysis of gene expression levels between samples. The X axis is the first principal component, and the Y axis is the second principal component, with a percentage of variances explained by each component approximately 53% and 29%, respectively. The ellipses indicate confidence ellipses around group mean points (large dot), (c) Principal component analysis of gene expression levels in each corresponding sample (age and brain region), (d) Gene expression correlations between Novalwt / Wtand Noval^tu^uin corresponding brain regions and age. Scattcrplots of gene expressions measured in average TPMs arc shown. The axes are shown in log2 scales. The red dot indicates a differentially expressed gene betweengenotypes ( ><0,05, FDR<0.1). The yellow dot indicates Noval gene. Pearson correlation is reported on top of the plots. The upper two plots (midbrain) are identical with Figure 2c. (e) Principal component analysis of gene expression levels between Noval^^° and Novalwt / wtmidbrain at E18.5. The RNA sequencing data are from GEO (GSE69711). The first principal component is 37.5%, and the second principal component is 18.6%, respectively. The ellipses indicate confidence ellipses around group mean points (large markers), (f) Gene expression correlations. Scatterplots of gene expressions measured in average TPMs are shown. The axes are shown in log2 scales. The yellow dot indicates Noval gene. Pearson correlation is reported on top of the plots, (left) Novalwt'wtand Noval^°''^° at El 8.5 midbrain. The red dots indicate differentially expressed genes between genotypes (FDR<0.05) (see Supplemental Table 18). (right) Novalwt^wtand Noval^u'^uat E18.5 midbrain (the same with top left panel in d). The green dots indicate differentially expressed genes in the comparison between Novalwt / Wtand Noval^0^0(left panel, corresponding to the red dot). The midbrain sample of E18.5, N
[0063]
[0064] oval^0^0N=3, Noval^^1N=3, Noval^l^'luN=6, Novalwt^wtN=6. The cortex, midbrain, and cerebellum samples of P21, Noval^11'^11N=4, Novalw^wN=4, respectively.
[0065]
[0041] Figure 8A and 8B: Comparison of NOVAI protein expression between the brains of Novalwl,w(mANoval^u, / ^l<mice, (a) Comparison ofNOVAl protein expression in dissected 3-week-old mouse brain tissues. Expression of NOVA proteins were analyzed by immunoblotted on cortex (ex), midbrain (mid), and cerebellum (cb) by panNOVA antibody, NOVAI antibody recognizing the N-terminus and the C-terminus, respectively. The predicted NOVA 1 / 2 protein isoforms are listed in the notes, (b) NOVAI protein expression in NovalHt / wand Noval^u^l‘ mice. Immunostaining for NOVAI protein in sagittal sections of the brain at 3 and 12 weeks of age, respectively.
[0066]
[0042] Figure 9: KH domain sequence alignment. Figure adapted from Lewis et al., 2000 with modifications^. Each RNA binding protein (Nova-2, Nova-1, FMR-1, hnRNP El, hnRNP E2, hnRNP K, GLD-1, Bicaudal-C, PNP and NusA) and comprising KH domain number (KHl, KH2, KH3) are listed on the left. Secondary structural elements were based on the X-ray structure. Color coding scheme: yellow, invariant GXXG motif: purple, hydrophobic core (aliphatic a / 0 platform). Functional classifications: A, aliphatic stacking interaction; S, side chain-base hydrogen bond, including water-mediated contacts; M, protein backbone-base hydrogen bond; *, van der Waals contact. Amino acids on the red background are those for which loss of protein function was reported due to substitution in the hydrophobic core. Amino acids on the green background indicate the 197^ valine ofNOVAl, which is unique to modern humans.
[0067]
[0043] Figure 10A-10J: NOVA1-CLIP analysis in 3-week-old mouse cortex and cerebellum, (a-e) cortex samples, (f-j) cerebellum samples, (a, f) Representative images of autoradiography in NOVA1-CLIP of 3-week- old Novalw^wtand Noval^11'^11mice. The yellow line indicates N0VA1 protein size, and the red enclosing line indicates where RNA was extracted by sectioning, (b, g) Distribution of N0VA1 CLIP peaks on the genome, (c, h) The most enriched binding sequence from N0VA1 CLIP peak (upper part) and frequency of that sequence (UCAU) present around the binding site (lower part), (d, i) Scatterplot of CLIP tag number per peak between Novalwt / w}and Noval^u^u. The axes are shown in log2 scales. R square value is shown, (e) Gene annotation analysis of N0VA1 bound transcripts. Transcripts with the top 100 peak height (read count) in each genotype were analyzed. Genes expressed in the P21 cortex were used as background for analysis. The term “behavior’’ is indicated with a black arrowhead, (j) Gene annotation analysis of NOVAI bound transcripts. Transcripts with the top 1% peak height (read count) in each genotype were analyzed. Genes expressed in the P21 cerebellum were used as background for analysis. The term “behavior” is indicated with a black arrowhead. The cortex and cerebellum samples of P21, Novant^luN=3, Novalwt / WtN=3, respectively.
[0068]
[0044] Figure 11A-11D: Predicted structural model caused by 1197V substitution in the KH2 domain of NOVAL (a) 3D structure prediction by AlphaFold2 (alphafold.ebi.ac.uk / ), showing the expanded KH2 domain of NOVA I. The 197inamino acid is centered, and its proximal amino acids (<5 A) are colored in pink. In this model, the change from isoleucine to valine results in the loss of contact with the amino acid residues at Hl and H3 due to the loss of one carbon chain of the amino acid side chain, (b) Illustration of the amino acids surrounding I197V relative to the secondary structure sequence of KH2. Amino acids structurally proximal to the 197^ amino acid (<5 A) predicted by AlphaFold2 are colored in pink. Corresponding to Fig. le. Amino acids structurally proximal (<5 A) to the 197^ amino acid predicted by AlphaFold2 are boxed in. The KH2 domain sequence of Noval ANQVKIIVPNSTAGLI IGKGGATVKAIMEQSGAWVQLSQKPDGINLQERWTVSGEPEQNRKAVELI IQK IQE (SEQ ID NO: 3) and of Nova J ANQVKI I VPNSTAGL 11 GKGGATVKAVMEQSGAWVQLSQKP DG INLQERWTVSGEPEQNRKAVEL I IQK I QE (SEQ ID NO:4) are depicted, (c-d) Structures of models based on crystallographic data, in which the 197^ amino acid is adjacent to amino acids involved in KH- domain-RNA interactions (Q32 in c, equivalent to the relative position of amino acid 198 inKHl) or KH-domain-protein interactions (M32 in d, equivalent to amino acid 198 in KH2). Themodels are from Figure 9 of Teplova et al., 2011
[0069]
[0045] Figure 12A-12F: Noval^u^umice exhibit alternative splicing changes in specific neuronal genes, (a) NOVAI immunostaining in P21 mouse brain. Scale bars represent 500pm. Abbreviations: CTX, cortex; HIP, hippocampus; CP, Caudate putamen; TH, thalamus; HY, hypothalamus; MB, midbrain; CB, cerebellum; SN, substantia nigra; PAG, periaqueductal gray, (b) Western blotting analysis of NOVAI in dissected mouse brain regions. The corresponding isoforms for each band is listed on the right. Each lane’sNOVAI band is quantified and normalized to the ACTB band signal. The cortex value is set at 1. (c) Alternative splicing (AS) changes in the midbrain of 3- week-old
[0070]
[0071] Nova mice. Differential AS events with p value and delta PSI (di). di >0.05 (in Noval^u'^ltvs. Novalwt,wt') events are shown in light green, those having NOVA1-CLIP peaks on the transcript are shown in dark green, di > -0.05 (in Noval^u'^uvs. Novalw( / fwt) events are shown in light orange, those having NOVA 1 -CLIP peaks on the transcript are shown in magenta. Representative differential AS events are labeled with each gene name. Differential AS events in vocal behavior related genes are labeled in each gene name with a yellow box (Fig. 15). Novalwt / wtN=4, Noval^11'^11N=4. (d) Examples of transcripts showing significant AS changes and having NOVA1-CLIP peaks on its transcript. Differential AS exons are colored in yellow. Information of each AS event is described below the IGV snapshots; Gene name, AS exon number, AS type, percent splice-in value (PSI, the percent of transcripts that include a specific AS exon), percent change (Al, APSI; Noval^w^uvs. Novalwf / wtand p- value. AS splicing events are classified into the following types: Cassette exon (cass), alternative 5’ splice site (alt5), alternative 3’ splice site (alt3), tandem cassette (taca), mutually exclusive exons (rnutx). (e) Gene annotation analysis for the transcripts with differential AS events. The expressed genes in the P21 midbrain were set as background for the analysis, (f) Percentage of genes with differential AS occupying each behavior-related gene ontology category. The number of transcripts with differential AS event in Novani'^urelative to the total number of genes comprising each category are shown.
[0072]
[0046] Figure 13: NOVA I immunostaining in the mouse brain. Immunostaining for NOVAI (green) and DA PI (blue) in postnatal day 21 (P21) and 0 (P0) mouse brain. The scale bars indicate 500pm. The corresponding brain regions are indicated in the orange characters. CTX: cortex, HIP: hippocampus, CP: Caudate putamen, TH: thalamus, HY: hypothalamus, MB: midbrain, CB: cerebellum. The images of NOVAI staining of P21 mouse brain are corresponding to Figure 12a.
[0073]
[0047] Figure 14: NOVAI protein expression in dissected mouse brain. Western blotting for NOVAI, panNOVA and ACTB proteins in dissected brain regions of adult mouse brain. Proteins and isoforms corresponding to each band are listed on the right. The highly expressed NOVAI in the hypothalamus, substantia nigra and periaqueductal gray is Exon4 minus isoforn Three biological replicates are shown. The NOVAI and ACTB plots in replicate 2 are corresponding to Fig3b.
[0074]
[0048] Figure 15: Vocal behavior related transcripts showing differential AS in Noval^11''^11mice. AS exons are colored in yellow. Information of each AS event is described below the IGV snapshots; Gene name, AS exon number, AS type, percent spliced-in value (PSI, the percent of transcripts that include a specific AS exon), percent change (Al, APSI; Noval^u^ vs. Novalw^wt) and p-valuc. AS splicing eventsare classified into the following types: Cassette exon (cass), alternative 5’ splice site (alt5), alternative 3’ splice site (alt3), tandem cassette (taca), mutually exclusive exons (mutx).
[0075]
[0049] Figure 16A and 16B: The resampling analysis for differential AS events in Noval^u'^umice, (a) The 650 random resampling were repeated 1000 times from a list of transcripts detected in the RNAseq dataset to calculate the number of transcripts those annotated in the behavior category in gene ontology database (https: / / geneontologv.org / ). The histogram shows the density of the number of transcripts detected for each resampling, the boxplots at the top show the distribution features (median and quartiles (box) and maximum minimum (whiskers) and outliers (dots)). The green triangle indicates mean value (21.6) of the resampling. The red triangle indicates the number of transcripts detected in this study (27). The number of trials that exceeded the number of 27 was 99, 9.9% probability, indicating the number of transcripts detected in this study is higher than the average number of transcripts detected by chance, (b) The 27 random resampling were repeated 1000 times from the 843 transcripts annotated as the behavior category in the gene ontology database to calculate the number of transcripts those annotated in the vocalization category. The green triangle indicates the mean value (0.711) of the resampling. The red triangle indicates the number of transcripts detected in this study (4). Four trials detected the same number of transcripts as the observed number of transcripts 4, and zero trials exceeded 4, the probability is less than 4%, indicating the number of transcripts detected in this study is higher than the number of transcripts detected by chance with the 5% level of significance.
[0076]
[0050] Figure 17A and 17B: NOVA 1 -CLIP binding peaks in vocalization related transcripts, (a) A list of genes classified to be involved in vocal behavior in gene ontology analysis. Genes for which N0VA1 binding was detected on the transcript in CLIP analysis were marked (check mark). The threshold for N0VA1 binding was a peak detected in all three biological replicates with peak height greater than 10. For P21 samples (cortex, midbrain, cerebellum), the gene was marked if it meets the above criteria in either Noval ^111or Novalwf'wt. (b) Examples of vocalization-related genes with N0VA1 CLIP peaks on their transcripts.
[0077]
[0051] Figure 18A-18I:
[0078]
[0079] mice show altered vocal patterns, (a) Isolation induced ultrasonic vocalization (USV) test for pups, (b) USV parameters and syllable classification, (c) Fqmax distribution and two Gaussian fit in pup USVs. Ashman's D score (a measure of separation of two distributions, above 2 means good separation) are shown. Each Gaussian center and weight are labeled. The intercept of the two Gaussian distributions (black triangle) were used as the cutoffbetween high and low Fqmax USVs. The dot plots at the bottom of the density plots show the mean (black dots) and standard deviation (whiskers) by peak for each genotype. There are no significant differences between genotypes, (d) Ratio of high or low Fqmax in syllables “d” and “m”. The ratio of syllables that belong to each distribution (high or low) are calculated for the total number of each syllable type, (e) Courtship induced USV test for adult mice, (f)
[0080] -loDuration distribution and two Gaussian fit for sy liable “s” in adult USVs. Ashman’s D score, each Gaussian center and weight are shown. The intercept of the two Gaussian distributions (black triangle) was used as the cutoff between long and short duration. The dot plot at the bottom of the density plot shows the mean (black dots) and standaid deviation (whiskers) by peak for each genotype. There are no significant differences between genotypes. Examples for short and long “s” are shown at the top of the plot, (g) Peak frequency parameters in long duration “s”. (h) Fqrnax distribution and two Gaussian fit in adult USVs. Ashman’s D score, Gaussian center and weight are labeled. The black star indicates 100kHz cutoff of high Fqmax and low Fqrnax. The dot plot at the bottom of the density plot shows the mean (black dots) and standard deviation (whiskers) in each USVs for genotype. There are no significant differences between genotypes. Examples for low and high Fqmax syllable are shown at the top of the plot, (i) Frequency variance (Fq variance) in high Fqmax in adult USVs. Data are represented as boxplots that include the minimum score, first (lower) quartile, median, third (upper) quartile, and maximum score, p-values wTere calculated by Wilcoxon rank sum test and corrected with Bonferroni method. * <0.05, ** / ?<0.01. For pup (c-d), each circle indicates data from a single pup. Nova l^lu^lu'^i=4\, Noval^u' / vt'tN=23, Nova
[0081]
[0082] pups. For adults (f-i), experiments were conducted three times in consecutive weeks, and the average of the three experiments was plotted as the value for the mouse (white circle). Novani / ^uN=13, Novalwt'w^ N=14, Novalwt / wtN=13 adults.
[0083]
[0052] Figure 19A-19E: USV characteristics in pups and adults, (a) Syllable composition (left) and amplitude (right) for each syllable in pup USVs. Data is represented as boxplots. Each open circle indicates data from a single pup. / ^-values were calculated by Wilcoxon rank sum test and corrected with Bonferroni method. * p value <0.05. Noval^u^uN=41, Noval^u'wtN=23, Novaiwt / wtN=40 pups, (b) Example of bimodal distributions in peak frequency (Fq) in pup USVs. Density plots of start, minimum, mean, and maximum Fq for syllable “d” observed in each genotype, (c) Density' plots of minimal and maximum Fq for jump syllables (“d”, “u”, “m”) observed in each genotype. Arrows indicate high Fqmax syllables above 100kHz, (d) Density plots of Fqmin and Fqmax in adult USVs. Two Gaussians (blue and orange lines) fitted are overlaid. The Fqmax plot is corresponding to Figure 4h. The Gaussian centers (green and red circles) and weights are labeled. The black star indicates 100kHz cutoff of high and low Fqmax USV s. (e) Examples of mouse USVs spanning the high- frequency regions from previous studies. Figures from Vogel et.al., 2019 7. and Grimsley et al., 2011 R. USVs whth signals above 100 kHz are indicated by orange arrows.
[0084]
[0053] Figure 20A-20E: Playback behavioral experiments for pup USVs. (a) Apparatus used for neonatal mouse vocal selection tests with nursing mother mice. The box consists of three rooms connected by a passageway through which the mouse can pass, and a speaker is attached to each room at each end. (b) Overview of the neonatal mouse vocal selection test with mother mice. Nursing mother mouse ofNova]wt^wtor Nova]^uwas placed in the center of the room and their behavior was recorded while vocal recordings of neonatal mice were played. Vocal recordings of NovaJwt / wtor Novam'^uneonatal mice were randomly played from speakers at both ends, respectively. To exclude direction preference, the recordings played from each speaker were switched after a one- minute break, (c) Vocal recordings of neonatal mice used in the experiments. The recordings were arranged from the data of the pup-USV test to reflect the overall parameters of each genotype (Novalwt / Wtor Noval^u^u). (d) Comparison of the time the mother mouse stayed in the room where the recording of each neonatal mouse was played. Bars indicate mean ± standard error; dots indicate values for each individual. The time spent in each room by the same individual is connected by a line, (e) Comparison of the number of times the mother mouse entered the room where each neonatal mouse recording was played. Bars indicate mean ± standard error; dots indicate values for each individual. The time spent in each room by the same individual is connected by a line. Novalwt^vtmother: N=32, Noviil^u^nimother: N=29.
[0085]
[0054] Figure 21A and 21B: (a) Rotarod performance test, (left) illustration of the test. To measure the locomotor performance (motor coordination), mice were placed on an elevated revolving rod that accelerates at a constant rate (4 to 40rpm in 300sec). The time it took the animals to fall were recorded. Tests were performed three times and the average value were calculated, (middle) Results in NOVAI deficiency mouse model (cKO mice: Tajima et al., 2023^). (right) Results in humanized NOVAI mice. The bar graph represent mean ± standard deviation. The dot indicates average time for one mouse. Gad2^reNovalfl' / wtN=4, Gad2~reNovalfl'fl N=5, Novaiwt / wtN=3, Noval^u'^uN=3. (b) Y-mazc test, (left) illustration of the test and calculation for the alternation rate. Mice were allowed to freely explore a Y-shaped maze for 8 minutes. The number of entries into the arms and the number of triads were recorded to calculate the percentage of alternation. Alternations are consecutive entries into each arm of the Y-maze without any repeats (e.g., aim 1 -> 2 ->3). (middle) results in NOVA 1 deficiency mouse model (cKO mice) The data is from Tajima et al., 2023^. (right) results in humanized NOVAI mice. The alteration rates and total number of entries into arms during the tests are represented as boxplots with the minimum score, first quartile, median, third quartile, and maximum score. Each dot indicates data from a single mouse. Gad2CreNovalfl / wtN=7, Gad2CreNovalfl / -f1N=8, Novalwt / wtN=18, Novalhli / wtN=23, Novalhu / huN=18.
[0086]
[0055] Corresponding reference characters indicate corresponding parts throughout the several views. The examples set out herein illustrate several embodiments of the invention but should not be construed as limiting the scope of the invention in any manner.DETAILED DESCRIPTION
[0087] Definitions:
[0088]
[0056] As used herein the term “about” refers to ± 10 %.
[0089]
[0057] The terms “comprises”, “comprising”, “includes”, “including”, “having” and their conjugates mean “including but not limited to”. It is understood that wherever aspects are described herein with the language "comprising," otherwise analogous aspects described in terms of "consisting of and / or "consisting essentially of’ are also provided.
[0090]
[0058] The term “consisting of’ means “including and limited to”.
[0091]
[0059] The term “consisting essentially of’ means that the composition, method or structure may include additional ingredients, steps and / or parts, but only if the additional ingredients, steps and / or parts do not materially alter the basic and novel characteristics of the claimed composition, method or structure.
[0092]
[0060] As used herein, the singular form “a”, “an” and “the” include plural references unless the context clearly dictates otherwise. For example, the term “a compound” or “at least one compound” may include a plurality of compounds, including mixtures thereof.
[0093]
[0061] Throughout this application, various embodiments of this disclosure may be presented in a range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of the disclosure. Accordingly, the description of a range should be considered to have specifically disclosed all the possible subranges as well as individual numerical values within that range. For example, description of a range such as from I to 6 should be considered to have specifically disclosed subranges such as from 1 to 3, from I to 4, from 1 to 5, from 2 to 4, from 2 to 6, from 3 to 6 etc., as well as individual numbers within that range, for example, 1, 2, 3, 4, 5, and 6. This applies regardless of the breadth of the range.
[0094]
[0062] Whenever a numerical range is indicated herein, it is meant to include any cited numeral (fractional or integral) within the indicated range. The phrases “ranging / ranges between” a first indicate number and a second indicate number and “ranging / ranges from” a first indicate number “to” a second indicate number are used herein interchangeably and are meant to include the first and second indicated numbers and all the fractional and integral numerals therebetween.
[0095]
[0063] As used herein the term “method” refers to manners, means, techniques and procedures for accomplishing a given task including, but not limited to, those manners, means, techniques and procedures either known to, or readily developed from known manners, means, techniques and procedures by practitioners of the chemical, pharmacological, biological, biochemical and medical arts.
[0096]
[0064] Units, prefixes, and symbols are denoted in their Systeme International de Unites (SI) accepted form. Numeric ranges are inclusive of the numbers defining the range. Unless otherwise indicated, amino acid sequences are written left to right in amino to carboxy orientation. The headings provided herein are not limitations of the various aspects of the disclosure, which can be had by reference to the specificationas a whole. Accordingly, the terms defined immediately below are more fully defined by reference to the specification in its entirety.
[0097]
[0065] The term “antibody” describes an immunoglobulin whether natural or partly or wholly synthetically produced. The term also covers any polypeptide or protein having a binding domain which is, or is homologous to, an antibody binding domain. CDR grafted antibodies are also contemplated by this term. An "antibody" is any immunoglobulin, including antibodies and fragments thereof, that binds a specific epitope. The term encompasses polyclonal, monoclonal, and chimeric antibodies, the last mentioned described in further detail in U. S. Patent Nos. 4,816,397 and 4,816,567. The term “antibody(ies)” includes a wild type immunoglobulin (Ig) molecule, generally comprising four full length polypeptide chains, two heavy (H) chains and two light (L) chains, or an equivalent Ig homologue thereof (e.g., a camelid nanobody, which comprises only a heavy chain); including full length functional mutants, variants, or derivatives thereof, which retain the essential epitope binding features of an Ig molecule, and including dual specific, bispecific, multispecific, and dual variable domain antibodies; Immunoglobulin molecules can be of any class (e.g., IgG, IgE, IgM, IgD, IgA, and IgY), or subclass (e.g., IgGl, IgG2, IgG3, IgG4, IgAl, and IgA2). Also included within the meaning of the term “antibody” are any “antibody fragment”.
[0098]
[0066] An “antibody fragment” means a molecule comprising at least one polypeptide chain that is not full length, including (i) a Fab fragment, which is a monovalent fragment consisting of the variable light (VL), variable heavy (VH), constant light (CL) and constant heavy 1 (CHI) domains; (ii) a F(ab')2 fragment, which is a bivalent fragment comprising two Fab fragments linked by a disulfide bridge at the hinge region; (iii) a heavy chain portion of an Fab (Fd) fragment, which consists of the VH and CHI domains; (iv) a variable fragment (Fv), which consists of the VL and VH domains of a single arm of an antibody, ( v) a domain antibody (dAb) fragment, which comprises a single variable domain (Ward, E. S. et al., Nature 341, 544-546 (1989)); (vi) a camelid antibody; (vii) an isolated complementarity determining region (CDR); (viii) a Single Chain Fv Fragment wherein a VH domain and a VL domain are linked by a peptide linker which allows the two domains to associate to form an antigen binding site (Bird et al, Science, 242, 423-426, 1988; Huston et al, PNAS USA, 85, 5879-5883, 1988); (ix) a diabody, which is a bivalent, bispecific antibody in which VH and VL domains are expressed on a single polypeptide chain, but using a linker that is too short to allow for pairing between the two domains on the same chain, thereby forcing the domains to pair with the complementarity domains of another chain and creating two antigen binding sites (WO94 / 13804; P. Holliger et al Proc. Natl. Acad. Sci. USA 90 6444-6448, (1993)); and (x) a linear antibody, which comprises a pair of tandem Fv segments (VH-CH1-VH-CH1) which, together with complementarity light chain polypeptides, form a pair of antigen binding regions; (xi) multivalent antibody fragments (scFv dimers, trimers and / or tetramers (Power and Hudson, J Immunol. Methods 242: 193-204 9 (2000)); (xii) a minibody, which is a bivalent molecule comprised of scFv fused to constant immunoglobulin domains, CH3 or CH4, wherein the constant CH3 or CH4 domains serve as dimerizationdomains (OlafsenT etal (2004) Prot Eng Des Sei 17(4):315-323; Hollinger P and Hudson PJ (2005) Nature Biotech 23(9): 1126-1136); and (xiii) other non-full length portions of heavy and / or light chains, or mutants, variants, or derivatives thereof, alone or in any combination. Chimeric molecules comprising an immunoglobulin binding domain, or equivalent, fused to another polypeptide are included.
[0099]
[0067] As antibodies can be modified in a number of ways, the term "antibody" should be construed as covering any specific binding member or substance having a binding domain with the required specificity. Thus, this term covers antibody fragments, derivatives, functional equivalents and homologues of antibodies, including any polypeptide comprising an immunoglobulin binding domain, whether natural or wholly or partially synthetic. Chimeric molecules comprising an immunoglobulin binding domain, or equivalent, fused to another polypeptide are therefore included. Cloning and expression of chimeric antibodies are described in EP-A-0120694 and EP-A-0125023 and U. S. Patent Nos. 4,816,397 and 4,816,567.
[0100]
[0068] An "antibody combining site" is that structural portion of an antibody molecule comprised of light chain or heavy and light chain variable and hypervariable regions that specifically binds antigen.
[0101]
[0069] The phrase "antibody molecule" in its various grammatical forms as used herein contemplates both an intact immunoglobulin molecule and an immunologically active portion of an immunoglobulin molecule. Exemplary antibody molecules are intact immunoglobulin molecules, substantially intact immunoglobulin molecules and those portions of an immunoglobulin molecule that contains the paratope, including those portions known in the art as Fab, Fab', F(ab')> and F(v), which portions are preferred for use in the therapeutic methods described herein.
[0102]
[0070] The term “specific” may be used to refer to the situation in which one member of a specific binding pair will not show any significant binding to molecules other than its specific binding partner(s). The term is also applicable where e.g. an antigen binding domain is specific for a particular epitope which is carried by a number of antigens, in which case the specific binding member carrying the antigen binding domain will be able to bind to the various antigens carrying the epitope.
[0103]
[0071] The term “comprise” is generally used in the sense of include, that is to say permitting the presence of one or more features or components.
[0104]
[0072] The term “consisting essentially of’ refers to a product, particularly a peptide sequence, of a defined number of residues which is not covalently attached to a larger product. In the case of the peptide of the invention referred to above, those of skill in the art will appreciate that minor modifications to the N- or C-terminal of the peptide may however be contemplated, such as the chemical modification of the terminal to add a protecting group or the like, e.g. the amidation of the C-terminus.
[0105]
[0073] The term “isolated” refers to the state in which specific binding members of the invention, or nucleic acid encoding such binding members will be, in accordance with the present invention. Members and nucleic acid will be free or substantially free of material with which they are naturally associated suchas other polypeptides or nucleic acids with which they are found in their natural environment, or the environment in which they are prepared (e.g. cell culture) when such preparation is by recombinant DNA technology practised in vitro or in vivo. Members and nucleic acid may be formulated with diluents or adjuvants and still for practical purposes be isolated - for example the members will normally be mixed with gelatin or other carriers if used to coat microtitre plates for use in immunoassays, or will be mixed with pharmaceutically acceptable carriers or diluents when used in diagnosis or therapy.
[0106]
[0074] As used herein, "pg" means picogram, "ng" means nanogram, "ng" or "pg" mean microgram, "mg" means milligram, ”ul" or "pl" mean microliter, "ml" means milliliter, "1" means liter.
[0107]
[0075] The amino acid residues described herein are preferred to be in the " L" isomeric form. However, residues in the " D" isomeric form can be substituted for any L-amino acid residue, as long as the desired functional property of immunoglobulin-binding is retained by the polypeptide. NH2refers to the free amino group present at the amino terminus of a polypeptide. COOH refers to the free carboxy group present at the carboxy terminus of a polypeptide.
[0108]
[0076] It should be noted that all amino-acid residue sequences are represented herein by formulae whose left and right orientation is in the conventional direction of amino-terminus to carboxy-terminus. Furthermore, it should be noted that a dash at the beginning or end of an amino acid residue sequence indicates a peptide bond to a further sequence of one or more amino-acid residues.
[0109]
[0077] A "replicon" is any genetic element (e.g., plasmid, chromosome, virus) that functions as an autonomous unit of DNA replication in vivo; i.e., capable of replication under its own control.
[0110]
[0078] A "vector" is a replicon, such as plasmid, phage or cosmid, to which another DNA segment may be attached so as to bring about the replication of the attached segment.
[0111]
[0079] A " DNA molecule" refers to the polymeric form of deoxyribonucleotides (adenine, guanine, thymine, or cytosine) in its either single stranded form, or a double-stranded helix. This term refers only to the primary and secondary' structure of the molecule, and does not limit it to any particular tertiary forms. Thus, this term includes double-stranded DNA found, inter alia, in linear DNA molecules (e.g., restriction fragments), viruses, plasmids, and chromosomes. In discussing the structure of particular double-stranded DNA molecules, sequences may be described herein according to the normal convention of giving only the sequence in the 5' to 3' direction along the nontranscribed strand of DNA (i.e., the strand having a sequence homologous to the mRNA).
[0112]
[0080] An "origin of replication" refers to those DNA sequences that participate in DNA synthesis.
[0113]
[0081] A DNA "coding sequence" is a double-stranded DNA sequence which is transcribed and translated into a polypeptide in vivo when placed under the control of appropriate regulatory sequences. The boundaries of the coding sequence are determined by a start codon at the 5' (amino) terminus and a translation stop codon at the 3' (carboxyl) terminus. A coding sequence can include, but is not limited to, prokaryotic sequences, cDNA from eukaryotic mRNA. genomic DNA sequences from eukaryotic (e.g.,
[0114] -Timammalian) DNA, and even synthetic DNA sequences. A polyadenyiation signal and transcription termination sequence will usually be located 3’ to the coding sequence.
[0115]
[0082] Transcriptional and translational control sequences are DNA regulatory sequences, such as promoters, enhancers, polyadenyiation signals, terminators, and the like, that provide for the expression of a coding sequence in a host cell.
[0116]
[0083] A "promoter sequence" is a DNA regulator}' region capable of binding RNA polymerase in a cell and initiating transcription of a downstream (3' direction) coding sequence. For purposes of defining the present invention, the promoter sequence is bounded at its 3' terminus by the transcription initiation site and extends upstream (5' direction) to include the minimum number of bases or elements necessary to initiate transcription at levels detectable above background. Within the promoter sequence will be found a transcription initiation site (conveniently defined by mapping with nuclease SI), as well as protein binding domains (consensus sequences) responsible for the binding of RNA polymerase. Eukaryotic promoters will often, but not always, contain " TATA" boxes and " CAT" boxes. Prokaryotic promoters contain Shine-Dalgarno sequences in addition to the -10 and -35 consensus sequences.
[0117]
[0084] An "expression control sequence" is a DNA sequence that controls and regulates the transcription and translation of another DNA sequence. A coding sequence is "under the control" of transcriptional and translational control sequences in a cell when RNA polymerase transcribes the coding sequence into mRNA, which is then translated into the protein encoded by the coding sequence.
[0118]
[0085] A "signal sequence" can be included before the coding sequence. This sequence encodes a signal peptide, N-terminal to the polypeptide, that communicates to the host cell to direct the polypeptide to the cell surface or secrete the polypeptide into the media, and this signal peptide is clipped off by the host cell before the protein leaves the cell. Signal sequences can be found associated with a variety of proteins native to prokaryotes and eukaryotes.
[0119]
[0086] The term "oligonucleotide," as used herein in referring to the probe of the present invention, is defined as a molecule comprised of two or more ribonucleotides, preferably more than three. Its exact size will depend upon many factors which, in ton, depend upon the ultimate function and use of the oligonucleotide.
[0120]
[0087] The term "primer" as used herein refers to an oligonucleotide, whether occurring naturally as in a purified restriction digest or produced synthetically, which is capable of acting as a point of initiation of synthesis when placed under conditions in which synthesis of a primer extension product, which is complementary to a nucleic acid strand, is induced, i.e., in the presence of nucleotides and an inducing agent such as a DNA polymerase and at a suitable temperature and pH. The primer may be either singlestranded or double-stranded and must be sufficiently long to prime the synthesis of the desired extension product in the presence of the inducing agent. The exact length of the primer will depend upon many factors, including temperature, source of primer and use of the method. For example, for diagnosticapplications, depending on the complexity of the target sequence, the oligonucleotide primer typically contains 15-25 or more nucleotides, although it may contain fewer nucleotides.
[0121]
[0088] The primers are selected to be "substantially” complementary to different strands of a particular target DNA sequence. This means that the primers must be sufficiently complementary to hybridize with their respective strands. Therefore, the primer sequence need not reflect the exact sequence of the template. For example, a non-complementary nucleotide fragment may be attached to the 5' end of the primer, with the remainder of the primer sequence being complementary to the strand. Alternatively, non-complementary bases or longer sequences can be interspersed into the primer, provided that the primer sequence has sufficient complementarity with the sequence of the strand to hybridize therewith and thereby form the template for the synthesis of the extension product.
[0122]
[0089] As used herein, the terms "restriction endonucleases" and "restriction enzymes" refer to bacterial enzymes, each of which cut double-stranded DNA at or near a specific nucleotide sequence.
[0123]
[0090] A cell has been "transformed" by exogenous or heterologous DNA when such DNA has been introduced inside the cell. The transforming DNA may or may not be integrated (covalently linked) into chromosomal DNA making up the genome of the cell. In prokaiyotes, yeast, and mammalian cells for example, the transforming DNA may be maintained on an episomal element such as a plasmid. With respect to eukaryotic cells, a stably transformed cell is one in which the transforming DNA has become integrated into a chromosome so that it is inherited by daughter cells through chromosome replication. This stability is demonstrated by the ability of the eukaryotic cell to establish cell lines or clones comprised of a population of daughter cells containing the transforming DNA. A "clone" is a population of cells derived from a single cell or common ancestor by mitosis. A "cell line" is a clone of a primary cell that is capable of stable growth in vitro for many generations.
[0124]
[0091] Two DNA sequences are "substantially homologous" when at least about 75% (preferably at least about 80%, and most preferably at least about 90 or 95%) of the nucleotides match over the defined length of the DNA sequences. Sequences that are substantially homologous can be identified by comparing the sequences using standard software available in sequence data banks, or in a Southern hybridization experiment under, for example, stringent conditions as defined for that particular system. Defining appropriate hybridization conditions is within the skill of the art.
[0125]
[0092] By "degenerate to" is meant that a different three-letter codon is used to specify a particular amino acid.
[0126]
[0093] Mutations can be made in the sequences encoding the amino acids, such that a particular codon is changed to a codon which codes for a different amino acid. Such a mutation is generally made by making the fewest nucleotide changes possible. A substitution mutation of this sort can be made to change an amino acid in the resulting protein in a non-conservativc maimer (for example, by changing the codon from an amino acid belonging to a grouping of amino acids having a particular size or characteristic to an aminoacid belonging to another grouping) or in a conservative manner (for example, by changing the codon from an amino acid belonging to a grouping of amino acids having a particular size or characteristic to an amino acid belonging to the same grouping). Such a conservative change generally leads to less change in the structure and function of the resulting protein. A non-conservative change is more likely to alter the structure, activity or function of the resulting protein.
[0127]
[0094] The present invention includes sequences containing amino acid changes and substitutions, including conservative changes, which do not significantly alter the activity or binding characteristics of the resulting protein. In the instant invention, human NOVAI particularly and specifically has a valine amino acid at residue 197. Non-human NOVAI protein does not have a valine at this position. Mouse NOVAI has an isoleucine. Isoleucine is encoded by ATA for example. Valine is encoded by any of GTA, GTT, GTC or GTG. In an aspect of the invention, the mouse ATA is mutated to GTA so that a valine is encoded and a human NO VAI with I197V is expressed, generated, encoded.
[0128]
[0095] Amino acid substitutions may also be introduced to substitute an amino acid with a particularly preferable property. For example, a Cys may be introduced a potential site for disulfide bridges with another Cys. A His may be introduced as a particularly "catalytic" site (z.e., His can act as an acid or base and is the most common amino acid in biochemical catalysis). Pro may be introduced because of its particularly planar structure, which induces p-turns in the protein's structure.
[0129]
[0096] Two amino acid sequences are "highly homologous" or "substantially homologous" when at least about 70% of the amino acid residues (preferably at least about 80%, and most preferably at least about 90% or 95% or 98% or 99%) are identical, or represent conservative substitutions.
[0130]
[0097] A "heterologous" region of the DNA construct is an identifiable segment of DNA within a larger DNA molecule that is not found in association with the larger molecule in nature. Thus, when the heterologous region encodes a mammalian gene, the gene will usually be flanked by DNA that does not flank the mammalian genomic DNA in the genome of the source organism. Another example of a heterologous coding sequence is a construct where the coding sequence itself is not found in nature (e.g., a cDNA where the genomic coding sequence contains introns, or synthetic sequences having codons different than the native gene). Allelic variations or naturally-occurring mutational events do not give rise to a heterologous region of DNA as defined herein.
[0131]
[0098] A DNA sequence is "operatively linked" to an expression control sequence when the expression control sequence controls and regulates the transcription and translation of that DNA sequence. The term "operatively linked" includes having an appropriate start signal (e.g., ATG) in front of the DNA sequence to be expressed and maintaining the correct reading frame to pennit expression of the DNA sequence under the control of the expression control sequence and production of the desired product encoded by the DNA sequence. If a gene that one desires to insert into a recombinant DNA molecule does not contain an appropriate start signal, such a start signal can be inserted in front of the gene.
[0099] The term "agent" means any molecule, including polypeptides, antibodies, polynucleotides, chemical compounds and small molecules. In particular the term agent includes compounds such as test compounds or drug candidate compounds.
[0132]
[0100] The term "agonist" refers to a ligand that stimulates the receptor the ligand binds to in the broadest sense.
[0133]
[0101] The term "assay" means any process used to measure a specific property of a compound. A "screening assay" means a process used to characterize or select compounds based upon their activity from a collection of compounds.
[0134]
[0102] The term "preventing" or "prevention" refers to a reduction in risk of acquiring or developing a disease or disorder (i.e., causing at least one of the clinical symptoms of the disease not to develop) in a subject that may be exposed to a disease-causing agent, or predisposed to the disease in advance of disease onset.
[0135]
[0103] The term "prophylaxis" is related to and encompassed in the term ‘prevention’, and refers to a measure or procedure the purpose of which is to prevent, rather than to treat or cure a disease. Non-limiting examples of prophylactic measures may include the administration of vaccines; the administration of low molecular weight heparin to hospital patients at risk for thrombosis due, for example, to immobilization; and the administration of an anti-malarial agent such as chloroquine, in advance of a visit to a geographical region where malaria is endemic or the risk of contracting malaria is high.
[0136]
[0104] " Therapeutically effective amount" means that amount of a drug, compound, antimicrobial, antibody, or pharmaceutical agent that will elicit the biological or medical response of a subject that is being sought by a medical doctor or other clinician. In particular, with regard to gram-positive bacterial infections and growth of gram-positive bacteria, the term “effective amount” is intended to include an effective amount of a compound or agent that will bring about a biologically meaningful decrease in the amount of or extent of tumor regression and or increase in length of a subject’s survival or period disease-free or in remission. The phrase "therapeutically effective amount" is used herein to mean an amount sufficient to prevent, and preferably reduce by at least about 30 percent, more preferably by at least 50 percent, most preferably by at least 90 percent, a clinically significant change in the growth or amount of tumor size, or enhanced survival or disease-free period by at least about 30 percent, more preferably by at least 50 percent, most preferably by at least 90 percent.
[0137]
[0105] The term "treating" or "treatment" of any disease or infection refers, in one embodiment, to ameliorating the disease or infection (i.e., arresting the disease or growth of the infectious agent or bacteria or reducing the manifestation, extent or severity of at least one of the clinical symptoms thereof). In another embodiment "treating" or "treatment” refers to ameliorating at least one physical parameter, which may not be discernible by the subject. In yet another embodiment, "treating" or "treatment" refers to modulating the disease or infection, either physically, (e.g., stabilization of a discernible symptom), physiologically,(e.g., stabilization of a physical parameter), or both. In a further embodiment, "treating" or "treatment" relates to slowing the progression of a disease or reducing an infection.
[0138]
[0106] The phrase "pharmaceutically acceptable" refers to molecular entities and compositions that are physiologically tolerable and do not typically produce an allergic or similar untoward reaction, such as gastric upset, dizziness and the like, when administered to a human.
[0139]
[0107] As used herein, "pg" means picogram, "ng" means nanogram, "ug" or "pg" mean microgram, "mg" means milligram, "ul" or "pl" mean microliter, "ml" means milliliter, "1" means liter.
[0140]
[0108] “Therapeutically effective amount” or “effective amount” as used herein refers to an amount that is effective to elicit the desired biological or medical response, including the amount of a compound that, when administered to a subject for treating a disease, is sufficient to affect such treatment for the disease. The effective amount will vary depending on the compound, the disease, and its severity and the age, weight, etc., of the subject to be treated. Tire effective amount can include a range of amounts. As is understood in the art, an effective amount may be in one or more doses, i.e., a single dose or multiple doses may be required to achieve the desired treatment endpoint. An effective amount may be considered in the context of administering one or more therapeutic agents, and a single agent may be considered to be given in an effective amount if, in conjunction with one or more other agents, a desirable or beneficial result may be or is achieved. Suitable doses of any co-administered compounds may optionally be lowered due to the combined action (e.g., additive or synergistic effects) of the compounds.
[0141]
[0109] The term “subject” or “animal” is meant any subject, particularly a mammalian subject, in need of treatment with a peptide or polypeptide provided herein. Mammalian subjects include, but are not limited to, humans, dogs, cats, guinea pigs, rabbits, rats, mice, horses, cattle, bears, cows, apes, monkeys, orangutans, and chimpanzees, and so on. In one aspect, the subject or animal is a non-human subject or animal. In one aspect, the non-human subject or animal is a mouse. In an aspect, the non-human subject or animal is a domesticated animal. In an aspect, the non-human subject or animal is a cat or a dog. Detailed Disclosure:
[0142] [HO] Disclosed herein are non-human animal cells and non- human animals comprising a human NOVAI protein or portion thereof comprising the human NOVAI I197V variant, and reagents for making the same. In some embodiments, the sequence encoding a human NOVA 1 protein or portion thereof comprising the human NOVA 1 1197V variant is incorporated in the endogenous locus of a gene.
[0143] [Hl] NOVAI is a neuronal RNA-binding protein expressed in the central nervous system and is essential for survival in mice and normal development in humans. A single amino acid change (1197V) in NOVA 1 ’s second RNA binding domain is present specifically in humans. As described herein, non-human animals (mice) carrying the human-specific 1197V variant (denoted Novalnu''m) have specific alterations in alternative splicing, including multiple binding peaks in mouse brain transcripts involved in vocalization.In particular, the human 1197V substitution results in behavioral differences in vocalization patterns in Novalm'i‘h‘ non-hunran animals (mice) as pups and adults. This human-specific N0VA1 substitution presents as being United to the development of spoken language via differential RNA regulation during brain development.
[0144]
[0112] In some embodiments, provided herein are non-human animal cells and non-human animals having and expressing a human NOVA 1 protein or portion thereof comprising the human N0VA1 I197V variant, in some aspects having the human NOVA! protein or portion thereof comprising the human N0VA1 II 97V variant in the genomes of the non-human animal cells or non-human animals.
[0145]
[0113] Nucleic acid encoding the human N0VA1 protein or portion thereof comprising the human N0VA1 II 97V variant can be inserted into an endogenous Noval locus, thus providing non-human animal cells and non-human animals having a genetically modified endogenous Noval locus.
[0146]
[0114] In some embodiments, provided herein are nucleic acids encoding human N0VA1 protein or portion thereof comprising the human N0VA1 1197V variant, and methods for making non-human animal cells and non-human animals with such nucleic acids. In some embodiments, such nucleic acids have sequences to facilitate the editing of the non-human animal (e.g., loxP sites) flanking the sequences encoding the Noval gene.
[0147]
[0115] In some embodiments, the disclosure provides methods that can be used for making such non-human animals (including such as a rodent, e.g., a rat or a mouse), cells and'' tissues derived from such non-human animals, and nucleotides (e.g., targeting vectors, genomes, etc.).
[0148] [H6[ In some embodiments, the disclosure also provides non-human animal genome comprising a genetically modified endogenous N0VA1 locus having sequence encoding a human N0VA1 protein or portion thereof comprising the human N0VA1 1197V variant.
[0149]
[0117] In some embodiments, non-human animals comprising a humanized N0VA1 locus and expressing a human or humanized N0VA1 protein from the humanized NOV Al locus are provided, as well as methods of using such non-human animals, cells and'tissues derived from such non-human animals, and nucleotides (e.g., targeting vectors, genomes, etc.) useful for making such animals.
[0150]
[0118] In some embodiments, a domain of the human NOVAI sequence is encoded by the segment of the endogenous non-human animal N0VA1 locus that has been deleted and replaced with a heterologous (human) sequence.
[0151] [H9] In some embodiments the non-human animal or non-human animal genome described herein is heterozygous for the genetically modified (human / humanized) endogenous N0VA1 locus. In some embodiments, the non-human animal or non-human animal genome is homozygous for the genetically modified endogenous (human / humanized) endogenous NOVAI locus.
[0152]
[0120] In some embodiments, segments of an endogenous NOV Al locus are deleted and replaced with an exogenous NOV Al sequence. In some of these cases, the endogenous NOV Al locus that hasbeen deleted can comprise a segment of the 3' untranslated region, a segment of one or more coding exon, a segment of one or more intron, or a combination of the aforementioned segments of the endogenous NOVAI locus.
[0153]
[0121] In some embodiments, a human N0VA1 sequence may be used to replace a locus within a non-human animal or non-human cell. In such embodiments the orthologous human NOVA 1 sequence that replaces the segment of the endogenous locus may comprise a segment of any one or more of the 3' untranslated region of the human NOV Al sequence, exon of the human NOVAI sequence, intron of the human N0VA1 sequence, or any combination thereof.
[0154]
[0122] In some embodiments, the non-human animal is a mammal, or the non-human animal genome is a mammalian genome. In some embodiments, the non-human animal can be a rodent, or the non-human animal genome can be a rodent genome. In some embodiments, the non-human animal can be a rat or mouse, or the non-human animal genome can be a rat genome or a mouse genome.
[0155]
[0123] In some embodiments, provided herein is a non-human animal or non-human animal cell comprising a genetically modified endogenous NOV Al locus encoding a modified N0VA1 protein, wherein the modified N0VA1 protein is a human or humanized N0VA1 protein comprising the human N0VA1 1197V variant, wherein a portion of the endogenous NOVAI locus sequence has been replaced with human N0VA1 sequence. In some embodiments, provided herein is a non-human animal or non-human animal cell comprising a genetically modified endogenous N0VA1 locus encoding a modified NOVAI protein, wherein the modified NOV Al protein is a human or humanized NOVAI protein comprising the human N0VA1 1197V variant, wherein nucleic acid encoding amino acid 197 of the N0VA1 protein has been replaced with nucleic acid encoding the human NOV Al sequence and amino acid 197 valine. In some embodiments, provided herein is a non-human animal ceil comprising a genetically modified endogenous NOV Al locus encoding a modified NOVAI protein, wherein the modified NO VAI protein is a human or humanized NOVAI protein comprising the human N0VA1 I197V variant, wherein a segment of the endogenous NOVAI locus has been deleted and replaced with an orthologous human NO VAI sequence. The non-human animal cell can be a neural cell, a pluripotent cell, an ES cell, or a germ cell.
[0156]
[0124] In some embodiments, the disclosure further provides methods for making any non-human animal, or reagents required for making the non-human animal as described herein.
[0157]
[0125] Optionally, a humanized N0VA1 locus in a non-human animal or cell can comprise other elements. Examples of such elements can include selection cassettes, reporter genes, recombinase recognition sites, or other elements. Alternatively, the humanized NOV Al locus can lack other elements (e g., can lack a selection marker or selection cassette). Examples of suitable reporter genes and reporter proteins are known and available. Examples of suitable selection markers includeneomycin phosphotransferase (near), hygromycin B phospho- transferase (hygr), puromycin-N-acetyltransferase (puror), blasticidin S deaminase (bsrr), xanthine / guanine phosphoribosyl transferase (gpt), and herpes simplex virus thymidine kinase (HSV-k). Examples of recombinases include Cre, Flp, and Dre recombinases. One example of a Cre recombinase gene is Crei, in which two exons encoding the Cre recombinase are separated by an intron to prevent its expression in a prokaryotic cell. Such recombinases can further comprise a nuclear localization signal to facilitate localization to the nucleus (e.g., NLS-Crei). Recombinase recognition sites include nucleotide sequences that are recognized by a site-specific recombinase and can serve as a substrate for a recombination event. Examples of recombinase recognition sites include FRT, FRT11, FRT71, attp, att, rox, and lox sites suchasloxP, lox511, lox2272, lox66, lox71, loxM2, and lox5171.
[0158]
[0126] Other elements such as reporter genes or selection cassettes can be self-deleting cassettes flanked by recombinase recognition sites. As an example, the self-deleting cassette can comprise a Crei gene (comprises two exons encoding a Cre recombinase, which are separated by an intron) operably linked to a mouse Prml promoter and a neomycin resistance gene operably linked to a human ubiquitin promoter. By employing the Prml promoter, the self-deleting cassette can be deleted specifically in male germ cells of FO animals. The polynucleotide encoding the selection marker can be operably linked to a promoter active in a cell being targeted. Examples of promoters are described elsewhere herein. As another specific example, a self-deleting selection cassette can comprise a hygromycin resistance gene coding sequence operably linked to one or more promoters (e.g., both human ubiquitin and EM7 promoters) followed by a polyadenylation signal, followed by a Crei coding sequence operably linked to one or more promoters (e.g., an mPrml promoter), followed by another polyadenylation signal, wherein the entire cassette is flanked by loxPsites.
[0159]
[0127] The non-human animal cells provided herein can be, for example, any non-human cell comprising an NOV Al locus or a genomic locus homologous or orthologous to the human N0VA1 locus, particularly wherein the encoded N0VA1 comprises a valine at amino acid 197. The cells can be eukaryotic cells, which include, for example, fungal cells (e.g., yeast), plant cells, non-human animal cells, non-human mammalian cells. An animal can be, for example, a mammal, fish, or bird. A mammalian cell can be, for example, a non-human mammalian cell, a rodent cell, a rat cell, a mouse cell, or a hamster cell. Other non-human mammals include, for example, non-human primates, monkeys, apes, orangutans, cats, dogs, rabbits, horses, bulls, deer, bison, livestock (e.g., bovine species such as cows, steer, and so forth; ovine species such as sheep, goats, and so forth; and porcine species such as pigs and boars). Birds include, for example, chickens, turkeys, ostrich, geese, ducks, and so forth. Domesticated animals and agricultural animals are also included. The term "non-human" excludes humans.
[0128] The cells can also be any type of undifferentiated or differentiated state. For example, a cell can be a totipotent cell, a pluripotent cell (e.g., a human pluripotent cell or a non-human pluripotent cell such as a mouse embryonicstem (ES) cell or a rat ES cell), or a non-pluripotent cell. Totipotent cells include undifferentiated cells that can give rise to any cell type, and pluripotent cells include undifferentiated cells that possess the ability to develop into more than one differentiated cell types. Such pluripotent and / or totipotent cells can be, for example, ES cells or ES-like cells, such as an induced pluripotent stem (iPS) cells. ES cells include embryo-derived totipotent or pluripotent cells that can contribute to any tissue of the developing embryo upon introduction into an embryo. ES cells can be derived from the inner cell mass of a blastocyst and can differentiate into cells of any of the three vertebrate germ layers (endoderm, ectoderm, and mesoderm).
[0160]
[0129] The cells provided herein can also be germ cells (e.g., sperm or oocytes). The cells can be mitotically competent cells or mitotically-inactive cells, meiotically competent cells or meiotically-inactive cells. Similarly, the cells disclosed herein can also be primary somatic cells or cells that are not a primary somatic cell. Somatic cells include any cell that is not a gamete, germ cell, gametocyte, or undifferentiated stem cell. For example, the cells disclosed herein can be neural cells. Suitable cells provided herein also include primary cells. Primary cells include cells or cultures of cells that have been isolated directly from an organism, organ, or tissue. Primary cells include cells that are neither transformed nor immortal. Primary cells include any cell obtained from an organism, organ, or tissue which was not previously passed in tissue culture or has been previously passed in tissue culture but is incapable of being indefinitely passed in tissue culture. Such cells can be isolated by conventional techniques.
[0161]
[0130] Other suitable cells provided herein include immortalized cells. Immortalized cells include cells from a multicellular organism that would normally not proliferate indefinitely but, due to mutation or alteration, have evaded normal cellular senescence and instead can keep undergoing division. Such mutations or alterations can occur naturally or be intentionally induced. Examples of immortalized cell lines are myofiber cell lines. Immortalized or primary cells include cells that can be used for culturing or for expressing recombinant genes or proteins.
[0162]
[0131] The cells provided herein also include one-cell stage embryos (i.e., fertilized oocytes or zygotes). Such one-cell stage embryos can be from any genetic background (e.g., BALB / c, C57BL / 6, 129, or a combination thereoffor mice), can be fresh or frozen, and can be derived from natural breeding or in vitro fertilization.
[0163]
[0132] Non-human animals comprising a humanized Noval locus as described herein can be made by the methods described elsewhere herein. An animal can be, for example, a mammal, fish, or bird. Non-human mammals include, for example, non-human primates, monkeys, apes, orangutans, cats, dogs, horses, bulls, deer, bison, sheep, rabbits, rodents (e.g., mice, rats, hamsters,
[0164] -SO-and guinea pigs), and livestock (e.g., bovine species such as cows and steer; ovine species such as sheep and goats; and porcine species such as pigs and boars). Birds include, for example, chickens, turkeys, ostrich, geese, and ducks. Domesticated animals and agricultural animals are also included. The term "non-human animal" excludes humans. Preferred non-human animals include, for example, rodents, such as mice and rats.
[0165]
[0133] The non-human animals can be from any genetic background. For example, suitable mice can be from a 129 strain, a C57BL / 6 strain, a mix of 129 and C57BL / 6, a BALB / c strain, or a Swiss Webster strain. Examples of 129 strains include 129P1, 129P2, 129P3, 129X1, 129S1 (e.g., 12951 / SV, 12951 / Svlm), 129S2, 129S4, 129S5, 12959 / SvEvH, 129S6 (129 / SvEvTac), 129S7, 129S8, 129T1, and 129T2. Examples of C57BL strains include C57BL / A, C57BL / An, C57BL / GrFa, C57BL / Kal_wN, C57BL / 6, C57BL / 6J, C57BL / 6ByJ, C57BL / 6NJ, C57BL / 10, C57BL / lOScSn, C57BL / 10Cr, and C57BL / Ola. Suitable mice can also be from a mix of an aforementioned 129 strain and an aforementioned C57BL / 6 strain (e.g., 50% 129 and 50% C57BL / 6). Likewise, suitable mice can be from a mix of aforementioned 129 strains or a mix of aforementionedBL / 6 strains (e.g., the 12956 (129 / SvEvTac) strain).
[0166]
[0134] Similarly, rats can be from any rat strain, including, for example, an ACI rat strain, a Dark Agouti (DA) rat strain, a Wistar rat strain, a LEA rat strain, a Sprague Dawley ( SD) rat strain, or a Fischer rat strain such as Fisher F344 or Fisher F6. Rats can also be obtained from a strain derived from a mix of two or more strains recited above. For example, a suitable rat can be from a DA strain or an ACI strain. Such strains are available from a variety of sources including Harlan Laboratories and Charles River Laboratories. Some suitable rats can be from an inbred rat strain.
[0167]
[0135] Various methods are provided for making a non-human animal comprising a human or humanized NO VAI locus as disclosed elsewhere herein. Any convenient method or protocol for producing a genetically modified organism is suitable for producing such a genetically modified non- human animal. Such genetically modified non-human animals can be generated, for example, through gene knock-in at a targeted NOVAI locus.
[0168]
[0136] For example, the method of producing a non-human animal comprising a humanized NOVAI locus can comprise: (1) modifying the genome of a pluripotent cell to comprise the humanized NOV Al locus or NOVAI encoding the human valine 197; (2) identifying or selecting the genetically modified pluripotent cell comprising the humanized NOV Al locus or NOVAI encoding the human valine 197; (3) introducing the genetically modified pluripotent cell into anon-human animal host embryo cells in vitro; and (4) implanting and gestating the host embryo cells in a surrogate mother. Optionally, the host embryo comprising modified pluripotent cell (e.g., a non- human ES cell) can be incubated until the blastocyst stage before being implanted into and gestated in the surrogate mother to produce an FO non-human animal. The surrogate mother can
[0169] -Sithen produce an FO generation non-human animal comprising the humanized NOV Al locus or NOVAI encoding the human valine 197.
[0170]
[0137] An example of a suitable pluripotent cell is an embryonic stem (ES) cell (e.g., a mouse ES cell or a rat ES cell). The modified pluripotent cell can be generated, for example, through recombination by (a) introducing into the cell one or more targeting vectors comprising an insert nucleic acidflanked by 5’and 3' homology arms corresponding to 5' and 3' target sites, wherein the insert nucleic acid comprises a heterologous / humanized NOV Al locus; and (b) identifying at least one cell comprising in its genome the insert nucleic acid integrated at the target genomic locus. Alternatively, the modified pluripotent cell can be generated by ( a) introducing into the cell: (i) a nuclease agent, wherein the nuclease agent induces a nick or double-strand break at a recognition site within the target genomic locus; and (ii) one or more targeting vectors comprising an insert nucleic acid flanked by 5' and 3' homology arms corresponding to 5' and 3' target sites located in sufficient proximity to the recognition site, wherein the insert nucleic acid comprises the heterologous / humanized NOVAI locus; and (c) identifying at least one cell comprising a modification (e.g., integration of the insert nucleic acid) at the target genomic locus. Any nuclease agent that induces a nick or double-strand break into a desired recognition site can be used. Examples of suitable nucleases include a Transcription Activator -Like Effector Nuclease (TALEN), a zinc-finger nuclease (ZFN), a meganuclease, and Clustered Regularly Interspersed Short Palindromic Repeats (CRISPR) / CRISPR-associated (Cas) systems or components of such systems (e.g., CRISPR / Cas9). The donor cell can be introduced into a host embryo at any stage, such as the blastocyst stage or the pre-morula stage (i.e., the 4 cell stage or the 8 cell stage). Progeny that are capable of transmitting the genetic modification though the germline are generated.
[0171]
[0138] Nuclear transfer techniques can also be used to generate the non-human mammalian animals. Briefly, meth- ods for nuclear transfer can include the steps of: (1) enucleating an oocyte or providing an enucleated oocyte; (2) isolating or providing a donor cell or nucleus to be combined with the enucleated oocyte; (3) inserting the cell or nucleus into the enucleated oocyte to form a reconstituted cell; (4) implanting the reconstituted cell into the womb of an animal to form an embryo; and (5) allowing the embryo to develop. In such methods, oocytes are generally retrieved from deceased animals, although they may be isolated also from either oviducts and / or ovaries of live animals. Insertion of the donor cell or nucleus into the enucleated oocyte to form a reconstituted cell can be by microinjection of a donor cell under the zona pellucida prior to fusion. Fusion may be induced by application of a DC electrical pulse across the contact / fusion plane (electrofusion), by exposure of the cells to fusion-promoting chemicals, such as polyethylene glycol, or by way of an inactivated virus, such as the Sendai virus. A reconstituted cell can be activated by electrical and / or non-electrical means before, during, and / or after fusion of the nuclear donor and recipient oocyte. Activation methodsinclude electric pulses, chemically induced shock, penetra- tion by sperm, increasing levels of divalent cations in the oocyte, and reducing phosphorylation of cellular proteinsfas by way of kinase inhibitors) in the oocyte. The activated reconstituted cells, or embryos, can be cultured in media and then transferred to the womb of an animal.
[0172]
[0139] The various methods provided herein allow for the generation of a genetically modified nonhuman animal wherein the cells of the genetically modified animal comprise the humanized N0VA1 locus. It is recognized that depending on the method used to generate the animal, the number of cells within the animal that have the heterologous NO VAI locus will vary. The introduction of the donor ES cells into a pre-morula stage embryo from a corresponding organism (e.g., an 8-cell stage mouse embryo) via for example, the VELOCIMOUSE® method allows for a greater percentage of the cell population of the animal to comprise cells having the nucleotide sequence of interest comprising the targeted genetic modification. For example, at least 50%, 60%, 65%, 70%, 75%, 85%, 86%, 87%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% of the cellular contribution of the non-human animal can comprise a cell population having the targeted modification.
[0173]
[0140] In accordance with the invention, methods are provided for altering or modifying vocalization in a non-human animal comprising expressing in said non-human animal a human or humanized NOVAI, whereby the NOVAI has a valine at amino acid 197. In accordance with the invention, methods are provided for altering or modifying vocal communication in a non-human animal comprising expressing in said non-human animal a human or humanized NOVAI, whereby the NOVAI has a valine at amino acid 197. In accordance with an aspect of the method, the native NOVAI encoding sequence in the non-human animal is modified or otherwise genetically engineered to encode a human or humanized NOVA 1 protein comprising a valine at amino acid 197.
[0174]
[0141] Vocalization may be assessed, evaluated and tested in accordance with standards and methods known in the art, including as provided and described herein.
[0175]
[0142] In an aspect of the invention, the neuronal transcriptome may be altered or modified and evaluated in the non-human animals of the invention. Thus, methods are provided for altering or modifying the neuronal transcriptome of a non-human animal comprising expressing in said non-human animal a human or humanized NOVAI, whereby the NOVAI has a valine at amino acid 197.
[0176]
[0143] In aspects of the invention, the native NOVAI encoding sequence in the non-human animal is modified or otherwise genetically engineered to encode a NOVAI protein comprising a valine at amino acid 197 and to also encode additional amino acids which are unique versus the native NOVAI gene. In an aspect, these additional amino acids may serve as a marker, tag, or ligand, including and such as for purposes of detection, isolation or characterization.
[0144] In a preferred aspect, the present invention provides a nucleic acid which codes for a polypeptide of the invention as defined above, including any polypeptides as set out herein. Nucleic acid includes DNA and RNA.
[0177]
[0145] The present invention also provides constructs in the form of plasmids, vectors, transcription or expression cassettes which comprise at least one polynucleotide as above. The present invention also provides a recombinant host cell which comprises one or more constructs as above. Expression may conveniently be achieved by culturing under appropriate conditions recombinant host cells containing the nucleic acid. Following production by expression a binding protein may be isolated and / or purified using any suitable technique, then used as appropriate.
[0178]
[0146] Systems for cloning and expression of a polypeptide in a variety of different host cells are well known. Suitable host cells include bacteria, mammalian cells, yeast and baculovirus systems. Suitable vectors can be chosen or constructed, containing appropriate regulatory sequences, including promoter sequences, terminator sequences, polyadenylation sequences, enhancer sequences, marker genes and other sequences as appropriate. Vectors may be plasmids, viral e.g. 'phage, or phagemid, as appropriate.
[0179]
[0147] Thus, a further aspect of the present invention provides a host cell containing nucleic acid as disclosed herein. A still further aspect provides a method comprising introducing such nucleic acid into a host cell. The introduction may employ any available technique. The introduction may be followed by causing or allowing expression from the nucleic acid, e.g. by culturing host cells under conditions for expression of the gene. The present invention also provides a method which comprises using a construct as stated above in an expression system in order to express a binding member polypeptide. Another feature of this invention is the expression of the DNA sequences disclosed herein. As is well known in the art, DNA sequences may be expressed by operatively linking them to an expression control sequence in an appropriate expression vector and employing that expression vector to transform an appropriate unicellular host. A wide variety of host / expression vector combinations may be employed in expressing the DNA sequences of this invention.
[0180]
[0148] It will be understood that not all vectors, expression control sequences and hosts will function equally well to express the DNA sequences of this invention. Neither will all hosts function equally well with the same expression system. However, one skilled in the art will be able to select the proper vectors, expression control sequences, and hosts without undue experimentation to accomplish the desired expression without departing from the scope of this invention.
[0181]
[0149] The invention may be better understood by reference to the following non-limiting Examples, which are provided as exemplary of the invention. The following examples are presented in order to more fully illustrate the preferred embodiments of the invention and should in no way be construed, however, as limiting the broad scope of the invention.
[0150] While the present disclosure has been described with reference to preferred embodiments, it will be understood by those skilled in the art that various changes may be made and equivalents may be substituted for elements thereof to adapt to particular situations without departing from the scope of the present disclosure. Therefore, it is intended that the present disclosure not be limited to the particular embodiments disclosed as the best mode contemplated for carrying out the present disclosure, but that the present disclosure will include all embodiments falling within the scope and spirit of the appended claims.
[0182] EXAMPLES EXAMPLE 1
[0183] Introduction
[0184]
[0151] Fossil records indicate that modem humans (Homo sapiens) emerged 200,000-300,000 years ago as the predominant species from several now-extinct hominid subclades11 ’ 7. Humans differ significantly from their closest living relatives, the great apes, particularly in their ability to communicate through complex learned vocal communication, a necessary component of spoken language. This complexity is driven by some anatomical adaptions of the vocal tract and intricate neural networks linking various brain 1 7
[0185] regions. However, the genetic basis underlying these specialized human traits remains to be fully identified.
[0186]
[0152] The closest evolutionary relatives of modern humans are two extinct lineages: Neanderthals and Denisovans. Genome sequencing from fossilized remains of these archaic humans has identified distinct genetic differences between them and modern humans, which may be relevant to recent human
[0187] O _ 11
[0188] evolution. Additionally, the availability of extensive human genome data over the past few decades, initially focused on European populations, has significantly expanded the scope of evolutionary studies^_14
[0189]
[0153] The transcription factor forkhead box P2 (FOXP2) is of particular interest as a potential driver of human language function, as it harbors two amino acid substitutions present in human but not in chimpanzee and many other mammal genomes. Families with FOXP2 mutations exhibit severe speech defects^-T6 while FOXP2 disruption in mice leads to vocalization abnormalities
[0190]
[0191] suggesting a role in spoken language function. Studies on mice with the two amino acids substituted to the human version have reported vocal changes both in the neonatal 10 and adult stages
[0192]
[0193] ’ I. While Hammerschmidt et al. observed minimal vocal changes, von Merten et al. reported qualitative changes under a more natural vocalization paradigm
[0194]
[0195] , suggesting the involvement of these two amino acids in vocalization.
[0196] However, these substitutions are also present in archaic humans, and comprehensive analyses usingdiverse human genome datasets have found no evidence of recent selection. This suggests that the FOXP2 substitutions occurred earlier than initially thought ’
[0197]
[0198] . Similarly, the TKTL1 gene contains a human-specific amino acid thought to influence greater ncurogcncsis in human than Neanderthal frontal cortex, though this finding is based on European ancestry genome datasets-0. Broader analyses of modem human genomes reveal that 0.03-0.2% of individuals possess the ‘putative Neanderthal variant’, indicating its presence in a significant portion of the population 9 These findings underscore the importance of incorporating diverse human samples to identify and validate the genetic background of modern human traits through genomic comparisons.
[0199]
[0154] Genomic comparisons between archaic humans, ape genomes, and the broader human population have identified 61 human-specific nonsynonymous coding variants that are fixed or nearly fixed in modern humans^. Notably, one of the genes includes an isoleucine to valine substitution at position 197 (1197V) in the RNA binding protein neuro-oncological ventral antigenl (NOVAI). NOVAl is highly expressed in neurons of the central nervous system (CNS) in both mice and humans
[0200]
[0201] , and its expression has also been observed in cultured human and rat
[0202]
[0203] cells-. NOVAl was first identified as an autoantigen targeted in the paraneoplastic neurologic disorder (PND) opsoclonus-myoclonus ataxia (POMA). PNDs develop when tumor cells ectopically express proteins normally restricted to the nervous system, triggering an anti-tumor immune response that breaches the blood-brain barrier, leading to autoimmune neurologic
[0204]
[0205] disease- --. In POMA, a robust immune response is mounted against NOVAl and its paralog, NOVA2
[0206]
[0207] . The autoimmune disorder is characterized by motor dysfunction due to failure of inhibition of midbrain neurons, which results in the hyperactivity associated with opsoclonus-myoclonus ataxia-^. In mice, homozygous deletion of the Noval gene results in an early postnatal lethal phenotype due to abnormal motor function. Therefore, N0VA1 plays a crucial role in neural development and neuromuscular control in mammals. NOVA proteins directly bind RNA in the mouse brain^^ ^to regLl]a epre-mRNA processing01 4,35, translation0^1and neurophysiology. Genetic studies mapping NOVA target RNAs in mice and humans have also linked it to autism. Interestingly, a human patient with a heterozygous deletion of NOVA] presented with delay of language development, learning disabilities, motor hyperactivity and behavioral dysregulation 900.
[0208]
[0155] Studies have explored the N0VA1 1197V variant by reverting the ancestral isoleucine 197 variant back into human iPSC-dcrivcd organoids, revealing morphological and electrophysiological changes in viiro^. However, these effects were not observed in another study that reintroduced the same substitutionin different iPSCs^ This discrepancy underscores challenges of obtaining consistent results with varying experimental methods and materials in vitro^ ’^.
[0209]
[0156] In this work, we used gene-editing to substitute the N0VA1 isoleucine (I) isoform present in most mammals and archaic hominids (Neanderthals and Denisovans) with the human-specific valine (V) variant at position 197 in mice. Comparison of these humanized N0VA1 mice (Noval^uf^u) with wild-type mice carrying the ancestral Nova I gene (Novalwt'wf) revealed specific transcriptomic and behavior differences related to vocalization. Taken together, the unique role of N0VA1 in neurons, its association with human disease, and evidence that the human-specific amino acid 197 variant confers vocalization changes in humanized mice suggest a role for NOVA! in the evolution of human-specific language.
[0210] RESULTS NOVAI I197V has characteristics of a variant that underwent a strong evolutionary selective sweep in modern humans
[0211]
[0157] NOV A proteins have three K homology (KH)-type RNA binding domains and have been shown to bind to YCAY repeat sequences on target transcripts^ 6,42-46 Comparison of the amino acid sequences of N0VA1 and N0VA2 in various organisms showed that N0VA1 is extremely highly conserved across the whole protein sequence, whereas N0VA2 shows higher variability between species (Fig. la). We compared eight human genomes with three high-coverage Neanderthal genomes and one high-coverage Denisovan genome. The only change between modern and ancient humans was a nonsynonymous nucleotide substitution encoding the 197^ amino acid of NOVAI, resulting in a valine in modern humans replacing an isoleucine in ancient humans (Fig. lb). An expanded genomic analysis of the dbSNP database revealed that this substitution was present in all but six of 650,058 human sequences, five of which were from individuals of Asian descent (Figure 2a). The samples are deidentified, so it is not possible to assess those individuals.
[0212]
[0158] We analyzed the mean allele frequencies (MAF) of NOVAI in 121,412 human genomes from the ExAC database (Fig. 1c, Fig. 2b). This analysis confirmed the extremely low frequency of variants encoding 197V in humans, consistent with strong selective pressure across the entire NOVA 1 coding sequence. Specifically, the upper bound of 95% confidence interval for the MAF of NOVAI was 0.00071, significantly lower than that of NO A2 (0.0099) or the average across all genes on chromosome 14, where NOVAI is located (0.0042; Fig. 1c). Moreover, evolutionary analysis on the NOVAI gene yielded a Tajima’s D statistic of -2.48 (Table 1). Normalizing Tajima's D values between genes by the theoretical minimum^, we found that NOVA / ’s normalized Tajima’s D statistic was exceedingly low. This suggeststhat NO VA has undergone strong purifying selection, particularly in comparison to N0VA2. neighboring genes (F0XG1, STXBP6), and a set of genes associated with 61 human-specific nonsynonymous coding variants^ (Fig. Id, Fig. 2, Table 1).
[0213]
[0159] To further examine whether a selective sweep occurred at the NOVAI gene locus, we performed the DH test — a robust method for detecting selection that remains insensitive to other perturbations. Using human genetic data from the 1000 Genomes Project, we calculated the DH value for the NOVAI gene locus, which was 0.42643, reaching the 5% significance level among the genes on the same chromosome (chr. 14). In contrast, adjacent genes did not reach significance (FOXGE. 0.90390, STXBP6'.
[0214] 0.77477). It has been reported that the human-specific variant in NO VAI resides on the third-largest human-specific haplotype among the fixed human-specific sites, featuring two high-frequency haplotypes Our results, together with those of Trujillo et al, support a model whereby the NOVA 1 197V variant distinguishes NOVAI in modern humans from ancient humans, primates, or more distant species (Fig. la-c).
[0215]
[0160] To further evaluate the hypothesis that the NOVA 1 197V variant may have been part of a selective sweep, we performed an analysis of the ancestral recombination graph (ARG) for modern and archaic A Q
[0216] hominins in the surrounding region, using ARGs previously inferred using ARGweaver- IX. These ARGs explicitly describe gene trees and accompanying recombination events throughout the region. Focusing on the NOVAI 197V variant, we analyzed the sampled ARGs using CLUES2^^, a method that estimates a selection coefficient to best explain observed changes in allele frequency over time, based on indirect information from the ARG.
[0217]
[0161] Our analysis revealed that selection at the NOVAI SNP is relatively strong and statistically significant, with an estimated selection coefficient of s = 0.00082 (p = 0.019 for the null hypothesis of no selection). While this estimate is an order of magnitude less than those observed at the strongest sweeps in the human genome (such as AC' / '. which has s = 0.01
[0218]
[0219] it is still substantial, corresponding to a population- scaled coefficient of 5 = 2Nes = 19, indicating strong selection relative to nearly neutral evolution (where |5|<= 1) (Fig. 4). For comparison, we applied the same CLUES2 analysis
[0220] to 38 other SNPs for which informative ARGs were available, previously identified as potentially selected in the human genome
[0221]
[0222] The results showed that the selection observed at the NOVAI SNP is relatively strong compared to these other genes, with 33 of the 38 showing either non- significant results or smaller selection coefficients (Table 2).
[0223]
[0162] Taken together, the observation that the NOVA! 197V allele became nearly fixed and is shared across human population groups suggests that it arose and increased to high frequency before theirdivergence. Our analyses support the idea that the N0VA1 197V variant was part of an ancient selective sweep in modern humans, predating many other known sweeps in the humangenome.
[0224] Humanized NOVAI mice are comparable to wild-type mice in development and gene expression in the brain
[0225]
[0163] To explore the physiologic and biological significance of the II 97V amino acid substitution in NOVA 1, we used CRISPR / Cas9-based gene editing to introduce nucleotide changes that result in humantype N0VA1 knock-in mice {Noval^u^nimice; Fig. le, Fig. 5). Noval^11'^11mice developed normally (Fig. 6a) and exhibited fertility similar to littermate controls (Nova
[0226]
[0227] mice. The brain to body weight ratio was comparable between Noval^u / ^uand Noval^^ mice (Fig. 6b). Comprehensive gene expression analysis of the midbrain at embryonic day 18.5 (E18.5) and at postnatal day 21 (P21), when fundamental neural circuits and behaviors have been established, revealed that the transcript levels, including Nova! itself, were nearly identical between genotypes throughout development (Fig. 6c) and across different brain regions (Fig.7a-d). Similarly, the expression pattern and levels of NOVA 1 protein were equivalent between genotypes in the brain (Fig. 8).
[0228]
[0164] The only gene that showed a significant steady-state difference va. Noval^u^umice was Gkn3 at P21, a secreted protein involved in endothelial cell proliferation (Figure 6c; y?-value<0.05, FDR<0.1). Gkn3 is thought to be involved in adaptive gene loss during recent human evolution’’3, and showed down regulation in the P21 midbrain of Nova
[0229]
[0230] mice (average TPM 18.6 in Novalwl, wl, 11.3 in Novalhu / hu, log2FC = -0.71, / ?-value = 2.3xl0'6, FDR = 0.037).
[0231] Modern human specific amino acid substitution does not affect sequence specific RNA- binding capacity' of NOVAI
[0232]
[0165] NOVA proteins harbor three KH domains that are responsible for sequence-specific RNA-
[0233]
[0234] binding^2>45,46 domain, found in many RNA binding proteins, includes common motifs: an invariant Gly-X-X-Gly motif, a hydrophobic core, and a variable loop. The 197^ amino acid in NOVAI is a part of the hydrophobic core. Several studies have shown that a single amino acid substitution in the hydrophobic core cause loss of function of RNA binding p
[0235]
[0236] roteins^^- ^6 pori\ova], jnvitro binding assays have demonstrated that single amino acid substitutions (lie > Thr in KH1, Leu— > Asn in KH3) in the hydrophobic cores cause loss of RNA binding capability" ^? ^pjg q
[0237]
[0166] To examine whether the I197V substitution affects NOVAl’s RNA binding ability, we performed CLIP (cross-linking and immunoprecipitation) to compare NOVAI genome-wide binding maps in Noval^u,'^uand Noval^^ midbrain at P21 (Fig. 6d). Across three biological replicates for each genotype, we identified 26,155 binding peaks in Noval^u'^ntand 27,720 in Novalwt'wtmice, respectively. NOVAI binding peaks were detected primarily on introns and 3’UTRs (Fig. 6e). Thebinding motifs were highly enriched for the known NOVAl binding sequences (UCAU), and the genomic distribution of NOVAI binding to this motif was highly similar between genotypes (Fig. 6f). The number of tags in the detected CLIP peaks were highly comparable (R =0.982), with only minor differences in low-count peaks (a total of 250 peaks: average CLIP tags / peak 16.3 v 't Noval^w^u, 13.1 in Novalwt / wt, p-value<0.01, |log2FC|>l) (Fig. 6g, Table 3). In both genotypes, NOVAI bound transcripts were strongly enriched for those encoding proteins involved in the synaptic signaling, transmission and secretion (Fig. 6h). The nearly identical characteristics between genotypes were also observed in the genomic distribution of NOVA 1 binding peaks, enriched motifs, and peak correlations in cortex and cerebellum at P21 (Fig. 10). Notably, NOVAl bound transcripts identified by CLIP analysis were enriched for behavior- and synapse-related categories across midbrain, cortex and cerebellum (Fig. 6h, Fig. lOe, l j).
[0238]
[0167] We performed in vitro RNA binding assays to compare the RNA binding characteristics of the NOVAI proteins. Full-length proteins were purified from E. coli, and RNA oligonucleotides with the 4?
[0239] NOVAl binding site (UCAU repeat sequence with stem-loop structure ) were used for gel shift assays. Each NOVAl protein bound ~ P-labeled UCAU-RNA and caused a dose-dependent shift inmigration (Fig. 6i), with indistinguishable binding dissociation constants (Kd for NOVA 1 protein was 105.83 ± 8.3, and for N0VAlwtprotein was 104.8 ± 11.6; (Fig. 6j). Taken together, these in vivo and in vitro studies reveal not only the I197V variant’s resilience in maintaining the biophysical features of RNA binding with minimal global disruption but also its remarkable conservation of overall function. However, this variant exerts specific effects on alternative splicing (AS) as explored in the following section.
[0240] Humanized NOVAl mice exhibit alternative splicing changes of specific genes
[0241]
[0168] The role of RNA binding proteins in mRNA processing is often influenced by various factors beyond their RNA binding capability. Several studies suggest that competing or cooperating proteins, as well as non-protein factors like metals or ATP, play roles in determining the downstream effects on target mRNAs3”-
[0242]
[0243] . The function of NOVA proteins on specific mRNA targets, such as splicing or stabilization, also depends on their association with other proteins^^’^’^. Although the I197V substitution has little effect on the RNA affinity and sequence specificity, X- ray crystallography analysis and protein structure prediction suggested subtle changes in amino acid interactions within the KH domain (Fig. 11, see Discussion). These changes may affect protein-protein interactions or NOVA dimerization^, leading us to investigate their impact on RNA regulation, particularly on AS.
[0244]
[0169] We first determined the brain regions for AS analysis based on the expression patterns of NOVAl. The immunostaining of mice brains at E18.5 and P21 shows that NOVAl is expressed - O-throughout brain development across various regions (Fig. 12a, Fig. 13). NOVAI is most highly expressed in midbrain, low-level across cortical layers, sparse in striatum and hippocampus, and intermediate level in the granular layer in cerebellum. Analysis of NOVA protein expression in P21 mouse brain using N0VA1 and panNOVA antibodies revealed that among the various isoforms of N0VA1 and NOV A2, a specific N0VA1 isoform (N0VA1 without ExonT^) is highly expressed in the midbrain, particularly in the periaqueductal gray (PAG) region (Fig. 12b, Fig. 14). The PAG is involved in a broad range of physiological and behavioral functions, including defense reaction, pain and anxiety,
[0245]
[0246] fear, micturition, and vocalization0~ ~. It is thought to integrate sensory signals from the periphery, acting as a control center for behavioral regulation
[0247]
[0248] . AS analysis was performed in the mouse midbrain, the region where N0VA1 is most highly expressed.
[0249]
[0170] AS analysis in the P21 midbrain revealed that several AS events were specifically altered in Novalu / f umice compared to Novalw^wtcontrols (Fig. 12c). Specifically, 720 events showed significant changes (delta PSI (di) more than 5% (\dI\>0.05, ><0.05)) (Table 4). Among the most intriguing changes were tandem cassette exons in Fnbpll (formin binding protein 1 -like, exons 10 and 11), a gene implicated in human intelligence 7Q RO and cassette exons in Itprl (inositol 1,4,5-trisphosphate receptor type 1, exon 41), a receptor that mediates calcium release from the endoplasmic reticulum (Fig. 12d). Although the effect of the I197V substitution on AS was smaller than the effects observed in previous NOVA knockout studies^,44,64, tesefindings highlight the variant’s subtle but specific molecular impact. We further investigated the NOVAI 1197V variant’s influence on AS by cross-referencing the CLIP and gene annotation datasets and assessing the statistical significance of the results using random resampling methods, as detailed in the section below.
[0250]
[0171] To assess the molecular impact and potential direct effect of the I197V substitution, we examined NOVAI binding on AS transcripts by analyzing the CLIP dataset. Among the 720 differential AS events, 258 (41%) had NOVAI binding peaks on their transcripts (Figure 12c). Gene annotation analysis revealed that the 630 transcripts with differential AS events are enriched in processes related to cell projection organization or chromatin remodeling, and the transcripts with NOVAI bindingpeaks show further enrichment in processes involving cell projection, morphogenesis, and synaptic function (Fig. 12e).
[0251]
[0172] To explore the potential effects on behavioral control vuNOVAl^11^11mice, we examined genes associated with behavior among transcripts showing differential AS events. Of the 630 transcripts with differential AS events, 27 transcripts were associated with the behavior category (Table 5). Remarkably, the vocal behavior term showed the highest ratio of gene coverage, with four of the 22 genes annotated for vocal behavior (Auts2, Myhl4, Nrxn2, Srpx2)^~^ being differentially spliced in Noval^u' ^umicerelative to Novaf^^ mice (Fig. 12f, Fig. 15). This number of vocalization-related genes was significantly higher than expected by chance, exceeding the 5% significance level in random resamplings tests (1000 resamplings, mean 0.71, median 1; Fig. 16).
[0252]
[0173] Interestingly, many genes involved in vocalization, including Foxp2, Celf6, Auts2, Nrxnl-3, and Shank! -3, showed reproducible N0VA1 binding peaks on their transcripts across multiple brain regions (Fig. 17). Given that Pasilla, the fly ortholog of mammalian NOVA1 / 2, regulates splicing of target transcript in an experience-dependent manner
[0253]
[0254] , it is plausible that these vocalization related transcripts are similarly affected in a context-dependent manner, such as in response to sensory cues from surrounding environment. Together, these findings indicate that Nova l^w^umice with 1197V substitution exhibit changes in specific splicing events in the brain, including those in genes involved in animal behavior, particularly vocalization.
[0255] Humanized NOVAI mice pups have altered vocalization
[0256]
[0174] We next examined whether Noval^u / ^umice with I197V substitution exhibit changes in vocalization behavior. Prior studies in birds, fish, and mammals have shown that all vocal species have a conserved midbrain / brainstem vocal motor pathway
[0257]
[0258] . Of particular importance in the vocal circuit in the mammalian brain is the periaqueductal gray (PAG), which plays a central role in the neural basis of primate vocal production'-'^. The PAG projects to brainstem respirator}- premotor and vocal motor nuclei, including the nucleus ambiguous (Amb), which directly innervates the l
[0259]
[0260] arynx^^O Inhibition of PAG or Amb function results in loss of innate v
[0261]
[0262] ocalization^-^whqestimulation of the PAG induces vocalizations in both primates and mice^^,suggesting that the PAG and downstream brainstem circuits are essential for vocalization. W7e found that NOVAI is highly expressed in the midbrain, including the PAG and Amb (Fig. 12a- b. Fig. 13-14). In addition to Vova7^M / ^zw-specific actions on RNA splicing, we also found that many of the genes reported to be involved in vocal behavior (14 out of 22 transcripts) were NOVAI binding targets (Fig. 17). These data str engthen the possibility of a relationship between Noval^u'^uand vocalization, suggesting that vocalization studies in these mice would be valuable.
[0263]
[0175] We first compared vocalizations from pups of Noval^u / ^u, Noval^u^wt(heterozygous) and Novalwt'wtmice. When pups are isolated from their mothers, they produce isolation induced ultrasonic vocalizations (USVs), which are distress calls that attract their mothcr^-^ '2 We recorded USVs from 7-day-old mice pups for 5 minutes in a dark sound isolation chamber (Fig. 18a, Table 6, Table 7). Following a previous protocol, we classified the syllables into four types (simple [s], upward [u],downward [d], multiple [m]), based on the direction and number of pitch jumps that separate notes within a syllable. Each syllable was analyzed for peak frequency measures (frequency value of start: Fqstart, minimum: Fqrnin, maximum: Fqmax, end: Fqend, mean: Fqmean), variance (Fqvariance), bandwidth, duration, amplitude and purity (Fig. 18b).
[0264]
[0176] The average number of USVs were 77.9 ± 6.0 per minute in Novalwt / wtand 62.7 ± 6.4 per minute inNoval^U / ^u(mean ± standard error), with no significant difference between genotypes (Table 8). However, several changes in USV features which are thought to be important for mouse vocal communication^ were observed. The “s” syllable (without pitch jumps), the most abundant syllable type, showed a trend of increased percentage within the total syllables in
[0265]
[0266] pups, while its amplitude was significantly lower
[0267]
[0268] in than that of Novalwt / w tpups. The amplitude of syllables containing pitch jumps (“u”, “d” and “m”) did not differ between genotypes, but these percentages were decreased in Noval^u^upups, particularly for the “u” and “m” syllables (Fig. 19a, Table 8).
[0269]
[0177] Pup isolation-induced USVs in mice show a bimodal distribution in their peak frequency (Fq), especially at 5-9 days of age, consolidating to a single peak as they grow^. In our study, 7 -day- old pups exhibited a clear bimodal distribution in the Fq parameter regardless of the genotypes (Fig. 19b). We assessed bimodality in Fqmax, showing clear separation using an Ashman’s D score
[0270]
[0271] This analysis indicated a clear separation in two peaks for the “s”, “d” and “m” syllables (D>2.0), but not for the “u” syllable (£>=1.603) (Fig. 18c). In order to separate these two peaks, we fitted two Gaussians
[0272]
[0273] for each syllable type showing bimodal distribution, comparing USV characteristics for each peak (Fig.
[0274] 18c, Table 9). Individual syllables were classified into those with Tow’ or ‘high’ Fqmax according to the cutoff value (the intersection point of each distribution). In the jump syllables of Novani / '^upups, the proportion of low Fqmax decreased, while the proportion of high Fqmax increased (Fig. 18d, Table 10). We also tested the bimodality and syllable ratio with Fqmin and confirmed the same trend (increased ratio in high Fq in Novan'^nipups). Heterozygous Novalwt^mpups showed intermediate values between Novai^u / ^uand Vow
[0275]
[0276] pups for these parameters, suggesting that the effect of the I197V substitution in NOVAI protein on pup USVs is dosage dependent. These observations demonstrate distinct changes in the vocalizations of AGvo / ^M / ^Mpups.
[0277]
[0178] The vocalizations of pups isolated from the nest are known to be able to influence maternal behavior in mice^, 102,10.3 p0 exp]orethe potential significance of vocal motor changes in Noval^l^lupups on their mothers, 'e tested whether mother mice were more attracted to recordedvocalizations from each genotype. We set up an experiment using a three-room box connected by passageways, with speakers placed in each of the two end rooms (Fig. 20a). Neonatal vocal recordings (from Noval^u / ^umice or Novalw^wtcontrol mice) were played from each speaker, and cither a Noval^u,'^uor a NovaJwl / wlmother mouse was placed in the center room (Fig. 20a-c). Regardless of the mother’s genotype or the neonatal vocal recordings, no significant differences were observed in the mother’s orientation toward the vocalizations of pups (Fig. 20d-e). Thus, changes in vocal quality in Noval^u^lupups had no impact on the behavior of the mother mice in this assay.
[0278] Humanized NOVAI adult mice have altered vocalization
[0279]
[0179] Adult mouse USVs are often produced in long, continuous sequences consisting of the same four major syllable types (Figure 18b), particularly during courtship^, 104,105 \( / eexplored whether changes in courtship vocalization behavior would occur in Noval^ni^umice, using a previously established paradigm^ J05, 106
[0280]
[0281] e|! Cj courtship USVs from adult male mice by exposing them to adult female mice in estrus (Fig. 18e, Table 11, Table 12). In this context, more than 90% of the vocalizations come from the m
[0282]
[0283] ale^?. The average number of courtship-induced USVs did not differ between genotypes: 144.8 ± 24.0 per minute in Novalwt'wtand 165.5 ± 27.0 per minute in Noval^u / ^u(mean ± standard error) (Table 13). Wenext examined short-duration and long-duration syllables1
[0284]
[0285] 1, which has been previously shown to be bimodal especially for “s” syllables^. Two Gaussian distributions were fitted to characterize each class of USVs (Fig. 18f). The cutoff duration (the intersection point of each distribution) between short and long USVs was determined to be 44 ms, with 29% ofu‘s” syllables categorized as long duration. The average duration of short- and long-“s” syllables was 22.6 ± 9.1 ms and 69.0 ± 40.1 ms, respectively. Novai^u / ^umice tended to have a lower frequency characteristic for long- duration “s” syllables, where the differences for Fqstart, Fqend and Fqmean were significantly lower compared to controls (Fig. 18g, Table 14). These changes were not observed in pup isolation induced USVs, indicating that this effect is developmentally specific and / or context dependent.
[0286]
[0180] We next examined peak frequencies. In adult USVs, the Fqmax parameter exhibited a unique distribution. High Fqmax syllables were those that exceeded 100 kHz, and constituted a distinct population from the main population (Fig. 19c, d). Syllables with High Fqmax of wild type animals are often observed in harmonic syllables, containing pitch jumps and some simple syllables^ 109
[0287]
[0288] 19e). Fitting two Gaussian distributions to the data, the High Fqmax was found to be 107.1 ± 2.1 kHz, whereas the main peak was 80.8 ± 11,5 kHz (mean± standard deviation) (Fig. 18h). Using this definition, the High Fqmax syllables accounted for roughly 8.5 % of total USVs observed. There were no significant differences in syllable composition with High Fqmax among genotypes (Table 15).However, the Fqvariance of High Fqmax syllables containing pitch jumps (“d”, “if ’ and “m”) was significantly greater n. NovaI^U / ^umice (Fig. 18i, Table 16). Fqvariance tends to increase with syllable complexity^®, with simple syllables like “s” having lower values and more complex syllables like “m” with multiple jumps having higher values. This suggests that Noval^u'^111mice produce more complex high frequency USVs than Novalw,'wtmice. These findings demonstrate that vocal behavior is altered in both pups and adults in Noval^u'^umice, resulting in unique vocal characteristics.
[0289] DISCUSSION
[0290]
[0181] In line with studies of genetic variants that have played a role in the evolution of modem humans^®’ HM12,weinvestigated the biological effect of a single amino acid substitution, I197V in N0VA1, which is unique to modem humans. By analyzing Noval^ni^umice carrying this allele, we identified molecular changes in alternative splicing in the brain, including brain regions associated with vocal behavior, and identified changes in vocalization patterns in pups and adult mice. These findings suggest that during human evolution, the II 97 V substitution in NO VAI protein may have contributed to the development of neural systems involved in more complex vocal communication.
[0291]
[0182] The importance of N0VA1 in mammals is evident from the lethal phenotype of Nova I 1 £ knockout mice31and the neurological symptoms caused by NOVAI haploinsufficiency in humans30. This significance is further highlighted by the high conservation of the NOVA! protein in mammals. Interestingly, the NOVAI gene harbors an Ultra Conserved Element (UCE; uc.359) at the end of the 3’
[0292]
[0293] UTRI 13,114 additional high conservation extending upstream from the NOVAI UCE to most of the 3' UTR and terminal exon encoding N0VA1 KH2 and KH3 domains. This underscores the unique nature of the I197V variant, which occurred within a region of the genome resistance to change.
[0294]
[0183] Previous studies have confirmed the evolutionary restriction of N0VA1 variants, including the I197V variant (term I200V in one study^). We support this analysis and have expanded upon it with larger human sequence datasets across diverse ethnic groups, and from methods that infer selection coefficients from ancient samples^®’-'®. These results confirm that N0VA1 has undergone strong positive selection, and that the 1197 V variant is part of an evolutionary selective sweep in the emergence of Homo sapiens.
[0295]
[0184] The observation that the II 97V N0VA1 allele is nearly fixed across human populations (Figure 2a) suggests that it emerged and increased in frequency well before the divergence of ancient human lineages. The earliest divergence among modern human groups - the San divergence- is currently- estimated to be around 250 - 300kya, prior to their migration to Eurasia^ J 15 Unlike more recentselective sweeps, such as the LCT locus, which is dated around lOkya, and is population-specific, the N0VA1 variant is part of an older, more widespread sweep. These older sweeps, shared across modem human populations, may leave subtler genetic signatures that require novel detection methods. This suggests that the ancient N0VA1 selective sweep may represent part of a broader set of undiscovered ancient sweeps.
[0296]
[0185] One possible explanation for the changes in vocal behavior observed in Noval^u'^umice could be molecular changes in midbrain and brainstem vocal pathways, which express high levels of NOVA 1 and are involved in regulating innate vocalizations (USVs), including breath coordination, timing, and amplitude^ ^-118
[0297]
[0298] alternative possibility is that changes occurred in more recently evolved cortical vocal regions, which control pitch, frequency modulation, and duration (the Kuypers / Jurgens hypothesising and the volitional articulatory motor network^ 120) Given that NOVA 1 is expressed in the mouse cortex, predominantly in inhibitory' neurons*^, it is plausible that the II 97V substitution affects cortical regulation of vocalization.
[0299]
[0186] Notably, Noval^u / ^umice exhibit qualitative changes in vocal characteristics compared to control mice both in pups and adults, despite producing a similar number of calls. Urese findings suggest that the vocalization changes in Noval^u'^nimice are not simply the result of alterations in general motor performance. This idea is supported by other observations showing that Noval^u^umice perform similarly to control mice in motor function tests, such as the rotarod, and display comparable locomotion activity levels in the Y-maze test (Fig. 21). Additionally, the Y-maze test results indicated that Noval^u,f^umice had spatial working memory' comparable to that of control mice.
[0300]
[0187] The changes in vocalization in Noval^u^umice varied in a development- or context-dependent manner. It has been reported that high-frequency (Fq) USVs are emitted more frequently by male adult i n i
[0301] mice during social interactions, and that female mice are attracted to male mice that emit more complex USV s. Given these reports, the increased the proportion of higher Fq USVs in pups, and the increased complexity of these USVs in adults, may potentially offer social advantages in mice. However, since auditory' frequency' resolution in mice has been reported to be limited
[0302]
[0303] and our preference experiment using humanized NOVAI pup USVs showed no significant preferences in the mother’s responses, it remains unclear to what extent other mice can recognize these vocalization changes. It should be noted that we were unable to address theeffect of the 1197V substitution on the vocalizations of adult female mice, as our study focused on a courtship-induced vocalization paradigm that predominantly elicits USVs from male mice. However, recent studies suggest that female mice also vocalize under certain experimental conditions or social contexts^ 24, 125 Future studies will benecessary to investigate the effects of the I197V substitution on USVs in female mice, as well as adult female preferences to USVs in adult Novai^u' ^umale mice.
[0304]
[0188] Interestingly, the vocalization changes observed in Noval^u,^umice share some similarities with those observed in humanized Foxp2 mice (with two human-specific substitutions). In both cases, the changes were developmental or context-dependent, and included a decrease in peak frequency in simple syllables and modulation of high frequency regions in complex syllables 1 —21 (Table ]7) Conversely, male mice with a humanized Foxp2 mutation produced simpler song bouts with more “s” 1 1
[0305] syllables >. These observations may indicate a common or related molecular alteration in the neural circuits involved in the USV production between humanized Noval mice and humanized Foxp2 mice. Future studies should aim to identify the molecular and neural basis of these alterations, as well as the physiological significance of these vocalization changes in the context of social behavior.
[0306]
[0189] Our molecular analysis showed that the sequence-specific RNA binding of N0VA1 was unaffected by the human substitution and that steady-state gene expression levels in the brains of Noval^u^umice were nearly identical to those of wild type mice. However, we detected alternative splicing changes in several transcripts associated with vocalization. The expression pattern of NOVA 1 in the brain and the enrichment of its target transcripts to specific biological pathway support a link between N0VA1 function and vocal behavior. Uncovering the precise molecular mechanisms underlying the phenotypes in Noval^u / ^umice will require further study of the neural circuits for vocalization, as well as on regulatory factors influencing NOVA protein function. This study sets the groundwork for understanding molecular mechanisms driving the evolution of human vocal communication.
[0307]
[0190] Biochemically, NOVA proteins harbor three KH domains responsible for sequence-specific RNA- b
[0308]
[0309] inding42>45,46,ainjno acK] 197 located in the KH2 domain. Although the 1197V substitution in NOVAI alters the hydrophobic core of the KH domain, it does not lead to a loss of RNA-binding capacity or functional attenuation, in contrast to other KH domain point mutations^^-- 6
[0310]
[0311] is supported by the unchanged global gene expression levels observed in the brains oiNoval^w^umice, while NOVA 1 knockout mice (which exhibit postnatal lethality3show significant expression changes of key neuronal genes in the midbrain at E18.5 (Fig. 7e, 7f, Table 18). KH domains harbor three alphahelices (H) and three beta-sheets (S) (S1-H1-H2-S2-S3-H3; Fig. l la-b); the first and second alpha helices (H1-H2) determine single-stranded RNA binding specificity^'3’^’3, and may also be involved in protein- protein dimerization^ (Fig. 1 lc-d). Protein structure predictions suggest that the addition of a single carbon atom in valine 197 extends its ability to interact with several nearby amino acids ( 185Ilein Hl; 232Ala, 235Leu, 236Ile, 239Lys in H3), allowing the KH2 domain to gain contacts with Hl and 239Lys in H3. Thus, while the I197V substitution does not change sequence specificity or binding affinity of N0VA1 with RNA (Fig. 6i-j), it may affect KH domain dimerization, or may have undiscovered effects on protein-protein interactions^. Interestingly, the amino acid corresponding to N0VA1 amino acid 197 is also an isoleucine in the related proteins FMRI and hnRNP E1 / E2 / K, but is a valine in both human and mouse N0VA2^^ (Fig. 9). Functional differences between N0VA1 and N0VA2 in mice^^Tmayrcf|cc( structural and functional differences in their respective KH domains.
[0312]
[0191] In summaiy, we analyzed a single amino acid unique to modem humans in the RNA binding protein NOVAE and examined its biological effects in vivo by introducing this amino acid in mice. N0VA1 is highly intolerant to changes in amino acid sequences during evolution with the exception of this single amino acid change in humans. We propose that this change was part of an evolutionary sweep associated with specific changes in the neuronal transcriptome and vocal communication.
[0313] Methods
[0314]
[0192] Animal experiments
[0315]
[0193] All procedures were performed according to the guidelines of the Institutional Animal Care and Use Committee (IACUC) at Rockefeller University. C57BL / 6J (stock no. 000664) mice were obtained from the Jackson Lab. Novai^u' ^umice generated in this study were backcrossed to C57BL / 6J strain at least 8 times. Mice were housed in a 12-h light / dark cycle, up to 5 mice per cage. Male or female mice aged 7 days to 20 weeks were used for animal experiments, as described. Littermates of the same-sex were randomly assigned to experimental groups.
[0316]
[0194] Generation of Noval^u^umice
[0317]
[0318]
[0195] Noval^u^nimice were generated by directly injecting the sgRNA'Cas9 RNP with a singlestranded repair template DNA (ssDNA) into C57BL6 zygote to substitute isoleucine to valine at amino acid 197 of mouse NOVAE gRNA and the ssDNA were designed as follows. gRNA (5’-TGCTACTGTGAAGGCTATAA-3’ (SEQ ID NO: 17)): overlapping the DNA sequence (mmlO / chrl2: 46,700,902- 46,700,904) of the mouse Noval genomic locus encoding the 197^ amino acid of NOVA 1. ssDNA: 140 nt length DNA homologous to the NOVAI locus with a nucleotide substitution (A to_G) to cause an amino acid change from isoleucine to valine at the 197^ position. Two silent mutations were also designed to create Btsal restriction enzyme recognition site for genotyping.
[0319]
[0196] Genomic DNA was extracted from the tail of the F0 animals, and the DNA corresponding to the area around the 197^ amino acid was amplified by PCR and subsequently cloned into a plasmid for determining the sequence of the modified allele. Genomic sequence analysis revealed that among 13FO animals, 8 animals harbored the designed allele (with three nucleotide substitutions: one causing 1197V amino acid substitution, two for restriction enzyme recognition site for genotyping (not causing amino acid changes)). Animals carrying the designed humanized Noval allele were crossed to the wildtype C57 / BL6 mice, and this process was continuously repeated for subsequent generations to eliminate possible off-target mutations.
[0320]
[0197] Genome sequencing analyses were performed on the possible off-target sites of the gRNA used (10 potential off-target loci with mismatches outside of the PAM+12mer core sequences; predicted by CRISPR direct: crispr.dbcls.jp) to compare sequences of control (wild type) and humanized NOVA! mouse. For all the potential off-target sites, the sequence in the humanized N0VA1 mice were identical to control mice and the reference genome, with I197V substitutions being the only detectable edits (Table 4d-e).
[0321]
[0198] For routine genotyping, sequences around the genomic DNA encoding the 197^ amino acid were amplified by PCR subsequently digested with the Btsal restriction enzyme. Each Mouse genotype; wild type Novalwt / Wf'). humanized Noval homozygous (Noval^u / ^ni), heterozygous (Noval^11^) was determined by band size obtained by electrophoresis. Siblings obtained by crossing heterozygous parents were used in the experiment. Primers used for typing are shown below'.
[0322] NOVAlhu-Fw'd: 5’- ccctcttttgacatgctggt -3’ (SEQ IDNO:18)
[0323] NOVAlhu-Rvs: 5’- cataaggagatccggttgga -3’ (SEQ ID NO: 19)
[0324] DNA band size after restriction enzyme treatment: wild type (613 bp), homozygous (389 bp and 224 bp), heterozygous (613 bp / 389 bp + 224 bp) (see Fig. 4c).
[0325]
[0199] Antibodies
[0326]
[0200] Primary antibodies used for immunohistochemistry and western blotting 'ere as follow's; rabbit anti-NOVAl [EPR13847] (ab 183024, abeam), rabbit anti-NOVAl C-terminal [EPR 13848] (ab!83723, abeam), human anti-pan NOVA (anti-Nova paraneoplastic human serum) and rabbit anti-ATCB (ab8227, abeam).
[0327]
[0201] Immunohistochemistry'
[0328]
[0202] Postnatal day 0, three or twelve-week-old mice w'ere perfused with PBS and 4% paraformaldehyde (PF A), and the brain was dissected. Dissected brain was further fixed. Overnight by 4% PFA at 4°C. The solution was sequentially replaced with 15% sucrose / PBS and 30% sucrose / PBS, embedded with OCT compound, and stored at -80°C until use. Frozen brains w'ere sliced into 30-50 pm thick sections in a cryostat (CM3050S, LEICA). Slices were w'ashed three times with PBS at room temperature (RT), incubated in 0.2% Triton X-100 / PBS for 15 min at RT, blocked in 1.5% normal donkey serum (NDS) / PBS for 1 h at RT, incubated overnight at4°C with primary antibody in 1.5% NDS / PBS, then incubated in Alexa Incubated with 488, 555 or 647 conjugated donkey secondaryantibody. The nuclei were stained using 4',6-diamidino-2- phenylindole (DAPI) solution (1 jig / ml). Images of specimens were collected with a BZ-X700 (KEYENCE) microscope.
[0329]
[0203] Western blotting
[0330]
[0204] Each dissected brain region (cortex, midbrain and cerebellum) of P21 mouse brains was lysed in RIPA buffer (50mM Tris-HCI; 150mM NaCl; 0.1% SDS; 0.5% sodium deoxycholate; 1%NP- 40). Extracts were separated by SDS-PAGE, and subjected to immunoblotting using the antibodies described above. Quantification of western blots were done with ImageJ (vl.53). Each band signal was quantified and normalized with ACTB signal to control for differences in loading.
[0331]
[0205] Electrophoretic Mobility Shift Assay (EMSA)
[0332]
[0206] Protein purification
[0333]
[0207] The genes encoding each NOVAI protein (NOVA1wtand NOVA1^U) were cloned into the pGEX6pl vector and expressed in E. coli BL21 strain. N-terminally GST (Glutathione S- Transferase) fused NOVAI was induced by the addition of IPTG (final cone. 0.1 mM) for 4hr. Pelleted cells were sonicated, and then incubated in the presence of Triton-X (final cone. 1%, 30 min). Cleared supernatant was collected after centrifugation (12,000 x 10 min 4°C). After incubating with Glutathione Sepharose beads (GE Healthcare Biosciences, 17075601) for 30 minutes, the mixture was washed three times with PBS. The GST tag was cleaved from the N0VA1 protein by PreScission Protease treatment (GE Healthcare, 27-0843-01: 50 mMTris-HCL 150mMNaCl, 1 mMEDTA, 1 mMDTT,pH7.5,4cCfor4hr) to obtain purified NOVA 1 protein. The concentration of each purified NOVAI protein was determined by SDS-PAGE followed by GelCode Blue staining (Thermo Fisher Scientific, PI24590) using BSA (Sigma-Aldrich, B8667) as standards.
[0334]
[0208] Single strand RNA probe preparation
[0335]
[0209] The single-stranded RNA probe was designed as previously^. The following single strand RNA were synthesized by IDT:
[0336] CCTTATCATGCTGACTCACGTCATTTCATCTCATCAAGGGAGTCAGTGGGATA (SEQ ID NO:20)
[0337]
[0210] Synthesized RNA was first incubated at 80°C for 10 min, then rapidly incubated on ice, mixed with [gamma-32P] ATP (3000 Ci / mmol, 10 mCi / ml) (Revvity, BLU502A), and labeled at the 5' end by T4 polynucleotide kinase treatment (New England BioLabs, M0201 S). The labeled probes were purified by G-25 column (VWR, 95017-621) and diluted to the appropriate concentration with water.
[0338]
[0211] Binding assay
[0339]
[0212] Purified NOVA! protein (0.5-2 pmol / reaction) and labeled RNA probe (0.08 pmol / reaction) were mixed under the following buffer conditions: lOmM HEPES, 3mM MgC12, 150mM NaCl, 5% -SO-Glycerol, O.lmM DTT, O.lU / pl RNase OUT (Thermo Fisher Scientific, 10777019)., lOmg / ml Yeast tRNA (Thermo Fisher Scientific, AM7119), Img / ml poly dl-dC (Thermo Fisher Scientific, 20148E). The binding reaction was performed at 22°C for 60 minutes, and the reaction solution was mixed with loading dye and run on an 8% acrylamide native gel at 150 V for 3 hours to separate RNA-protein binding. After the gel electrophoresis, the gels were dried by a gel dryer and autoradiographs were detected. Quantification of each signal was performed by ImageJ, and the Kd value of each NOVA1 protein for the RNA probe w'as calculated using Prism software (https: / / www.graphpad.com / features).
[0340]
[0213] RNA-seq library preparation and analysis
[0341]
[0214] For RNA-seq, samples included dissected cortex, midbrain and cerebellum at P21, or dissected midbrain at El 8.5 of Noval^u, / ^uand Novalw^wtmice. The mRNA-seq library was prepared from RNA extracted with Trizol following the Illumina TruSeq protocol of poly A selection, fragmentation, and adapter ligation. Multiplex libraries were sequenced as 125 nt paired- end runs on the HiSeq-2500 platform at Rockefeller University Genomic Core. These raw datasets and processed data files have been deposited with Gene Expression Omnibus (GSE253297).
[0342]
[0215] Cross-linking Immunoprecipitation (CLIP)
[0343]
[0216] NOVA 1 -CLIP was performed in P21 dissected cortex, midbrain and cerebellum of Nova1hu / huand Novalwt'wtusing each three biological replicates. Tissues were dissected in PBS, triturated using 20G needle and crosslinked three times on ice for 400 mJ / cm2 using Stratalinker. Crosslinked material was collected by centrifugation, resuspended in wash buffer (IX PBS, 0.5% NP-40, 0.5% deoxycholate and 0.1% SDS with protease inhibitor), and subjected to DNase (RQ1 DNase: Promega) and RNase (RNase A: Affymetrix) treatment at a final dilution of 1: 20,000 for 5 min. The lysate was clarified by centrifugation at 20,000 x g for 20 min. The supernatant was used for immunoprecipitation with 200 uL of Protein A Dynabeads (Invitrogen) loaded with 18 pg anti -Nova 1 antibody (abeam) for 2 hours at 4°C. The samples were washed as follows: twice with wash buffer, twice with Nelson stringent wash buffer (15mM Tris pH 7.4, 5mM EDTA, 2.5 mM EGTA, 1% Triton X-100, 1% Sodium deoxycholate, 0.1% SDS, 120 mM NaCl, 25mM KC1), twice with Nelson high salt buffer (15mM Tris pH 7.4, 5mM EDTA, 2.5 mM EGTA, 1% Triton X-100, 1% Sodium deoxycholate, 0.1% SDS, 1M NaCl), twice with Nelson low' salt buffer (15mM Tris pH 7.4, 5mM EDTA), and twice with PNK wash buffer (50mM Tris pH 7.4, 10 mM MgC12, 0.5% NP-40). RNA fragments were dephosphorylated using FastAP Alkaline phosphatase (Thermo Fisher Scientific) and subjected to 3' ligation overnight at 16°C with a preadenylated linker (preA-L32) using truncated KQ T4 RNA Ligase 2 (NEB). The RNA-protein complexes were labeled with32P-γ-ATP using T4 PNK (NEB) and subjected to SDS-PAGE and transfer to nitrocellulose membrane. Appropriate regions of the membrane w'ere cut out and RNA w'as extracted according to the following conditions: 100mM Tris pH7.5, 50mM NaCl, 10mM EDTA, 7M Urea withproteinase K. RNA was purified by phenol-chloroform extraction method. Cloning
[0344] was performed using the BrdU-CLIP protocol. Briefly, the reverse transcription reaction was performed using Superscript III (Thermo Fisher Scientific), and the cDNA was BrdU-labeled by including BrdU in the reaction solution. Immunoprecipitation was performed with 5 pg anti-BrdU antibody (abeam) and 25 pg protein G Dynabeads per reaction (45 min at room temperature), followed by washing with the following solutions (including Denhardt’s solution): once with IP buffer (0.3x SSPE, ImM EDTA, 0.05% Tween 20), twice with Nelson low salt buffer, twice with Nelson stringent wash buffer, twice with IP buffer. After eluting the cDNA, BrdU- immunoprecipitation was performed again under the same conditions. cDNA was circularized on beads using CircLigase II (Epicentre) and PCR was performed using Accuprime Pfx supermix (Thermo Fisher Scientific) and Syber Green until RFU 250-500. PCR products were purified using Agencourt AMPure XP (Beckman Coulter) and concentrations were measured by TapeStation. High-throughput sequencing was performed at the Rockefeller University Genome Resource Center. These raw datasets and processed data files have been deposited with Gene Expression Omnibus (GSE253296).
[0345]
[0217] Bioinformatics
[0346]
[0218] Paired-end reads from RNA sequencing were aligned to the mouse genome (mm 10) builds of the mouse genome using OLego (v1.1.7) (zhanglab.c2b2.columbia.edu / index.php / OLego)176. Mapped reads were counted using gapless (for inference of transcript structure) and countit (for quantification of gene expression and alternative splicing) in Quantas (v1.0.9) (zhanglab.c2b2.columbia.edu / index.php / Quantas)177. All reads mapping to transcripts were included in the differential expression analysis using edgeR
[0347]
[0348] . The data set for Noval knockout mouse (El 8.5 midbrain) was kindly provided by Dr. Yuhki Saito. The data are available from GEO submission GSE69711. Data visualizations were done using R (v4.2.0). Correlation matrix was visualized using corrplot package. PCA analysis was performed using FactoMineR and factoextra packages and visualized using ggplot2 package. Sequencing tracks were visualized using Integrative Genomic Viewer (IGV, v2.13.0).
[0349]
[0219] Quantification of splicing for annotated cassette exons was performed using the Quantas pipeline (v1.0.9)177. In brief, the inclusion level of each cassette exon (percent-spliced-in: PSI) was calculated from the number of supporting exon junction reads for the inclusion and skipping isoforms. Only quantifications with > 20 supporting junction reads were used for downstream analysis. Gene annotation analysis was performed using Metascape (v3.5.20240101)
[0350]
[0351] . The expressed genes in the P21 midbrain (filtering lowly expressed genes by edgeR) were set as background for the analysis. Attribution for genes was performed using Gene Ontology (GO) resource in MGI (6.24) (informatics.jax.org / ).
[0352]
[0220] CLIP reads were processed using the CLIP Tool Kit (CTK, vl.1.3) as described previously I 70. Briefly, raw reads were filtered for quality and demultiplexed using indexes introduced during the reverse transcription reaction. PCR duplicates were collapsed and adapter sequences removed. Reads were mapped to the mmlO build of mouse genome using novoalign (v3.09.02) (www.novocraft.com). Mapped reads were further collapsed for potential PCR duplicates by coordinates and taking into consideration the degenerate barcodes introduced during the reverse transcription. Only unique CLIP tags were used for subsequent analyses. We performed three biological replicates per sample. All scripts used in the analysis including the peak finding algorithm and more information can be publicly obtained at (zhanglab.c2b2.columbia.edu / index.php / Standard / BrdU-CLIP_data_analysis_using_CTK). De novo motif analysis and motif density analysis were done using findMotifsGenome.pl and annotatePeaks.pl commands in HOMER (v4.11).
[0353]
[0221] Ultrasonic vocalization (USV) tests
[0354]
[0222] Isolation induced pup USV test
[0355]
[0223] To elicit isolation-induced USV, 7-day-old pups were isolated from their mother and littermates. Each pup was placed quietly on a small open-faced plastic plate in the sound attenuating chamber (15” x24” x!2” Igloo beach cooler with a tube for pumped air circulation input, no light). An ultrasonic microphone was suspended a small distance from the pup, and the USVs were recorded for 5 minutes. Between trials, the recording box was cleaned with 70% alcohol and distilled water, and allowed to fully dry before the next experiments. Vocalizations were recorded with UltraSoundGateCM16 / CMPA ultrasonic microphones connected to an Ultrasound Gate USGH amplifier. Recordings were saved using the Avisoft Recorder USG software (Sampling frequency: 250 kHz; FFT-length: 1024 points; 16-bits). All acoustic hardware was obtained from Avisoft Bioacoustics® (Berlin, Germany).
[0356]
[0224] Playback behavioral experiment
[0357]
[0225] Mothers rearing 7-day-old offspring were used in the playback experiment. We used a three-chamber box ( 12” x 23.5” x 15.5”) connected by a passageway through which a mouse could pass for the test. Each chamber at both ends was equipped wdth a speaker (Vifa ultrasonic speaker, Avisoft Bioacoustics), and a camera (Firefly S, FFY-U3- 16S2C-S, FLIR) was placed on the ceiling of the chamber to record the behavior of the mouse. The speakers were connected to an UltraSoundGate Player 216H (Avisoft Bioacoustics), using Avisoft Recorder USGH and had a frequency range (±12 dB as the maximum deviation from the average sound volume) of 25-125 kHz. We adjusted the loudness between the channels by controlling the level of the peak power before the experiment. Using two microphones, we made sure that both songs were audible at the entrance of both rooms so that the mother can respond to the songs but not loud enough that the microphones could detect the song being played in the otherroom. After the 10 min habitnation period, playbacks were triggered when the mouse broke an infrared sensor located in the center of the three-chamber box. One speaker on one side played one pup-USV recording and the other speaker simultaneously played another pup-USV recording both of which were previously recorded during the pup isolation induced USV test for 5 min.
[0358]
[0226] Pup-USV recording was prepared in Audacity® by stitching vocalizations from 4-5 pups for each genotype. These recording files contained an equivalent number of pup-USVs (Novalw^wt1689 USVs, Noval^u / '^u1561 USVs) and were confirmed to reflect the vocal characteristics of each genotype. After the first 5 min playback session, a second 5 min playback session was conducted after 1 min quiet period. In the second playback experiment, the two recordings playing from the speakers were switched to eliminate the possible preference by the location. During each 5 min period, the mother was allowed to explore freely in the box and the time she spent in each room was counted. The box was cleaned between experiments with 70% alcohol and distilled water and allowed to fully dry' before the next experiments.
[0359]
[0227] ( 'ourtship induced adult USV test
[0360]
[0228] The protocol for the courtship-induced vocalization test in adult male mice has been performed QO
[0361] according to previous studies^0with minor modifications. Briefly, adult male mice (8-12 weeks old) were sexually socialized by spending one night with a sexually mature female to enhance the male's motivational state to exhibit courtship USV behavior^. On the next day, female mouse was removed from the cage, and the male mice housed in the same cage until the test day. On the day of the recording, the males were placed in a new cage and then singly habituated in the sound recording environment (as described for pup USV test) for 15 min. The males were then exposed to adult female mice for 5 min. We used the females (8-12 weeks old) in estrus (selected visually for wide vaginal opening and pink surrounding). The test was conducted three times per mouse, one week apart, and a different female mouse was used as the stimulus each time to avoid familiarity effects. The order of mice tested each time was shuffled to avoid the possible order effects. Between trials, the mouse cage was cleaned with 70% alcohol and distilled w ater, and allowed to fully dry before the next experiments.
[0362]
[0229] USV analysis
[0363]
[0230] Acoustic waveforms were processed using a custom Python program called “Mouse Song Analyzer 2” (MSA2): available on the w'ebsite (github.com / Neurogenetics- Jarvis / MSA2^^’ ^5, 110, 1 aO, and analysis as previously described85,110Briefly, the software computed the sonograms from each waveform, threshold to eliminate the white noise component of the signal, and truncated for frequencies outside the USV song range (35-125 kHz). We used a criterion of 10 ms minimum to separate two syllables and 3 ms as the minimum duration of a syllable. The identified syllables werethen classified by presence or absence of instantaneous “pitch jumps” separating notes within a syllable into four categories: (1) simple syllable without any pitch jumps (“s”); (2) complex syllables containing two notes separated by a single upward (“u”) or (3) downward (“d”) pitch jump; and (4) more complex syllables containing a series of multiple pitch jumps (type “m”). Any sounds the software could not classify were put into “not identified: notIDd” category. The following spectral features were calculated automatically by MSA2 from the sonograms of each of the classified syllable types: Syllable duration, inter-syllable interval, standard deviation of pitch distribution, pitch (mean frequency), frequency modulation, spectral purity, and bandwidth. The fitting of Gaussian in the duration and peak frequency maximum (Fqmax) distributions in pup isolation induced USVs and in adult courtship USVs were performed using R package mixtools: tools for analyzing finite mixture modelsref. The bimodality of the USV duration distributions were assessed using Ashman’s D Score
[0364]
[0365] where μ and σ are the center and standard deviation of each Gaussian, respectively. Cutoff Fqmax in pup USVs between low and high USVs were defined as the intercept of the two Gaussian fits to the distribution to the nearest frequency (kHz). Cutoff durations between short and long USVs were defined as the intercept of the two Gaussian fits to the USV duration distribution to the nearest millisecond. Statistical analysis was performed by pairwise Wilcoxon rank sum tests, with correction (Bonferroni) for multiple comparisons between genotypes. Correction was not applied for the parameters in call structure. This is because individual properties are assumed to be related to each other, which increases type 2 error caused by overcorrection.
[0366]
[0231] Rotarod test
[0367]
[0232] The tests were performed using the elevated revolving rod (Stoelting, Cat#. 57624). Mice w ere placed on the apparatus and habituated for few minutes. The rod accelerated at a constant rate (4 to 40rpm in 300sec) and the time it took the animals to fall was recorded. Tests were performed three times and the average value are calculated. Statistical analysis was performed by Wilcoxon rank sum test.
[0368]
[0233] Y-maze test
[0369]
[0234] The Y-maze tests were conducted according to the described procedure ‘ \ The tests were performed in a Y-maze with three arms of equal length at 120° angles to each other (Stoelting, Cat#. 60180). Mice were placed in the center of the maze and had free access to all three arms. If the animal chooses an arm different from the arm it arrived in, this choice is called an alteration. This is considered a correct response; conversely, returning to the previous arm is considered an error. The number of times and the order in which the animals entered the arms are recorded and used to calculate the alternation rate. The behavior of the mice was recorded for 8 min. Statistical analysis was performed by Wilcoxon rank sum test.
[0235] Variant identification and annotations of Neanderthal and Denisovan
[0370]
[0236] The sequencing data of three Neanderthal genomes were obtained from The Draft Neanderthal Genome Project (ebi.ac.uk / ena / browser / view / PRJEB2065). The sequencing data of Denisovan genome accompanied with nine modern human genomes were obtained from Denisovan Genome Project (cdna.eva.mpg.de / denisova / ). The fastq files of each sample were aligned to the human genome (hg 19) by Burrows-Wheeler Aligner (BWA, v0.7.17-r1188). The aligned SAM files were processed into BAM files by Picard (v2.18.7) and Genome Analysis Toolkit (GATK, v4.4.0.0). Variant calling for each sample was processed with Mutect2 of GATK4. The variants were annotated by using ANNOVAR (v2).
[0371]
[0237] Variants in NOVA 1 loci
[0372]
[0238] The variants in modem human populations were obtained from ExAC database (vO.3.1) (gnomad.broadinstitute.org / downloads), which contains 60,706 exomes mapped to hgl9. Then, the variants in NOVA1 loci (chromosome 14: 26912296-27067239) of Neanderthal, Denisovan, and modern human populations were subsetted by using bcftools (v1.19). The frequency of minor alleles for each position in N0VA1 loci was calculated.
[0373]
[0239] Tajima’s D statistical analysis
[0374]
[0240] Corrected allelic frequencies of each SNP site from ExAC (v0.3.1) were extracted from the dbSNP database (version 2022-11-16). The sites with very low total allele count were filtered out (cutoff for total count: 200). pi, theta and Tajima’s D values were calculated for all genes examined
[0375]
[0376] . The normalized Tajima’s D value were calculated as the ratio of Tajima’s D to its theoretical minimum value (D(rnin))47.
[0377]
[0241] DH test
[0378]
[0242] Allelic frequency for each mutation site were extracted from refsnp files (version bl 56) downloaded from NCBI (version of November 2022). In-house program was used to calculate the evolutionary statistic, Tajima’s D is calculated by according (Tajima, 1989), Fay and Wu's Ef-statistic is defined as Oa: - 0 / / (Fay and Wu, 2000), we calculate normalized H (formula 11. 12), E-test (formula 13, 14), and DH-test (formula 15) (Zeng et al., 2006). Our in-house program is shared to public: github.com / cafcbluc / popgcn_dbsnp”
[0379]
[0243] Selection analysis on NOVA1 197V
[0380]
[0244] The ancestral recombination graph (ARG) analysis was performed using ARGweaver-D^. These ARGs explicitly describe gene trees and accompanying recombination events throughout the region. They were sampled from an approximate posterior distribution using Markov chain Monte Carlo methods to capture uncertainty in the ARG, given the sequence data and evolutionary' parameters. The sampled ARGs included two Yoruba (HGDP00927, SS6004475), two Mbuti (SS6004471, HGDP00456), and two San(HGDP01029, SS6004473) individuals, all sequenced to high coverage, as well as the Altai Neanderthal and Denisovan sequences and a chimpanzee outgroup (panTro4). Because the N0VA1 197V variant is nearly fixed in modem humans and likely predates the separation of major continental population groups, we do not anticipate significant additional insights from including more modern human samples, as most would coalesce well after the allele reached high frequency. Based on information from the ARG, the sampled ARGs were analyzed using CLUES250.
[0381]
[0245] Statistical information
[0382]
[0246] Information of statistical methods and the number of biological replicates in the analysis are in the figure legends and methods section of each analysis as appropriate.
[0383] References
[0384] 1. Scally, A. & Durbin, R. Revising the human mutation rate: implications for understanding human evolution. Nat. Rev. Genet. 13, 745-753 (2012).
[0385] 2. Mounier, A. & Mirazon Lahr, M. Deciphering African late middle Pleistocene hominin diversity and the origin of our species. Nat. Commun. 10, 1-13 (2019).
[0386] 3. Jarvis, E. D. Evolution of vocal learning and spoken language. Science 366, 50-54 (2019).
[0387] 4. Fitch, W. T., de Boer, B., Mathur, N. & Ghazanfar, A. A. Monkey vocal tracts are speech-ready. Sci Adv 2, e1600723 (2016).
[0388] 5. Boe, L.-J. et al. Evidence of a Vocalic Proto-System in the Baboon (Papio papio) Suggests Pre-Hominin Speech Precursors. PLoS One 12, e0169321 (2017).
[0389] 6. Nishimura, T. et al. Evolutionary loss of complexity in human vocal anatomy as an adaptation for speech. Science 377, 760-763 (2022).
[0390] 7. Jurgens, U. Neural pathways underlying vocal control. Neurosci. Biobehav. Rev. 26, 235-258 (2002).
[0391] 8. Meyer, M. et al. A High-Coverage Genome Sequence from an Archaic Denisovan
[0392] Individual. Science 338, 222-226 (2012)..
[0393] 9. Prtifer, K. et al. The complete genome sequence of a Neanderthal from the Altai Mountains. Nature 505, 43-49 (2013).
[0394] 10. Green, R. E. etal. A draft sequence of the Neandertal genome. Science 328, 710-722
[0395] (2010).
[0396] 11. Prtifer, K. et al. A high-coverage Neandertal genome from Vindija Cave in Croatia. Science 358, 655-658 (2017).
[0397] 12. Atkinson, E. G. et al. No Evidence for Recent Selection at FOXP2 among Diverse Human Populations. Cell 174, 1424-1435.e15 (2018).
[0398] 13. Trujillo, C. A. et al. Reintroduction of the archaic variant of NOVA1 in cortical organoids alters neurodevelopment. Science 371, (2021).
[0399] 14. Hcrai, R. H., Scmcndcfcri, K. & Muotri, A. R. Comment on ‘Human TKTL1 implies greater neurogenesis in frontal neocortex of modem humans than Neanderthals’. Science 379, eadft)602 (2023).
[0400] 15. Lai, C. S. L., Fisher, S. E., Hurst, J. A., Vargha-Khadem, F. & Monaco, A. P. A forkheaddomain gene is mutated in a severe speech and language disorder. Nature 413, 519-523 (2001).
[0401] 16. MacDermot, K. D. et al. Identification of FOXP2 truncation as a novel cause ofdevelopmental speech and language deficits. Am. J. Hum. Genet. 76, 1074—1080(2005).
[0402] 17. Chabout, J. et al. A Foxp2 mutation implicated in human speech deficits alters sequencing of ultrasonic vocalizations in adult male mice. Front. Behav. Neurosci. 10, 197 (2016).
[0403] 18. Castellucci, G. A., McGinley, M. J. & McCormick, D. A. Knockout of Foxp2 disrupts vocal development in mice. Sci. Rep. 6, 23305 (2016).
[0404] 19. Enard, W. et al. A Humanized Version of Foxp2 Affects Cortico-Basal Ganglia Circuits in Mice. Cell 137, 961-971 (2009).
[0405] 20. Hammerschmidt, K. et al. A humanized version of Foxp2 does not affect ultrasonic vocalization in adult mice. Genes Brain Behav. 14, 583-590 (2015).
[0406] 21. von Merten, S., Pfeifle, C., Kunzel, S., Hoier, S. & Tautz, D. A humanized version of
[0407] Foxp2 affects ultrasonic vocalization in adult female and male mice. Genes Brain Behav. 20,
[0408] e 12764 (2021).
[0409] 22. Krause, J. et al. The derived FOXP2 variant of modem humans was shared with Neandertals. Curr. Biol. 17, 1908-1912 (2007).
[0410] 23. Pinson, A. et al. Human TKTL1 implies greater neurogenesis in frontal neocortex of
[0411] modem humans than Neanderthals. Science 377, eabl6422 (2022).
[0412] 24. Buckanovich, R. J., Posner, J. B. & Darnell, R. B. Nova, the paraneoplastic Ri antigen, is homologous to an RNA-binding protein and is specifically expressed in the developing motor system. Neuron 11, 657—672 (1993).
[0413] 25. Eizirik, D. L. et al. The human pancreatic islet transcriptome: expression of candidate genes for type 1 diabetes and the impact of pro-inflammatory cytokines. PLoS Genet. 8, e1002552 (2012).
[0414] 26. Villate, O. et al. Noval is a master regulator of alternative splicing in pancreatic betacells. Nucleic Acids Res. 42, 11818-11830 (2014).
[0415] 27. Yang, Z. et al. NOVAI prevents overactivation of the unfolded protein response and facilitates chromatin access during human white adipogenesis. Nucleic Acids Res. 51, 6981- 6998 (2023).
[0416] 28. Darnell, R. B. & Posner, J. B. Paraneoplastic syndromes involving the nervous system. N.
[0417] Engl. J. Med. 349, 1543-1554 (2003).
[0418] 29. Darnell, R. B. & Posner, J. B. Paraneoplastic Syndromes. (Oxford University' Press, USA, 2011).
[0419] 30. Yang, Y. Y., Yin, G. L. & Darnell, R. B. The neuronal RNA-binding protein Nova-2 is implicated as the autoantigen targeted in POMA patients with dementia. Proc. Natl. Acad. Sci. U. S. A. 95, 13254-13259 (1998).
[0420] 31. Jensen, K. B. et al. Nova-1 regulates neuron-specific alternative splicing and is essential for neuronal viability. Neuron 25, 359-371 (2000).
[0421] 32. Ule, J. et al. CLIP identifies Nova-regulated RNA networks in the brain. Science 302,
[0422] 1212-1215 (2003).
[0423] 33. Licatalosi, D. D. et al. HITS-CLIP yields genome-wide insights into brain alternative RNA processing. Nature 456, 464-469 (2008).
[0424] 34. Dredge, B. K. & Darnell, R. B. Nova regulates GABA(A) receptor gamma2 alternative splicing via a distal downstream UCAU-rich intronic splicing enhancer. Mol. Cell. Biol. 23, 4687-4700 (2003).
[0425] 35. Darnell, R. B. RNA protein interaction in neurons. Ann. Rev. Neurosci. 36, 243—270
[0426] (2013).
[0427] 36. Tajima, Y. et al. NOVAI acts on Impact to regulate hypothalamic function and translation in inhibitory neurons. Cell Rep. 42, 112050(2023).
[0428] 37. Huang, C. S. et al. Common molecular pathways mediate long-term potentiation of synaptic excitation and slow synaptic inhibition. Cell 123, 105-118 (2005).
[0429] 38. Zhang, C. et al. Integrative modeling defines the Nova splicing-regulatory network and itscombinatorial controls. Science 329, 439-443 (2010).
[0430] 39. Riesenberg, S. et al. Efficient high-precision homology-directed repair-dependent genome editing by HDRobust. Nat. Methods 20, 1388-1399 (2023).
[0431] 40. Maricic, T. et al. Comment on ‘Reintroduction of the archaic variant of NOVA 1 in cortical organoids alters neurodevelopment’. Science 374, eabi6060(2021).
[0432] 41. Herai, R. H., Szeto, R. A., Trujillo, C. A. & Muotri, A. R. Response to Comment on ‘Reintroduction of the archaic variant of NOVA 1 in cortical organoids alters neurodevelopment’. Science 374, eabi9881 (2021).
[0433] 42. Buckanovich, R. J. & Darnell, R. B. The neuronal RNA binding protein Nova-1 recognizes specific RNA targets in vitro and in vivo. Mol. Cell. Biol. 17, 3194-3201 (1997).
[0434] 43. Racca, C. et al. The Neuronal Splicing Factor Nova Co-Localizes with Target RNAs inthe Dendrite. Front. Neural Circuits 4, 5 (2010).
[0435] 44. Saito, Y. et al. NOVA2-mediated RNA regulation is required for axonal pathfinding during development. Elife 5, (2016).
[0436] 45. Lewis, H. A. et al. Sequence-specific RNA binding by a Nova KH domain: implications for paraneoplastic disease and the fragile X syndrome. Cell 100, 323-332 (2000).
[0437] 46. Lewis, H. A. et al. Crystal structures of Nova- 1 and Nova-2 K-homology RNA-binding domains. Structure 7, 191-203 (1999).
[0438] 47. Schaeffer, S. W. Molecular population genetics of sequence length diversity in theAdh region of Drosophila pseudoobscura. Genet. Res. 80, 163-175 (2002).
[0439] 48. Zeng, K., Fu, Y.-X., Shi, S. & Wu, C.-I. Statistical tests for detecting positive selection by utilizing high-frequency variants. Genetics 174, 1431-1439(2006).
[0440] 49. Hubisz, M. J., Williams, A. L. & Siepel, A. Mapping gene flow between ancient hominins through demography-aware inference of the ancestral recombination graph. PLoS Genet. 16, e1008895 (2020).
[0441] 50. Vaughn, A. H. & Nielsen, R. Fast and accurate estimation of selection coefficientsand allele histories from ancient and modem DNA. Mol. Biol. Evol. 41, (2024).
[0442] 51. Hejase, H. A., Mo, Z., Campagna, L. & Siepel, A. A deep-learning approach for Inference of selective sweeps from the ancestral recombination graph. Mol. Biol. Evol. 39, (2022).
[0443] 52. Zhou, Y. et al. Metascape provides a biologist-oriented resource for the analysis of systems-level datasets. Nat. Commun. 10, 1-10 (2019).
[0444] 53. Menheniott, T. R. et al. A novel gastrokine, Gkn3, marks gastric atrophy and shows evidence of adaptive gene loss in humans. Gastroenterology 138, 1823—1835 (2010).
[0445] 54. De Boulle, K. et al. A point mutation in the FMR-1 gene associated with fragile Xmental retardation. Nat. Genet. 3, 31-35 (1993).
[0446] 55. Siomi, H., Choi, M., Siomi, M. C., Nussbaum, R. L. & Dreyfuss, G. Essential role for KH domains in RNA binding: impaired RNA binding by a mutation in the KH domain of FMRI that causes fragile X syndrome. Cell 11, 33-39 (1994).
[0447] 56. Jones, A. R. & Schedl, T. Mutations in gid- 1, a female germ cell-specific tumor suppressor gene in Caenorhabditis elegans, affect a conserved domain also found in Src -associated protein Sam68. Genes Dev. 9, 1491-1504 (1995).
[0448] 57. Buckanovich, R. J., Yang, Y. Y. & Darnell, R. B. The onconeural antigen Nova-1 is a neuron-specific RNA-binding protein, the activity of which is inhibited by paraneoplastic antibodies. J. Neurosci. 16, 1114-1122 (1996).
[0449] 58. Kashima, I. et al. Binding of a novel SMG-l-Upfl-eRFl-eRF3 complex (SURF) tothe exon junction complex triggers Upfl phosphorylation and nonsense-mediated mRNA decay.
[0450] Genes Dev. 20, 355-367 (2006).
[0451] 59. Fiorini, F., Bagchi, D., Le Hir, H. & Croquette, V. Human Upfl is a highly processiveRNA helicase and translocase with RNP remodelling activities. Nat. Commun. 6, 1-10(2015).60. Jackson, R. J., Hellen, C. U. T. & Pestova, T. V. The mechanism of eukaryotic translation initiation and principles of its regulation. Nat. Rev. Mol. Cell Biol. 11, 113-127 (2010).
[0452] 61. Ping, X.-L. et al. Mammalian WTAP is a regulatory subunit of the RNAN6-methyladenosine methyltransferase. Cell Res. 24, 177—189(2014).
[0453] 62. Liu, J. et al. A METTL3— METTL14 complex mediates mammalian nuclear RNAN6-adenosine methylation. Nat. Chem. Biol. 10, 93-95 (2013).
[0454] 63. Linder, P. & Jankowsky, E. From unwinding to clamping — the DEAD box RNAhelicase family. Nat. Rev. Mol. Cell Biol. 12, 505-516 (2011).
[0455] 64. Saito, Y. et al. Differential NOV A2 -Mediated Splicing in Excitatory and Inhibitory Neurons Regulates Cortical Development and Cerebellar Function. Neuron 101, 707-720.e5 (2019).
[0456] 65. Teplova, M. et al. Protein-RNA and protein-protein recognition by dual KH1 / 2 domains of the neuronal splicing factor Nova-1. Structure 19, 930-944 (2011).
[0457] 66. Dredge, B. K., Stefani, G., Engelhard, C. C. & Darnell, R. B. Nova autoregulation reveals dual functions in neuronal splicing. EMBO J. 24, 1608-1620 (2005).
[0458] 67. Subramanian, H. H., Balnave, R. J. & Holstege, G. Microstimulation in Different Parts of the Periaqueductal Gray Generates Different Types of Vocalizations in the Cat. J. Voice 35, 804.e9-804.e25 (2021).
[0459] 68. Subramanian, H. H. & Holstege, G. Stimulation of the midbrain periaqueductal gray modulates preinspiratory neurons in the ventrolateral medulla in the rat in vivo. J. Comp. Neurol.
[0460] 521, 3083-3098 (2013).
[0461] 69. Subramanian, H. H., Balnave, R. J. & Holstege, G. The midbrain periaqueductal gray control of respiration. J. Neurosci. 28, 12274-12283 (2008).
[0462] 70. Zhang, S. P., Davis, P. J., Bandler, R. & Carrive, P. Brain stem integration of vocalization: role of the midbrain periaqueductal gray. J. Neurophysiol. 72, 1337-1356 (1994).
[0463] 71. Holstege, G. Anatomical study of the final common pathway for vocalization in the cat. J. Comp. Neurol. 284, 242-252 (1989).
[0464] 72. Bandler, R. & Carrive, P. Integrated defence reaction elicited by excitatory amino acid microinjection in the midbrain periaqueductal grey region of the unrestrained cat. Brain Res. 439, 95-106 (1988).
[0465] 73. Zhang, S. P., Bandler, R. & Carrive, P. Flight and immobility evoked by excitatory amino acid microinjection within distinct parts of the subtentorial midbrain periaqueductal gray of the cat. Brain Res. 520, 73-82 (1990).
[0466] 74. Holstege, G. The emotional motor system and micturition control. Neurourol. Urodyn. 29, 42-48 (2010).
[0467] 75. Stone, E., Coote, J. H., Allard, J. & Lovick, T. A. GABAergic control of micturition within the periaqueductal grey matter of the male rat. J. Physiol. 589, 2065-2078 (2011).
[0468] 76. Jurgens, U. The role of the periaqueductal grey in vocal behaviour. Behav. Brain Res. 62, 107-117 (1994).
[0469] 77. Faull, O. K., Subramanian, H. H., Ezra, M. & Pattinson, K. T. S. The midbrain periaqueductal gray as an integrative and interoceptive neural structure for breathing. Neurosci.
[0470] Biobehav. Rev. 98, 135-144 (2019).
[0471] 78. Zhang, H. et al. The contribution of periaqueductal gray in the regulation of physiological and pathological behaviors. Front. Neurosci. 18, 1380171 (2024).
[0472] 79. Davies, G. et al. Genome-wide association studies establish that human intelligence is highly heritable and polygenic. Mol. Psychiatry 16, 996-1005 (2011).
[0473] 80. Benyamin, B. et al. Childhood intelligence is heritable, highly polygenic and associated with FNBP1L. Mol. Psychiatry 19, 253-258 (2013).81. Sia, G. M., Clem, R. L. & Huganir. R. L. The human language-associated gene SRPX2 regulates synapse formation and vocalization in mice. Science 342, 987-991 (2013).
[0474] 82. Sotcros, B. M., Cong, Q., Palmer, C. R. & Sia, G.-M. Sociability and synapse subtypespecific defects in mice lacking SRPX2, a language-associated gene. PLoS One S, e0199399 (2018).
[0475] 83. Hori, K. el al. AUTS2 Regulation of Synapses for Proper Synaptic Inputs and Social Communication. iScience 23, 101183 (2020).
[0476] 84. Gauthier, J. et al. Truncating mutations in NRXN2 and NRXNI in autism spectrum disorders and schizophrenia. Hum. Genet. 130, 563-573 (2011).
[0477] 85. Gill, J. et al. Regulated Intron Removal Integrates Motivational State and Experience. Cell 169, 836-848.el5 (2017).
[0478] 86. Arriaga, G. & Jarvis. E. D. Mouse vocal communication system: Are ultrasounds leamedor innate? Brain Lang. 124, 96-116 (2013).
[0479] 87. Jurgens, U. The Neural Control of Vocalization in Mammals: A Review. J. Voice 23, 1— 10 (2009).
[0480] 88. Kittelberger, J. M., Land, B. R. & Bass, A. H. Midbrain periaqueductal gray andvocal patterning in a teleost fish. J. Neurophysiol. 96, 71-85 (2006).
[0481] 89. Grijseels, D. M., Prendergast, B. J., Gorman, J. C. & Miller, C. T. The neurobiology of vocal communication in marmosets. Ann. N. Y. Acad. Sci. 1528, 13-28 (2023).
[0482] 90. Ennis, M., Xu, S. J. & Rizvi, T. A. Discrete subregions of the rat midbrain periaqueductal gray project to nucleus ambiguus and the periambigual region. Neuroscience 80, 829-845 (1997).
[0483] 91. Floody, O. R. & DeBold, J. F. Effects of midbrain lesions on lordosis and ultrasound production. Physiol. Behav. 82, 791-804 (2004).
[0484] 92. Jurgens, U. & Pratt, R. Role of the periaqueductal grey in vocal expression of emotion. Brain Res.
[0485] 167, 367-378 (1979).
[0486] 93. Kirzinger, A. & Jurgens, U. The effects of brainstem lesions on vocalization in the squirrel monkey. Brain Res. 358, 150-162 (1985).
[0487] 94. Tschida, K. et al. A Specialized Neural Circuit Gates Social Vocalizations in the Mouse. Neuron 103, 459-472.e4 (2019).
[0488] 95. Branchi, I., Santucci, D. & Alieva, E. Ultrasonic vocalisation emitted by infant rodents: A tool for assessment of neurobehavioural development. Behav. Brain Res. 125, 49-56 (2001).
[0489] 96. Branchi, I., Santucci, D., Puopolo, M. & Alieva, E. Neonatal behaviors associated with ultrasonic vocalizations in mice (mus musculus): A slow-motion analysis. Dev. Psychobiol. 44, 37-44 (2004).
[0490] 97. Zimmer, M. R., Fonseca, A. H. O., lyilikei, O., Pra, R. D. & Dietrich, M. O. Functional Ontogeny of Hypothalamic Agrp Neurons in Neonatal Mouse Behaviors. Cell 178, 44- 59.e7 (2019).
[0491] 98. Chabout, J., Sarkar, A., Dunson, D. B. & Jarvis, E. D. Male mice song syntax depends on social contexts and influences female preferences. Front. Behav. Neurosci. 9, 76(2015).
[0492] 99. Grimsley, J. M. S., Monaghan, J. J. M. & Wenstrup, J. J. Development of social vocalizations in mice. PLoS One 6, el 7460 (2011).
[0493] 100. Ashman, K. M., Bird, C. M. & Zepf, S. E. Detecting Bimodality in Astronomical Datasets. arXiv [astro-ph] (1994).
[0494] 101. Benaglia, T., Chauveau, D., Hunter, D. R. & Young, D. S. mixtools: An RPackagefor Analyzing Mixture Models. J. Stat. Softw. 32, 1-29 (2010).
[0495] 102. D’Amato, F. R., Scalera, E., Sarli, C. & Moles, A. Pups Call, Mothers Rush: Does Maternal Responsiveness Affect the Amount of Ultrasonic Vocalizations in Mouse Pups? Behav. Genet. 35, 103-112 (2005).103. Cohen-Salmon, C. Differences in patterns of pup care in Mus musculus domesticus. VIII. Effects of previous experience and parity in XLII inbred mice. Physiol. Behav. 40, 177-180 (1987).
[0496] 104. Fischer, J. & Hammerschmidt, K. Ultrasonic vocalizations in mouse models for speech and socio-cognitive disorders: insights into the evolution of vocal communication. Genes Brain Behav. 10, 17-27 (2011).
[0497] 105. Holy, T. E. & Guo, Z. Ultrasonic songs of male mice. PLoS Biol. 3, e386(2005).
[0498] 106. Chabout, J., Jones-Macopson, J. & Jarvis, E. D. Eliciting and Analyzing Male Mouse Ultrasonic Vocalization (USV) Songs. J. Vis. Exp. (2017) doi: 10.3791 / 54137.
[0499] 107. Waidmann, E. N., Yang, V. H. Y., Doyle, W. C. & Jarvis, E. D. Mountable miniature microphones to identify and assign mouse ultrasonic vocalizations. bioRxiv
[0500] 2024.02.05.579003 (2024) doi: 10.1101 / 2024.02.05.579003.
[0501] 108. Castellucci, G. A., Calbick, D. & McCormick, D. The temporal organization ofmouse ultrasonic vocalizations. PLoS One 13, eO 199929 (2018).
[0502] 109. Vogel, A. P., Tsanas, A. & Scattoni, M. L. Quantifying ultrasonic mouse vocalizations using acoustic analysis in a supervised statistical machine learning framework. Sei. Rep.9, 1-10 (2019). 110. Arriaga, G., Zhou, E. P. & Jarvis, E. D. Of mice, birds, and men: the mouse ultrasonic song system has some features similar to humans and song-learning birds. PLoS One 7, e46610 (2012). 111. Mora-Bermudez, F. et al. Longer metaphase and fewer chromosome segregation errors in modern human than Neanderthal brain development. Science Advances 8, eabn7702(2022).
[0503] 112. Stepanova, V. et al. Reduced purine biosynthesis in humans after their divergence from Neandertals. Elife 10, e58741 (2021).
[0504] 113. Bejerano, G. et al. Ultraconserved elements in the human genome. Science 304, 1321-1325 (2004).
[0505] 114. Liu, A. et al. GC-biased gene conversion drives accelerated evolution of ultraconserved elements in mammalian and avian genomes. Genome Res. 33, 1673-1689(2023).
[0506] 115. Gronau, L, Hubisz, M. J., Gulko. B., Danko, C. G. & Siepel. A. Bayesian inference of ancient human demography from individual genome sequences. Nat. Genet. 43, 1031—1034 (2011).
[0507] 116. Wei, X. P., Collie, M., Dempsey, B., Fortin, G. & Yackle, K. A novel reticular node in the brainstem synchronizes neonatal mouse crying with breathing. Neuron 110, 644— 657. e6 (2022). 117. Hartmann, K. & Brecht, M. A Functionally and Anatomically Bipartite Vocal Pattern Generator in the Rat Brain Stem. iScience 23, 101804 (2020).
[0508] 118. Concha-Miranda, M., Tang, W., Hartmann, K. & Brecht, M. Large-Scale Mappingof Vocalization-Related Activity in the Functionally Diverse Nuclei in Rat Posterior Brainstem. J.
[0509] Neurosci. 42, 8252-8261 (2022).
[0510] 119. Fitch, W. T., Huber, L. & Bugnyar, T. Social cognition and the evolution of language: constructing cognitive phylogenies. Neuron 65, 795-814 (2010).
[0511] 120. Oren, G. et al. Vocal labeling of others by nonhuman primates. Science 385,996-1003 (2024).
[0512] 121. Lefebvre, E., Granon, S. & Chauveau, F. Social context increases ultrasonic vocalizations during restraint in adult mice. Anim. Cogn. 23, 351—359 (2020).
[0513] 122. Portfors, C. V., Mayko, Z. M., Jonson, K., Cha, G. F. & Roberts, P. D. Spatial organization of receptive fields in the auditory midbrain of awake mouse. Neuroscience 193, 429M39 (2011).
[0514] 123. Whitney, G., Coble, J. R., Stockton, M. D. & Tilson, E, F. Ultrasonic emissions: Dothey facilitate courtship of mice? J. Comp. Physiol. Psychol. 84, 445-452 (1973).
[0515] 124. Neunuebel, J. P., Taylor, A. L., Arthur, B. J. & Egnor, S. E. R. Female mice ultrasonically interact with males during courtship displays. Elife 4, e06203 (2015).
[0516] 125. Moles, A., Costantini, F., Garbugino, L., Zanettini, C. & D’Amato, F. R. Ultrasonic vocalizations emitted during dyadic interactions in female mice: a possible index of sociability? Behav. Brain Res. 182, 223-230 (2007).126. Wu, J., Anczukow, 0., Kraincr. A. R., Zhang, M. Q. & Zhang, C. OLego: fast and sensitive mapping of spliced mRNA-Seq reads using small seeds. Nucleic Acids Res. 41, 5149-5163 (2013). 127. Yan, Q. et al. Systematic discovery’ of regulated and conserved alternative exons in the mammalian brain reveals NMD modulating chromatin regulators. Proc. Natl. Acad. Sci. U. S. A. 112.
[0517] 3445-3450 (2015).
[0518] 128. Robinson, M. D., McCarthy, D. J. & Smyth, G. K. edgeR: a Bioconductor package for differential expression analysis of digital gene expression data. Bioinformatics 26, 139-140 (2010). 129. Shah, A., Qian, Y., Weyn-Vanhentenryck, S. M. & Zhang, C. CLIP Tool Kit (CTK): a flexible and robust pipeline to analyze CLIP sequencing data. Bioinformatics 33, 566-567 (2017). 130. Stoumpou, V. et al. Analysis of Mouse Vocal Communication (AMVOC): a deep, unsupcrviscd method for rapid detection, analysis and classification ofultrasonic
[0519] vocalisations. Bioacoustics 32, 199-229 (2023).
[0520] 131. Kraeuter, A.-K., Guest, P. C. & Samyai, Z. The Y-Maze for Assessment of Spatial Working and Reference Memory in Mice. Methods Mol. Biol. 1916, 105-111 (2019).
[0521] 132. Tajima, F. Statistical method for testing the neutral mutation hypothesis byDNA polymorphism. Genetics 123, 585-595 (1989).
[0522] TABLES
[0523]
[0247] TABLE I
[0524] Tajima’s D and normalized Tajima’s D.
[0525] Tajima’s D values and their normalized values were calculated based on the ExAC data sets. The genes on datasets 3 and 4 are based on the reports of Meyer et al-, and the genes on dataset 5 are based on Trujillo et al. See Methods for details on calculation of the Tajima’s D values. A subset of data is presented here for readability.
[0526] gene. chr pi. thetaW. i Tajima.s. D Tajima.s. D. Gene set normalized.
[0527] NOVA1 chr14 i 1.68E-06 0.00191534 i -2.4807974 -0.999292729 i NOVA1 and its related FOXG1 chr14 i 1.00E-05 0.00267173 | -2.4017538 -0.9963984 i NOVA1 and its related NOVA2 chr19 i 5.26E-05 0.00129301 i -2.3220406 -0.959435228 i NOVA1 and its related STXBP6 chr14 9.66E-06 0.00127155 i -2.3828803 -0.992563164 NOVA1 and its related ADSL chr22 i 3.81 E-05 0.00103031 i -2.2524394 -0.963218649 Meyer2012_neuron.related ARHGAP32 chr11 i 9.87E-05 0.00136933 i -2.2933292 -0.928104343; Meyer2012_neuron.related CNTNAP2 chr7 i 9.02E-05 0.00272893 i -2.4524935 -0.967131933 Meyer2012_neuron.related HTR2B chr2 i 3.68E-05 0.00097537 i -1.8752014 -0.96248593 i Meyer2012_neuron.related KATNA1 chr6 0.00022786 0.00342952 -2.1938463 -0.93374683 i Meyer2012 neuron.related LUZP1 chr1 0.00020082 0.00015172 0.57781166 0.323680306 Meyer2012 neuron.related NOVA1 chr14 1.68E-06 0.00191534 i -2.4807974 -0.999292729 I Meyer2012_neuron.related SLITRK1 chr13 i 2.22E-07 1.57E-05 i -0.5603988 -0.986039139 Meyer2012 neuron.related
[0528]
[0529]
[0248] TABLE 2
[0530] Selection analysis for human-specific SNPs using CLUES2.
[0531] The columns display the log-likelihood ratio (logLR), the negative loglO-transformed p-value (-loglO(p-vahie)), and the selection coefficient (SelectionMLEl) for each SNP across a single epoch(0 to 200,000). The table is sorted by -value, with SNPs highlighted in bold, including N0VA1 (bold and underlined), showing stronger selection signals. Gene list and SNP information are from Trujillo et al., 20213.
[0532] Gene name Ensembl chr position logLR -Iog10 Epochl Epochl SelectionMLEI Gene ID (hg19) (hg19) (p-value) _start _end HERC5 ENSG00000138646 4 89410317 5.5077 3.04 0 200000 0.00135 AHR ENSG00000106546 7 17375392 6.2143 3.37 0 200000 0.00127 LAG3 ENSG00000089692 12 6883790 3.9207 2.29 0 200000 0.00097 C3 ENSG00000125730 19 6685111 3.8964 2.28 0 200000 0.00089 ZNF2 ENSG00000163067 2 95831534 3.8143 2.24 0 200000 0.00087 NOVA1 ENSG00000139910 14 26918100 2.7454 1.72 0 200000 0.00082 SSH2 ENSG00000141298 17 27959034 2.739 1.72 0 200000 0.00057 SSH2_2 ENSG00000141298 17 27959258 2.7372 1.71 0 200000 0.00057 IF144L ENSG00000137959 1 79106805 2.5711 1.63 0 200000 0.00054 KIAA1199 ENSG00000103888 15 81173308 2.0899 1.39 0 200000 0.00052 FAM166A ENSG00000188163 9 140139881 2.4931 1.59 0 200000 000051 KIF18A ENSG00000121621 11 28119295 2.3184 1.5 0 200000 0.00051 ZNF106 ENSG00000103994 15 42742312 2.2816 1.49 0 200000 000051 CDH16 ENSG00000166589 16 66947064 2.1992 1.44 0 200000 0.00051 KIF26B ENSG00000162849 1 245582905 1.6157 1.14 0 200000 0.00049 ITGB4 ENSG00000132470 17 73753035 1.7434 1.21 0 200000 0.00048 PRDM10 ENSG00000170325 11 129772293 1.8097 1.24 0 200000 0.00038 NLRP2 ENSG00000022556 19 55489189 1.8316 1.25 0 200000 0.00037 ADSL ENSG00000239900 22 40760978 1.6985 1.18 0 200000 0.00036 ADAM18 ENSG00000168619 8 39537618 1.6707 1.17 0 200000 0.00035 VCAN ENSG00000038427 5 82837946 1.6271 1.15 0 200000 0.00034 SCAP ENSG00000114650 3 47469149 1.5695 1.12 0 200000 0.00033 OR1K1 ENSG00000165204 9 125563200 1.5636 1.11 0 200000 0.00033 DNAJC11 ENSG00000007923 1 6694660 1.52 1.09 0 200000 000032 ADAM18_2 ENSG00000168619 8 39564352 1.2405 0.94 0 200000 0.00031 DCHS1 ENSG00000166341 11 6654769 1.3738 1.01 0 200000 0.0003 PIGZ ENSG00000119227 3 196674495 1.0181 0.81 0 200000 0.00024 ZNHIT2 ENSG00000174276 11 64884957 1.0153 0.81 0 200000 0.00021 MRPL49 ENSG00000149792 11 64893151 1.0113 0.81 0 200000 0.00021 PIEZO1 ENSG00000103335 16 88804443 0.9977 0.8 0 200000 0.00021 AC074212.3 ENSG00000237452 19 46265288 0.9664 0.78 0 200000 0.00021 CASC5_2 ENSG00000137812 15 40915640 0.9432 0.77 0 200000 0.00021 CASC5 ENSG00000137812 15 40912860 0.943 0.77 0 200000 0.00021 NOTO ENSG00000214513 2 73438011 0.9422 0.77 0 200000 0.00021 NCOA6 ENSG00000198646 20 33337529 0.9277 0.76 0 200000 0.0002
[0533]
[0534] ZNF726 ENSG00000213967 19 24116551 0.732 0.65 0 200000 0.00018 CCT5 ENSG00000150753 5 10250094 0.7118 0.63 0 200000 0.00013 TEX2 ENSG00000136478 17 62290457 0.6899 0.62 0 200000 0.00012 FRMD8 ENSG00000126391 11 65154602 0.5303 0.52 0 200000 0.00011
[0535]
[0536]
[0249] TABLE 3
[0537] Differential CLIP peaks between Novalwt / wt and Novalhu / hu.
[0538] Peaks are extracted that fulfill the criteria of being detected in all three biological replicates and having tags per peak at least 10 (peak heigh t>l 0) for either Novalhu / hu or Novalwt / wt mice for NOVA 1 -CLIP analysis in P21 midbrain. Peaks on lowly expressed transcripts in P21 midbrain are excluded using RNA sequencing data. NOVA! binding peaks with / ?-value less than 0.01 and absolute value of log2FC more than 1 were shown. The table contains the number of read counts for each genotype along with the genomic location of each peak: gene name, chromosome, start, end, strand, and annotation on the transcript and peak ID. The analysis was performed using edgeR. Novalwt / wt N=3, Novalhu / hu N=3.Gene_name chr start end strand annotation Noval Noval PValue peak.lD
[0539] (wt / wt) (hu / hu)
[0540] Tln2 chr9 67375090 67375105 - intron 1 52 1.51E-08 218276[gene=chr9_r_c45452][PH=53][PH0=15.39][P=3.92e-12] Wfs1 chr5 36968335 36968337 - CDS] 13 45 6.81E-05 81327[gene=chr5_r_c18062][PH=58][PH0=16.54][P=2.50e-13] downstream
[0541] 10K
[0542] Cog5 chr12 31782366 31782368 + intron 68 135 7.19E-05 63195[gene=chr12_f_c15280][PH=203][PH0=57.50i[P=9 17e-14] TbHx chrX 77660566 77660573 + 3'UTR 24 4 8.81E-05 70789[gene=chrX_f_c12881][PH=28][PH0=12.95][P=2.92e-03] Usp25 chr16 77100491 77100495 + intron 3 20 0.00012122 347232[gene=chr16 J_c56828][PH=23][PH0=11,04][P=1,01e-02] Rhobtb2 chr14 69796341 69796343 - CDS 42 14 0.00012782 206997[gene=chr14_r_c41627][PH=56][PH0=11.81 ][P=6.44e-15] Sgcz chr8 38483559 38483564 - intron 3 20 0.00022142 168310[gene=chr8_r_c28981][PH=23][PH0=6.24][P=1.37e-05] Celf2 chr2 6543160 6543166 - 3'UTR 17 2 0.00027482 9779[gene=chr2_r_c2189][PH=19][PH0=9.79][P=1.87e-01] Gripl chr1O 119494678 119494680 + intron 6 27 0.00032527 266625[gene=chr10_f_c59146][PH=33][PH0=10.02][P=987e-07] Nacc2 chr2 26056875 26056876 - 3'UTR 26 5 0.00032527 59469[gene=chr2_r_c13933][PH=31][PH0=1003][P=6.53e-06] Opcml chr9 28827554 28827558 + intron 19 47 0.00037902 82659[gene=chr9_f_c16295][PH=66][PH0=18.33][P=1.63e-14] 8030462N17Rik chr18 77673267 77673277 - intron 0 12 0.00048846 164709[gene=chr18_r_c36578][PH=12][PH0=7.05][P=2.36e-01] Zbtb20 chr16 43176802 43176828 + intron 0 12 0.00048846 302113[gene=chr16_f_c47580][PH=12][PH0=5.81][P=1.13e-01] Cadps chr14 12435532 12435538 - intron 1 16 0.00051907 7826[gene=chr14_r_c2170][PH=17][PH0=4.20][P=1.18e-04] Fgf14 chr14 124326934 124326936 - intron 12 36 0.00053645 336266[gene=chr14_r_c64873][PH=48][PH0=9.19][P=6.55e-15] Fam 155a chr8 9238881 9238887 - intron 37 14 0.00053645 5870[gene=chr8_r_c1204][PH=51][PH0=18.12][P=1.09e-08] Nudcdl chr15 44411046 44411054 - intron 1 15 0.00061696 43887[gene=chr15 _r_c 11355][PH= 16][PH0=8.73] [P=7.86e-02] Dmxl2 chr9 54418345 54418348 - intron 5 29 0.00065465 171082[gene=chr9_r_c34491][PH=34][PH0=14.60][P=1.14e-03] Rgs8 chr1 153693792 153693796 + 3'UTR 38 16 0.00071843 283660[gene=chr1_f_c62451][PH=54][PH0=13.66][P=5.87e-14] Rgs9 chr11 109234998 109234999 - downstream 2 19 0.00072905 394173[gene=chr11_r_c76058][PH=21][PH0=6.77][P=3.64e-04]
[0543] 10K|intron
[0544] Plxna2 chr1 194815094 194815101 + 3'UTR 32 11 0.00087261 388504[gene=chr1_f_c84665][PH=43][PH0=14.55][P=6.66e-08] 2210408l21Rik chr13 77207945 77207947 + intron 12 0 0.00097687 137388[gene=chr13_f_c30891][PH=12][PH0=4.83][P=4.22e-02] A330076H08Rik chr7 61941615 61941618 - non-coding 15 2 0.000977 155846[gene=chr7_r_c28589][PH=17][PH0=5.75][P=2.48e-03] Snx24 chr18 53311185 53311186 + intron 8 43 0.00099762 135693[gene=chr18_f_c28621][PH=51][PH0=14.58][P=936e-12]
[0545]
[0546] Ctnna2 chr6 77349600 77349611 - intron 2 14 0.00105834 128839[gene=chr6_r_c28203][PH=16][PH0=10.71][P=1.99e-01] Wwox chr8 115234745 115234746 + intron 11 0 0.00106582 263063[gene=chr8_f_c59865][PH=11][PH0=4.26][P=6.04e-02] Tmem64 chr4 15285827 15285831 + 3'UTR 9 30 0.00119457 17194[gene=chr4J_c4304][PH=39][PH0=17.82][P=5.20e-04] Gpr165 chrX 96718486 96718491 + 3'UTR 17 44 0.001215 90685[gene=chrXJ_c18065][PH=61][PH0=22.69][P=2.16e-09] Bmprlb chr3 141846816 141846821 - downstream 2 16 0.0012544 235944[gene=chr3_r_c49354][PH=18][PH0=9.26][P=5.49e-02]
[0547] 10K|intron
[0548] Tmem68 chr4 3552881 3552890 - downstream 35 15 0.00126174 224[gene=chr4_r_c18][PH=50][PH0=13.71][P=2.82e-12]
[0549] 10K|intron
[0550] Cstf2t chr1S 31085227 31085229 + 3'UTR 34 14 000126174 81632[gene=chr19_f_c16468][P H=48] [P H0= 1934][P= 1 51 e-06] Fam155a chr8 9736285 9736292 - intron 16 3 0.00131294 40897[gene=chr8_r_c4373][PH=19][PH0=6.23][P=3.22e-03] Mbp chr18 82571209 82571218 + intron 5 21 0.00141581 203306[gene=chr18_f_c44190][PH=26][PH0=12.91][P=1.10e-01] Ptprt chr2 162427558 162427559 - intron 9 38 0.00141669 403860[gene=chr2_r_c88859][PH=47][PHO=14.31][P=1.01e-09] Dab1 chr4 104734817 104734822 + downstream 3 19 0.00149054 187819[gene=chr4_f_c43783][PH=22][PH0=8.83][P=652e-03]
[0551] 10K|intron
[0552] Fgf9 chr14 58109986 58109988 + 3'UTR 41 20 0.0014987 114977[gene=chr14_f_c28269][PH=61][PH0=16.01][P=2.34e-14] Tcf3 chr1O 80412537 80412539 - intron 15 40 0.00151809 183602[gene=chr10_r_c42408][PH=54][PH0=14.15][P=1.44e-13] Bacel chr9 45863180 45863196 + 3'UTR 4 20 0.00154495 170783[gene=chr9_f_c27569][PH=23][PH0=10.19][P=4.45e-03] Gas5 chr1 161037408 161037414 + non-coding 4 21 0.00154495 319034[gene=chr1_f_c68139][PH=25][PH0=15.49][P= 7.28e-02] Foxn3 chr12 99301581 99301584 - intron 7 36 0.00159816 156172[gene=chr12_r_c33299][PH=43][PH0=12.33][P=1,49e-09] Lingol chr9 56619908 56619909 - CDS 13 2 0.00183175 179080[gene=chr9_r_c36383][PH=15][PH0=5.89][P=2.49e-02] Ipo11 chr13 106878281 106878293 - intron 2 14 000183175 216140[gene=chr13_r_c48769][PH= 16][PH0=647] [P=306e-02] Repsi chr1O 18119017 18119032 + 3'UTR|CDS|intron 13 2 0.00183175 26215[gene=chr10_f_c5525][PH= 15][PH0=12.27][P=3.81e-01] Eif4e chr3 138555579 138555580 + 3'UTR 13 2 0.00183175 269987[gene=chr3_f_c59205][PH=15][PH0=9.33][P=4.15e-01] Zbtb20 chr16 43421583 43421585 + intron 2 15 0.00183175 304091 [gene=chr16 J_c48276][PH=17] [PH0=7.59][P= 1.83e-02] Rasgrfl chr9 89931789 89931799 + intron 13 2 0.00183175 312891[gene=chr9_f_c54416][PH=15][PH0=6.14][P=2.41e-02] Phactr3 chr2 178219367 178219377 + intron 9 29 0.00188011 501468[gene=chr2_f_c104787][PH=38][PH0=11,48][P=363e-08] Fam120a chr13 48922185 48922188 - CDS 28 8 0.00188011 92863[gene=chr13_r_c21897][PH=36][PH0=22.72][P=5.48e-02] Femlc chr18 46504648 46504654 - 3'UTR 20 4 00019522 105803[gene=chr18_r_c22842][PH=24][PH0= 11.98][P=1 67e-01] Nrxn3 chr12 89641569 89641574 + intron 0 11 0.00195362 230588[gene=chr12_Lc53385][PH=11][PH0=5.81][P=2.07e-01] Pls3 chrX 75786155 75786156 - 3'UTR 8 27 0.00210428 74940[gene=chrX_r_c15617][PH=35j[PH0=1570][P=3.44e-03]
[0553]
[0554] Frmd6 chr12 70833161 70833167 + intron 35 12 0.00215546 137754[gene=chr12 J_c33175j[PH=47j[PH0=12.99][P=1 92e-11] Scrtl chr15 76517795 76517799 - 3'UTR 11 31 0.00222392 145440[gene=chr15_r_c33056][PH =42][PH0= 10.15][P=6.17e-12] Shisa6 chr11 66479786 66479811 - intron 15 3 0.00235091 188045[gene=chr11_r_c40003][PH=18][PH0=10.15][P=8.56e-02] Sntgl chr1 9193556 9193578 - intron 3 15 0.00235091 22497[gene=chr1_r_c4951 ][PH=18][PH0=8.09][P=2.31 e-02] Nrxn3 chr12 89835988 89836001 + intron 3 16 0.00235091 241655[gene=chr12_f_c54827][PH=19][PH0=9.62][P=3.87e-02] Smap2 chr4 120968800 120968801 - 3'UTR 15 3 000235091 254548[gene=chr4_r_c53732][PH=18][PH0=10 17][P=281 e-01 ] Kirrel chr3 87081635 87081644 - 3'UTR 22 7 0.00249566 131406[gene=chr3_r_c28517][PH=29][PH0=7.26][P=9.15e-08] Nudt11 chrX 6053068 6053074 + 3'UTR 22 7 0.00249566 47[gene=chrX_f_d 1 ][PH=29] [PH0= 12.13][P= 1,30e-03] Arhgap44 chr11 65023153 65023155 - CDS 29 9 0.00256579 176033[gene=chr11 _r_c37417][PH = 37][PH0=22.80] [P=3.11 e-02] Sema4f chr6 82924400 82924405 - intron 18 4 0.00257819 147696[gene=chr6_r_c31467][PH=22][PH0= 11,52][P= 1.12e-01] Stxbp6 chr12 44854700 44854711 - 3'UTR 20 6 0.00260107 48761 [gene=chr12_r_c11800][PH =26][PH0= 13.67] [P=1.37e-02] Mapk8 chr14 33391156 33391165 - intron 5 20 0.00260107 98954[gene=chr14_r_c21513][PH=25][PH0=12.82][P=1.42e-02] Kirrel3 chr9 34697558 34697562 + intron 18 3 000262142 119765[gene=chr9_f_c19812][PH=21][PH0=9.25][P=1.93e-02] Numb chr12 83815750 83815772 - intron 6 22 0.00265304 136438[gene=chr 12_r_c29075][PH=28][PH0= 17.42] [P=5.06e-02] Auts2 chr5 131865939 131865958 - intron 2 13 0.00274509 356276[gene=chr5_r_c74056][PH=15][PH0=511][P=6.80e-03j Farpl chr14 121094837 121094842 + intron 2 20 0.00288105 234673[gene=chr14_f_c54729][PH=22][PH0=10.27][P=1.57e-02] Auts2 chr5 131615099 131615122 - intron 3 21 0.00297168 345719[gene=chr5_r_c72487][PH=24][PH0=14.14][P=4.85e-02] Sertad4 chrd 192845131 192845133 - 3'UTR 18 3 0.00311834 455652[gene=chr1_r_c99472][PH=21j[PH0=8.61][P=3.53e-03] Phf20 chr2 156227793 156227808 + intron 7 23 0.00313858 435302[gene=chr2_f_c91425][PH=30][PH0=8.65][P=6.36e-07] Scamp5 chr9 57442182 57442183 - 3'UTR 45 21 0003202 184401[gene=chr9_r_c37743][PH=66j[PH0=2356][P=581e-10] Pcsk2 chr2 143616863 143616871 + intron 14 2 0.00328831 397207[gene=chr2_f_c82251][PH=16][PH0=8.83][P=8.39e-02] Plec chr15 76172043 76172062 - 3'UTR 10 31 0.00338126 142998[gene=chr15_r_c32635][PH=41][PH0=20.82][P=6.94e-04] Aqp4 chr18 15392136 15392137 - 3'UTR 12 30 0.00338126 24755[gene=chr18_r_c6655] [PH=42] [PH0= 19.36] [P=3.14e-04] Kcnip4 chr5 48800208 48800209 - intron 12 2 0.00341905 105704[gene=chr5_r_c24161][PH=14][PH0=3.77][P=1.17e-03] Kcnip4 chr5 49515481 49515486 - intron 12 2 0.00341905 146850[gene=chr5_r_c29186][PH=14][PH0=7.45][P=9.25e-02] Erbb4 chr1 68849423 68849428 - intron 1 13 0.00341905 152763[gene=chr1_r_c33551][PH=14][PH0=688][P=9.44e-02j 4933431E20Rik chr3 107891945 107891963 - NA 12 2 000341905 172994[gene=chr3_r_c36764][PH=14][PH0=439][P=414e-03] Hdlbp chr1 93413864 93413874 - CDS 12 2 0.00341905 203697[gene=chr1_r_c45482][PH=14][PH0=8.59][P=1.68e-01]
[0555]
[0556] Plekha6 chr1 133297020 133297041 + intron 12 2 0.00341905 252614[gene=chr1_f_c56079][PH=14][PH0=5.07][P=5.71e-02] Hecwl chr13 14427719 14427739 - intron 2 13 0.00341905 26548[gene=chr13 _r_c5850] [PH= 15] [PH0=9.06] [P=1,58e-01 ] Paqr7 chr4 134508409 134508424 + 3'UTR 12 2 0.00341905 288802[gene=chr4_f_c67294][PH=14][PH0=6.19][P=1,58e-01] Vps33a chr5 123529952 123529966 - 3'UTR 12 1 0.00341905 307338[gene=chr5_r_c63396][PH=13][PH0=6.31 ] [P=8.60e-02] Slc38a3 chr9 107651950 107651966 - CDS 12 2 0.00341905 314736[gene=chr9_r_c67228][PH=14][PH0=6.19] [P=3.84e-02] St7 chr6 17916162 17916175 + downstream 13 2 000341905 36822[gene=chr6_f_c7347][PH=15][PH0=532][P=832e-03]
[0557] 10K|intron
[0558] Fam175b chr7 132883689 132883696 + 3'UTR 12 2 0.00341905 386934[gene=chr7_f_c81165][PH=14][PH0=4.69][P=7.92e-03] Actn4 chr7 28893487 28893492 - 3'UTR 12 2 0.00341905 39944[gene=chr7_r_c7021][PH=14][PH0=6.95][P=6.57e-02] Camtai chr4 151801066 151801069 - intron 12 2 0.00341905 407690[gene=chr4_r_c82368][PH=14][PH0=4.24][P=6.64e-03] Arhgap18 chr10 26786392 26786418 + intron 12 1 000341905 43473[gene=chr10_f_c9557][PH=13][PH0=4.87][P=1.84e-02] Zfp266 chr9 20499284 20499285 - CDS 13 2 0.00341905 44383[gene=chr9_r_c10451][PH=15][PH0=6.65][P=6.27e-02] Lingo2 chr4 36151380 36151389 - intron 1 12 0.00341905 49162[gene=chr4_r_c10673][PH=13][PH0=4.47][P=1.44e-02] Gridl chr14 35375525 35375530 + intron 1 13 0.00341905 91769[gene=chr14_f_c22328][PH=14][PH0=5.87][P=3.27e-02] Nlgn3 chrX 101324631 101324639 + 3'UTR 12 2 0.00341905 98346[gene=chrXJ_c19340][PH=13][PH0=4.91][P=2.31e-02] Malatl chr19 5802563 5802564 - non-coding 22 49 0.00355118 18052[gene=chr19_r_c2896] [PH=71 ] [PH0=33.56] [P=3.30e-06] Ank3 chr10 69988359 69988364 + CDSjintron 23 6 0.00372178 121240[gene=chr10_f_c27580][PH=29][PH0=10.48][P=329e-04] Lnx1 chr5 74671385 74671393 - intron 7 23 000372178 200642[gene=chr5_r_c41295][PH=30][PH0=18 15][P=326e-02] Rab6b chr9 103183763 103183767 + 3'UTR 54 26 0.00375134 338504[gene=chr9_f_c60832][PH=79][PH0=28.47][P=4.08e-13] Pcdh9 chr14 93886975 93886977 - CDS 10 28 0.00393673 277078[gene=chr14_r_c51949][PH=38][PH0=11.81][P=8.19e-08] Pcdh7 chr5 58029071 58029094 + intron 3 19 0.00401138 199109[gene=chr5_f_c39064][PH=22][PH0=14.29][P=1.12e-01] Gpr83 chr9 14861784 14861797 + intron 5 20 0.00407982 22298[gene=chr9_f_c4710][PH=25][PH0=9.32][P=3.15e-04] Osbp!6 chr2 76428749 76428750 + intron 24 55 0.00413442 216993[gene=chr2_f_c43929][PH=79][PH0=30.78][P=1,56e-11 ] Mobp chr9 120156451 120156452 + intron 22 52 0.00416529 381075[gene=chr9_f_c70113][PH=74][PH0=15.09][P=1.62e-14] Clmn chr12 104863094 104863116 - intron 3 15 00041825 178069[gene=chr12_r_c38617][PH=18][PH0=982][P=596e-02] Prepl chr17 85086906 85086920 - intron 3 15 0.0041825 179404[gene=chr17_r_c37901 ][PH= 18][PH0=6.94] ]P=4,48e-03] Rnf152 chr1 105278568 105278571 - 3'UTR 2 14 0.0041825 212352[gene=chr1_r_c47602)[PH=16][PH0=9.75][P=4.67e-01] Dlg2 chr7 91644519 91644525 + intron 15 3 0.0041825 215217[gene=chr7_f_c43677][PH=18][PH0=7.43][P=7.63e-03] Cntn5 chr9 10496870 10496871 - intron 2 14 0.0041825 22971[gene=chr9_r_c5639][PH=16][PH0=5.91][P=6.04e-03]
[0559]
[0560] D17Wsu92e chr17 27764200 27764208 - downstream 3 16 0.0041825 48244[gene=chr17_r_c10255][PH=19][PH0=10.90][P=7.15e-02] lOKjintron
[0561] Cacnb2 chr2 14973623 14973637 + downstream 2 14 0.0041825 60322[gene=chr2_f_c8199][PH=16][PH0=8.73][P=7.86e-02]
[0562] 10K[intron
[0563] Lyrm4 chr13 35979561 35979563 - 3'UTR 2 15 0.0041825 63572[gene=chr13_r_c15403][PH= 17][PH0=7.22] [P=3.22e-02] Fam189a1 chr7 64954199 64954211 - intron 12 36 0.00418453 166946[gene=chr7_r_c30976][PH=48][PH0=13.08][P=8.64e-12] Azi2 chr9 118050716 118050718 + intron 51 110 0.00432514 374034[gene=chr9_f_c68395][PH=160][PH0=19.50][P=391e-14] Grm7 chr6 111566704 111566712 + 3'UTR 31 13 0.00432839 357967[gene=chr6_f_c73851][PH=44][PH0=13.25][P=1.53e-09] Ctnna2 chr6 77359325 77359333 - intron 4 19 000434582 129354[gene=chr6_r_c28274][PH=23][PH0=7.41][P=1.02e-04] HiatH chr13 65066382 65066390 - 3'UTR 4 18 0.00434582 136248[gene=chr13 r c31593][PH =22][PH0=9.15][P=5.15e-03] CoblH chr2 65106073 65106080 - intron 19 4 0.00434582 186842[gene=chr2_r_c42177][PH=23][PH0=9.32] [P= 1,53e-03] Grlfl chr7 16561768 16561774 - CDS 19 6 0.00434582 20382[gene=chr7_r_c3277][PH=25][PH0=6.29][P=1.12e-06] Sec24c chr14 20684727 20684733 + intron 4 18 0.00434582 29427[gene=chr14_f_c7604][PH=22][PH0=9.55][P=4.88e-03] Tm9sf3 chr19 41223167 41223169 - CDS 18 6 0.00434582 89658[gene=chr19_r_c20096][PH=24][PH0=11,21][P=9.64e-03] Syn3 chr10 86293222 86293244 - 3'UTRjintrcn 1 12 0.00436048 201864[gene=chr10_r_c46311][PH=13][PH0=6.06][P=6.43e-02] Rapgef6 chr11 54676088 54676103 + intron 3 18 000442706 104111 [gene=chr11 J_c22599][PH=21][PH0=12.71][P=8.00e-02] Atg3 chr16 45176241 45176246 + intron 3 17 0.00442706 312386[gene=chr16 _f_c49796][PH=20][PH0=7.49][P=2.80e-03] Tex2 chr11 106502353 106502364 - 3’UTR 17 4 0.00442706 373752[gene=chr11_r_c71697][PH=21][PH0=6.85][P=5.16e-04] Atxn7l1 chr12 33250212 33250213 + 3'UTRjintron 16 4 0.00442706 67938[gene=chr12_f_c16456][PH=20][PH0=9.27][P=4.67e-02] Stx18 chr5 38056112 38056123 + intron 12 2 0.00444495 176536[gene=chr5_f_c33679][PH=14][PH0=8.03][P=1.29e-01] Msi2 chr11 88449510 88449511 - intron 13 40 0.00459909 321418[gene=chr11_r_c61284][PH=53][PH0=10.94][P=1.00e-100] Pcdh7 chr5 57994529 57994562 + intron 3 16 0.00489756 198369[gene=chr5_f_c38908][PH=19][PH0=16.46][P=4.05e-01] Ank2 chr3 126968902 126968907 - intron 11 30 000510227 209256[gene=chr3_r_c43622][PH=41][PH0=12.74][P=1.99e-08] Ltn1 chr16 87377669 87377671 - 3'UTR 50 23 0.00519474 204946[gene=chr16_r_c42303][PH=73][PH0=21.43][P=1.00e-100] Grial chr11 57186931 57186933 + intron 8 25 0.00540195 121657[gene=chr11 _f_c25669][PH=33][PH0=11,32][P=491 e-06] Fam155a chr8 9587638 9587640 - intron 14 35 0.00544983 25500[gene=chr8_r_c3385][PH=48][PH0=18.26][P=2.47e-07] Cacnb4 chr2 52430329 52430330 - 3'UTR 39 18 0.00549324 161975[gene=chr2_r_c35753][PH=57][PH0=22.65][P=1.89e-07] Dab1 chr4 103991619 103991624 + intron 15 36 0.00552146 173200[gene=chr4_f_c40210][PH=51][PH0=12.29][P=3.12e-14] Gria4 chr9 4534702 4534716 - intron 6 21 0.00555689 5520[gene=chr9_r_c737][PH=27][PH0=893][P=3.47e-05] Rap2a chr14 120482218 120482236 + intron 13 1 000557665 233224[gene=chr14_f_c54321][PH=14][PH0=754][P=978e-02]
[0564]
[0565] Ubr4 chr4 139425322 139425323 + CDSjintron 22 6 0.00592828 305380[gene=chr4_f_c70955][PH=28][PH0=13.17][P=3.65e-03] Nlgnl chr3 26028563 26028565 - intron 7 23 0.00592828 42304[gene=chr3_r_c9019][PH=30][PH0=13.83][P=1.77e-03] Plbd2 chr5 120484878 120484879 - 3'UTR 122 232 0.00603636 288717[gene=chr5_r_c59873][PH=354][PH0=50.08][P=4.03e-13] Ildr2 chr1 166302739 166302763 + intron 3 15 0.00615219 325630[gene=chr1_f_c69530][PH=18][PH0=12.62][P=2.11 e-01 ] Csnklal chr18 61589749 61589750 + 3'UTR 21 44 0.00622624 151980[gene=chr18_f_c32696][PH=65][PH0=20.45][P=9.85e-13] Ptprs chr17 56440551 56440553 - intron 1 11 00063493 100879[gene=chr17_r_c21306][PH=12][PH0=409][P=262e-02] Rbfox.1 chr16 7190257 7190264 + intron 1 12 0.0063493 109859[gene=chr16_f_c12377][PH=13][PH0=4.95][P=2.03e-02] Nrg3 chr14 38658892 38658905 - intron 1 11 0.0063493 110477[gene=chr14_r_c23874][PH=12][PH0=4.75][P=3.48e-02] Ank3 chr10 69885184 69885188 + CDSjintron 12 1 0.0063493 116735[gene=chr10_f_c27021 ][P H= 13][PH0=5.52][P=3.52e-02] Kcnn2 chr18 45431465 45431479 + intron 1 11 0.0063493 121451[gene=chr18_f_c24885][PH=12][PH0=6.31][P=1.18e-01] Megf9 chr4 70491002 70491012 - intron 2 11 0.0063493 134622[gene=chr4_r_c28731][PH=13][PH0=807j[P=2.73e-01] Epha6 chr16 60393163 60393166 - intron 1 12 0.0063493 135852[gene=chr16_r_c28511][PH=13][PH0=6.68][P=9.17e-02] Foxn3 chr12 99266900 99266907 - intron 2 11 00063493 154841 [gene=chr12 _r_c33073][PH= 13][PH0=665] [P=1.00e-01 ] Kcnj3 chr2 55535452 55535464 + intron 11 1 0.0063493 162751 [gene=chr2_f_c31996][PH=12][PH0=6.39][P=1.24e-01] Atg9a chr1 75186254 75186265 - CDS 11 2 0.0063493 166510[gene=chr1_r_c36915][PH=13][PH0=661][P=1.71e-01j Chst11 chr1O 83098086 83098111 + intron 1 11 0.0063493 167378[gene=chr10_f_c38091][PH=12][PH0=7.68][P=2.31e-01] Trio chr15 27779612 27779616 - intron 2 11 0.0063493 18083[gene=chr15_r_c5239][PH=13][PH0=3.93][P=6.92e-03] Scn2a1 chr2 65724305 65724322 + intron 11 2 0.0063493 182203[gene=chr2_f_c36629][PH=13][PH0=9.73][P=3.30e-01] Pde4d chr13 109143085 109143091 + intron 2 11 0.0063493 182350[gene=chr13_f_c42168][PH=13][PH0=6.85][P=2.85e-01] Dlgapl chr17 70184422 70184424 + intron 1 11 00063493 183681[gene=chr17_f_c42440][PH=12][PH0=673][P=324e-01] Rgs6 chr12 82949776 82949778 + intron 2 11 0.0063493 184798[gene=chr12_f_c43840][PH = 13] [PH0=5.15][P=2.61 e-02] Usp46 chr5 74000749 74000754 - 3'UTR 11 2 0.0063493 197455[gene=chr5_r_c40249][PH=13][PH0=4.56][P=2.12e-02] Gpr56 chr8 95011791 95011805 CDS 12 2 0.0063493 204796[gene=chr8_f_c45814j[PH=14j[PH0=6.38][P=4.12e-02] Pafahlbl chr11 74720203 74720207 - intron 12 2 0.0063493 209928[gene=chr11 _r_c44123][PH= 14][PH0=7.62] [P=1.13e-01] Abr chr11 76463642 76463643 - CDS 11 2 0.0063493 217102[gene=chr11 _r_c45929][PH= 13][PH0=5.18][P=2.94e-02] Carlo chr11 93268234 93268252 + intron 1 12 0.0063493 239264[gene=chr11_f_c50822][PH=13][PH0=5.89][P=5.54e-02] Rora chr9 68757549 68757550 + intron 11 2 00063493 242264[gene=chr9_f_c42091][PH=13][PH0=468][P=1 70e-02] Svop chr5 114026972 114026985 - 3'UTR 12 2 0.0063493 273107[gene=chr5 r c56412][PH = 14][PH0=493] [P=2.05e-02]
[0566]
[0567] Nfasc chr1 132624330 132624341 - intron 1 11 0.0063493 284961 [gene=chr1_r_c64083][PH=12][PH0=6.31][P=1.18e-01] Parva chr7 112588604 112588615 + CDS 11 1 0.0063493 305980[gene=chr7_f_c64662][PH=12][PH0=5.42][P=6.83e-02] Spen chr4 141512222 141512235 - intron 1 11 0.0063493 316763[gene=chr4_r_c67497][PH=12][PH0=5.05][P=4.32e-02] Vamp4 chr1 162597669 162597670 + 3'UTR 12 2 0.0063493 320164[gene=chr1_f_c68357][PH=14][PH0=6.87][P=1,99e-01 ] Prune2 chr1S 17221018 17221028 + 3'UTR 1 11 0.0063493 32655[gene=chr19_f_c6658][PH=12][PH0=7.07][P=2.57e-01] Pou2f1 chr1 165866880 165866885 - 3'UTR 11 1 00063493 369368[gene=chr1_r_c82165][PH=12][PH0=4.10][P=1.62e-02] Slc35c2 chr2 165282831 165282850 - CDS|downstream 11 2 0.0063493 414986[gene=chr2_r_c91440][PH=13][PH0=6.31][P=7.06e-02]
[0568] 10K
[0569] Nsg2 chr11 32037173 32037177 + downstream 12 2 0.0063493 44784[gene=chr11_f_c10024][PH=14][PH0=7.49][P=1.43e-01]
[0570] 10K|intron
[0571] Mcf2 chrX 60056436 60056445 - 3'UTR 1 11 0.0063493 53094[gene=chrX_r_c12000][PH=12][PH0=5.62][P=1.46e-01] Bin! chr18 32425420 32425422 + intron 1 11 0.0063493 55756[gene=chr18_f_c13157][PH=12][PH0=4.39][P=3.38e-02] Ntm chr9 29932789 29932816 - intron 1 11 0.0063493 96131 [gene=chr9_r_c 19286] [P H= 12] [PH0=6.97] [P= 1.71 e-01 ] Ptprg chr14 11794267 11794272 intron 11 2 0.0063493 9921 [gene=chr14_f_c2352][PH=13][PH0=4.50][P=1,40e-02] Scn8a chr15 100948358 100948361 + downstream 12 31 0.00643244 243072[gene=chr15_f_c53079][PH=43][PH0=11,25][P=5,52e-11]
[0572] 10K|intron
[0573] AW549877 chr15 3983814 3983821 - 3'UTR 25 8 0.00652143 1863[gene=chr 15_r_c536] [PH =33] [PH0= 12.19][P=3.08e-05] Gria2 chr3 80688704 80688706 - 3'UTR 17 37 0.00660797 116522[gene=chr3_r_c25424][PH=54][PH0=8.27][P=3.87e-14] Cdh23 chr10 60323354 60323357 - CDS 20 7 0.00661429 139347[gene=chr10_r_c32807][PH=27][PHO=11.30][P=1.00e-03] Ppp1r13b chr12 111833650 111833654 - CDS 20 7 0.00661429 186042[gene=chr12_r_c40978][PH=26][PH0=8.32][P=2.31e-05] Gpr158 chr2 21532663 21532669 + intron 6 21 000661429 82137[gene=chr2_f_c13433][PH=27][PH0=1242][P=254e-03] Plcxd2 chr16 45962796 45962807 - 3'UTR 34 14 0.00661558 115904[gene=chr16_r_c23848][PH=48][PH0=23.97] [P=2.79e-03] Slc9a6 chrX 56662332 56662338 + 3'UTR 15 33 0.00661558 36506[gene=chrXJ_c7704][PH=48][PH0=20.98][P=2.33e-05] Nrg3 chr14 39347080 39347081 - intron 6 21 0.00662165 137104[gene=chr14_r_c28013][PH=27][PH0= 13.13][P=1 00e-02] Asxl2 chr12 3465161 3465179 + downstream 1 11 0.00674359 2824[gene=chr12_f_c352][PH=12][PH0=6.88][P=1.63e-01]
[0574] 10K|intron
[0575] Taccl chr8 25157244 25157245 - 3'UTR 30 14 0.00677474 96870[gene=chr8_r_c14876][PH=44][PH0=15.65][P=1.14e-06] Pcdh17 chr14 84535997 84536000 + 3'UTR 9 25 0.00700535 179904[gene=chr14_f_c40472][PH=34][PHO=21.49][P=4.65e-01] Plekhm2 chr4 141626063 141626064 - 3'UTR 25 9 0.00700535 316985[gene=chr4_r_c67569][PH=34][PH0=15.33][P=5.32e-04] Bai3 chr1 25639840 25639841 - intron 25 10 0.00700535 74076[gene=chr1_r_c16357][PH=35][PHO=13.91][P=3.84e-05] Prexl chr2 166673382 166673384 - intron 16 41 0.00718109 422880[gene=chr2_r_c93126][PH=57][PH0= 11,64][P=2.30e- 14]
[0576]
[0577] Rab11fip2 chr19 59904994 59905002 - 3‘UTR 18 6 0.00720067 131541[gene=chr19_r_c30851][PH=24][PH0=10.26][P=1.17e-02] Prkce chr17 86189466 86189468 + intron 18 6 0.00720067 215446[gene=chr17_f_c50287][PH=23][PH0=7.79][P=3.89e-04] MapkIO chr5 102911709 102911719 - 3‘UTR 18 5 0.00720067 235805[gene=chr5_r_c48873][PH=23][PH0=10.96][P=1.36e-02] Snx5 chr2 144250967 144250972 - 3'UTR 4 18 0.00720067 351896[gene=chr2_r_c77937][PH=22][PH0=6.94][P=3.44e-04] Grial chr11 57100290 57100293 + intron 14 3 0.00738761 118324[gene=chr11 _f_c25112][PH=17][PH0=6.18][P=6.25e-03] Macrodl chr19 7109426 7109433 + intron 3 13 000738761 13464[gene=chr19_f_c2895][PH=15][PH0=521][P=747e-03] Zfp458 chr13 67255954 67255964 - 3'UTR 3 13 0.00738761 138590[gene=chr13_r_c32167][PH=16][PH0=6.95][P=4.00e-02] MtmdO chr7 64317457 64317458 + downstream 2 13 0.00738761 142217[gene=chr7_f_c26384][PH=15][PH0=5.48][P=9.18e-03]
[0578] WKjintron
[0579] Ubxn7 chr16 32385529 32385531 + 3‘UTR 13 3 0.00738761 184185[gene=chr16_f_c28439][PH=16][PH0=4.88][P=3.35e-03] Pum2 chr12 8750486 8750493 + 3'UTR 13 3 000738761 21380[gene=chr12_f_c5115][PH=16] [PH0=573] [P=727e-03] Spared chr5 104088440 104088454 - CDS 13 2 0.00738761 243556[gene=chr5_r_c50477][PH=15][PH0=9.18][P=1.66e-01] Cntn5 chr9 10821898 10821905 - intron 2 14 0.00738761 26787[gene=chr9_r_c6611][PH=16][PH0=10.29][P=1,86e-01 ] Lsamp chr16 41655904 41655910 + intron 3 13 0.00738761 273646[gene=chr16_f_c43757][PH=16][PH0=5.06][P=4.35e-03] Zfp280d chr9 72343292 72343308 + downstream 14 3 0.00738761 280799[gene=chr9_f_c47244][PH=17][PH0=5.63][P=2.49e-03]
[0580] 10K|intron
[0581] Ksr2 chr5 117657054 117657056 + intron 13 2 0.00738761 322715[gene=chr5_f_c67714][PH=15][PH0=5.95][P=2.05e-02] Fgf14 chr14 124004855 124004856 - intron 14 3 0.00738761 327855[gene=chr14_r_c63184][PH=17][PH0=8.31][P=5.00e-02] Ptprt chr2 162532839 162532844 - intron 2 14 0.00738761 406485[gene=chr2_r_c89439][PH=16][PH0=6.36][P=1.17e-02] Aatk chr11 120011690 120011693 - CDS 3 13 0.00738761 448858[gene=chr11_r_c86316][PH=16][PH0=4.06][P=3.66e-04] Csmdl chr8 16865361 16865362 - intron 3 14 0.00738761 65016[gene=chr8_r_c9802] [PH = 17][PH0=5.42][P= 1,23e-03] Stard4 chr18 33201593 33201613 - 3'UTR 3 14 0.00738761 68258[gene=chr18_r_c14133][PH=17][PH0=4.53][P=2.48e-04] Ralgpsl chr2 33151900 33151920 - intron 2 14 000738761 84741 [gene=chr2_r_c18587] [PH=16][PH0=6.83][P=1.84e-02] 1810013L24Rik chr16 8843517 8843523 + intron 15 4 0.00754075 120910[gene=chr16 _f_c14078][PH= 19] [PH0=6.00][P= 1,34e-03] Arhgef12 chr9 42965228 42965231 - 3'UTR 15 4 0.00754075 126616[gene=chr9_r_c26635][PH=19][PH0=6.78] [P=2.22e-03] Bail chr15 74553164 74553174 + intron 4 16 0.00754075 142801[gene=chr15_f_c33450][PH=20][PH0=6.61][P=1.28e-03] Nrxnl chr17 90702013 90702019 - intron 3 16 0.00754075 204180[gene=chr17_r_c42399][PH=19][PH0=8.83][P=1.87e-02] Ralgps2 chr1 156806788 156806794 - 3'UTR 16 4 0.00754075 339859[gene=chr1_r_c75429][PH=20][PH0=6.67][P=5.76e-04] Mapk8ip3 chr17 24897630 24897639 - 3'UTR]intron 3 16 0.00754075 34751[gene=chr17_r_c7362][PH=19][PH0=7.69][P=7.86e-03] Rab11fip3 chr17 25990143 25990149 - 3'UTR 15 4 000754075 41880[gene=chr17_r_c8747][PH=19][PH0=552][P=274e-04]
[0582]
[0583] Arhgap5 chr12 52559290 52559308 + intron 15 4 0.00754075 99670[gene=chr12_f_c24798][PH=19][PH0=6.82][P=1.92e-03] Dpysl3 chr18 43323323 43323339 - 3'UTR 13 3 0.00763313 97858[gene=chr18_r_c20773][PH=16][PH0=8.10][P=4.27e-01j Ctnna2 chr6 77507171 77507177 - intron 28 13 0.00763827 137154[gene=chr6_r_c29339][PH=41][PH0=6.72][P=1.67e-14] Ptprd chr4 76059952 76059955 - intron 11 27 0.00763827 139529[gene=chr4_r_c29682][PH=38][PH0=10.17][P=1.71 e-09] Akt3 chr1 177022933 177022934 - 3'UTR|intron 28 10 0.00763827 419681 [gene=chr1 _r_c91076][PH=38][PH0= 12.53] [P=6,49e-07] Trpm3 chr19 22989168 22989171 + 3'UTR 11 28 000763827 62303[gene=chr19_f_c12516][PH=39][PH0=1940][P=1 16e-02] Speed 1 chr1O 75276900 75276917 + CDS 15 3 0.00778858 149493[gene=chr10 _f_c34045][PH=18][PH0=10.331[P=805e-02] Nrxn3 chr12 89679467 89679470 + intron 13 31 0.00792315 233329[gene=chr12_f_c53675][PH=44][PH0=14.79j[P=502e-08] PedhIO chr3 45422808 45422814 + 3'UTR|intron 13 32 0.00792315 79265[gene=chr3_f_c17900][PH=45][PH0=28.22][P= 1 27e-02] Plekha6 chr1 133272277 133272284 + CDS 36 15 0.00795088 251699[gene=chr1 _f_c55952][P H=51 ][PH0= 11,40][P= 1.00e-100] Kcnip4 chr5 48667122 48667124 - intron 8 24 0.00813512 101305[gene=chr5_r_c23377][PH=32][PH0=15.96][P=4.20e-03] Kirrei3 chr9 34838356 34838357 + intron 7 24 0.00813512 131287[gene=chr9_f_c20770][PH=31][PH0=6.25][P=1.56e-10] Nrg3 chr14 39392839 39392846 - intron 23 9 000813512 140336[gene=chr14_r_c28343][PH=32][PH0=7.09][P=1.22e-09] Trp53bp1 chr2 121198311 121198313 - 3'UTR|intron 23 8 0.00813512 303681[gene=chr2_r_c67553][PH=31][PH0=7.88][P=3.85e-08] Pitpnd chr11 107212257 107212278 - 3’UTR 8 23 0.00813512 379186[gene=chr11_r_c72624][PH=31][PH0=9.97][P=3.63e-06] Mobp chr9 120160123 120160125 + intron 24 9 0.00813512 381850[gene=chr9_f_c70122][PH=33][PH0=13.57][P=1.72e-03] Fgf14 chr14 124574783 124574784 - intron 20 43 0.00865247 348422[gene=chr14_r_c66365][PH=63][PH0=17.75][P=3.87e-14] Nvl chr1 181103308 181103316 - intron 20 44 0.00865247 436813[gene=chr1_r_c94918][PH=64][PH0=26.03][P=9.24e-09] Slc25a12 chr2 71272613 71272614 - 3'UTR 26 10 0.00904766 201327[gene=chr2_r_c45292][PH=36][PH0=15.57][P=3.06e-04] Grik2 chrtO 49262892 49262893 - intron 17 38 000908375 119788[gene=chr10_r_c28184][PH=55][PH0=31 37][P=267e-03] Frmd3 chr4 74158541 74158545 + intron 2 14 0.00923366 99602[gene=chr4_f_c23253][PH=16][PH0=7.85][P=4.29e-02] Snx27 chr3 94526463 94526464 - CDS|downstream 7 20 0.00936035 143799[gene=chr3_r_c30588][PH=27][PH0=11,56][P=1,07e-03]
[0584] 10K|intron
[0585] Scamp5 chr9 57443829 57443830 - CDS 21 8 0.00936035 184996[gene=chr9_r_c37744][PH=29][PH0=9.47][P=1.78e-05] Pbx3 chr2 34172068 34172072 - 3'UTR 21 6 000936035 97894[gene=chr2_r_c21409][PH=27][PH0=1229][P=419e-03] Hspa12a chr19 58840035 58840041 - intron 12 1 0.0093801 125266[gene=chr19_r_c29378][PH= 13][PH0=5.63] [P=1.17e-01] Mycbp2 chr14 103295678 103295687 - intron 3 14 0.00939383 299560[gene=chr14_r_c56989][PH=17][PH0=10.60][P= 1,49e-01] Ephbl chr9 102163510 102163516 - intron 36 16 0.00940893 292238[gene=chr9_r_c62047][PH=52][PH0=13.48][P=1,48e-13] Gria2 chr3 80691959 80691960 - intron 12 28 0.00948303 116867[gene=chr3_r_c25446][PH=40][PH0=10.25][P=2.96e-10]
[0586]
[0587] Negri chr3 156892225 156892228 + intron 11 30 0.00948303 310769[gene=chr3_f_c67188][PH=41][PH0=12.88][P=3.88e-08] Rgs9 chr11 109230439 109230441 - downstream 12 29 0.00948303 393390[gene=chr11_r_c76035][PH=41][PH0=21.08][P=1.87e-02]
[0588] 10K|intron
[0589] Purg chr8 33404377 33404378 + intron 32 14 000956867 61592[gene=chr8_f_c14691][PH=46][PH0=14.36][P=1.07e-08] Pgbd5 chr8 124369837 124369844 - 3‘UTR 20 8 0.00978145 355417[gene=chr8_r_c73946][PH=28][PH0=7.95][P=2.76e-06]
[0590]
[0591]
[0250] TABLE 4
[0592] Differential alternative splicing (AS) events in Novalhu / hu mice
[0593] List of events for which differences were detected in AS analysis in the midbrain of Novalhu / hu and Novalwt / wt mice (p-value<0.05, |d7]>0.05). For each event, the gene name, chromosome number, genomic start and end position of the event, strand, splicing type, PSI values and its differences (APSI, di), / 7-value are listed. PSI: percent spliced-in value, the percent of
[0594] transcripts that include a specific AS exon. APSI: percent change in Novalhu / hu vs. Novalwt / wt. AS splicing events are classified into the following types: Cassette exon (cass), alternative 5’ splice site (alt5), alternative 3’ splice site (alt3), tandem cassette (taca), mutually exclusive exons (mutx), intron retention (het). Novalwt / wt N=4, Novalhu / hu N=4.Gene_name chr start end strand AS type PSI_NOVA1(wt / wt) PSI_NOVA1(hu / hu) delta. PSI Fnbpll chr3 122549155 122558105 - taca 0 152 0.361 -0.209 Ccp110 chr7 118732393 118737023 + cass 0.637 0.754 -0.117 Itprl chr6 108431460 108438431 + cass 0864 0986 -0 123 Tle6 chr10 81595344 81595554 alt3 0.895 0.360 0.535 Tle6 chr10 81595322 81595554 - alt3 0.895 0.360 0.535 Ccdc163 chr4 116712722 116714148 + alt3 0.102 0.625 -0.523 Rev3l chrdO 39732159 39783331 + cass 0.194 0.627 -0.434 Mrps18b chr17 35911418 35914401 - cass 0.924 0.995 -0.071 Rhbdd3 chr11 5105180 5105716 alt5 0.813 0.929 -0.117 Dvl1 chr4 155856613 155859302 + cass 0.381 0289 0.091 Impa2 chr18 67302175 67306774 alt3 0.020 0.164 -0.144 Tcf3 chr1O 80415364 80415925 alt3 0.399 0.572 -0.173 AbhdIO chr16 45737523 45742890 - cass 0.972 0.866 0.106 Htr3a chr9 48899212 48900781 alt3 0.004 0.278 -0.274 Zfp655 chr5 145233440 145235862 + cass 0.715 1.000 -0.285 Col4a5 chrX 141656684 141666974 + taca 0.076 0.375 -0.299 Stabl chr14 31142957 31143266 - iret 0295 0075 0220 Rpain chr11 70973014 70977801 + taca 0.707 0.882 -0.175 likap chr1 91373860 91376398 - alt3 0.143 0.857 -0.714 Cenpa chr5 30666896 30672569 + taca 0.073 0.700 -0.627 Atpaf2 chr11 60405775 60407390 - alt3 0.110 0.167 -0.057 BC029214 chr2 25459774 25460073 alt3 0.139 0.208 -0.069 Fam73b chr2 30381938 30383018 + iret 0.134 0.081 0.053 Snapc4 chr2 26363637 26366036 - alt3 0219 0331 -0 112 6720401G13Rik chrX 50608477 50608813 alt3 0.571 0.827 -0.255 Ctc1 chr11 69021029 69022619 + alt3 0.810 0.586 0.224
[0595]
[0596] Ctc1 chr11 69021029 69022613 + alt3 0.810 0.586 0.2242700062C07Rik chr18 24471061 24473148 + alt5 0.021 0.098 -0.078 Pex5l chr3 33014964 33143276 taca 0.495 0.363 0.131 Setd2 chr9 110573880 110590435 + cass 0086 0034 0052 Gnbll chr16 18540971 18544272 + alt3 0.609 0.966 -0.357 Dlg1 chr16 31842769 31853883 + mutx 0.250 0.174 0.076 Amd1 chr10 40290117 40290614 - iret 0.130 0.364 -0.233 Lrrc16b chr14 55501458 55502432 + fret 0.114 0.055 0.058 Fam21 chr6 116208206 116209117 + alt5 0.976 0.925 0.051 Rtell chr2 181335892 181339093 + cass 0.747 0.445 0.301 Abtbl chr6 88839626 88840899 cass 0642 0776 -0 134 Rps6kl1 chr12 85147745 85149907 alt5 0.828 0.938 -0.109 Atrx chrX 105879908 105884419 alt3 0.933 0.718 0.215 Pfkfb3 chr2 11471421 11478057 - cass 0.289 0.468 -0.179 Zfp445 chr9 122856698 122857160 - alt5 0.298 0.370 -0.071 Camtai chr4 151061378 151074061 - cass 0.637 0.565 0.072 Rnf111 chr9 70475775 70503726 - cass 0.422 0.791 -0.369 Zfp239 chr6 117863083 117869272 + taca 0075 0200 -0 125 Gphn chr12 78492037 78504750 + cass 0.187 0.263 -0.076 Itprl chr6 108417888 108438431 + taca 0.729 0.601 0.128 Trp63 chr16 25866074 25868248 + alt5 0.667 0.192 0.474 Samp chr10 128821922 128833410 + cass 0.586 0.776 -0.190 Hsf4 chr8 105269800 105270895 + cass 0.868 0.989 -0.121 Cadml chr9 47818733 47848297 + cass 0.909 0.825 0.084 Snapc3 chr4 83418705 83435294 cass 0965 0832 0133 Hsd3b7 chr7 127801833 127802112 4- iret 0.345 0.540 -0.195 Mprip chr11 59771616 59772216 iret 0.092 0.146 -0.054 Crispld2 chr8 120018852 120023701 + alt3 0.143 0.527 -0.384
[0597]
[0598] A830018L16Rik chr1 11414183 11518708 + cass 1.000 0.899 0.101Dyrkla chr16 94659496 94663860 + alt3 0.404 0.541 -0.136 Hmqn3 chr9 83109941 83111105 iret 0.583 0.688 -0.105 Vidir chr19 27239621 27240636 + cass 0994 0927 0067 Naa40 chr19 7230067 7230282 - iret 0.431 0.584 -0.153 Naa40 chr19 7230023 7230282 iret 0.431 0.584 -0.153 Naa40 chr19 7230073 7230282 - iret 0.431 0.584 -0.153 Mast2 chr4 116330390 116337496 cass 0.774 0.870 -0.096 Gak chr5 108569418 108569955 - iret 0.866 0.917 -0.051 Gak chr5 108569422 108569955 - iret 0.866 0.917 -0.051 Ccdc50 chr16 27406650 27409413 + cass 0048 0122 -0074 Cwc22 chr2 77936592 77946181 cass 0.843 0.506 0.337 Ppp3cc chr14 70256359 70267505 cass 1.000 0.941 0.059 Ppan chr9 20889337 20889678 + iret 0.049 0.168 -0.119 Appl2 chr10 83610977 83611198 - iret 0.300 0.386 -0.086 Suv39h1 chrX 8063695 8064056 - iret 0.016 0.074 -0.057 Galnt7 chr8 57540019 57545407 - mutx 0.678 0.466 0.212 Wdr70 chr15 8093675 8099020 - alt3 0794 0994 -0200 Cyp4f13 chr17 32932628 32941199 - cass 0.132 0.042 0.090 Cabinl chr10 75725258 75725754 alt3 0.916 0.983 -0.067 Ptprd chr4 76091421 76099604 - cass 0.316 0.478 -0.162 Osbp!9 chr4 109087394 109098552 - cass 0.659 0.913 -0.254 Ccdc50 chr16 27435561 27436601 + cass 0.925 0.990 -0.065 Phldbl chr9 44696036 44697841 alt5 0.289 0.351 -0.063 Foxml chr6 128371967 128372602 alt3 0417 0000 0417 Dnajc6 chr4 101508103 101597950 4- cass 0.449 0.376 0.073 Cacnb2 chr2 14963905 14971641 mutx 0.397 0.485 -0.088 Banp chr8 122007752 122024145 + cass 0.653 0.494 0.160
[0599]
[0600] Thada chr17 84464356 84466207 - alt3 0.458 0.880 -0.422Btbd6 chr12 112976675 112977370 + cass 0.986 0.912 0.073 Rasqrp2 chr19 6400546 6401862 + alt3 0.977 0.577 0.400 Lig3 chr11 82785314 82787827 + alt3 0672 0842 -0 171 Bdp1 chr13 100098952 100103877 - alt3 0.847 1.000 -0.153 Prrc2c chr1 162674098 162676816 cass 0.910 0.834 0.075 Bap1 chr14 31256660 31257766 + iret 0.923 0.974 -0.051 MbnH chr3 60501264 60595763 + cass 0.716 0.905 -0.189 Ubr2 chr17 46975926 46981441 - mutx 0.378 0.299 0.079 Gas7 chr11 67629609 67652911 + cass 0.907 0.848 0.060 Zfp7 chr15 76881111 76888346 + alt3 0705 0937 -0233 Nme6 chr9 109833195 109835413 alt3 0.724 0.976 -0.252 Wdr52 chr16 44413647 44416052 4- cass 0.053 0.440 -0.388 Zfp12 chr5 143239953 143240414 + alt5 0.178 0.273 -0.095 Mttnr2 chr9 13782276 13785926 + cass 0.134 0.083 0.051 Cars chr7 143559648 143563040 - cass 0.920 0.974 -0.054 Cgn chr3 94777993 94778269 - iret 0.385 0.000 0.385 Fbxw9 chr8 85064383 85064671 + iret 0444 0286 0158 Chrd chr16 20735745 20736142 + iret 0.023 0.169 -0.146 Snrk chr9 122117319 122137600 + taca 0.625 0.738 -0.113 Rimsl chr1 22346292 22373509 - alt3 0.798 0.890 -0.092 Hr chr14 70552244 70555106 + cass 0.242 0.577 -0.334 Slc25a19 chr11 115624198 115628111 - cass 0.068 0.128 -0.060 Fam161b chr12 84345316 84346934 - iret 0.425 0.264 0.161 Vldlr chr19 27240465 27241452 alt3 0968 0894 0075 Slmap chr14 26428241 26438737 cass 0.852 0.774 0.078 Cln3 chr7 126582999 126583446 alt3 0.020 0.157 -0.136 Baz2b chr2 59933670 59936750 - cass 0.716 0.893 -0.178
[0601]
[0602] Spopl chr2 23537460 23543323 - cass 0.244 0.145 0.099Psma3 chr12 70983318 70984593 + alt3 0.971 0.915 0.056 Atg9a chr1 75190822 75191954 alt3 0.398 0.491 -0.093 Dlg1 chr16 31846833 31853883 + cass 0536 0645 -0 109 Zfp273 chr13 67813815 67822366 + alt3 1.000 0.118 0.882 Inpp5e chr2 26399232 26399573 iret 0.200 0.293 -0.093 Saft>2 chr17 56562938 56563430 - iret 0.163 0.229 -0.066 Clybl chr14 122181716 122371428 + taca 0.642 0.861 -0.219 My!6 chr10 128492601 128493352 - cass 0.151 0.242 -0.091 Nemf chr12 69353793 69356131 - cass 0.871 0.969 -0.098 Drosha chr15 12881622 12883264 + alt5 0859 0915 -0055 Brf1 chr12 112964183 112965991 alt3 0.908 0.811 0.097 Mtg2 chr2 180084450 180085901 4- alt5 0.924 0.998 -0.074 Nfib chr4 82310302 82327912 - taca 0.751 0.965 -0.215 Zfp87 chr13 67515780 67520720 - cass 0.070 0.001 0.069 Gpatch2l chr12 86268975 86291369 + cass 0.740 0.854 -0.114 Med16 chr10 79907490 79908918 - alt3 0.839 0.974 -0.135 Naa60 chr16 3884695 3897379 + taca 0817 0904 -0086 Ppfibpl chr6 147012454 147016392 + cass 0.049 0.140 -0.091 Ube4a chr9 44959978 44965548 cass 0.552 0.742 -0.190 N28178 chr4 42916603 42933666 + cass 0.812 0.698 0.113 Gm 16596 chr12 108536401 108539472 - cass 0.000 0.067 -0.067 Robo2 chr16 73897015 73904555 - cass 0.243 0.346 -0.103 Lrrcd chr3 14550250 14551526 + alt3 0.046 0.200 -0.154 B3gaint2 chr13 13966309 13987462 mutx 0482 0364 0118 C2cd3 chr7 100374263 100380124 4- alt3 0.786 0.979 -0.193 Trmt11 chr10 30590842 30591101 iret 0.000 0.110 -0.110 Ip6k1 chr9 108044724 108045538 + cass 0.253 0.193 0.059
[0603]
[0604] Slc35a4 chr18 36681527 36682053 + iret 0.453 0.533 -0.080Slc35a4 chr18 36681527 36683856 + iret 0.453 0.533 -0.080 Slc35a4 chr18 36681527 36683863 + iret 0.453 0.533 -0.080 Sic35a4 chr18 36681527 36682030 + iret 0453 0533 -0080 Ptprq chr14 12166744 12190772 + cass 0.559 0.388 0.171 Cadps chr14 12467022 12473514 - cass 0.775 0.825 -0.050 Pak3 chrX 143712277 143715288 + cass 0.861 0.971 -0.110 Pari chr16 20287856 20293401 alt5 0.017 0.073 -0.056 Tmem70 chr1 16665225 16667772 + alt3 0.106 0.041 0.065 Tmem70 chr1 16665225 16667769 + alt3 0.106 0.041 0.065 Cdk13 chr13 17719355 17721292 alt3 0367 0442 -0075 Myo7a chr7 98098186 98102701 cass 0.419 0.204 0.215 Kdm6a chrX 18246974 18248358 4- cass 0.443 0.699 -0.257 Tcergl chr18 42573259 42575554 + cass 0.073 0.142 -0.069 Clipl chr5 123627293 123631168 - alt5 0.555 0.794 -0.240 Mgl2 chr11 70134118 70135161 + alt3 0.600 0.000 0.600 Suds3 chr5 117091679 117093030 - iret 0.602 0.530 0.072 Suds3 chr5 117092682 117093030 - iret 0602 0530 0072 Suds3 chr5 117092434 117093030 - iret 0.602 0.530 0.072 Bicd2 chr13 49383042 49387024 + iret 0.397 0.479 -0.083 Sic15a2 chr16 36771876 36774640 - cass 0.938 0.816 0.122 Dock11 chrX 36009338 36011809 + cass 0.712 0.333 0.378 Eftudl chr7 82672809 82674581 + alt3 0.969 0.825 0.145 Rtell chr2 181355923 181356615 + iret 0.975 0.862 0.113 Rtell chr2 181355923 181356427 iret 0975 0862 0113 Cenpt chr8 105849625 105849922 iret 0.109 0.231 -0.122 Rtell chr2 181356157 181356615 alt3 0.333 0.014 0.319 Rtell chr2 181356157 181356614 + alt3 0.333 0.014 0.319
[0605]
[0606] Hnrnpk chr13 58395573 58396898 - cass 0.227 0.304 -0.077Stk30 chr12 110807798 110808404 - alt3 0.998 0.885 0.114 Dcaf17 chr2 71078059 71082052 + cass 0.689 0.881 -0.192 Snph chr2 151597108 151601073 cass 0801 0860 -0060 Golgal chr2 39047072 39047765 iret 0.164 0.280 -0.117 Golgal chr2 39047022 39047765 - iret 0.164 0.280 -0.117 Ern1 chr11 106434914 106459027 - cass 1.000 0.715 0.285 Trdmtl chr2 13523425 13525693 cass 0.176 0.060 0.116 Wdr47 chr3 108618489 108623448 + alt3 0.815 0.733 0.082 Ptpmtt chr2 90916853 90917508 - alt3 0.286 0.000 0.286 D3Ertd751e chr3 41756048 41758929 + iret 0925 1 000 -0075 Nudtl chr5 140331908 140334618 cass 0.735 0.941 -0.206 Trmt2a chr16 18250819 18251659 4- taca 0.815 0.912 -0.097 Mapk8ip3 chr17 24892151 24894222 - taca 0.906 0.973 -0.067 Vtila chr19 55380941 55391867 + cass 0.663 0.765 -0.102 Plxnb2 chr15 89170568 89180787 - cass 0.928 1.000 -0.072 Smekl chr12 101051530 101053583 - alt5 0.134 0.200 -0.066 Arap3 chr18 37974364 37974680 - iret 0000 0089 -0089 Kif22 chr7 127027728 127027989 - iret 0.067 0.197 -0.130 Zcchc6 chr13 59771879 59782342 taca 0.048 0.106 -0.058 Apbb3 chr18 36676864 36677255 - alt3 0.370 0.267 0.103 Myh11 chr16 14246726 14250623 - cass 0.938 0.500 0.438 Med12l chr3 59037559 59042412 + cass 0.776 0.560 0.216 Sec23ip chr7 128778428 128779224 + alt3 0.940 0.871 0.069 Fance chr17 28317467 28318155 alt3 0368 0227 0141 Slc19a2 chr1 164256746 164261016 4- alt5 0.887 0.975 -0.088 Trmt13 chr3 116590229 116592258 cass 0.000 0.083 -0.083 Fbxwl 1 chr11 32642795 32711996 + cass 0.215 0.368 -0.152
[0607]
[0608] Pcgf5 chr19 36437154 36437361 + alt3 0.931 0.999 -0.068Smox chr2 131520094 131522218 + taca 0.751 0.805 -0.054 Man2b1 chr8 85084400 85084813 + alt3 0.105 0.173 -0.068 Man2b1 chr8 85084400 85084735 + alt3 0105 0173 -0068 Man2b1 chr8 85084400 85084724 + alt3 0.105 0.173 -0.068 Wrn chr8 33353294 33385526 cass 0.000 0.158 -0.158 Chrd chr16 20738552 20739163 + cass 0.686 0.921 -0.235 Crlfl chr8 70503304 70503638 + alt3 0.467 0.077 0.390 Comtdl chr14 21847872 21848119 - iret 0.092 0.148 -0.057 Med15 chr16 17655677 17673397 - cass 0.887 0.779 0.108 Acini chr14 54653213 54653680 alt3 0320 0731 -0411 Osbpl6 chr2 76555013 76560259 cass 0.932 0.993 -0.061 Pde9a chr17 31459900 31461726 4- cass 0.912 0.990 -0.078 Zcchc12 chrX 36195985 36196577 + alt3 0.864 0.740 0.124 Fam120b chr17 15401744 15405934 + iret 0.962 0.721 0.241 Ccdc91 chr6 147475870 147507826 + cass 0.035 0.095 -0.060 Armcxl chrX 134718919 134719724 + alt5 0.769 0.691 0.079 Golga4 chr9 118572917 118577227 + cass 0033 0084 -0051 Pickl chr15 79247264 79248295 + iret 0.249 0.185 0.064 Alg8 chr7 97386801 97388587 + alt3 0.536 0.788 -0.252 2010107G23Rik chr1O 62109883 62111012 - cass 0.762 0.864 -0.102 Phykpl chr11 51593603 51594140 + iret 0.687 0.450 0.236 FralOad chr19 38207262 38214456 - cass 0.924 0.987 -0.063 Gfra2 chr14 70890129 70966315 + cass 0.905 0.968 -0.063 Mms19 chr19 41955418 41955960 - alt3 0025 0085 -0060 Nr1h2 chr7 44552034 44552619 alt3 0.743 0.656 0.087 Svil chr18 5059229 5062404 cass 0.440 0.187 0.253 Klh!32 chr4 24649539 24682269 - alt3 0.988 0.895 0.093
[0609]
[0610] Whsc1 chr5 33864659 33867657 + alt3 0.886 0.810 0.076Tra2b chr16 22247189 22248494 - alt3 0.707 0.827 -0.120 Arfgapl chr2 180977319 180979226 + cass 0.719 0.789 -0.069 Tnfsf13b chr8 10014168 10031421 + cass 0600 0953 -0353 Golga3 chr5 110184337 110185897 + alt5 0.326 0.522 -0.196 Sft2d1 chr17 8321799 8323334 + alt5 0.000 0.143 -0.143 Rmadl chr3 87924343 87925430 - iret 0.538 0.662 -0.124 Rnf215 chr11 4139738 4140077 + fret 0.219 0.149 0.070 Ankrd54 chr15 79057954 79061210 - alt3 0.964 0.853 0.111 Immt chr6 71851771 71857047 + taca 0.871 0.821 0.050 Zc3h13 chr14 75331705 75336073 + alt3 0640 0539 0.102 Kif21a chr15 90943764 90952817 taca 0.485 0.546 -0.061 Sema4c chr1 36555931 36557533 alt3 1.000 0.000 1.000 Pmp22 chr11 63128981 63133243 + alt3 0.671 0.724 -0.052 Ap3s1 chr18 46780660 46790622 + cass 0.588 0.678 -0.090 Cpsf3l chr4 155872849 155885195 + cass 0.899 0.975 -0.076 Sh2b3 chr5 121818464 121818853 - iret 0.149 0.251 -0.101 Wiz chr17 32378839 32381155 - alt5 0356 0053 0303 Luc7l chr17 26252909 26255124 + cass 0.110 0.177 -0.067 Slc4a9 chr18 36531099 36531525 + iret 0.000 1.000 -1.000 Nnat chr2 157560447 157562106 + cass 0.769 0.569 0.200 Navi chr1 135464705 135467788 - cass 0.907 0.828 0.079 Arfip2 chr7 105637120 105638326 cass 0.836 0.921 -0.085 Zfyve16 chr13 92508121 92509045 alt3 0.786 0.935 -0.149 AnkrdlO chr8 11619061 11628529 - cass 0302 0445 -0 143 Inip chr4 59775446 59783851 cass 0.807 0.931 -0.123 Mtg2 chr2 180071793 180078822 mutx 0.667 0.718 -0.051 Gak chr5 108613512 108623373 - taca 0.674 0.785 -0.111
[0611]
[0612] Pacrgl chr5 48374179 48374537 + alt3 0.001 0.114 -0.113Ift88 chr14 57434794 57437282 + alt3 0.075 0.149 -0.074 Ino80d chr1 63093284 63093729 alt3 0.800 0.970 -0.170 Adcy3 chr12 4206912 4208660 + alt3 0774 0678 0096 Cbx5 chr15 103215021 103215351 - alt3 0.333 0.241 0.093 Ogq1 chr6 113328364 113329407 + alt3 0.457 0.341 0.116 Erbb2ip chr13 103824719 103830330 - cass 0.750 0.831 -0.081 Acsm5 chr7 119534233 119534914 + alt5 0.100 0.667 -0.567 Mark2 chr19 7275395 7280011 - cass 0.937 0.877 0.060 Srpx2 chrX 133908425 133910730 + cass 0.599 0.286 0.313 Edc4 chr8 105887781 105888044 + iret 0097 0047 0050 Dcaf11 chr14 55561209 55563018 alt5 0.737 0.658 0.079 Dcafl 1 chr14 55561074 55563018 4- alt5 0.737 0.658 0.079 Slc44a2 chr9 21342742 21343046 + iret 0.091 0.167 -0.076 Srrm4 chr5 116453408 116467587 - cass 0.156 0.098 0.058 Ripk2 chr4 16129022 16132861 - cass 0.462 0.172 0.289 2700094K13Rik chr2 84669217 84670115 - alt5 0.772 0.846 -0.074 Ppip5k2 chr1 97740823 97741194 - alt3 0813 0.921 -0 108 Dcxr chr11 120726087 120726310 - iret 0.203 0.099 0.104 Htr4 chr18 62437382 62464641 + cass 0.000 0.400 -0.400 Enoxl chr14 77615438 77637799 + cass 0.113 0.251 -0.138 Tinf2 chr14 55680014 55680706 - iret 0.131 0.226 -0.096 Col11a1 chr3 114212072 114213285 + cass 0.080 0.000 0.080 Cacna2d4 chr6 119345043 119347250 + cass 0.129 0.440 -0.311 Iqce chr5 140669021 140670566 - alt3 0477 0848 -0370 Gale chr4 135965565 135966115 4- alt3 0.925 0.999 -0.075 Efccl chr6 87749131 87751904 cass 0.699 0.490 0.208 Cenpt chr8 105846983 105848898 - alt3 0.941 1.000 -0.059
[0613]
[0614] Ap3s1 chr18 46754370 46758113 + cass 0.069 0.144 -0.075Cog1 chr11 113659268 113661051 + alt3 0.956 0.901 0.055 Supv3l1 chr10 62430470 62432464 alt5 0.925 0.987 -0.063 Map3k5 chr10 20000565 20023800 + cass 0803 0962 -0 159 Miip chr4 147862896 147865269 - cass 0.423 0.289 0.133 Mtr chr13 12247892 12250690 cass 0.628 0.925 -0.297 Fam69c chr18 84720155 84730688 + cass 0.313 0.000 0.313 Acsl6 chr11 54314951 54315761 + cass 0.426 0.286 0.140 Repsi chr10 18104153 18114485 + cass 0.815 0.733 0.083 Chd2 chr7 73454328 73455640 - cass 0.072 0.123 -0.050 Ift57 chr16 49699347 49703267 + cass 0979 0926 0053 Mterfdl chr13 66928169 66930222 alt3 0.978 0.902 0.076 Pum1 chr4 130752575 130763020 4- cass 0.760 0.830 -0.070 Grina chr15 76246789 76247736 + alt3 0.380 0.220 0.160 Ccdc60 chr5 116225268 116288984 - alt3 1.000 0.060 0.940 Naa25 chr5 121426726 121430755 + cass 0.068 0.004 0.065 Ctc1 chr11 69031099 69031603 + alt3 0.668 0.523 0.144 Gnmt chr17 46726060 46726410 - iret 0440 0110 0330 Ablim2 chr5 35836947 35841400 + cass 0.468 0.535 -0.067 C2cd5 chr6 143017882 143020373 cass 0.925 0.774 0.150 Ccdc137 chr11 120458611 120460110 + alt3 0.999 0.941 0.058 Tctex1d2 chr16 32425240 32426920 + iret 0.109 0.190 -0.081 Trim30a chr7 104412214 104429350 - cass 0.238 0.810 -0.571 Arfgapl chr2 180967244 180971113 + cass 0.659 0.758 -0.099 Ptpn4 chr1 119765877 119783588 taca 0868 0953 -0086 Gigyfl chr5 137525171 137525558 4- alt3 1.000 0.947 0.053 Aifml chrX 48499771 48513397 mutx 0.634 0.739 -0.105 Sybu chr15 44746245 44748397 - alt3 0.885 0.966 -0.081
[0615]
[0616] Arnt chr3 95491022 95493859 + alt3 0.793 0.900 -0.107Pbrml chr14 31019148 31025649 + mutx 0.274 0.116 0.158 Casp2 chr6 42276694 42279925 + cass 0.292 0.484 -0.192 Shroom2 chrX 152623125 152657523 alt3 0859 0697 0.163 Ggtal chr2 35422179 35432631 cass 1.000 0.722 0.278 Zfp655 chr5 145235305 145238321 + cass 1.000 0.808 0.192 Ski! chr3 31113407 31117038 + alt5 0.597 0.395 0.202 Shq1 chr6 100632238 100637121 alt5 0.773 1.000 -0.227 Papd5 chr8 88247508 88250852 + cass 0.893 0.963 -0.069 Aars chr8 111033878 111037072 + cass 0.609 0.673 -0.064 Dus3l chr17 56768347 56768627 + iret 0408 0478 -0070 Gpsm3 chr17 34589805 34590940 cass 1.000 0.873 0.127 Ccser2 chr14 36874935 36896343 cass 0.878 0.767 0.111 Dock? chr4 98991363 99001201 - cass 0.810 1.000 -0.190 Asnsdl chr1 53346579 53348557 - alt5 0.831 0.907 -0.076 Stx3 chr19 11789556 11791844 - cass 0.521 0.695 -0.174 Lrrc16b chr14 55507684 55508263 + alt3 0.759 0.689 0.069 Gemin5 chr11 58122281 58125417 - alt3 0446 0613 -0 167 Wbp1 chr6 83120259 83120875 - iret 0.129 0.185 -0.056 BC003331 chr1 150388522 150390355 alt3 0.813 0.733 0.080 Ppip5k2 chr1 97755885 97759363 - cass 1.000 0.941 0.059 Fam 188a chr2 12386608 12405919 - taca 0.929 0.992 -0.063 Mettl25 chr10 105832850 105841377 cass 0.404 0.754 -0.351 Ip6k2 chr9 108796060 108797689 + cass 0.109 0.029 0.079 Ptcd3 chr6 71905097 71907845 - cass 0897 0962 -0065 Gsptl chr16 11238871 11240706 alt3 0.492 0.582 -0.090 Gsptl chr16 11239031 11240706 alt3 0.492 0.582 -0.090 Rbbp5 chr1 132497929 132505664 + alt5 0.212 0.316 -0.104
[0617]
[0618] Tec chr5 72773830 72782174 - cass 0.000 1.000 -1.000L3mbti1 chr2 162959515 162961081 + cass 0.760 0.914 -0.154 Zscan29 chr2 121165641 121170241 cass 0.167 0.320 -0.153 Tpd52l2 chr2 181508165 181511614 + cass 0804 0861 -0057 Slc25a25 chr2 32420304 32421389 - cass 0.844 0.792 0.051 Piekha5 chr6 140536647 140544186 + cass 0.716 0.854 -0.139 Agfgl chr1 82883180 82886186 + alt3 0.748 0.817 -0.069 Pdcd11 chr19 47119792 47120423 + fret 0.114 0.203 -0.089 Ccdc171 chr4 83604021 83635741 + cass 0.280 0.846 -0.566 Bratl chr5 140710137 140711652 + iret 0.876 0.467 0.409 Ccnd3 chr17 47505050 47578791 + cass 0.154 0000 0.154 Thnsl2 chr6 71141219 71144346 cass 0.200 0.000 0.200 Rbm39 chr2 156178558 156179239 iret 0.559 0.669 -0.110 Robo2 chr16 73928023 73933867 - cass 0.387 0.281 0.107 Pcsk6 chr7 66025227 66031850 + cass 0.678 0.543 0.136 2210018M11Rik chr7 98600671 98610864 - cass 0.744 0.844 -0.100 Fn1 chr1 71597310 71598462 - alt3 0.960 0.910 0.050 Phf21b chr15 84791353 84791940 - iret 0265 0470 -0205 Exocl chr5 76537697 76542237 + cass 0.911 0.978 -0.067 Ss18 chr18 14636495 14640360 cass 0.615 0.517 0.098 Fam 188a chr2 12403997 12405919 - alt3 0.982 0.913 0.069 Usp1 chr4 98923970 98928353 + cass 0.685 0.878 -0.192 Erall chr11 78074279 78074626 iret 0.292 0.214 0.078 Pxk chr14 8152080 8155344 + alt3 0.562 0.491 0.070 FbrsH chr5 110378037 110379113 - cass 0329 0202 0127 Atxn2 chr5 121811384 121814719 4- cass 0.529 0.479 0.050 Gtpbp2 chr17 46164193 46164832 iret 0.124 0.176 -0.052 Taz chrX 74288177 74288512 + iret 0.937 0.838 0.099
[0619]
[0620] Slc25a10 chr11 120497013 120497460 + iret 0.197 0.126 0.070Ube2e1 chr14 18330949 18331790 - alt3 1.000 0.837 0.163 Txlng chrX 162786616 162787677 iret 0.495 0.304 0.191 Nek1 chr8 61049783 61054621 + cass 0202 0316 -0 115 Asapl chr15 64152829 64159003 - alt5 0.385 0.335 0.050 Mtbp chr15 55571294 55572967 + alt3 0.000 0.357 -0.357 Tarbp2 chr15 102522848 102523673 + iret 0.486 0.626 -0.141 Ezh2 chr6 47540676 47542405 cass 0.183 0.323 -0.139 Mrohl chr15 76432148 76433616 + cass 0.917 0.968 -0.051 Hmgxb4 chr8 74993708 74999707 + cass 1.000 0.698 0.302 Hnrnph3 chr10 63017523 63018224 alt5 0677 0752 -0075 Qpctl chr7 19144659 19147150 cass 0.974 0.908 0.066 Dok1 chr6 83032766 83033429 alt3 0.667 1.000 -0.333 Dok1 chr6 83032804 83033429 - alt3 0.667 1.000 -0.333 Sema4d chr13 51725223 51748788 - cass 0.448 0.353 0.095 Wdr4 chr17 31503509 31509911 - alt5 0.933 0.839 0.094 3130024G19Rik chr7 70365382 70388458 + cass 1.000 0.231 0.769 Dcaf11 chr14 55560491 55561441 + alt3 0583 0900 -0317 Zfp12 chr5 143239953 143244318 + cass 0.674 0.811 -0.136 Nmi chr2 51960570 51973003 cass 0.700 0.250 0.450 Atf6 chr1 170787345 170788674 - alt3 0.929 0.824 0.104 Jmjd4 chr11 59453510 59454058 + alt3 0.440 0.321 0.119 Myef2 chr2 125095911 125098074 cass 0.805 0.751 0.054 Trim26 chr17 36837177 36851127 + cass 0.846 0.943 -0.097 Tcf3 chr10 80418765 80419598 - alt3 0120 0 186 -0066 Wdr17 chr8 54690003 54696396 cass 0.858 1.000 -0.142 Brox chr1 183292474 183294449 alt3 0.175 0.014 0.161 Ddx26b chrX 56493038 56496733 + cass 0.332 0.594 -0.262
[0621]
[0622] Medag chr5 149422133 149427505 + cass 0.913 0.782 0.131Taf6l chr19 8773504 8774452 - alt3 0.925 0.776 0.148 Nrxn2 chr19 6443630 6450598 + cass 0.703 0.793 -0.090 Stk35 chr2 129801403 129827986 + cass 0765 0649 0.116 Eps15l1 chr8 72340998 72358444 cass 0.864 0.924 -0.060 Ml!t1 chr17 56899768 56905854 - cass 0.387 0.314 0.072 Ildr2 chr1 166270498 166294736 + cass 0.838 0.914 -0.075 Smpd4 chr16 17625719 17626537 + alt3 0.020 0.500 -0.480 Nr1h2 chr7 44551954 44552619 - iret 0.671 0.580 0.091 Casp9 chr4 141793842 141796833 alt3 0.982 0.857 0.125 Casp9 chr4 141793842 141796690 + alt3 0982 0857 0 125 Rapgefl chr2 29679132 29686282 cass 0.422 0.477 -0.055 Pin chr1O 53337704 53344271 4- cass 0.000 0.067 -0.067 Fdxr chr11 115271876 115272281 - alt3 0.587 0.728 -0.141 Fdxr chr11 115271920 115272281 - alt3 0.587 0.728 -0.141 Cops4 chr5 100518442 100528689 + taca 0.772 0.844 -0.072 Mutyh chr4 116807733 116814435 + cass 0.000 0.217 -0.217 Uty chrY 1168089 1170183 - cass 0366 0730 -0364 Zfp384 chr6 125030797 125033317 + cass 0.668 0.565 0.103 Slc12a7 chr13 73809905 73813671 + cass 0.385 0.765 -0.381 Hax1 chr3 89997737 89998665 - alt3 0.727 0.937 -0.210 Acd chr8 105700006 105700357 - alt3 0.142 0.221 -0.079 Tjapl chr17 46260094 46261201 - cass 0.849 0.933 -0.084 Ank3 chr1O 69980271 69982196 + cass 0.423 0.476 -0.053 Fam 179b chr12 64965939 64976850 cass 0022 0088 -0066 MfsdIO chr5 34636592 34637145 - alt3 0299 0.045 0.254 Abcblb chr5 8806007 8812845 + cass 1.000 0.773 0.227 Tnntl chr7 4512264 4513612 - alt3 0.276 0.164 0.112
[0623]
[0624] Cecr2 chr6 120756548 120757690 + alt5 0.450 0.875 -0.425Zfp346 chr13 55113677 55122514 + cass 0.912 0.982 -0.069 Dars2 chr1 161046776 161051437 - cass 0.777 0.878 -0.102 Pax6 chr2 105692736 105697363 + cass 0934 1 000 -0066 1110007C09Rik chr13 49203717 49205344 cass 0.864 0.921 -0.057 Lmbrll chr15 98908519 98908920 - iret 0.035 0.109 -0.074 Cux1 chr5 136312642 136314402 - alt3 0.962 0.787 0.174 Cenpt chr8 105849186 105849520 iret 0.079 0.163 -0.083 Mlh1 chr9 111249219 111255715 - cass 1.000 0.906 0.094 Usp24 chr4 106370983 106372763 iret 0.103 0.035 0.068 Oprll chr2 181715689 181718731 + cass 0.251 0170 0082 Wbpll chr19 46599136 46644454 cass 0.480 0.362 0.118 Orc5 chr5 22526362 22526582 iret 0.064 0.117 -0.053 L3mbtl1 chr2 162966545 162967084 + iret 0.198 0.319 -0.121 Slc38a10 chr11 120106433 120109535 - alt3 0.762 0.868 -0.106 Zfp58 chr13 67494602 67500452 - cass 0.069 0.182 -0.113 BC053749 chr7 30549606 30552271 - cass 0.333 0.722 -0.389 Apitdl chr4 149132258 149137581 - alt3 0083 0029 0054 Fam149b chr14 20375509 20378522 + cass 0.604 0.726 -0.123 Kcnh6 chr11 106023724 106025877 + alt3 0.071 0.167 -0.095 Cadml chr9 47813772 47848297 + taca 0.299 0.363 -0.064 Atg2b chr12 105648988 105649631 - alt3 0.885 0.950 -0.065 Soatl chr1 156457945 156474239 - cass 0.855 1.000 -0.145 Akap13 chr7 75743964 75746809 + alt3 0.829 0.657 0.172 Trmt2a chr16 18252401 18253497 cass 0928 0991 -0063 Grikl chr16 87950027 87957596 - cass 0456 0.317 0.138 Nckap5 chr1 125913619 125981696 - cass 1.000 0.692 0.308 MrpslO chr17 47372426 47375108 + cass 0.842 0.752 0.090
[0625]
[0626] Vps54 chr11 21263198 21264870 + alt3 0.343 0.500 -0.157Stxbp5l chr16 37139843 37174361 - taca 0.836 0.655 0.182 Timeless chr10 128249967 128250461 alt3 0.889 0.589 0.300 Dlq3 chrX 100767726 100771664 cass 0260 0312 -0052 Exod chr5 76554099 76559166 cass 0.702 0.550 0.153 Cxcl16 chr11 70458744 70459166 iret 0.026 0.220 -0.194 Capn3 chr2 120502386 120502592 iret 0.077 0.228 -0.150 9030617O03Rik chr12 100779094 100829724 cass 0.000 0.115 -0.115 RapgefS chr11 54691237 54699285 cass 0.212 0.358 -0.146 Trpm3 chr19 22732999 22766807 cass 0.550 0.473 0.078 Unc13b chr4 43115102 43165977 cass 0532 0250 0282 Pdlim5 chr3 142304704 142312213 cass 0.719 0.446 0.273 Mynn chr3 30603503 30607829 alt5 0.675 0.879 -0.204 Rrp36 chr17 46670110 46674286 taca 0.702 0.816 -0.114 Smarca2 chr19 26749851 26752040 cass 0.094 0.175 -0.081 Plekha5 chr6 140552714 140569414 taca 0.840 0.673 0.166 Porcn chrX 8201443 8203288 mutx 0.606 0.518 0.088 Auts2 chr5 131439297 131445481 mutx 0229 0301 -0072 Ngb chr12 87097530 87100114 alt5 0.155 0.098 0.057 Hnrnpk chr13 58396826 58399211 cass 0.876 0.802 0.074 Rpi12 chr2 32962984 32963852 cass 0.932 0.992 -0.060 Pex2 chr3 5560499 5563271 cass 0.510 0.613 -0.103 Fance chr17 28316858 28317610 cass 0.500 0.333 0.167 Gphn chr12 78412340 78454850 cass 0.297 0.210 0.086 Usp37 chr1 74441531 74450554 cass 0907 1 000 -0093 QricM chr9 108517717 108528917 cass 0955 0.853 0.102 Sh2b2 chr5 136224154 136224495 alt3 0.982 0.868 0.114 Proml chr5 44000776 44001900 cass 0.924 1.000 -0.076
[0627]
[0628] Cacnala chr8 84638611 84640248 alt3 0.978 0.896 0.082Flna chrX 74230505 74233337 - cass 0.568 0.652 -0.085 Hax1 chr3 89997386 89998002 - iret 0.736 0.554 0.182 Chd8 chr14 52212562 52213022 - alt3 0320 0264 0056 lfi203 chr1 173936517 173942293 cass 0.500 1.000 -0.500 Pixna3 chrX 74335769 74336190 + iret 0.024 0.112 -0.088 Faxc chr4 21948691 21982501 + cass 0.760 0.869 -0.109 Gria2 chr3 80690403 80692531 cass 0.411 0.471 -0.060 Ssh3 chr19 4267746 4268581 - iret 0.174 0.270 -0.096 Kbtbd3 chr9 4309898 4313805 alt3 0.000 0.107 -0.107 Slx4ip chr2 137000183 137044076 + cass 0737 0417 0320 Slc22a5 chr11 53876008 53891677 cass 0.828 0.938 -0.110 Pou2f1 chr1 165931661 166002633 taca 0.565 0.354 0.211 Nuak2 chr1 132324952 132327847 + cass 0.097 0.429 -0.332 Lca5 chr9 83426592 83441098 - cass 0.210 0.538 -0.328 Ccdc33 chr9 58033219 58033714 - alt3 0.308 0.857 -0.549 Sema6c chr3 95171557 95173652 + cass 0.215 0.130 0.085 Lrrfipl chr1 91107297 91112296 + taca 0869 0930 -0061 Ncaph2 chr15 89370394 89370642 + alt3 0.227 0.176 0.051 Zfp821 chr8 109717761 109721342 + alt5 0.485 0.597 -0.112 Fopnl chr16 14311046 14317332 - cass 0.914 1.000 -0.086 Serac 1 chr17 6067528 6070888 - cass 0.937 0.697 0.240 Ermard chr17 15059343 15061181 cass 0.957 0.861 0.096 Rgs12 chr5 35020313 35021258 + alt3 0.999 0.946 0.053 Kankl chr19 25422918 25425983 cass 0538 0605 -0067 Zkscan17 chr11 59502919 59503809 - alt5 1 000 0.938 0.062 0rmdl2 chr10 128820260 128821597 - alt3 0.557 0.741 -0.184 2310022A10Rik chr7 27571563 27574739 + alt5 0.006 0.062 -0.056
[0629]
[0630] Gnbll chr16 18499045 18548190 + cass 0.742 1.000 -0.2581700086006Rik chr18 38238404 38250197 taca 0.500 0.214 0.286 Eps8 chr6 137539321 137591491 alt3 0.868 0.968 -0.100 Tmem191c chr16 17277668 17277890 iret 0327 0390 -0064 Arhgap4 chrX 73906668 73911270 taca 0.108 0.250 -0.142 Gripl chr10 119819511 119930035 cass 0.895 0.630 0.265 Lyplal chr1 4886743 4889609 alt3 0.780 0.703 0.077 Lyplal chr1 4886743 4889601 alt3 0.780 0.703 0.077 Lyplal chr1 4886743 4889559 alt3 0.780 0.703 0.077 Lyplal chr1 4886743 4889508 alt3 0.780 0.703 0.077 Ankrd27 chr7 35620542 35622348 alt3 0907 0964 -0057 Eva 1c chr16 90830858 90876190 taca 0.591 1.000 -0.409 Setdbl chr3 95339906 95340311 alt3 0.841 0.911 -0.070 Sfswap chr5 129543213 129549683 alt5 0.949 0.881 0.068 1700066M21 Rik chr1 57377643 57380144 cass 0.067 0.579 -0.512 Mettl23 chr11 116843556 116845944 alt3 0.991 0.500 0.491 Kdm5c chrX 152237632 152240195 cass 0.976 0.912 0.064 Tcf7 chr11 52257683 52260620 cass 0210 0077 0133 Pgap2 chr7 102235661 102236345 alt5 0.193 0.740 -0.547 Itgb4 chr11 115997949 115999954 alt5 0.876 0.939 -0.062 Fhodl chr8 105337317 105337768 alt3 0.429 0.000 0.429 Acads chr5 115111856 115112390 iret 0.073 0.167 -0.094 Ocrl chrX 47948124 47960568 cass 0.765 0.685 0.081 Gtf2ird2 chr5 134191179 134192790 alt5 0.182 0.304 -0.123 Zbtb20 chr16 43569680 43577192 cass 0446 0261 0186 Kcntl chr2 25909204 25909676 iret 0234 0.294 -0.060 Usp47 chr7 112077791 112082560 cass 0.358 0.451 -0.092 Zkscan2 chr7 123484939 123490261 cass 0.292 0.122 0.170
[0631]
[0632] Dph6 chr2 114519711 114535585 cass 0.556 0.725 -0.170Nsun5 chr5 135374926 135375484 + alt5 0.383 0.535 -0.152 Rft1 chr14 30676849 30677816 + alt5 0.241 0.148 0.093 St18 chr1 6730050 6752367 + cass 0000 0750 -0750 Stagl chr9 100643622 100705253 + cass 0.022 0.081 -0.059 Ddx41 chr13 55535729 55536044 - iret 0.098 0.159 -0.061 Ddx41 chr13 55535762 55536044 - iret 0.098 0.159 -0.061 Clasrp chr7 19603189 19604460 alt3 0.905 0.986 -0.081 Cdc14b chr13 64196642 64205416 - taca 1.000 0.804 0.196 Zdhhc12 chr2 30091666 30092064 - alt3 1.000 0.924 0.076 Nfatc2ip chr7 126382853 126390571 - taca 1 000 0835 0.165 Cdanl chr2 120825241 120850438 cass 0.333 0.157 0.177 Bicdl chr6 149518905 149556900 4- cass 0.794 0.880 -0.086 Ulk3 chr9 57593750 57594045 + iret 0.114 0.204 -0.090 Kif17 chr4 138254272 138255697 + cass 0.941 0.882 0.059 5530601H04Rik chrX 105044001 105066876 - cass 0.641 0.769 -0.127 Sh2d3c chr2 32737518 32744878 + cass 0.667 0.463 0.204 Cyba chr8 122426185 122427302 - cass 0988 0.916 0072 Fbxl6 chr15 76537089 76537419 - iret 0.078 0.134 -0.055 Rars2 chr4 34623410 34630569 + cass 0.965 0.889 0.076 Sitm chr9 70559036 70572175 + cass 0.934 0.986 -0.052 Izumo4 chr10 80704417 80704712 + alt3 0.018 0.118 -0.101 Fam 195a chr17 25864593 25868542 - cass 0.249 0.161 0.088 Plekha6 chr1 133273857 133280406 + cass 0.846 0.920 -0.074 2410004N09Rik chr18 33794891 33795988 cass 0962 0906 0056 Ccdc64 chr5 115648174 115651938 - cass 0914 0.855 0.059 Erp44 chr4 48219342 48279451 - taca 1.000 0.896 0.104 D430042009Rik chr7 125707922 125753048 + cass 0.467 0.805 -0.338
[0633]
[0634] Htra2 chrtS 83052697 83053069 - iret 0.286 0.214 0.072Polr3gl chr3 96579797 96580088 - iret 0.349 0.258 0.091 Ppip5k2 chr1 97719833 97723833 cass 0.392 0.282 0.110 Hpca chr4 129118363 129121666 cass 0169 0239 -0070 RteH chr2 181354352 181355551 cass 0.765 0.452 0.313 Ube2q2 chr9 55162967 55176246 cass 0.780 0.716 0.064 D930015E06Rik chr3 83900326 83901465 alt3 1.000 0.860 0.140 Mms19 chr19 41962949 41963469 alt3 0.932 1.000 -0.068 Smpd4 chr16 17625719 17626018 iret 0.057 0.121 -0.065 Slx4ip chr2 137043999 137066229 taca 1.000 0.647 0.353 Prrxl chr1 163248255 163257941 cass 0420 0645 -0225 Epb4 chr10 25495444 25501630 cass 0.816 0.665 0.151 Nktr chr9 121741597 121742786 alt3 0.626 0.693 -0.067 Otud5 chrX 7873839 7875283 alt3 0.254 0.204 0.050 Otud5 chrX 7873839 7874861 alt3 0.254 0.204 0.050 Dpp7 chr2 25353152 25353515 iret 0.227 0.304 -0.077 Rabepk chr2 34785566 34790654 alt3 0.868 0.788 0.080 Fam126b chr1 58557976 58565953 alt5 0035 0098 -0063 Ythdcl chr5 86804536 86815729 cass 0.506 0.284 0.222 Tbc1d19 chr5 53830454 53833044 cass 0.267 0.393 -0.126 Ascc2 chr11 4656268 4664302 taca 1.000 0.941 0.059 Nae1 chr8 104527082 104528216 cass 0.955 0.880 0.075 Sltm chr9 70572148 70574624 cass 0.943 0.880 0.063 Mthfd2l chr5 90974322 91021367 cass 0.830 0.926 -0.096 Tle1 chr4 72158212 72169195 alt5 0259 0 100 0159 Aasdh chr5 76904208 76905451 cass 0846 0.481 0.365 Sorbs 1 chr19 40318021 40324831 cass 0.589 0.518 0.072 Pla2g3 chr11 3491891 3492241 iret 0.138 0.269 -0.132
[0635]
[0636] Nemf chr12 69340977 69341480 alt3 0.963 0.904 0.059Nemf chr12 69340994 69341480 - alt3 0.963 0.904 0.059 Znf512b chr2 181589372 181590180 - iret 0.872 0.778 0.095 Med7 chr11 46436970 46442720 + cass 0671 0837 -0166 Rpain chr11 70973014 70974131 + cass 0.846 0.926 -0.080 Focad chr4 88185879 88229426 + cass 0.879 0.966 -0.086 Coro6 chr11 77462624 77464109 + cass 0.120 0.030 0.090 Tfdp2 chr9 96287600 96295126 + cass 0.619 0.752 -0.133 8430427H17Rik chr2 153417959 153420808 - alt5 0.710 0.590 0.120 Itsn2 chr12 4639654 4650118 cass 0.405 0.497 -0.092 Dpf1 chr7 29313081 29314401 + cass 0788 0694 0094 Chd6 chr2 160969338 160970205 alt3 0.982 0.932 0.050 Mtssl chr15 58941233 58945523 alt5 0.275 0.193 0.082 Traf7 chr17 24516503 24518798 - alt3 0.296 0.415 -0.119 Traf7 chr17 24516504 24518798 - alt3 0.296 0.415 -0.119 Traf7 chr17 24516510 24518798 - alt3 0.296 0.415 -0.119 Traf7 chr17 24516526 24518798 - alt3 0.296 0.415 -0.119 Traf7 chr17 24516550 24518798 - alt3 0296 0415 -0119 Traf7 chr17 24516562 24518798 - alt3 0.296 0.415 -0.119 Pisd chr5 32764772 32785625 - cass 0.894 0.970 -0.076 Toplmt chr15 75669286 75670183 - alt3 0.968 0.884 0.084 Dpp7 chr2 25355813 25356141 - iret 0.054 0.110 -0.057 Ecscr chr18 35713087 35715214 - cass 0.890 1.000 -0.110 Gpr137b chr13 13359125 13361421 cass 0.168 0.103 0.065 Mars chr10 127296540 127296943 cass 0596 0462 0135 Zcchc2 chr1 106023663 106027530 + alt3 0 104 0.200 -0.096 Trub2 chr2 29776110 29779894 - cass 0.974 0.912 0.062 Art13b chr16 62827163 62846997 - cass 0.999 0.843 0.156
[0637]
[0638] Smtn chr11 3517526 3521967 - cass 0.983 0.906 0.077Zmym3 chrX 101416908 101417329 alt3 0.463 0.538 -0.075 Zfp62 chr11 49214222 49215492 alt5 0.889 0.981 -0.093 Exoc7 chr11 116295543 116300432 cass 0773 0698 0075 Crtc2 chr3 90262450 90263356 cass 0.474 0.346 0.128 Map3k12 chr15 102509289 102510004 alt3 0.130 0.077 0.053 Myo19 chr11 84892075 84894651 cass 0.440 0.720 -0.280 Ntmtl chr2 30807967 30819831 mutx 0.300 0.214 0.086 Cdkl3 chr11 52033513 52084494 cass 0.167 0.500 -0.333 Hdac7 chr15 97798223 97802124 mutx 0.147 0.210 -0.062 Porcn chrX 8201443 8203042 cass 0896 0956 -0060 Ccnk chr12 108179801 108186621 cass 0.960 0.903 0.057 Mpv17 chr5 31144699 31145775 mutx 0.605 0.681 -0.076 Zfp788 chr7 41633530 41647616 taca 0.312 0.425 -0.113 Cflar chr1 58711507 58713658 alt3 0.000 0.111 -0.111 Fam126a chr5 23965255 23979646 cass 0.129 0.039 0.090 Mybptf chr10 88518281 88523152 cass 0.250 0.802 -0.552 Ccdc103 chr11 102883073 102884522 alt3 0154 0063 0091 D11Wsu47e chr11 113687773 113692501 cass 0.980 0.863 0.117 Brcc3 chrX 75449968 75455700 alt5 0.714 0.881 -0.167 Sorbsl chr19 40364993 40373616 cass 0.501 0.564 -0.063 Fam229b chr10 39122176 39132377 alt3 0.765 0.865 -0.100 Cep72 chr13 74037631 74040161 iret 0.462 0.191 0.270 Kctd20 chr17 28952715 28961686 cass 0.278 0.440 -0.162 Map3k4 chr17 12239969 12243595 cass 0947 0892 0055 Mdm2 chr10 117705154 117710027 cass 0725 0.605 0.120 Acyl chr9 106434964 106435168 iret 0.071 0.130 -0.059 Atp9b chr18 80738652 80739803 alt3 0.967 0.917 0.050
[0639]
[0640] Atp2b4 chr1 133702673 133711821 alt5 0.387 0.273 0.114Zfp688 chr7 127419003 127421517 cass 0.403 0.557 -0.154 Lyplall chr1 186089443 186114362 cass 0.895 0.765 0.130 Abca7 chr10 80007135 80007450 alt3 0755 0896 -0141 Pbrml chr14 31107108 31114003 cass 0.470 0.335 0.135 Sp100 chr1 85679058 85692048 cass 1.000 0.591 0.409 Cggbp! chr16 64852798 64855895 alt3 0.655 0.724 -0.070 Cggbp-I chr16 64852798 64855872 alt3 0.655 0.724 -0.070 Cggbpl chr16 64852798 64855632 alt3 0.655 0.724 -0.070 Pnpti chr11 29148277 29153337 alt3 0.909 0.990 -0.081 Myh14 chr7 44634358 44637878 cass 0411 0535 -0 124 Pds5a chr5 65627998 65630054 alt3 0.917 0.978 -0.061 Lairl chr7 4028706 4055952 alt5 0.643 0.397 0.245 Slc35f5 chr1 125587364 125589964 alt3 0.890 0.958 -0.068 Trp53inp1 chr4 11165089 11174376 cass 0.350 0.459 -0.110 Zscan21 chr5 138116951 138125689 taca 1.000 0.708 0.292 Sic7a3 chrX 101083870 101085371 cass 0.176 0.248 -0.073 Gyk chrX 85737313 85740360 cass 0745 0623 0123 Camk4 chr18 32939171 33107941 cass 0.866 0.948 -0.082 Ankzfl chr1 75195800 75196380 iret 0.449 0.311 0.138 Prrgl chrX 78449612 78483910 cass 0.141 0.294 -0.153 N28178 chr4 42917250 42933666 cass 0.934 0.866 0.068 Lrp6 chr6 134520395 134566964 cass 0.901 0.986 -0.085 Slc44a2 chr9 21352466 21355027 cass 0.672 0.600 0.072 Cep164 chr9 45823645 45828584 cass 0128 0000 0128 Dnm2 chr9 21505463 21507145 alt3 0642 0.565 0.077 Dnm2 chr9 21505463 21506688 alt3 0.642 0.565 0.077 Dnm2 chr9 21505463 21506389 alt3 0.642 0.565 0.077
[0641]
[0642] Sh3pxd2a chr19 47314034 47343467 cass 0.854 0.690 0.164Atxn3 chr12 101948003 101948438 - alt3 0.677 0.876 -0.199 Cd97 chr8 83734295 83741182 - alt3 0.000 0.114 -0.114 Dec chr18 71378654 71384222 - alt3 0935 0996 -0061 Srr chr11 74912961 74925676 cass 0.841 0.896 -0.055 Miefl chr15 80234079 80236152 + cass 0.143 0.267 -0.124 Zfp945 chr17 22861505 22865335 - cass 0.481 0.279 0.202 Nup214 chr2 31989072 31991424 + cass 0.919 0.838 0.080 Rbms3 chr9 116636368 116681226 - cass 0.116 0.174 -0.059 6330408A02Rik chr7 13258966 13269646 - alt5 0.837 0.957 -0.120 Cd33 chr7 43527455 43529931 - cass 0231 0074 0 157 Ccdc74a chr16 17648066 17650073 taca 0.767 0.591 0.175 Adamts6 chr13 104399942 104427067 4- cass 0.238 0.429 -0.190 Eda chrX 100395018 100400759 + cass 0.048 0.385 -0.337 Trpm3 chr19 22897669 22901334 + cass 0.828 0.886 -0.058 Fcld2 chr17 29360931 29363941 + alt3 0.300 0.002 0.298 Zdhhc3 chr9 123089031 123091131 - alt5 0.333 0.002 0.331 Repsi chr10 18104153 18107747 + alt3 0796 0727 0069 Pter chr2 12924040 12978616 + cass 0.529 0.852 -0.322 Agrn chr4 156167281 156168568 - alt5 0.683 0.597 0.086 Maltl chr18 65448921 65451572 + cass 0.077 0.000 0.077 Dnajc24 chr2 105966709 105981119 - cass 0.909 0.967 -0.058 Lrrfipl chr1 91079030 91085004 cass 0.324 0.444 -0.121 Acapl chr11 69881566 69882030 - iret 1.000 0.111 0.889 Lilrb4 chr10 51493169 51494177 iret 0000 0600 -0600 1112a chr3 68695189 68695345 + iret 0667 0.000 0.667 Tmem232 chr17 65486471 65517236 - cass 1.000 0.200 0.800 Mtif2 chr11 29526407 29530153 + cass 0.918 1.000 -0.082
[0643]
[0644] Phf23 chr11 69997745 69999131 + cass 0.629 0.729 -0.100Usp43 chr11 67876372 67880143 - alt5 1.000 0.774 0.226 Zfp182 chrX 21060502 21062021 - alt3 0.250 1.000 -0.750 Baz2b chr2 59978542 59983969 - alt3 0510 0718 -0207 Ambral chr2 91772290 91810208 + cass 0.725 0.660 0.065 Ube2v2 chr16 15581058 15594500 - cass 0.000 0.086 -0.086 Agapl chr1 89743727 89789323 + cass 0.716 0.654 0.062 Plekha5 chr6 140556803 140569414 + taca 0.880 0.977 -0.097 Pld5 chr1 176044834 176074524 - alt5 0.958 0.865 0.093 Slc50a1 chr3 89269826 89270138 - iret 0.344 0.285 0.059 Cdk11b chr4 155625530 155626888 + alt5 0399 0464 -0065 Myolh chr5 114361042 114361323 iret 0.231 0.000 0.231 Hacel chr10 45618497 45648686 4- cass 0.913 0.974 -0.061 Gtf2ird1 chr5 134363901 134380019 - taca 0.921 0.980 -0.059 Usp37 chr1 74456078 74461729 - cass 0.584 0.726 -0.142 Stxbp5l chr16 37142275 37174361 - cass 0.284 0.159 0.126 Ybx3 chr6 131370322 131379479 - cass 0.431 0.355 0.076 Arhgap12 chr18 6069830 6135930 - cass 0133 0024 0109 Upf1 chr8 70339742 70340098 - alt5 0.194 0.278 -0.084 Pusl1 chr4 155889433 155889722 - iret 0.061 0.111 -0.050 Bak1 chr17 27021166 27022581 - cass 0.414 0.296 0.118
[0645]
[0646] Rapgef3 chr15 97757682 97758104 - alt3 0.081 0.027 0.054
[0251] TABLE 5
[0647] Classification in transcripts belonging to behavioral gene ontology categories.
[0648] Detailed classification of 27 transcripts with differential AS events in the midbrain of P21 Novalhu / hu mice. The 27 transcripts that were classified into behavioral categories in the gene ontology were divided according to the minor classification to which each transcript belongs. The major gene ontology term. ID, minor classification, and category are shown.
[0649] Gene_name Name Minor Classification GO term ID Category ATP-binding cassette,
[0650] memory learning or memory G0:0007611 Learning Abca7 sub-family A member
[0651] 7 visual learning learning or memory GG:0007611 Learning Adcy3 adenylate cyclase 3 olfactory learning learning or memory G0:0007611 Learning Atxn3 ataxin 3 exploration behavior exploration behavior GG:0035640 Exploration autism susceptibility innate vocalization vocalization
[0652] Auts2 candidate 2 behavior behavior GG:0071625 Vocalization calcium channel, adult walking
[0653] voltage-dependent, behavior locomotory behavior GG:0007626 Locomotor Cacnala
[0654] P / Qtype, alpha 1A behavioral response
[0655] subunit to pain Other calcium / calmodulin- dependent protein
[0656] Camk4 kinase IV long-term memory learning or memory G0:0007611 Learning chromodomain
[0657] social behavior social behavior G0:0035176 Sociability Chd8 helicase DNA binding
[0658] protein 8 long-term memory learning or memory G0:0007611 Learning exploration behavior exploration behavior G0:0035640 Exploration Chrd chordin
[0659] visual learning learning or memory G0:0007611 Learning CLN3
[0660] associative learning learning or memory G0:0007611 Learning lysosomal / endosomal
[0661] Cln3
[0662] transmembrane
[0663] protein, battenin learning or memory learning or memory GG:0007611 Learning dishevelled segment
[0664] Dvl1 polarity protein 1 social behavior social behavior G0:0035176 Sociability behavioral response
[0665] epidermal growth
[0666] to ethanol Other
[0667] Eps8 factor receptor
[0668] adult locomotory
[0669] pathway substrate 8
[0670] behavior locomotory behavior G0:0007626 Locomotor guanine nucleotide
[0671] binding protein (G
[0672] protein), beta
[0673] Gnbll polypeptide 1-like social behavior social behavior GG:0035176 Sociability behavioral response
[0674] glutamate receptor,
[0675] Grikl to pain Other ionotropic, kainate 1
[0676] adult behavior Other adult locomotory
[0677] HtrA serine peptidase behavior locomotory behavior G0:0007626 Locomotor Htra2
[0678] 2 adult walking
[0679] behavior locomotory behavior G0:0007626 Locomotor intraflagellar transport regulation of feeding
[0680] Ift88 88 behavior feeding behavior G0:0007631 Feeding mannosidase 2, alpha
[0681] Man2b1 B1 learning or memory learning or memory G0:0007611 Learning myosin, heavy vocalization vocalization
[0682]
[0683] Myh14 polypeptide 14 behavior behavior G0:0071625 Vocalizationadult behavior Other vocalization vocalization
[0684] Nrxn2 neurexin II behavior behavior G0:0071625 Vocalization social behavior social behavior G0:0035176 Sociability vocal learning learning or memory G0:0007611 Learning eating behavior feeding behavior G0:0007631 Feeding regulation of
[0685] locomotor rhythm locomotory behavior G0:0007626 Locomotor Opril opioid receptor-like 1
[0686] conditioned place
[0687] preference Other behavior Other learned vocalization
[0688] behavior or vocal
[0689] Pax6 paired box 6 learning learning or memory G0:0007611 Learning adult locomotory
[0690] behavior locomotory behavior G0:0007626 Locomotor peripheral myelin
[0691] Pmp22
[0692] protein 22 motor behavior motor behavior G0:0061744 Motor adult walking
[0693] behavior locomotory behavior G0:0007626 Locomotor pumilio RNA-binding adult locomotory
[0694] Pum1 family member 1 behavior locomotory behavior G0:0007626 Locomotor Rap guanine
[0695] nucleotide exchange
[0696] Rapgef3 factor (GE F) 3 associative learning learning or memory G0:0007611 Learning solute carrier family 22
[0697] (organic cation
[0698] Slc22a5 transporter), member 5 locomotory behavior locomotory behavior G0:0007626 Locomotor sushi-repeat- containing protein, X- vocalization vocalization
[0699] Srpx2 linked 2 behavior behavior G0:0071625 Vocalization VPS54 GARP complex
[0700] Vps54 subunit motor behavior motor behavior G0:0061744 Motor locomotory behavior locomotory behavior G0:0007626 Locomotor Wdr47 WD repeat domain 47 motor behavior motor behavior G0:0061744 Motor
[0701] adult locomotory
[0702]
[0703] behavior locomotory behavior G0:0007626 Locomotor
[0704]
[0252] TABLE 6
[0705] Syllables detected in isolation induced USV test in pups.
[0706] The acoustic waveform data for each pup was processed by the Mouse Song Analyzer (from Erich lands lab) to obtain values for each syllable: syllable type, duration (time per syllable [sec]), ISI (intersyllable interval), fqVariance (degree of variance), purity, amplitude (magnitude of loudness), bandwidth (width of peak frequency) and peak frequency (Fq) variabilities; fqmin (minimum), fqmean (mean), fqmax (maximum), fqstart (start), fqend (end). notIDd: not identified, (partial exemplary data shown)u 1
[0707] $
[0708] R
[0709] g $ I
[0710] $
[0711] 1.
[0712]
[0713] I i § $
[0714] I
[0715] & R I R $ R I
[0716] 8 $ § £ I I s £ $ 3 i
[0717] 3
[0718] £ S £i I
[0719] «i 5 w s, } j J X J J J J J I J j I t i t t $ S. u t § g 'S i'1$ s g | I s i £ ■ $
[0720]
[0721] au
[0722] c
[0723] & » I» i w &
[0724] 8 & & b'• $
[0725] $ g % x y <?> >»
[0726] I i g % 1 u>
[0727] Ci I
[0728] ►ft
[0729] I M I ►»
[0730] g 5 § 3$ ►5 <2 i
[0731]
[0732] s>K*
[0733] I 9 i I 3 3 f f
[0734] P-* § § 1 8 $
[0735] o
[0736] $ 8 gj g >« f:i 8 $ 8 8 8 § 8 g g 8 t 0 I b!
[0737] & a § § S i;< § I K S?: ® 8 & £• $ & I I I hS n
[0738]
[0739] I I I % % $ I I £ 8 t?l ■# u< i ik'> W i g > (p M5 r:?.2< I
[0740] 8 8 g & 2? $ 8
[0741] $ K 8 ’•> 8 g M 8 g O 8 g M S34% I a •» §
[0742] i
[0743] I • oTJ § H H;§ i» iif & # fey & $ <> & H & is- I a $
[0744] ys y? wj i,<? « J J I > 5•• 5 I } 5 5 5 5 *» 5 5 I 5 I <5 >3 i g s i & i. r ■s 1 $ i i f I
[0745]
[0746] Si $ § <*• v' •t>
[0253] TABLE 7
[0747] USV features in each pup.
[0748] In isolation induced USV test for pup, the following USV features are calculated for each pup: number of USV s, call rate, percent.starting (percent of starting syllable type of sequence (continuous syllables)), percent. composition (percent of syllable), sequence length. For each syllable type (“s”, “u”, “d”, “m”), following parameter are calculated: Bw (bandwidth), Amp (amplitude), fqVar (Fq variance), Purity, Dur (duration), and Fq variabilities; fqMin (minimum), fqMean (mean), fqMax (maximum), fqStart (start), fqEnd (end), (partial exemplary data shown)
[0749]
[0750]
[0254] TABLE 8
[0751] USV features in each genotype of pup.USV features in Novalhu / hu, Novalhu / wt and Novalwt / wt pups. Bw (bandwidth), Amp (amplitude), fqVar (Fq variance), Dur (duration). The values represent the mean value for each genotype. The standard error (se) values for each parameter are inserted in the adjacent mean columns. The
[0752] parameters statistically different from control (Novalwt / wf) are indicated with asterisk (*p < 0.05).
[0753] / ’-values were calculated by Wilcoxon rank sum test and were corrected with Bonferroni method.
[0754] Novalwt / wt pup N=40, Novalhu / wt pup N=23, Novalhu / hu pup N=41.
[0755] ft ft ft ft
[0756] IsqsMMA IJs?) 41L« UU tw| 135
[0757] y / r?.?■) w ■3g) frb as? LS3| TP
[0758]
[0759] 1^1^ r= J3h IP OS.: a® f:
[0760] L’i. i)L
[0761] ft ft ft: ft ft st jMg ft pSUlSiijii id w| aiK (>&• V.15? PT. §5.5g3| PTA )&$ JaPs TIP Z Z w Ui? 1. P5 1 Wife) os JSJ) Oq ftllK t® iw ‘?> i O sU3| 13» 05
[0762]
[0763] 8 d-hfe S= aW ft nW ft: ft st l&h ft a=i?® ws ^.10 W 0.00?®'4s| ill® w 32^335 Will liisL a® £02 aw <32^3^ s?a2®| HO;
[0764]
[0765] w aa®ai *wl?2B?a
[0766] ifssti® ft n. Dft ft ft a aaggs aa«2?) saw) ITS] m a?g® £$Hi) 83® G® asis? £$ii?) oit ws » awi
[0767]
[0768] pgni; W 8W3 WS a*60i| a® PB SW?|
[0769]
[0255] TABLE 9
[0770] Bimodal distribution parameters at Fqmax (maximum frequency) in pup USVs
[0771] For each syllable type, the bimodality in Fqmax was assessed by Ashman’s D test (Fig. 4c). Two Gaussians were fitted to calculate each distribution parameters: component (weight), Fq mean
[0772] [Hz], Sd (standard deviation) and cutoff value (intersection point) [kHz]. The distributions are
[0773] classified as Low Fqmax and High Fqmax by the cutoff.
[0774] Bimudal distribution analysis and twO Gaussians f at maximum frequency of Pup USV
[0775] Syllable type Fqmax ^Component (peak weight) Mean iSd ^Ashman’s D score leutoff (im$tsecti®t)tkHz] Low 0,56 57275,73 I 9345'lli ZZ i z
[0776] 2,57098 81,0
[0777] High 0.44 i 92589.37 8772.521
[0778] Low 0.30 ZZZ'
[0779] High 0.7 nsnS-U758i82.5
[0780] 0 99013.49
[0781] Low 0.77 87754.79
[0782] 97,5
[0783] High 0.23 100146.58 4258.26) _
[0784] Lew 0.62 S7205.49 7632131!g.g76gg398.5
[0785]
[0786] 0.38 101813.27 3767.82)>
[0256] TABLE 10
[0787] Proportion of high / Iow Fqmax in pup-USVs in each genotype.
[0788] The ratio of low Fqmax and high Fqmax in each syllable type are shown for each genotype. The values represent the mean value for each genotype. The standard deviation (sd) and the standard error (se) values for each parameter are inserted in the adjacent mean columns. The parameters statistically different from control (Novalwt / wt) are indicated with asterisk (* / ><0.05, ** / ;<0.01). rvalues were calculated by Wilcoxon rank sum test and were corrected with Bonferroni method.
[0789] Novalwt / wt pup N=40, Novalhu / wt pup N=23. Novalhu / hu pup N=41.
[0790] (*) p=0.039; (**) p=0.0069
[0791] genotype ratio, lowFqmax sd se rati«xhighFc|max sd se Novel (wt / wt) 0.361 0.2S9 0,042 0.635 0.259 0.042 Novel (hij / wt) d 0.272 0.236 0.050 0.728 0.236 0.050 Novel (hu / hu) 0,254 0.269 0.044 0.745 0.268 0,044 Novel (wt / wt) 0.676 6.292 0.049 6.324 6.292 0.049 Novel (hu / wt) m *0.498 0.301 0.067 *0.502 0301 0.067 Novel (hu / hu) **&407 0360 0.067 **0.593 0.360 0,067 Novel (wt / wt) 0.553 0157 0,025 0.432 0.158 0.025 Novel (hu / wt) s 0.522 0.145 0.030 0.467 0143 0.030
[0792] 0149
[0793]
[0794] Novel (hu / hu) 0.541 0155 0.024 0.442 0.023
[0795]
[0257] TABLE 11
[0796] Syllables detected in courtship induced USV test in adults.
[0797] The acoustic waveform data for each adult mouse was processed by the Mouse Song Analyzer (from Erich Jarvis lab) to obtain values for each syllable: syllable type, duration (time per syllable [sec]), ISI (intersyllable interval), fqVariance (degree of variance), purity, amplitude (magnitude of loudness), bandwidth (width of peak frequency) and peak frequency (Fq) variabilities; fqmin (minimum), fqmean (mean), fqmax (maximum), fqstart (start), fqend (end). Day: recording day 1-3. The test was conducted three times per mouse, one week apart. LF: live female for stimulation. notIDd: not identified, (partial exemplary data shown)u cu
[0798] ft ft ft « i I I ft I t I, ft fl ft
[0799]
[0800] I 5 It i f f O
[0801]
[0802] W
[0803] fc
[0804]
[0805]
[0258] TABLE 12
[0806] USV features in each adult mouse.
[0807] In courtship induced USV test for adult mouse, the following USV features are calculated: number of USVs, call rate, percent. starting (percent of starting syllabic type of sequence (continuous syllables)), percent. composition (percent of syllable), sequence length. For each syllable type (“s”, ‘u”, “d”, “m”), following parameter are calculated: Bw (bandwidth). Amp (amplitude), fqVar (Fq variance), Purity, Dur (duration), and Fq variabilities: fqMin (minimum), fqMean (mean), fqMax (maximum), fqStart (start), fqEnd (end). The values were shown by each recording day per mouse (dl-d3). (partial exemplary data shown)u
[0808] co
[0809] $
[0810] fl W 1 I f f $ 9 < S
[0811]
[0812]
[0813]
[0814]
[0815]
[0816]
[0259] TABLE 13
[0817] USV features in each genotype of adult mouse.
[0818] USV features in Novalhu / hu, Novalhu / wt and Novalwt / wt pups. Bw (bandwidth), Amp (amplitude), fqVar (Fq variance), Dur (duration, sec). The values represent the mean value for each genotype. The standard error (se) values for each parameter are inserted in the adjacent mean columns. The parameters statistically different from control {Novalwt / wt) are indicated with asterisk (*p < 0.05, ** p < 0.01). / 7-valucs were calculated by Wilcoxon rank sum test and were corrected with Bonferroni method. Novalwt / wtN=l, Novalhu / wt i=\, Novalhu / hu N=l3.w caiS.rst®
[0819] 13 723. S74 119.998 144.755 3.359 tow! Ow / wt).14 712.405 197.494 143.583 21.789 8,525
[0820]
[0821] 13 827,718 13S, 354 27.033 S.77S3
[0822] ,c«m{ Fere® nr.
[0823] 77.249 10.97X 2.408 7.918 0.773
[0824] 2,462 30,495 3.371
[0825]
[0826] Bitawl (ho / iwjj 7# 3.137 13. X0& 1.5D8?.? G~ aw? genotype s4«jVar S© slfqVsr s<s u.fqVar
[0827] 0.0.3253 0.00229 0,11553 0.00247 0.12644 G. GS515 0.0:3.18.2 0.1208.9 0.00475 8.133.37 ft, W82
[0828]
[0829] ©.©3221 S.00298 042477 <■ 00295 0.1283® 000454 s.&ur U.£Hw S-i> ®>w3 (wt / Wt) O. O3482 0.00275 O.-OS-irtO 0.00422 0.06292 O. OG442 Bttawl 0. O327S ©.00219 0.03057 0.00339 0,90353
[0830]
[0831] 0.0297? 0.0504© ©.©©277 ©.©©478
[0832]
[0260] TABLE 14
[0833] USV characteristics in long / short duration “s” in adult-USVs in each genotype.
[0834] The “s” syllables were classified by the cutoff (44ms) into short or long duration (Fig. 4f). The values represent the mean value for each genotype. The standard error (se) values for each parameter are inserted in the adjacent mean columns. The parameters statistically different from control (Novalwt / wt) are indicated with asterisk (*p < 0.05, **p < 0.01). -values were calculated by Wilcoxon rank sum test and were corrected with Bonferroni method. Novalwt / wt N=13, Novalhu / wt N=14, Novalhu / hu N=13. (partial exemplary data shown)
[0835] Long duration
[0836] genotype Duration |$ei se FqVsriance [se Purity ]se Amplitude [se Sandwidth h® Nwal |wt / wt} 0.97713 0.00371 a®209| '0.00243 0.76702; 0.01419 695.75; 62.74 15342.43] 780.31 Na*81 (hu / wt) 0.07599 0.0024S 0,03825^ 0.00395 G77677f 0.01380 526.17i 53.02 13995.66 i 939.81
[0837]
[0838] Novel (hu / hu) 0.07439 0.00277 0.9598O| 0.00333 076734i 0,01218 699.031 46.81 14343.48 | 863.80 genmype Fqmax [Ht] SB Fqmin [Kt] Ise Fqineafi [Hz] se Prosit [Hz[ i se Fqend [Hzi he I Natal (wVwt| 81988.90 989,29 65746.45] 698.81 73371.371 1163.65 73567,92 i 837.22 728401 1594.97 Naval (hu / wt) 7859172 2064.12 xiiw.'l 1314.17 70751021 1670.48 71297.281 1649.24 69923661 1740.56
[0839]
[0840] Naval (hu / hu) 77818.34 100834 53474,861 56102 **69752.041 717.77 *70161111 743.10 *69336.54^ 728.83] Short duration *s"
[0841] genotype _ Durahenjsa se FqVarbTO Ise _ Purity jse _ Amplitude ise _ Sandwidth he _ Naval t wt / wl j '”^183 0079 '"'0^24651 0145 WWI Mm SOS. Szl 5074 5221HT M24 Natal ( hn / wt[ 0.02186 0,00051 6.02488^ 0.00140 0.731501 0,01181 443.83 i 20.96 521026 ^ 330.85
[0842]
[0843] Nwal |hu / w'; 0.02131 0.00069 0.023981 0.00131 0728211 0.01193 450.361 24.01 4994.891 288.37 genotype |F©w[fe] SB FqmiM&l Ise Fqmean (Hz] se fqsisrt (Hz| he Fqend (Hzj he Naval (wt / wrt| j 79553.49 955,63 74332.27] 745.17 76035.791 81842 76G3S.33I 798.91 77586.921 896.80 Nwal (bu / wQ | 776S352 1042.43 72443.25| 1388,72 75017.341 1499.® 74202.301 1450,41 75750.95] 1569.99
[0844]
[0845]
[0261] TAB E 15
[0846] Proportion of syllable composition in high / low Fqmax in adult-USVs in each genotype.
[0847] The syllables were classified by the cutoff (100 kHz) into low or high Fqmax (Fig. 4h). The ratio of low Fqmax and high Fqmax in each syllable type are shown for each genotype. notIDd: not identified. The values represent the mean value for each genotype. The standard error (se) for each parameter is inserted in the adjacent mean columns. The parameters statistically different from control (Novalwt / wt) are indicated with asterisk (*p < 0.05). / ?-values were calculated by Wilcoxon rank sum test and were corrected with Bonferroni method. Novalwt / wt N=13, Novalhu / wt N=14, Novalhu / hu N=13.
[0848] Syllable Composition in low / high Fqmax US Vs
[0849] genotype Syl high. Fqmax_component se low. Fqm ax ^component se Naval (wt / wt) s 6,416 0.032 0.772 0.021 Naval (hu / wt) s 0.339 6.035 0.755 6.62 a Naval (hu / hu) s 6,394 0.035 0.769 0.023 Naval (wt / wt) u 6139 6.616 0.064 6.066 Naval (hu / wt) u 0.143 0.014 0-067 6.00S Naval (hu / hu) u 6,162 6.617 0.074 6.664 Naval (wt / wt) d 0.229 0.012 0.101 6.611 Naval (hu / wt) d 6,212 6.617 6.116 6.614 Naval (hu / hu) d 0.236 6.616 0.107 6.616 Neva! (wt / wt) m 0.117 0.017 0.046 0.008 Noval (hu / wt) m 6,129 6.626 0.044 6.666 Naval (hu / hu) m 0.100 0.011 0.040 0.007 Naval (wt / wt) notIDd 0.126 6.615 0.029 6.066 Naval (hu / wt) notIDd 0.201 0.064 0.028 0.004
[0850]
[0851] Naval (hu / hu) notIDd 0.171 6.636 0.038 6.066
[0852]
[0262] TABLE 16
[0853] USV characteristics in low / high Fqmax in adult-USVs in each genotype.
[0854] The syllables were classified by the cutoff ( 100kHz) into low or high Fqmax (Fig. 4h). The values represent the mean value for each genotype. The standard error (se) values for each parameter are inserted in the adjacent mean columns. The parameters statistically different from control (Novalwt / wf) are indicated with asterisk (*p < 0.05, < 0.01). p- values were calculated by Wilcoxon rank sum test and were corrected with Bonferroni method. Novalwt / wt N=13, Novalhu / wt N= 14, Novalhu / hu N= 13.8*W» W
[0855] W
[0856] «
[0857] ■i’Xi'V W
[0858]
[0859] W W
[0860] U
[0861] -yr DM M / wt}
[0862]
[0863] fH
[0864] m FWwe
[0865] &
[0866] W
[0867] m
[0868]
[0869]
[0263] TABLE 17
[0870] Comparison of vocalization tests between humanized mouse models.
[0871] Comparison on USV tests between this study and three studies using humanized Foxp2 mice (Enard et al., 2009, Hammerschmidts et al., 2015, von Merten et al., 2021)10-12. The table contains experimental conditions (methods and analysis) and findings in each study. To avoid changing the nuance of the words in each report, the terms in each paper were quoted verbatim in the table (e.g., calls, elements, vocalizations).
[0872] Comparison of USV sasiysis for ttosHMd moass
[0873] B;a;i et;■!, 2003
[0874] ii»aW2 2 ate ate sfefe ® mottse FCXP2
[0875] isolate pvp fate tetr n»ter ate 2 irtog Pilfer PIS- ypraSaadGafe tw G2te-n5iA¥tsafiSASLa6PFO*i.33c AvtSOFT Recorder 22? cate soteare paxpan LMA 2035 tteter s' salts, peas ft (mean, s&t end, sax ate m®}, iocsiton of te max ffemcy, greatest dtfesence [Hz; is peas freqaeaey,
[0876] Sope et te raSSan site to te end oea fewy sspe of a Ite trend Sw# te peak ttesjsentes <sf a a#, modurs&gi of casts gifearap
[0877] rat type!: aS es& 8tat ®ffes i» or os» rater 'tefety Ms, short cte at rate tet feM no a crty tet farpency jtsrps, long cat i>50msi:
[0878] ^t^:s%tAAtregwyiU!p£ it8!;^
[0879] t-to sigtffe alters ■« te natter at rails
[0880] Agferst ite s sigfeFiSy kw sfat <m tenm ate waten peak fa^ette to st;s; s
[0881] Far short ‘s’, te ife onto caRs dated fess in tepenty in Foxfefen
[0882]
[0883] Cats tei taag fes feetert Icnger, had toatger gays. ate siaated anti fed tet Mgber peak in Fo^itetBn trfat
[0884] HaSFilSFSChFMe 5; a:, g) 15 |A tte^ eaten ft>p2 doss oat ted dfente wcteteti » tet^ A? j 13.1:2237j
[0885] hwteedKm srrirs; ate ttetette In ■ mt F0XF2 Ait ■ 3-15 itfei mfefflfaatefa _
[0886] tacorsteg 33# <
[0887] ■Bts "?;s’ ^tetea^jtetggfeatetejetw 12tteec fa tateSm, ^fagpt^teagt fate.^**8
[0888] CifeUSfe CMtt Sate tWcftSASL® Provo 52 Afe! Refer 4.2 ferrseteMe ptagatn:^ 20-2 feer a eaR te^etes, inter-itefe iffe^;£ft. gte?te iaratsFi. esemete sope _
[0889] Svtta&te c&sstet&n
[0890] tfe sigfeBt feeices fe Ste itttteer o' elates
[0891]
[0892] idtees tarcaysis fa htmatesec! Fct^a pfesd ctei etemetes teh siffiiy owe pronouncet! xsr yiaiiy jtfe ate 3 s&ftfy eater Ww ntete
[0893]
[0894]
[0264] TABLE 18
[0895]
[0265] Expression changes in Novalko / ko midbrain at E18.5.
[0896] Transcripts whose expression was affected in El 8.5 midbrain in Novai knockout mice were shown (Novalwt / wt vs. Novalko / ko). RNA sequencing data are from Saito et al., 201613. Corresponding values on the same transcripts in humanized NOVAI mice at El 8.5 midbrain are shown on the right column (Novalwt / wt vs. Novalhu / hu). Transcripts Per Kilobase Million (tpm) values represent the average value for each genotype. The analysis was performed using edgeR. NovaIwt / wt'N=, Novalko / ko N=3 for Novai knockout mice comparison, Nova Iwt / wt N=6, Nova Ihu / hu N=6 for humanized Novai mice comparison.
[0897] Hawi,< Vt> M iwt / wt} fiaval (fin Afowi MMi
[0898] M M
[0899] A
[0900] W
[0901] M
[0902] _ y_>2
[0903] ^
[0904] W
[0905] M-M
[0906] w
[0907]
[0908] A2 '6 1
[0909] X
[0910] j
[0911] | 1
[0912] 1 j O
[0913] V
[0914] |
[0915] 1:J,33^S5V$ -K. <
[0916] SL T
[0917] . 4-_ >' 1
[0918] ]
[0919] 1.7460
[0920] I
[0921] to!
[0922]
[0923] _..
[0924] _
[0925] M
[0926]
[0927]
[0928] _
[0929]
[0930]
[0931] W
[0932] TABLE REFERENCES
[0933] 1. Schaeffer, S. W. Molecular population genetics of sequence length diversity in the Adh region of Drosophila pseudoobscura. Genet. Res. 80, 163-175 (2002).
[0934] 2. Meyer, M. et al. A High-Coverage Genome Sequence from an Archaic Denisovan Individual. Science 338, 222-226 (2012).
[0935] 3. Trujillo, C. A. et al. Reintroduction of the archaic variant of NOVA 1 in cortical organoids alters neurodevelopment. Science 371, (2021).
[0936] 4. Lewis, H. A. et al. Sequence-specific RNA binding by a Nova KH domain: implications for paraneoplastic disease and the fragile X syndrome. Cell 100, 323-332 (2000).
[0937] 5. Teplova, M. et al. Protein-RNA and protein-protein recognition by dual KH1 / 2 domains of the neuronal splicing factor Nova-1. Structure 19, 930-944 ( 2011).
[0938] 6. Dredge, B. K., Stefani, G., Engelhard, C. C. & Darnell, R. B. Nova autoregulation reveals dual functions in neuronal splicing. EMBO J. 24, 1608-1620 (2005).
[0939] 7. Vogel, A. P., Tsanas, A. & Scattoni, M. L. Quantifying ultrasonic mouse vocalizations using acoustic analysis in a supervised statistical machine learning framework. Sci. Rep. 9, 1-10 (2019). 8. Grimsley, J. M. S., Monaghan, J. J. M. & Wenstrup, J. J. Development of social vocalizations in mice. PLoS One 6, e17460 (2011).
[0940] 9. Tajima, Y. et al. NOVAI acts on Impact to regulate hypothalamic function and translation in inhibitory neurons. Cell Rep. 42, 112050 (2023).
[0941] 10. Enard, W. et al. A Humanized Version of Foxp2 Affects Cortico-Basal Ganglia Circuits in Mice. Cell 137, 961-971 (2009).
[0942] 11. Hammerschmidt, K. et al. A humanized version ofFoxp2 does not affect ultrasonic vocalization in adult mice. Genes Brain Behav. 14, 583-590 (2015).
[0943] 12. von Merten, S., Pfeifle, C., Kunzel, S., Hoier, S. & Tautz, D. A humanized version of Foxp2 affects ultrasonic vocalization in adult female and male mice. Genes Brain Behav. 20, el2764 (2021).
[0944] 13. Saito, Y. et al. NOVA2-mediated RNA regulation is required for axonal pathfinding during development. Elife 5, (2016).
[0266] While the inventions have been described with reference to preferred embodiments, it will be understood by those skilled in the art that various changes may be made and equivalents may be substituted for elements thereof to adapt to particular situations without departing from the scope of the invention. Therefore, it is intended that the inventions not be limited to the particular embodiments disclosed as the best mode contemplated for carrying out this invention, but that the inventions will include all embodiments falling within the scope and spirit of the appended claims.
Claims
Claims:
1. A genetically modified non-human animal or non-human animal cell encoding a human neuro-oncological ventral antigen! (NOVA1) protein, or a portion thereof, comprising the human N0VA1 1197V variant, wherein amino acid 197 is a valine.
2. The non-human animal or cell of claim 1, wherein the genome of the animal is modified to encode a human NOVA! protein, or a portion thereof, comprising the human NOVA! I197V variant, wherein amino acid 197 is a valine.
3. The non-human animal or cell of claim 1 or 2 wherein the animal genome is homozygously altered for NOVA! to only encode a human NOVA! protein comprising a valine at amino acid 197.
4. The non-human animal or cell of claim 1 or 2 wherein the animal genome is altered to encode a human NOVA 1 protein comprising a valine at amino acid 1 7, such that the animal is heterozygous at the N0VA1 locus.
5. The non-human animal or cell of any of claims 1-4, further comprising a human or humanized N0VA2 gene.
6. The non-human animal or cell of any of claims 1-5, further comprising a human or humanized FoxP2 gene.
7. The non-human animal or non-human animal cell of any of claims 1-6, wherein the non-human animal is a rodent, dog, horse, cattle, bird, or a non-human primate such as a macaque or monkey or chimpanzee8. The non-human animal or non-human animal cell of claim 7, wherein the non-human animal is a rodent and is a mouse or a rat.
9. A non-human animal comprising a nucleic acid sequence encoding a human NOVA! protein comprising a valine at amino acid 197.
10. The non-human animal of claim 9, wherein the nucleic acid sequence is integrated in the genome of the animal at the native NOVAI locus.
11. A method for altering or modifying vocalization or vocal communication in a non-human animal comprising expressing in a non-human animal the human NOVAI protein, or a portion thereof, comprising the human NOVA! Il 97V variant, wherein amino acid 197 is a valine.
12. The method of claim 9, wherein the human NOVAI protein is homozygously expressed.
13. A method for evaluating the neuronal transcriptome in a non-human animal comprising expressing in a non-human animal the human NOVAI protein, or a portion thereof, comprising the human NOVAI I197V variant, wherein amino acid 197 is a valine and assessing neuronal transcripts and RNA processing in the non-human animal.