Engineered G-protein coupled receptors and uses thereof

CN121464153APending Publication Date: 2026-02-03ALPHELIX BIOTECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480042809.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-06-27
Filing Date
2024-06-26
Publication Date
2026-02-03

AI Technical Summary

Technical Problem

The prior art has difficulty stably existing and resolving the structure of G protein-coupled receptors in a ligand-free state, especially in cryo-electron microscopy technology, which has difficulties in analyzing the structure of G protein-coupled receptors in a non-agonistic state.

Method used

By inserting fragments of fluorescent proteins (such as GFP10-11) into specific locations of G-protein coupled receptors and binding to other fragments of fluorescent proteins (such as GFP1-9 and clamping proteins), a stable complex is improved The stability and structural rigidity of the receptor are suitable for cryo-electron microscopy structural analysis.

Benefits of technology

The stability and structural analysis of G protein-coupled receptors in the ligand-free state can be realized, and the complex structure of GPCR and agonists or inhibitors can be analyzed, reducing the cost of ligand use, and improving drug discovery and screening efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121464153A_ABST
    Figure CN121464153A_ABST
Patent Text Reader

Abstract

Provided are an engineered G protein-coupled receptor, a complex comprising the engineered G protein-coupled receptor, and uses thereof. The modified G protein coupled receptor sequentially comprises an N terminal, a first transmembrane region, a first intracellular ring, a second transmembrane region, a first extracellular ring, a third transmembrane region, a second intracellular ring, a fourth transmembrane region, a second extracellular ring, a fifth transmembrane region, a third intracellular ring, a sixth transmembrane region, a third extracellular ring, a seventh transmembrane region and a C terminal from an N terminal to a C terminal, wherein at least a portion of the first intracellular loop, the second intracellular loop and / or the third intracellular loop is replaced with an optional first linker, a first portion of a fluorescent protein and an optional second linker, and wherein the first linker and the second linker independently comprise one or more amino acids.
Need to check novelty before this filing date? Find Prior Art

Description

Modified G protein-coupled receptor and its use

[0001] Cross-references

[0002] This application refers to Chinese Patent Application No. 2023107655060, filed on June 27, 2023, entitled “A Method and Application for Modifying G Protein-Coupled Receptors,” which is incorporated herein by reference in its entirety. Technical Field

[0003] The present application belongs to the field of protein engineering technology, and specifically relates to a modified G protein-coupled receptor, a complex comprising the modified G protein-coupled receptor, and uses thereof. Background Art

[0004] Most drug molecules exert their effects in the body by binding to their corresponding target proteins. They activate or inhibit the function of the target proteins, thereby altering physiological processes and achieving the goal of treating diseases. G-protein coupled receptors (GPCRs) are a class of target proteins in the human body with important drug development potential. Approximately 40% of drugs currently on the market target GPCRs, and many more are under development. A large number of GPCRs exist in the human body, with over 800 discovered. These GPCRs share certain structural similarities: they all contain seven transmembrane helices; the N-terminus and three interhelical loops are located on the extracellular side, while the C-terminus and another three interhelical loops are located on the intracellular side. These proteins reside on the cell membrane and act as receptors for signaling molecules, transmitting signals to the cell interior. External signaling molecules bind to the extracellular portion of the GPCR, triggering conformational changes in the receptor's transmembrane helices, allowing the intracellular portion of the receptor to bind to the G protein and trigger physiological effects within the cell. GPCRs participate in a wide variety of physiological processes and play a crucial role in maintaining normal life in the human body. Drug molecules can bind to GPCRs, stimulating or inhibiting them, thereby regulating physiological processes within cells. Therefore, in theory, GPCRs have broad applications in drug discovery and screening.

[0005] However, many GPCRs cannot exist stably in the absence of ligands, which limits their application in drug discovery and screening. The recent rapid development of DNA-encoded compound library screening and the use of mass spectrometry to identify compounds that bind to specific target proteins from compound mixtures have placed high demands on the target proteins used for screening, including: good stability, ligand binding properties similar to those of natural wild-type proteins, and especially unoccupied ligand binding sites. For the discovery of therapeutic antibody drugs, whether using hybridoma technology, single B cell technology, or various display technologies, purified target proteins with these characteristics have obvious advantages over cells overexpressing target proteins, can greatly accelerate the research and development process, and avoid the influence of other background proteins. On the other hand, using these stable pure proteins, combined with biophysical techniques such as surface plasmon resonance (SPR), to accurately measure the affinity between target GPCRs and compounds or antibodies is very beneficial for drug optimization and improvement.

[0006] On the other hand, analyzing the complex structures of drug molecules or potential drug molecules with their target GPCR proteins is of great guiding significance for studying their mechanisms of action and further drug development. Currently, cryo-electron microscopy has been widely used in GPCR structural analysis, and a large number of GPCR structures have been obtained. However, the vast majority of these structures are complex structures formed by GPCRs bound to agonists and G proteins. Such complexes meet the molecular size requirements of target proteins for cryo-electron microscopy. However, when GPCRs are bound to antagonists, inverse agonists, or negative regulators and are in a non-agonized state, they do not form stable complexes with G proteins. The molecular weight of individual GPCR proteins is mostly in the range of 30-60 kD, making it difficult to analyze their structures using cryo-electron microscopy.

[0007] Therefore, there is an urgent need in this field to improve the stability of G protein-coupled receptors in the ligand-free state and to solve the problem of using cryo-electron microscopy to analyze the structure of G protein-coupled receptors in the non-excited state.

[0008] Summary of the Invention

[0009] In order to solve the above problems, the present application provides a modified G protein-coupled receptor, a complex comprising the modified G protein-coupled receptor and uses thereof.

[0010] In the first aspect, the present application provides a modified G protein coupled receptor, which comprises, from N-terminus to C-terminus, an N-terminus, a first transmembrane region, a first intracellular loop, a second transmembrane region, a first extracellular loop, a third transmembrane region, a second intracellular loop, a fourth transmembrane region, a second extracellular loop, a fifth transmembrane region, a third intracellular loop, a sixth transmembrane region, a third extracellular loop, a seventh transmembrane region, and a C-terminus, wherein at least a portion of the first intracellular loop, the second intracellular loop and / or the third intracellular loop is replaced with an optional first linker, a first part of a fluorescent protein and an optional second linker, wherein the first linker and the second linker independently comprise one or more amino acids.

[0011] In a second aspect, the present application provides a complex comprising the modified G protein-coupled receptor according to the first aspect of the present application, and a second part of a fluorescent protein, wherein the first part of the fluorescent protein is combined with the second part of the fluorescent protein.

[0012] In a third aspect, the present application provides the use of the modified G protein-coupled receptor according to the first aspect of the present application or the complex according to the second aspect of the present application in ligand affinity determination, drug discovery or screening.

[0013] In a fourth aspect, the present application provides the use of the complex described in the second aspect of the present application in analyzing three-dimensional structures using cryo-electron microscopy.

[0014] The modified G protein-coupled receptors and corresponding complexes of the present application have good stability in the ligand-free state, their ligand binding sites are unoccupied, and they have ligand binding activity similar to that of wild-type G protein-coupled receptors, thus enabling their use in drug discovery and screening. In addition, the complexes of the present application can be used to analyze their three-dimensional structures using cryo-electron microscopy, thereby providing further guidance for drug development and optimization. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 is a schematic diagram of GPCR protein engineering;

[0016] Figure 2 is an identification diagram of the purified GFP-clamp protein; wherein, Figure 2 A is a molecular sieve chromatography result diagram, and Figure 2 B is a SDS gel electrophoresis identification result diagram;

[0017] FIG3 is a diagram showing the screening results of replacing at least a portion of the third intracellular loop of the A2A adenosine receptor with an optional first linker, a first portion of a fluorescent protein (the 10th to 11th β-sheets of GFP, referred to as GFP10-11), and an optional second linker, wherein the replacement positions are shown;

[0018] FIG4 is a diagram showing the screening results in which at least a portion of the third intracellular loop of the A2A adenosine receptor is replaced with an optional first linker, GFP10-11, and an optional second linker, wherein the amino acid sequence information of the optional first linker and the optional second linker is shown;

[0019] FIG5 is a map of the pFastBac4R insect cell expression vector modified based on the pFastBac Dual vector;

[0020] Figure 6 shows the purification and identification of the A2A adenosine receptor modified protein complex and the structural analysis results of the complex with the small molecule inhibitor ZM241385; wherein, Figure 6 A is the molecular sieve chromatography result, Figure 6 B is the SDS gel electrophoresis identification result, Figure 6 C is the two-dimensional classification result of the cryo-electron microscopy data particles, Figure 6 D is the FSC curve of the three-dimensional reconstruction of the cryo-electron microscopy data, Figure 6 E is the overall density of the cryo-electron microscopy three-dimensional reconstruction protein complex, and Figure 6 F is the density of the cryo-electron microscopy three-dimensional reconstruction ZM241385 molecule;

[0021] FIG7 is a map of the mammalian cell expression vector pBacMam4R;

[0022] FIG8 shows the purification and identification of the CNR1 cannabinoid receptor modified protein complex and the structural analysis of the complex with the small molecule inhibitor taranabant; FIG8A shows the molecular sieve chromatography results, and FIG8B shows the SDS gel electrophoresis identification results;

[0023] FIG9 shows the results of molecular sieve chromatography and SDS gel electrophoresis identification of the clamp protein with a 186-Cys mutation (GFP clamp-Cys) after purification; FIG9A shows the molecular sieve chromatography results, and FIG9B shows the SDS gel electrophoresis results;

[0024] Figure 10 shows the purification and identification of the disulfide-crosslinked CNR1 cannabinoid receptor modified protein complex and the structural analysis of the complex with the inhibitor taranabant; Figure 10 A shows the molecular sieve chromatography results, Figure 10 B shows the SDS gel electrophoresis identification results; Figure 10 C shows the FSC curve of the three-dimensional reconstruction of the cryo-electron microscopy data; Figure 10 D shows the overall density of the three-dimensional reconstruction of the cryo-electron microscopy protein complex; and Figure 10 E shows the density of the taranabant molecule.

[0025] Figure 11 shows the purification and identification of the disulfide-crosslinked NK1R neurokinin receptor modified protein complex and the structural analysis of the complex with the inhibitor aprepitant; Figure 11 A shows the molecular sieve chromatography results, Figure 11 B shows the SDS gel electrophoresis identification results; Figure 11 C shows the FSC curve of the three-dimensional reconstruction of the cryo-electron microscopy data; Figure 11 D shows the overall density of the three-dimensional reconstruction of the cryo-electron microscopy protein complex; and Figure 11 E shows the aprepitant molecular density.

[0026] Figure 12 shows the purification and identification of the disulfide-crosslinked A2A adenosine receptor modified protein complex and the structural analysis of the complex with the agonist adenosine molecule; wherein, Figure 12 A shows the molecular sieve chromatography results, Figure 12 B shows the SDS gel electrophoresis identification results, Figure 12 C shows the FSC curve of the three-dimensional reconstruction of the cryo-electron microscopy data, Figure 12 D shows the overall density of the three-dimensional reconstructed protein complex by cryo-electron microscopy, and Figure 12 E shows the density of the adenosine molecule;

[0027] Figure 13 shows the purification, identification, and structural analysis results of the disulfide-crosslinked CCR8 chemokine receptor modified protein complex; Figure 13 A shows the molecular sieve chromatography results, Figure 13 B shows the SDS gel electrophoresis identification results, Figure 13 C shows the FSC curve of the three-dimensional reconstruction of the cryo-electron microscopy data, and Figure 13 D shows the overall density of the three-dimensional reconstructed protein complex by cryo-electron microscopy;

[0028] Figure 14 shows the purification, identification, and structural analysis results of the disulfide-crosslinked GPRC5D receptor modified protein complex; Figure 14 A shows the molecular sieve chromatography results, Figure 14 B shows the SDS gel electrophoresis identification results, Figure 14 C shows the FSC curve of the three-dimensional reconstruction of the cryo-electron microscopy data, and Figure 14 D shows the overall density of the three-dimensional reconstructed protein complex by cryo-electron microscopy;

[0029] FIG15 is a graph showing the results of molecular sieve chromatography and SDS gel electrophoresis identification of the GPR52-Cys-GFP10-11 protein after purification; FIG15A is a graph showing the results of molecular sieve chromatography, and FIG15B is a graph showing the results of SDS gel electrophoresis identification;

[0030] FIG16 is a diagram showing the results of molecular sieve chromatography and SDS gel electrophoresis identification of the purified MC4R-Cys-GFP10-11 protein; FIG16A is a diagram showing the results of molecular sieve chromatography, and FIG16B is a diagram showing the results of SDS gel electrophoresis identification;

[0031] FIG17 is a diagram showing the results of molecular sieve chromatography and SDS gel electrophoresis identification of the GNRHR-Cys-GFP10-11 protein after purification; wherein, FIG17A is a diagram showing the results of molecular sieve chromatography, and FIG17B is a diagram showing the results of SDS gel electrophoresis identification;

[0032] FIG18 is a diagram showing the results of molecular sieve chromatography and SDS gel electrophoresis identification of the purified PTGDR2-Cys-GFP10-11 protein; wherein FIG18A is a diagram showing the results of molecular sieve chromatography, and FIG18B is a diagram showing the results of SDS gel electrophoresis identification;

[0033] FIG19 is a diagram showing the results of molecular sieve chromatography and SDS gel electrophoresis identification of the purified HTR2C-Cys-GFP10-11 protein; FIG19A is a diagram showing the results of molecular sieve chromatography, and FIG19B is a diagram showing the results of SDS gel electrophoresis identification;

[0034] FIG20 is a diagram showing the results of molecular sieve chromatography and SDS gel electrophoresis identification of the purified ADRA2B-Cys-GFP10-11 protein; wherein FIG20A is a diagram showing the results of molecular sieve chromatography, and FIG20B is a diagram showing the results of SDS gel electrophoresis identification;

[0035] Figure 21 is a diagram showing the results of molecular sieve chromatography and SDS gel electrophoresis identification of the purified ADRB1-Cys-GFP10-11 protein; wherein, Figure 21 A is a diagram showing the results of molecular sieve chromatography, and Figure 21 B is a diagram showing the results of SDS gel electrophoresis identification;

[0036] FIG22 is a diagram showing the results of molecular sieve chromatography and SDS gel electrophoresis identification of the purified ADRB2-Cys-GFP10-11 protein; wherein, FIG22A is a diagram showing the results of molecular sieve chromatography, and FIG22B is a diagram showing the results of SDS gel electrophoresis identification;

[0037] FIG23 is a diagram showing the results of molecular sieve chromatography and SDS gel electrophoresis identification of the purified C5AR1-Cys-GFP10-11 protein; wherein FIG23A is a diagram showing the results of molecular sieve chromatography, and FIG23B is a diagram showing the results of SDS gel electrophoresis identification;

[0038] FIG24 shows the results of molecular sieve chromatography and SDS gel electrophoresis identification of the purified CCR2-Cys-GFP10-11 protein; FIG24A shows the molecular sieve chromatography results, and FIG24B shows the SDS gel electrophoresis results;

[0039] FIG25 shows the results of molecular sieve chromatography and SDS gel electrophoresis identification of the purified CCR5-Cys-GFP10-11 protein; FIG25A shows the molecular sieve chromatography results, and FIG25B shows the SDS gel electrophoresis identification results;

[0040] FIG26 shows the results of molecular sieve chromatography and SDS gel electrophoresis identification of the purified CCR6-Cys-GFP10-11 protein; FIG26A shows the molecular sieve chromatography results, and FIG26B shows the SDS gel electrophoresis results;

[0041] FIG27 shows the results of molecular sieve chromatography and SDS gel electrophoresis identification of the purified CCR7-Cys-GFP10-11 protein; FIG27A shows the molecular sieve chromatography results, and FIG27B shows the SDS gel electrophoresis results;

[0042] FIG28 is a diagram showing the molecular sieve chromatography and SDS gel electrophoresis identification results of the CHRM2-Cys-GFP10-11 protein after purification; wherein, FIG28A is a diagram showing the molecular sieve chromatography results, and FIG28B is a diagram showing the SDS gel electrophoresis identification results;

[0043] FIG29 is a diagram showing the results of molecular sieve chromatography and SDS gel electrophoresis identification of the purified CLTR2-Cys-GFP10-11 protein; FIG29A is a diagram showing the results of molecular sieve chromatography, and FIG29B is a diagram showing the results of SDS gel electrophoresis identification;

[0044] FIG30 is a diagram showing the results of molecular sieve chromatography and SDS gel electrophoresis identification of the purified CXCR2-Cys-GFP10-11 protein; wherein FIG30A is a diagram showing the results of molecular sieve chromatography, and FIG30B is a diagram showing the results of SDS gel electrophoresis identification;

[0045] FIG31 is a diagram showing the results of molecular sieve chromatography and SDS gel electrophoresis identification of the purified CXCR4-Cys-GFP10-11 protein; FIG31A is a diagram showing the results of molecular sieve chromatography, and FIG31B is a diagram showing the results of SDS gel electrophoresis identification;

[0046] FIG32 is a diagram showing the molecular sieve chromatography and SDS gel electrophoresis identification results of the purified DRD2-Cys-GFP10-11 protein; wherein, FIG32A is a diagram showing the molecular sieve chromatography results, and FIG32B is a diagram showing the SDS gel electrophoresis identification results;

[0047] FIG33 is a diagram showing the molecular sieve chromatography and SDS gel electrophoresis identification results of the purified DRD3-Cys-GFP10-11 protein; wherein, FIG33A is a diagram showing the molecular sieve chromatography results, and FIG33B is a diagram showing the SDS gel electrophoresis identification results;

[0048] FIG34 is a graph showing the molecular sieve chromatography and SDS gel electrophoresis identification results of the GPBAR-Cys-GFP10-11 protein after purification; wherein, FIG34A is a graph showing the molecular sieve chromatography results, and FIG34B is a graph showing the SDS gel electrophoresis identification results;

[0049] FIG35 is a diagram showing the results of molecular sieve chromatography and SDS gel electrophoresis identification of the purified HRH1-Cys-GFP10-11 protein; wherein FIG35A is a diagram showing the results of molecular sieve chromatography, and FIG35B is a diagram showing the results of SDS gel electrophoresis identification;

[0050] FIG36 is a diagram showing the results of molecular sieve chromatography and SDS gel electrophoresis identification of the purified HTR1A-Cys-GFP10-11 protein; FIG36A is a diagram showing the results of molecular sieve chromatography, and FIG36B is a diagram showing the results of SDS gel electrophoresis identification;

[0051] FIG37 is a diagram showing the results of molecular sieve chromatography and SDS gel electrophoresis identification of the purified HTR1B-Cys-GFP10-11 protein; FIG37A is a diagram showing the results of molecular sieve chromatography, and FIG37B is a diagram showing the results of SDS gel electrophoresis identification;

[0052] FIG38 is a diagram showing the results of molecular sieve chromatography and SDS gel electrophoresis identification of the purified HTR2B-Cys-GFP10-11 protein; FIG38A is a diagram showing the results of molecular sieve chromatography, and FIG38B is a diagram showing the results of SDS gel electrophoresis identification;

[0053] FIG39 is a diagram showing the results of molecular sieve chromatography and SDS gel electrophoresis identification of the purified LPAR1-Cys-GFP10-11 protein; FIG39A is a diagram showing the results of molecular sieve chromatography, and FIG39B is a diagram showing the results of SDS gel electrophoresis identification;

[0054] FIG40 is a diagram showing the results of molecular sieve chromatography and SDS gel electrophoresis identification of the purified MTNR1B-Cys-GFP10-11 protein; wherein, FIG40A is a diagram showing the results of molecular sieve chromatography, and FIG40B is a diagram showing the results of SDS gel electrophoresis identification;

[0055] FIG41 is a diagram showing the results of molecular sieve chromatography and SDS gel electrophoresis identification of the purified NPY1R-Cys-GFP10-11 protein; FIG41A is a diagram showing the results of molecular sieve chromatography, and FIG41B is a diagram showing the results of SDS gel electrophoresis identification;

[0056] FIG42 is a diagram showing the results of molecular sieve chromatography and SDS gel electrophoresis identification of the purified OPRD-Cys-GFP10-11 protein; wherein FIG42A is a diagram showing the results of molecular sieve chromatography, and FIG42B is a diagram showing the results of SDS gel electrophoresis identification;

[0057] FIG43 is a diagram showing the results of molecular sieve chromatography and SDS gel electrophoresis identification of the purified OX2R-Cys-GFP10-11 protein; wherein FIG43A is a diagram showing the results of molecular sieve chromatography, and FIG43B is a diagram showing the results of SDS gel electrophoresis identification;

[0058] Figure 44 shows the results of molecular sieve chromatography and SDS gel electrophoresis identification of the purified PTGDR-Cys-GFP10-11 protein; wherein, A in Figure 44 shows the results of molecular sieve chromatography, and B in Figure 44 shows the results of SDS gel electrophoresis identification.

[0059] FIG45 is a diagram showing the results of molecular sieve chromatography and SDS gel electrophoresis identification of the GPR146-Cys-GFP10-11 protein after purification; FIG45A is a diagram showing the results of molecular sieve chromatography, and FIG45B is a diagram showing the results of SDS gel electrophoresis identification;

[0060] FIG46 shows the results of molecular sieve chromatography and SDS gel electrophoresis identification of the purified MCHR1-Cys-GFP10-11 protein; FIG46A shows the molecular sieve chromatography results, and FIG46B shows the SDS gel electrophoresis identification results;

[0061] FIG47 is a diagram showing the results of molecular sieve chromatography and SDS gel electrophoresis identification of the purified TAAR1-Cys-GFP10-11 protein; FIG47A is a diagram showing the results of molecular sieve chromatography, and FIG47B is a diagram showing the results of SDS gel electrophoresis identification;

[0062] FIG48 is a diagram showing the results of molecular sieve chromatography and SDS gel electrophoresis identification of the purified AGTR1-Cys-GFP10-11 protein; FIG48A is a diagram showing the results of molecular sieve chromatography, and FIG48B is a diagram showing the results of SDS gel electrophoresis identification;

[0063] FIG49 is a diagram showing the results of molecular sieve chromatography and SDS gel electrophoresis identification of the purified FPR1-Cys-GFP10-11 protein; FIG49A is a diagram showing the results of molecular sieve chromatography, and FIG49B is a diagram showing the results of SDS gel electrophoresis identification;

[0064] FIG50 is a diagram showing the results of molecular sieve chromatography and SDS gel electrophoresis identification of the purified GALR1-Cys-GFP10-11 protein; wherein, FIG50A is a diagram showing the results of molecular sieve chromatography, and FIG50B is a diagram showing the results of SDS gel electrophoresis identification;

[0065] FIG51 is a diagram showing the results of molecular sieve chromatography and SDS gel electrophoresis identification of the purified GHSR-Cys-GFP10-11 protein; FIG51A is a diagram showing the results of molecular sieve chromatography, and FIG51B is a diagram showing the results of SDS gel electrophoresis identification;

[0066] FIG52 is a diagram showing the results of molecular sieve chromatography and SDS gel electrophoresis identification of the purified CCKAR-Cys-GFP10-11 protein; wherein FIG52A is a diagram showing the results of molecular sieve chromatography, and FIG52B is a diagram showing the results of SDS gel electrophoresis identification;

[0067] FIG53 is a diagram showing the results of molecular sieve chromatography and SDS gel electrophoresis identification of the purified MTLR-Cys-GFP10-11 protein; wherein, FIG53A is a diagram showing the results of molecular sieve chromatography, and FIG53B is a diagram showing the results of SDS gel electrophoresis identification;

[0068] FIG54 is a diagram showing the results of molecular sieve chromatography and SDS gel electrophoresis identification of the purified EDNRA-Cys-GFP10-11 protein; FIG54A is a diagram showing the results of molecular sieve chromatography, and FIG54B is a diagram showing the results of SDS gel electrophoresis identification;

[0069] FIG55 is a diagram showing the results of molecular sieve chromatography and SDS gel electrophoresis identification of the purified PRLHR-Cys-GFP10-11 protein; wherein FIG55A is a diagram showing the results of molecular sieve chromatography, and FIG55B is a diagram showing the results of SDS gel electrophoresis identification;

[0070] FIG56 is a diagram showing the results of molecular sieve chromatography and SDS gel electrophoresis identification of the purified NPFF1-Cys-GFP10-11 protein; FIG56A is a diagram showing the results of molecular sieve chromatography, and FIG56B is a diagram showing the results of SDS gel electrophoresis identification;

[0071] FIG57 shows the results of molecular sieve chromatography and SDS gel electrophoresis identification of the purified CXCR1-Cys-GFP10-11 protein; FIG57A shows the molecular sieve chromatography results, and FIG57B shows the SDS gel electrophoresis identification results;

[0072] FIG58 is a diagram showing the results of molecular sieve chromatography and SDS gel electrophoresis identification of the purified TRFR-Cys-GFP10-11 protein; FIG58A is a diagram showing the results of molecular sieve chromatography, and FIG58B is a diagram showing the results of SDS gel electrophoresis identification;

[0073] FIG59 is a diagram showing the results of molecular sieve chromatography and SDS gel electrophoresis identification of the purified CML1-Cys-GFP10-11 protein; FIG59A is a diagram showing the results of molecular sieve chromatography, and FIG59B is a diagram showing the results of SDS gel electrophoresis identification;

[0074] FIG60 is a diagram showing the results of molecular sieve chromatography and SDS gel electrophoresis identification of the purified QRFPR-Cys-GFP10-11 protein; wherein FIG60A is a diagram showing the results of molecular sieve chromatography, and FIG60B is a diagram showing the results of SDS gel electrophoresis identification;

[0075] FIG61 is a diagram showing the results of molecular sieve chromatography and SDS gel electrophoresis identification of the purified SSTR2-Cys-GFP10-11 protein; FIG61A is a diagram showing the results of molecular sieve chromatography, and FIG61B is a diagram showing the results of SDS gel electrophoresis identification;

[0076] FIG62 is a diagram showing the results of molecular sieve chromatography and SDS gel electrophoresis identification of the purified FFAR1-Cys-GFP10-11 protein; FIG62A is a diagram showing the results of molecular sieve chromatography, and FIG62B is a diagram showing the results of SDS gel electrophoresis identification;

[0077] FIG63 is a diagram showing the results of molecular sieve chromatography and SDS gel electrophoresis identification of the purified PTAFR-Cys-GFP10-11 protein; FIG63A is a diagram showing the results of molecular sieve chromatography, and FIG63B is a diagram showing the results of SDS gel electrophoresis identification;

[0078] FIG64 shows the results of molecular sieve chromatography and SDS gel electrophoresis identification of the purified P2RY1-Cys-GFP10-11 protein; FIG64A shows the molecular sieve chromatography results, and FIG64B shows the SDS gel electrophoresis identification results;

[0079] FIG65 is a diagram showing the results of molecular sieve chromatography and SDS gel electrophoresis identification of the purified HCAR2-Cys-GFP10-11 protein; FIG65A is a diagram showing the results of molecular sieve chromatography, and FIG65B is a diagram showing the results of SDS gel electrophoresis identification;

[0080] FIG66 shows the results of molecular sieve chromatography and SDS gel electrophoresis identification of the purified SUCR1-Cys-GFP10-11 protein; FIG66A shows the molecular sieve chromatography results, and FIG66B shows the SDS gel electrophoresis identification results;

[0081] FIG67 is a diagram showing the results of molecular sieve chromatography and SDS gel electrophoresis identification of the APJ-Cys-GFP10-11 protein after purification; FIG67A is a diagram showing the results of molecular sieve chromatography, and FIG67B is a diagram showing the results of SDS gel electrophoresis identification;

[0082] FIG68 shows the results of molecular sieve chromatography and SDS gel electrophoresis identification of the purified GPR39-Cys-GFP10-11 protein; FIG68A shows the molecular sieve chromatography results, and FIG68B shows the SDS gel electrophoresis identification results;

[0083] FIG69 shows the results of molecular sieve chromatography and SDS gel electrophoresis identification of the purified GPR75-Cys-GFP10-11 protein; FIG69A shows the molecular sieve chromatography results, and FIG69B shows the SDS gel electrophoresis identification results;

[0084] FIG70 shows the results of molecular sieve chromatography and SDS gel electrophoresis identification of the purified CCR1-Cys-GFP10-11 protein; FIG70A shows the molecular sieve chromatography results, and FIG70B shows the SDS gel electrophoresis results;

[0085] FIG71 shows the results of molecular sieve chromatography and SDS gel electrophoresis identification of the purified CCR4-Cys-GFP10-11 protein; FIG45A shows the molecular sieve chromatography results, and FIG44B shows the SDS gel electrophoresis results;

[0086] FIG72 shows the results of molecular sieve chromatography and SDS gel electrophoresis identification of the purified PTGER4-Cys-GFP10-11 protein; FIG72A shows the molecular sieve chromatography results, and FIG72B shows the SDS gel electrophoresis identification results;

[0087] FIG73 shows the purification, identification, and structural analysis results of the disulfide-crosslinked glucagon receptor (GCGR) modified protein complex; FIG73A shows the molecular sieve chromatography results, FIG73B shows the SDS gel electrophoresis identification results, FIG73C shows the FSC curve of the three-dimensional reconstruction of cryo-electron microscopy data, and FIG73D shows the overall density of the three-dimensional reconstructed protein complex by cryo-electron microscopy; FIG73A shows the molecular sieve chromatography results, FIG73B shows the SDS gel electrophoresis identification results, FIG73C shows the FSC curve of the three-dimensional reconstruction of the cryo-electron microscopy data, and FIG73D shows the overall density of the protein complex reconstructed by cryo-electron microscopy;

[0088] FIG74 shows the purification, identification, and molecular sieve chromatography of wild-type CCR8 and five different CCR8 fusion proteins; FIG74A shows the SDS gel electrophoresis identification results; FIG74BG shows the molecular sieve chromatography results;

[0089] FIG75 shows the results of molecular sieve chromatography and SDS gel electrophoresis identification of the purified CRFR1-ICL1-GFP10-11 protein; FIG75A shows the molecular sieve chromatography results, and FIG75B shows the SDS gel electrophoresis identification results;

[0090] Figure 76 shows the results of molecular sieve chromatography and SDS gel electrophoresis identification of the purified CRFR1-ICL2-GFP10-11 protein; wherein, A in Figure 76 shows the results of molecular sieve chromatography, and B in Figure 76 shows the results of SDS gel electrophoresis identification.

[0091] Figure 77 is a diagram showing the SDS gel electrophoresis identification results of GFP-clamp-avi-biotin protein and GFP-clamp-Cys-avi-biotin; wherein, Figure 77 A is a diagram showing the SDS gel electrophoresis identification results of GFP-clamp-avi-biotin protein, and Figure 77 B is a diagram showing the SDS gel electrophoresis identification results of GFP-clamp-Cys-avi-biotin protein;

[0092] FIG78 shows the results of purification, identification, and affinity determination of the MC4R-GFP10-11 modified protein complex with its ligand; FIG78A shows the results of molecular sieve chromatography, FIG78B shows the results of SDS gel electrophoresis identification, and FIG78CF shows the results of affinity determination;

[0093] FIG79 shows the results of purification, identification, and affinity determination of the GPR75-GFP10-11 modified protein complex with its ligand; FIG79A shows the results of molecular sieve chromatography, FIG79B shows the results of SDS gel electrophoresis identification, and FIG79C shows the results of affinity determination;

[0094] Figure 80 shows the immunoassay results of the A2A adenosine receptor modified protein complex and the CNR1 cannabinoid receptor modified protein complex; wherein, A in Figure 80 shows the detection result of the A2A adenosine receptor modified protein complex on the serum titer of A2A-immunized mice, B in Figure 80 shows the detection result of the CNR1 cannabinoid receptor modified protein complex on the serum titer of CNR1-immunized mice, C in Figure 80 shows the ELISA detection result of the A2A adenosine receptor modified protein complex and the supernatant of the anti-A2A hybridoma 96-well plate, and D in Figure 80 shows the ELISA detection result of the CNR1 cannabinoid receptor modified protein complex and the supernatant of the anti-CNR1 hybridoma 96-well plate. DETAILED DESCRIPTION

[0095] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application relates. For example, Concise Dictionary of Biomedicine and Molecular Biology, Juo, Pei-Show, 2nd edition, 2002, CRC Press; The Dictionary of Cell and Molecular Biology, 5th edition, 2013, Academic Press; and Oxford Dictionary Of Biochemistry And Molecular Biology, 2nd edition, 2006, Oxford University Press, provide a general dictionary for many of the terms used in this application.

[0096] As used herein, the expression "A and / or B" or "A and / or B" includes three cases: (1) A; (2) B; and (3) A and B. The expression "A, B and / or C" or "A, B and / or C" includes seven cases: (1) A; (2) B; (3) C; (4) A and B; (5) A and C; (6) B and C; and (7) A, B, and C. The meanings of similar expressions can be deduced analogously.

[0097] In the first aspect, the present application provides a modified G protein coupled receptor, which comprises, from N-terminus to C-terminus, an N-terminus, a first transmembrane region, a first intracellular loop, a second transmembrane region, a first extracellular loop, a third transmembrane region, a second intracellular loop, a fourth transmembrane region, a second extracellular loop, a fifth transmembrane region, a third intracellular loop, a sixth transmembrane region, a third extracellular loop, a seventh transmembrane region, and a C-terminus, wherein at least a portion of the first intracellular loop, the second intracellular loop and / or the third intracellular loop is replaced with an optional first linker, a first part of a fluorescent protein and an optional second linker, wherein the first linker and the second linker independently comprise one or more amino acids.

[0098] The modified G protein-coupled receptor of the present application has good stability in the ligand-free state, its ligand binding site is unoccupied, and has ligand binding activity similar to that of the wild-type G protein-coupled receptor, so it can be used for drug discovery and screening.

[0099] In this article, "G protein-coupled receptor" should be understood in its broadest sense. It is a general term for a large class of membrane protein receptors. The common features of this type of receptors are: they are composed of a polypeptide chain containing 7 transmembrane α-helices, and there are G protein (guanylate binding protein) binding sites on the C-terminus of the peptide chain and the intracellular loop (the third intracellular loop) connecting the 5th and 6th transmembrane helices (counting from the N-terminus of the peptide chain).

[0100] As used herein, reference to a "G protein coupled receptor" encompasses wild-type G protein coupled receptors as well as engineered G protein coupled receptors, unless otherwise indicated or clearly contradicted by the context.

[0101] Since G protein-coupled receptors have similar overall structures and secondary structure combinations, and the embodiments of the present application have selected representative members from many subfamilies (e.g., A2A, CNR1, NK1R, CCR8, GPRC5D, GCGR, MC4R, GPR75, GPR52, GNRHR, PTGDR2, HTR2C, ADRA2B, ADRB1, ADRB2, C5AR1, CCR2, CCR5, CCR6, CCR7, CHRM2, CLTR2, CXCR2, CXCR4, DRD2, DRD3, GPBAR, HRH1, HTR1A, HTR1B, HTR2B, LPAR1, MTNR1B, NPY 1R, OPRD, OX2R, PTGDR, GPR146, MCHR1, TAAR1, AGTR1, FPR1, GALR1, GHSR, CCKAR, MTLR, EDNRA, PRLHR, NPFF1, CXCR1, TRFR, CML1, QRFPR, SSTR2, FFAR1, PTAFR, P2RY1, HCAR2, SUCR1, APJ, GPR39, GPR75, PTGER4, CCR1, CCR4), and the results showed that stable proteins were obtained. Therefore, it can be reasonably inferred that the modified GPCR of the present application is applicable to and covers a wide range of GPCRs, including but not limited to 5-HT1A, 5 -HT1B, 5-HT1D, -5-HT1E, 5-HT1F, 5-HT2A, 5-HT2B, 5-HT2C, 5-HT4, 5HT5A, 5-HT6, 5-HT7, M1, M2, M3, M4, M5, A1, A2A, A2B, A3, ADA1A, ADA1B, ADA1D, ADA2 A, ADA2B, ADA2C, ADRB1, ADRB2, ADRB3, C3a, C5a, C5L2, AT1, AT2, APJ, GPBA, BB1, BB2, BB3, B1, B2, CB1, CB2, CCR1, CCR2, CCR3, CCR4, CCR5, CCR6, CCR7, CC R8, CCR9, CCR10, CXCR1, CXCR2, CXCR3, CXCR4, CXCR5, CXCR6, CXCR7, CX3CR1, XCR1, CCK1, CCK2, D1, D2, D3, D4, D5, ETA, ETB, GPER, FPR1, FPR2 / ALX, FPR3 , FFA1, FFA2, FFA3, GPR42, GAL1, GAL2, GAL3, GHSR, FSH, LH, TSH, GnRH, GnRH2, H1, H2, H3, H4, HCA1, HCA2, HCA3, KISSR, BLT1, BLT2, CysLT1, CysLT2, OXE,FPR2 / ALX、LPA1、LPA2、LPA3、LPA4、LPA5、S1P1、S1P2、S1P3、S1P4、S1P5、MCH1、MCH2、MC1、MC2、MC3、MC4、MC5、MT1、MT2、MTLR、NMU1、NMU2、NPFF1、NPFF2、NPS、NPBW1、NPBW2、Y1、Y2、Y4、Y5、NTS1、NTS2、delta、kappa、mu、NOP、OX1、OX2、P2Y1、P2Y2、P2Y4、P2Y6、P2Y11、P2Y12、P2Y13、P2Y14、QRFP、PAF、PKR1、PKR2、PRRP、DP1、DP2、EP1、EP2、EP3、EP4、FP、IP1、TP、PAR1、PAR2、PAR3、PAR4、RXFP1、RXFP2、RXFP3、RXFP4、SST1、SST2、SST3、SST4、SST5、NK1、NK2、NK3、TRH1、TA1、UT、V1A、V1B、V2、OT、CCRL2、CMKLR1、GPR1、GPR3、GPR4、GPR6、GPR12、GPR15、GPR17、GPR18、GPR19、GPR20、GPR21、GPR22、GPR25、GPR26、GPR27、GPR31、GPR32、GPR33、GPR34、GPR35、GPR37、GPR37L1、GPR39、GPR42、GPR45、GPR50、GPR52、GPR55、GPR61、GPR62、GPR63、GPR65、GPR68、GPR75、GPR78、GPR79、GPR82、GPR83、GPR84、GPR85、GPR87、GPR88、GPR101、GPR119、GPR120、GPR132、GPR135、GPR139、GPR141、GPR142、GPR146、GPR148、GPR149、GPR150、GPR151、GPR152、GPR153、GPR160、GPR161、GPR162、GPR171、GPR173、GPR174、GPR176、GPR182、GPR183、LGR4、LGR5、LGR6、LPAR6、MAS1、MAS1L、MRGPRD、MRGPRE、MRGPRF、MRGPRG、MRGPRX1、MRGPRX2、MRGPRX3、MRGPRX4、OPN3、OPN5、OXGR1、P2RY8、P2RY10、SUCNR1、TAAR2、TAAR3、TAAR4、TAAR5、TAAR6、TAAR8、TAAR9、CCPB2, CCRL1, FY, CT, CALRL, CRF1, CRF2, GHRH, GIP, GLP-1, GLP-2, GCGR, SCTR, PTH1, PTH2, PAC1, VPAC1, V PAC2, BAI1, BAI2, BAI3, CD97, CELSR1, CELSR2, CELSR3, ELTD1, EMR1, EMR2, EMR3, EMR4P, GPR56, GPR64, GPR 97. GPR98, GPR110, GPR111, GPR112, GPR113, GPR114, GPR115, GPR116, GPR123, GPR124, GPR125, GPR126, G PR128, GPR133, GPR143, GPR144, GPR157, LPHN1, LPHN2, LPHN3, CaS, GPRC6, GABAB1, GABAB2, mGlu1, mGlu2, mGlu3, mGlu4, mGlu5, mGlu6, mGlu7, mGlu8, GPR156, GPR158, GPR179, GPRC5A, GPRC5B, GPR C5C, GPRC5D, frizzled, FZD1, FZD2, FZD3, FZD4, FZD5, FZD6, FZD7, FZD8, FZD9, FZD10, SMO. ,

[0102] In some embodiments, the first linker and the second linker can be optionally present independently of each other. That is, there can be only the first linker, or only the second linker, or neither the first linker nor the second linker exists, or both the first linker and the second linker exist.

[0103] In some embodiments, the first linker and the second linker can be independently selected from a peptide segment comprising 1-50 amino acids, a peptide segment comprising 1-40 amino acids, a peptide segment comprising 1-30 amino acids, a peptide segment comprising 1-20 amino acids, a peptide segment comprising 1-10 amino acids, a peptide segment comprising 1-9 amino acids, a peptide segment comprising 1-8 amino acids, a peptide segment comprising 1-7 amino acids, a peptide segment comprising 1-6 amino acids, or a peptide segment comprising 1-5 amino acids.

[0104] In some embodiments, the first linker and the second linker can independently be 1 amino acid, a peptide consisting of 2 amino acids, a peptide consisting of 3 amino acids, a peptide consisting of 4 amino acids, a peptide consisting of 5 amino acids, a peptide consisting of 6 amino acids, a peptide consisting of 7 amino acids, a peptide consisting of 8 amino acids, a peptide consisting of 9 amino acids, a peptide consisting of 10 amino acids or more amino acids.

[0105] In some embodiments, the first linker and the second linker may be the same or different.

[0106] Herein, "fluorescent protein" encompasses wild-type fluorescent proteins as well as modified fluorescent proteins (eg, obtained by truncation, extension, insertion, deletion, or substitution).

[0107] In some embodiments, the fluorescent protein comprises one or more of green fluorescent protein (GFP), red fluorescent protein, yellow fluorescent protein, blue fluorescent protein, cyan fluorescent variant, and enhanced green fluorescent protein.

[0108] In some embodiments, the fluorescent protein comprises the amino acid sequence shown in SEQ ID NO.81 or an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% sequence identity to the amino acid sequence shown in SEQ ID NO.81.

[0109] In some embodiments, the first portion of the fluorescent protein includes the 10th to 11th β-sheet, the 1st to 2nd β-sheet, the 8th to 11th β-sheet, or the 1st to 4th β-sheet of the fluorescent protein.

[0110] In some embodiments, the first portion of the fluorescent protein comprises the amino acid sequence shown in SEQ ID NO.1 or an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% sequence identity to the amino acid sequence shown in SEQ ID NO.1.

[0111] As used herein, the term "identity" refers to the degree of similarity between a pair of sequences (nucleotides or amino acids). Identity is determined by dividing the number of identical residues by the total number of residues and multiplying the quotient by 100 to obtain a percentage. Gaps are not counted when evaluating identity. Thus, two copies of identical sequences have 100% identity, but sequences with deletions, additions, or substitutions may have a lower degree of identity. It is known to those skilled in the art that there are computer programs that can be used to determine sequence identity, such as those using algorithms such as BLAST. BLAST nucleotide searches are performed using the NBLAST program, and BLAST protein searches are performed using the BLASTP program, using the default parameters of each program.

[0112] In some embodiments, the amino acid at position 34.51 of the Ballesteros-Weinstein numbering system in the second intracellular region is cysteine.

[0113] In some embodiments, the amino acid at position 34.51 in the second intracellular region according to the Ballesteros-Weinstein numbering system is naturally cysteine. In some embodiments, the amino acid at position 34.51 in the second intracellular region according to the Ballesteros-Weinstein numbering system is mutated to cysteine ​​(e.g., represented as GPCR-Cys-GFP10-11).

[0114] In some embodiments, the engineered G protein-coupled receptor comprises an amino acid sequence as shown in any one of SEQ ID NOs.3-71, or an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% sequence identity to an amino acid sequence as shown in any one of SEQ ID NOs.3-71.

[0115] In a second aspect, the present application provides a complex comprising the modified G protein-coupled receptor according to the first aspect of the present application, and a second part of a fluorescent protein, wherein the first part of the fluorescent protein is combined with the second part of the fluorescent protein.

[0116] The complex of the present application has good stability in the ligand-free state, its ligand binding site is unoccupied, and it has ligand binding activity similar to that of wild-type G protein-coupled receptors, making it suitable for drug discovery and screening. In addition, the complex of the present application can be used to analyze the three-dimensional structure using cryo-electron microscopy, thereby providing further guidance for drug development and optimization.

[0117] In some embodiments, the first portion of the fluorescent protein and the second portion of the fluorescent protein do not contain a linker group. In other words, the first portion of the fluorescent protein and the second portion of the fluorescent protein are bound together by van der Waals forces.

[0118] In some embodiments, the first portion of the fluorescent protein and the second portion of the fluorescent protein together constitute a complete fluorescent protein.

[0119] In some embodiments, the first portion of the fluorescent protein and the second portion of the fluorescent protein are combined by co-expressing the engineered G protein-coupled receptor and the second portion of the fluorescent protein.

[0120] In some embodiments, co-expression is performed in a eukaryotic cell expression system.

[0121] In some embodiments, the eukaryotic cell is an insect cell or a mammalian cell. In some embodiments, the insect cell is a Spodoptera frugiperda cell Sf-9. In some embodiments, the mammalian cell is a HEK293F cell.

[0122] In some embodiments, the engineered G protein-coupled receptor and the second portion of the fluorescent protein (eg, GFP1-9 protein) are expressed by different promoters, such as the polyhedrin gene promoter and the P10 promoter, respectively.

[0123] In some embodiments, the N-terminus of the vector for co-expression carries a signal peptide sequence and / or a maltose binding protein fusion tag. In some embodiments, the downstream of the vector for co-expression carries an internal ribosome entry site.

[0124] In some embodiments, the gene sequence encoding the second portion of the fluorescent protein (eg, GFP1-9 protein) is cloned into the internal ribosome entry site of the expression vector.

[0125] In some embodiments, the complete fluorescent protein can be a wild-type fluorescent protein or an engineered fluorescent protein.

[0126] In some embodiments, the second portion of the fluorescent protein includes the 1st to 9th β-sheets, the 3rd to 11th β-sheets, the 1st to 7th β-sheets, or the 5th to 11th β-sheets of the fluorescent protein.

[0127] In some embodiments, the first portion of the fluorescent protein comprises the 10th to 11th beta-sheet of the fluorescent protein, and the second portion of the fluorescent protein comprises the 1st to 9th beta-sheet of the fluorescent protein. In some embodiments, the first portion of the fluorescent protein comprises the 1st to 2nd beta-sheet of the fluorescent protein, and the second portion of the fluorescent protein comprises the 3rd to 11th beta-sheet of the fluorescent protein. In some embodiments, the first portion of the fluorescent protein comprises the 8th to 11th beta-sheet of the fluorescent protein, and the second portion of the fluorescent protein comprises the 1st to 7th beta-sheet of the fluorescent protein. In some embodiments, the first portion of the fluorescent protein comprises the 1st to 4th beta-sheet of the fluorescent protein, and the second portion of the fluorescent protein comprises the 5th to 11th beta-sheet of the fluorescent protein.

[0128] In some embodiments, the second portion of the fluorescent protein comprises the amino acid sequence shown in SEQ ID NO.2 or an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% sequence identity to the amino acid sequence shown in SEQ ID NO.2.

[0129] In some embodiments, the first portion of the fluorescent protein comprises the amino acid sequence as shown in SEQ ID NO.1, or an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% sequence identity to the amino acid sequence as shown in SEQ ID NO.1, and the second portion of the fluorescent protein comprises the amino acid sequence as shown in SEQ ID NO.2, or an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% sequence identity to the amino acid sequence as shown in SEQ ID NO.2.

[0130] In some embodiments, the first portion of the fluorescent protein includes the amino acid sequence shown in SEQ ID NO.1, and the second portion of the fluorescent protein includes the amino acid sequence shown in SEQ ID NO.2.

[0131] In some embodiments, the complex further comprises a clamping protein that binds to the first portion of the fluorescent protein and the second portion of the fluorescent protein.

[0132] In some embodiments, there is no linker between the clamped protein and the first portion of the fluorescent protein, and between the clamped protein and the second portion of the fluorescent protein. In other words, the clamped protein binds to the first portion of the fluorescent protein and the second portion of the fluorescent protein via van der Waals forces.

[0133] In some embodiments, the clamping protein comprises a biomarker. The biomarker can be any biomarker known to those skilled in the art, for example, any pair of molecules capable of specific binding, including but not limited to biotin / avidin (e.g., biotin / streptavidin, biotin / neutravidin), antibody / antigen, antibody / hapten, DNA / complementary DNA or RNA, enzyme / substrate, enzyme / inhibitor, enzyme / cofactor, receptor / ligand (e.g., hormone / hormone receptor, folate / folate receptor), lectin / carbohydrate, Staphylococcus protein A / IgG, cation / anion, spy-tag / spy-catcher, strep-tag (and its mutants) / streptavidin (and its mutants). In some embodiments, the clamping protein comprises a biotin label. In some embodiments, the clamping protein comprises an avi tag. For compound drug screening and therapeutic antibody drug discovery using purified GPCRs, it would be very convenient if the target GPCR was labeled with biotin. The target protein can be fixed or modified through the very strong binding between biotin and streptavidin.

[0134] In some embodiments, the complex is a ternary complex formed by the engineered G protein-coupled receptor, the second portion of the fluorescent protein, and the clamping protein.

[0135] In some embodiments, the molecular weight of the ternary complex is above 85 kD.

[0136] In some embodiments, the ternary complex can be purified by affinity chromatography and molecular sieve chromatography.

[0137] The clamping protein can be any protein that can bind to both the first portion and the second portion of the fluorescent protein, thereby securing or enclosing the first and second portions of the fluorescent protein. For example, when the fluorescent protein is GFP, the first portion of the fluorescent protein can include the amino acid sequence set forth in SEQ ID NO. 1, the second portion of the fluorescent protein can include the amino acid sequence set forth in SEQ ID NO. 2, and the clamping protein can include the amino acid sequence set forth in SEQ ID NO. 72 (GFP clamp).

[0138] In some embodiments, the clamping protein comprises an amino acid sequence as shown in any one of SEQ ID NOs.72-74, SEQ ID NO.82, or an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% sequence identity with the amino acid sequence as shown in any one of SEQ ID NOs.72-74, SEQ ID NO.82.

[0139] In some embodiments, position 186 of the clamping protein is cysteine, and the position is based on SEQ ID NO. 72.

[0140] In some embodiments, position 186 of the clamp protein is naturally cysteine, with the position being referenced to SEQ ID NO. 72. In some embodiments, position 186 of the clamp protein is mutated to cysteine, with the position being referenced to SEQ ID NO. 72, and the resulting clamp protein may include the amino acid sequence shown in SEQ ID NO. 73 (GFP clamp-Cys).

[0141] In some embodiments, a disulfide bond is formed between the amino acid at position 186 of the clamped protein and the amino acid at position 34.51 in the second intracellular region according to the Ballesteros-Weinstein numbering system. In some embodiments, no linker groups other than disulfide bonds are present between the clamped protein and the first portion of the fluorescent protein, and between the clamped protein and the second portion of the fluorescent protein.

[0142] GPCRs share certain structural similarities: they all contain seven transmembrane helices; the N-terminus and three interhelical loops are located on the extracellular side, while the C-terminus and another three interhelical loops are located on the intracellular side. These proteins reside on the cell membrane and act as receptors for signaling molecules, transmitting signals to the cell interior. External signaling molecules bind to the extracellular portion of the GPCR, triggering conformational changes in the receptor's transmembrane helices, allowing the intracellular portion of the receptor to bind to G proteins and trigger physiological effects within the cell. GPCRs participate in a wide variety of physiological processes and play a vital role in maintaining normal life in the human body. Drug molecules can regulate physiological processes within the cell by binding to these receptor proteins, exerting either agonistic or inhibitory effects on them. Analyzing the structures of these drugs or potential drugs in complex with their target GPCR proteins is crucial for understanding the mechanisms of action of these compounds and for further drug development.

[0143] In a third aspect, the present application provides the use of the modified G protein-coupled receptor according to the first aspect of the present application or the complex according to the second aspect of the present application in ligand affinity determination, drug discovery or screening.

[0144] For use in ligand affinity determination, drug discovery, or screening, the following conditions are not essential: (1) the amino acid at position 34.51 in the second intracellular loop according to the Ballesteros-Weinstein numbering system is cysteine; (2) position 186 of the clamp protein is cysteine; (3) a disulfide bond is formed between the amino acid at position 186 of the clamp protein and the amino acid at position 34.51 in the second intracellular loop according to the Ballesteros-Weinstein numbering system. In other words, regardless of whether the above conditions are met, the engineered G protein-coupled receptor according to the first aspect of the present application or the complex according to the second aspect of the present application can be used in ligand affinity determination, drug discovery, or screening.

[0145] In a fourth aspect, the present application provides the use of the complex described in the second aspect of the present application in analyzing three-dimensional structures using cryo-electron microscopy.

[0146] In some embodiments, the amino acid at position 34.51 of the Ballesteros-Weinstein numbering system in the second intracellular loop of the engineered G protein-coupled receptor is a cysteine, the amino acid at position 186 of the clamping protein is a cysteine, and a disulfide bond is formed between the amino acid at position 186 of the clamping protein and the amino acid at position 34.51 of the Ballesteros-Weinstein numbering system in the second intracellular loop.

[0147] In some embodiments, use of the complex according to the second aspect of the present application in analyzing the three-dimensional structure of a modified G protein-coupled receptor using cryo-electron microscopy may be provided.

[0148] In some embodiments, use of the complex according to the second aspect of the present application in analyzing the three-dimensional structure of the complex using cryo-electron microscopy can be provided.

[0149] In some embodiments, the use of the complex according to the second aspect of the present application in analyzing the three-dimensional structure of a complex formed by the combination of a compound molecule and the complex using cryo-electron microscopy can be provided.

[0150] In some embodiments, the compound molecule is an inhibitor or agonist of a G protein-coupled receptor.

[0151] The complex of the present application has good stability in the ligand-free state, its ligand binding site is unoccupied, and has ligand binding activity similar to that of wild-type G protein-coupled receptors, thereby being able to be used for drug discovery and screening. In addition, the complex of the present application can be used to analyze the three-dimensional structure of the modified G protein-coupled receptor, the complex, or the complex formed by the combination of compound molecules and the complex using cryo-electron microscopy, thereby providing further guidance for drug development and optimization.

[0152] The modified GPCRs provided herein incorporate a short fusion fragment inserted into a specific position within the GPCR. This fusion fragment, combined with two other partner proteins, results in a complex particle size exceeding 85 kD. Furthermore, the partner proteins are positioned sufficiently close to the GPCR to form a relatively rigid structure. Mutating amino acids in adjacent positions on the GPCR and one of the partner proteins to cysteine ​​creates a disulfide bond between them, further reducing the flexibility of the complex. This allows for the three-dimensional structure of the GPCR-ligand complex to be elucidated using cryo-electron microscopy.

[0153] In a specific embodiment, the green fluorescent protein (GFP) can be divided into two parts, GFP1-9, comprising the first to ninth β-sheets (i.e., β-sheets), and GFP10-11, comprising the tenth and eleventh β-sheets, and these two parts can spontaneously associate. Through genetic engineering, the first intracellular loop between the first transmembrane helix (TM1) and the second transmembrane helix (TM2), the second intracellular loop between the third transmembrane helix (TM3) and the fourth transmembrane helix (TM4), or the third intracellular loop between the fifth transmembrane helix (TM5) and the sixth transmembrane helix (TM6) of a GPCR can be replaced with GFP10-11. For the A2A adenosine receptor, GFP10-11 can be inserted between L208 (Ballesteros-Weinstein numbering system 5.69a) and L225 (Ballesteros-Weinstein numbering system 6.27a), and connecting peptides can optionally be added at both ends as linkers. For other GPCRs, GFP10-11 can be inserted into the position corresponding to the A2A adenosine receptor (between 5.69a and 6.27 in the Ballesteros-Weinstein numbering system), or one, two, three, four, or five amino acids at the linker can be replaced with other amino acid sequences, such as the corresponding sequence in the A2A adenosine receptor. This GFP10-11 insertion fragment flexibly connects to the first and second helices, the third and fourth helices, or the fifth and sixth helices of the GPCR, without affecting the correct folding of the GPCR structure. It also does not form a continuous helix with the GPCR sequence at the insertion position, thereby reducing structural constraints on the GPCR. The modified GPCR is still able to bind ligand molecules. This GPCR and GFP10-11 fragment fusion protein (GPCR-GFP10-11, i.e., the modified GPCR of the present application) is co-expressed in cells with the GFP1-9 fragment. The GFP10-11 fragment in the GPCR-GFP10-11 fusion protein binds to GFP1-9 to form a complete, fluorescent GFP molecule. During the subsequent purification process, a GFP-clamp protein with a specific sequence was added to bind to the GFP moiety. This modified GPCR, such as the A2A adenosine receptor, formed a complex with its partner proteins GFP1-9 and GFP-clamp, exhibiting strong overall rigidity, enabling cryo-electron microscopy structure analysis, which revealed the high-resolution structure of the complex with the inhibitor compound.In addition, in order to further improve the overall rigidity of the complex formed by the modified GPCR and the partner proteins GFP1-9 and GFP-clamp, the amino acid at position 34.51 of the Ballesteros-Weinstein numbering system in the second intracellular loop of the GPCR can be mutated to cysteine, and the amino acid at position 186 on the GFP-clamp can be mutated to cysteine, so that a disulfide bond is formed between the GPCR and the GFP-clamp. The formation of the disulfide bond causes cross-linking between the GPCR and the GFP-clamp, and a band with a larger molecular weight is shown in non-reducing SDS electrophoresis, while in reducing SDS electrophoresis, the disulfide bond is reduced and the bands of the two proteins are still shown, which can be used to identify the formation of the disulfide bond. The disulfide bond designed in this way greatly reduces the flexibility of the entire complex, and the overall structure has strong rigidity, which is suitable for cryo-electron microscopy structure analysis. A schematic diagram of the modified G protein-coupled receptor described in this application is shown in Figure 1.

[0154] Beneficial effects of this application:

[0155] (1) The modified GPCR obtained in the present application not only has a molecular weight increased to above 85 kD, but also because the replacement fragment inserted into the GPCR (the optional first linker, the first part of the fluorescent protein, and the optional second linker) is simultaneously connected to two sites on the GPCR and simultaneously binds to the second part of the fluorescent protein and the two partner proteins of the clamping protein, the spatial structure of the entire complex is relatively rigid, which is conducive to cryo-electron microscopy structure analysis. The formation of disulfide bonds makes this beneficial effect more obvious.

[0156] (2) The first part of the fluorescent protein (e.g., GFP10-11) is very short, and its insertion has little effect on the expression and correct folding of the GPCR. It does not form a continuous helical structure with the helices on the GPCR, and has little restriction on the GPCR conformation. It is suitable for the structural analysis of various types of GPCRs with different conformations.

[0157] (3) The stable complex formed by the modified GPCR of the present application and the second part of the fluorescent protein is well expressed in both insect cells and mammalian cells.

[0158] (4) The formed complex is not easy to dissociate during the purification process, and there is no need to conduct extensive experimental exploration of the conditions for complex formation and purification.

[0159] (5) The structures of GPCR-agonist complexes as well as GPCR-inhibitor complexes can be elucidated.

[0160] (6) The formation of a complex with the partner protein (i.e., the second part of the fluorescent protein, the clamping protein) does not require the addition of a ligand. Expensive ligand compounds can be added to the complex after the sample is purified and frozen during sample preparation, greatly reducing costs.

[0161] (7) Purifying a batch of protein complexes can prepare a variety of complexes with different small molecule compounds, reducing time and workload and lowering costs.

[0162] (8) The fluorescent properties of fluorescent proteins facilitate the use of fluorescence detection size exclusion chromatography for expression screening of small samples, saving costs.

[0163] (9) There is no need for a time-consuming and low-certainty antibody screening process.

[0164] (10) It has good universality and can obtain stable proteins for many types of GPCRs.

[0165] (11) GPCR proteins in a natural conformation without ligands can be obtained, which have natural ligand binding activity similar to that of wild-type GPCRs and are suitable for compound drug screening and functional monoclonal antibody discovery.

[0166] Example

[0167] The experimental methods in the following examples are conventional methods unless otherwise specified. The pharmaceutical reagents used in the following examples are purchased from conventional biochemical reagent stores unless otherwise specified.

[0168] Example 1: Cryo-electron microscopy structure analysis of the complex between A2A adenosine receptor and ZM241385

[0169] (1) GFP-clamp expression and purification

[0170] The DNA sequence encoding the GFP-clamp protein (SEQ ID NO.72) was cloned into the E. coli expression plasmid pMALc2x, with a maltose binding protein fusion tag and a TEV protease cleavage site at the N-terminus and a histidine fusion tag at the C-terminus. TM2(DE3) competent cells. Select a positive single colony and transfer it to 20 mL of LB medium supplemented with resistance. Cultivate the cells in small amounts until supersaturated. Transfer 20 mL of cells to 800 mL of culture medium containing ampicillin and chloramphenicol resistance for amplification. When the absorbance at 600 nm is around 0.6, add 0.5 mM isopropyl-β-D-thiogalactoside and induce overnight at 20°C. Collect the culture into a large centrifuge bottle and centrifuge at 4000 rpm for 30 minutes at 4°C. Remove the supernatant and retain the pellet. Resuspend the pellet in a buffer solution (500 mM NaCl, 5% glycerol, 20 mM imidazole, 50 mM Tris-HCl, pH 7.8). Add appropriate amounts of lysozyme, nuclease, and phenylmethylsulfonyl fluoride, and disrupt the cells using a cell sonicator. Centrifuge the sonicated suspension again at high speed at 20,000 g for 30 minutes at 4°C, and retain the supernatant. Take an appropriate volume of Ni-NTA resin and combine it with the supernatant, and rotate and mix at 4°C for 1.5 hours. After sufficient binding, centrifuge at 4°C, 2000rpm for 2 minutes, discard the supernatant, and retain the resin bound to the protein. Transfer the Ni-NTA resin to a gravity column and wash the impurities with the above buffer for about 15-20 column volumes. Elute the protein with elution buffer (500mM NaCl, 5% glycerol, 500mM imidazole, 50mM Tris-HCl, pH 7.8), collect and measure the protein concentration, and add TEV protease at a mass ratio of 1 / 10 and digest overnight. Dilute with 500mM NaCl, 5% glycerol, 50mM Tris-HCl, pH 7.8 buffer to an imidazole concentration of 20mM and combine with an appropriate amount of Ni-NTA resin again. The resin was washed with 500 mM NaCl, 5% glycerol, 20 mM imidazole, 50 mM Tris-HCl, pH 7.8, and the protein was eluted with 500 mM NaCl, 5% glycerol, 500 mM imidazole, 50 mM Tris-HCl, pH 7.8. The aggregation state and purity of the protein were analyzed by molecular sieve chromatography and SDS-PAGE (see Figure 2).

[0171] (2) Optimization of the fusion protein formed by A2A adenosine receptor and GFP10-11

[0172] The amino acid sequence of the human A2A adenosine receptor was obtained from GenBank, accession number NM_000675.6. The first amino acid at the N-terminus and the C-terminal sequence after 318 were removed. The third intracellular loop was located between L208 and L225. The DNA sequence encoding the GFP10-11 fragment (SEQ ID NO. 1) was inserted between these two amino acids by overlapping PCR, replacing the original third intracellular loop sequence. To optimize the connecting peptide to maintain fusion protein stability and reduce flexibility, a series of combinations of linker positions and lengths were prepared, and the proteins of various combinations were identified and analyzed using fluorescence size exclusion chromatography (FSEC).

[0173] First, the connection position between the A2A adenosine receptor and GFP10-11 was optimized: GFP10-11 was connected between the 208th amino acid leucine L in the TM5 transmembrane region of the A2A adenosine receptor and the 220th amino acid arginine R in the TM6 transmembrane region of the A2A adenosine receptor, with a connecting peptide segment (i.e., linker) containing 5 amino acids on each side, glycine-serine-glycine-glycine-glycine (GSGGG) on the left (i.e., the first linker), and glycine-glycine-serine-glycine-glycine (GGSGG) on the right (i.e., the second linker). This combination is recorded as: L208_GSGGG_GFP10-11_GGSGG_R220(a). Similarly, the modifications at other linker positions are: L208_GSGGG_GFP10-11_GGSGG_A221 (b); L208_GSGGG_GFP10-11_GGSGG_R222 (c); L208_GSGGG_GFP10-11_GGSGG_T224 (d); and L208_GSGGG_GFP10-11_GGSGG_L225 (e). Figure 3 shows the results of a screening in which at least a portion of the third intracellular loop of the A2A adenosine receptor was replaced with an optional first linker, the first portion of a fluorescent protein (the 10th to 11th β-sheets of GFP, referred to as GFP10-11), and an optional second linker, with the replacement positions indicated. According to the screening results, GFP10-11 was selected to be connected between the 208th amino acid leucine L of the TM5 transmembrane region of the A2A adenosine receptor and the 225th amino acid leucine L of the TM6 transmembrane region of the A2A adenosine receptor.

[0174] Then, the linker between A2A adenosine receptor and GFP10-11 was optimized, and different length linkers were designed on both sides of GFP10-11: L208_GSGG_GFP10-11_GGSG_L225; L208_GGG_GFP10-11_GGG_L225; L208_GG_GFP10-11_GGG_L225; L208_GSGGG_GFP10-11_GGSGG_L225; L208_G_GFP10 -11_G_L225;L208_GGG_GFP10-11_GG_L225;L208_GG_GFP10-11_GGG_L225;L208_GG_GFP10-11_G_L225 ;L208_G_GFP10-11_GG_L225; L208_G_GFP10-11_L225; L208_GFP10-11_G_L225;

[0175] The modified A2A adenosine receptor sequences described above were cloned into the mammalian cell expression vector pBacMam4R, which carries a signal peptide sequence and a maltose binding protein fusion tag at the N-terminus, as well as a TEV protease cleavage site. The vector also contains an internal ribosome entry site. Cloning the gene sequence encoding the GFP1-9 fragment (SEQ ID NO. 2) into this internal ribosome entry site allows for the simultaneous expression of the A2A-GFP10-11 fusion protein and the GFP1-9 fragment in the cells, where they spontaneously associate to form a complex. The constructed vector DNA was transiently transfected into 293T cells using lipofectamine 2000 transfection reagent. After 40 hours, the cells were harvested, weighed, and sonicated. Cell membranes were solubilized by adding 1% lauryl maltose neopentyl glycol (LMNG) and 0.1% cholesteryl hemisuccinate (CHS), and 50 μg of GFP-clamp protein was added per 100 mg of cells. After 1.5 hours of incubation, the supernatant was centrifuged at 17,000 g and analyzed by size exclusion chromatography (SEC) using fluorescence detection. Fluorescence in the column effluent was detected using 488 nm excitation and 510 nm emission (see Figure 4). The position and shape of the fluorescence peaks were used to determine whether the protein was correctly folded and in what state it was aggregated. The L208_G_GFP10-11_L225 sequence was ultimately selected for expression and structural analysis. The amino acid sequence of the engineered A2A adenosine receptor is SEQ ID NO. 3: (The underlined portion is the GFP10-11 fragment sequence, and the wavy line is the short peptide sequence at the junction)

[0176] (3) Cryo-electron microscopy structure analysis of the complex between A2A adenosine receptor and ZM241385

[0177] The modified pFastbacDual plasmid was modified to add a maltose binding protein sequence with a hemagglutinin signal sequence at the N-terminus to the downstream sequence of the polyhedrin promoter (see Figure 5). The nucleotide sequence encoding the A2A adenosine receptor fused with the GFP10-11 fragment determined in (2) (SEQ ID NO.3) was cloned into the C-terminus of the maltose binding protein, and the nucleotide sequence encoding the GFP1-9 fragment was cloned downstream of the P10 gene promoter. The recombinant baculovirus was produced using the standard Bac-to-Bac baculovirus expression system operation process: the above plasmid was transformed into competent DH10Bac competent cells and plated with LB agar medium containing kanamycin, gentamicin and tetracycline, as well as IPTG and Bluo-gal, and incubated at 37°C in a 37°C incubator for 48 hours; after blue-white screening, white positive single colonies were picked and incubated in 5 mL LB liquid medium overnight. The bacteria were centrifuged the next day and the recombinant bacmid DNA was extracted from them. Purified recombinant bacmid DNA is transfected into Spodoptera frugiperda (Sf-9) cells using Cellfectin transfection reagent. Four to five days later, recombinant baculovirus of generation P0 can be harvested. High-titer baculovirus can be obtained by amplifying the P0 virus in Sf-9 cells.

[0178] The cell density in good growth condition was 2×10 6Sf-9 cells (1000 mM, ... The amylose resin was transferred to a gravity column. Contaminants were washed with wash buffer (500 mM NaCl, 5% glycerol, 50 mM Tris-HCl pH 7.8, 0.01% LMNG, 0.001% CHS). The target protein was eluted with elution buffer (500 mM NaCl, 5% glycerol, 20 mM maltose, 50 mM Tris-HCl pH 7.8, 0.01% LMNG, 0.001% CHS). The fraction was collected and the concentration was measured. The protein was then concentrated to 500 μL using a centrifugal filter-10K. The concentrated protein was separated using a Superdex 200 Increase 10 / 300 GL molecular sieve chromatography column with a mobile phase of buffer (500 mM NaCl, 5% glycerol, 50 mM Tris-HCl pH 7.8, 0.01% LMNG, 0.001% CHS). Proteins were collected from the apex of the UV absorbance peak corresponding to the target protein complex. An appropriate excess of GFP-clamp protein was added to the collected protein to form a complex, which was then separated again using a Superdex 200 Increase 10 / 300 GL molecular sieve chromatography column (see Figure 6, A). The mobile phase consisted of a buffer solution (150 mM NaCl, 20 mM HEPES pH 7.4, 0.005% LMNG, 0.0005% CHS). Proteins were collected from the peak corresponding to the target protein complex and concentrated to 10-15 mg / mL using a centrifugal filter. The resulting sample components were analyzed and identified by SDS-PAGE gel electrophoresis (see Figure 6, B). Molecular sieve chromatography and SDS gel electrophoresis demonstrated good protein properties and high purity, with the components contained as expected. The purified protein formed a non-aggregated UV absorption peak during molecular sieve chromatography, demonstrating its good stability.

[0179] ZM241385 was dissolved in DMSO at a concentration of 100mM, diluted 100-fold with water to a concentration of 1mM, and added to the purified and concentrated protein sample at 10% of the volume. The final concentration of ZM241385 was 100μM. The Quantifoil Cu 300 1.2 / 1.3 electron microscope grid was hydrophilized using a PELCO easiGlow glow discharger. Using a Vitrobot cryo-sampler, 3μL of protein sample was added to the hydrophilized grid. After a 4-second waiting time, it was blotted at a force of 1-6 and a time of 3-5 seconds. It was then quickly inserted into liquid ethane pre-cooled by liquid nitrogen for freezing and then transferred to liquid nitrogen for storage. The prepared frozen sample was loaded into a Titan Krios cryo-electron microscope equipped with a Gatan K2 direct electron detection camera, and image data of protein particles in the sample well was collected at an accelerating voltage of 300 kV. Data collection used the super-resolution mode, and the specific parameters were pixel size. Each image is composed of 40 frames, and the total electron dose is

[0180] The collected data were corrected for electron beam motion using MotionCor2. Contrast transfer function parameters were estimated for each image in CryoSPARC software. Particles were then selected and localized images of the protein particles extracted. After several rounds of 2D classification and removal of erroneous particles (see Figure 6, C), initial 3D density models were reconstructed ab initio. These different reconstructions were then subjected to heterogeneous refinement, and the highest-quality species with correct shape were selected. This 2D classification and 3D heterogeneous refinement process was repeated multiple times to obtain the highest-quality protein particle images and density models. Finally, after homogeneous refinement and non-uniform refinement (see Figure 6, D), high-quality 3D densities were obtained (see Figure 6, E). Based on the 3D densities obtained from the cryo-EM data, structural models of protein complexes were constructed using Coot software, and molecular structural models of small molecules were constructed based on the densities (see Figure 6, F). Then, Phenix software was used to perform structural modification and obtain a high-quality structural model.

[0181] Example 2: Human cannabinoid receptor type I (CNR1) was modified, expressed, and purified according to the method of Example 1

[0182] The human cannabinoid receptor type 1 (CNR1) sequence was derived from GenBank accession number NM_001160226.3. The N-terminal and C-terminal regions were removed, stabilizing mutations were made based on existing structures, and the GFP10-11 fragment (underlined) was fused to the third intracellular loop. The linker is marked with a wavy line. The amino acid sequence of the modified CNR1 receptor-GFP10-11 fusion protein is SEQ ID NO. 4:

[0183] The nucleotide sequence encoding the CNR1 receptor-GFP10-11 fusion protein was cloned into the mammalian cell expression vector pBacMam4R (see Figure 7), with a signal peptide sequence and a maltose binding protein fusion tag at the N-terminus. An internal ribosome entry site (IRES) downstream of the sequence allowed for co-expression of the GFP1-9 fragment. The cloned plasmid was amplified, and large amounts of plasmid DNA were extracted using an endotoxin-free plasmid extraction kit. The plasmid was then transfected into healthy HEK293F cells using PEI-40K. Twenty-four hours after transfection, 5 mM sodium butyrate was added. After an additional 24 hours, cell fluorescence was observed and cells were harvested by centrifugation at 2000 rpm for 2 minutes. Cells were resuspended in a buffer solution (500 mM NaCl, 5% glycerol, 50 mM Tris-HCl, pH 7.8) and disrupted using a cell sonicator. Lauryl maltose neopentyl glycol (LMNG) and 0.1% cholesterol hemisuccinate (CHS) were added to a final concentration of 1%, and the cells were rotated at 4°C for 2 hours to dissolve cell membranes. 4 ℃ high speed centrifugation, 20000g, centrifugation for 30min, the supernatant is mixed with an appropriate volume of amylose resin, and the mixture is rotated at 4 ℃ for 1.5 hours. After sufficient binding, centrifuge at 4 ℃, 2000rpm for 2min, discard the supernatant, and retain the resin bound to the protein. The amylose resin is transferred to a gravity column and washed with washing buffer (500mM NaCl, 5% glycerol, 50mM Tris-HCl pH 7.8, 0.01% LMNG, 0.001% CHS). The target protein is eluted with elution buffer (500mM NaCl, 5% glycerol, 20mM maltose, 50mM Tris-HCl pH 7.8, 0.01% LMNG, 0.001% CHS), and the concentration is collected and measured. Then, the protein is concentrated to 500μL using a centrifugal filter-10K. The concentrated proteins were separated using a Superdex 200 Increase 10 / 300 GL molecular sieve column with a mobile phase of buffer (500 mM NaCl, 5% glycerol, 50 mM Tris-HCl pH 7.8, 0.01% LMNG, 0.001% CHS). Proteins were collected at the apex of the UV absorption peak corresponding to the target protein complex. An appropriate excess of GFP-clamp protein was added to the collected proteins to form a complex. The complex was then separated again using a Superdex 200 Increase 10 / 300 GL molecular sieve column with a mobile phase of buffer (150 mM NaCl, 20 mM HEPES pH 7.4, 0.005% LMNG, 0.0005% CHS). Proteins were collected at the apex of the UV absorption peak corresponding to the target protein complex and concentrated to 10-15 mg / mL using a centrifugal filter. The resulting sample components were analyzed and identified by SDS-PAGE gel electrophoresis.

[0184] Molecular sieve chromatography and SDS gel electrophoresis showed that the protein had good properties (see A and B in Figure 8), high purity, and the components contained were consistent with expectations; the purified protein formed a non-aggregated UV absorption peak in molecular sieve chromatography, thus demonstrating good stability.

[0185] Example 3: Expression and purification of GFP-clamp cysteine ​​mutation and GFP-clamp-Cys mutant proteins

[0186] The glutamic acid at position 186 of GFP-clamp was mutated to cysteine ​​(marked by the border), and the rest of the fragments, including the N-terminal maltose binding protein and TEV protease cleavage sites, and the C-terminal histidine tag, remained unchanged. The amino acid sequence of GFP-clamp-Cys is SEQ ID NO. 73:

[0187] The plasmid carrying the nucleotide sequence encoding the above GFP-clamp-Cys was transformed into Rosetta TM2(DE3) competent cells. Select a positive single colony and transfer it to 20 mL of LB medium supplemented with resistance strains. Cultivate the cells in small amounts until supersaturated. Transfer 20 mL of cells to 800 mL of culture medium containing ampicillin and chloramphenicol resistance strains for amplification. When the absorbance at 600 nm is around 0.6, add 0.5 mM isopropyl-β-D-thiogalactopyranoside and induce overnight at 20°C. Collect the culture into a large centrifuge bottle and centrifuge at 4000 rpm for 30 minutes at 4°C. Remove the supernatant and retain the pellet. Resuspend the pellet in a buffer solution containing 500 mM NaCl, 5% glycerol, 20 mM imidazole, 2 mM tris(2-carbonylethyl)phosphine hydrochloride (TCEP), and 50 mM Tris-HCl, pH 7.8. Add appropriate amounts of lysozyme, nuclease, and phenylmethylsulfonyl fluoride, and disrupt the cells using a cell sonicator. The sonicated suspension was centrifuged again at high speed at 20,000 g for 30 min at 4°C, and the supernatant was retained. An appropriate volume of Ni-NTA resin was combined with the supernatant and mixed by rotation at 4°C for 1.5 h. After complete binding, the mixture was centrifuged at 2000 rpm at 4°C for 2 min, and the supernatant was discarded, retaining the protein-bound resin. The Ni-NTA resin was transferred to a gravity column and washed with the above buffer for approximately 15-20 column volumes. The protein was eluted with elution buffer (500 mM NaCl, 5% glycerol, 500 mM imidazole, 50 mM Tris-HCl, 2 mM TCEP, pH 7.8), and the protein concentration was collected and measured. TEV protease was added at a 1 / 10 mass ratio and digested overnight. The mixture was diluted to an imidazole concentration of 20 mM with a buffer containing 500 mM NaCl, 5% glycerol, 50 mM Tris-HCl, 2 mM TCEP, pH 7.8, and re-bound to an appropriate amount of Ni-NTA resin. The resin was washed with 500 mM NaCl, 5% glycerol, 20 mM imidazole, 2 mM TCEP, and 50 mM Tris-HCl, pH 7.8, and the protein was eluted with 500 mM NaCl, 5% glycerol, 500 mM imidazole, 50 mM Tris-HCl, and 2 mM TCEP, pH 7.8. The protein buffer was exchanged to 500 mM NaCl, 5% glycerol, and 50 mM Tris-HCl, pH 7.8, using a desalting column. The results of molecular sieve chromatography and SDS gel electrophoresis following purification of the GFP-clamp-Cys protein are shown in Figure 9.

[0188] Example 4: Structural analysis of the complex between a cannabinoid receptor type I (CNR1) cysteine ​​mutant and the inhibitor taranabant

[0189] Based on the CNR1 and GFP10-11 fusion protein sequence in Example 2, the leucine at position 222 (Ballesteros-Weinstein numbering system 34.51) in the second intracellular loop was mutated to cysteine ​​(border mark). The amino acid sequence of the modified CNR1-Cys_GFP10-11 is SEQ ID NO. 5:

[0190] The nucleotide sequence encoding SEQ ID NO. 5 was cloned into the mammalian cell expression vector pBacMam4R, with a signal peptide sequence and a maltose binding protein fusion tag at the N-terminus. An internal ribosome entry site (IRES) downstream of the sequence allowed for co-expression of the GFP1-9 fragment. The cloned plasmid was amplified and large amounts of plasmid DNA were extracted using an endotoxin-free plasmid extraction kit. The plasmid was then transfected into well-grown HEK293F cells using PEI-40K. Twenty-four hours after transfection, 5 mM sodium butyrate was added. After an additional 24 hours, cell fluorescence was observed and the cells were collected by centrifugation at 2000 rpm for 2 minutes. The cells were resuspended in a buffer solution (500 mM NaCl, 5% glycerol, 50 mM Tris-HCl, pH 7.8), the purified GFP-clamp-Cys protein was added, and the cells were disrupted using a cell sonicator. Add lauryl maltose neopentyl glycol (LMNG) and 0.1% cholesteryl hemisuccinate (CHS) to a final concentration of 1% and rotate at 4°C for 2 hours to dissolve cell membranes. Centrifuge at 20,000g for 30 minutes at 4°C. Mix the supernatant with an appropriate volume of Ni-NTA resin and rotate at 4°C for 1.5 hours. After sufficient binding, centrifuge at 2000 rpm at 4°C for 2 minutes. Discard the supernatant and retain the protein-bound resin. The resin was transferred to a gravity column, and impurities were washed away with a wash buffer (500 mM NaCl, 5% glycerol, 20 mM imidazole, 50 mM Tris-HCl, pH 7.8, 0.01% LMNG, 0.001% CHS), and the target protein was eluted with an elution buffer (500 mM NaCl, 5% glycerol, 500 mM imidazole, 50 mM Tris-HCl, pH 7.8, 0.01% LMNG, 0.001% CHS). The protein eluted from the Ni-NTA resin was then combined with an appropriate volume of amylose resin, rotated and mixed at 4°C for 1.5 hours, and centrifuged at 2000 rpm for 2 minutes at 4°C. The supernatant was discarded, and the amylose resin was transferred to a gravity column. The impurities were washed with washing buffer (500 mM NaCl, 5% glycerol, 50 mM Tris-HCl, pH 7.8, 0.01% LMNG, 0.001% CHS), and the target protein was eluted with elution buffer (500 mM NaCl, 5% glycerol, 20 mM maltose, 50 mM Tris-HCl, pH 7.8, 0.01% LMNG, 0.001% CHS).After adding 1 mg of GFP-clamp-Cys, the protein was concentrated to 500 μL using a centrifugal filter-10K. The protein was further purified using a Superdex 200 Increase 10 / 300 GL molecular sieve column (see Figure 10, A). The mobile phase consisted of a buffer (150 mM NaCl, 20 mM HEPES pH 7.4, 0.005% LMNG, 0.0005% CHS). Proteins were collected from the apex of the UV absorbance peak corresponding to the target protein complex. The samples were mixed with reducing and non-reducing SDS loading buffers, respectively, and treated at 37°C for 30 minutes. The resulting sample fractions were analyzed by SDS gel electrophoresis (see Figure 10, B; sample #1 was treated under reducing conditions, and sample #2 was treated under non-reducing conditions. In subsequent examples, SDS gel electrophoresis of different GPCR samples also showed sample #1 treated under reducing conditions, and sample #2 treated under non-reducing conditions). Molecular sieve chromatography and SDS gel electrophoresis showed that the protein had good properties and high purity, and the components it contained were consistent with expectations; the purified protein formed a non-aggregated ultraviolet absorption peak in molecular sieve chromatography, thus proving its good stability.

[0191] Taranabant, a CNR1 receptor antagonist, was dissolved in DMSO to a 200 mM stock solution. The stock solution was diluted 8-fold with DMSO and then further diluted 100-fold with water to obtain a 250 μM taranabant solution. A 10% volume fraction of the taranabant solution was added to the CNR1 protein sample purified by molecular sieve chromatography. The protein was concentrated to 10-15 mg / mL using a centrifugal filter-10K. Quantifoil Cu 300 1.2 / 1.3 electron microscope grids were hydrophilized using a PELCO easiGlow glow discharger. Using a Vitrobot cryostat, 4 μL of protein sample was applied to the hydrophilized grids. After a 4-second wait time, the grids were blotted at a force of 1-6 and a time of 3-5 seconds. The grids were then quickly frozen in liquid ethane pre-cooled with liquid nitrogen and transferred to liquid nitrogen for storage. The prepared frozen sample was loaded into a Titan Krios cryo-electron microscope equipped with a Gatan K2 direct electron detection camera, and image data of the protein particles in the sample well were collected at an accelerating voltage of 300 kV. The data were collected in super-resolution mode, with the specific parameters being pixel size. Each image is composed of 40 frames, and the total electron dose is

[0192] The collected data were corrected for electron beam motion using MotionCor2. Contrast transfer function parameters were estimated for each image in CryoSPARC software. Particles were then selected and localized images of the protein particles extracted. After several rounds of 2D classification and removal of erroneous particles, initial 3D density models were reconstructed ab initio. These different reconstructions were then subjected to heterogeneous refinement, and the highest-quality species with correct shape were selected. This 2D classification and 3D heterogeneous refinement process was repeated multiple times to obtain the highest-quality protein particle images and density models. Finally, after homogeneous refinement and non-uniform refinement (see Figure 10, C), high-quality 3D densities were obtained (see Figure 10, D). Based on the 3D densities obtained from the cryo-EM data, structural models of protein complexes were constructed using Coot software, and molecular structural models of small molecule compounds were constructed based on the densities (see Figure 10, E). Then, Phenix software was used to perform structural modification and obtain a high-quality structural model.

[0193] Example 5: Structural analysis of the complex between the neurokinin receptor NK1R and the inhibitor aprepitant

[0194] The NK1R amino acid sequence was derived from GenBank accession number NM_001058.4. The C-terminal region was removed from this sequence, stabilizing mutations were made based on existing structures, and the GFP10-11 fragment (underlined) was fused to the third intracellular loop. The linker is marked with a wavy line. The leucine at position 138 (Ballesteros-Weinstein numbering system 34.51) in the second intracellular loop was mutated to cysteine ​​(marked with a border). The amino acid sequence of the modified NK1R-Cys_GFP10-11 fusion protein is shown in SEQ ID NO. 6:

[0195] The purification and identification methods for the modified NK1R-Cys_GFP10-11 fusion protein are described in Example 4. The molecular sieve chromatography results for the modified NK1R-Cys_GFP10-11 fusion protein are shown in Figure 11A, and the SDS gel electrophoresis identification results are shown in Figure 11B. Molecular sieve chromatography and SDS gel electrophoresis demonstrated good protein properties and high purity, with the components contained consistent with expectations. The purified protein exhibited a non-aggregated UV absorption peak during molecular sieve chromatography, demonstrating good stability.

[0196] Aprepitant, an antagonist of the NK1R receptor, was dissolved in DMSO to prepare a stock solution with a concentration of 187 mM. The stock solution was diluted 16-fold with DMSO and then further diluted 100-fold with water to obtain a taranabant solution with a concentration of 116 μM. This solution was added to the NK1R protein sample purified by molecular sieve chromatography at a volume ratio of 10% by volume, and the protein was concentrated to 10-15 mg / ml using a centrifugal filter-10K. The specific method for structural elucidation is described in Example 4. After homogeneity correction and heterogeneity correction (see Figure 11, C), a high-quality three-dimensional density was obtained (see Figure 11, D). Based on the three-dimensional density obtained from the cryo-electron microscopy data, a structural model of the protein complex was constructed using Coot software, and a molecular structure model was constructed based on the density of the aprepitant molecule (see Figure 11, E). Structural refinement was then performed using Phenix software to obtain a high-quality structural model.

[0197] Example 6: Structural analysis of the complex between the A2A adenosine receptor and the agonist adenosine

[0198] The leucine at position 110 (Ballesteros-Weinstein numbering system 34.51) in the second intracellular loop of the A2A adenosine receptor sequence fused to the GFP10-11 fragment determined in Example 1 was mutated to cysteine ​​(marked by the border). The amino acid sequence of the modified A2A-Cys_GFP10-11 fusion protein is SEQ ID NO. 7:

[0199] The purification and identification methods for the modified A2A-Cys_GFP10-11 fusion protein are described in Example 4. The molecular sieve chromatography results for the modified A2A-Cys_GFP10-11 fusion protein are shown in Figure 12, A, and the SDS gel electrophoresis results are shown in Figure 12, B. Molecular sieve chromatography and SDS gel electrophoresis demonstrated good protein properties and high purity, with the components contained as expected. The purified protein exhibited a non-aggregated UV absorption peak during molecular sieve chromatography, demonstrating its excellent stability.

[0200] The A2A adenosine receptor complex purified by molecular sieve chromatography was concentrated to 10-15 mg / ml using a centrifugal filter (molecular weight cutoff 10 kDa), and adenosine was dissolved in pure water to prepare a mother solution with a concentration of 10 mM. It was added to the concentrated protein solution at a volume of 10%, and the final concentration was 1 mM. The specific method of structural analysis is shown in Example 4. After homogeneity correction and heterogeneity correction (see C in Figure 12), a high-quality three-dimensional density was obtained (see D in Figure 12). Based on the three-dimensional density obtained from the cryo-electron microscopy data, the structural model of the protein complex was built using Coot software, and its molecular structure model was built based on the density of the adenosine molecule (see E in Figure 12). The structure was then corrected using Phenix software to obtain a high-quality structural model.

[0201] Example 7: Structural analysis of chemokine receptor CCR8

[0202] The chemokine receptor CCR8 sequence was derived from the amino acid sequence of GenBank accession number NM_005201.4. The first amino acid was removed from this sequence, and the GFP10-11 fragment (underlined) was fused to the third intracellular loop. The linker is marked with a wavy line. V139 (Ballesteros-Weinstein numbering system 34.51) in the second intracellular loop was mutated to cysteine ​​(border mark). The amino acid sequence of the modified fusion protein is SEQ ID NO. 8:

[0203] The purification and identification methods for the modified fusion protein are described in Example 4. The results of molecular sieve chromatography of the modified fusion protein are shown in Figure 13A, and the results of SDS gel electrophoresis identification are shown in Figure 13B. Molecular sieve chromatography and SDS gel electrophoresis demonstrated good protein properties and high purity, with the components contained consistent with expectations. The purified protein formed a non-aggregated UV absorption peak in molecular sieve chromatography, demonstrating its good stability.

[0204] The CCR8 protein complex, purified by molecular sieve chromatography, was concentrated to 6 mg / ml using a centrifugal filter (molecular weight cutoff 10 kDa). The detailed structural elucidation procedures are described in Example 4. After homogeneity and heterogeneity corrections (see Figure 13, C), a high-quality three-dimensional density (see Figure 13, D) was obtained. Based on the three-dimensional density obtained from the cryo-electron microscopy data, a structural model of the protein complex was constructed using Coot software. This was then refined using Phenix software to obtain a high-quality structural model.

[0205] Thus, the modified chemokine receptor CCR8 according to the present application exhibits excellent stability without the addition of a ligand. To further illustrate the superiority of the modified GPCRs of the present application, for comparison, an unmodified wild-type CCR8 protein and modified CCR8 proteins were constructed in which the third intracellular loop was fused with several other commonly used GPCR fragments. The amino acid sequences of these proteins are as follows, with the fusion fragment sequences underlined:

[0206] Wild-type CCR8 (CCR8wt, SEQ ID NO. 75)

[0207] CCR8-T4L (T4lysozyme, SEQ ID NO.76)

[0208] CCR8-Rubredoxin (SEQ ID NO. 77)

[0209] CCR8-BRIL(Cyrochrom b562RIL,SEQ ID NO.78)

[0210] CCR8-Flavodoxin (SEQ ID NO. 79)

[0211] CCR8-PGS(Pyrococcus glycogen synthase,SEQ ID NO.80)

[0212] These proteins were tagged with an MBP tag at the N-terminus and a His tag at the C-terminus. The expression and purification methods were the same as in Example 4, except that GFP-clamp was not added during the purification process. The results (see Figure 74) showed that the obtained proteins exhibited a broad peak on the molecular sieve chromatography column, indicating that these native or modified CCR8 proteins were unable to form their native conformation, resulting in aggregation and, therefore, were unable to exist stably in the absence of ligand.

[0213] Example 8: Structural analysis of the orphan receptor GPRC5D (GPCR C family)

[0214] The orphan receptor GPRC5D sequence was derived from the amino acid sequence of GenBank accession number NM_018654.2. The first amino acid was removed from the sequence, and the GFP10-11 fragment (underlined) was fused to the third intracellular loop; the linker is marked with a wavy line. Amino acid 121 in the second intracellular loop is cysteine ​​(marked with a border) and was not mutated. The amino acid sequence of the modified fusion protein is SEQ ID NO. 9:

[0215] The purification and identification methods for the modified fusion protein are described in Example 4. The molecular sieve chromatography results for the modified fusion protein are shown in Figure 14A. A protein sample was collected at the second UV absorption peak indicated by the arrow, and the SDS gel electrophoresis results are shown in Figure 14B. Molecular sieve chromatography and SDS gel electrophoresis demonstrated good protein properties and high purity, with the components contained as expected. The purified protein formed a non-aggregated UV absorption peak during molecular sieve chromatography, demonstrating good stability.

[0216] The CCR8 protein complex, purified by molecular sieve chromatography, was concentrated to 6 mg / ml using a centrifugal filter (molecular weight cutoff 10 kDa). The detailed structural elucidation procedures are described in Example 4. After homogeneity and heterogeneity corrections (see Figure 14, C), a high-quality three-dimensional density (see Figure 14, D) was obtained. Based on the three-dimensional density obtained from the cryo-EM data, a structural model of the protein complex was constructed using Coot software. This was then refined using Phenix software to obtain a high-quality structural model.

[0217] Example 9: (GPCR B1 family) Glucagon receptor GCGR structure analysis

[0218] The glucagon receptor GCGR sequence was derived from the amino acid sequence of GenBank accession number NM_000160.5. The N-terminal signal peptide sequence was removed and the GFP10-11 fragment (underlined) was fused to the third intracellular loop. The linker is marked with a wavy line. T234 (34.51 in the Ballesteros-Weinstein numbering system) in the second intracellular loop was mutated to cysteine ​​(bordered). The resulting fusion protein has the amino acid sequence of SEQ ID NO. 10:

[0219] The purification and identification methods for the modified fusion protein are described in Example 4. The results of molecular sieve chromatography of the modified fusion protein are shown in Figure 73A, and the results of SDS gel electrophoresis identification are shown in Figure 73B. Molecular sieve chromatography and SDS gel electrophoresis demonstrated good protein properties and high purity, with the components contained consistent with expectations. The purified protein formed a non-aggregated UV absorption peak in molecular sieve chromatography, demonstrating good stability.

[0220] The GCGR protein complex, purified by molecular sieve chromatography, was concentrated to 8 mg / ml using a centrifugal filter (molecular weight cutoff 10 kDa). The specific method for structural elucidation is described in Example 4. After homogeneity correction and heterogeneity correction (see Figure 73, C), a high-quality three-dimensional density was obtained (see Figure 73, D). Based on the three-dimensional density obtained from the cryo-electron microscopy data, a structural model of the protein complex was constructed using Coot software, and then structural correction was performed using Phenix software to obtain a high-quality structural model.

[0221] Example 10: Other forms of modified GPCRs according to the present application

[0222] The CRFR receptor sequence was derived from the amino acid sequence of GenBank No. NP_001138618. The N-terminal signal peptide sequence was removed and the GFP10-11 fragment (underlined) was fused to the first intracellular loop. The amino acid sequence of the modified fusion protein is SEQ ID NO. 11:

[0223] The purification and identification methods for the modified fusion protein are described in Example 4. The results of molecular sieve chromatography of the modified fusion protein are shown in Figure 75A, and the results of SDS gel electrophoresis identification are shown in Figure 75B. Molecular sieve chromatography and SDS gel electrophoresis demonstrated good protein properties and high purity, with the components contained consistent with expectations. The purified protein formed a non-aggregated UV absorption peak in molecular sieve chromatography, demonstrating good stability.

[0224] The CRFR receptor N-terminal signal peptide sequence was removed and the GFP10-11 fragment (underlined) was fused to the second intracellular loop. The amino acid sequence of the modified fusion protein is SEQ ID NO.12:

[0225] The purification and identification methods for the modified fusion protein are described in Example 4. The results of molecular sieve chromatography of the modified fusion protein are shown in Figure 76A, and the results of SDS gel electrophoresis identification are shown in Figure 76B. Molecular sieve chromatography and SDS gel electrophoresis demonstrated good protein properties and high purity, with the components contained consistent with expectations. The purified protein formed a non-aggregated UV absorption peak in molecular sieve chromatography, demonstrating good stability.

[0226] The results showed that fusion of GFP10-11 fragments into the first or second intracellular loop of GPCR could form stable complexes with GFP1-9 and GFP-clamp proteins.

[0227] Example 11: Preparation of biotinylated GFP-clamp

[0228] Leveraging the ability of the engineered GPCR-GFP10-11 fusion protein to form a stable complex with GFP19 and GFP-clamp, an Avi tag was added to GFP-clamp and co-expressed with the biotin ligase BirA to biotinylate GFP-clamp-avi within cells. The purified GFP-clamp-avi-biotin was then used to purify GPCR-GFP10-11. The resulting complex, now bearing a biotin tag, can be used for compound and antibody drug screening or affinity determination using biophysical techniques.

[0229] The amino acid sequence of GFP-clamp-avi is as follows (SEQ ID NO. 74), with the Avi tag (marked by a wavy line) at its C-terminus:

[0230] GFP-clamp-avi and biotin ligase BirA in Rosetta TM 2(DE3) competent cells and purified in the same manner as in Example 1(1) to obtain biotinylated GFP-clamp-avi-biotin. The binding experiment with streptavidin magnetic beads determined that the degree of biotinylation of the protein was very high, and almost all of the protein was bound to the streptavidin magnetic beads. The SDS gel electrophoresis results of the GFP-clamp-avi-biotin protein are shown in Figure 77A. SDS gel electrophoresis showed that the protein had good properties and high purity, and the various components contained were consistent with expectations; the purified protein formed a non-aggregated ultraviolet absorption peak in the molecular sieve chromatography, thus proving its good stability.

[0231] The glutamic acid at position 186 of GFP-clamp-Avi was mutated to cysteine ​​(marked by the border), and the rest of the fragments, including the N-terminal maltose binding protein and TEV protease cleavage sites, and the C-terminal histidine tag, remained unchanged. The amino acid sequence of GFP-clamp-Cys-Avi is SEQ ID NO.82:

[0232] GFP-clamp-Cys-avi and biotin ligase BirA in Rosetta TM2(DE3) competent cells and purified in the same manner as in Example 3 to obtain biotinylated GFP-clamp-avi-biotin. Binding experiments with streptavidin magnetic beads confirmed that the degree of biotinylation of the protein was very high, with almost all of the protein bound to the streptavidin magnetic beads. The SDS gel electrophoresis results of the GFP-clamp-Cys-avi-biotin protein are shown in Figure 77B. SDS gel electrophoresis showed that the protein had good properties and high purity, and the various components contained were consistent with expectations; the purified protein formed a non-aggregated UV absorption peak in molecular sieve chromatography, thus demonstrating good stability.

[0233] Example 12: Binding properties of the modified MC4R according to the present application and its ligand

[0234] The GFP10-11 fragment (underlined) was fused to the third intracellular loop of melanocortin receptor 4 (MC4R), and the linker was marked with a wavy line. The modified sequence is SEQ ID NO.13:

[0235] The purification and characterization methods for the modified MC4R fusion protein are described in Example 4, except that GFP-clamp-avi-biotin was added during the purification process instead of GFP-clamp. The molecular sieve chromatography results of the modified fusion protein are shown in Figure 78, A, and the SDS gel electrophoresis results are shown in Figure 78, B. Both molecular sieve chromatography and SDS gel electrophoresis demonstrated good protein properties and high purity, with the components contained as expected. The purified protein exhibited a non-aggregated UV absorption peak during molecular sieve chromatography, demonstrating its excellent stability.

[0236] The purified MC4R was used for surface plasmon resonance (SPR) analysis to determine its affinity for four known ligands. The instrument used was a Biacore 8K, using a SAD200M SPR sensor chip. Both the protein immobilization buffer and the running buffer consisted of 20 mM HEPES, 150 mM NaCl, 2 mM CaCl2, 1 mM MgCl2, 1 mM EDTA, 0.005% LMNG, 0.0005% CHS, pH 7.4. The affinity results are shown in Figures 78, C-F, respectively.

[0237] The results showed that the modified MC4R protein according to the present application was stable in nature, and the affinity test results with several known compounds were close to those reported, indicating that the ligand binding site of the modified GPCR has a structure consistent with that of the natural protein and can be used for screening compound libraries.

[0238] Example 13: Binding properties of the modified GPR75 according to the present application and its ligand

[0239] The GFP10-11 fragment (underlined) was fused to the third intracellular loop of the obesity-related receptor GPR75 (1-395). The modified sequence is SEQ ID NO.14:

[0240] The purification and characterization methods for the modified GPR75 fusion protein are described in Example 4, except that GFP-clamp-avi-biotin was added during the purification process instead of GFP-clamp. The results of molecular sieve chromatography of the modified fusion protein are shown in Figure 79A, and the results of SDS gel electrophoresis are shown in Figure 79B. Both molecular sieve chromatography and SDS gel electrophoresis demonstrated good protein properties and high purity, with the components contained as expected. The purified protein exhibited a non-aggregated UV absorption peak during molecular sieve chromatography, demonstrating its excellent stability.

[0241] Purified GPR75 was used for surface plasmon resonance (SPR) analysis to determine its affinity for the CCL5 ligand. The instrument used was a Biacore 8K, using a Series S Sensor SA Chip: SA-2429. Both the protein immobilization buffer and the running buffer were 20 mM HEPES, 150 mM NaCl, 0.005% LMNG, 0.0005% CHS, pH 7.4. The affinity results are shown in Figure 79, C.

[0242] The results showed that the modified GPR75 protein according to the present application had stable properties, and the affinity determination results with CCL5 were close to those reported, indicating that the ligand binding site of the modified GPCR had a structure consistent with that of the natural protein and could be used for screening compound libraries.

[0243] Example 14: Use of the modified A2A adenosine receptor and cannabinoid type I receptor (CNR1) according to the present application for antibody discovery

[0244] The modified sequences of the A2A receptor and CNR1 receptor were the same as those in Example 1 and Example 2, respectively. The two modified receptors were co-expressed with GFP1-9, and after adding GFP-clamp-avi-biotin, they were purified by the same method as in Example 1 to obtain a biotin-tagged complex.

[0245] ELISA assay for antibody titers in the serum of BALB / C mice immunized with virus-like particles (VLPs) of A2A receptor and CNR1 receptor: Streptavidin (0.1 M carbonate buffer, pH 9.6, 2.5 μg / well) was added to a 96-well ELISA plate and coated overnight at 4°C. 350 μL of 40 mg / ml bovine serum albumin (BSA) was added to each well for blocking. The plate was then washed with PBSL solution (10 mM The wells were washed once with PBS (pH 7.4, 0.01% LMNG, 0.001% CHS); 1 μg of purified biotinylated GPCR protein complex was added to each well, incubated at room temperature for 1 hour, and washed once with PBSL. Each well was blocked with PBSL containing 40 mg / ml bovine serum albumin (BSA). Mouse serum was diluted in PBS at various times and added to each well, incubated at room temperature for 1 hour, and washed four times with PBSL. A goat anti-mouse antibody conjugated to horseradish peroxidase was added and incubated at room temperature for 1 hour, and washed four times with PBSL. 100 μL of TMB substrate was added to each well and incubated in the dark at room temperature for 10-20 minutes for color development. Once the color development stabilized, 50 μL of stop solution was added to each well, and the absorbance of each well was read at 450 nm on a microplate reader. The titer of sera from A2A- and CNR1-immunized mice is shown in Figures 80, A and B, respectively.

[0246] BALB / C mice immunized with virus-like particles (VLPs) encoding the target proteins A2A and CNR1 were sacrificed and immersed in 75% ethanol for 5 minutes before removal. The abdomen of the mice was wiped dry, the limbs were fixed, and the abdominal skin, muscle, and peritoneum were cut open. The spleen was removed and temporarily stored in PBS containing the double-antibody. The spleen was washed in serum-free RPMI-1640 medium and placed on a 70μm cell sieve. The spleen was ground with a grinding rod, with RPMI-1640 medium continuously added during the grinding process. Splenocytes that passed through the sieve were collected in a centrifuge tube and centrifuged at 300g for 10 minutes. The supernatant was discarded, and the cells were resuspended and washed twice in RPMI-1640 medium and counted. Simultaneously, well-grown mouse myeloma SP2 / 0 cells were collected from a culture dish, centrifuged at 300g for 5 minutes, resuspended in serum-free RPMI-1640 medium, and counted.

[0247] Mix 1-2x10^8 resuspended mouse spleen cells with SP2 / 0 cells at a 10:1 ratio, centrifuge at 300g for 10 minutes, and aspirate the supernatant. Place the centrifuge tube in a 37°C water bath. Use a pipette to draw up 1 ml of 50% PEG 1450 solution preheated to 37°C. Insert the pipette tip into the cell pellet and slowly add the PEG solution, gently stirring to thoroughly mix the cells and PEG solution. Add the PEG solution within 1 minute and continue stirring gently for 1 minute. Then, draw up 1 ml of serum-free RPMI-1640 medium and slowly add the cell-PEG mixture in the same manner, stirring gently. Slowly add 10 ml of serum-free RPMI-1640 medium, stirring continuously. Centrifuge at 300g for 10 minutes to collect the cell pellet, resuspend the cells in 1 ml of RPMI-1640 medium supplemented with 10% FBS, and remove the supernatant. Resuspend the cells in 120 ml of HAT semisolid selection medium (RPMI-1640 medium, containing 10% FBS, 1x HAT selection reagent, 1x Hybridoma feeder supplement, 1x penicillin / streptomycin, 1.68% methylcellulose), mix well, and pour into eight 100 mm cell culture dishes. Incubate at 37°C in an incubator containing 5% carbon dioxide. After 10-14 days, the successfully fused hybridoma cells will grow into visible cell clusters.

[0248] Add 200 μl of RPMI-1640 medium (containing 10% FBS, 0.5x Hybridoma feeder supplement, and 1x penicillin / streptomycin) to each well of a 96-well cell culture plate. Transfer the single cell clusters from the semi-solid medium to each well of the 96-well plate sequentially and allow them to grow until they cover most of the well bottom. During growth, the hybridoma cells secrete antibodies into the culture supernatant.

[0249] Follow the same steps as the serum titer test above to perform ELISA to detect target-specific antibodies in each well:

[0250] Streptavidin (0.1 M carbonate buffer, pH 9.6, 2.5 μg / well) was added to a 96-well ELISA plate and coated overnight at 4°C. 350 μL of 40 mg / ml bovine serum albumin (BSA) was added to each well for blocking, and the wells were washed once with PBSL solution (10 mM PBS, pH 7.4, 0.01% LMNG, 0.001% CHS). 1 μg of purified biotinylated GPCR protein complex was added to each well, incubated at room temperature for 1 hour, and washed once with PBSL solution. Each well was blocked with PBSL containing 40 mg / ml bovine serum albumin (BSA). The supernatant of hybridoma cells cultured in each well of the 96-well cell culture plate was transferred to the corresponding wells of the 96-well ELISA plate and incubated at room temperature for 1 hour. The wells were washed four times with PBSL solution. 100 μL of horseradish peroxidase-conjugated goat anti-mouse antibody was added to each well and incubated at room temperature for 1 hour. The wells were washed four times with PBSL solution. Incubate the TMB substrate in the dark at room temperature for 10-20 minutes to allow color development. Once color development stabilizes, add 50 μL of stop solution to each well. The absorbance of each well is then read at 450 nm on a microplate reader. The ELISA results for the supernatant from one 96-well plate containing anti-A2A hybridomas and one from one 96-well plate containing anti-CNR1 hybridomas are shown in Figures 80, C and D, respectively.

[0251] The results show that the modified GPCRs described in this application can be effectively used for antibody discovery. The difficulty in discovering therapeutic GPCR antibodies using mouse immunization and hybridoma technology lies in the ability of the antibodies generated and identified to recognize GPCRs in their native conformation on the surface of human cells. First, mice must be immunized with GPCRs in their native conformation to stimulate the mouse immune system to produce antibodies that recognize GPCRs in their native conformation. Hybridomas are then prepared and identified using GPCRs in their native conformation. Considering that a large proportion of antibodies, while capable of binding to the target protein, are unsuitable for use as therapeutic antibodies due to factors such as affinity and stability, hybridoma preparation and screening are crucial for obtaining as many hybridoma cell lines as possible that secrete specific antibodies. However, multiply transmembrane proteins like GPCRs have a very small extracellular domain, and the proportion of B cells producing GPCR antibodies in the spleen after immunization is typically very low (as shown by the percentage of positive cells in Figure 80). Therefore, it is often necessary to generate a large number (at least several thousand) of hybridoma cells for screening. ELISA has the characteristics of high sensitivity and high throughput, and is suitable for quickly identifying cells that secrete antibodies that recognize GPCRs in their native conformation. However, high-purity proteins are required to identify specific antibodies to avoid screening out hybridoma cells that secrete other non-specific antibodies. The amount of purified protein required to identify a large number of hybridoma cells is also relatively large. The present invention can easily obtain high-purity GPCR proteins, which can be used in ELISA to eliminate the interference of other background proteins and quickly screen out hybridoma cells that secrete specific antibodies against the target GPCR (as shown in Figure 80). In particular, the present invention obtains GPCR proteins with a native structure on the extracellular side, and screens out antibodies that recognize the native structure of GPCRs, which is crucial for therapeutic antibodies.

[0252] Example 15: Other modified G protein-coupled receptors according to the present application

[0253] The amino acid sequences of the following GPCRs were modified using the same method as described for CNR1 and NK1R: 1. At least a portion of the third intracellular loop of these GPCRs was replaced with the GFP10-11 fragment (SEQ ID NO. 1) flanked by linkers; 2. Amino acids 34.51 (Ballesteros-Weinstein numbering system) in the second intracellular loop were mutated to cysteine. The amino acid sequences of the modified GPCRs are as follows:

[0254] GPR52-Cys-GFP10-11 (SEQ ID NO.15):

[0255] GNRHR-Cys-GFP10-11 (SEQ ID NO.16):

[0256] PTGDR2-Cys-GFP10-11(SEQ ID NO.17):

[0257] HTR2C-Cys-GFP10-11(SEQ ID NO.18):

[0258] ADRA2B-Cys-GFP10-11(SEQ ID NO.19):

[0259] ADRB1-Cys-GFP10-11(SEQ ID NO.20):

[0260] ADRB2-Cys-GFP10-11(SEQ ID NO.21):

[0261] C5AR1-Cys-GFP10-11(SEQ ID NO.22):

[0262] CCR2-Cys-GFP10-11(SEQ ID NO.23):

[0263] CCR5-Cys-GFP10-11(SEQ ID NO.24):

[0264] CCR6-Cys-GFP10-11(SEQ ID NO.25):

[0265] CCR7-Cys-GFP10-11(SEQ ID NO.26):

[0266] CHRM2-Cys-GFP10-11(SEQ ID NO.27):

[0267] CLTR2-Cys-GFP10-11(SEQ ID NO.28):

[0268] CXCR2-Cys-GFP10-11(SEQ ID NO.29):

[0269] CXCR4-Cys-GFP10-11(SEQ ID NO.30):

[0270] DRD2-Cys-GFP10-11(SEQ ID NO.31):

[0271] DRD3-Cys-GFP10-11(SEQ ID NO.32):

[0272] GPBAR-Cys-GFP10-11(SEQ ID NO.33):

[0273] HRH1-Cys-GFP10-11(SEQ ID NO.34):

[0274] HTR1A-Cys-GFP10-11(SEQ ID NO.35):

[0275] HTR1B-Cys-GFP10-11(SEQ ID NO.36):

[0276] HTR2B-Cys-GFP10-11(SEQ ID NO.37):

[0277] LPAR1-Cys-GFP10-11(SEQ ID NO.38):

[0278] MTNR1B-Cys-GFP10-11(SEQ ID NO.39):

[0279] NPY1R-Cys-GFP10-11(SEQ ID NO.40):

[0280] OPRD-Cys-GFP10-11(SEQ ID NO.41):

[0281] OX2R-Cys-GFP10-11(SEQ ID NO.42):

[0282] PTGDR-Cys-GFP10-11(SEQ ID NO.43):

[0283] GPR146(SEQ ID NO.44)

[0284] MCHR1(SEQ ID NO.45)

[0285] TAAR1(SEQ ID NO.46)

[0286] AGTR1(SEQ ID NO.47)

[0287] FPR1(SEQ ID NO.48)

[0288] GALR1(SEQ ID NO.49)

[0289] GHSR(SEQ ID NO.50)

[0290] CCKAR(SEQ ID NO.51)

[0291] MTLR(SEQ ID NO.52)

[0292] EDNRA(SEQ ID NO.53)

[0293] PRLHR(SEQ ID NO.54)

[0294] NPFF1(SEQ ID NO.55)

[0295] CXCR1(SEQ ID NO.56)

[0296] TRFR(SEQ ID NO.57)

[0297] CML1(SEQ ID NO.58)

[0298] QRFPR(SEQ ID NO.59)

[0299] SSTR2(SEQ ID NO.60)

[0300] FFAR1(SEQ ID NO.61)

[0301] PTAFR(SEQ ID NO.62)

[0302] P2RY1(SEQ ID NO.63)

[0303]

[0304] HCAR2(SEQ ID NO.64)

[0305] SUCR1(SEQ ID NO.65)

[0306] APJ(SEQ ID NO.66)

[0307] GPR39(SEQ ID NO.67)

[0308] GPR75(SEQ ID NO.68)

[0309] PTGER4(SEQ ID NO.69)

[0310] CCR1(SEQ ID NO.70)

[0311] CCR4(SEQ ID NO.71)

[0312] The nucleotide sequences encoding the modified GPCR and GFP10-11 fusion proteins were cloned into the mammalian cell expression vector pBacMam4R and transfected into HEK293F cells using the same procedures as in the previous example to co-express the GPCR and GFP10-11 fusion proteins and the GFP1-9 fragment. After harvesting, the cells were added with the GFP-clamp-Cys protein obtained in Example 3 and purified using the same procedures as in Example 7. The purified proteins were analyzed by molecular sieve chromatography and SDS gel electrophoresis, as shown in Figures 15-74. These results demonstrate that these modified GPCRs exhibited excellent properties and high purity, with the components contained therein consistent with expectations. The purified proteins exhibited a non-aggregated UV absorption peak on molecular sieve chromatography, demonstrating excellent stability.

Claims

1. A modified G protein coupled receptor, which comprises, from N-terminus to C-terminus, an N-terminus, a first transmembrane region, a first intracellular loop, a second transmembrane region, a first extracellular loop, a third transmembrane region, a second intracellular loop, a fourth transmembrane region, a second extracellular loop, a fifth transmembrane region, a third intracellular loop, a sixth transmembrane region, a third extracellular loop, a seventh transmembrane region, and a C-terminus, wherein at least a portion of the first intracellular loop, the second intracellular loop and / or the third intracellular loop is replaced with an optional first linker, a first portion of a fluorescent protein and an optional second linker, wherein the first linker and the second linker independently comprise one or more amino acids.

2. The engineered G protein-coupled receptor according to claim 1, wherein the fluorescent protein comprises one or more of green fluorescent protein, red fluorescent protein, yellow fluorescent protein, blue fluorescent protein, cyan fluorescent variant, and enhanced green fluorescent protein.

3. The engineered G protein coupled receptor according to claim 1 or 2, wherein the first portion of the fluorescent protein comprises the 10th-11th β-sheet, the 1st-2nd β-sheet, the 8th-11th β-sheet or the 1st-4th β-sheet of the fluorescent protein.

4. The engineered G protein-coupled receptor according to any one of claims 1 to 3, wherein the first portion of the fluorescent protein comprises the amino acid sequence as shown in SEQ ID NO.1 or an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% sequence identity to the amino acid sequence as shown in SEQ ID NO.

1.

5. The engineered G protein-coupled receptor according to any one of claims 1 to 4, wherein the amino acid at position 34.51 of the Ballesteros-Weinstein numbering system in the second intracellular loop is cysteine.

6. The engineered G protein coupled receptor according to any one of claims 1-5, wherein the engineered G protein coupled receptor comprises an amino acid sequence as shown in any one of SEQ ID NOs.3-71 or an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% sequence identity with an amino acid sequence as shown in any one of SEQ ID NOs.3-71.

7. A complex comprising the engineered G protein coupled receptor according to any one of claims 1 to 6, and a second part of a fluorescent protein, wherein the first part of the fluorescent protein is combined with the second part of the fluorescent protein.

8. The complex according to claim 7, wherein the second portion of the fluorescent protein comprises the 1st to 9th β-sheets, the 3rd to 11th β-sheets, the 1st to 7th β-sheets or the 5th to 11th β-sheets of the fluorescent protein.

9. The complex according to claim 7 or 8, wherein the second part of the fluorescent protein comprises the amino acid sequence as shown in SEQ ID NO.2 or an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% sequence identity with the amino acid sequence as shown in SEQ ID NO.

2.

10. The complex according to any one of claims 7-9, wherein the complex further comprises a clamping protein, which binds to the first part of the fluorescent protein and the second part of the fluorescent protein.

11. The complex according to claim 10, wherein the clamping protein comprises an amino acid sequence as shown in any one of SEQ ID NOs.72-74, SEQ ID NO.82, or an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% sequence identity with the amino acid sequence as shown in any one of SEQ ID NOs.72-74, SEQ ID NO.

82.

12. The complex according to claim 11, wherein the 186th position of the clamping protein is cysteine, and the position is based on SEQ ID NO.

72.

13. The complex according to claim 11 or 12, wherein a disulfide bond is formed between the amino acid at position 186 of the clamping protein and the amino acid at position 34.51 of the Ballesteros-Weinstein numbering system in the second intracellular loop.

14. Use of the engineered G protein-coupled receptor according to any one of claims 1 to 6 or the complex according to any one of claims 7 to 13 in ligand affinity determination, drug discovery or screening.

15. Use of the complex according to any one of claims 7 to 13 in analyzing three-dimensional structures using cryo-electron microscopy.