Genomics knowledge graph construction method and system

By combining character slices and chromosome segment numbers from gene ID set data, the problem of cross-species homologous gene identification was solved, achieving highly consistent construction of genomics knowledge graphs and improving the accuracy and completeness of semantic connections between genes and phenotypes.

CN121303293APending Publication Date: 2026-01-09NEIJIANG NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511506772.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-21
Publication Date
2026-01-09

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately identify homologous genes across species when constructing genomics knowledge graphs. They also lack in-depth analysis of the relationship between local variation features in gene sequences and chromosomal locations, leading to semantic ambiguity and insufficient label consistency, which makes it difficult to support high-quality semantic modeling and reasoning.

Method used

By using a set of gene numbering data based on species, character slicing is performed to extract prefix identifiers. This is combined with chromosome segment numbering to determine attribution, generating a cross-species gene entity node set. Furthermore, semantic normalization of the tags is achieved by judging the consistency of tag source frequency and field meaning, thus establishing a semantic connection between genes and phenotypes.

Benefits of technology

It enables accurate identification and aggregation of homologous genes across species, improves the structural continuity and semantic coherence between genes, enhances the semantic connection accuracy and contextual integrity between genes and phenotypes, and forms a set of map fragments with structural consistency and semantic coherence.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121303293A_ABST
    Figure CN121303293A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of bioinformatics, in particular to a genomics knowledge graph construction method and system.The genomics knowledge graph construction method comprises the following steps that gene numbers are obtained, prefix identifiers are extracted, affiliation is judged in combination with chromosome numbers, homologous number pairs are recognized, a connection edge set is generated, tags are normalized to form main semantics, symptom entries are extracted, and semantic paths are established; and constructing a graph node to obtain a structure fragment set. According to the method, the prefix identification is extracted through the character slices, affiliation judgment is carried out in combination with the chromosome segment numbers, accurate recognition and aggregation of cross-species homologous genes can be achieved, and connectable gene pairs are screened in a combined judgment mode of the prefix character set and the chromosome segment sequence; coupling labeling of structural continuity and semantic continuity between genes is achieved, a semantic path is established through logic chain mapping of co-occurrence symptom entries in literatures, and high-consistency structured expression of cross-species gene functional characteristics and phenotypic expressions is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of bioinformatics technology, and in particular to a method and system for constructing genomics knowledge graphs. Background Technology

[0002] Bioinformatics is an interdisciplinary field that combines knowledge from biology, computer science, and mathematics. It aims to acquire, store, analyze, and interpret bioinformatics data using computational methods. Its core aspects include genome sequence alignment and annotation, prediction of biomolecular structure and function, evolutionary relationship analysis, biological network modeling and simulation, and knowledge mining and integration. This field systematically covers the entire process from raw biological data acquisition to structured processing, semantic understanding, and application support, and is widely used in disease research, drug development, and genetic analysis. Traditional genomics knowledge graph construction methods involve extracting concepts and constructing relationships from genomic information through manual annotation or rule-based methods to form a data network with semantic structure. This primarily addresses the problem of unifying the structure and representing the relationships among multi-source, heterogeneous genomic data, such as gene function annotation information, sequence feature description information, protein interaction information, and disease association information. Traditional methods employ semantic extraction based on keyword matching, entity alignment strategies based on preset templates, and methods that rely on manually defined relationship rules to generate graph relationships for this purpose.

[0003] Existing technologies rely on manual annotation or fixed rules for concept extraction and relationship construction, lacking in-depth analysis of the relationship between local variation features in gene sequences and chromosome positions. This makes it difficult to accurately identify the affiliation of homologous genes between species. Keyword matching and template alignment methods lack dynamic adaptability to the semantic expression of tags, easily leading to semantic ambiguity or unification failure. When faced with heterogeneous tags from complex sources, manual rules cannot cover all field representations, reducing tag consistency and system scalability. In knowledge graph construction, there is a lack of symptom-term linkage mechanisms based on real semantic chains, resulting in insufficient connection between genes and phenotypes, making it difficult to support high-quality semantic modeling and reasoning. Summary of the Invention

[0004] To address the technical problems existing in the prior art, embodiments of the present invention provide a method for constructing a genomics knowledge graph, including the following steps: To achieve the above objectives, the present invention adopts the following technical solution: a method for constructing a genomics knowledge graph, comprising the following steps: S1: Based on the numbered set data of species genes, perform numbered character slicing operation, extract prefix identifiers, perform one-to-one comparison of prefix fragments according to the standard prefix structure template, and record the offset character set at the position of character deviation in the comparison results to obtain the cross-species gene entity node set; S2: Call the assigned number pairs in the cross-species gene entity node set, sort them according to the chromosome segment number to which the entity belongs, and combine the number of prefix similar characters between each group of entities with the chromosome segment number order to obtain the gene number recombination connection edge set; S3: Use the gene number to recombine the connection edge set, extract the corresponding functional annotation tag set, and perform tag source frequency recording according to the source field, perform sorting operation on the tags, and form a gene entity master tag list; S4: Call the main tag list of the gene entity, obtain the set of symptom term groups co-occurring in the corresponding literature, arrange the symptom terms in the order of co-occurrence to generate a linear chain, extract any combination of core terms and modifier terms in the chain, mark the positions where the logical chain terms and tag fields have semantic overlap, and obtain the set of gene phenotype semantic mapping paths.

[0005] As a further embodiment of the present invention, the cross-species gene entity node set includes entity number, corresponding chromosome segment number, and prefix offset character position mapping; the gene number recombination connection edge set includes recombination connection pairs, linkable markers, prefix character matching information, and chromosome sequence continuity identifiers; the gene entity main label list includes main semantic description content, source field frequency records, semantically normalized label groups, and label repetition filtering results; and the gene phenotype semantic mapping path set includes core terms, modifier term combination groups, semantic overlap identifiers, and field-level mapping corresponding points.

[0006] As a further aspect of the present invention, the specific steps of S1 are as follows: S101: Based on the numbered set data of species genes, slice the numbered character sequence, extract the starting fragment of the character sequence as a prefix identifier, compare the extracted prefix identifier with the preset character fragment in the standard prefix structure template one by one, compare the positions where there are character differences during the matching process, record the offset characters corresponding to the positions, and generate a character offset mapping set. S102: Call the offset character positions recorded in the character offset mapping set, obtain the chromosome segment number information corresponding to the numbered entry, retrieve the distribution of each offset character position in the chromosome segment number, and perform cross mapping between the offset character position and the chromosome segment number. Identify the number combinations that have the same offset characteristics and belong to the same chromosome segment among the differentiated species through the mapping relationship, and generate a homologous number mapping relationship table. S103: Call the homologous number pairs and corresponding chromosome segment numbers recorded in the homologous number mapping table, aggregate the homologous number and the chromosome segment information to which each group belongs, and generate a cross-species gene entity node set.

[0007] As a further aspect of the present invention, the specific steps of S2 are as follows: S201: Call the assigned number pairs in the cross-species gene entity node set, extract the corresponding chromosome segment numbers, sort the entities in the number pairs in ascending order according to the chromosome segment numbers, and keep the original pairing relationship of the number content included in the sorted entity pairs unchanged, and generate a chromosome order arrangement reference table. S202: Call the chromosome sequence arrangement comparison table, count the number of similar characters for the prefix characters of the two numbers in the same number pair, calculate the similarity value of the prefix characters, and jointly judge the number of similar prefix characters of each number pair with the chromosome segment number sequence information. If both contents satisfy the conditions of character group overlap and segment sequence continuity, then assign a link mark to the number pair and generate an entity link mark list. S203: Call the entity pair numbers that have been assigned link tags in the entity link tag list, treat the corresponding entity pairs as connected edges in the graph structure, construct the edge information set according to the original number pair relationship, and generate the gene number recombination connection edge set.

[0008] As a further aspect of the present invention, the specific steps of S3 are as follows: S301: Call the numbering information of entity pairs in the gene number recombination connection edge set, extract the associated functional annotation tag set, retrieve the source field content corresponding to the tag, count the source field of the tag, sort it in descending order according to the source frequency, and generate the main tag placement candidate set. S302: Call the main tag placement candidate set, perform tag intersection filtering on entity pairs with related relationships, retain the tag groups that are repeated between differentiated entities, and for tag pairs with the same field meaning but different text descriptions, perform character-by-character comparison and semantic aggregation operations based on the semantic structure of the tag description field to generate a tag semantic normalization matching table. S303: Call the tag semantic normalization matching table to replace the original tag content between differentiated entities, use the normalized tag as the unified main semantic description of the entity, construct the mapping pair between the corresponding entity number and the main tag, and generate a list of gene entity main tags.

[0009] As a further aspect of the present invention, the specific steps of S4 are as follows: S401: Call the main tag list of the gene entity, retrieve the set of symptom term groups that co-occur in the associated literature, arrange them according to the order of the symptom terms in the literature, construct the sequential sequence of symptom terms, and maintain the original adjacency relationship in the text to generate a linear chain sequence of symptom terms; S402: Call the continuous entries in the linear chain sequence of the symptom entries, extract the core entries and modifying entries with physiological meaning or clinical orientation respectively, construct multiple core and modifying combination groups, and perform field-level one-to-one mapping judgment on the functional category field, response participation field, and organ field in the gene entity main label list to generate a set of semantic overlap annotations for the tag fields; S403: Call the positions marked as having semantic overlap in the label field semantic overlap annotation set, combine them with the sequential structure in the linear chain sequence of symptom terms, construct a semantic path with logical direction, and combine the symptom term sequence corresponding to the semantic path with the entity label field as path elements to construct a semantic correspondence pathway between entities and phenotypes, and generate a gene phenotype semantic mapping path set.

[0010] As a further aspect of the present invention, in the process of constructing the sequential sequence of symptom entries, only symptom entries with an adjacent distance of no more than three entries in the same related document are retained as components of the sequential sequence that can be constructed; In the process of extracting core terms and modifier terms with physiological meaning or clinical relevance, only symptom terms that appear at least twice in the linear chain sequence of symptom terms are extracted, and the part of speech of the symptom terms is used as an auxiliary criterion for determining whether they are modifiers or core terms. During the field-level one-to-one mapping judgment process, when the core term in the symptom term and the root word matching degree of the functional category field in the gene entity main label list are not less than 80%, it is determined to be semantic overlap. In the process of constructing a semantic path with logical direction, if the interval between two positions marked as having overlapping fields in the linear chain sequence of the symptom terms exceeds five, they will not be included in the generation range of the semantic path.

[0011] As a further aspect of the present invention, the method further includes step S5: S5: Using the gene phenotype semantic mapping path set, map and combine the symptom word chain sequence and corresponding gene main label information contained in the path, extract each pair of gene and symptom combinations in the path and establish a node structure, combine the nodes of the path in the same entity into graph semantic nodes according to the main label normalization logic, and obtain a set of knowledge graph structure fragments. The knowledge graph structure fragment set includes graph semantic nodes, gene-symptom combination pairs, and path node normalization structures.

[0012] As a further aspect of the present invention, the specific steps of S5 are as follows: S501: Call the symptom term chains and corresponding gene main label field contents included in each path of the gene phenotype semantic mapping path set, establish a one-to-one mapping combination relationship with the main label information according to the order of the symptom term chains in the path, regard each pair of combinations as an independent semantic unit, and generate a set of gene and symptom semantic combination pairs. S502: Call each pair of combinations in the gene and symptom semantic combination set, construct a node entity structure based on the gene main label content included in the combination, represent each pair of combinations as a set of node configurations with semantic connection relationship, and maintain the original order relationship within the path to generate a path node structure set. S503: Call the node configurations belonging to the same gene entity in the path node structure group, and perform node aggregation according to the main label unification logic marked by the gene entity in the gene entity main label list to generate a knowledge graph structure fragment set.

[0013] A genomics knowledge graph construction system, including: The number parsing module obtains the species gene number set data, calls the character content in the number field, compares the character position of the number slice with the standard template field in turn, extracts the corresponding character at the position where there is a character offset and records it, performs mapping judgment on the offset position of the number character slice and the chromosome segment number, and generates a cross-species gene entity node set. The homology comparison module, based on the cross-species gene entity node set, calls the assigned number pairs, performs the same character quantity extraction on the character prefix fragment of each number, obtains the same character quantity, and then performs combination judgment according to the sorting result of chromosome segment number to generate gene number recombination connection edge set. The connection reconstruction module calls the gene number to reorganize the connection edge set, extracts the functional annotation tag content in the annotation field, records the source frequency of the tag source field, performs a sorting operation on the tag set according to the source frequency, and generates a list of gene entity master tags. The main tag classification module uses the gene entity main tag list to collect gene literature content corresponding to the tag entities, extracts symptom entries that co-occur with the tags in the literature paragraphs, arranges the entries according to the order of their appearance in the original text to construct a linear combination chain, divides the core entries and modifier entries into the combination group from each chain, and generates a gene phenotype semantic mapping path set. The semantic mapping module uses the gene phenotype semantic mapping path set to extract the corresponding tag field information and symptom chain content, constructs the node path group between gene number and symptom entry, merges the path nodes under the same tag field, and generates a set of knowledge graph structure fragments.

[0014] Compared with the prior art, the advantages and positive effects of the present invention are as follows: In this invention, prefix identifiers are extracted by character slicing and combined with chromosome segment numbers for attribution determination, enabling accurate identification and aggregation of homologous genes across species. Connectable gene pairs are screened by combining prefix character groups with chromosome segment order, achieving coupled labeling of structural continuity and semantic coherence between genes. Semantic normalization of tags is completed by judging the consistency of tag source frequency and field meaning, making the main semantic identification between heterogeneous data more stable and consistent. Semantic paths are established by mapping the logical chains of co-occurring symptom entries in the literature, improving the accuracy and contextual completeness of semantic connections between genes and phenotypes. Tag normalization and integration of path nodes form a set of map fragments with structural consistency and semantic coherence, achieving a highly consistent structured expression of gene functional characteristics and phenotypic performance across species. Attached Figure Description

[0015] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0016] Figure 1 This is a schematic diagram of the steps of the present invention; Figure 2 This is a detailed schematic diagram of S1 of the present invention; Figure 3 This is a detailed schematic diagram of S2 of the present invention; Figure 4 This is a detailed schematic diagram of S3 of the present invention; Figure 5 This is a detailed schematic diagram of S4 of the present invention; Figure 6 This is a detailed schematic diagram of S5 of the present invention; Figure 7 This is a system module diagram of the present invention. Detailed Implementation

[0017] The technical solution of the present invention will now be described with reference to the accompanying drawings.

[0018] In embodiments of the present invention, words such as "exemplarily," "for example," etc., are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the word "exemplary" is intended to present the concept in a concrete manner. Furthermore, in embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one.

[0019] In the embodiments of this invention, the terms "image" and "picture" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning. Similarly, the terms "of," "corresponding (relevant)," and "corresponding" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning.

[0020] In this embodiment of the invention, sometimes a subscript such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meaning they express is the same.

[0021] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.

[0022] Please see Figure 1 This invention provides a method for constructing a genomics knowledge graph, including the following steps: S1: Based on the number set data of species genes, perform number character slicing operation, extract prefix identifiers, perform one-to-one comparison of prefix fragments according to the standard prefix structure template, and record the offset character set where there is character deviation in the comparison results. Obtain the chromosome segment number corresponding to the numbered item, perform cross-mapping classification judgment on the offset character position and chromosome segment number, and combine the numbered items identified as homologous number pairs with chromosome numbers to form entity classification pairs, and obtain the cross-species gene entity node set; S2: Call the assigned number pairs in the cross-species gene entity node set, sort them according to the chromosome segment number to which the entity belongs, combine the number of prefix similar characters between each group of entities with the chromosome segment number order to judge, generate linkable tags for entity pairs that meet the conditions of sequential continuity and character group overlap, and obtain the gene number recombination connection edge set. S3: Use gene numbering to recombine and connect edge sets, extract the corresponding functional annotation tag set, and perform tag source frequency recording according to the source field. Sort the tags, filter the tag set with repetition as the main tag placement candidate, perform semantic normalization processing on tags with consistent field meanings in the main tag candidate group among differentiated entities, and use the normalization result as the main semantic description content of the entity to form a gene entity main tag list. S4: Call the gene entity main label list, obtain the set of symptom term groups co-occurring in the corresponding literature, arrange the symptom terms in the co-occurrence order to generate a linear chain, extract any core term and modifier term combination group in the chain, perform field-level mapping judgment on the functional category field, response participation field, and organ field in the entity main label of the combination group, mark the positions where the logical chain terms and label fields have semantic overlap, and obtain the gene phenotype semantic mapping path set; S5: Utilize the gene phenotype semantic mapping path set, map and combine the symptom term chains contained in the path according to their order and corresponding gene main label information, extract each pair of gene and symptom combinations in the path and establish a node structure, combine the nodes of the path in the same entity into graph semantic nodes according to the main label normalization logic, and obtain a set of knowledge graph structure fragments. The cross-species gene entity node set includes entity ID, corresponding chromosome segment ID, and prefix offset character position mapping; the gene ID recombination link set includes recombination link pairs, linkable markers, prefix character matching information, and chromosome sequence continuity identifiers; the gene entity main label list includes main semantic description content, source field frequency records, semantically normalized label groups, and label repetition filtering results; the gene phenotype semantic mapping path set includes core terms, modifier term combinations, semantic overlap identifiers, and field-level mapping correspondence points; and the knowledge graph structure fragment set includes graph semantic nodes, gene and symptom combination pairs, and path node normalization structure.

[0023] Please see Figure 2 The specific steps of S1 are as follows: S101: Based on the numbered set data of species genes, slice the numbered character sequence, extract the starting fragment of the character sequence as a prefix identifier, compare the extracted prefix identifier with the preset character fragment in the standard prefix structure template one by one, compare the positions where there are character differences during the matching process, record the offset characters corresponding to the positions, and generate a character offset mapping set. Obtain the set of gene number sequences of the target species, such as the human gene number "HSA001", the mouse gene number "MMU001", etc. The character sequence can be in the form of "HSA00123ABC" or "MMU00456XYZ". Perform slicing operation on the number character sequence, set a fixed slice length, and set the first 6 characters of the extracted number prefix character as the prefix identifier (such as "HSA001" or "MMU004"). The string truncation function can be used to achieve this, such as prefix=code[6]. The slice length set here should be matched according to the length of the standard prefix template. The extracted prefix identifier is compared with the preset standard prefix template (such as "AAA000" or "BBB111"). When comparing, the character matching method is used. Record the inconsistent character sites. Set the difference sites after comparing "HSA001" and "AAA000" to be the 1st, 2nd, and 3rd characters. The 3rd, 4th, 5th, and 6th bits are recorded as "H≠A", "S≠A", "A≠A", "0≠0", "0≠0", and "1≠0" respectively. A character offset mapping set is constructed. For each differing character, a record such as {position:1, deviation:'H→A'} is generated. This operation can be achieved by looping through the string. Each offset mapping record contains three parameters: character position, original character, and target character. At the encoding level, it can be structured as a dictionary, such as offset-map=[{"pos":1, "orig":"H", "target":"A"}, ...], for quick lookup later. In the example, if the operation is performed on "MMU00456XYZ", the prefix "MMU004" is extracted and compared with "AAA000" to generate a character offset mapping set.

[0024] S102: Call the offset character positions recorded in the character offset mapping set to obtain the chromosome segment number information corresponding to the numbered entry. For each offset character position, retrieve the distribution in the chromosome segment number and perform cross mapping between the offset character position and the chromosome segment number. Identify the number combinations that have the same offset characteristics and belong to the same chromosome segment among the different species through the mapping relationship, and generate a homologous number mapping relationship table. If the extracted offset positions are from the 1st to the 6th position, obtain the chromosome segment number information corresponding to each prefix number in the numbered entries. Set the chromosome segment corresponding to "HSA00123ABC" to chr3:12000-15000, and "MMU00456XYZ" to chr3:12500-15500. The chromosome segment numbers here can be obtained through species databases such as Ensembl and NCBI platform interfaces. For each offset character position, retrieve the distribution of the corresponding character in the chromosome segment number. Operationally, this can be achieved by establishing a position-chromosome segment index dictionary to map the relationship, such as position-map={1:[chr3],2:[chr3],...}, which will show the distribution of positions in multiple species. Cross-mapping is performed, setting human numbers as "H" and mouse numbers as "M" at the same position, but both belonging to the chromosome segment chr3. This determines that although the offset characters differ between species, they belong to the same segment. This mapping relationship is stored in a homologous number mapping table, with the structure [{"human":"HSA00123ABC", "mouse":"MMU00456XYZ", "chrom":"chr3"}]. This constructs a cross-species offset feature mapping relationship. In practical applications, if a third species, such as "rat number RNO003", is introduced, its prefix "RNO003" has the same offset position as "HSA001" and also maps to the chr3:13000-16000 region. In this case, it is also added to the homologous combination, generating a homologous number mapping table.

[0025] S103: Call the homologous number pairs and corresponding chromosome segment numbers recorded in the homologous number mapping table, aggregate the homologous number and the chromosome segment information to which each group belongs, and generate a cross-species gene entity node set; Aggregate each homologous number group, such as "{HSA00123ABC,MMU00456XYZ,RNO00378DEF}", with the shared chromosomal segment chr3. Gene entity nodes are generated after aggregating the numbers. Each entity node is uniquely identified by the number combination, with the node ID set to "Entity_chr3_001". Node attributes include the relevant number and corresponding species information, original offset position and character difference information, chromosomal segment coordinates, etc. The structured storage format can be JSON, such as {"id":"Entity-chr3-001","species":["HSA","MMU","RNO"]],"o ffsets":[1,2,3],"region":"chr3:12000-16000"}, the node set is the basic building block of the graph database, used to establish the gene structure map connection relationship between species. In actual operation, the starting number can be selected as "MMU00456XYZ", and "HSA00123ABC" can be matched by offset comparison in the chr3 segment to which it belongs, and then extended to "RNO00378DEF", and aggregated to form the node "Entity-chr3-001". At the same time, the node has an additional field to record the offset character position, such as offset:1→M→H→R, etc., to generate a cross-species gene entity node set.

[0026] Please see Figure 3 The specific steps of S2 are as follows: S201: Call the assigned number pairs in the cross-species gene entity node set, extract the corresponding chromosome segment numbers, sort the entities in the number pairs in ascending order according to the chromosome segment numbers, and keep the original pairing relationship of the number content included in the sorted entity pairs unchanged, and generate a chromosome order arrangement reference table. The retrieval node set retrieves the chromosome segment number information corresponding to each pair of numbers. For the number pairs "HSA00123ABC" and "MMU00456XYZ", the corresponding chromosome numbers and chromosome start and end positions are extracted respectively. "HSA00123ABC" is located in the 12000-15000 segment of chromosome 3, and "MMU00456XYZ" is located in the 12500-15500 segment of chromosome 3. After extraction, the number pairs are sorted according to chromosome number and start position. During the sorting process, chromosome numbers are compared first. If the chromosome numbers are the same, the numerical value of the start position is compared. "chr1" is placed in the order of "c". Before "hr2", the number starting at position 10000 is placed before the number starting at position 20000. After the sorting operation is completed, for each pair of numbers, the original pairing relationship remains unchanged, that is, "HSA00123ABC" is still paired with "MMU00456XYZ". Only the numbers are rearranged according to chromosome order in the overall structure. If there are five pairs of numbers, located on chromosomes 1, 2, 3, 3, and X respectively, the sorting result is chromosome 1 first, chromosome 2 second, the two pairs of numbers in chromosome 3 are placed after, and the chromosome X pair is placed last. This sorting operation ensures a consistent basis for segment arrangement in subsequent analysis and generates a chromosome order arrangement reference table.

[0027] S202: Call the chromosome sequence arrangement comparison table, count the number of similar characters for the prefix characters of the two numbers in the same number pair, calculate the similarity value of the prefix characters, and jointly judge the number of similar prefix characters of each number pair with the chromosome segment number sequence information. If both contents satisfy the conditions of character group overlap and segment sequence continuity, then assign a link mark to the number pair and generate an entity link mark list. The similarity value of prefix characters is determined by the formula: ; in, Representative number pair The degree of similarity of the prefix characters, Representative number The The encoded value of the bit prefix character. Representative number The The encoded value of the bit prefix character. The maximum comparison length for the prefix character. For the prefix of the first The importance weight of each character Representative number In the The stability score of the chromosome segment corresponding to the prefix character position. Representative number In the Stability score of the chromosome segment corresponding to the prefix character position; Formula calculation logic: Compare the prefix characters one by one, convert the character differences into numerical distances using absolute values, and then... A matching score is obtained; if they are completely identical, the score is 1, and the greater the difference, the lower the score. The matching score is then multiplied by the importance weight of the corresponding character position. This is used to reflect the differences in matching contribution at different locations, and also introduces a chromosome segment stability score. After averaging, the scores are multiplied to emphasize that characters in stable regions are more representative. The weighted scores of each character position are then summed to obtain the overall similarity value. A higher value indicates a more consistent prefix. Prefix character similarity score is a numerical indicator that measures the similarity of two numbers in the prefix characters. Parameter meaning and calculation logic: : Indicates number and number In the prefix The encoded value (numerical type) corresponding to a digit character, for example, if the prefix character comes from A, C, G, T, or numbers, letters, etc., a mapping can be defined in advance, such as A→1, C→2, G→3, T→4. Numeric characters are mapped according to their numerical values, or a mapping of ASCII encoding minus a constant is used to make the encoded values ​​all positive integers. This mapping is the standard for quantizing non-numerical characters. Example: If the number... The first prefix character is "A" (mapping 1), numbered The first digit is "C" (mapping 2), then ; The maximum length of the prefix character to be compared (positive integer) can be taken as follows: That is, only the first 5 prefix characters are compared. The value is a threshold selected based on the common lengths in the design of the prefix number. In practice, it can be determined based on the statistical distribution of prefix lengths. : No. The importance weight of each character (dimensionless coefficient) is used to assign different importance to the differences between characters at different positions. The sum of the weights can be normalized to 1. For example, if the earlier characters are considered to be more important, we can set... ,make ; :serial number ,serial number In the The "stability score" (numerical, positive real number) of the chromosomal segment corresponding to each prefix character can be obtained through methods such as mutation frequency, regional repetition rate, sequencing coverage, and mutation rate statistics. For example, in past monitoring, for each number, the character mutation rate within each prefix segment was monitored. The stability score is defined as Or a suitable mapping, for example: if the segment variation rate ,but For numbering If the variation rate of the second segment is measured to be 0.05, then ; Threshold / Baseline Value / Coefficient Settings: A threshold needs to be set when subsequently determining "character group overlap". ,like This is considered "character group overlap"; the threshold can be set with reference to the 90th percentile of the similarity distribution in the numbered pairs, setting the similarity values ​​in the sample to be distributed between [0.2, 0.9], and taking... As a benchmark; Set number pairs The first 5 bits of the prefix are mapped as follows:

[0028] The first two characters are set to be completely identical (same mapping value), and the fourth character is slightly different (4 vs 5). The weight is set to... ; Then calculate according to the formula: when : Average stability Multiply by the weighted term: ; : Average stability Multiply by the weighted term: ; : Average stability Multiply by the weighted term: ; : → The entire term is 0 (regardless of stability); : Average stability Multiply by the weighted term: ; Substitute into the formula to calculate: ; Compare the results with the threshold Comparison, because If the characters are considered to overlap, the result can be substituted into the original judgment process, and then the segment order can be used to determine whether to assign a link tag. The advantage of the formula lies in the introduction of weights. and stability score The weighted summation of the product of the three factors results in a greater contribution from characters that appear earlier in the same position and have higher stability in the corresponding region, thus improving the ability to identify truly high-trust matching of similar characters.

[0029] S203: Call the entity pair numbers that have been assigned link tags in the entity link tag list, treat the corresponding entity pairs as connected edges in the graph structure, construct the edge information set according to the original number pair relationship, and generate the gene number recombination connection edge set; The number pairs are treated as edges in a graph structure. If the number pairs “HSA00123ABC” and “MMU00456XYZ” have a link marker, then in the graph structure, this indicates that there is a connected edge between these two numbers. The starting point of the edge is “HSA00123ABC”, and the ending point is “MMU00456XYZ”. This edge is not directional; it only indicates that there is a relationship between the two numbered entities. The set of edges is a list of records consisting of multiple number pairs with link markers. Each record includes additional attribute fields such as the two endpoint numbers of the edge, the corresponding chromosome segment, the number of prefix characters of the same type, and the sorting difference. In the graph structure, multiple number pairs with edges form a connected graph branch. Setting the number “HSA00123ABC” ↔ “MMU00456XYZ” ↔ “RNO00378DEF” indicates that there is a transitive connection relationship between the three numbered entities. Each edge can exist independently or serve as a basic unit in topological structure analysis, resulting in the set of gene number recombination connection edges.

[0030] Please see Figure 4 The specific steps of S3 are as follows: S301: Call the numbering information of entity pairs in the gene number recombination connection edge set, extract the associated functional annotation tag set, retrieve the source field content corresponding to the tag, count the source field of the tag, sort it in descending order according to the source frequency, and generate the main tag placement candidate set. Extract the set of functional annotation tags associated with each pair of numbers. Set the tags associated with the number "HSA00123ABC" to include "protein kinase", "intracellular signal transduction", and "ATP binding", and set the tags associated with the number "MMU00456XYZ" to include "protein kinase activity", "signal transduction", and "ATP binding site". Each tag comes from a specific source field, such as GO (GeneOntology), KEGG pathway, Reactome annotation, UniProt description, etc. The source field corresponding to each tag needs to be extracted and the occurrence frequency of the source field is counted. Set the "GO" tag to appear 340 times, "KEGG" to appear 210 times, and "Reactome" to appear 190 times. Sort the source fields in descending order of occurrence frequency to obtain the source frequency ranking list. That is, the source field "GO" is the source field with the highest annotation frequency. Based on this, it is determined to occupy the dominant position in the tag system. Record the set of the top N high-frequency source fields after sorting to generate the main tag placement candidate set.

[0031] S302: Call the main tag placement candidate set, perform tag intersection filtering on entity pairs with related relationships, retain the tag groups that are repeated between differentiated entities, and for tag pairs with the same field meaning but different text descriptions, perform character-by-character comparison and semantic aggregation operations based on the semantic structure of the tag description field to generate a tag semantic normalization matching table. For each pair of related entities in the edge set, a label intersection screening is required. This involves extracting the label sets associated with the two numbers, calculating the intersection, and setting the label set "HSA00123ABC" to {protein kinase, intracellular signal transduction, ATP binding}, and the label set "MMU00456XYZ" to {protein kinase activity, signal transduction, ATP binding site}. The intersection, after semantic comparison, is "protein kinase" and "ATP binding". Although "intracellular signal transduction" and "signal transduction" are different in wording, they have similar semantics. Consistency should be confirmed through semantic comparison and aggregation analysis. For near-synonymous expressions... The tag pairs undergo semantic structure analysis, which involves comparing each tag pair character by character, recording the degree of character overlap and the location of differences, and combining this with the definition text of the meaning in the tag description field to extract the core semantic components and perform aggregation matching. For example, "protein kinase activity" and "protein kinase" have more than 90% character overlap and both emphasize their association with catalytic phosphorylation reactions in their annotations, thus classifying them as belonging to the same semantic class. Similarly, although "signal transduction" and "intracellular signal transduction" differ significantly in character structure, they are essentially identified as having the same pathway process through the semantic core words "signal" and "transduction / transduction." In this way, a tag semantic normalization matching table is constructed.

[0032] S303: Call the tag semantic normalization matching table to replace the original tag content between differentiated entities, use the normalized tags as the unified main semantic description of the entities, construct the corresponding mapping pair between entity number and main tag, and generate a list of gene entity main tags. In each pair of entity IDs, the original differential labels existing in the normalized matching table need to be replaced with the corresponding unified normalized labels. For example, the original label of ID "HSA00123ABC" is set to "protein kinase", and the original label of the counterpart ID "MMU00456XYZ" is "protein kinase activity". The unified label found through the normalized matching table is "protein kinase", so both labels are replaced with the normalized label. For example, "ATP binding" and "ATP binding site" are uniformly classified as "ATP binding". After the replacement, a mapping relationship between each entity ID and the main label is established. The mapping pair structure is ID → label set. Each ID corresponds to multiple normalized main labels. Set "HSA00123ABC" → {protein kinase, ATP binding, signal transduction} to generate a list of gene entity main labels.

[0033] Please see Figure 5 The specific steps of S4 are as follows: S401: Call the list of gene entity main tags, retrieve the set of symptom term groups co-occurring in the associated literature, arrange them according to the order of symptom terms in the literature, construct the sequential sequence of symptom terms, and maintain the original adjacency relationship in the text to generate a linear chain sequence of symptom terms; Using gene IDs as indexes, the system retrieves contextual content from a literature corpus, focusing on identifying sets of co-occurring symptom terms. For example, in literature describing the gene "HSA00123ABC," terms such as "fever," "muscle aches," and "enhanced inflammatory response" appear. During extraction, the order of terms within the original text paragraphs must be maintained; they should not be disassembled or recombinated. The sequential relationship of their appearance in the text must be recorded to construct a linear chain sequence following the original text's flow. For instance, if the original sentence is "The patient experienced persistent fever, accompanied by muscle aches and enhanced inflammatory response," the linear chain sequence would be {fever → muscle aches → enhanced inflammatory response}. The adjacent order of terms constitutes the basic structure of the semantic chain; adjusting the order or ignoring logical relationships is not allowed. This process is repeated to process symptom groups of terms co-occurring in the literature, resulting in a linear chain sequence of symptom terms.

[0034] S402: Call the continuous terms in the linear chain sequence of symptom terms, extract the core terms and modifier terms with physiological meaning or clinical relevance respectively, construct multiple core and modifier combination groups, and perform field-level one-to-one mapping judgment on the functional category field, response participation field, and organ field in the gene entity main label list to generate a set of semantic overlap annotations for the tag fields; Semantic component analysis is required for each consecutive symptom term in the sequence to extract core terms with physiological or clinical significance, such as "fever," "ache," "swelling," and "difficulty breathing," as well as terms that constitute a modifying relationship, such as "persistent," "severe," "mild," and "acute." During extraction, the semantic role of each term in the sentence is used as the criterion. Verbs and nouns in subject-verb-object structures are used as core terms, and adjectives and adverbs as modifying terms. For example, in "persistent fever," "fever" is the core term, and "persistent" is the modifier, forming the combination {fever: persistent}. Each linear chain sequence can be solved... After extracting multiple such combinations, it is necessary to perform field-level mapping matching between the combinations and the functional category fields (such as "signal transduction", "inflammatory factor secretion"), response participation fields (such as "immune regulation", "apoptosis"), and organ fields (such as "lung", "liver", "muscle") in the gene entity master label list. Compare whether the semantic content of the core term appears in the label field or has a similar expression. Set "muscle" in "muscle soreness" to be mapped to the label organ field "muscle", and "soreness" has semantic overlap with the functional field "inflammation perception", and generate a label field semantic overlap annotation set.

[0035] S403: Call the positions marked in the semantic overlap annotation set of the label fields that have semantic overlap, combine them with the sequential structure in the linear chain sequence of symptom words, construct a semantic path with logical direction, and combine the symptom word sequence corresponding to the semantic path with the entity label fields as path elements to construct a semantic correspondence pathway between entities and phenotypes, and generate a set of gene phenotype semantic mapping paths. It is necessary to filter out combinations that have been marked as having semantic overlap in their fields, and then construct semantic logical paths based on their order of appearance in the original linear chain sequence of symptom terms. The starting point of each path is the first semantically overlapping term combination, and the ending point is the term combination in the sequence that still has an overlap with the entity label. In the sequence {persistent fever → muscle soreness → enhanced inflammatory response}, "fever" is mapped to "heat stress", and "muscle soreness" is mapped to "muscle-inflammatory response". The path structure is "fever → muscle soreness" forming a first-level pointing logic. In the path construction, the position number of each term combination in the sequence must be retained, and a coherent logical link structure must be formed in the path. Each path element consists of two parts: one is the sequence of symptom term combinations, and the other is the corresponding entity label field group, forming a structure such as: path segment 1 = {fever, heat stress}, path segment 2 = {muscle soreness, inflammation perception}, which are then combined to form the mapping path "fever-heat stress → muscle soreness-inflammation perception", generating a set of gene phenotype semantic mapping paths.

[0036] Please see Figure 6 The specific steps of S5 are as follows: S501: Call the symptom term chains and corresponding gene main label fields included in each path of the gene phenotype semantic mapping path set, establish a one-to-one mapping combination relationship with the main label information according to the order of the symptom term chains in the path, treat each pair of combinations as an independent semantic unit, and generate a set of gene and symptom semantic combination pairs. Each symptom term fragment and its corresponding primary label field are extracted from the path. In the path "Persistent Fever → Muscle Soreness → Enhanced Inflammatory Response", "Persistent Fever" corresponds to the primary label field "Heat Stress Response", "Muscle Soreness" corresponds to "Skeletal Muscle Damage-Related Protein Function", and "Enhanced Inflammatory Response" corresponds to "Upregulation of Inflammatory Factors". Following the natural order of the symptom term chain in the path, a one-to-one mapping relationship is established between each symptom term and its corresponding primary label field, generating combinations such as {Persistent Fever, Heat Stress Response}, {Muscle Soreness, Skeletal Muscle Damage-Related Protein Function}, {Enhanced Inflammatory Response, Upregulation of Inflammatory Factors}, etc. Each combination is an independent semantic unit, and each unit consists of a symptom expression and a gene function label, with a clear semantic mapping meaning between clinical phenotype and gene function. The combination unit maintains complete consistency with the order of appearance of the original symptom terms in the path, and the order cannot be adjusted or the combination content cannot be recombined, generating a set of gene and symptom semantic combination pairs.

[0037] S502: Call each pair of gene and symptom semantic combination pairs in the set, construct the node entity structure based on the gene main label content included in the combination, represent each pair of combinations as a set of node configurations with semantic connection relationship, and maintain the original order relationship within the path to generate a set of path node structure groups. Based on the content of the main label fields included in the combination, a node entity structure is constructed. Each combination pair is represented as a node configuration with semantic connection relationship. The node includes two main parts: one is a phenomenon node representing the symptom entry, such as "persistent fever", and the other is a functional node representing the gene main label, such as "heat stress response". The nodes form a one-to-one semantic connection. The link direction can be set from the symptom node to the functional node, representing causal or participation relationship. The sequential relationship of the combination pair in the original path remains unchanged. If the original path is {fever → muscle soreness → increased inflammation}, then the constructed path node structure should be: node 1 (fever → heat stress) → node 2 (muscle soreness → muscle function) → node 3 (increased inflammation → inflammatory response). Multiple nodes are combined in a linear structure to construct a path node structure set.

[0038] S503: Call the node configurations of the path node structure group that belong to the same gene entity, and perform node aggregation according to the main label unification logic marked by the gene entity in the gene entity main label list to generate a set of knowledge graph structure fragments. All elements must be uniformly identified as belonging to the same gene entity. The main label fields are categorized and aggregated using the main label normalization logic recorded in the gene entity's main label list. If the number "HSA00123ABC" is uniformly classified under the "Inflammation Response Regulation" category in the main label list, then node configurations related to the main label logic, such as "fever → heat stress," "muscle soreness → muscle inflammation," and "inflammation enhancement → inflammatory response," all belong to "Inflammation Regulation." After aggregating these node configurations, a unified and normalized graph structure fragment is formed. Nodes within the fragment are connected radially or in a chain-like manner with genes as the core, representing multiple phenotypic symptoms mapped to the gene's main functional unit through different paths. Each fragment uses the gene number as the primary key and includes multiple node configurations, label field groups, symptom chain structures, semantic connection types, node aggregation labels, and other constituent fields, generating a set of knowledge graph structure fragments.

[0039] Please see Figure 7 A genomics knowledge graph construction system, including: The number parsing module obtains the species gene number set data, calls the character content in the number field, compares the character position of the number slice with the standard template field in turn, extracts the corresponding character at the position where there is a character offset and records it, performs mapping judgment on the offset position of the number character slice and the chromosome segment number, and generates a cross-species gene entity node set. The homology comparison module is based on a cross-species gene entity node set. It calls the assigned number pairs, extracts the number of similar characters from the character prefix fragments of each number, and then combines and judges them according to the sorting results of chromosome segment numbers to generate a gene number recombination connection edge set. The connection reconstruction module calls the gene number to recombine the connection edge set, extracts the functional annotation tag content from the annotation field, records the source frequency of the tag source field, performs a sorting operation on the tag set according to the source frequency, and generates a list of gene entity master tags. The main tag classification module uses a list of gene entity main tags, collects gene literature content corresponding to the tag entities, extracts symptom terms that co-occur with the tags from the literature paragraphs, arranges the terms in the order of their appearance in the original text to construct a linear combination chain, divides the core terms and modifier terms combination groups from each chain, and generates a set of gene phenotype semantic mapping paths. The semantic mapping module utilizes the gene phenotype semantic mapping path set to extract the corresponding label field information and symptom chain content, constructs node path groups between gene numbers and symptom entries, merges path nodes under the same label fields, and generates a set of knowledge graph structure fragments.

[0040] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for constructing a genomics knowledge graph, characterized in that, Includes the following steps: S1: Based on the numbered set data of species genes, perform numbered character slicing operation, extract prefix identifiers, perform one-to-one comparison of prefix fragments according to the standard prefix structure template, and record the offset character set at the position of character deviation in the comparison results to obtain the cross-species gene entity node set; S2: Call the assigned number pairs in the cross-species gene entity node set, sort them according to the chromosome segment number to which the entity belongs, and combine the number of prefix similar characters between each group of entities with the chromosome segment number order to obtain the gene number recombination connection edge set; S3: Use the gene number to recombine the connection edge set, extract the corresponding functional annotation tag set, and perform tag source frequency recording according to the source field, perform sorting operation on the tags, and form a gene entity master tag list; S4: Call the main tag list of the gene entity, obtain the set of symptom term groups co-occurring in the corresponding literature, arrange the symptom terms in the order of co-occurrence to generate a linear chain, extract any combination of core terms and modifier terms in the chain, mark the positions where the logical chain terms and tag fields have semantic overlap, and obtain the set of gene phenotype semantic mapping paths.

2. The method for constructing a genomics knowledge graph according to claim 1, characterized in that, The cross-species gene entity node set includes entity ID, corresponding chromosome segment ID, and prefix offset character position mapping; the gene ID recombination link set includes recombination link pairs, linkable markers, prefix character matching information, and chromosome sequence continuity identifiers; the gene entity main label list includes main semantic description content, source field frequency records, semantically normalized label groups, and label repetition filtering results; and the gene phenotype semantic mapping path set includes core terms, modifier term combinations, semantic overlap identifiers, and field-level mapping corresponding points.

3. The method for constructing a genomics knowledge graph according to claim 1, characterized in that, The specific steps of S1 are as follows: S101: Based on the numbered set data of species genes, slice the numbered character sequence, extract the starting fragment of the character sequence as a prefix identifier, compare the extracted prefix identifier with the preset character fragment in the standard prefix structure template one by one, compare the positions where there are character differences during the matching process, record the offset characters corresponding to the positions, and generate a character offset mapping set. S102: Call the offset character positions recorded in the character offset mapping set, obtain the chromosome segment number information corresponding to the numbered entry, retrieve the distribution of each offset character position in the chromosome segment number, and perform cross mapping between the offset character position and the chromosome segment number. Identify the number combinations that have the same offset characteristics and belong to the same chromosome segment among the differentiated species through the mapping relationship, and generate a homologous number mapping relationship table. S103: Call the homologous number pairs and corresponding chromosome segment numbers recorded in the homologous number mapping table, aggregate the homologous number and the chromosome segment information to which each group belongs, and generate a cross-species gene entity node set.

4. The method for constructing a genomics knowledge graph according to claim 3, characterized in that, The specific steps of S2 are as follows: S201: Call the assigned number pairs in the cross-species gene entity node set, extract the corresponding chromosome segment numbers, sort the entities in the number pairs in ascending order according to the chromosome segment numbers, and keep the original pairing relationship of the number content included in the sorted entity pairs unchanged, and generate a chromosome order arrangement reference table. S202: Call the chromosome sequence arrangement comparison table, count the number of similar characters for the prefix characters of the two numbers in the same number pair, calculate the similarity value of the prefix characters, and jointly judge the number of similar prefix characters of each number pair with the chromosome segment number sequence information. If both contents satisfy the conditions of character group overlap and segment sequence continuity, then assign a link mark to the number pair and generate an entity link mark list. S203: Call the entity pair numbers that have been assigned link tags in the entity link tag list, treat the corresponding entity pairs as connected edges in the graph structure, construct the edge information set according to the original number pair relationship, and generate the gene number recombination connection edge set.

5. The method for constructing a genomics knowledge graph according to claim 4, characterized in that, The specific steps for S3 are as follows: S301: Call the numbering information of entity pairs in the gene number recombination connection edge set, extract the associated functional annotation tag set, retrieve the source field content corresponding to the tag, count the source field of the tag, sort it in descending order according to the source frequency, and generate the main tag placement candidate set. S302: Call the main tag placement candidate set, perform tag intersection filtering on entity pairs with related relationships, retain the tag groups that are repeated between differentiated entities, and for tag pairs with the same field meaning but different text descriptions, perform character-by-character comparison and semantic aggregation operations based on the semantic structure of the tag description field to generate a tag semantic normalization matching table. S303: Call the tag semantic normalization matching table to replace the original tag content between differentiated entities, use the normalized tag as the unified main semantic description of the entity, construct the mapping pair between the corresponding entity number and the main tag, and generate a list of gene entity main tags.

6. The method for constructing a genomics knowledge graph according to claim 5, characterized in that, The specific steps of S4 are as follows: S401: Call the main tag list of the gene entity, retrieve the set of symptom term groups that co-occur in the associated literature, arrange them according to the order of the symptom terms in the literature, construct the sequential sequence of symptom terms, and maintain the original adjacency relationship in the text to generate a linear chain sequence of symptom terms; S402: Call the continuous entries in the linear chain sequence of the symptom entries, extract the core entries and modifying entries with physiological meaning or clinical orientation respectively, construct multiple core and modifying combination groups, and perform field-level one-to-one mapping judgment on the functional category field, response participation field, and organ field in the gene entity main label list to generate a set of semantic overlap annotations for the tag fields; S403: Call the positions marked as having semantic overlap in the label field semantic overlap annotation set, combine them with the sequential structure in the linear chain sequence of symptom terms, construct a semantic path with logical direction, and combine the symptom term sequence corresponding to the semantic path with the entity label field as path elements to construct a semantic correspondence pathway between entities and phenotypes, and generate a gene phenotype semantic mapping path set.

7. The method for constructing a genomics knowledge graph according to claim 6, characterized in that, In the process of constructing the sequential sequence of symptom entries, only symptom entries that are no more than three adjacent entries in the same related document are retained as components of the sequential sequence that can be constructed. In the process of extracting core terms and modifier terms with physiological meaning or clinical relevance, only symptom terms that appear at least twice in the linear chain sequence of symptom terms are extracted, and the part of speech of the symptom terms is used as an auxiliary criterion for determining whether they are modifiers or core terms. During the field-level one-to-one mapping judgment process, when the core term in the symptom term and the root word matching degree of the functional category field in the gene entity main label list are not less than 80%, it is determined to be semantic overlap. In the process of constructing a semantic path with logical direction, if the interval between two positions marked as having overlapping fields in the linear chain sequence of the symptom terms exceeds five, they will not be included in the generation range of the semantic path.

8. The method for constructing a genomics knowledge graph according to claim 1, characterized in that, The method further includes step S5: S5: Using the gene phenotype semantic mapping path set, map and combine the symptom word chain sequence and corresponding gene main label information contained in the path, extract each pair of gene and symptom combinations in the path and establish a node structure, combine the nodes of the path in the same entity into graph semantic nodes according to the main label normalization logic, and obtain a set of knowledge graph structure fragments. The knowledge graph structure fragment set includes graph semantic nodes, gene-symptom combination pairs, and path node normalization structures.

9. The method for constructing a genomics knowledge graph according to claim 8, characterized in that, The specific steps of S5 are as follows: S501: Call the symptom term chains and corresponding gene main label field contents included in each path of the gene phenotype semantic mapping path set, establish a one-to-one mapping combination relationship with the main label information according to the order of the symptom term chains in the path, regard each pair of combinations as an independent semantic unit, and generate a set of gene and symptom semantic combination pairs. S502: Call each pair of combinations in the gene and symptom semantic combination set, construct a node entity structure based on the gene main label content included in the combination, represent each pair of combinations as a set of node configurations with semantic connection relationship, and maintain the original order relationship within the path to generate a path node structure set. S503: Call the node configurations belonging to the same gene entity in the path node structure group, and perform node aggregation according to the main label unification logic marked by the gene entity in the gene entity main label list to generate a knowledge graph structure fragment set.

10. A genomics knowledge graph construction system, characterized in that, The system is used to implement the genomics knowledge graph construction method according to any one of claims 1-9, the system comprising: The number parsing module obtains the species gene number set data, calls the character content in the number field, compares the character position of the number slice with the standard template field in turn, extracts the corresponding character at the position where there is a character offset and records it, performs mapping judgment on the offset position of the number character slice and the chromosome segment number, and generates a cross-species gene entity node set. The homology comparison module, based on the cross-species gene entity node set, calls the assigned number pairs, performs the same character quantity extraction on the character prefix fragment of each number, obtains the same character quantity, and then performs combination judgment according to the sorting result of chromosome segment number to generate gene number recombination connection edge set. The connection reconstruction module calls the gene number to reorganize the connection edge set, extracts the functional annotation tag content in the annotation field, records the source frequency of the tag source field, performs a sorting operation on the tag set according to the source frequency, and generates a list of gene entity master tags. The main tag classification module uses the gene entity main tag list to collect gene literature content corresponding to the tag entities, extracts symptom entries that co-occur with the tags in the literature paragraphs, arranges the entries according to the order of their appearance in the original text to construct a linear combination chain, divides the core entries and modifier entries into the combination group from each chain, and generates a gene phenotype semantic mapping path set. The semantic mapping module uses the gene phenotype semantic mapping path set to extract the corresponding tag field information and symptom chain content, constructs the node path group between gene number and symptom entry, merges the path nodes under the same tag field, and generates a set of knowledge graph structure fragments.