A method and system for constructing a knowledge graph for autism spectrum disorder

By constructing an autism spectrum disorder knowledge graph through a hybrid convolutional neural network-long short-term memory network-hidden Markov model, the problems of insufficient information extraction and dynamic updating in existing technologies are solved, efficient and reliable knowledge graph construction is achieved, and the research and diagnosis and treatment level of autism spectrum disorder is improved.

CN120163222BActive Publication Date: 2025-09-19LUZHOU VOCATIONAL & TECHN COLLEGE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510638252.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-09-19
Estimated Expiration
2045-05-19

AI Technical Summary

Technical Problem

When constructing autism spectrum disorder knowledge graphs, existing technologies have problems such as relying on a single model for information extraction, insufficient ability to capture semantic associations, inability to adapt to rapid iteration and dynamic updates, lack of adaptive adjustment mechanisms, and difficulty in handling multi-source heterogeneous data and complex network association analysis.

Method used

A hybrid convolutional neural network-long short-term memory network-hidden Markov model is used for information extraction, a dynamic evolution mechanism of credibility is designed, knowledge topology self-organization is constructed, and a knowledge graph of autism spectrum disorder is generated.

Benefits of technology

It improves the accuracy and efficiency of information extraction, realizes the real-time and reliability of knowledge graphs, enhances the adaptive ability and intelligence level, and promotes the research and diagnosis and treatment of autism spectrum disorders.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163222B_ABST
    Figure CN120163222B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of artificial intelligence, and specifically to a method and system for constructing a knowledge graph for autism spectrum disorder. The method comprises: obtaining multi-source heterogeneous autism spectrum disorder information; constructing an original corpus of autism spectrum disorder; extracting information using a hybrid convolutional neural network-long short-term memory network-hidden Markov model; designing a dynamic evolution mechanism for credibility; constructing and optimizing knowledge topology self-organization; and generating an autism spectrum disorder knowledge graph. The present invention constructs a hybrid model to extract information efficiently and accurately, significantly improving the parsing efficiency of unstructured medical texts; designs a dynamic evolution mechanism for credibility to achieve dynamic iterative optimization of triple credibility; introduces an adversarial verification mechanism to efficiently and accurately identify triple polarity, significantly reducing the need for manual intervention; and constructs a multi-granularity interpretation engine to achieve cross-level association analysis of gene pathways-neural circuits-behavioral phenotypes, thereby improving analysis accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and in particular to a method and system for constructing a knowledge graph for autism spectrum disorder. Background Art

[0002] Autism spectrum disorder is a highly heterogeneous neurodevelopmental disease, and its research involves the multidisciplinary intersection of genetics, neuroscience, and behavior. With the explosive growth of medical data, multi-source heterogeneous data from authoritative literature, electronic medical records, patient records, and the internet have provided a rich information resource for autism spectrum disorder research. However, these data have significant differences in form, structure, and credibility, making knowledge integration more difficult. Although traditional knowledge graph technology can partially solve the association problem of structured data, it has obvious bottlenecks in processing unstructured text, dynamic update mechanisms, and deep semantic analysis, making it difficult to meet the urgent needs of multi-dimensional and cross-scale knowledge integration for autism spectrum disorder.

[0003] Existing methods for constructing knowledge graphs generally face the following problems: First, information extraction relies on a single model, which is insufficient in capturing the semantic associations of unstructured text, resulting in limited accuracy in entity recognition and relationship extraction. Second, credibility assessments often rely on static thresholds or manual annotation, which cannot adapt to the rapid iteration of medical knowledge and the dynamic evolution of evidence, resulting in poor graph timeliness. Third, knowledge topology optimization relies heavily on manual intervention and lacks an adaptive adjustment mechanism, making it difficult to address the dynamic correlation analysis of complex gene-neural-behavioral networks in autism spectrum disorder research. Furthermore, existing methods lack systematic solutions for multi-source data fusion, heterogeneous feature collaborative computing, and multi-level decision path modeling, severely limiting the practical value of autism spectrum disorder knowledge graphs.

[0004] At present, there is not much research on the construction of knowledge graphs for autism spectrum disorders, and there is no specific multi-dimensional, multi-level, and multi-scale method for constructing knowledge graphs for autism spectrum disorders. Summary of the Invention

[0005] In response to the deficiencies in the prior art, the present invention provides a method and system for constructing an autism spectrum disorder knowledge graph.

[0006] In a first aspect, the present invention provides a method for constructing an autism spectrum disorder knowledge graph, comprising the following steps: obtaining multi-source heterogeneous autism spectrum disorder information; using the autism spectrum disorder information to construct an autism spectrum disorder original corpus; based on the autism spectrum disorder original corpus, using a hybrid convolutional neural network-long short-term memory network-hidden Markov model, performing information extraction and obtaining information extraction results; based on the information extraction results, designing a credibility dynamic evolution mechanism; according to the credibility dynamic evolution mechanism, constructing and optimizing knowledge topology self-organization; and generating an autism spectrum disorder knowledge graph through the knowledge topology self-organization. The present invention uses a hybrid convolutional neural network-long short-term memory network-hidden Markov model for information extraction, thereby achieving efficient parsing and precise extraction of complex texts, overcoming the shortcomings of traditional methods in processing unstructured data, and greatly improving the accuracy and efficiency of information extraction; by designing a dynamic evolution mechanism of credibility, real-time evaluation of information credibility is achieved, ensuring the real-time and reliability of the knowledge graph, and solving the problem that traditional knowledge graphs are difficult to adapt to rapid information updates; by constructing and optimizing knowledge topology self-organization, automatic optimization and dynamic adjustment of knowledge structure are achieved, the adaptability and intelligence level of knowledge graphs are improved, and the limitations of relying on manual intervention in the construction process of traditional knowledge graphs are broken through; by generating a knowledge graph for autism spectrum disorders, the research and diagnosis and treatment of autism spectrum disorders are greatly promoted.

[0007] Optionally, the acquisition of multi-source heterogeneous autism spectrum disorder information includes: obtaining structured, semi-structured and unstructured initial information on autism spectrum disorder from authoritative medical books, medical research papers, reports of professional autism research institutions, hospital electronic medical records, patient self-report records and Internet data resources; preprocessing the initial information on autism spectrum disorder to obtain processed multi-source heterogeneous autism spectrum disorder information, wherein the preprocessing includes removing duplicate data, filling missing values, correcting erroneous data and unifying data formats. The present invention obtains initial information from multi-source heterogeneous data to construct a comprehensive data pool covering structured, semi-structured and unstructured data, significantly improving data diversity and comprehensiveness, and providing a rich foundation for knowledge graph construction; by preprocessing the initial information to remove duplicate data, fill missing values, correct erroneous data and unify data formats, efficient integration and cleaning of multi-source heterogeneous data is achieved, and the reliability and availability of the data are greatly improved; by combining multi-source heterogeneous data with preprocessing technology, deep mining and precise processing of autism spectrum disorder information are achieved.

[0008] Optionally, the use of the autism spectrum disorder information to construct an autism spectrum disorder original corpus includes: classifying and storing the pre-processed autism spectrum disorder information according to data type; based on the classified storage, performing natural language processing on text data to form structured text data, the natural language processing including word segmentation, part-of-speech tagging and named entity recognition; integrating the structured text data, semi-structured data and unstructured data to construct an autism spectrum disorder original corpus. The present invention achieves efficient organization and management of multi-source heterogeneous data by classifying and storing the pre-processed autism spectrum disorder information according to data type, breaking through the limitations of traditional methods in data storage and retrieval, and significantly improving the systematicness and scalability of data processing; by performing natural language processing operations such as word segmentation, part-of-speech tagging and named entity recognition on text data, unstructured text is converted into structured text data, solving the inefficiency problem of traditional methods in processing unstructured data, greatly improving the availability and analysis accuracy of text data, and providing high-quality input for knowledge graph construction; by integrating structured text data, semi-structured data and unstructured data, a comprehensive and multi-level autism spectrum disorder original corpus is constructed, realizing deep fusion and unified management of multi-source data, significantly improving the coverage and application value of the corpus, and providing strong data support for the research and diagnosis and treatment of autism spectrum disorders.

[0009] Optionally, the method of extracting information and obtaining information extraction results based on the original corpus of autism spectrum disorder using a hybrid convolutional neural network-long short-term memory network-hidden Markov model includes: establishing a bidirectional gated graph convolutional network, a bidirectional long short-term memory network and an improved hidden Markov-conditional random field joint model based on the original corpus of autism spectrum disorder; and extracting information and obtaining information extraction results using the bidirectional gated graph convolutional network, the bidirectional long short-term memory network and the improved hidden Markov-conditional random field joint model. The present invention establishes a bidirectional gated graph convolutional network and a bidirectional long short-term memory network to achieve efficient capture and modeling of complex semantic relationships in the original corpus of autism spectrum disorders, breaking through the limitations of traditional methods in processing graph structure data, significantly improving the depth and accuracy of information extraction, and providing more accurate semantic support for the construction of knowledge graphs; by improving the hidden Markov-conditional random field joint model, the accuracy of named entity recognition and relationship extraction is greatly improved, providing a more reliable technical guarantee for the automated extraction of autism spectrum disorder knowledge; by combining the bidirectional gated graph convolutional network, the bidirectional long short-term memory network and the improved hidden Markov-conditional random field joint model, the collaborative analysis and joint optimization of multi-source heterogeneous data are achieved, breaking through the performance bottleneck of a single model in information extraction, significantly improving the comprehensiveness and robustness of information extraction, and providing an efficient and intelligent technical path for the construction of autism spectrum disorder knowledge graphs.

[0010] Optionally, the use of the bidirectional gated graph convolutional network, the bidirectional long short-term memory network and the improved hidden Markov-conditional random field joint model to extract information and obtain information extraction results includes: using the bidirectional gated graph convolutional network to classify the preprocessed text data to obtain classification results; based on the classification results, using the bidirectional long short-term memory network to perform enhanced classification to obtain enhanced classification results; based on the enhanced classification results, using the improved hidden Markov-conditional random field joint model to reclassify the text data to obtain reclassification results; based on the reclassification results, executing classification iterations of the text data to complete information extraction and obtain a set of triples, wherein the information extraction includes entity extraction, attribute extraction and relationship extraction. The present invention uses a bidirectional gated graph convolutional network to classify the preprocessed text data, and uses a bidirectional long short-term memory network to perform enhanced classification on the classified text data, thereby achieving deep extraction and efficient classification of multi-level features of text data, breaking through the limitations of traditional classification models in feature extraction capabilities, significantly improving the accuracy and robustness of classification results, and providing high-quality input for subsequent information extraction; using an improved hidden Markov-conditional random field joint model to reclassify text data, it solves the shortcomings of traditional models in long-distance dependence and semantic association, greatly improves the accuracy and consistency of reclassification results, and lays a solid foundation for entity, attribute and relationship extraction; through the reclassification results, the classification iteration of text data is performed, which realizes dynamic optimization of the information extraction process, breaks through the performance bottleneck of traditional static extraction methods, significantly improves the integrity and accuracy of triple sets, and provides an efficient and intelligent technical path for the construction of knowledge graphs for autism spectrum disorders.

[0011] Optionally, the improved hidden Markov-conditional random field joint model satisfies the following expression:

[0012]

[0013] in, The observation sequence The state transition energy function of the hidden state sequence, is the sequence length, is the hidden state transfer matrix, which represents the Status Transfer to Status The probability of is a new position-sensitive feature function, is the original characteristic function, 、 、 is the adaptive weight coefficient, is the total number of position-sensitive eigenfunctions, is the total number of original characteristic functions, 、 、 are the indexes of position-sensitive feature function, original feature function and interactive feature function respectively, is the total number of interactive characteristic functions, The present invention accurately captures the influence of position information on state transition in sequence data by fusing position-sensitive feature functions, improves the recognition accuracy of the model in complex scenarios, and realizes effective modeling of data sequences with strong position dependence; by introducing adaptive weight coefficients, it can dynamically adjust the contribution of each feature function to the state transition energy according to the data characteristics, enhances the flexibility and generalization ability of the model, and enables the model to perform excellent performance in different fields and data sets; by adding interactive feature functions and considering their complex interactive relationships with observation sequences and hidden states, it comprehensively describes the potential laws and patterns in sequence data, improves the model's ability to parse complex structured data, and provides new ideas and methods for processing high-dimensional, nonlinear sequence data.

[0014] Optionally, the designing of a dynamic credibility evolution mechanism based on the information extraction result includes: establishing a multi-dimensional credibility evaluation system based on the information extraction result; and iteratively optimizing triple credibility through a Markov decision process based on the multi-dimensional credibility evaluation system. The present invention introduces a multi-dimensional credibility evaluation system and combines it with the information extraction results to achieve a comprehensive dynamic evaluation of the credibility of triples, significantly improving the accuracy and adaptability of the evaluation, especially showing stronger robustness when processing complex, multi-source heterogeneous data, and providing a more refined and scientific basis for credibility evaluation; through the iterative optimization mechanism of the Markov decision process, the credibility of the triples is dynamically adjusted, and real-time updating and optimization of the credibility is achieved, overcoming the defect that traditional static evaluation methods are difficult to adapt to dynamic changes in data, so that the credibility evaluation can be continuously optimized as the data evolves, greatly improving the timeliness and reliability of the evaluation results; by deeply integrating multi-dimensional evaluation with the Markov decision process, a closed-loop mechanism for the dynamic evolution of credibility is constructed, realizing the automation of the entire process from data extraction to credibility evaluation to optimization, significantly reducing labor costs, while improving evaluation efficiency and consistency, and providing a new technical path for credibility management in large-scale data application scenarios.

[0015] Optionally, the construction and optimization of knowledge topology self-organization according to the dynamic evolution mechanism of credibility includes: developing a heterogeneous graph growth algorithm according to the dynamic evolution mechanism of credibility; constructing knowledge topology self-organization using the heterogeneous graph growth algorithm; introducing an adversarial verification mechanism based on the knowledge topology self-organization to identify the polarity of triples and obtain recognition results; and optimizing the knowledge topology self-organization based on the recognition results. The present invention develops a heterogeneous graph growth algorithm and combines it with a dynamic evolution mechanism of credibility to achieve adaptive growth and expansion of knowledge topology, significantly improving the dynamics and flexibility of the graph, especially showing stronger adaptability when processing multi-source heterogeneous data, and providing a more intelligent and automated solution for the construction of knowledge topology; by introducing an adversarial verification mechanism, the polarity of triples is identified, and accurate analysis and verification of the semantics of triples in the knowledge topology is achieved, overcoming the defect that traditional methods have difficulty in effectively identifying complex semantic relationships, greatly improving the accuracy and reliability of knowledge topology, and providing more solid semantic support for knowledge reasoning and decision-making; by dynamically optimizing the self-organization of knowledge topology, continuous iteration and improvement of knowledge topology is achieved, significantly improving the timeliness and practicality of knowledge topology.

[0016] Optionally, generating an autism spectrum disorder knowledge graph through the knowledge topology self-organization includes: establishing a multi-level decision path through the knowledge topology self-organization; and designing a tracking algorithm based on the multi-level decision path, wherein the tracking algorithm is expressed as follows:

[0017] ,

[0018] in, Indicates gene pathway , neural circuits , behavioral phenotype The correlation between represents the number of associated paths, Indicates the The weight of the path, Indicates the Gene pathways , neural circuits The correlation coefficient between Indicates the Neural circuits on the pathway and behavioral phenotypes The correlation coefficient between 、 、 Represents gene pathways , neural circuits and behavioral phenotypes variance; based on the tracking algorithm, a segmentation algorithm is designed; through the tracking algorithm and the segmentation algorithm, an autism spectrum disorder knowledge graph is generated. The present invention realizes the systematic analysis and modeling of the complex associations between gene pathways, neural circuits and behavioral phenotypes in autism spectrum disorders by establishing a multi-level decision path and combining knowledge topology self-organization, breaking through the limitation of traditional single-level analysis methods that are difficult to capture multi-level interactions, significantly improving the comprehensiveness and depth of the knowledge graph, and providing a more refined theoretical framework for the study of the mechanism of autism spectrum disorders; by designing a tracking algorithm, the correlation between gene pathways, neural circuits and behavioral phenotypes is dynamically calculated based on the multi-level decision path, realizing efficient analysis and tracking of complex biological data, overcoming the defect that traditional methods are difficult to quantify multi-dimensional correlation relationships, greatly improving the accuracy and interpretability of the knowledge graph, and providing more reliable data support for the precise diagnosis and intervention of autism spectrum disorders; by combining the tracking algorithm and the segmentation algorithm, the autism spectrum disorder knowledge graph is dynamically generated, realizing the adaptive update and optimization of the knowledge graph.

[0019] In a second aspect, the present invention provides a system for constructing a knowledge graph for autism spectrum disorders, which includes an input device, a processor, an output device, and a memory, wherein the input device, the processor, the output device, and the memory are interconnected, wherein the memory is used to store a computer program, the computer program includes program instructions, the processor is configured to call the program instructions, and the system uses the method for constructing a knowledge graph for autism spectrum disorders. The system provided by the present invention has high integration and smooth information transmission between various components. By integrating multi-source heterogeneous data, a comprehensive and high-quality original corpus of autism spectrum disorder is constructed, which significantly improves the data coverage and reliability, and provides rich and accurate basic data support for the construction of knowledge graphs. By adopting a bidirectional gated graph convolutional network, a bidirectional long short-term memory network and an improved hidden Markov-conditional random field joint model, multi-level feature extraction of text data, context-aware classification iteration and efficient extraction of entities, attributes and relationships are achieved, which solves the shortcomings of traditional information extraction methods in semantic understanding, long-distance dependence and dynamic optimization, greatly improves the accuracy and intelligence level of information extraction, and provides an efficient and reliable technical path for the construction of knowledge graphs. By designing a dynamic evolution mechanism of credibility and optimizing knowledge topology self-organization, real-time updating, dynamic adjustment and adaptive optimization of the knowledge graph are achieved, which significantly improves the real-time, reliability and application value of the knowledge graph, and provides a comprehensive, accurate and dynamically updated intelligent knowledge platform for the research and diagnosis of autism spectrum disorders. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1This is a flowchart of a method for constructing a knowledge graph for autism spectrum disorders according to an embodiment of the present invention;

[0021] Figure 2 A schematic diagram of a complete autism spectrum disorder knowledge graph constructed in an embodiment of the present invention;

[0022] Figure 3 This is a schematic diagram of the structure of the autism spectrum disorder knowledge graph construction system according to an embodiment of the present invention. DETAILED DESCRIPTION

[0023] Specific embodiments of the present invention will be described in detail below. It should be noted that the embodiments described herein are for illustrative purposes only and are not intended to limit the present invention. In the following description, numerous specific details are set forth to provide a thorough understanding of the present invention. However, it will be apparent to one of ordinary skill in the art that these specific details are not necessarily required to practice the present invention. In other instances, well-known circuits, software, or methods are not specifically described to avoid obscuring the present invention.

[0024] Throughout this specification, references to "one embodiment," "an embodiment," "an example," or "an example" mean that a particular feature, structure, or characteristic described in connection with the embodiment or example is included in at least one embodiment of the present invention. Therefore, appearances of the phrases "in one embodiment," "in an embodiment," "an example," or "an example" in various places throughout this specification are not necessarily all referring to the same embodiment or example. Furthermore, the particular features, structures, or characteristics may be combined in any suitable combinations and / or subcombinations in one or more embodiments or examples. Furthermore, those of ordinary skill in the art will appreciate that the figures provided herein are for illustrative purposes only and are not necessarily drawn to scale.

[0025] See Figure 1 , an embodiment of the present invention provides a method for constructing an autism spectrum disorder knowledge graph, the method comprising the following steps:

[0026] S1. Obtain multi-source heterogeneous information on autism spectrum disorders.

[0027] Among them, S1 includes the following steps:

[0028] S11. Obtain structured, semi-structured, and unstructured initial information on autism spectrum disorders from authoritative medical books, medical research papers, reports from professional autism research institutions, hospital electronic medical records, patient self-report records, and Internet data resources.

[0029] Specifically, authoritative medical books were selected, including the Diagnostic and Statistical Manual of Mental Disorders (DSM-5) and the International Classification of Diseases (ICD); structured information was extracted through manual reading, including the definition, diagnostic criteria, symptoms and epidemiological data of autism spectrum disorder (ASD); and the extracted information was organized into a database format for subsequent processing.

[0030] Furthermore, scientific research papers on ASD were searched through the CNKI academic database; semi-structured information was extracted using text mining tools or manual reading, including the research methods, experimental data, and conclusions in the papers; and the extracted information was organized according to the paper title, author, publication year, journal name, and research content fields.

[0031] Furthermore, pay attention to reports released by the Autism Research Foundation and autism science organizations; download and read the reports, extract structured and semi-structured information, including the latest research progress, treatment strategies and policy recommendations for ASD; organize the extracted information according to the report title, publishing organization, publication date and main content fields.

[0032] Furthermore, we cooperate with multiple hospitals to obtain the electronic medical record data of ASD patients; extract unstructured information in the medical records through data interfaces or data exports, including basic patient information, diagnosis records, treatment records and follow-up records; anonymize the extracted information and store it in a secure database.

[0033] Furthermore, self-reports from ASD patients and their families were collected through online questionnaires and face-to-face interviews. Unstructured information from patients' self-reports, including symptoms, treatment experiences, and life impacts, was organized. The self-reports were converted into text and stored in text files.

[0034] Furthermore, the post and comment data of the Internet platform are crawled, and the Internet platform includes autism-related forums, social media and blogs; natural language processing technology is used to extract unstructured information, including ASD discussion hotspots, patient experiences and treatment recommendations; the extracted information is organized according to the source, release time and content fields, and stored in the database.

[0035] S12. Preprocessing the initial autism spectrum disorder information to obtain processed multi-source heterogeneous autism spectrum disorder information, wherein the preprocessing includes removing duplicate data, filling missing values, correcting erroneous data, and unifying data formats.

[0036] In one embodiment, a data deduplication algorithm is first used to identify and remove duplicate records or data items. For structured data, such as records in a table or database, duplicates are removed by comparing key fields, such as patient IDs and paper titles. For unstructured data, such as text files or social media posts, duplicate content is identified and removed by calculating text similarity.

[0037] Furthermore, appropriate imputation methods are selected based on the nature and missingness of the data, including mean imputation, median imputation, mode imputation, interpolation imputation, and imputation methods based on machine learning prediction models. For numerical data such as age and weight, the mean or median is used for imputation; for categorical data such as gender and diagnosis type, the mode is used for imputation; for time series data such as follow-up records, interpolation is used for imputation; for complex unstructured data such as missing information in patient self-report records, machine learning prediction models are used for imputation.

[0038] Furthermore, data validation rules and machine learning algorithms are used to identify and correct errors in the data. For structured data, validation rules, such as age ranges and gender values, are set to automatically detect and correct errors. For semi-structured or unstructured data, natural language processing technology is used to identify and correct errors, such as spelling errors and unclear semantics.

[0039] Furthermore, based on the data's purpose and analysis requirements, a unified data format and storage structure were designed. Structured data was converted to the standard CSV file format; semi-structured data was converted to XML; and unstructured data was converted to text and stored in a unified database. Furthermore, all data adhered to the same encoding standards and date formats.

[0040] S2. Using the autism spectrum disorder information, construct an autism spectrum disorder original corpus.

[0041] In one embodiment, a database system is first used to classify and store the pre-processed autism spectrum disorder information according to the types of text structured data, semi-structured data and unstructured data, and an index or label is created for each type of data to facilitate subsequent retrieval and processing.

[0042] Furthermore, a word segmentation tool is used to segment the text data, remove stop words, and retain meaningful words.

[0043] Furthermore, the words after word segmentation are tagged with parts of speech, such as nouns, verbs, and adjectives. The part of speech tagging helps to understand the role and meaning of the words in the sentence.

[0044] Furthermore, named entities in the text, such as names of people, places, institutions, and professional terms related to ASD, are identified. This helps extract key information from the text and provides a basis for subsequent analysis.

[0045] Furthermore, the text data after word segmentation, part-of-speech tagging and named entity recognition processing is integrated into structured text data, and the structured text data is stored in XML format.

[0046] Furthermore, structured text data, semi-structured data and unstructured data are integrated together to form the original corpus of autism spectrum disorder.

[0047] S3. Based on the autism spectrum disorder original corpus, a hybrid convolutional neural network-hidden Markov model is used to extract information and obtain information extraction results.

[0048] In one embodiment, a bidirectional gated graph convolutional network is established based on the autism spectrum disorder original corpus, and the bidirectional gated graph convolutional network includes an input layer, a convolutional layer, a pooling layer, a fully connected layer, and a SoftMax layer.

[0049] The input layer is responsible for converting raw text data into a format that the model can process.

[0050] Specifically, a pre-trained word vector model, such as Word2Vec, is used to convert each word in the text into a fixed-dimensional word vector. , then a length of The sentence is represented as a matrix .

[0051] Furthermore, each word is regarded as a node in the graph, and the node feature is the corresponding word vector.

[0052] Furthermore, based on the co-occurrence relationship and syntactic dependency between words, an adjacency matrix is ​​constructed. If the word and words If they are adjacent or have a dependent relationship in a sentence, then ,otherwise .

[0053] The convolutional layer is used to extract local features from the graph structure.

[0054] Specifically, a gated recurrent unit is used as a gating mechanism to control the propagation of information in the graph. , the first The feature vector of the layer Calculated by the following formula:

[0055]

[0056] in, For the Nodes in the layer The eigenvector of is the gate activation function, is the domain order, For nodes The adjacency set of For the Layer The weight matrix of the neighborhood, For the Nodes in the layer The eigenvector of For the Layer The bias term of the neighborhood; respectively from the node The information is propagated to the forward and backward neighbors of , and then the feature vectors of the two are concatenated.

[0057] It is important to note that the network contains two convolutional layers:

[0058] The first convolutional layer: input is the word vector matrix and the adjacency matrix , the output is the feature matrix ,in, is the output dimension.

[0059] Second convolutional layer: input is the output of the first convolutional layer and the adjacency matrix , the output is the feature matrix ,in, is the output dimension.

[0060] The lexicalization layer is used to reduce feature dimensions and extract important features.

[0061] Specifically, for each node’s feature vector, the maximum value is taken along the feature dimension. The formula is as follows:

[0062]

[0063] in, Representation node The maximum value of the eigenvector, Indicates that the feature dimension is The output is the feature matrix after pooling .

[0064] It is important to note that the network contains two pooling layers:

[0065] The first pooling layer: the input is the output of the first convolutional layer , the output is the feature matrix .

[0066] Second pooling layer: input is the output of the second convolutional layer , the output is the feature matrix .

[0067] The fully connected layer is used to map the pooled features to the classification label space.

[0068] Specifically, the feature matrix after two layers of pooling and Splicing to get the final feature vector

[0069] Furthermore, the feature vector is mapped to the classification label space:

[0070]

[0071] in, is the weight matrix of the fully connected layer, is the bias term of the fully connected layer, is the output of the fully connected layer.

[0072] The SoftMax layer is used to output classification probabilities.

[0073] Specifically, the output of the fully connected layer Perform SoftMax operation to obtain the probability of each category:

[0074]

[0075] in, For samples Middle category categories The probability of The fully connected layer Category output, is the total number of categories, The fully connected layer Category output.

[0076] Furthermore, the category with the highest probability is selected as the final classification result:

[0077]

[0078] It's important to note that the bidirectionally gated graph convolutional network (BGCN) uses two convolutional layers and two pooling layers to gradually reduce feature dimensionality, significantly reducing computational effort. Pooling extracts important features, reduces redundant information, and accelerates model convergence. The two convolutional layers capture more complex local features, while the two pooling layers enhance the model's generalization capabilities and prevent overfitting. Therefore, the BGCN can efficiently extract features from text data and, through its two convolutional and pooling layers, improve training speed and classification accuracy.

[0079] Furthermore, the bidirectional gated graph convolutional network is used to classify the preprocessed text data to obtain a classification result.

[0080] Furthermore, a bidirectional long short-term memory (LSTM) network is introduced to enhance the classification of the text data. The bidirectional LSTM network not only incorporates the traditional forward and backward LSTM unit structures, but also integrates a temporal attention mechanism, multi-granularity feature fusion, domain knowledge-guided regularization, and a dynamic error correction mechanism. The bidirectional LSTM network simultaneously captures contextual information before and after each time step in the sequence by stacking two LSTM layers together—one that processes the input sequence forward and the other that processes it backward. The outputs of the two LSTM layers are then concatenated.

[0081] Specifically, the classification results output by the bidirectional gated graph convolutional network are used as initial features and input into the bidirectional long short-term memory network. A temporal attention mechanism is introduced into the bidirectional long short-term memory network to dynamically calculate the attention weight of each time step and enhance the capture of key time step features in the classification results.

[0082] Furthermore, a hierarchical feature extraction module is set up: local temporal patterns, such as phrase-level features of symptom descriptions, are captured at the low level of the bidirectional long short-term memory network; and global temporal patterns, such as paragraph-level features of disease progression, are captured at the high level of the bidirectional long short-term memory network.

[0083] Furthermore, a gated fusion unit is used to adaptively weight the fusion of multi-granularity features, enhancing the model's ability to express complex semantics in medical text. Multi-granularity features refer to information features at multiple scales and levels that can be captured when processing sequence data, including temporal and contextual information.

[0084] Furthermore, medical knowledge graph embedding is introduced as external prior knowledge, and the implicit state space of the bidirectional long short-term memory network is constrained by the knowledge alignment loss function to ensure that the features learned by the model are semantically consistent with the medical field.

[0085] Furthermore, a course learning strategy is designed to gradually increase the complexity of training samples, such as from simple symptom descriptions to complex comorbidity relationships, to improve the model's classification performance on samples.

[0086] Furthermore, an error perception module is added to the output layer of the bidirectional long short-term memory network. By calculating the error distribution between the classification results and the true labels, the implicit state update rule of the bidirectional long short-term memory network is dynamically adjusted.

[0087] Furthermore, an adversarial training strategy is adopted to generate adversarial samples to enhance the robustness of the model and reduce classification bias caused by data noise.

[0088] Furthermore, the final hidden state of the bidirectional long short-term memory network is mapped to the classification space through a fully connected layer, and the enhanced classification result is output.

[0089] It should be noted that the bidirectional long short-term memory network introduced in this invention dynamically weights key time step features through the temporal attention mechanism, significantly improving the ability to capture important information in long texts; through hierarchical feature extraction and gated fusion units, it achieves effective fusion of multi-granularity features, enhancing the model's ability to understand complex medical texts.

[0090] Furthermore, based on the enhanced classification results, an improved hidden Markov-conditional random field joint model is established, and the improved hidden Markov-conditional random field joint model satisfies the following expression:

[0091]

[0092] in, The observation sequence Hidden state sequence The state transition energy function, is the sequence length, is the hidden state transfer matrix, which represents the Status Transfer to Status The probability of is a new position-sensitive feature function, is the original feature function, which is used to capture the relationship between the observation sequence and the hidden state. 、 、 is the adaptive weight coefficient, which is used to adjust the contribution of the characteristic function. is the total number of position-sensitive eigenfunctions, is the total number of original characteristic functions, 、 、 are the indexes of position-sensitive feature function, original feature function and interactive feature function respectively, is the total number of interactive characteristic functions, is the interaction feature function.

[0093] The improvements of the improved hidden Markov-conditional random field joint model are as follows:

[0094] The first is the dynamic state transfer matrix. The hidden state transfer matrix of the existing model is It is static and can be improved by The purpose of the improvement is to make the state transition probability depend on the local features of the observation sequence, such as the context information of the current position.

[0095] The second is the nonlinear interaction of the characteristic function. The improved characteristic function The contribution of Modulation, forming nonlinear dependence. The role of improvement is to allow competition or cooperation between features, for example, if Represents the interference characteristics, the denominator can suppress The weight of .

[0096] The third is position-sensitive features. Used to capture modeling positions Observational characteristics of the text, such as the statistical regularity of the beginning or end of a sentence.

[0097] Furthermore, the text data is reclassified according to the improved hidden Markov-conditional random field joint model to obtain a reclassification result.

[0098] Specifically, we first obtain the classified text data based on the bidirectional gated graph convolutional network and the bidirectional long short-term memory network, and express it as a word vector sequence. .

[0099] Furthermore, the trained improved hidden Markov-conditional random field joint model is used to reason about the text data and find the energy function Minimum tag sequence :

[0100]

[0101] Furthermore, the Viterbi algorithm is used to efficiently find the optimal tag sequence.

[0102] Furthermore, the output ,in, For the The reclassification results of the words.

[0103] Furthermore, based on the reclassification results, the bidirectional gated graph convolutional network, the bidirectional long short-term memory network and the improved hidden Markov-conditional random field joint model are used to perform classification iterations to complete information extraction and obtain a set of triples, wherein the information extraction includes entity extraction, attribute extraction and relationship extraction.

[0104] Specifically, first, based on the reclassification results, named entities in the text are confirmed and extracted, and the extracted entities are summarized and organized to form an entity list. The named entities include clinical symptoms, treatment methods and related drugs of autism spectrum disorder.

[0105] Furthermore, attribute information of the identified entities is extracted, and the extracted attribute information is associated with the corresponding entities to form entity-attribute pairs. The attribute information includes a specific description of the rehabilitation treatment method, the severity of clinical symptoms, and the dosage and usage of the drug.

[0106] Furthermore, based on the entity and attribute extraction, the association relationships between entities are extracted, and the extracted association relationships are expressed in the form of triples. The association relationships include the relationship between clinical symptoms and treatment methods, the relationship between drugs and side effects, and the relationship between different clinical symptoms.

[0107] Furthermore, the extracted entities, attributes and relationships are combined into a set of triples to form a knowledge unit and verified to ensure its accuracy and completeness. The verification method includes manual verification and comparison verification with an existing knowledge base.

[0108] Furthermore, the verified triple set is used as new training data and added to the training set.

[0109] Furthermore, the bidirectional gated graph convolutional network model, the bidirectional long short-term memory network and the improved hidden Markov-conditional random field joint model are retrained, and the model parameters are optimized using the new training data.

[0110] Furthermore, the classification, entity extraction, attribute extraction, relationship extraction, and triple generation steps are repeated for multiple iterations. After each iteration, the model performance is evaluated, and the accuracy and recall metrics are recorded. The optimal model parameters are then selected. After multiple iterations of optimization, the final information extraction results are obtained.

[0111] Furthermore, the triple set is stored in the knowledge base to facilitate the subsequent knowledge graph construction and application.

[0112] S4. Based on the information extraction results, a dynamic evolution mechanism of credibility is designed.

[0113] In one embodiment, based on the information extraction results, a multi-dimensional credibility assessment system is established, and the multi-dimensional credibility assessment system includes constructing a dynamic weight matrix by combining the authority of the data source, the number of clinical verifications, and cross-modal consistency.

[0114] Specifically, first evaluate the authority of the data source, such as whether it comes from an authoritative medical journal, clinical trial, or expert consensus, and obtain a corresponding score. The scoring criteria are as follows: High authority, such as top journals: ; Medium authority, such as ordinary journals: ; Low authority, such as non-academic sources: .

[0115] Furthermore, the number of times the knowledge described by the triplet has been verified in clinical practice is evaluated to obtain the corresponding score. The scoring criteria are as follows: High number of verifications : ;Number of verifications : ; Low verification times : .

[0116] Furthermore, the consistency of the triples in different data modalities, such as text and images, is evaluated to obtain corresponding scores. The scoring criteria are as follows: Completely consistent: ; Partial agreement: ; Inconsistency: .

[0117] Furthermore, we define the weight matrix ,in, , , The weights represent the authority of the data source, the number of clinical validations, and cross-modal consistency, respectively. The weights are dynamically adjusted based on the evaluation results and feedback.

[0118] Furthermore, the triple credibility is calculated as follows:

[0119]

[0120] in, is the credibility of the triple, Score the authority of the data source, Score the number of clinical validations, Scoring cross-modal consistency, , , They represent the weights of data source authority, clinical verification times, and cross-modal consistency, respectively.

[0121] Furthermore, based on the multi-dimensional credibility evaluation system, a reinforcement learning optimizer is designed. The reinforcement learning optimizer realizes iterative optimization of triple credibility through a Markov decision process and constructs a reward function based on the clinical diagnosis gold standard.

[0122] Specifically, based on the multi-dimensional credibility assessment system, a Markov decision process is used to model the triple credibility optimization problem. The Markov decision process includes a state space, an action space, a reward function, and a transition probability. The state space represents the credibility assessment result of the current triple. The action space represents the operation on the triple, such as retention, modification, or deletion. The reward function is defined based on the clinical diagnosis gold standard. The reward rules are as follows:

[0123] If the triplet after the operation is consistent with the clinical diagnosis gold standard, the reward is +1;

[0124] If the triplet after the operation is inconsistent with the clinical diagnosis gold standard, a penalty of -1 is applied;

[0125] If the operation is invalid, such as deleting a correct triple, a penalty of -0.5 is applied.

[0126] The transition probability represents the probability of transitioning to a new state after performing an action in a state.

[0127] Furthermore, the Q-learning algorithm is used for optimization, and the Q value update formula is as follows:

[0128]

[0129] in, is the updated Q value, is the Q value before updating, is the learning rate, is the discount factor, For reward, For the new state, During the optimization process, when the Q value converges or reaches the maximum number of iterations, the optimization is stopped.

[0130] Furthermore, a set of triples with significantly improved credibility is output.

[0131] This method realizes the dynamic evolution and optimization of triple credibility through the application of a multi-dimensional credibility evaluation system, a dynamic weight matrix, a reinforcement learning optimizer, a reward function based on the clinical diagnosis gold standard, and a Q-learning algorithm. It not only improves the accuracy and reliability of credibility evaluation, but also makes the evaluation results more in line with actual needs. It has high practical value and promotion prospects.

[0132] S5. Based on the dynamic evolution mechanism of credibility, construct and optimize knowledge topology self-organization.

[0133] Among them, S5 includes the following steps:

[0134] S51. Construct knowledge topology self-organization based on the dynamic evolution mechanism of credibility.

[0135] In one embodiment, a heterogeneous graph growth algorithm was developed based on the dynamic evolution mechanism of credibility and used to construct a self-organizing knowledge topology. This process involves constructing a hierarchical topology structure using the heterogeneous graph growth algorithm combined with symptom severity gradients, and then using a graph neural network propagation algorithm to autonomously cluster knowledge nodes, thereby forming a self-organizing knowledge topology.

[0136] It is important to note that knowledge topology self-organization is a hierarchical knowledge network structure automatically constructed through heterogeneous graph growth algorithms and graph neural network propagation algorithms. The core construction process includes hierarchical topology construction and autonomous clustering of knowledge nodes.

[0137] The effect of this method is to form a knowledge organization system that is hierarchical, adaptive, and does not require human intervention, providing dynamic support for the modeling of complex medical relationships.

[0138] Specifically, in the process of constructing the self-organized knowledge topology, we first define the symptom severity gradient to guide the hierarchical construction of the graph. Let the symptom node set be ,in, For the Symptoms, for which the severity of a symptom is associated ,and A hierarchical function is constructed to map symptom nodes to different levels, where the number of levels is proportional to the severity. The hierarchical function is defined as:

[0139]

[0140] in, is the layered function, For the The severity of the symptoms, For the The severity of the symptoms, is the preset maximum number of levels, Indicates rounding down.

[0141] Furthermore, let the heterogeneous graph be ,in, is a collection of nodes, is the edge set, Is a collection of node types.

[0142] Furthermore, a graph neural network is used to propagate node features and achieve autonomous clustering. The node feature matrix is ,in, is the feature dimension. The update formula of the graph neural network is as follows:

[0143] ,

[0144] in, For nodes In the The hidden state of the layer, For nodes In the The hidden state of the layer, For nodes The neighbor set of For nodes The neighbor set of For the The weight matrix of the layer, For the The bias vector of the layer, is the activation function.

[0145] It should be noted that, compared with the prior art, the present invention implements autonomous clustering by introducing a clustering loss function. The clustering loss function is as follows:

[0146]

[0147] in, is the clustering loss function, is the cluster center set, is the cluster center The expression, is the number of layers of the graph neural network.

[0148] Furthermore, knowledge topology self-organization is formed;

[0149] Furthermore, multi-source heterogeneous knowledge is integrated to preliminarily optimize the knowledge topology self-organization.

[0150] Specifically, let the multi-source heterogeneous knowledge source be , , No. Knowledge Source Provides a knowledge subgraph .

[0151] Furthermore, a fusion function is established, and its relationship is:

[0152]

[0153] in, is the fusion function, is the total number of knowledge sources, For the Knowledge nodes, For the edge sets, For the Node types, For the edge set across knowledge sources, add by calculating the similarity between nodes:

[0154]

[0155] in, is the similarity threshold, 、 Represent two nodes respectively. 、 Represent two different sets of nodes, is the similarity function, defined as:

[0156]

[0157] in, In the graph neural network, the node In the The hidden state of the layer, In the graph neural network, the node In the The hidden state of the layer.

[0158] It should be noted that in the knowledge fusion process, the invention compared to the existing technology lies in: introducing a cross-knowledge source similarity metric, combining the hidden state of the graph neural network to calculate the similarity between nodes, and dynamically adding cross-knowledge source edges. By introducing the cross-knowledge source similarity metric, the similarity between nodes or entities in different knowledge sources can be calculated more accurately. Traditional similarity calculation methods are often limited to a single knowledge source, while the cross-knowledge source similarity metric breaks this limitation and realizes cross-domain similarity comparison. Graph neural networks are good at learning node and edge representations from complex graph structures and can capture the potential relationships between nodes. Therefore, combining the hidden state of the graph neural network to calculate the similarity between nodes can fully utilize the learning ability of the graph neural network and improve the accuracy and robustness of the similarity calculation.

[0159] S52. Based on the heterogeneous graph growth algorithm, an adversarial verification mechanism is introduced to optimize the knowledge topology self-organization. The adversarial verification mechanism detects and repairs logically contradictory edges in the graph by generating an adversarial network.

[0160] In knowledge graphs, the polarity of a triple typically refers to whether the relationship within the triple is positive or negative. A positive triple indicates the existence of a relationship between two entities, while a negative triple indicates the absence of a relationship. This invention, by introducing an adversarial verification mechanism, can identify the polarity of triples.

[0161] In one embodiment, based on the heterogeneous graph growth algorithm, an adversarial network is generated to detect logical contradiction edges. Generate potential contradictory edges, discriminator Determine whether the edge is a logical contradiction. The loss function of the generator is:

[0162]

[0163] Furthermore, the loss function of the discriminator is:

[0164]

[0165] in, is the loss function of the generator, is the loss function of the discriminator, Represents the distribution Sample Seek hope, Represents the distribution Sample Seek hope, is the distribution of real edges, is the distribution of edges generated by the generator, For samples The discriminant output.

[0166] It should be noted that the goal of the generator is to generate some possible logically contradictory edges, which may be negative triplets. The goal of the discriminator is to distinguish between real positive triplets and negative triplets generated by the generator, that is, to play a role in identifying the polarity of triplets.

[0167] Furthermore, for edges that the discriminator identifies as logical contradictions :The present invention innovatively proposes a repair function, which satisfies the following conditions:

[0168]

[0169] in, To repair the function, is the repair threshold, 、 Node and nodes The neighbor set of represents the empty set, situation, that is, no repair in other cases, For nodes The set of neighbor nodes Any node in For nodes The set of neighbor nodes Any node in .

[0170] It's important to note that this invention combines an adversarial verification mechanism with a similarity metric to dynamically detect and repair logically contradictory edges in the graph, a method not previously disclosed in the art. Furthermore, if edge connection errors occur in the self-organization of the knowledge topology, this repair algorithm can also be used to repair them, thereby achieving correct edge connections.

[0171] Furthermore, through the above-mentioned adversarial verification mechanism and repair function, triple polarity is dynamically detected and repaired, thereby optimizing the knowledge topology self-organization.

[0172] S6. Generate an autism spectrum disorder knowledge graph through the knowledge topology self-organization.

[0173] In one embodiment, a multi-granularity interpretation engine is first constructed through the self-organization of the knowledge topology.

[0174] Specifically, a multi-level decision path is established, which includes a gene pathway layer, a neural circuit layer and a behavioral phenotype layer. The gene pathway layer is used to analyze gene-gene interactions and gene-protein interactions and identify key gene pathways. The neural circuit layer constructs a brain region-brain region connection model based on neuroimaging data and identifies key neural circuits. The behavioral phenotype layer is used to analyze behavioral data and identify behavioral phenotypes related to ASD.

[0175] Furthermore, a tracking algorithm is designed to track the association path between gene pathways, neural circuits, and behavioral phenotypes. The tracking algorithm is expressed as follows:

[0176] ,

[0177] in, Indicates gene pathway , neural circuits , behavioral phenotype The correlation between represents the number of associated paths, Indicates the The weight of the path, Indicates the Gene pathways , neural circuits The correlation coefficient between Indicates the Neural circuits on the pathway and behavioral phenotypes The correlation coefficient between 、 、 Represents gene pathways , neural circuits and behavioral phenotypes The variance of .

[0178] It is important to note that compared to existing technologies, this paper, by building a multi-granular interpretation engine, enables cross-level association analysis, from gene pathways to neural circuits to behavioral phenotypes, improving analytical flexibility and accuracy. Based on the knowledge topology, a traceable decision path is established, providing a powerful tool for in-depth understanding of the gene-neuron-behavior relationship.

[0179] Furthermore, individual developmental stages are divided according to age and gender factors, and data from different developmental stages are labeled to ensure the accuracy of personalized diagnosis.

[0180] Furthermore, a segmentation algorithm is designed to segment the knowledge graph in real time according to the individual development stage to generate personalized knowledge graph subgraphs. The segmentation algorithm is expressed as follows:

[0181]

[0182] in, Indicates the individuals (or vertices), Represents an individual and The edges between Represents an edge The weight of Indicates the use of individual and The age factor of the correlation change caused by age difference, Indicates the use of individual and The gender factor of the correlation change caused by gender differences, Represents an individual The total number of edges of The number of directly connected edges, It represents the balance factor used to measure the load balancing degree of each subgraph after segmentation, which is determined by calculating the number of individuals and the number of edges in the subgraph. Represents the load in the subgraph, including the number of individuals and the number of edges.

[0183] Traditional graph segmentation algorithms only consider the graph's topology or edge weights, ignoring individual attributes such as age and gender. Our algorithm, by introducing age and gender factors, comprehensively considers multiple attributes of individuals, making the segmentation results more consistent with the needs of practical application scenarios. It can also dynamically adjust the segmentation strategy based on an individual's real-time data, ensuring that the generated subgraphs always accurately reflect the individual's current state.

[0184] Furthermore, the high-performance graph database Neo4j is used to generate the final autism spectrum disorder knowledge graph, which includes a complete graph and sub-graphs. The graph is mainly composed of entity node information and relationship information. The entity nodes are linked through relationships to form a visual mesh structure and continue to expand outward. As the data source is supplemented and updated, the number of related entity nodes and the number of link relationships will also increase, and the generated autism spectrum disorder knowledge graph will also change dynamically. For the complete graph in the embodiment of the present invention, please refer to Figure 2 Each circle in the diagram represents an entity, and entities are connected by relationships. These relationships include symptoms (SYM), treatments (TR), methods (MT), and confusions (EC). This paper utilizes the high-performance graph database Neo4j to implement knowledge storage and visualization.

[0185] See Figure 3 , Figure 3 This is a schematic diagram of the structure of a system for constructing a knowledge graph for autism spectrum disorders according to an embodiment of the present invention. The system includes an input device, a processor, an output device, and a memory, wherein the input device, the processor, the output device, and the memory are interconnected. The memory is used to store a computer program, which includes program instructions. The processor is configured to invoke the program instructions. The system utilizes the aforementioned method for constructing a knowledge graph for autism spectrum disorders.

[0186] In this embodiment, the input device includes a multi-source data acquisition module for acquiring structured, semi-structured, and unstructured data from authoritative medical books, medical research papers, reports from specialized autism research institutions, hospital electronic medical records, patient self-reports, and internet data resources. The function of the input device is to input multi-source, heterogeneous initial information on autism spectrum disorders into the system through a data interface, ensuring the comprehensiveness and diversity of the data and providing a foundation for subsequent processing.

[0187] Furthermore, the processor is the core computing unit of the system, including a data preprocessing module, an information extraction module and a knowledge graph construction module.

[0188] Specifically, the data preprocessing module is responsible for cleaning the input data, including removing duplications, filling gaps, correcting errors, and unifying the format. The information extraction module uses a bidirectional gated graph convolutional network, a bidirectional long short-term memory network, and an improved hidden Markov-conditional random field joint model to perform text classification, entity extraction, attribute extraction, and relationship extraction to generate a set of triples. The knowledge graph construction module dynamically updates and optimizes the knowledge graph based on a dynamic evolution mechanism of credibility and self-organizing optimization of knowledge topology. The function of the processor is to realize the full-process automated processing from raw data to structured knowledge, ensuring the accuracy and real-time performance of the knowledge graph.

[0189] Furthermore, the output device includes a visual interface and an API interface.

[0190] The visualization interface graphically displays the autism spectrum disorder knowledge graph, supporting interactive query and exploration. The API provides access to the knowledge graph data for external systems, such as medical diagnostic platforms and scientific research analysis tools. The visualization interface outputs the constructed knowledge graph in an intuitive and actionable format to meet the needs of diverse users.

[0191] Furthermore, the memory includes a raw corpus, an intermediate data repository, and a knowledge graph database. The raw corpus stores preprocessed multi-source heterogeneous data; the intermediate data repository is used to store classification process data during the information extraction process; and the knowledge graph database stores the final constructed autism spectrum disorder knowledge graph and its dynamic evolution record data. The memory's function is to provide efficient data storage and management, supporting the system's full-process data processing and the long-term maintenance and updating of the knowledge graph.

[0192] In summary, the present invention achieves efficient and accurate information extraction through a hybrid convolutional neural network-long short-term memory network-hidden Markov model, optimizes triple weights in real time in combination with a dynamic evolution mechanism of credibility, and realizes intelligent growth and verification of the graph based on a knowledge topology self-organizing algorithm. Compared with the existing technology, this method significantly improves the parsing efficiency of unstructured medical texts, realizes dynamic iteration of credibility through a Markov decision process, and reduces the need for manual intervention by using an adversarial verification mechanism. The knowledge graph finally constructed supports multi-level association tracking of gene pathways, neural circuits, and behavioral phenotypes, providing full-dimensional, highly reliable knowledge support for the precise diagnosis and treatment of ASD and mechanism research.

[0193] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some or all of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present invention, and they should all be included in the scope of the claims and description of the present invention.

Claims

1. A method for constructing a knowledge graph for autism spectrum disorder, characterized in that: The method comprises the following steps: Obtaining multi-source heterogeneous information on autism spectrum disorders; Using the autism spectrum disorder information, constructing an autism spectrum disorder original corpus; Based on the autism spectrum disorder original corpus, information extraction is performed using a hybrid convolutional neural network-long short-term memory network-hidden Markov model, and the information extraction results obtained include: Based on the autism spectrum disorder original corpus, a bidirectional gated graph convolutional network, a bidirectional long short-term memory network and an improved hidden Markov-conditional random field joint model are established; Using the bidirectional gated graph convolutional network, classify the preprocessed text data to obtain a classification result; Based on the classification result, using the bidirectional long short-term memory network to perform enhanced classification to obtain an enhanced classification result; Based on the enhanced classification result, the text data is reclassified using the improved hidden Markov-conditional random field joint model to obtain a reclassification result; According to the reclassification result, performing classification iteration of the text data, completing information extraction and obtaining a triple set, wherein the information extraction includes entity extraction, attribute extraction and relationship extraction; Based on the information extraction results, the dynamic evolution mechanism of design credibility includes: Based on the information extraction results, a multi-dimensional credibility evaluation system is established; According to the multi-dimensional credibility evaluation system, the credibility of the triples is iteratively optimized through a Markov decision process; According to the dynamic evolution mechanism of credibility, building and optimizing knowledge topology self-organization includes: Develop a heterogeneous graph growth algorithm based on the dynamic evolution mechanism of credibility; Using the heterogeneous graph growth algorithm, constructing knowledge topology self-organization; Based on the knowledge topology self-organization, an adversarial verification mechanism is introduced to identify the polarity of the triples and obtain the identification results; Optimizing the knowledge topology self-organization according to the recognition result; Through the self-organization of the knowledge topology, an autism spectrum disorder knowledge graph is generated.

2. The method for constructing a knowledge graph for autism spectrum disorder according to claim 1, characterized in that: The obtaining of multi-source heterogeneous autism spectrum disorder information includes: Initial information on autism spectrum disorder (ASD) was obtained from authoritative medical books, medical research papers, reports from professional autism research institutions, hospital electronic medical records, patient self-reports, and Internet data resources. The initial autism spectrum disorder information is preprocessed to obtain processed multi-source heterogeneous autism spectrum disorder information, wherein the preprocessing includes removing duplicate data, filling missing values, correcting erroneous data and unifying data formats.

3. The method for constructing a knowledge graph for autism spectrum disorder according to claim 1, characterized in that: The step of constructing an autism spectrum disorder original corpus using the autism spectrum disorder information includes: Classify and store the pre-processed autism spectrum disorder information according to data type; Based on the classified storage, natural language processing is performed on the text data to form structured text data, wherein the natural language processing includes word segmentation, part-of-speech tagging and named aspect recognition; The structured text data, semi-structured data and unstructured data are integrated to construct an autism spectrum disorder original corpus.

4. The method for constructing a knowledge graph for autism spectrum disorder according to claim 1, characterized in that: The improved hidden Markov-conditional random field joint model satisfies the following expression: , in, The observation sequence Hidden state sequence The state transition energy function, is the sequence length, is the hidden state transfer matrix, which represents the Status Transfer to Status The probability of is a new position-sensitive feature function, is the original characteristic function, 、 、 is the adaptive weight coefficient, is the total number of position-sensitive eigenfunctions, is the total number of original characteristic functions, 、 、 are the indexes of position-sensitive feature function, original feature function and interactive feature function respectively, is the total number of interactive characteristic functions, is the interaction feature function.

5. The method for constructing a knowledge graph for autism spectrum disorder according to claim 1, characterized in that: Generating an autism spectrum disorder knowledge graph through the knowledge topology self-organization includes: Establishing a multi-level decision path through the self-organization of the knowledge topology; According to the multi-level decision path, a tracking algorithm is designed. The expression of the tracking algorithm is as follows: , in, Indicates gene pathway , neural circuits , behavioral phenotype The correlation between represents the number of associated paths, Indicates the The weight of the path, Indicates the Gene pathways , neural circuits The correlation coefficient between Indicates the Neural circuits on the pathway and behavioral phenotypes The correlation coefficient between 、 、 Represents gene pathways , neural circuits and behavioral phenotypes variance; Based on the tracking algorithm, a segmentation algorithm is designed; By using the tracking algorithm and the segmentation algorithm, an autism spectrum disorder knowledge graph is generated.

6. A system for constructing a knowledge graph for autism spectrum disorders, the system using the method for constructing a knowledge graph for autism spectrum disorders according to any one of claims 1 to 5, characterized in that: The system includes an input device, a processor, an output device and a memory, wherein the input device, the processor, the output device and the memory are connected to each other, wherein the memory is used to store a computer program, the computer program includes program instructions, and the processor is configured to call the program instructions.

Citation Information

Patent Citations

  • Knowledge graph construction method and system capable of distinguishing uniphasic and biphasic affective disorder

    CN115630697A

  • Text classification method and device, equipment and storage medium

    CN115640399A