Autism spectrum disorder knowledge graph construction method and system

Information extraction is performed through the hybrid convolutional neural network-long and short-term memory network-hidden Markov model, and combined with the dynamic evolution mechanism of credibility and self-organization of knowledge topology, the problem of insufficient information extraction accuracy and static credibility evaluation in the construction of autism spectrum disorder knowledge graph is solved, and efficient, real-time and adaptive knowledge graph construction is achieved, which promotes the research and diagnosis and treatment of autism spectrum disorder.

CN120163222AActive Publication Date: 2025-06-17LUZHOU VOCATIONAL & TECHN COLLEGE

Patent Information

Application Number
CN202510638252.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-06-17
Estimated Expiration
2045-05-19

AI Technical Summary

Technical Problem

When building a knowledge graph for autism spectrum disorders, the existing technology faces problems such as insufficient information extraction accuracy, static credibility assessment, and relying on manual intervention, and it is difficult to meet the needs of multi-dimensional and cross-scale knowledge integration.

Method used

The hybrid convolutional neural network-long and short-term memory network-hidden Markov model is used for information extraction, and a dynamic evolution mechanism of credibility is designed, and autism spectrum disorder knowledge graph is generated through knowledge topology self-organization.

Benefits of technology

It improves the accuracy and efficiency of information extraction, realizes the real-time and reliability of the knowledge graph, improves the adaptability and intelligence level, and promotes the research and diagnosis and treatment of autism spectrum disorders.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163222A_ABST
    Figure CN120163222A_ABST
Patent Text Reader

Abstract

The invention relates to the field of artificial intelligence, in particular to an autism spectrum disorder knowledge graph construction method and system. The method comprises the following steps: acquiring multi-source heterogeneous autism spectrum disorder information; constructing an autism spectrum disorder original corpus; performing information extraction by using a hybrid convolutional neural network-long and short-term memory network-hidden Markov model; designing a credibility dynamic evolution mechanism; constructing and optimizing knowledge topology self-organization; and generating an autism spectrum disorder knowledge graph. According to the method, information is efficiently and accurately extracted by constructing the hybrid model, so that the analysis efficiency of the unstructured medical text is remarkably improved; by designing a credibility dynamic evolution mechanism, dynamic iterative optimization of the triad credibility is realized; by introducing an adversarial verification mechanism, the polarity of the triple is efficiently and accurately identified, and the manual intervention requirement is greatly reduced; by constructing a multi-granularity interpretation engine, cross-level correlation analysis of gene pathway-neural circuit-behavior phenotype is realized, and the analysis accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and specifically relates to a method and system for constructing an autism spectrum disorder knowledge graph. Background Art

[0002] As a highly heterogeneous neurodevelopmental disorder, the research on autism spectrum disorder involves multidisciplinary fields such as genetics, neuroscience, and behavior. With the explosive growth of medical data, multi-source heterogeneous data from authoritative literature, electronic medical records, patient records, and the Internet provide rich information resources for the study of autism spectrum disorder. However, these data have significant differences in form, structure, and credibility, resulting in a sharp increase in the difficulty of knowledge integration. Although traditional knowledge graph technologies can partially solve the association problem of structured data, there are obvious bottlenecks in processing unstructured text, dynamic update mechanisms, and semantic depth parsing, making it difficult to meet the urgent needs of multi-dimensional and cross-scale knowledge fusion of autism spectrum disorder.

[0003] In the prior art, the knowledge graph construction methods generally face the following problems: First, information extraction relies on a single model, and the ability to capture semantic associations of unstructured text is insufficient, resulting in limited accuracy of entity recognition and relationship extraction; Second, credibility assessment mostly uses static thresholds or manual annotation, which cannot adapt to the characteristics of rapid iteration of medical knowledge and dynamic evolution of evidence, resulting in poor timeliness of the graph; Third, knowledge topology optimization highly depends on manual intervention and lacks an adaptive adjustment mechanism, making it difficult to cope with the dynamic association analysis of the gene-neuro-behavior complex network in the study of autism spectrum disorder. In addition, the existing methods lack a systematic solution in multi-source data fusion, heterogeneous feature collaborative calculation, and multi-level decision path modeling, seriously restricting the practical value of the autism spectrum disorder knowledge graph.

[0004] Currently, there are not many research works on the construction of the autism spectrum disorder knowledge graph, and there is no specific multi-dimensional, multi-level, and multi-scale method for constructing the autism spectrum disorder knowledge graph. Summary of the Invention

[0005] In view of the deficiencies in the prior art, the present invention provides a method and system for constructing an autism spectrum disorder knowledge graph.

[0006] In a first aspect, a method for constructing an autism spectrum disorder knowledge graph provided by the present invention includes the following steps: obtaining multi-source heterogeneous autism spectrum disorder information; constructing an original corpus of autism spectrum disorder using the autism spectrum disorder information; based on the original corpus of autism spectrum disorder, using a hybrid convolutional neural network-long short-term memory network-hidden Markov model to perform information extraction and obtain an information extraction result; designing a credibility dynamic evolution mechanism based on the information extraction result; constructing and optimizing a knowledge topology self-organization according to the credibility dynamic evolution mechanism; and generating an autism spectrum disorder knowledge graph through the knowledge topology self-organization. By using a hybrid convolutional neural network-long short-term memory network-hidden Markov model for information extraction, the present invention realizes efficient parsing and accurate extraction of complex texts, overcomes the deficiencies of traditional methods in processing unstructured data, and greatly improves the accuracy and efficiency of information extraction; by designing a credibility dynamic evolution mechanism, it realizes real-time evaluation of information credibility, ensures the timeliness and reliability of the knowledge graph, and solves the problem that traditional knowledge graphs are difficult to adapt to rapid information updates; by constructing and optimizing a knowledge topology self-organization, it realizes automatic optimization and dynamic adjustment of the knowledge structure, improves the adaptive ability and intelligent level of the knowledge graph, and breaks through the limitation of relying on manual intervention in the construction process of traditional knowledge graphs; by generating an autism spectrum disorder knowledge graph, it greatly promotes the research and diagnosis and treatment level of autism spectrum disorder.

[0007] Optionally, the obtaining of multi-source heterogeneous autism spectrum disorder information includes: respectively obtaining structured, semi-structured, and unstructured initial information of autism spectrum disorder from medical authoritative books, medical research papers, reports of professional autism research institutions, hospital electronic medical records, patient self-report records, and Internet data resources; preprocessing the initial information of autism spectrum disorder to obtain processed multi-source heterogeneous autism spectrum disorder information, and the preprocessing includes removing duplicate data, filling in missing values, correcting incorrect data, and unifying data formats. By obtaining initial information from multi-source heterogeneous data, the present invention constructs a comprehensive data pool covering structured, semi-structured, and unstructured data, significantly improving data diversity and comprehensiveness, and providing a rich foundation for knowledge graph construction; by preprocessing the initial information to remove duplicate data, fill in missing values, correct incorrect data, and unify data formats, it realizes the efficient integration and cleaning of multi-source heterogeneous data, greatly improving the reliability and usability of the data; by combining multi-source heterogeneous data with preprocessing techniques, it realizes in-depth mining and accurate processing of autism spectrum disorder information.

[0008] Optionally, constructing the original corpus of autism spectrum disorder by using the autism spectrum disorder information includes: classifying and storing the preprocessed autism spectrum disorder information according to data types; based on the classified storage, performing natural language processing on text data to form structured text data, and the natural language processing includes word segmentation, part-of-speech tagging, and named entity recognition; integrating the structured text data, semi-structured data, and unstructured data to construct the original corpus of autism spectrum disorder. By classifying and storing the preprocessed autism spectrum disorder information according to data types, the present invention realizes the efficient organization and management of multi-source heterogeneous data, breaks through the limitations of traditional methods in data storage and retrieval, and significantly improves the systematicness and scalability of data processing; by performing natural language processing operations of word segmentation, part-of-speech tagging, and named entity recognition on text data, the unstructured text is converted into structured text data, solving the problem of low efficiency in processing unstructured data by traditional methods, greatly improving the availability and analysis accuracy of text data, and providing high-quality input for knowledge graph construction; by integrating structured text data, semi-structured data, and unstructured data, a comprehensive and multi-level original corpus of autism spectrum disorder is constructed, realizing the deep integration and unified management of multi-source data, significantly improving the coverage and application value of the corpus, and providing strong data support for the research and diagnosis and treatment of autism spectrum disorder.

[0009] Optionally, based on the original corpus of autism spectrum disorder, using a hybrid convolutional neural network-long short-term memory network-hidden Markov model for information extraction and obtaining the information extraction result includes: based on the original corpus of autism spectrum disorder, establishing a bidirectional gated graph convolutional network, a bidirectional long short-term memory network, and an improved hidden Markov-conditional random field joint model; using the bidirectional gated graph convolutional network, the bidirectional long short-term memory network, and the improved hidden Markov-conditional random field joint model for information extraction and obtaining the information extraction result. By establishing a bidirectional gated graph convolutional network and a bidirectional long short-term memory network, the present invention realizes the efficient capture and modeling of complex semantic relationships in the original corpus of autism spectrum disorder, breaks through the limitations of traditional methods in processing graph-structured data, significantly improves the depth and accuracy of information extraction, and provides more accurate semantic support for knowledge graph construction; by improving the hidden Markov-conditional random field joint model, the accuracy of named entity recognition and relationship extraction is greatly improved, providing a more reliable technical guarantee for the automatic extraction of autism spectrum disorder knowledge; by combining the bidirectional gated graph convolutional network, the bidirectional long short-term memory network with the improved hidden Markov-conditional random field joint model, the collaborative analysis and joint optimization of multi-source heterogeneous data are realized, breaking through the performance bottleneck of a single model in information extraction, significantly improving the comprehensiveness and robustness of information extraction, and providing an efficient and intelligent technical path for the construction of the autism spectrum disorder knowledge graph.

[0010] Optionally, using the bidirectional gated graph convolutional network, the bidirectional long short-term memory network, and the improved hidden Markov - conditional random field joint model for information extraction and obtaining the information extraction result includes: using the bidirectional gated graph convolutional network to classify the preprocessed text data to obtain a classification result; based on the classification result, using the bidirectional long short-term memory network for enhanced classification to obtain an enhanced classification result; based on the enhanced classification result, using the improved hidden Markov - conditional random field joint model to reclassify the text data to obtain a reclassified result; according to the reclassified result, performing the classification iteration of the text data, completing information extraction and obtaining a triple set, where the information extraction includes entity extraction, attribute extraction, and relationship extraction. The present invention uses the bidirectional gated graph convolutional network to classify the preprocessed text data and uses the bidirectional long short-term memory network to perform enhanced classification on the classified text data, achieving deep extraction and efficient classification of multi-level features of the text data, breaking through the limitations of traditional classification models in feature extraction ability, significantly improving the accuracy and robustness of the classification result, and providing high-quality input for subsequent information extraction; using the improved hidden Markov - conditional random field joint model to reclassify the text data, solving the deficiencies of traditional models in long-distance dependence and semantic association, greatly improving the accuracy and consistency of the reclassified result, and laying a solid foundation for entity, attribute, and relationship extraction; performing the classification iteration of the text data through the reclassified result, realizing the dynamic optimization of the information extraction process, breaking through the performance bottleneck of traditional static extraction methods, significantly improving the integrity and accuracy of the triple set, and providing an efficient and intelligent technical path for the construction of the autism spectrum disorder knowledge graph.

[0011] Optionally, the improved hidden Markov - conditional random field joint model satisfies the following expression:

[0012] where is the state transition energy function of the hidden state sequence in the observation sequence , is the sequence length, is the hidden state transition matrix, indicating the probability of transitioning from the -th state to the -th state , is the newly added position-sensitive feature function, is the original feature function, , , are adaptive weight coefficients, is the total number of position-sensitive feature functions, is the total number of original feature functions, , , are the indices of the position-sensitive feature function, the original feature function, and the interaction feature function respectively, is the total number of interaction feature functions, is the interaction feature function. By integrating the position-sensitive feature function, the present invention accurately captures the influence of position information in sequence data on state transition, improves the recognition accuracy of the model in complex scenarios, and realizes effective modeling of data sequences with strong position dependence; by introducing an adaptive weight coefficient, it can dynamically adjust the contribution of each feature function to the state transition energy according to data characteristics, enhances the flexibility and generalization ability of the model, and enables the model to exhibit excellent performance in different fields and datasets; by increasing the interaction feature function and considering its complex interaction relationship with the observation sequence and the hidden state, it comprehensively describes the potential laws and patterns in sequence data, improves the parsing ability of the model for complex structured data, and provides new ideas and methods for processing high-dimensional and non-linear sequence data.

[0013] Optionally, designing a credibility dynamic evolution mechanism based on the information extraction result includes: establishing a multi-dimensional credibility evaluation system based on the information extraction result; and iteratively optimizing the triple credibility through a Markov decision process according to the multi-dimensional credibility evaluation system. By introducing a multi-dimensional credibility evaluation system and combining the information extraction result, the present invention realizes a comprehensive and dynamic evaluation of the triple credibility, significantly improves the accuracy and adaptability of the evaluation, especially shows stronger robustness when dealing with complex and multi-source heterogeneous data, and provides a more refined and scientific basis for credibility evaluation; through the iterative optimization mechanism of the Markov decision process, it dynamically adjusts the triple credibility, realizes real-time update and optimization of the credibility, overcomes the defect that traditional static evaluation methods are difficult to adapt to the dynamic changes of data, enables the credibility evaluation to be continuously optimized with the evolution of data, and greatly improves the timeliness and reliability of the evaluation result; by deeply integrating the multi-dimensional evaluation with the Markov decision process, a closed-loop mechanism for credibility dynamic evolution is constructed, realizing full-process automation from data extraction to credibility evaluation and then to optimization, significantly reducing the labor cost, while improving the evaluation efficiency and consistency, and providing a new technical path for credibility management in large-scale data application scenarios.

[0014] Optionally, constructing and optimizing the knowledge topology self-organization according to the credibility dynamic evolution mechanism includes: developing a heterogeneous graph spectrum growth algorithm according to the credibility dynamic evolution mechanism; constructing the knowledge topology self-organization by using the heterogeneous graph spectrum growth algorithm; introducing an adversarial verification mechanism based on the knowledge topology self-organization to identify the polarity of triples and obtain the identification result; and optimizing the knowledge topology self-organization according to the identification result. By developing a heterogeneous graph spectrum growth algorithm and combining it with the credibility dynamic evolution mechanism, the present invention realizes the adaptive growth and expansion of the knowledge topology, significantly improves the dynamics and flexibility of the graph spectrum, especially shows stronger adaptability when dealing with multi-source heterogeneous data, and provides a more intelligent and automated solution for the construction of the knowledge topology; by introducing an adversarial verification mechanism to identify the polarity of triples, it realizes the accurate analysis and verification of the semantics of triples in the knowledge topology, overcomes the defect that traditional methods are difficult to effectively identify complex semantic relationships, greatly improves the accuracy and reliability of the knowledge topology, and provides stronger semantic support for knowledge reasoning and decision-making; by dynamically optimizing the knowledge topology self-organization, it realizes the continuous iteration and improvement of the knowledge topology, significantly improves the timeliness and practicality of the knowledge topology.

[0015] Optionally, generating an autism spectrum disorder knowledge graph through the knowledge topology self-organization includes: establishing a multi-level decision-making path through the knowledge topology self-organization; designing a tracking algorithm according to the multi-level decision-making path, and the expression of the tracking algorithm is as follows: , where represents the gene pathway , the neural circuit , and the behavioral phenotype the degree of association between them, represents the number of association paths, represents the th path weight, represents the th path on the gene pathway , the neural circuit the correlation coefficient between them, represents the th path on the neural circuit and the behavioral phenotype the correlation coefficient between them, , , respectively represent the gene pathway , the neural circuit and the behavioral phenotype The variance; based on the tracking algorithm, a segmentation algorithm is designed; through the tracking algorithm and the segmentation algorithm, an autism spectrum disorder knowledge graph is generated. By establishing a multi-level decision-making path and combining knowledge topology self-organization, the present invention realizes the systematic analysis and modeling of the complex associations among gene pathways, neural circuits, and behavioral phenotypes in autism spectrum disorder, breaks through the limitation that traditional single-level analysis methods are difficult to capture multi-level interactions, significantly improves the comprehensiveness and depth of the knowledge graph, and provides a more refined theoretical framework for the mechanism research of autism spectrum disorder; by designing a tracking algorithm, the association degree among gene pathways, neural circuits, and behavioral phenotypes is dynamically calculated based on the multi-level decision-making path, realizing the efficient analysis and tracking of complex biological data, overcoming the defect that traditional methods are difficult to quantify multi-dimensional association relationships, and greatly improving the accuracy and interpretability of the knowledge graph, providing more reliable data support for the accurate diagnosis and intervention of autism spectrum disorder; by combining the tracking algorithm with the segmentation algorithm, an autism spectrum disorder knowledge graph is dynamically generated, realizing the adaptive update and optimization of the knowledge graph.

[0016] In a second aspect, an autism spectrum disorder knowledge graph construction system provided by the present invention includes an input device, a processor, an output device, and a memory. The input device, the processor, the output device, and the memory are interconnected. Among them, the memory is used to store a computer program, the computer program includes program instructions, the processor is configured to call the program instructions, and the system uses the provided autism spectrum disorder knowledge graph construction method. The system provided by the present invention has a high degree of integration, and the information transmission among various components is smooth. By integrating multi-source heterogeneous data, a comprehensive and high-quality autism spectrum disorder raw corpus is constructed, significantly improving the data coverage and reliability, and providing rich and accurate basic data support for knowledge graph construction; by adopting a bidirectional gated graph convolutional network, a bidirectional long short-term memory network, and an improved hidden Markov - conditional random field joint model, multi-level feature extraction of text data, context-aware classification iteration, and efficient extraction of entities, attributes, and relationships are realized, solving the deficiencies of traditional information extraction methods in semantic understanding, long-distance dependence, and dynamic optimization, and greatly improving the accuracy and intelligence level of information extraction, providing an efficient and reliable technical path for knowledge graph construction; by designing a credibility dynamic evolution mechanism and optimizing knowledge topology self-organization, real-time update, dynamic adjustment, and adaptive optimization of the knowledge graph are realized, significantly improving the real-time performance, reliability, and application value of the knowledge graph, and providing a comprehensive, accurate, and dynamically updated intelligent knowledge platform for the research and diagnosis and treatment of autism spectrum disorder. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1Flowchart of the method for constructing an autism spectrum disorder knowledge graph according to an embodiment of the present invention; Figure 2 Schematic diagram of the complete autism spectrum disorder knowledge graph constructed according to an embodiment of the present invention; Figure 3 Schematic diagram of the structure of the autism spectrum disorder knowledge graph construction system according to an embodiment of the present invention. Detailed implementation manners

[0018] The specific embodiments of the present invention will be described in detail below. It should be noted that the embodiments described here are only for illustrative purposes and are not used to limit the present invention. In the following description, in order to provide a thorough understanding of the present invention, a large number of specific details are set forth. However, it will be apparent to those of ordinary skill in the art that the present invention does not have to employ these specific details. In other instances, well-known circuits, software, or methods have not been specifically described in order to avoid obscuring the present invention.

[0019] Throughout the specification, references to "one embodiment", "an embodiment", "an example", or "an example" mean that a particular feature, structure, or characteristic described in connection with the embodiment or example is included in at least one embodiment of the present invention. Thus, the phrases "in one embodiment", "in an embodiment", "an example", or "an example" appearing throughout the specification do not necessarily all refer to the same embodiment or example. Additionally, the particular features, structures, or characteristics may be combined in any suitable combination and / or sub-combination in one or more embodiments or examples. Moreover, those of ordinary skill in the art should understand that the diagrams provided herein are for illustrative purposes only and are not necessarily drawn to scale.

[0020] Please refer to Figure 1 , an embodiment of the present invention provides a method for constructing an autism spectrum disorder knowledge graph, the method comprising the following steps: S1. Obtain multi-source heterogeneous autism spectrum disorder information.

[0021] Among them, S1 further includes the following steps: S11. Respectively obtain the initial information of autism spectrum disorder in structured, semi-structured, and unstructured forms from authoritative medical books, medical research papers, reports of professional autism research institutions, hospital electronic medical records, patient self-report records, and Internet data resources.

[0022] Specifically, select authoritative medical books, including "Diagnostic and Statistical Manual of Mental Disorders (DSM-5)" and "International Classification of Diseases (ICD)"; through manual reading, extract structured information, including the definition, diagnostic criteria, symptom manifestations, and epidemiological data of autism spectrum disorder (ASD); organize the extracted information into a database form for subsequent processing.

[0023] Further, search for scientific research papers on ASD through the CNKI academic database; use text mining tools or manual reading to extract semi-structured information, including research methods, experimental data, and conclusions in the papers; organize the extracted information according to the fields of paper title, author, publication year, journal name, and research content.

[0024] Further, pay attention to the reports released by the Autism Research Foundation and autism science organizations; download and read the reports, and extract structured and semi-structured information, including the latest research progress, treatment strategies, and policy recommendations for ASD; organize the extracted information according to the fields of report title, releasing agency, release date, and main content.

[0025] Further, cooperate with multiple hospitals to obtain the electronic medical record data of ASD patients; extract the unstructured information in the medical records through data interfaces or data export methods, including patient basic information, diagnosis records, treatment records, and follow-up records; anonymize the extracted information and store it in a secure database.

[0026] Further, collect the self-reported records of ASD patients and their families through online questionnaires and face-to-face interviews; organize the unstructured information self-reported by the patients, including symptom manifestations, treatment experiences, and life impacts; process the self-reported records into text and store them in a text file.

[0027] Further, crawl the post and comment data on Internet platforms, where the Internet platforms include autism-related forums, social media, and blogs; use natural language processing technology to extract unstructured information, including discussion hotspots, patient experiences, and treatment suggestions for ASD; organize the extracted information according to the fields of source, release time, and content, and store it in a database.

[0028] S12. Preprocess the initial information of the autism spectrum disorder to obtain the processed multi-source heterogeneous autism spectrum disorder information, and the preprocessing includes removing duplicate data, filling in missing values, correcting incorrect data, and unifying the data format.

[0029] In one embodiment, first use a data deduplication algorithm to identify and delete duplicate records or data items. For structured data, such as records in a table or database, remove duplicates by comparing key fields, such as patient ID, paper title; for unstructured data, such as text files or social media posts, identify and remove duplicate content through text similarity calculation.

[0030] Furthermore, according to the nature and missing situation of the data, select appropriate filling methods, including mean filling, median filling, mode filling, interpolation filling, and prediction model filling methods based on machine learning. For numerical data, such as age and weight, use the mean or median for filling; for categorical data, such as gender and diagnosis type, use the mode for filling; for time series data, such as follow-up records, use interpolation for filling; for complex unstructured data, such as missing information in patient self-report records, use a prediction model based on machine learning for filling.

[0031] Furthermore, identify and correct errors in the data through data verification rules and machine learning algorithms. For structured data, set verification rules, such as age range and gender value, to automatically detect and correct errors; for semi-structured or unstructured data, use natural language processing techniques to identify and correct errors, such as spelling mistakes and semantic ambiguity.

[0032] Furthermore, design a unified data format and storage structure according to the use and analysis requirements of the data. For structured data, convert it into the standard CSV file format; for semi-structured data, convert it into XML format; for unstructured data, perform text processing on it and store it in a unified database. At the same time, ensure that all data follows the same coding standard and date format.

[0033] S2. Utilize the autism spectrum disorder information to construct an original corpus of autism spectrum disorder.

[0034] In one embodiment, first use a database system to classify and store the preprocessed autism spectrum disorder information according to the types of text structured data, semi-structured data, and unstructured data, and establish indexes or labels for each type of data for subsequent retrieval and processing.

[0035] Furthermore, use a word segmentation tool to perform word segmentation on the text data, remove stop words, and retain meaningful vocabulary.

[0036] Furthermore, perform part-of-speech tagging on the segmented vocabulary, such as nouns, verbs, and adjectives. The part-of-speech tagging helps to understand the role and meaning of the vocabulary in the sentence.

[0037] Furthermore, identify named entities in the text, such as personal names, place names, and organization names, as well as professional terms related to ASD. This helps to extract key information in the text and provides a basis for subsequent analysis.

[0038] Furthermore, integrate the text data processed by word segmentation, part-of-speech tagging, and named entity recognition into structured text data, and store the structured text data in XML format.

[0039] Furthermore, structured text data, semi-structured data, and unstructured data are integrated together to form an original corpus of autism spectrum disorder.

[0040] S3. Based on the original corpus of autism spectrum disorder, a hybrid convolutional neural network-hidden Markov model is used to perform information extraction and obtain an information extraction result.

[0041] In one embodiment, a bidirectional gated graph convolutional network is established based on the original corpus of autism spectrum disorder. The bidirectional gated graph convolutional network includes an input layer, a convolutional layer, a pooling layer, a fully connected layer, and a SoftMax layer.

[0042] The input layer is responsible for converting the original text data into a format that can be processed by the model.

[0043] Specifically, a pre-trained word vector model, such as Word2Vec, is used to convert each word in the text into a word vector of a fixed dimension. Among them, the dimension of the word vector is , then a sentence of length is represented as a matrix .

[0044] Furthermore, each word is used as a node in the graph, and the node feature is the corresponding word vector.

[0045] Furthermore, according to the co-occurrence relationship and syntactic dependency relationship between words, an adjacency matrix is constructed. If word and word are adjacent or have a dependency relationship in the sentence, then , otherwise .

[0046] The convolutional layer is used to extract local features from the graph structure.

[0047] Specifically, a gated recurrent unit is used as the gating mechanism to control the propagation of information in the graph. For each node , its feature vector of the -th layer is calculated by the following formula:

[0048] Among them, is the feature vector of node in the -th layer, is the gating activation function, is the domain order, is the adjacency set of node , is the The weight matrix of the -th neighborhood in the layer, is the feature vector of the node in the layer, and is the bias term of the -th neighborhood in the layer; information is propagated from the forward neighborhood and the backward neighborhood of the node respectively, and then the feature vectors of the two are concatenated.

[0049] It should be noted that this network contains two convolutional layers: The first convolutional layer: The input is the word vector matrix and the adjacency matrix , and the output is the feature matrix , where is the output dimension.

[0050] The second convolutional layer: The input is the output of the first convolutional layer and the adjacency matrix , and the output is the feature matrix , where is the output dimension.

[0051] The tokenization layer is used to reduce the feature dimension and extract important features.

[0052] Specifically, for the feature vector of each node, the maximum value is taken along the feature dimension, and the formula is as follows:

[0053] where represents the maximum value of the feature vector of the node , and represents the feature vector when the feature dimension is . Its output is the pooled feature matrix .

[0054] It should be noted that this network contains two pooling layers: The first pooling layer: The input is the output of the first convolutional layer, and the output is the feature matrix .

[0055] The second pooling layer: The input is the output of the second convolutional layer, and the output is the feature matrix .

[0056] The fully connected layer is used to map the pooled features to the classification label space.

[0057] Specifically, the feature matrices after two layers of pooling​ and concatenate them to obtain the final feature vector

[0058] Furthermore, map the feature vector to the classification label space:

[0059] where is the weight matrix of the fully connected layer, is the bias term of the fully connected layer, is the output of the fully connected layer.

[0060] The SoftMax layer is used to output the classification probability.

[0061] Specifically, perform the SoftMax operation on the output of the fully connected layer to obtain the probability of each category:

[0062] where is the probability of the th category in the sample , is the output of the th category of the fully connected layer, is the total number of categories, is the output of the th category of the fully connected layer.

[0063] Furthermore, select the category with the highest probability as the final classification result:

[0064] It should be noted that the bidirectional gated graph convolutional network gradually reduces the feature dimension through two convolutional layers and two pooling layers, significantly reducing the computational amount; extracts important features through the pooling operation, reduces redundant information, and accelerates the model convergence; the two convolutional layers can capture more complex local features, and the two pooling layers can enhance the generalization ability of the model and prevent overfitting. Therefore, the bidirectional gated graph convolutional network can efficiently extract features from text data and improve the training speed and classification accuracy through two convolutional layers and two pooling layers.

[0065] Furthermore, use the bidirectional gated graph convolutional network to classify the preprocessed text data to obtain the classification result.

[0066] ​Furthermore, a bidirectional long short-term memory network is introduced to perform enhanced classification on the text data. The bidirectional long short-term memory network not only includes the structures of traditional forward and backward long short-term memory units, but also integrates a temporal attention mechanism, multi-granularity feature fusion, domain knowledge-guided regularization, and a dynamic error correction mechanism. The bidirectional long short-term memory network stacks two long short-term memory network layers together, with one processing the input sequence forward and the other processing the input sequence backward, and then concatenating the outputs of the two long short-term memory network layers, so as to be able to capture the context information before and after each time step in the sequence simultaneously.

[0067] Specifically, take the classification result output by the bidirectional gated graph convolutional network as the initial feature and input it into the bidirectional long short-term memory network. Introduce a temporal attention mechanism in the bidirectional long short-term memory network to dynamically calculate the attention weights at each time step, and strengthen the capture of the key time step features in the classification result.

[0068] Furthermore, set up a hierarchical feature extraction module: capture local temporal patterns at the lower layer of the bidirectional long short-term memory network, such as phrase-level features of symptom descriptions; capture global temporal patterns at the higher layer of the bidirectional long short-term memory network, such as paragraph-level features of disease course evolution.

[0069] Furthermore, adaptively weight and fuse multi-granularity features through a gated fusion unit to enhance the model's ability to express complex semantics in medical texts. Among them, multi-granularity features refer to various information features of different scales and levels that can be captured when processing sequence data, including information on time scale and context.

[0070] Furthermore, introduce medical knowledge graph embedding as external prior knowledge, and constrain the hidden state space of the bidirectional long short-term memory network through a knowledge alignment loss function to ensure the semantic consistency of the features learned by the model with the medical domain.

[0071] Furthermore, design a curriculum learning strategy to gradually increase the complexity of training samples, such as from simple symptom descriptions to complex comorbidity relationships, to improve the model's classification performance on samples.

[0072] Furthermore, add an error perception module to the output layer of the bidirectional long short-term memory network. By calculating the error distribution between the classification result and the true label, dynamically adjust the hidden state update rule of the bidirectional long short-term memory network.

[0073] Furthermore, adopt an adversarial training strategy to generate adversarial samples to enhance the model's robustness and reduce classification biases caused by data noise.

[0074] Furthermore, map the final hidden state of the bidirectional long short-term memory network to the classification space through a fully connected layer to output the enhanced classification result.

[0075] It should be noted that the bidirectional long short-term memory network introduced in the present invention significantly improves the ability to capture important information in long texts by dynamically weighting the features of key time steps through a temporal attention mechanism; through hierarchical feature extraction and a gated fusion unit, it realizes the effective fusion of multi-granularity features and enhances the model's understanding ability of complex medical texts.

[0076] Furthermore, based on the enhanced classification result, an improved hidden Markov - conditional random field joint model is established, and the improved hidden Markov - conditional random field joint model satisfies the following expression:

[0077] where is the observation sequence and is the state transition energy function of the hidden state sequence is the sequence length, is the hidden state transition matrix, indicating the probability of transitioning from the th state to the th state , is the newly added position-sensitive feature function, is the original feature function, used to capture the relationship between the observation sequence and the hidden state, , , are the adaptive weight coefficients, used to adjust the contribution degree of the feature functions, is the total number of position-sensitive feature functions, is the total number of original feature functions, , , are the indices of the position-sensitive feature function, the original feature function, and the interaction feature function respectively, is the total number of interaction feature functions, is the interaction feature function.

[0078] The improvements of the improved hidden Markov - conditional random field joint model are as follows: First, a dynamic state transition matrix. The hidden state transition matrix of the existing model is static, and after improvement, it is dynamically adjusted through . The role of the improvement is to make the state transition probability depend on the local features of the observation sequence, such as the context information of the current position.

[0079] Second, the non-linear interaction of the feature functions. The contribution of the improved feature function is divided by the denominator Modulation forms a non-linear dependence. The improvement allows competition or cooperation between features. For example, if represents interfering features, the denominator can suppress the weights.

[0080] Thirdly, position-sensitive features. Used to capture the observed characteristics of the modeling position such as the statistical laws at the beginning or end of a sentence in the text.

[0081] Furthermore, according to the improved Hidden Markov-Conditional Random Field joint model, the text data is reclassified to obtain a reclassification result.

[0082] Specifically, first obtain the text data classified based on the bidirectional gated graph convolutional network and the bidirectional long short-term memory network, represented as a sequence of word vectors .

[0083] Furthermore, use the trained improved Hidden Markov-Conditional Random Field joint model to infer the text data and find the label sequence that minimizes the energy function :

[0084] Furthermore, use the Viterbi algorithm to efficiently find the optimal label sequence.

[0085] Furthermore, output , where is the reclassification result of the th word.

[0086] Furthermore, based on the reclassification result, use the bidirectional gated graph convolutional network, the bidirectional long short-term memory network, and the improved Hidden Markov-Conditional Random Field joint model for classification iteration to complete information extraction and obtain a triple set. The information extraction includes entity extraction, attribute extraction, and relationship extraction.

[0087] Specifically, first based on the reclassification result, confirm and extract the named entities in the text, and summarize and organize the extracted entities to form an entity list. The named entities include the clinical symptoms, treatment methods, and related drugs of autism spectrum disorder.

[0088] Furthermore, for the identified entities, extract their attribute information and associate the extracted attribute information with the corresponding entities to form entity-attribute pairs. The attribute information includes the specific description of the rehabilitation treatment method, the severity of the clinical symptoms, the dosage and usage of the drug.

[0089] Furthermore, based on entity and attribute extraction, extract the association relationships between entities, and represent the extracted association relationships in the form of triples. The association relationships include the relationship between clinical symptoms and treatment methods, the relationship between drugs and side effects, and the relationship between different clinical symptoms.

[0090] Furthermore, combine the extracted entities, attributes, and relationships into a triple set to form a knowledge unit and verify it to ensure its accuracy and integrity. The verification methods include manual verification and comparison verification with existing knowledge bases.

[0091] Furthermore, use the verified triple set as new training data and add it to the training set.

[0092] Furthermore, retrain the bidirectional gated graph convolutional network model, bidirectional long short-term memory network, and improved hidden Markov - conditional random field joint model, and optimize the model parameters using the new training data.

[0093] Furthermore, repeatedly execute the above classification, entity extraction, attribute extraction, relationship extraction, and triple generation steps for multiple iterations. After each iteration, evaluate the model performance, record the accuracy and recall metrics, and select the optimal model parameters. After multiple iterations of optimization, obtain the final information extraction result.

[0094] Furthermore, store the triple set in the knowledge base for subsequent knowledge graph construction and application.

[0095] S4. Design a credibility dynamic evolution mechanism based on the information extraction result.

[0096] In one embodiment, based on the information extraction result, establish a multi-dimensional credibility evaluation system. The multi-dimensional credibility evaluation system includes constructing a dynamic weight matrix by combining the authority of data sources, the number of clinical verifications, and cross-modal consistency.

[0097] Specifically, first evaluate the authority of the data source, such as whether it comes from authoritative medical journals, clinical trials, or expert consensus, to obtain the corresponding score. . The scoring criteria are as follows: high authority, such as top journals: ; medium authority, such as general journals: ; low authority, such as non-academic sources: .

[0098] Furthermore, evaluate the number of times the knowledge described by the triples has been verified clinically to obtain the corresponding score. . The scoring criteria are as follows: high number of verifications : ; medium number of verifications : ; Low verification times : .

[0099] Furthermore, evaluate the consistency of triples in different data modalities, such as text and image, to obtain corresponding scores . The scoring criteria are as follows: Completely consistent: ; Partially consistent: ; Inconsistent: .

[0100] Furthermore, define a weight matrix , where , , represent the weights of data source authority, clinical verification times, and cross-modal consistency respectively. Dynamically adjust the weights according to the evaluation results and feedback.

[0101] Furthermore, calculate the credibility of triples, and the formula is as follows:

[0102] where is the credibility of the triple, is the score of data source authority, is the score of clinical verification times, is the score of cross-modal consistency, , , represent the weights of data source authority, clinical verification times, and cross-modal consistency respectively.

[0103] Furthermore, according to the multi-dimensional credibility evaluation system, design a reinforcement learning optimizer. The reinforcement learning optimizer realizes the iterative optimization of triple credibility through a Markov decision process and constructs a reward function based on the gold standard of clinical diagnosis.

[0104] Specifically, according to the multi-dimensional credibility evaluation system, model the triple credibility optimization problem using a Markov decision process; the Markov decision process includes a state space, an action space, a reward function, and a transition probability; the state space represents the credibility evaluation result of the current triple. The action space represents operations on the triple, such as retaining, correcting, or deleting. The reward function is defined based on the gold standard of clinical diagnosis. The reward rules are as follows: If the triple is consistent with the gold standard of clinical diagnosis after the operation, the reward is +1; If the triple is inconsistent with the gold standard of clinical diagnosis after the operation, the penalty is -1; If the operation is invalid, such as deleting a correct triple, the penalty is -0.5.

[0105] The transition probability represents the probability of transitioning to a new state after performing an action in a state.

[0106] Furthermore, the Q-learning algorithm is used for optimization, and its Q-value update formula is as follows:

[0107] Where, is the updated Q-value, is the Q-value before update, is the learning rate, is the discount factor, is the reward, is the new state, is the new action. During the optimization process, when the Q-value converges or reaches the maximum number of iterations, the optimization stops.

[0108] Furthermore, a set of triples with significantly improved credibility is output.

[0109] Through the application of a multi-dimensional credibility evaluation system, a dynamic weight matrix, a reinforcement learning optimizer, a reward function based on the clinical diagnosis gold standard, and the Q-learning algorithm, this method realizes the dynamic evolution and optimization of the credibility of triples. It not only improves the accuracy and reliability of credibility evaluation, but also makes the evaluation results more in line with actual needs, having high practical value and promotion prospects.

[0110] S5. According to the credibility dynamic evolution mechanism, construct and optimize the knowledge topology self-organization.

[0111] Among them, S5 further includes the following steps: S51. According to the credibility dynamic evolution mechanism, construct the knowledge topology self-organization.

[0112] In one embodiment, based on the credibility dynamic evolution mechanism, a set of heterogeneous graph spectrum growth algorithms is developed, and the knowledge topology self-organization is constructed by using the heterogeneous graph spectrum growth algorithm. The process of constructing the knowledge topology self-organization is as follows: The heterogeneous graph spectrum growth algorithm combines the symptom severity gradient to construct a hierarchical topology structure, and uses the graph neural network propagation algorithm to realize the autonomous clustering of knowledge nodes, forming the knowledge topology self-organization.

[0113] It should be noted that the knowledge topology self-organization is a hierarchical knowledge network structure automatically constructed by the heterogeneous graph spectrum growth algorithm and the graph neural network propagation algorithm. The core process of construction includes hierarchical topology construction and autonomous clustering of knowledge nodes.

[0114] The effect of this method is to form a knowledge organization system with hierarchy, self-adaptability and no need for manual intervention, providing dynamic support for complex medical relationship modeling.

[0115] Regarding the construction process of knowledge topology self-organization, specifically, first define the symptom severity gradient to guide the hierarchical construction of the atlas. Let the symptom node set be , where is the th symptom, associate it with a symptom severity , and . Construct a hierarchical function to map symptom nodes to different levels, where the number of levels is proportional to the severity. The hierarchical function is defined as:

[0116] where is the hierarchical function, is the severity of the th symptom, is the severity of the th symptom, is the preset maximum number of levels, represents rounding down.

[0117] Furthermore, let the heterogeneous graph be , where is the node set, is the edge set, is the node type set.

[0118] Furthermore, use a graph neural network to propagate node features and achieve autonomous clustering. The node feature matrix is , where is the feature dimension. The update formula of the graph neural network is as follows: , where is the hidden state of node at the th layer, is the hidden state of node at the th layer, is the neighbor set of node , is the neighbor set of node , is the weight matrix of the th layer, is the bias vector of the th layer, is the activation function.

[0119] It should be noted that compared with the prior art, the present invention realizes autonomous clustering by introducing a clustering loss function. The clustering loss function is as follows:

[0120] Wherein, is the clustering loss function, is the set of clustering centers, is the representation of the clustering center , is the number of layers of the graph neural network.

[0121] Furthermore, a knowledge topology self-organization is formed; Furthermore, multi-source heterogeneous knowledge fusion is performed to preliminarily optimize the knowledge topology self-organization.

[0122] Specifically, let the multi-source heterogeneous knowledge sources be , , the th knowledge source provides a knowledge subgraph .

[0123] Furthermore, a fusion function is established, and its relational expression is:

[0124] Wherein, is the fusion function, is the total number of knowledge sources, is the th knowledge node, is the th edge set, is the th node type, is the edge set across knowledge sources, which is added by calculating the similarity between nodes:

[0125] Wherein, is the similarity threshold, , respectively represent two nodes, , respectively represent two different node sets, is the similarity function, defined as:

[0126] Wherein, is in the graph neural network, the hidden state of node at the th layer, is in the graph neural network, the node at the layer's hidden state.

[0127] It should be noted that, in the process of knowledge fusion, compared with the prior art, the inventive point lies in: introducing a cross-knowledge-source similarity metric, combining the hidden state of the graph neural network to calculate the similarity between nodes, and dynamically adding edges across knowledge sources. By introducing the cross-knowledge-source similarity metric, the similarity between nodes or entities in different knowledge sources can be calculated more precisely. Traditional similarity calculation methods are often limited within a single knowledge source, while the cross-knowledge-source similarity metric breaks this limitation and realizes cross-domain similarity comparison. The graph neural network is good at learning the representations of nodes and edges from complex graph structures and can capture the potential relationships between nodes. Therefore, combining the hidden state of the graph neural network to calculate the similarity between nodes can make full use of the learning ability of the graph neural network and improve the accuracy and robustness of similarity calculation.

[0128] S52. Based on the heterogeneous graph spectrum growth algorithm, introduce an adversarial verification mechanism to optimize the self-organization of knowledge topology. The adversarial verification mechanism detects and repairs logical contradictory edges in the spectrum through a generative adversarial network.

[0129] In a knowledge graph, the polarity of a triple usually refers to whether the relationship in the triple is positive or negative. A positive triple indicates that there is a certain relationship between two entities, while a negative triple indicates that there is no certain relationship between two entities. And the present invention can identify the polarity of a triple by introducing an adversarial verification mechanism.

[0130] In one embodiment, based on the heterogeneous graph spectrum growth algorithm, the generative adversarial network detects logical contradictory edges. Let the generator generate potential contradictory edges, and the discriminator judges whether the edge is a logical contradiction. The loss function of the generator is:

[0131] Furthermore, the loss function of the discriminator is:

[0132] where is the loss function of the generator, is the loss function of the discriminator, denotes taking the expectation over samples from the distribution , denotes taking the expectation over samples from the distribution , is the distribution of real edges, is the distribution of edges generated by the generator. is the sample of the discriminant output.

[0133] It should be noted that the goal of the generator is to generate some possible logically contradictory edges, which may be negative triples. The goal of the discriminator is to distinguish between real positive triples and negative triples generated by the generator, that is, to play a role in identifying the polarity of triples.

[0134] Furthermore, for the edges identified as logically contradictory by the discriminator : The present invention innovatively proposes a repair function, and the repair function satisfies the following conditions:

[0135] Among them, is the repair function, is the repair threshold, and are the neighbor sets of nodes and node respectively, represents the empty set, indicating that in the case, that is, in other cases, no repair is performed, is any node in the neighbor node set of node , is any node in the neighbor node set of node .

[0136] It should be noted that the present invention combines an adversarial verification mechanism and a similarity metric to dynamically detect and repair logically contradictory edges in the graph, and this method is not disclosed in the prior art. In addition, if there are incorrect edge connections in the knowledge topology self-organization, the connection can also be repaired through this repair algorithm, so as to achieve correct edge connection.

[0137] Furthermore, through the above adversarial verification mechanism and repair function, the polarity of triples is dynamically detected and repaired, so as to optimize the knowledge topology self-organization.

[0138] S6. Generate an autism spectrum disorder knowledge graph through the knowledge topology self-organization.

[0139] In one embodiment, first, a multi-granularity interpretation engine is constructed through the knowledge topology self-organization.

[0140] Specifically, a multi-level decision-making path is established. The multi-level includes a gene pathway layer, a neural circuit layer, and a behavioral phenotype layer. The gene pathway layer is used to analyze gene-gene interactions and gene-protein interactions and identify key gene pathways. The neural circuit layer constructs a brain region-brain region connection model based on neuroimaging data and identifies key neural circuits. The behavioral phenotype layer is used to analyze behavioral data and identify behavioral phenotypes related to ASD.

[0141] Furthermore, a tracing algorithm is designed to trace the association path between gene pathways, neural circuits, and behavioral phenotypes. The expression of the tracing algorithm is as follows: , where represents the gene pathway , the neural circuit , and the behavioral phenotype ; the degree of association between them, represents the number of association paths, represents the th weight of the path, represents the th correlation coefficient between the gene pathway and the neural circuit on the th path, represents the th correlation coefficient between the neural circuit and the behavioral phenotype ; ; respectively represent the variances of the gene pathway , the neural circuit , and the behavioral phenotype .

[0142] It should be noted that compared with the prior art, the present invention realizes cross-level association analysis from gene pathways to neural circuits and then to behavioral phenotypes by constructing a multi-granularity interpretation engine, improving the flexibility and accuracy of the analysis. On the basis of the knowledge topology, a traceable decision-making path is established, providing a powerful tool for in-depth understanding of the association between genes, nerves, and behaviors.

[0143] Furthermore, the individual development stages are divided according to age and gender factors, and the data of different development stages are labeled to ensure the accuracy of personalized diagnosis.

[0144] Furthermore, a segmentation algorithm is designed to segment the knowledge graph in real time according to the individual development stage to generate a personalized knowledge graph subgraph. The expression of the segmentation algorithm is as follows:

[0145] Among them, represents the th individual (or vertex), represents the individual and the edge between, represents the edge weight, represents the age factor used to measure the change in the degree of association between individuals and due to age differences, represents the gender factor used to measure the change in the degree of association between individuals and due to gender differences, represents the total number of edges of the individual , that is, the number of edges directly connected to , represents the balance factor used to measure the load balance degree of each sub-graph after segmentation, which is determined by calculating the number of individuals and the number of edges in the sub-graph, represents the load in the sub-graph, including the number of individuals and the number of edges.

[0146] Traditional graph segmentation algorithms only consider the topological structure of the graph or the weights of the edges, while ignoring the attributes of the individuals themselves, including age and gender. The algorithm of the present invention comprehensively considers multiple attributes of the individuals by introducing age factors and gender factors, making the segmentation result more in line with the requirements of the actual application scenario; it can dynamically adjust the segmentation strategy according to the real-time data of the individuals to ensure that the generated sub-graphs can always accurately reflect the current state of the individuals.

[0147] Furthermore, using the high-performance graph database Neo4j, the final autism spectrum disorder knowledge graph is generated. The graph includes a complete graph and sub-graphs, and the graph is mainly composed of entity node information and relationship information. The entity nodes are finally linked by relationships to form a visualized networked intertwined structure and continuously expand outwards. As the data source is supplemented and updated, the number of relevant entity nodes and the number of link relationships will also increase, and the generated autism spectrum disorder knowledge graph will also change dynamically. For the complete graph in the embodiment of the present invention, please refer to Figure 2 , each circle in the figure represents 1 entity, and the entities are connected by relationships. The relationships include symptom (SYM), treatment (TR), method (MT), and confusion (EC). The present invention uses the high-performance graph database Neo4j to achieve knowledge storage and visualization.

[0148] Please refer to Figure 3 , Figure 3This is a schematic structural diagram of the autism spectrum disorder knowledge graph construction system in an embodiment of the present invention. The system includes an input device, a processor, an output device, and a memory. The input device, the processor, the output device, and the memory are interconnected. Among them, the memory is used to store a computer program, and the computer program includes program instructions. The processor is configured to call the program instructions, and the system uses the autism spectrum disorder knowledge graph construction method described above.

[0149] In this embodiment, the input device includes a multi-source data acquisition module, which is used to obtain structured, semi-structured, and unstructured data from medical authoritative books, medical research papers, reports of professional autism research institutions, hospital electronic medical records, patient self-report records, and Internet data resources. The function of the input device is to input the initial information of autism spectrum disorder with multi-source heterogeneity into the system through a data interface, ensuring the comprehensiveness and diversity of the data and providing a basis for subsequent processing.

[0150] Furthermore, the processor is the core computing unit of the system and includes a data preprocessing module, an information extraction module, and a knowledge graph construction module.

[0151] Specifically, the data preprocessing module is responsible for cleaning the input data, including removing duplicates, filling in missing values, correcting errors, and unifying formats; the information extraction module uses a bidirectional gated graph convolutional network, a bidirectional long short-term memory network, and an improved hidden Markov - conditional random field joint model to perform text classification, entity extraction, attribute extraction, and relationship extraction, generating a set of triples; the knowledge graph construction module completes the dynamic update and optimization of the knowledge graph based on a credibility dynamic evolution mechanism and knowledge topology self-organization optimization. The function of the processor is to realize the full-process automation processing from raw data to structured knowledge, ensuring the accuracy and real-time nature of the knowledge graph.

[0152] Furthermore, the output device includes a visualization interface and an API interface.

[0153] The visualization interface displays the autism spectrum disorder knowledge graph in a graphical manner, supporting user interactive queries and explorations; the API interface provides knowledge graph data access services for external systems, such as medical diagnosis platforms and scientific research analysis tools. The function of the visualization interface is to output the constructed knowledge graph in an intuitive and operable form to meet the needs of different users.

[0154] Further, the memory includes an original corpus, an intermediate data repository, and a knowledge graph database. The original corpus stores preprocessed multi-source heterogeneous data; the intermediate data repository is used to save the classification process data during the information extraction process; the knowledge graph database stores the finally constructed autism spectrum disorder knowledge graph and its dynamic evolution record data. The function of the memory is to provide efficient data storage and management, support the full-process data processing of the system, and the long-term maintenance and update of the knowledge graph.

[0155] In summary, the present invention realizes efficient and accurate information extraction through a hybrid convolutional neural network-long short-term memory network-hidden Markov model, combines a credibility dynamic evolution mechanism to optimize the triple weights in real time, and realizes the intelligent growth and verification of the graph based on the knowledge topology self-organization algorithm. Compared with the prior art, this method significantly improves the parsing efficiency of unstructured medical texts, realizes the dynamic iteration of credibility through the Markov decision process, and uses the adversarial verification mechanism to reduce the need for manual intervention. The finally constructed knowledge graph supports multi-level association tracking of gene pathways, neural circuits, and behavioral phenotypes, providing all-dimensional and highly reliable knowledge support for the precise diagnosis and treatment and mechanism research of ASD.

[0156] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention, and they should all be covered by the scope of the claims and the description of the present invention.

Claims

1. A method for constructing a knowledge graph for autism spectrum disorders, characterized in that: The method comprises the following steps: Obtaining multi-source heterogeneous information on autism spectrum disorders; Using the autism spectrum disorder information, constructing an autism spectrum disorder original corpus; According to the autism spectrum disorder original corpus, a hybrid convolutional neural network-long short-term memory network-hidden Markov model is used to extract information and obtain information extraction results; Based on the information extraction results, a credibility dynamic evolution mechanism is designed; According to the dynamic evolution mechanism of credibility, construct and optimize knowledge topology self-organization; Through the knowledge topology self-organization, an autism spectrum disorder knowledge graph is generated.

2. The method for constructing a knowledge graph for autism spectrum disorders according to claim 1, characterized in that: The obtaining of multi-source heterogeneous autism spectrum disorder information comprises: Obtain structured, semi-structured and unstructured initial information on autism spectrum disorders from authoritative medical books, medical research papers, reports from professional autism research institutions, hospital electronic medical records, patient self-report records and Internet data resources; The initial information of autism spectrum disorder is preprocessed to obtain processed multi-source heterogeneous autism spectrum disorder information, wherein the preprocessing includes removing duplicate data, filling missing values, correcting erroneous data and unifying data formats.

3. The method for constructing a knowledge graph for autism spectrum disorders according to claim 1, characterized in that: The step of using the autism spectrum disorder information to construct an autism spectrum disorder original corpus comprises: Classify and store the pre-processed autism spectrum disorder information according to data types; Based on the classified storage, natural language processing is performed on the text data to form structured text data, wherein the natural language processing includes word segmentation, part-of-speech tagging and named body recognition; The structured text data, semi-structured data and unstructured data are integrated to construct an autism spectrum disorder original corpus.

4. The method for constructing a knowledge graph for autism spectrum disorders according to claim 1, characterized in that: The method of extracting information based on the original corpus of autism spectrum disorder and obtaining the information extraction result by using a hybrid convolutional neural network-long short-term memory network-hidden Markov model includes: Based on the autism spectrum disorder original corpus, a bidirectional gated graph convolutional network, a bidirectional long short-term memory network and an improved hidden Markov-conditional random field joint model are established; The bidirectional gated graph convolutional network, the bidirectional long short-term memory network and the improved hidden Markov-conditional random field joint model are used to extract information and obtain information extraction results.

5. The method for constructing a knowledge graph for autism spectrum disorders according to claim 4, characterized in that: The method of using the bidirectional gated graph convolutional network, the bidirectional long short-term memory network and the improved hidden Markov-conditional random field joint model to extract information and obtain the information extraction result includes: Using the bidirectional gated graph convolutional network, classify the preprocessed text data to obtain a classification result; Based on the classification result, using the bidirectional long short-term memory network to perform enhanced classification to obtain an enhanced classification result; Based on the enhanced classification result, the text data is reclassified using the improved hidden Markov-conditional random field joint model to obtain a reclassification result; According to the reclassification result, the classification iteration of the text data is performed to complete information extraction and obtain a triple set, wherein the information extraction includes entity extraction, attribute extraction and relationship extraction.

6. The method for constructing a knowledge graph for autism spectrum disorders according to claim 5, characterized in that: The improved hidden Markov-conditional random field joint model satisfies the following expression: , in, For the observation sequence The state transition energy function of the hidden state sequence, is the sequence length, is the hidden state transfer matrix, indicating the Status Transfer to Status The probability of is a newly added position-sensitive feature function, is the original characteristic function, , , is the adaptive weight coefficient, is the total number of position-sensitive eigenfunctions, is the total number of original characteristic functions, , , are the indexes of position-sensitive feature function, original feature function and interactive feature function respectively, is the total number of interactive characteristic functions, is the interaction feature function.

7. The method for constructing a knowledge graph for autism spectrum disorders according to claim 1, characterized in that: The design of the dynamic evolution mechanism of credibility based on the information extraction result includes: Based on the information extraction results, a multi-dimensional credibility evaluation system is established; According to the multi-dimensional credibility evaluation system, the credibility of the triples is iteratively optimized through a Markov decision process.

8. The method for constructing a knowledge graph for autism spectrum disorders according to claim 1, characterized in that: The construction and optimization of knowledge topology self-organization according to the dynamic evolution mechanism of credibility includes: Develop a heterogeneous graph growth algorithm based on the dynamic evolution mechanism of credibility; Using the heterogeneous graph growth algorithm, constructing knowledge topology self-organization; Based on the knowledge topology self-organization, an adversarial verification mechanism is introduced to identify the polarity of the triples and obtain the identification results; The knowledge topology self-organization is optimized according to the recognition result.

9. The method for constructing a knowledge graph for autism spectrum disorders according to claim 1, characterized in that: The generating of the autism spectrum disorder knowledge graph through the knowledge topology self-organization includes: Through the self-organization of the knowledge topology, a multi-level decision path is established; According to the multi-level decision path, a tracking algorithm is designed, and the expression of the tracking algorithm is as follows: , in, Indicates gene pathway , neural circuits , behavioral phenotype The correlation between represents the number of associated paths, Indicates The weight of the path, Indicates Gene pathways , neural circuits The correlation coefficient between Indicates Neural circuits on the path and behavioral phenotypes The correlation coefficient between , , Represents gene pathway , neural circuits and behavioral phenotypes The variance of Based on the tracking algorithm, a segmentation algorithm is designed; By means of the tracking algorithm and the segmentation algorithm, an autism spectrum disorder knowledge graph is generated.

10. A system for constructing a knowledge graph for autism spectrum disorders, the system using a method for constructing a knowledge graph for autism spectrum disorders according to any one of claims 1 to 9, characterized in that: The system includes an input device, a processor, an output device and a memory, wherein the input device, the processor, the output device and the memory are connected to each other, wherein the memory is used to store a computer program, the computer program includes program instructions, and the processor is configured to call the program instructions.

Citation Information

Patent Citations

  • Knowledge graph construction method and system capable of distinguishing uniphasic and biphasic affective disorder

    CN115630697A

  • Text classification method and device, equipment and storage medium

    CN115640399A

  • Complex action recognition method and device based on learnable Markov logic network

    CN116469155A

Cited By

  • Rare disease knowledge graph construction method based on modal injection and multi-modal fusion

    CN120806104A

  • Forest fire knowledge modeling method based on named entity recognition and relation extraction

    CN121524352A