Method, device, and medium for mining copy number variation data for input text

By splicing the context and querying information, using pre-trained language model and multi-head attention network, the problem of inefficient extraction of complex CNV description forms in traditional methods is solved, and accurate and efficient copy number variation data extraction and normalization processing is achieved.

CN119903177BActive Publication Date: 2025-07-08BERRYGENOMICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510408379.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-08
Estimated Expiration
2045-04-02

AI Technical Summary

Technical Problem

Traditional methods are difficult to accurately and efficiently extract copy number variation data in complex CNV description forms (including chromosome bands and coordinates), and rely on manual intervention and regular expression rules, which is inefficient.

Method used

By splicing the context information and query information of copy number variation to generate input text sequences, using pre-trained language model and multi-headed attention network to generate output vectors, predict the start and end positions of copy number variation entities, and normalized processing is performed by combining semantic analysis and regular matching to generate standard CNV result representations.

Benefits of technology

Accurate and efficient extraction of complex CNV description forms is achieved, reducing noise interference, improving extraction efficiency, and supporting data consistency and biological interpretation of multigenomic versions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119903177B_ABST
    Figure CN119903177B_ABST
Patent Text Reader

Abstract

The present invention relates to a method, device, and medium for mining copy number variation data for an input text. The method includes: splicing the obtained context information and query information regarding copy number variation to generate an input text sequence; converting the context vocabulary sequence and query vocabulary sequence into word index sequences to generate an input vector sequence based at least on the word index sequences; inputting the input vector sequence into a pre-trained language model to generate output vectors through stacked multi-head attention and feed-forward networks, where the output vectors include vector context representations of each vocabulary; and predicting the probability of each vocabulary as the start or end of an entity based on the context representation of each vocabulary to determine copy number variation entities for generating a result representation regarding copy number variation. The present invention can accurately and efficiently extract CNV data for an input text with complex CNV description forms.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention generally relates to data mining of bioinformatics, and specifically, to a method, a computing device, and a computer storage medium for mining copy number variation data for input text. Background Art

[0002] Copy number variation (abbreviated as "CNV") refers to a deletion or amplification of a DNA fragment with a length of not less than 1 kbp compared to the reference genome. CNV detection and normalization provide an important basis for precision medicine.

[0003] Traditional methods for mining CNV data for input text mainly rely on rule matching and / or traditional machine learning, and need to perform conversions for different genome versions and data formats. Among them, traditional methods include, for example: mainly focusing on the description form of point mutations, and normalizing the chromosomal coordinate form of CNV by combining machine learning and regular expressions. However, through the analysis of the titles and abstracts of more than 30 million relevant documents, it can be seen that there are about 6 million documents with descriptions of CNV. Among them, the number of documents with coordinate descriptions in the titles and abstracts does not exceed 100,000. Most of the documents basically describe CNV in the form of cytobands (also known as "cell bands"). The above traditional methods that highly rely on regular expression rules for mutation extraction and normalization are difficult to accurately extract complex CNV description forms (including cytobands and coordinates), and require manual intervention, with low efficiency.

[0004] In summary, the deficiencies of traditional methods for mining copy number variation data for input text are as follows: it is difficult to accurately and efficiently extract input text with complex CNV description forms (including cytobands and coordinates). Summary of the Invention

[0005] The present invention provides a method, a computing device, and a computer storage medium for mining copy number variation data for input text, which can accurately and efficiently extract CNV data for input text with complex CNV description forms.

[0006] According to a first aspect of the present invention, there is provided a method for mining copy number variation data for input text. The method includes: splicing the obtained context information and query information regarding copy number variation to generate an input text sequence, the input text sequence at least including: a context vocabulary sequence in the context information and a query vocabulary sequence in the query information; converting the context vocabulary sequence and the query vocabulary sequence into word index sequences to generate an input vector sequence at least based on the word index sequences, the input vector sequence including a lexical embedding vector, a position embedding vector, and a segment embedding vector; inputting the input vector sequence into a pre-trained language model to generate an output vector through stacked multi-head attention and feed-forward networks, the output vector including a vector context representation of each vocabulary; and predicting the probability of each vocabulary as the start or end of an entity based on the context representation of each vocabulary to determine copy number variation entities, so as to generate a result representation regarding copy number variation.

[0007] In some embodiments, determining the copy number variation entities includes: calculating a first probability of each vocabulary as the start position of a copy number variation entity to select the vocabulary index corresponding to the maximum value of the first probability as the target start position of the copy number variation entity; calculating a second probability of each vocabulary as the end position of a copy number variation entity to select the vocabulary index corresponding to the maximum value of the second probability as the target end position of the copy number variation entity; and determining the entity range and entity content of the copy number variation entity based on the determined target start position and target end position of the copy number variation entity.

[0008] In some embodiments, generating a result representation regarding copy number variation includes: extracting mutation type entities and genetic pattern entities from the input text based on regular matching; recording the position indexes of each of the mutation type entities, genetic pattern entities, and copy number variation entities in the input text; calculating the text distance between entities; and determining the matching relationship between entities at least based on the calculated text distance between entities.

[0009] In some embodiments, determining the matching relationship between entities at least based on the calculated text distance between entities includes: determining a candidate matching relationship between entities based on the text distance between entities; calculating matching degree characterization data for each candidate matching relationship between entities based on the dependency path weight and the text distance weight; and selecting the candidate matching relationship corresponding to the maximum value of the matching degree characterization data as the target matching relationship set between entities.

[0010] In some embodiments, determining the matching relationship between entities includes: in response to determining that there are candidate matching relationships between multiple variant type entities or multiple genetic pattern entities and the same copy number variant entity, preferentially selecting the candidate matching relationship of the direct dependency modification relationship as the target matching relationship; in response to determining that there is no candidate matching relationship of the direct dependency modification relationship, calculating the matching degree characterization correction data based on the dependency relationship characterization data and the context matching characterization data, so as to select the candidate matching relationship corresponding to the highest value of the matching degree characterization correction data as the target matching relationship; and in response to determining that multiple matching degree characterization data are the same, selecting the candidate matching relationship with the earlier text order as the target matching relationship.

[0011] In some embodiments, the result representation regarding copy number variation includes: regarding the chromosome number, position interval, variant type, genetic pattern, and genome version number of the copy number variant entity, generating the result representation regarding copy number variation includes multiple items among the following: normalizing the genome version number; normalizing the variant type entity; normalizing the genetic pattern entity; and normalizing the position interval unit.

[0012] In some embodiments, normalizing the genome version number includes: mapping the copy number variant entity to the first human reference genome and the second human reference genome respectively; obtaining the information of the corresponding reference genome version extracted from the input text, so as to match the entity coordinates of the copy number variant entity corresponding to the reference genome version; and based on the entity coordinates of the copy number variant entity, generating the normalized result representation regarding copy number variation, the normalized result representation regarding copy number variation includes: the chromosome number regarding copy number variation, the position interval relative to the first human reference genome, and the position interval relative to the second human reference genome.

[0013] In some embodiments, normalizing the variant type entity includes: generating the normalized result representation regarding copy number variation based on the entity coordinates of the copy number variant entity and the matching relationship between the variant type entity and the copy number variant entity; the matching relationship between the variant type entity and the copy number variant entity is generated through the following items: using semantic analysis or natural language processing method based on rule matching to extract the variant type entity from the input text; calculating the text distance between the variant type entity and the copy number variant entity in the input text; determining the matching relationship between the variant type entity with the closest text distance and the copy number variant entity; and in response to determining that there are multiple matching relationships between the variant type entity and the copy number variant entity, selecting among the multiple matching relationships in combination with the context information.

[0014] In some embodiments, normalizing for a position interval unit includes: using a regular expression to extract a position interval description containing a unit from the input text; calculating a standardized position interval length for the extracted position interval description; and combining the standardized position interval length, the variant type, and the reference genome version to generate a normalized result representation for the copy number variation.

[0015] In some embodiments, normalizing for a genetic pattern entity includes: extracting a genetic pattern entity from the input text; classifying the extracted genetic pattern entity into a predetermined genetic pattern classification, where the predetermined classification is one of de novo, paternal inheritance, and maternal inheritance; calculating a text distance between the genetic pattern entity and a copy number variation entity to match the genetic pattern entity corresponding to the minimum text distance with the copy number variation entity; and combining the standardized position interval length, the variant type, the reference genome version, and the predetermined genetic pattern classification to generate a normalized result representation for the copy number variation.

[0016] According to a second aspect of the present invention, there is also provided a computing device, which includes: a memory configured to store one or more computer programs; and a processor coupled to the memory and configured to execute one or more programs to cause the device to perform the method of the first aspect of the present invention.

[0017] According to a third aspect of the present invention, there is also provided a non-transitory computer-readable storage medium. Machine-executable instructions are stored on the non-transitory computer-readable storage medium, and when executed, the machine-executable instructions cause the machine to perform the method of the first aspect of the present invention.

[0018] The summary of the invention is provided to introduce, in a simplified form, a selection of concepts that will be further described in the detailed description below. The summary of the invention is not intended to identify the key features or main features of the present invention, nor is it intended to limit the scope of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 A schematic diagram of a system for implementing a method for mining copy number variation data from an input text according to an embodiment of the present invention is shown.

[0020] Figure 2 A flowchart of a method for mining copy number variation data from an input text according to an embodiment of the present invention is shown.

[0021] Figure 3 A schematic diagram of a method for mining copy number variation data from an input text according to an embodiment of the present invention is shown.

[0022] Figure 4The flowchart of a method for determining copy number variation entities according to an embodiment of the present invention is shown.

[0023] Figure 5 The flowchart of a method for generating a result representation regarding copy number variation according to an embodiment of the present invention is shown.

[0024] Figure 6 The flowchart of a method for determining a matching relationship between entities according to an embodiment of the present invention is shown.

[0025] Figure 7 The flowchart of a method for correcting and disambiguating the matching relationship between entities according to an embodiment of the present invention is shown.

[0026] Figure 8 The flowchart of a method for normalizing mutation interval units according to an embodiment of the present invention is shown.

[0027] Figure 9 The flowchart of a method for normalizing genetic pattern entities according to an embodiment of the present invention is shown.

[0028] Figure 10 The block diagram of an electronic device suitable for implementing the embodiments of the present invention is schematically shown.

[0029] In the respective drawings, the same or corresponding reference numerals denote the same or corresponding parts. Detailed Description of the Embodiments

[0030] The preferred embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although the preferred embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to make the present invention more thorough and complete, and to fully convey the scope of the present invention to those skilled in the art.

[0031] The term "comprising" and its variations used herein mean open inclusion, i.e., "including but not limited to". Unless otherwise stated, the term "or" means "and / or". The term "based on" means "at least partially based on". The term "an example embodiment" and "an embodiment" mean "at least one example embodiment". The term "another embodiment" means "at least one additional embodiment". The terms "first", "second", etc. may refer to different or the same objects.

[0032] As described above, traditional methods for mining copy number variation data from input text in the prior art mainly focus on the description form of point mutations, making it difficult to comprehensively cover more complex genetic variations such as CNVs. Additionally, their description of chromosomal band forms lacks effective normalization capabilities, so they can only process CNV descriptions in coordinate form, resulting in incomplete information extraction. Moreover, traditional methods highly rely on regular expression rules for mutation extraction and normalization, making it difficult to capture cross-sentence or paragraph-level context information in complex texts and requiring manual intervention, with low efficiency. Therefore, the deficiencies of traditional methods for mining copy number variation data from input text are as follows: it is difficult to accurately and efficiently extract input text with complex CNV description forms (including chromosomal bands and coordinates).

[0033] To at least partially address one or more of the above problems and other potential problems, example embodiments of the present invention propose a method for mining copy number variation data from input text. By splicing the obtained context information and query information regarding copy number variation to generate an input text sequence, converting the context vocabulary sequence and query vocabulary sequence into word index sequences, and generating an input vector sequence including word embedding vectors, position embedding vectors, and segment embedding vectors based at least on the word index sequences, and inputting the input vector sequence into a pre-trained language model to generate output vectors through stacked multi-head attention and feed-forward networks, the present invention can make each output vector contain both the semantic information of the word itself and the global information of the context and query. Moreover, it can capture the position information of the word in the sentence, enhancing the model's perception of order. Therefore, even for entities with rare or complex expressions, noise interference can be effectively reduced. Additionally, by predicting the probability of each word as the start or end of an entity, and then determining the entity range of the copy number variation entity to generate a result representation of the copy number variation, the present invention can accurately extract relevant data of the copy number variation entity. Therefore, the present invention can accurately and efficiently extract CNV data from input text with complex CNV description forms.

[0034] Figure 1 FIG. shows a schematic diagram of a system 100 for implementing a method for mining copy number variation data from input text according to an embodiment of the present invention. As Figure 1 shown, the system 100 includes: a computing device 110, a server 130. In some embodiments, the computing device 110 and the server 130 perform data interaction directly or via a network (not shown).

[0035] Regarding the server 130, it is used, for example, to provide input text. The input text is, for example, biomedical text.

[0036] Regarding computing device 110, it is used to mine copy number variation data for input text. Specifically, computing device 110 is used to splice the obtained context information and query information regarding copy number variation to generate an input text sequence, where the input text sequence at least includes: the context vocabulary sequence in the context information and the query vocabulary sequence in the query information; and to convert the context vocabulary sequence and the query vocabulary sequence into a word index sequence, so as to generate an input vector sequence based at least on the word index sequence, where the input vector sequence includes a vocabulary embedding vector, a position embedding vector, and a segment embedding vector. Computing device 110 is further used to input the input vector sequence into a pre-trained language model to generate an output vector through stacked multi-head attention and feed-forward networks, where the output vector includes a vector context representation of each vocabulary; and to predict the probability of each vocabulary as the start or end of an entity based on the context representation of each vocabulary, so as to determine a copy number variation entity for generating a result representation regarding copy number variation.

[0037] In some embodiments, computing device 110 may have one or more processing units, including dedicated processing units such as GPUs, FPGAs, and ASICs, as well as general-purpose processing units such as CPUs. Additionally, one or more virtual machines may also be running on each computing device. Computing device 110 includes, for example: an input text sequence generation unit 112, an input vector sequence generation unit 114, an output vector generation unit 116, and a result representation generation unit 118 for copy number variation. The above input text sequence generation unit 112, input vector sequence generation unit 114, output vector generation unit 116, and result representation generation unit 118 for copy number variation may be configured on one or more computing devices 110.

[0038] Regarding input text sequence generation unit 112, it is used to splice the obtained context information and query information regarding copy number variation to generate an input text sequence, where the input text sequence at least includes: the context vocabulary sequence in the context information and the query vocabulary sequence in the query information.

[0039] Regarding input vector sequence generation unit 114, it is used to convert the context vocabulary sequence and the query vocabulary sequence into a word index sequence, so as to generate an input vector sequence based at least on the word index sequence, where the input vector sequence includes a vocabulary embedding vector, a position embedding vector, and a segment embedding vector.

[0040] Regarding output vector generation unit 116, it is used to input the input vector sequence into a pre-trained language model to generate an output vector through stacked multi-head attention and feed-forward networks, where the output vector includes a vector context representation of each vocabulary.

[0041] A result representation generation unit 118 for copy number variation, which is used to predict the probability of each word as the start or end of an entity based on the context representation of each word, so as to determine copy number variation entities and generate a result representation regarding copy number variation.

[0042] Figure 2 FIG. 200 shows a flowchart of a method 200 for mining copy number variation data based on an input text according to an embodiment of the present invention. It should be understood that the method 200 can be executed, for example, at Figure 10 the described electronic device 1000. It can also be executed at Figure 1 the described computing device 110. It should be understood that the method 200 may further include additional actions not shown and / or may omit the actions shown, and the scope of the present invention is not limited in this regard.

[0043] At step 202, the computing device 110 splices the obtained context information and query information regarding copy number variation to generate an input text sequence, which at least includes: a context word sequence in the context information and a query word sequence in the query information.

[0044] The context information regarding copy number variation, which is, for example, a context sentence in a biomedical text regarding copy number variation.

[0045] The query information regarding copy number variation, which is, for example, a query sentence regarding copy number variation. Regarding the query word sequence, it is, for example, a word sequence in the query sentence, which is used to guide the model to focus on a specific entity (e.g., a copy number variation entity).

[0046] Regarding the input text sequence, it at least includes: a context word sequence in the context information and a query word sequence in the query information. In some embodiments, the input text sequence includes: position identification information, a context word sequence, segmentation identification information, and a query word sequence. For example, the following expression (1) shows an example of the input text sequence.

[0047] {[CLS], x1, x2, ..., xN, [SEP], q1, q2, ..., qM, [SEP]} (1)

[0048] In the above expression (1), [CLS] represents the position identification information of the input text sequence. [SEP] represents the segmentation identification information used to separate the context information and the query information. x1, x2, ..., xN represent the context word sequence. q1, q2, ..., qM represent the query word sequence.

[0049] At step 204, computing device 110 converts the context vocabulary sequence and the query vocabulary sequence into word index sequences, so as to generate an input vector sequence based at least on the word index sequences. The input vector sequence includes word embedding vectors, position embedding vectors, and segment embedding vectors. For example, computing device 110 maps each word in the context vocabulary sequence and the query vocabulary sequence to a discrete word index sequence through a vocabulary.

[0050] For example, as Figure 3 shown, computing device 110 converts the context vocabulary sequence 310 (e.g., Token_X1, Token_X2, … Token_XN) and the query vocabulary sequence 312 (e.g., Token_Q1, Token_Q2, … Token_QN)) into word index sequences. As Figure 3 shown, the word index sequences also include: position identification information 314, segmentation identification information 316.

[0051] Regarding the input vector sequence, it is, for example, as Figure 3 indicated by the marker 320 in. The input vector sequence is generated based on the word index sequence. The input vector sequence includes, for example, word embedding vectors, position embedding vectors, and segment embedding vectors. For example, the following expression (2) illustrates an example of the input vector sequence.

[0052] Embeddings = [E [CLS] , E Tx1 , E Tx2 … E TxN , E [SEP] , E Tq1 , E Tq2 … E TqM , E [SEP] (2)

[0053] In the above expression (2), E Tx1 , E Tx2 … E TxN and E Tq1 , E Tq2 … E TqM represent word embeddings, which are used to provide semantic information of words. E [SEP] represents segment embedding, which is used to distinguish the context part and the query part in the word embeddings. E [CLS] represents position embedding, which is used to capture the position information of words in a sentence to enhance the model's perception of order. Embeddings represents the input vector sequence.

[0054] At step 206, the computing device 110 inputs the input vector sequence into the pre-trained language model to generate output vectors through stacked multi-head attention and feed-forward networks, and the output vectors include the vector context representations of each vocabulary.

[0055] By adopting the above means, each output vector contains both the semantic information of the vocabulary itself and the global information of the context and the query. Moreover, the long-range dependencies can be captured through the multi-head attention mechanism, which is beneficial to improving the model's ability to understand complex sentence structures.

[0056] Regarding the pre-trained language model, it is, for example but not limited to, the Bidirectional Encoder Representations from Transformers for Biomedical Text Mining (BioBERT) model. The BioBER model is mainly a pre-trained language representation model constructed based on the BERT (Bidirectional Encoder Representations from Transformers) model. The BioBERT model is composed of a stack of multiple layers of Transformer encoders. Each layer contains a self-attention mechanism and a feed-forward neural network, which can help the BioBERT model capture complex dependencies in the text. It should be understood that the BioBERT model can consider the context information on both the left and the right during pre-training. Therefore, it is beneficial to understand the context of each word more comprehensively. By using the pre-trained BioBERT model to generate output vectors, the present invention is beneficial to further improving the representation ability of biomedical terms in the input text.

[0057] In some embodiments, the pre-trained language model can also be one of the ClinicalBERT model, the SciBERT model, the BlueBERT model, and the MedBERT model. Among them, the ClinicalBERT model is a model pre-trained with a large amount of medical data, which can more accurately capture and understand the semantic information in medical texts. The SciBERT model is a BERT model pre-trained based on 1.14 million scientific papers (covering the fields of biomedicine and computer science), which focuses on scientific text processing. The BlueBERT model is further optimized and trained on biomedical corpora such as Public Medicine (PubMed) abstracts and clinical notes based on the pre-trained parameters of BERT. The MedBERT model is a pre-trained model designed specifically for medical text processing and is deeply trained on medical text datasets such as PubMed and WebMD (an online medical health information service platform in the United States).

[0058] For example, the computing device 110 uses the BioBERT model to process the input vector sequence 320 encoded at step 204, and generates an output vector 322 through a stacked multi-head attention and feed-forward network. The output vector 322 includes the vector context representation of each word. It should be understood that each output vector contains both the semantic information of the vocabulary itself and the global information of the context and query. It should be understood that in the few-shot scenario, the generalization ability and localization accuracy of the model are improved through the BioBERT and query enhancement methods.

[0059] Regarding the output vector, it is exemplified as shown in the following expression (3).

[0060] [H [CLS] ,H Tx1 ,H Tx2 …H TxN ,H [SEP] ,H Tq1 ,H Tq2 …H TqM ,H [SEP] (3)

[0061] In the above expression (3), H Tx1 ,H Tx2 …H TxN represents the vector context representation of the context vocabulary. H Tq1 ,H Tq2 …H TqM represents the vector context representation of the query vocabulary. H [SEP] represents the segment vector. [CLS]Represents a position vector. It should be understood that this output vector contains both the semantic information of the vocabulary itself and integrates the global information of the context and the query. Additionally, long-range dependencies are captured through the multi-head attention mechanism to improve the pre-trained language model's ability to understand complex sentence structures. The query sentence can guide the model to focus on specific entity types, thereby enhancing the ability to identify rare or ambiguous entities.

[0062] At step 208, computing device 110 predicts the probability of each vocabulary as the start or end of an entity based on the context representation of each vocabulary to determine copy number variation entities for generating a result representation regarding copy number variation.

[0063] Regarding the method for determining copy number variation entities, for example, it includes: calculating the first probability of each vocabulary as the start position of a copy number variation entity to select the vocabulary index corresponding to the maximum value of the first probability as the target start position of the copy number variation entity; calculating the second probability of each vocabulary as the end position of a copy number variation entity to select the vocabulary index corresponding to the maximum value of the second probability as the target end position of the copy number variation entity; and determining the entity range and entity content of the copy number variation entity based on the determined target start position and target end position of the copy number variation entity. The following will be combined with Figure 4 Specifically illustrate method 400 for determining the entity range and entity content of copy number variation entities. Here, it will not be elaborated further.

[0064] It should be understood that normalizing the result representation of copy number variation is a crucial task, whose goal is to convert the diverse result representation formats of copy number variation into a consistent and standard format. This normalization process not only improves the usability and comparability of data but also lays a solid foundation for subsequent functional analysis and biological interpretation.

[0065] Regarding the result representation of copy number variation entities, for example, it includes: the chromosome number, position interval, variation type, inheritance pattern, and genome version number of the copy number variation entity. For example, the result representation regarding copy number variation is: [chromosome number: position interval_variation type_inheritance pattern_hg19, chromosome number: position interval_variation type_inheritance pattern_hg38], where hg19 and hg38 respectively represent different genome version numbers.

[0066] Methods for generating result representations regarding copy number variations, for example, include: normalization of result representations regarding copy number variations. Specifically, it includes multiple of the following items: normalization processing for genome version numbers; normalization processing for variant type entities; normalization processing for inheritance pattern entities; and normalization processing for position interval units. It should be understood that the standardization and multi-dimensional (for example, including multi-dimensions such as genome version numbers, variant types, variant inheritance patterns, and variant interval units) normalization of result representations regarding copy number variations not only enhance the data sharing and reuse capabilities but also promote cross-domain collaborative research, providing important support for precision medicine and population genetics research.

[0067] The following explains the method for normalization processing for genome version numbers. It should be understood that the genome version (such as hg19 or hg38) may not be explicitly mentioned in biomedical texts, and even the descriptions of some copy number variations are far from the genome version description. Additionally, there are significant differences in the coordinate systems of different genome versions, and the coordinates of the same genomic region may change in the coordinate systems of different genome versions, which will directly affect the precise positioning of variants. Therefore, it is necessary to normalize the genome version numbers of copy number variation entities.

[0068] Methods for normalizing genomic version numbers, for example, include: mapping copy number variation entities to a first human reference genome and a second human reference genome respectively; obtaining information on the corresponding reference genome version extracted from the input text to match the entity coordinates of the copy number variation entities corresponding to the reference genome version; and generating a normalized result representation of the copy number variation based on the entity coordinates of the copy number variation entities, where the normalized result representation of the copy number variation includes: the chromosome number of the copy number variation, the position interval relative to the first human reference genome, and the position interval relative to the second human reference genome. In some embodiments, the first human reference genome is, for example, hg19. The second human reference genome is, for example, hg38. It should be understood that for copy number variation entities in chromosomal band description format, they describe genomic regions rather than precise entity coordinates and do not depend on genomic version information. Therefore, during normalization processing, the identified CNV entities need to be mapped to two common genomic versions, hg19 and hg38, respectively, to support CNV analysis across genomic versions. For chromosomal coordinate format, since the chromosomal coordinates directly depend on genomic version information, the computing device 110 preferentially matches the chromosomal coordinate format of the corresponding version based on the reference genome version mentioned in the literature (such as explicitly noted as hg19 or hg38). It should be understood that by adopting the above means, the present invention can support both hg19 and hg38 genomic versions simultaneously, avoid the deviation of the result representation of copy number variation caused by inconsistent reference genome versions, and thus enhance the generality of the result representation of the copy number variation after normalization. Moreover, through the normalization processing with dual reference genome versions, the present invention can also ensure data integrity and provide reliable positioning support even in the absence of version information.

[0069] For example, the normalized result representation after genomic version number normalization is: [chromosome number: position interval_hg19, chromosome number: position interval_hg38]. In some embodiments, the extracted CNV entity is, for example: 5q35.2 --> q35.3. Then the normalized result representation after genomic version number normalization is: ['chr5:172800000-180915260_hg19', 'chr5:173300000-181538259_hg38'].

[0070] The following describes the method for normalizing the variant type.

[0071] Methods for normalizing variant types, for example, include: generating a normalized result representation of copy number variations based on the entity coordinates of copy number variation entities and the matching relationship between variant type entities and copy number variation entities; the matching relationship between the variant type entities and the copy number variation entities is generated through the following: using semantic analysis or natural language processing methods based on rule matching to extract variant type entities from the input text; calculating the text distance between the variant type entities and the copy number variation entities in the input text; determining the matching relationship between the variant type entity with the closest text distance and the copy number variation entity; and in response to determining that there are multiple matching relationships between the variant type entity and the copy number variation entity, selecting among the multiple matching relationships in combination with context information. For example, the normalized result representation formed after variant type normalization is, for example: [chromosome number: position interval_variant type_hg19, chromosome number: position interval_variant type_hg38]. In some embodiments, the input text is, for example, "A deletion was detected in chromosome region 5q35.2 --> q35.3." The extracted CNV entity is, for example, "5q35.2 --> q35.3". The variant type is, for example, deletion (i.e., "deletion"). The normalized result representation formed after variant type normalization is, for example: ['chr5:172800000-180915260_del_hg19', 'chr5:173300000-181538259_del_hg38'].

[0072] It should be understood that the CNV entities and their variant type entities (e.g., "deletion" or "duplication") described in biomedical texts are often far apart, and directly matching the CNV entity with its variant type may lead to incorrect associations. Additionally, in some biomedical texts, the variant type entity is not explicitly indicated, but is expressed through context or implicitly, which increases the difficulty of extracting the variant type entity. In large-scale biomedical text data processing, efficient and automated methods are needed to accurately associate the variant type entity with the CNV entity.

[0073] In the above approach, by independently extracting variant type entities and using a distance judgment strategy for variant types and CNV entities, the present invention significantly reduces the probability of incorrectly associating variant types with CNV entities due to dispersed or implicitly expressed variant type information. Additionally, in cases where variant type information is missing or there are multiple possibilities, optimizing in combination with context information can ensure the biological relevance and accuracy of the results.

[0074] Regarding the method for normalizing the variant genetic pattern and the variant interval unit, the following will be described separately in conjunction with Figure 8 , Figure 9 and will not be elaborated here.

[0075] In the above solution, by splicing the obtained context information and query information about copy number variation to generate an input text sequence; converting the context vocabulary sequence and the query vocabulary sequence into word index sequences to generate an input vector sequence including at least a lexical embedding vector, a position embedding vector, and a segment embedding vector based on the word index sequences; and inputting the input vector sequence into a pre-trained language model to generate an output vector through stacked multi-head attention and feed-forward networks, the present invention can make each output vector contain both the semantic information of the vocabulary itself and integrate the global information of the context and the query. Moreover, it can capture the position information of the vocabulary in the sentence and enhance the model's perception of the order. Therefore, even for entities with rare or complex expressions, it can effectively reduce noise interference. Additionally, by predicting the probability of each vocabulary as the start or end of an entity, and then determining the entity range of the copy number variation entity to generate a result representation of the copy number variation, the present invention can accurately extract the relevant data of the copy number variation entity. Therefore, the present invention can accurately and efficiently extract CNV data for input texts with complex CNV description forms.

[0076] Figure 4 FIG. shows a flowchart of a method 400 for determining a copy number variation entity according to an embodiment of the present invention. It should be understood that the method 400 can be executed, for example, at the Figure 10 described electronic device 1000. It can also be executed at the Figure 1 described computing device 110. It should be understood that the method 400 may further include additional actions not shown and / or may omit the actions shown, and the scope of the present invention is not limited in this regard.

[0077] In step 402, the computing device 110 calculates a first probability of each vocabulary as the start position of a copy number variation entity to select the vocabulary index corresponding to the maximum value of the first probability as the target start position of the copy number variation entity.

[0078] As Figure 3As shown, based on the vector context representation of each term in the output vector 322, the computing device 110 predicts, via the linear connection network layer 340 and the Softmax network layer 350, the first probability of each term being the starting position of a copy number variation entity; selects the term index corresponding to the maximum value of the first probability as the target starting position of the copy number variation entity. For example, the output vector includes the vector context representations of N terms. The corresponding N predicted first probabilities for these N terms being the starting positions of the copy number variation entity respectively (the N first probabilities are, for example, L(start)1, L(start)2 … L(start) N ). For example, it is determined through the argmax function that the value of the second first probability (i.e., L(start)2) is the largest. Then the computing device 110 selects the term index corresponding to L(start)2 as the target starting position of the copy number variation entity. The argmax function is a function that finds the argument (set) of a function.

[0079] The function for predicting the first probability of each term being the starting position of a copy number variation entity will be described below in conjunction with Expression (4).

[0080]

[0081] In the above Expression (4), represents the context representation of the i-th token in the output vector. represents the weight of the function for predicting the first probability. represents the bias function of the function for predicting the first probability. represents the function for predicting the first probability that the i-th token in the predicted output vector is the starting position of a copy number variation entity. It should be understood that the i-th token corresponds to the i-th term.

[0082] At step 404, the computing device 110 calculates the second probability of each term being the ending position of a copy number variation entity in order to select the term index corresponding to the maximum value of the second probability as the target ending position of the copy number variation entity.

[0083] For example, the computing device 110 predicts, via the linear connection network layer 340 and the Softmax network layer 350, the second probability of each term being the ending position of a copy number variation entity; selects the term index corresponding to the maximum value of the second probability as the target starting position of the copy number variation entity. For example, the N predicted second probabilities for these N terms being the starting positions of the copy number variation entity respectively (the N second probabilities are, for example, L(end)1, L(end)2 … L(end) N). For example, the Nth second probability is selected through the argmax function (i.e., L(end) N ) and the corresponding vocabulary index is used as the target end position of the copy number variation entity.

[0084] The function for predicting the second probability of each vocabulary as the end position of the copy number variation entity will be described below in conjunction with Expression (12).

[0085]

[0086] In the above Expression (12), represents the context representation of the jth token in the output vector. represents the weight of the function for predicting the second probability. represents the bias function of the function for predicting the second probability. represents the function for predicting the second probability that the jth token in the predicted output vector is the end position of the copy number variation entity.

[0087] At step 406, the computing device 110 determines the entity range and entity content of the copy number variation entity based on the determined target start position and target end position of the copy number variation entity.

[0088] For example, the computing device 110 can obtain the entity range and entity content of the copy number variation entity based on the vocabulary index corresponding to the determined target start position of the copy number variation entity and the vocabulary index corresponding to the target end position. In some embodiments, the prediction of the start position and end position can be jointly trained with shared parameters to reduce overfitting.

[0089] Regarding the entity range of the copy number variation entity, it is determined by the vocabulary index corresponding to the predicted target start position (e.g., ) and the vocabulary index corresponding to the target end position (e.g., ). Among them is less than or equal to .

[0090] Regarding the entity content, it is the corresponding substring in the input sequence .

[0091] In some embodiments, the input sequence is, for example, "{[CLS], 'A', 'deletion', 'in', 'chromosome', '5q35.2', '-->', 'q35.3', [SEP], 'What', 'is', 'the', 'CNV','?', [SEP]}". The vocabulary index corresponding to the target start position is, for example: 5 (labeled '5q35.2'). The vocabulary index corresponding to the target end position is, for example: 7 (labeled 'q35.3'). The entity range is: '5q35.2 --> q35.3'.

[0092] By adopting the above means, the present invention can achieve precise boundary recognition and accurate positioning of copy number variation entities.

[0093] Figure 5 A flowchart of a method 500 for generating a result representation regarding copy number variation according to an embodiment of the present invention is shown. It should be understood that the method 500 can be executed, for example, at Figure 10 the described electronic device 1000. It can also be executed at Figure 1 the described computing device 110. It should be understood that the method 500 may further include additional actions not shown and / or may omit the shown actions, and the scope of the present invention is not limited in this regard.

[0094] At step 502, the computing device 110 extracts variation type entities and genetic pattern entities from the input text based on regular matching.

[0095] Regarding the input text, in some embodiments, it is, for example, "Maternally derived terminal deletion of chromosome 7(q35-qter) and a de novo duplication of 11q12."

[0096] Regarding the variation type entities, they include, for example: "deletion" and "duplication".

[0097] Regarding the genetic pattern entities, they include, for example: de novo (i.e., De Novo), paternal inheritance (i.e., Paternal Inheritance), and maternal inheritance (i.e., Maternal Inheritance).

[0098] For example, the copy number variation entities are, for example, 7(q35-qter) (which indicates the part of the long arm (q) of human chromosome 7 from region q35 to the end (qter)), 11q12, etc.

[0099] At step 504, computing device 110 records the position indices of each of the variant type entity, the inheritance pattern entity, and the copy number variation entity in the input text.

[0100] Regarding the position indices of the entities in the input text, it is characterized below in conjunction with expression (5).

[0101]

[0102] In the above expression (5), represents the position index of the entity in the input text. represents the start position index. represents the end position index.

[0103] For example, in the above input text, the extracted variant type entity and its position index are "deletion (position index: [28, 36]), duplication (position index: [77, 88])". The copy number variation entity and its position index are "7(q35 - qter) (position index: [51, 62]), 11q12 (position index: [92, 97])". The inheritance pattern entity and its position index are "Maternally (position index: [0, 10]), de novo (position index: [69, 76])"

[0104] At step 506, computing device 110 calculates the text distance between the entities.

[0105] Regarding the text distance between the entities, its algorithm is described below in conjunction with expression (6).

[0106]

[0107] In the above expression (6), represents the text distance between entities x and y. represents the start position of the index of entity x. represents the end position of the index of entity y.

[0108] At step 508, computing device 110 determines the matching relationship between the entities based at least on the calculated text distance between the entities.

[0109] Regarding the method for determining the matching relationship between the entities, in some embodiments, for example, it includes: if computing device 110 determines that the matching relationship between the entities is less than a predetermined text distance threshold, then it determines that there is a matching relationship between the entities. For example, if the following expression (7) is satisfied, it indicates that there is a matching relationship between the corresponding entities.

[0110]

[0111] In the above expression (7), represents the text distance between entities x and y. In some embodiments, entity x represents, for example, a copy number variation entity; entity y represents, for example, a mutation type entity or a genetic pattern entity. δ represents a predetermined text distance threshold.

[0112] By adopting the above means, the present invention can determine the matching relationship between entities through the text order and index position of the entities.

[0113] Figure 6 FIG. shows a flowchart of a method 600 for determining a matching relationship between entities according to an embodiment of the present invention. It should be understood that method 600 can be executed, for example, at Figure 10 the described electronic device 1000. It can also be executed at Figure 1 the described computing device 110. It should be understood that method 600 may further include additional actions not shown and / or may omit the shown actions, and the scope of the present invention is not limited in this regard.

[0114] At step 602, the computing device 110 determines a candidate matching relationship between entities based on the text distance between the entities.

[0115] At step 604, the computing device 110 calculates match degree characterization data for each candidate matching relationship between entities based on the dependency path weight and the text distance weight.

[0116] Regarding the dependency path weight, the algorithm for calculating the dependency path weight is described below in conjunction with expression (8).

[0117]

[0118] In the above expression (8), represents the dependency path weight. In the above expression (8), represents the dependency path weight.

[0119] Regarding the text distance weight, the algorithm for calculating the dependency path weight is described below in conjunction with expression (9).

[0120]

[0121] In the above expression (9), represents the text distance weight. represents the text distance between entities x and y.

[0122] The following describes the algorithm for calculating the matching degree representation data in conjunction with Expression (10).

[0123]

[0124] In the above Expression (10), represents the matching degree representation data. represents the dependency path weight. represents the text distance weight. represents the weight hyperparameters, and the two satisfy .

[0125] For example, the input text is "Maternally derived terminal deletion of chromosome 7(q35-qter) and a de novo duplication of 11q12". The target matching relationships between the finally determined copy number variation entities and the variation type entities are: "- deletion <-> 7(q35-qter)" and "- duplication<-> 11q12". The target matching relationships between the inheritance pattern entities and the copy number variation entities are: "- maternally <->7(q35-qter)" and "- de novo <-> 11q12". That is, "- maternal inheritance <-> 7(q35-qter)" and "- de novo<-> 11q12".

[0126] At step 606, the computing device 110 selects the candidate matching relationship corresponding to the maximum value of the matching degree representation data as the target matching relationship between the entities.

[0127] By adopting the above means, the present invention can use dependency syntactic parsing to identify grammatical modifications and entity dependency relationships, and optimize the matching relationships between entities.

[0128] Figure 7 shows a flowchart of a method 700 for correcting and disambiguating the matching relationship between entities according to an embodiment of the present invention, that is, a method for further determining the matching relationship between entities. It should be understood that the method 700 can be executed, for example, at Figure 10 the described electronic device 1000. It can also be executed at Figure 1 the described computing device 110. It should be understood that the method 700 may further include additional actions not shown and / or may omit the shown actions, and the scope of the present invention is not limited in this regard.

[0129] At step 702, if the computing device 110 determines that there are multiple variant type entities or multiple genetic pattern entities that have candidate matching relationships with the same copy number variant entity, and there are candidate matching relationships with direct dependency modification relationships, the candidate matching relationships with direct dependency modification relationships are preferentially selected as the target matching relationships.

[0130] At step 704, if the computing device 110 determines that there are no candidate matching relationships with direct dependency modification relationships, match degree characterization correction data is calculated based on dependency relationship characterization data and context matching characterization data, so as to select the candidate matching relationship corresponding to the highest value of the match degree characterization correction data as the target matching relationship.

[0131] Regarding the calculation method of the match degree characterization correction data, it will be described below in conjunction with Expression (11).

[0132]

[0133] In the above Expression (11), represents the dependency relationship characterization data. represents the context matching characterization data. represents the match degree characterization correction data. represents the balance coefficient.

[0134] At step 706, if the computing device 110 determines that there are multiple match degree characterization data that are the same, the candidate matching relationship with the earlier text order is selected as the target matching relationship.

[0135] In the above solution, through the algorithm based on entity position, dependency syntactic analysis, and context relationship, not only can the matching relationship between entities be accurately matched, but also the robustness of the matching algorithm can be improved to adapt to the entity relationship matching of complex text descriptions.

[0136] Figure 8 FIG. shows a flowchart of a method 800 for normalizing a variant interval unit according to an embodiment of the present invention. It should be understood that the method 800 can be executed, for example, at Figure 10 the electronic device 1000 described. It can also be executed at Figure 1 the computing device 110 described. It should be understood that the method 800 may further include additional actions not shown and / or may omit the actions shown, and the scope of the present invention is not limited in this regard.

[0137] It should be understood that when describing the variant interval length of a CNV entity in a biomedical text, multiple units may be used, such as KB and MB. There are inconsistencies in the conversion of these units among different data sources. In addition, the lack of unit uniformity leads to incomparability between data sources, which has a negative impact on the accuracy of the analysis results. Therefore, an efficient method is needed to standardize the interval lengths in different units into a unified format.

[0138] At step 802, the computing device 110 extracts the location interval description containing the unit from the input text using a regular expression.

[0139] In some embodiments, the input text is, for example, "chr19:54.6-54.7 Mb del". The extracted location interval is, for example: 54.6-54.7 Mb.

[0140] At step 804, the computing device 110 calculates the standardized location interval length for the extracted location interval description. For example, all units are converted to BP (base pairs), which is the most commonly used unit in genomic data. For example, the following conversion rules are adopted: 1 KB = 1,000 BP; 1 MB = 1,000,000 BP. It should be understood that by uniformly converting to the BP unit, the data deviation caused by unit differences is eliminated, ensuring the reliability of the analysis results.

[0141] For example, the starting position is: 54.6 MB × 1,000,000 = 54,600,000 BP. The ending position is: 54.7 MB × 1,000,000 = 54,700,000 BP.

[0142] At step 806, the computing device 110 combines the standardized location interval length, the variant type, and the reference genome version to generate a normalized result representation for copy number variation. For example, the normalized result identifier after the normalization process of the variant interval unit is: [chromosome number: location interval_variant type_hg19, chromosome number: location interval_variant type_hg38]. Specifically, it is "['chr19:54600000-54700000_del_hg19', 'chr19:54600000-54700000_del_hg38']".

[0143] In the above solution, the efficient and automatic normalization process of the variant interval unit is achieved through the combination of regular expressions and mathematical calculations, reducing manual intervention; and it is applicable to variant interval descriptions in multiple units and formats, and can handle complex literature data scenarios.

[0144] Figure 9The flowchart of method 900 for normalizing genetic pattern entities according to an embodiment of the present invention is shown. It should be understood that method 900 can be executed, for example, at Figure 10 the electronic device 1000 described. It can also be executed at Figure 1 the computing device 110 described. It should be understood that method 900 may further include additional actions not shown and / or may omit the actions shown, and the scope of the present invention is not limited in this regard.

[0145] It should be understood that CNV is usually closely related to specific genetic patterns (such as de novo, paternal inheritance, maternal inheritance), and the standardization of genetic patterns is crucial for clarifying their biological significance. However, the genetic patterns in biomedical texts may be scattered in different parts of the text, and the direct relevance description with CNV entities is relatively low, which increases the difficulty of extracting genetic pattern entities and matching CNV entities with genetic pattern entities. Therefore, it is necessary to normalize genetic pattern entities.

[0146] At step 902, the computing device 110 extracts genetic pattern entities from the input text.

[0147] At step 904, the computing device 110 classifies the extracted genetic pattern entities into a predetermined genetic pattern classification, and the predetermined classification is one of de novo, paternal inheritance, and maternal inheritance. For example, the computing device 110 standardizes the genetic pattern description extracted from the text into a predetermined genetic pattern classification, which is predefined and is one of de novo (i.e., De Novo), paternal inheritance (i.e., Paternal Inheritance), and maternal inheritance (i.e., Maternal Inheritance).

[0148] At step 906, the computing device 110 calculates the text distance between the genetic pattern entity and the copy number variation entity to match the genetic pattern entity corresponding to the minimum text distance with the copy number variation entity.

[0149] For example, the computing device 110 preferentially selects the genetic pattern entity and the CNV entity corresponding to the nearest distance according to the distance between the genetic pattern entity and the CNV entity in the text. If there are multiple genetic pattern entities in the same text, the genetic pattern entity with a higher semantic relevance to the CNV entity is preferentially selected for association. For example: If both the "maternal inheritance" and "de novo" genetic pattern entities are described in a text, but semantically "de novo" is closer to the CNV entity, then "de novo" is selected as the matching result of the genetic pattern entity associated with the CNV entity. By adopting the above means, the present invention can combine context information and further optimize the matching result by using semantic similarity to ensure that the pattern and the entity have biological relevance.

[0150] At step 908, the computing device 110 combines the standardized position interval length, the variant type, the reference genome version, and the predetermined genetic pattern classification to generate a normalized result representation of the copy number variation.

[0151] The normalized result representation of the copy number variation processed by the genetic pattern entity normalization is, for example: [chromosome number: position interval_variant type_genetic pattern_hg19, chromosome number: position interval_variant type_genetic pattern_hg38]. In some embodiments, the input text is, for example, "Molecular cytogenetic analysis showed apaternally derived 5q35.2 --> q35.3 direct duplication". The extracted CNV entity is, for example, "5q35.2 --> q35.3". The genetic pattern entity is, for example, "paternally". The normalized result representation processed by the genetic pattern entity normalization is, for example: "['chr5:172800000-180915260_dup_Paternal Inheritance_hg19', 'chr5:173300000-181538259_dup_Paternal Inheritance_hg38']".

[0152] In the above solution, by extracting and standardizing the genetic pattern, the present invention can further enrich the annotation of CNV and improve the accuracy of its functional classification. And the standardized genetic pattern information provides reliable data support for subsequent analyses such as disease association studies and genetic risk assessments. In addition, by combining the distance principle and semantic analysis, the matching efficiency is improved while ensuring the biological relevance of the result; and by means of a consistent format and predefined classification, the complexity of multi-source data integration is simplified, which is applicable to the diverse requirements across studies.

[0153] Figure 10A block diagram schematically shows an electronic device 1000 suitable for implementing the embodiments of the present invention. The electronic device 1000 can be used to implement the execution of Figure 2 , Figures 4 to 9 the methods 200, 400 to 900 shown. As Figure 10 shown, the electronic device 1000 includes a central processing unit (i.e., CPU 1001), which can execute various appropriate actions and processes according to computer program instructions stored in a read-only memory (i.e., ROM 1002) or computer program instructions loaded from a storage unit 1008 into a random access memory (i.e., RAM 1003). In the RAM 1003, various programs and data required for the operation of the electronic device 1000 can also be stored. The CPU 1001, ROM 1002, and RAM 1003 are connected to each other via a bus 1004. An input / output interface (i.e., I / O interface 1005) is also connected to the bus 1004.

[0154] Multiple components in the electronic device 1000 are connected to the I / O interface 1005, including: an input unit 1006, an output unit 1007, a storage unit 1008, and the CPU 1001 executes the various methods and processes described above, such as executing the methods 200, 400 to 900. For example, in some embodiments, the methods 200, 400 to 900 can be implemented as computer software programs, which are stored in a machine-readable medium, such as the storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 1000 via the ROM 1002 and / or the communication unit 1009. When the computer program is loaded into the RAM 1003 and executed by the CPU 1001, one or more operations of the methods 200, 400 to 900 described above can be executed. Alternatively, in other embodiments, the CPU 1001 can be configured to execute one or more actions of the methods 200, 400 to 900 in any other suitable manner (e.g., by means of firmware).

[0155] It should be further noted that the present invention can be a method, apparatus, system, and / or computer program product. The computer program product can include a computer-readable storage medium having computer-readable program instructions thereon for performing various aspects of the present invention.

[0156] A computer-readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. A computer-readable storage medium can be, for example, but is not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punched card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium as used herein is not construed to be a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.

[0157] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to respective computing / processing devices, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include a copper transmission cable, an optical fiber transmission, a wireless transmission, a router, a firewall, a switch, a gateway computer, and / or an edge server. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.

[0158] The computer program instructions for performing the operations of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, firmware instructions, state - setting data, or source code or object code written in any combination of one or more programming languages, including object - oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer - readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand - alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer - readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field - programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer - readable program instructions to implement various aspects of the present invention.

[0159] Aspects of the present invention are described herein with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer - readable program instructions.

[0160] These computer - readable program instructions can be provided to a processing unit of a processor in a voice interaction device, a general - purpose computer, a special - purpose computer, or other programmable data - processing devices, thereby producing a machine such that when these instructions are executed by the processing unit of the computer or other programmable data - processing devices, a device is produced that implements the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams. These computer - readable program instructions can also be stored in a computer - readable storage medium, and these instructions cause the computer, programmable data - processing devices, and / or other devices to work in a specific manner. Thus, the computer - readable medium storing the instructions includes a manufacture that includes instructions for implementing various aspects of the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams.

[0161] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process such that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.

[0162] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of devices, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of instructions, which comprises one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the boxes may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending upon the functionality involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by special purpose hardware-based systems that perform the specified functions or acts, or combinations of special purpose hardware and computer instructions.

[0163] The embodiments of the present invention have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or improvements made to the technology in the market, or to enable other ordinary skilled in the art to understand the embodiments disclosed herein.

[0164] The above are only optional embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for mining copy number variation data for input text, characterized in that, Including: Concatenate the obtained context information and query information regarding copy number variations to generate an input text sequence, where the input text sequence at least includes: the context vocabulary sequence in the context information and the query vocabulary sequence in the query information; Convert the context vocabulary sequence and the query vocabulary sequence into word index sequences, so as to generate an input vector sequence based at least on the word index sequences, where the input vector sequence includes word embedding vectors, position embedding vectors, and segment embedding vectors; Input the input vector sequence into a pre-trained language model to generate output vectors through stacked multi-head attention and feed-forward networks, where the output vectors include the vector context representations of each vocabulary; and Based on the context representations of each vocabulary, predict the probability of each vocabulary as the start or end of an entity, so as to determine copy number variation entities for generating a result representation regarding copy number variations.

2. The method according to claim 1, wherein Determining copy number variation entities includes: Calculate the first probability of each vocabulary as the start position of a copy number variation entity, so as to select the vocabulary index corresponding to the maximum value of the first probability as the target start position of the copy number variation entity; Calculate the second probability of each vocabulary as the end position of a copy number variation entity, so as to select the vocabulary index corresponding to the maximum value of the second probability as the target end position of the copy number variation entity; and Based on the determined target start position and target end position of the copy number variation entity, determine the entity range and entity content of the copy number variation entity.

3. The method according to claim 1, characterized in that, Generating a result representation regarding copy number variations includes: Extract variation type entities and inheritance pattern entities in the input text based on regular matching; Record the position index of each entity among the variation type entities, inheritance pattern entities, and copy number variation entities in the input text; Calculate the text distance between entities; and Determine the matching relationship between entities based at least on the calculated text distance between entities.

4. The method according to claim 3, wherein Determining the matching relationship between entities based at least on the calculated text distance between entities includes: Based on the text distance between entities, determine the candidate matching relationship between entities; Based on the dependency path weight and text distance weight, calculate the matching degree characterization data for each candidate matching relationship between entities; and Select the candidate matching relationship corresponding to the maximum value of the matching degree characterization data as the target matching relationship between entities.

5. The method according to claim 4, wherein Determining the matching relationship between entities includes: In response to determining that multiple variation type entities or multiple inheritance pattern entities have candidate matching relationships with the same copy number variation entity, preferentially select the candidate matching relationship with a direct dependency modification relationship as the target matching relationship; In response to determining that there is no candidate matching relationship with a direct dependency modification relationship, calculate the matching degree characterization correction data based on the dependency relationship characterization data and context matching characterization data, so as to select the candidate matching relationship corresponding to the highest value of the matching degree characterization correction data as the target matching relationship; and In response to determining that multiple matching degree characterization data are the same, select the candidate matching relationship with a higher text order as the target matching relationship.

6. The method according to claim 1, wherein The result representation regarding copy number variation includes: chromosome number, position interval, variation type, inheritance pattern, and genome version number of the copy number variation entity. Generating the result representation regarding copy number variation includes multiple ones of the following: Normalizing with respect to the genome version number; Normalizing with respect to the variation type entity; Normalizing with respect to the inheritance pattern entity; and Normalizing with respect to the position interval unit.

7. The method according to claim 6, characterized in that, Normalizing with respect to the genome version number includes: Mapping the copy number variation entity to the first human reference genome and the second human reference genome respectively; Obtaining the information of the corresponding reference genome version extracted from the input text to match the entity coordinates of the copy number variation entity corresponding to the reference genome version; and Based on the entity coordinates of the copy number variation entity, generating a normalized result representation regarding copy number variation, the normalized result representation regarding copy number variation includes: chromosome number of the copy number variation, position interval relative to the first human reference genome, and position interval relative to the second human reference genome.

8. The method according to claim 6, characterized in that Normalizing with respect to the variation type entity includes: Based on the entity coordinates of the copy number variation entity and the matching relationship between the variation type entity and the copy number variation entity, generating a normalized result representation regarding copy number variation; The matching relationship between the variation type entity and the copy number variation entity is generated through the following: Using semantic analysis or natural language processing method based on rule matching to extract the variation type entity from the input text; Calculating the text distance between the variation type entity and the copy number variation entity in the input text; Determining the matching relationship between the variation type entity with the closest text distance and the copy number variation entity; and In response to determining that there are multiple matching relationships between the variation type entity and the copy number variation entity, selecting from multiple matching relationships in combination with context information.

9. The method according to claim 6, wherein Normalizing with respect to the position interval unit includes: Using regular expressions to extract the position interval description containing the unit from the input text; Calculating the standardized position interval length for the extracted position interval description; and Combining the standardized position interval length, variation type, and reference genome version to generate a normalized result representation regarding copy number variation.

10. The method according to claim 9, wherein Normalizing with respect to the inheritance pattern entity includes: Extracting the inheritance pattern entity from the input text; Classifying the extracted inheritance pattern entity into a predetermined inheritance pattern classification, the predetermined inheritance pattern classification being one of de novo, paternal inheritance, and maternal inheritance; Calculating the text distance between the inheritance pattern entity and the copy number variation entity to match the inheritance pattern entity corresponding to the minimum text distance with the copy number variation entity; and Combining the standardized position interval length, variation type, reference genome version, and predetermined inheritance pattern classification to generate a normalized result representation regarding copy number variation.

11. A computing device, characterized in that, Includes: At least one processing unit; At least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions when executed by the at least one processing unit cause the device to perform the steps of the method according to any one of claims 1 to 10.

12. A computer-readable storage medium, characterized in that, A computer program is stored on a computer-readable storage medium, and when the computer program is executed by a machine, it implements the method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Methods for estimating genome-wide copy number variations

    CN103201744A

  • Automatic analysis method and system for genome copy number variation

    CN118588157A