Structural variation filtering method, device and equipment based on multi-modal fusion

By adopting a multimodal fusion method in structural variation detection, combining feature pictures and ESF values, and using CLIP multimodal model for filtering, the existing detection tools are solved in terms of sensitivity and accuracy, and high-precision structural variation recognition and filtering are achieved.

CN120219902AActive Publication Date: 2025-06-27NAT UNIV OF DEFENSE TECH
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510689615.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-06-27
Estimated Expiration
2045-05-27

AI Technical Summary

Technical Problem

Existing structural variation detection tools have shortcomings in detection sensitivity and accuracy, especially in long-read sequencing data. It is difficult for statistical information-based methods to fully capture complex patterns, while deep learning-based methods have the problem of incomplete acquisition of feature signals.

Method used

A structural variation filtering method based on multimodal fusion is adopted to encode the variable information of structural variation sites, generate feature pictures, and calculate the ESF value to form an image text pair. Then, the CLIP multimodal model is used to train these data and build a multimodal fusion model to filter structural variation sites and improve detection accuracy.

Benefits of technology

It significantly improves the accuracy of gene sequence structural variation detection, reduces false positive results, enhances the adaptability and robustness of the model to different data, and can integrate with existing detection tools.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219902A_ABST
    Figure CN120219902A_ABST
Patent Text Reader

Abstract

The invention relates to a structural variation filtering method, device and equipment based on multi-modal fusion. The method comprises the following steps: encoding variable information of a structural variation site to generate a feature picture of the structural variation site; and after calculating the ESF value of the variable information, setting a label for an image text pair formed by the structure variation site feature picture and the ESF value according to the structure variation type to obtain an image text pair with a label. And training the CLIP multi-modal model according to preset configuration parameters to obtain a trained multi-modal fusion model. And filtering variation sites of the preprocessed structure variation data and the labeled image text pair through the trained multi-modal fusion model to obtain filtering result data. By adopting the method, the detection precision of the structural variation in the gene sequence can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of gene sequence structural variation detection, and particularly relates to a structural variation filtering method, device and equipment based on multimodal fusion. Background Art

[0002] Structural variations generally refer to large-scale sequence variations in the genome, usually involving DNA fragments with a sequence length greater than 50 bases, including: deletions (DEL), insertions (INS), inversions (INV), duplications (DUP), and translocations (TRAN), etc. Compared with short insertions and deletions (INDELs) and single nucleotide variations (SNV), structural variations have a higher probability of causing diseases such as cancer, genetic diseases, and neurodevelopmental disorders. For example, the translocation between chromosome 9 and chromosome 22 will form the so-called "Philadelphia chromosome", causing chronic myeloid leukemia. The deletion on chromosome 7 will trigger Werdnig-Hoffmann disease. The trisomy of chromosome 21 will lead to Down syndrome.

[0003] Existing structural variation detection tools can be classified into statistical data-based methods and deep learning-based methods according to their implementation principles. The former infers structural variations by judging which features in the data are abnormal, and the latter relies on the feature extraction and pattern recognition capabilities of deep neural networks to discover structural variations.

[0004] Early structural variation detection tools were mainly based on short reads, including DELY, LUMPY, Manta, and SvABA, etc. These tools usually identify and detect structural variations based on statistical information in the alignment results, such as read depth (RD), discordant read pairs (RP), split read alignments (SR), local assembly, or a combination thereof. DELY comprehensively uses two signals of discordant read pairs and split read alignments to achieve the detection of structural variations corresponding to short reads. LUMPY comprehensively detects structural variations of short reads by combining multiple signals such as discordant read pairs, split read alignments, and read depth. Manta locally assembles suspected structural variation regions by combining discordant read pairs and split read alignments to verify structural variation events of short reads. SvABA comprehensively uses discordant read pair and split read alignment signals, and at the same time combines local assembly methods to enhance the detection ability of complex structural variations of short reads. However, the short read lengths of short reads limit the sensitivity of these tools during detection, resulting in a large number of false positives in the results.

[0005] The rapid development of third-generation sequencing, also known as long-read sequencing technologies such as the Pacific Biosciences (PacBio) and Oxford Nanopore Technology (ONT) platforms, can provide alignment information over a longer span, which offers opportunities for more comprehensive and accurate detection of structural variations. However, the higher error rate and longer read lengths of long-read sequencing technologies also render short-read-based structural variation detection algorithms inapplicable, prompting researchers to develop new structural variation detection tools.

[0006] Many existing long-read structural variation detection tools are mostly based on statistical methods, such as PBSV, Sniffles2, SVIM, and cuteSV. These structural variation detection tools follow the idea of short-read sequence-based structural variation detection tools that rely on statistical information in the alignment results. By integrating various statistical information and leveraging the advantages of long-read sequencing, high-precision structural variation detection is achieved. PBSV is designed specifically for PacBio long-read data and uses statistical signals in the alignment results (such as CIGAR strings, paired-end information, etc.) to detect structural variations. Sniffles2 generates candidate regions for potential structural variation events based on discordant read pairs and split-read alignment signals in long-read sequences, and verifies the authenticity of candidate structural variation events through further analysis and statistical evaluation. SVIM analyzes discordant read pairs and split-read alignments in long-read sequences to identify potential structural variation breakpoints, uses a probability model to evaluate the likelihood of different structural variation types, and distinguishes true structural variations from noise signals. cuteSV is compatible with both short-read and long-read sequencing data, comprehensively analyzes signals such as discordant read pairs, split-read alignments, and read depth in the alignment results, and identifies different types of structural variations.

[0007] With the rapid development of deep learning technology, a batch of deep learning-based structural variant detection tools have emerged, including DeepSVFilter, SVision, SVcnn, and cnnLSV, etc. The core of such methods is to convert various signals in the alignment results into two-dimensional images, and then utilize the powerful feature extraction and pattern recognition capabilities of deep neural networks to achieve high-precision detection of structural variants. DeepSVFilter encodes the read depth, inconsistent read pairs, and split read alignments in the alignment data as RGB images, and uses a CNN network for classification, which can achieve the filtering of DEL and INV types in the short-read structural variant results. SVision generates pictures of VAR-REF and REF-REF based on the "variant feature sequence" (VAR) and reference sequence (REF) in the long-read sequences, and uses CNN to learn and extract the variant features in the pictures to achieve the detection of complex structural variants. SVcnn designs a candidate variant region algorithm, generates feature pictures according to the CIGAR string and split read alignment information in the candidate region, and uses a CNN network to classify the types of long-read sequence structural variants. cnnLSV fuses the detection results of multiple long-read detection tools, proposes a picture encoding strategy based on the CIGAR string and split read alignment information, and uses a CNN network to remove false positives in the results and improve the detection performance.

[0008] The methods based on statistical information rely on predefined features, which may lead to the inability to comprehensively capture the complex patterns and potential non-linear relationships existing in the data, limiting the detection ability. And this method is usually based on a fixed detection pattern, with poor generalization ability. While the methods based on deep learning are usually trained based on general CNN models, such as MobileNet, ResNet models, etc., there is a problem of incomplete acquisition of feature signal information. In addition, the recognition results of the model may be affected by the dataset or the alignment tool, and the performance of the model may decline for cross-platform data. Summary of the Invention

[0009] Based on this, it is necessary to provide a structural variant filtering method, device, and equipment based on multi-modal fusion that can improve the detection accuracy of gene sequence structural variants for the above technical problems.

[0010] A structural variant filtering method based on multi-modal fusion, the method includes: Encode the variable information of the structural variant sites to generate structural variant site feature pictures.

[0011] After calculating the ESF value of the variable information, set labels for the image-text pairs composed of the structural variant site feature images and the ESF values according to the structural variant type to obtain labeled image-text pairs.

[0012] Train the CLIP multi-modal model according to the preset configuration parameters to obtain a trained multi-modal fusion model, and filter the variant sites of the preprocessed structural variant data and the labeled image-text pairs through the trained multi-modal fusion model to obtain filtered result data.

[0013] A structural variant filtering device based on multi-modal fusion, the device includes: A feature image generation module, configured to encode the variable information of the structural variant site to generate a structural variant site feature image.

[0014] A multi-modal data annotation module, configured to set labels for the image-text pairs composed of the structural variant site feature images and the ESF values according to the structural variant type after calculating the ESF value of the variable information to obtain labeled image-text pairs.

[0015] A filtering module, configured to train the CLIP multi-modal model according to the preset configuration parameters to obtain a trained multi-modal fusion model, and filter the variant sites of the preprocessed structural variant data and the labeled image-text pairs through the trained multi-modal fusion model to obtain filtered result data.

[0016] A computer device, including a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented: Encode the variable information of the structural variant site to generate a structural variant site feature image.

[0017] After calculating the ESF value of the variable information, set labels for the image-text pairs composed of the structural variant site feature images and the ESF values according to the structural variant type to obtain labeled image-text pairs.

[0018] Train the CLIP multi-modal model according to the preset configuration parameters to obtain a trained multi-modal fusion model, and filter the variant sites of the preprocessed structural variant data and the labeled image-text pairs through the trained multi-modal fusion model to obtain filtered result data.

[0019] The above-mentioned structural variant filtering method, device, and equipment based on multimodal fusion generate feature pictures by encoding variable information of structural variant sites, convert them into two-dimensional feature pictures, break through the constraints of traditional fixed feature patterns, and present complex relationships and potential patterns in data in a visual and structured manner. Compared with simply relying on predefined features, it can capture data features more comprehensively. At the same time, calculate the ESF value of the variable information to form one-dimensional data, and form an image-text pair with the two-dimensional feature picture. Through this multimodal data combination, the structural variant sites are characterized from different dimensions, making up for the deficiency of the general CNN model based on deep learning in obtaining feature signals. In addition, set labels for the image-text pairs according to the structural variant types, combine the preprocessed structural variant data, and use the trained CLIP multimodal model to construct a multimodal fusion model. Then, use the trained multimodal fusion model to filter variant sites, significantly enhancing the adaptability and robustness of the model to different data, achieving high-precision structural variant recognition of gene sequences, reducing false positive results, improving detection accuracy, and being able to be integrated with existing detection tools. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 It is a schematic flowchart of a structural variant filtering method based on multimodal fusion in an embodiment; Figure 2 It is a schematic flowchart of realizing the structural variant recognition and filtering of gene sequences by fusing the structural variant feature picture information and the ESF value based on multimodal technology in an embodiment; Figure 3 It is a structural block diagram of a structural variant filtering device based on multimodal fusion in an embodiment; Figure 4 It is an internal structure diagram of a computer device in an embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0021] In order to make the objectives, technical solutions, and advantages of this application clearer, the following further elaborates on this application in combination with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.

[0022] In one embodiment, as Figure 1 shown, a structural variant filtering method based on multimodal fusion is provided, including the following steps: Step 102: Encode the variable information of the structural variant sites to generate a structural variant site feature picture.

[0023] Step 104: After calculating the ESF value of the variable information, set labels for the image-text pair composed of the structural variant site feature picture and the ESF value according to the structural variant type to obtain a labeled image-text pair.

[0024] Step 106: Train the CLIP multi-modal model according to preset configuration parameters to obtain a trained multi-modal fusion model. Filter the variant sites of the preprocessed structural variant data and the labeled image-text pairs through the trained multi-modal fusion model to obtain filtered result data.

[0025] In the above structural variant filtering method based on multi-modal fusion, by encoding the variable information of the structural variant sites to generate feature pictures and converting them into two-dimensional feature pictures, it breaks through the bondage of the traditional fixed feature mode and presents the complex relationships and potential patterns in the data in a visual and structured way. Compared with simply relying on predefined features, it can capture data features more comprehensively. At the same time, calculate the ESF value of the variable information to form one-dimensional data, and combine it with the two-dimensional feature pictures to form image-text pairs. Through this multi-modal data combination, the structural variant sites are characterized from different dimensions, making up for the deficiency of the general CNN model based on deep learning in obtaining feature signals. In addition, set labels for the image-text pairs according to the structural variant types, combine the preprocessed structural variant data, use the trained CLIP multi-modal model to construct a multi-modal fusion model, and then use the trained multi-modal fusion model to filter the variant sites, significantly enhancing the adaptability and robustness of the model to different data, achieving high-precision identification of gene sequence structural variants, reducing false positive results, improving detection accuracy, and being able to be integrated with existing detection tools.

[0026] In one embodiment, extract the depth information of the structural variant sites in the alignment file, and record the coverage depth of each chromosome location according to the depth information; and preprocess the read sequences of the alignment file to obtain preprocessing information. Use the coverage depth and the preprocessing information as the variable information of each structural variant site, and generate a structural variant site feature picture according to the variable information. The height of each pixel point in the structural variant site feature picture is equal to the value of the variable in the variable information.

[0027] In one embodiment, calculate the ESF value of the preprocessing information in the variable information, and check whether the total number of the structural variant site feature pictures and the ESF values is consistent with the total number of the structural variant sites. If they are consistent, set labels for the image-text pairs composed of the structural variant site feature pictures according to the structural variant types to obtain labeled image-text pairs. Otherwise, re-obtain the variable information to generate structural variant site feature pictures.

[0028] In one embodiment, train the CLIP multi-modal model according to preset configuration parameters, standard picture data, and standard text data to obtain a trained multi-modal fusion model.

[0029] In one embodiment, by encoding the variable information of the structural variation sites in the structural variation data, a to-be-processed feature image is generated, and the ESF value is calculated to obtain preprocessed structural variation data.

[0030] In one embodiment, the preprocessed structural variation data and the feature images and ESF values respectively corresponding to the labeled image-text pairs are classified at each to-be-processed variation site by a trained multi-modal fusion model to obtain a model classification result, and the false positive results of the structural variation data are filtered according to the model classification result to obtain filtered result data.

[0031] In one embodiment, the entries consistent with the model classification result are retained in the structural variation data, and the entries inconsistent with the model classification result are deleted or modified to obtain filtered result data.

[0032] In one embodiment, as Figure 2 shown, a method for identifying and filtering structural variations of gene sequences by fusing structural variation feature image information and ESF values based on multi-modal technology is provided. Based on multi-modal technology, the structural variation feature image information and ESF (Embedding Sequence Feature) text information are fused to realize the identification and filtering of structural variations of gene sequences.

[0033] Specifically, based on the RD, RP, and SR statistical information in the alignment (BAM / SAM, Binary / Sequence Alignment / Map format) file, as well as the M (Match, insertion), I (Insert, insertion), D (Delete, deletion), and S (Soft clipping) string information of the CIGAR signal in the alignment file, a feature image is generated, and the ESF features of the CIGAR signal (the length, minimum value, first quartile, median, third quartile, maximum value, root mean square value, harmonic mean, mean, standard deviation, and coefficient of variation of each CIGAR signal) are fused. The multi-modal model CLIP is used for training to realize the identification of structural variations of gene sequences, and the false positive results in the structural variation result file (VCF format, Variant Call Format) are filtered. The specific content is as follows: Step 1: Obtain the original data file and preprocess the data file; Step 1.1: VCF file preprocessing: Preprocess the structural variation data; Step 1.2: Alignment file preprocessing: Preprocess the sequence alignment file; Step 2: Preprocess the data file based on the structural variation sites; Step 2.1: Obtaining Mutation Site Information: Traverse the structural variation result file to obtain chromosomal mutation site information; Step 2.2: Preprocessing Alignment Information: Preprocess the alignment information according to the mutation sites; Step 2.2.1: Apply for variable spaces all_img, all_img_mids, and all_list to store statistical information, picture feature information, CIGAR picture feature information, and ESF variable information; Step 2.2.2: Apply for memory spaces split_read_left, split_read_right, and rd_count for three types of statistical information: left split reads, right split reads, and read depth; Step 2.2.3: Apply for memory spaces conjugate_m, conjugate_i, conjugate_d, and conjugate_s for four CIGAR signals: Match, Insert, Delete, and Soft Clipping; Step 2.2.4: Load the chromosomal depth information file, record the read sequence coverage depth at each structural variation position, and store the result in the variable rd_count; Step 2.3: Reads Screening; Step 2.3.1: Traverse the read sequences in the alignment result file, judge and count the left (soft clipped) and right (soft clipped) reads, and store the results in the variables split_read_left and split_read_right; Step 2.3.2: Traverse the CIGAR signals in the read alignment result, count the occurrences of the Match, Insert, Delete, and SoftClipping signals, and store them in the variables conjugate_m, conjugate_i, conjugate_d, and conjugate_s; Step 3: Perform Structural Variation Site Feature Picture Encoding and Calculate ESF Values Based on the Preprocessed Information; Step 3.1: Initialize the Picture: Initialize the preprocessed information; Step 3.1.1: Apply for the variable args_list to store the generated picture index, start and end positions; Step 3.1.2: Prepare variables for different structural variation types; Step 3.1.3: Count the quantities corresponding to different structural variation types; Step 3.2: Generate Feature Pictures: Generate feature pictures according to different mutation types and quantities; Step 3.2.1: Load variables split_read_left, split_read_right, rd_count, conjugate_m, conjugate_i, conjugate_d, and conjugate_s; Step 3.2.2: Calculate the minimum value of the dimensions of each image; Step 3.2.3: Define the image height as the maximum value of each dimension minus the minimum value plus 2; Step 3.2.4: Fill the initial image (with a length equal to the length of each dimension and a height as defined above) with 0; Step 3.2.5: Generate an initial image according to the variable information. The length of the image is the length of the mutation region, and the height of each pixel is equal to the value of the variable; Step 3.2.5.1: Obtain the length of each dimension in the variable information and use this length as the abscissa of the image pixel; Step 3.2.5.2: Calculate the difference from the minimum value according to the value of each locus in the variable information, and this difference is used as the ordinate corresponding to each position in the picture; Step 3.2.5.3: Traverse each locus, calculate the values of all loci, and generate an initial image; Step 3.2.6: Transform the image and resize it to 244×244; Step 3.2.7: Save the generated statistical information picture into the all_img.pt file, and save the generated CIGAR signal picture into the all_img_mids file; Step 3.3: Calculate the ESF value: A specific embodiment is to calculate the ESF value of the CIGAR signal; Step 3.3.1: Apply memory space conjugate_esf_list for the ESF variable; Step 3.3.2: Load variables conjugate_m, conjugate_i, conjugate_d, and conjugate_s; Step 3.3.3: Calculate the lengths of the four CIGAR signals according to each variable, and calculate their minimum value, first quartile, median, third quartile, maximum value, root mean square value, harmonic mean, mean, standard deviation, and coefficient of variation, and save the results into the variable conjugate_esf_list; Step 3.3.4: Generation completed: Save the results of the variable conjugate_esf_list into the all_list.pt file; Step 3.4: Generation completed: Check whether the number of feature pictures and ESF values is consistent with the number of variant sites. If they are not consistent, return to Step 3.2; if they are consistent, proceed to Step 4; Step 4: Set labels for the corresponding feature pictures according to the structural variant type; Step 4.1: Create a new picture storage folder: Create a folder for each structural variant type; Step 4.2: Set the generated picture / ESF training labels: Traverse each picture and assign a structural variant type label to it according to the index and variant type in the args_list for later model training; Step 4.3: Store the relevant files in the folder: Store the labeled pictures in the corresponding folders according to the variant type; Step 5: Train using the CLIP multi-modal model to obtain a trained model; Step 5.1: Configure the model parameters: Configure the model training parameters; Step 5.2: Prepare the dataset: Prepare the dataset and randomly select; Step 5.2.1: Set the random number generator seed; Step 5.2.2: Randomly shuffle the elements of the training list according to the random number generator seed order; Step 5.2.3: Randomly select 80% of the data as training data and 20% of the data as validation and test data; Step 5.3: Load the CLIP model for training; Step 5.4: Save the trained model; Step 6: Process the VCF to be filtered to generate feature pictures / ESF: Obtain the structural variant results to be processed, generate feature pictures based on the structural variant sites, and calculate the ESF values; Step 6.1: Process the VCF file to be processed in the same way as in Step 1.2; Step 6.2: Process the (BAM / SAM) file in the same way as in Steps 1.2 - 1.3 to extract depth information; Step 6.3: Preprocess the data file in the same way as in Step 2; Step 6.4: Generate feature pictures and calculate ESF values in the same way as in Step 3; Step 7: Classify the model results and write to the filtered results: Use the trained model for filtering to obtain the filtered VCF file; Step 7.1: Load the trained model to classify the pictures and ESF of each variant site to be processed; Step 7.1.1: Load the trained model and set the model parameters; Step 7.1.2: Classify the pictures and ESFs of the variant sites to be processed using the model; Step 7.1.3: Retain the entries that are consistent with the model classification results according to the model classification results, and delete / modify the inconsistent entries; Step 7.2: Write the filtered results into a VCF file.

[0034] It should be noted that based on multimodal technology, the picture information of structural variation features and ESF text information are fused, and the CLIP model is used for training, which is applicable to second-generation and third-generation sequencing, realizes high-precision identification of structural variations in gene sequences, and can be integrated with existing structural variation detection tools, significantly reducing false positive results.

[0035] It should be understood that although Figure 1 - Figure 2 the steps in the flowchart of Figure 1 - Figure 2 are shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover,

[0036] In one embodiment, as Figure 3 shown, a structural variation filtering device based on multimodal fusion is provided, including: a feature picture generation module 302, a multimodal data annotation module 304, and a filtering module 306, where: The feature picture generation module 302 is used to encode the variable information of the structural variation sites to generate structural variation site feature pictures.

[0037] The multimodal data annotation module 304 is used to calculate the ESF value of the variable information, and then set labels for the image-text pairs composed of the structural variation site feature pictures and the ESF values according to the structural variation types to obtain labeled image-text pairs.

[0038] The filtering module 306 is used to train the CLIP multimodal model according to the preset configuration parameters to obtain a trained multimodal fusion model, and filter the variant sites of the preprocessed structural variation data and the labeled image-text pairs through the trained multimodal fusion model to obtain filtered result data.

[0039] For the specific limitations of the structural variation filtering device based on multimodal fusion, reference may be made to the limitations of the structural variation filtering method based on multimodal fusion in the foregoing text, which will not be elaborated herein. Each module in the above-mentioned structural variation filtering device based on multimodal fusion can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in the processor of the computer device in hardware form or be independent of it, or be stored in the memory of the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to each of the above modules.

[0040] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as Figure 4 shown. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a structural variation filtering method based on multimodal fusion. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device may be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.

[0041] Those skilled in the art can understand that Figure 3 - Figure 4 the structure shown in

[0042] is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements. Encode the variable information of the structural variation sites to generate structural variation site feature pictures.

[0043] After calculating the ESF value of the variable information, set labels for the image-text pair composed of the structural variation site feature picture and the ESF value according to the structural variation type to obtain a labeled image-text pair.

[0044] Train the CLIP multimodal model according to the preset configuration parameters to obtain a trained multimodal fusion model. Filter the mutation sites of the preprocessed structural variation data and the labeled image-text pairs through the trained multimodal fusion model to obtain filtered result data.

[0045] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0046] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0047] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. A structural variant filtering method based on multimodal fusion, characterized in that, The method includes: Encoding the variable information of the structural variation sites to generate structural variation site feature pictures; After calculating the ESF value of the variable information, setting labels for the image-text pairs composed of the structural variation site feature pictures and the ESF value according to the structural variation type to obtain labeled image-text pairs; Training the CLIP multi-modal model according to preset configuration parameters to obtain a trained multi-modal fusion model, and filtering the variant sites of the preprocessed structural variation data and the labeled image-text pairs through the trained multi-modal fusion model to obtain filtered result data.

2. The method according to claim 1, wherein Encoding the variable information of the structural variation sites to generate structural variation site feature pictures, including: Extracting the depth information of the structural variation sites in the alignment file, recording the coverage depth at each chromosomal position according to the depth information; and preprocessing the read sequences of the alignment file to obtain preprocessing information; Taking the coverage depth and the preprocessing information as the variable information of each structural variation site, and generating structural variation site feature pictures according to the variable information; the height of each pixel point in the structural variation site feature pictures is equal to the value of the variable in the variable information.

3. The method according to claim 2, characterized in that, After calculating the ESF value of the variable information, setting labels for the image-text pairs composed of the structural variation site feature pictures and the ESF value according to the structural variation type to obtain labeled image-text pairs, including: Calculating the ESF value of the preprocessing information in the variable information, checking whether the total number of the structural variation site feature pictures and the ESF value is consistent with the total number of the structural variation sites. If they are consistent, setting labels for the image-text pairs composed of the structural variation site feature pictures according to the structural variation type to obtain labeled image-text pairs; otherwise, re-obtaining the variable information to generate structural variation site feature pictures.

4. The method according to claim 3, wherein Training the CLIP multi-modal model according to preset configuration parameters to obtain a trained multi-modal fusion model, including: Training the CLIP multi-modal model according to preset configuration parameters, standard picture data, and standard text data to obtain a trained multi-modal fusion model.

5. The method according to claim 4, characterized in that, By encoding the variable information of the structural variation sites of the structural variation data, generating to-be-processed feature pictures, and calculating the ESF value, obtaining the preprocessed structural variation data.

6. The method according to claim 5, characterized in that, Filtering the variant sites of the preprocessed structural variation data and the labeled image-text pairs through the trained multi-modal fusion model to obtain filtered result data, including: Classifying the feature pictures and ESF values corresponding to the preprocessed structural variation data and the labeled image-text pairs respectively at each to-be-processed variant site through the trained multi-modal fusion model to obtain a model classification result, and filtering the false positive results of the structural variation data according to the model classification result to obtain filtered result data.

7. The method according to claim 6, wherein Filtering the false positive results of the structural variation data according to the model classification result to obtain filtered result data, including: Retain the entries in the structural variation data that are consistent with the model classification results, and delete or modify the entries that are inconsistent with the model classification results to obtain filtered result data.

8. A structural variation filtering device based on multimodal fusion, characterized in that, The device includes: A feature image generation module, configured to encode the variable information of the structural variation site to generate a feature image of the structural variation site; A multi-modal data annotation module, configured to calculate the ESF value of the variable information, and then set labels for the image-text pair composed of the feature image of the structural variation site and the ESF value according to the structural variation type to obtain a labeled image-text pair; A filtering module, configured to train the CLIP multi-modal model according to preset configuration parameters to obtain a trained multi-modal fusion model, and filter the variation sites of the preprocessed structural variation data and the labeled image-text pair through the trained multi-modal fusion model to obtain filtered result data.

9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Genome structure variation detection method based on filtering and noise reduction

    CN113436678A

  • Bioinformatics system, device and method for performing secondary and / or tertiary processing

    CN118016151A

  • IGV genetic structure variation image recognition method and equipment based on multi-task learning, and medium

    CN118486371A

  • Genome structure variation detection method based on long read sequencing data and false positive filtering model

    CN118824360A

  • False positive structural variation filtering method, storage medium, and computing device

    WO2022011855A1