Respiratory syncytial virus typing information processing method, device and equipment

By constructing the reference genome and stratified one hot encoding, combining CNN and Transformer technology, the problems of unstable virus classification results and poor robustness to indel variants in the prior art are solved, and higher genotype classification accuracy and robustness are achieved.

CN119942202AActive Publication Date: 2025-05-06四川脉得影深信息技术有限公司
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510027001.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-05-06
Estimated Expiration
2045-01-08

AI Technical Summary

Technical Problem

Among the existing virus classification methods, the classification method based on evolutionary tree has the problem that different reference evolution trees lead to different classification results, and the zero padding strategy has clustering errors and poor robustness to indel variants when genotypes of the same species.

Method used

An information processing method for respiratory syncytial virus typing is proposed. By constructing reference genome and stratified one hot encoding, combining CNN networks to extract features and fusion, and genotyping is performed using the Transformer architecture to improve classification accuracy and robustness.

Benefits of technology

It improves the accuracy of respiratory syncytial virus genotype classification, reduces clustering error rate, enhances the robustness of insertion variants, and is better than the evolutionary tree-based classification method.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942202A_ABST
    Figure CN119942202A_ABST
Patent Text Reader

Abstract

The invention discloses an information processing method, device and equipment for respiratory syncytial virus typing, and relates to the technical field of virus classification. According to the information processing method for the respiratory syncytial virus typing, the hierarchical one hot coding strategy is innovatively provided on the basis of the original one hot coding and the zeropadding strategy, so that the interference of the traditional zeropadding mode on one hot coding is reduced, and the robustness of insertion variation is enhanced. Based on a hierarchical one hot coding strategy and a neural network image classification technology, doctors and scientific researchers can be helped to make accurate diagnosis and treatment guidance, the working efficiency is improved, and the efficiency and diagnosis and treatment quality of the medical and public safety industries are helped to be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of virus classification, and in particular to an information processing method, device and equipment for respiratory syncytial virus typing. Background Art

[0002] Respiratory syncytial virus genotyping is of great significance. It helps to gain a deeper understanding of the epidemiological characteristics of the virus. By accurately determining different genotypes, it can clearly track the transmission path and epidemic trends of the virus in different regions and different populations, providing a key basis for the formulation of public health prevention and control strategies. In clinical practice, genotyping can assist doctors in more accurately determining the severity of the disease, because different genotypes of respiratory syncytial virus may differ in pathogenicity, clinical manifestations after infection, and the rate of disease progression, so that personalized treatment plans can be formulated to improve treatment efficacy and improve patient prognosis.

[0003] In the Nextrain typing technology, an evolutionary tree is constructed for the G gene sequence or the whole genome sequence. In the process of constructing the evolutionary tree, different strains of respiratory syncytial virus are carefully classified and typed based on various characteristics such as sequence similarities and differences. This typing method has greatly improved our understanding of the diversity and evolutionary relationships of respiratory syncytial virus, and helps to clearly track the source, transmission path and mutation trend of the virus in epidemiological studies. In terms of clinical diagnosis, it can provide doctors with more accurate virus typing information, assist in determining the specific type of infection, and thus develop more targeted treatment strategies. However, in the typing method based on the evolutionary tree, there are problems such as different reference evolutionary trees leading to different classification results.

[0004] In the research and practice of virus classification, deep learning technology has shown great potential and application value. Deepvirusclassifier software focuses on the whole genome sequence and uses one hot encoding technology to effectively extract features from the whole genome sequence data of the virus. This encoding method is better than directly converting A, T, C, G to 1, 2, 3, 4, because the difference between A and G is not the difference between T and G, which is 2 and 3. Based on the deep learning network framework, the software accurately identifies the characteristic patterns unique to the whole genome sequence of the new coronavirus, which is in sharp contrast to the sequence characteristics of other viruses, so as to efficiently and accurately determine whether an unknown virus sample belongs to the new coronavirus. However, in this software one hot encoding, there is no fixed reference gene, base substitution is represented by ATCG, N represents the unknown base here, and the zero padding strategy is used at the end to represent the insertion and deletion between sequences. However, when genotyping the same species, this zero padding strategy has some shortcomings, such as clustering errors, zero padding interference with onehot encoding, and poor robustness to indel (insertion) mutations.

[0005] In view of this, the present invention is proposed. Summary of the invention

[0006] The object of the present invention is to provide an information processing method, device and equipment for respiratory syncytial virus typing to solve the above technical problems.

[0007] The present invention is achieved in that:

[0008] In a first aspect, the present invention provides a method for constructing a respiratory syncytial virus typing model, comprising the following steps:

[0009] S1: Construction of reference genome: Select the respiratory syncytial virus genome used to construct the reference genome from the respiratory syncytial virus genome sequence database, and divide the sequence into training set, validation set and test set; perform sequence alignment on the genomes in some training sets, find the similar regions and variant regions between all genomes in the training sets, and construct a pan-genome map containing all genome information based on the alignment results; merge the variant regions to obtain the reference genome;

[0010] S2: Perform hierarchical one hot encoding: The length of the reference genome is X. The whole genome sequences of the training set and the validation set are aligned with the reference genome and one hot encoded to obtain a picture of query vector*key vector*value vector (X+1)*7*1; the key vector in the picture refers to the following seven types of information: whether there are A bases, T bases, C bases, G bases, missing or unknown bases, insertion mutations, and whether the bases at each position are located in the coding gene region of the G protein after alignment with the reference genome;

[0011] The whole genome sequences of the training set and the validation set are once again ont hot encoded to obtain a query vector*key vector*value vector of 50*120~200*1 images; the key vector in the 50*120~200*1 image refers to the number of insertion mutations; if there is at least one insertion mutation, it is encoded in the front row of the 50*120~200*1 image;

[0012] S3: Fusing hierarchical features: Using the CNN network, extract features from the (X+1)*7*1 image and the 50*120~200*1 image, respectively, obtain features of the target dimension, splice and fuse the obtained images, and obtain a fused feature map;

[0013] S4: RSV genotyping based on neural network: input fusion feature map, use neural network to perform RSV genotyping and output;

[0014] S5: Input the full genome of the respiratory syncytial virus in the test set and perform steps S2-S4.

[0015] In a second aspect, the present invention provides an information processing method for respiratory syncytial virus typing, which comprises the following steps:

[0016] A respiratory syncytial virus typing model is constructed according to the above-mentioned method for constructing a respiratory syncytial virus typing model, and then the full genome of the respiratory syncytial virus of the sample to be tested is input to perform steps S2-S4.

[0017] In a third aspect, the present invention also provides a device for typing respiratory syncytial virus, comprising: an input module, a control module and an output module;

[0018] The input module was configured to: input the respiratory syncytial virus genome;

[0019] The control modules include: reference genome construction module, hierarchical one hot encoding module, fusion feature module and genotyping module;

[0020] The output module includes: outputting respiratory syncytial virus typing results;

[0021] The reference genome construction module is configured as follows: based on the input respiratory syncytial virus genome, the sequence is divided into a training set, a validation set, and a test set; the genomes in some training sets are sequenced, similar regions and variant regions between all genomes in some training sets are found, and a pan-genome map containing all genome information is constructed based on the alignment results; the variant regions are merged to output the reference genome, training set, validation set, and test set;

[0022] The hierarchical one hot encoding module is configured as follows: input the whole genome sequences of the reference genome, training set and validation set, align the whole genome sequences of the training set and validation set with the reference genome, and perform one hot encoding to obtain a picture with query vector * key vector * value vector of (X+1)*7*1, where X is the length of the reference genome; the key vector in the picture refers to the following seven types of information: whether there are A bases, T bases, C bases, G bases, missing or unknown bases, inserted mutations and whether the bases at the position are located in the coding gene region of the G protein at each position after alignment with the reference genome; output a picture with query vector * key vector * value vector of (X+1)*7*1; perform one hot encoding on the whole genome sequences of the training set and validation set again. Hot encoding, output query vector * key vector * value vector is 50*120~200*1 picture; the key vector in the 50*120~200*1 picture refers to the number of insertion mutations; if there is at least one insertion mutation, it is encoded in the front column of the 50*120~200*1 picture;

[0023] The fusion feature module is configured as follows: input (X+1)*7*1 pictures and 50*120~200*1 pictures, use CNN network to extract features, obtain the features of target dimensions respectively, splice and fuse the obtained images, and output fusion feature maps;

[0024] The genotyping module is configured to input the fused feature map, perform RSV genotyping using a neural network, and output the RSV genotyping results.

[0025] In a fourth aspect, the present invention also provides a respiratory syncytial virus typing device, the device includes a processor and a memory, the memory stores at least one instruction, at least one program, code set or instruction set, and the at least one instruction, at least one program, code set or instruction set is loaded and executed by the processor to implement the above-mentioned information processing method for respiratory syncytial virus typing.

[0026] In a fifth aspect, the present invention also provides a computer-readable storage medium, which stores at least one instruction, at least one program, code set or instruction set, and the at least one instruction, at least one program, code set or instruction set is loaded and executed by a processor to implement the above-mentioned information processing method for respiratory syncytial virus typing.

[0027] The present invention has the following beneficial effects:

[0028] The present invention innovatively proposes an information processing method for respiratory syncytial virus typing. Based on the original onehot encoding and zero padding strategies, an innovative hierarchical one hot encoding strategy is proposed, which reduces the interference of the traditional zero padding method on the one hot encoding and enhances the robustness to indel (insertion) mutations.

[0029] The present invention obtains two pictures based on one hot encoding, a picture of (X+1)*7*1 and a picture of 50*120~200*1, based on proposing a CNN network to extract the features of the two pictures and align them, enhance the ability to capture mutation information in RSV genome images, make CNN easier to capture detailed information about RSV genome sequences, there is a stronger correspondence between the features extracted by CNN, and therefore it is easier to capture the supplementary information for insertion mutations in the second picture during the fusion process, and improve the final RSV genotype classification accuracy. By contrast, the classification accuracy of the RSV genotype classification method provided by the present invention is better than that of the classification method based on the evolutionary tree, and the clustering error rate is lower.

[0030] The development of respiratory syncytial virus typing strategies will help doctors and researchers to diagnose and evaluate more accurately. By optimizing the feature extraction of respiratory genome sequences and applying deep learning technology, more accurate diagnostic tools can be provided for the public safety field, helping doctors and researchers to better understand respiratory syncytial virus and provide more effective treatment options. Based on the hierarchical one hot encoding strategy and neural network image classification technology, it can help doctors and researchers make accurate diagnoses and treatment guidance, improve work efficiency, and help improve the efficiency and quality of diagnosis and treatment in the medical and public safety industries. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments are briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without creative work.

[0032] Figure 1 The first picture of hierarchical ont hot encoding;

[0033] Figure 2 The second picture for hierarchical onto hot coding;

[0034] Figure 3A schematic diagram of the network architecture. DETAILED DESCRIPTION

[0035] References to embodiments of the present invention will now be provided in detail, one or more examples of which are described below. Each example is provided as an explanation rather than a limitation of the present invention. In fact, it will be apparent to those skilled in the art that various modifications and variations may be made to the present invention without departing from the scope or spirit of the present invention. For example, a feature illustrated or described as part of one embodiment may be used in another embodiment to produce a further embodiment.

[0036] In a first aspect, the present invention provides a method for constructing a respiratory syncytial virus typing model, comprising the following steps:

[0037] S1: Construction of reference genome: Select the respiratory syncytial virus genome used to construct the reference genome from the respiratory syncytial virus genome sequence database, and divide the sequence into training set, validation set and test set; perform sequence alignment on the genomes in some training sets, find out the similar regions and variant regions between all the genomes in some training sets, and construct a pan-genome map containing all genome information based on the alignment results; merge the variant regions to obtain the reference genome;

[0038] S2: Perform hierarchical one hot encoding: The length of the reference genome is X. The whole genome sequences of the training set and the validation set are aligned with the reference genome and one hot encoded to obtain a picture of query vector*key vector*value vector (X+1)*7*1; the key vector in the picture refers to the following seven types of information: whether there are A bases, T bases, C bases, G bases, missing or unknown bases, insertion mutations, and whether the bases at each position are located in the coding gene region of the G protein after alignment with the reference genome;

[0039] The whole genome sequences of the training set and the validation set are once again ont hot encoded to obtain a query vector*key vector*value vector of 50*120~200*1 images; the key vector in the 50*120~200*1 image refers to the number of insertion mutations; if there is at least one insertion mutation, it is encoded in the front row of the 50*120~200*1 image;

[0040] S3: Fusing hierarchical features: Using the CNN network, extract features from the (X+1)*7*1 image and the 50*120~200*1 image, respectively, obtain features of the target dimension, splice and fuse the obtained images, and obtain a fused feature map;

[0041] S4: RSV genotyping based on neural network: input fusion feature map, use neural network to perform RSV genotyping and output;

[0042] S5: Input the full genome of the respiratory syncytial virus in the test set and perform steps S2-S4.

[0043] The present invention first constructs a pan-genome based on a graphical pan-genome strategy as a reference genome. Based on the reference genome, the respiratory syncytial virus genome sequence is encoded, including A, T, C, G bases, as well as missing and unknown bases, inserted bases and the gene coding sequence of G protein (base weight). In order to more accurately and conveniently capture insertion variation information, a multi-graph strategy is used to represent possible insertions. Two images are obtained based on one hot encoding, and the ability to capture mutation information in RSV genome images is enhanced based on the proposed CNN network to extract features and align alignment, so that CNN is easier to capture detailed information about RSV genome sequences. There is a stronger correspondence between the features extracted by CNN, so it is easier to capture the supplementary information for insertion mutations in the second image during the fusion process, and improve the final RSV genotype classification accuracy.

[0044] After the RSV typing model is established using the training set and the validation set, the model is tested using the test set and can then be used to type the respiratory syncytial virus in the sample to be tested.

[0045] In one embodiment, deletions and unknowns are coded as N, insertions are coded as I, and G protein coding genes are coded as S. In other embodiments, substitutions can be made as needed, and are not limited to the above coding naming methods.

[0046] The base weights of the G protein gene sequence are added to perform accurate classification based on the G protein gene sequence variation information.

[0047] The ratio of the training set to the validation set is, for example, 6-8:2-4, such as 8:2, 6:4 or 7:3. The number of test sets can be adjusted as needed.

[0048] In a preferred embodiment of the present invention, step S1 includes merging the variant regions according to the following rules: 1) if a certain position is only a substitution variant, only the base with the highest variant frequency is retained; 2) if a certain position has a deletion variant, the non-deleted base is retained; 3) if a certain position has an insertion variant, the insertion variant is retained.

[0049] In a preferred embodiment of the present invention, the rules of ont hot encoding when obtaining a (X+1)*7*1 picture are as follows:

[0050] The 7 lines in the random picture are used as the coding lines in the coding gene region of A base, T base, C base, G base, insertion variation and G protein respectively;

[0051] After comparison with the reference genome, whether there is an A base at each position, if there is an A base, the coding line of the A base is coded as 1 at this position; otherwise, it is coded as 0; whether there is a T base at each position after comparison with the reference genome, if there is a T base, the coding line of the T base is coded as 1 at this position, otherwise it is 0; whether there is a C base at each position after comparison with the reference genome, if there is a C base, the coding line of the C base is coded as 1 at this position, otherwise it is 0; whether there is a G base at each position after comparison with the reference genome, if there is a G base, the coding line of the G base is coded as 1 at this position, otherwise it is 0; whether there is a missing or unknown base at each position after comparison with the reference genome, if there is a missing or unknown base, the coding line of the missing or unknown base is coded as 1 at this position, otherwise it is 0; whether there is an insertion variation at each position after comparison with the reference genome, if there is an insertion variation, the coding line is coded as 1 at this position, otherwise it is 0; whether each position is located in the coding gene region of the G protein, if so, it is coded as 1, otherwise it is 0.

[0052] In a preferred implementation of the present invention, the rules for obtaining the ont hot encoding of a 50*120-200*1 picture are as follows:

[0053] If there is no insertion variation in the sequence, all 0s are used for padding; if there is at least one insertion variation in the sequence, it is encoded in the front column of the 50*120~200*1 picture; if an insertion variation is 1bp long, the inserted base is one-hot encoded in rows 1 to 4; if an insertion variation is 2bp long, the inserted base is one-hot encoded in rows 1 to 8, and so on, until an insertion variation is 50bp long, in which case the inserted base is one-hot encoded in rows 1 to 200.

[0054] In the 50*120~200*1 picture, 120~200 refers to the number of bases of insertion mutation, which is 30~50. There are 4 possible bases in each insertion mutation. Therefore, the key vector is a multiple of 4.

[0055] In a preferred embodiment of the present invention, when obtaining a 50*120~200*1 picture, firstly perform a sequence comparison between the whole genome sequence of the training set that did not participate in the construction of the reference genome and / or the RSV whole genome sequence in the verification set and the reference genome, determine the number of inserted mutations in the training set that did not participate in the construction of the reference genome and / or the verification set, and adjust the key vector according to the number of inserted mutations.

[0056] This setting helps prevent the key vector of the second picture from being set up so that the number of inserted mutations is too small, resulting in the number of inserted mutations in some sequences in the genome sequence being too large to be captured, resulting in missing information on inserted mutations, and thus making it impossible to classify RSV more accurately. By comparing the whole genome sequences of the training set that did not participate in the construction of the reference genome and / or the whole genome sequences of RSV in the validation set with the reference genome, the number of inserted mutations in the training set and / or validation set that did not participate in the construction of the reference genome is determined, and the key vector is adjusted according to the number of inserted mutations, so that the key vector can be adaptively adjusted according to the information in the database.

[0057] In a preferred implementation of the present invention, when step S3 fuses the hierarchical features, the ResNet50 network backbone layer is used to extract features, and each layer has a 3×3 convolution layer, a Relu activation function, and a batch normalization process to extract features.

[0058] Based on the one hot encoding step, two images are obtained. The second image is an effective supplement to the first image. Therefore, fusing these two images can provide more comprehensive and rich features, thereby improving the final RSV genotype classification accuracy. The cross-modal features are fused using the feature weighted sum fusion method. The fusion layer weight matrix is ​​trainable. The initial weight value of multi-layer feature fusion is set to 1. The feature fusion weight value will be back-propagated and updated along with the classification loss.

[0059] In a preferred embodiment of the present invention, the neural network is selected from a BP neural network, a RNN neural network, a GRU neural network, a LSTM neural network or a Transformer architecture.

[0060] The Transformer architecture is preferred. Compared with BP neural network, RNN neural network and LSTM neural network, the Transformer architecture is more accurate in classification.

[0061] The present invention applies the Transformer model to virus classification, breaking through the limitations of traditional methods. The Transformer self-attention mechanism has the ability to establish global relationships, improving the model's ability to perceive the genome's full genome information. Through Transformer position encoding, the model can take into account the relative position of features in the RSV full genome sequence, thereby better capturing the spatial information of the image and improving the accuracy of RSV genotype classification.

[0062] In a preferred embodiment of the present invention, in step S4, the fused feature map is input, RSV genotyping is performed and output using the Transformer architecture; the feature encoding layer Tranformer encoder is composed of Multi-head Self-Attention (MSA) and MLP layers;

[0063] In a preferred embodiment of the present invention, Layer Norm normalization is used before executing the MSA and MLP layers, and residual connection is used after the layers;

[0064] In a preferred embodiment of the present invention, if the maximum probability of typing after the whole genome sequence of the tested sample is compared with the reference genome is less than 0.95, the tested sample is judged to be a new RSV genotype;

[0065] In a preferred embodiment of the present invention, when RSV genotyping, RSV type A is divided into GA1, GA2, GA2.1, GA2.2, GA2.3, GA2.3.1, GA2.3.2, GA2.3.2a, GA2.3.2b, GA2.3.3, GA2.3.5, GA3, GA3.0.1 and GA3.0.2; RSV type B is divided into GB1, GB2, GB4, GB5, GB5.0.1, GB5.0.2, GB5.0.3, GB5.0.4a and GB5.0.5a.

[0066] In one embodiment, the respiratory syncytial virus genome sequence database is selected from the nextstrain website, such as https: / / nextstrain.org / rsv / a / genome / 6y. The whole genome sequences of RSVA and RSVB were downloaded, and the sequences with a length of less than 12,000 were deleted using seqkit v2.2.0 software, and the sequences were retained for subsequent analysis.

[0067] In a second aspect, the present invention provides an information processing method for respiratory syncytial virus typing, which comprises the following steps:

[0068] A respiratory syncytial virus typing model is constructed according to the above-mentioned method for constructing a respiratory syncytial virus typing model, and then the full genome of the respiratory syncytial virus of the sample to be tested is input to perform steps S2-S4.

[0069] In a third aspect, the present invention also provides a device for typing respiratory syncytial virus, comprising: an input module, a control module and an output module;

[0070] The input module was configured to: input the respiratory syncytial virus genome;

[0071] The control modules include: reference genome construction module, hierarchical one hot encoding module, fusion feature module and genotyping module;

[0072] The output module includes: outputting respiratory syncytial virus typing results;

[0073] The reference genome construction module is configured as follows: based on the input respiratory syncytial virus genome, the sequence is divided into a training set, a validation set, and a test set; the genomes in some training sets are sequenced, similar regions and variant regions between all genomes in some training sets are found, and a pan-genome map containing all genome information is constructed based on the alignment results; the variant regions are merged to output the reference genome, training set, validation set, and test set;

[0074] The hierarchical one hot encoding module is configured as follows: input the whole genome sequences of the reference genome, training set and validation set, align the whole genome sequences of the training set and validation set with the reference genome, and perform one hot encoding to obtain a picture with query vector * key vector * value vector of (X+1)*7*1, where X is the length of the reference genome; the key vector in the picture refers to the following seven types of information: whether there are A bases, T bases, C bases, G bases, missing or unknown bases, inserted mutations and whether the bases at the position are located in the coding gene region of the G protein at each position after alignment with the reference genome; output a picture with query vector * key vector * value vector of (X+1)*7*1; perform one hot encoding on the whole genome sequences of the training set and validation set again. Hot encoding, output query vector * key vector * value vector is 50*120~200*1 picture; the key vector in the 50*120~200*1 picture refers to the number of insertion mutations; if there is at least one insertion mutation, it is encoded in the front column of the 50*120~200*1 picture;

[0075] The fusion feature module is configured as follows: input (X+1)*7*1 pictures and 50*120~200*1 pictures, use CNN network to extract features, obtain the features of target dimensions respectively, splice and fuse the obtained images, and output fusion feature maps;

[0076] The genotyping module is configured to input the fused feature map, perform RSV genotyping using a neural network, and output the RSV genotyping results.

[0077] In a fourth aspect, the present invention also provides a device for establishing a respiratory syncytial virus typing model, comprising: an input module, a control module and an output module;

[0078] The input module was configured to: input the RSV genome in the database;

[0079] The control modules include: reference genome construction module, hierarchical one hot encoding module, fusion feature module and genotyping module;

[0080] The output module includes: outputting respiratory syncytial virus typing model;

[0081] The reference genome construction module is configured as follows: based on the input respiratory syncytial virus genome, the sequence is divided into a training set, a validation set, and a test set; the genomes in some training sets are sequenced, similar regions and variant regions between all genomes in some training sets are found, and a pan-genome map containing all genome information is constructed based on the alignment results; the variant regions are merged to output the reference genome, training set, validation set, and test set;

[0082] The hierarchical one hot encoding module is configured as follows: input the whole genome sequences of the reference genome, training set and validation set, align the whole genome sequences of the training set and validation set with the reference genome, and perform one hot encoding to obtain a picture with query vector * key vector * value vector of (X+1)*7*1, where X is the length of the reference genome; the key vector in the picture refers to the following seven types of information: whether there are A bases, T bases, C bases, G bases, missing or unknown bases, inserted mutations and whether the bases at the position are located in the coding gene region of the G protein at each position after alignment with the reference genome; output a picture with query vector * key vector * value vector of (X+1)*7*1; perform one hot encoding on the whole genome sequences of the training set and validation set again. Hot encoding, output query vector * key vector * value vector is 50*120~200*1 picture; the key vector in the 50*120~200*1 picture refers to the number of insertion mutations; if there is at least one insertion mutation, it is encoded in the front column of the 50*120~200*1 picture;

[0083] The fusion feature module is configured as follows: input (X+1)*7*1 pictures and 50*120~200*1 pictures, use CNN network to extract features, obtain the features of target dimensions respectively, splice and fuse the obtained images, and output fusion feature maps;

[0084] The genotyping module is configured to input the fused feature map, perform RSV genotyping using a neural network, and output a respiratory syncytial virus typing model.

[0085] In a fifth aspect, the present invention also provides a respiratory syncytial virus typing device, the device includes a processor and a memory, the memory stores at least one instruction, at least one program, code set or instruction set, and the at least one instruction, at least one program, code set or instruction set is loaded and executed by the processor to implement the above-mentioned information processing method for respiratory syncytial virus typing.

[0086] In a sixth aspect, the present invention also provides a device for establishing a respiratory syncytial virus typing model, the device comprising a processor and a memory, the memory storing at least one instruction, at least one program, code set or instruction set, and the at least one instruction, at least one program, code set or instruction set is loaded and executed by the processor to implement the above-mentioned method for constructing a respiratory syncytial virus typing model.

[0087] Specifically, the electronic device may include a memory, a processor, a bus, and a communication interface, and the memory, processor, and communication interface are electrically connected to each other directly or indirectly to achieve data transmission or interaction. For example, these elements can be electrically connected to each other via one or more buses or signal lines. The processor can process information and / or data related to target identification to perform one or more functions described in this application.

[0088] The memory can be, but is not limited to, random access memory (RAM), read only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable read-only memory (EEPROM), etc.

[0089] The processor can be an integrated circuit chip with signal processing capabilities. The processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0090] In the seventh aspect, the present invention also provides a computer-readable storage medium, which stores at least one instruction, at least one program, code set or instruction set, and the at least one instruction, at least one program, code set or instruction set is loaded and executed by a processor to implement the above-mentioned information processing method for respiratory syncytial virus typing or the method for constructing a respiratory syncytial virus typing model.

[0091] In order to make the purpose, technical scheme and advantages of the embodiments of the present invention clearer, the technical scheme in the embodiments of the present invention will be described clearly and completely below. If the specific conditions are not specified in the embodiments, they are carried out according to conventional conditions or conditions recommended by the manufacturer. If the manufacturer of the reagents or instruments used is not specified, they are all conventional products that can be purchased commercially.

[0092] The features and performance of the present invention are further described in detail below in conjunction with the embodiments.

[0093] Example 1

[0094] The embodiment provides a respiratory syncytial virus typing method based on hierarchical one hot coding and transformer. First, a pan-genome is constructed based on a graphical pan-genome strategy as a reference genome. Based on the pan-genome, the sequence is encoded, where both missing and unknown are encoded using N, a multi-graph strategy is used to represent possible insertions, and the gene sequence base weights of the G protein are added. The features of multiple graphs are extracted based on CNN, and feature fusion is performed. Multimodal image classification technology based on Transformer.

[0095] 1. Construction of reference genome based on pan-genome concept

[0096] For subsequent hierarchical one hot coding, the present invention first constructs a reference genome based on the pan-genome concept to provide a reference basis for subsequent coding.

[0097] First, the whole genome sequences of RSVA and RSVB were downloaded from the nextstrain website (https: / / nextstrain.org / rsv / a / genome / 6y). Sequences with a length of less than 12,000 were removed using seqkit v2.2.0 software, and 11,325 sequences were retained for subsequent analysis. The sequencing sequences were divided into a training set and a validation set at an 8:2 ratio, that is, 8,000 sequences were used for the training set, 2,000 sequences were used for the validation set, and 1,325 sequences were used for the test set.

[0098] Then, a pan-genome was constructed based on the whole genome alignment strategy, that is, the multiple sequence alignment of 6000 respiratory syncytial viruses in the training set was performed using the mafft v7.526 software, without relying on a specific reference genome, and then the similar and different regions between all genomes were found, including various types of variations such as SNPs and Indels. Then, a pan-genome map containing all genomic information was constructed based on these alignment results. Finally, these variant regions were merged to obtain a reference genome based on the maximum frequency and maximum degree principles, that is, 1) if a certain position is only a substitution variation, only the base with the maximum frequency is retained; 2) if a certain position has a deletion variation, the non-deleted base is retained; 3) if a certain position has an insertion variation, the insertion variation is retained. Finally, the length of the reference genome is 26739bp.

[0099] Among the remaining 2000 RSV whole genome sequences in the training set and the 2000 RSV whole genome sequences in the validation set, 384 sequences had insertion mutations compared to the reference genome obtained in the present invention.

[0100] 2. Hierarchical one hot encoding strategy

[0101] Coding strategy for the first picture: The present invention innovatively proposes a hierarchical one hot coding strategy. After determining the reference genome, the RSV full genome sequences of the training set and the validation set are compared with the reference genome and encoded to obtain a 26740*7*1 picture. Among them, 26740 is the width of the picture, the previous 26739 refers to the length of the reference genome, and the last 1 is at the end of the picture to show that the RSV full genome sequence still has insertion mutations at the end of the gene compared with the reference genome. It is worth noting that the original RSV full genome sequence is about 15,000bp. The gap here is caused by the addition of a large number of insertion mutations based on the longest principle when constructing the reference genome. Among them, 7 refers to the height of the picture. The first row refers to whether there is an A base at a certain position after comparison with the reference genome. If there is an A base, it is encoded as 1, otherwise it is 0; the second row refers to whether there is a T base at a certain position after comparison with the reference genome. If there is a T base, the encoding row of the T base is encoded as 1 at this position, otherwise it is 0; the third row refers to whether there is a C base at a certain position after comparison with the reference genome. If there is a C base, the encoding row of the C base is encoded as 1 at this position, otherwise it is 0; the fourth row refers to whether there is a G base at a certain position after comparison with the reference genome. If there is a G base, the encoding row of the G base is encoded as 1 at this position, otherwise it is 0; At the G base, the coding line of the G base is coded as 1 at this position, otherwise it is 0; the fifth line refers to whether there is a missing or unknown base at a certain position after comparison with the reference genome. If there is a missing or unknown base, the coding line of the missing or unknown base is coded as 1 at this position, otherwise it is 0; the sixth line refers to whether there is an insertion at a certain position after comparison with the reference genome. If there is an insertion variation, the coding line at this position is coded as 1, otherwise it is 0 and coded in another figure; the seventh line refers to whether the base is located in the coding gene region of the G protein. If it is, it is coded as 1, otherwise it is 0. The first picture of the hierarchical ont hot coding is the coding principle diagram. Figure 1 shown.

[0102] Coding strategy for the second figure: In step 1, through comparison, it was found that in the 384 sequences with insertion mutations, the length of the insertion was less than 50bp and the number of insertions was less than 40. Therefore, the present invention creates a 50*200*1 figure to represent insertion mutations. If there is no mutation, all are filled with 0; if there is 1 insertion mutation, it is encoded in the first column of this figure, and the remaining columns are filled with 0. If there are two mutations, they are encoded in the first two columns, and so on, up to 50 insertion mutations can be included. If there are more than 50 mutations, it may be a new genotype or the sequence of another virus. If an insertion mutation is 1bp in length, then the inserted base is one hot encoded in rows 1 to 4; if an insertion mutation is 2bp in length, then the inserted base is one hot encoded in rows 1 to 8, and so on, until 50bp. Refer to the second figure of the encoding principle diagram of hierarchical ont hot encoding. Figure 2 shown.

[0103] 3. Hierarchical feature fusion

[0104] The present invention obtains two pictures based on the one hot encoding step. The second picture is an effective supplement to the first picture. Therefore, the fusion of these two pictures can provide more comprehensive and rich features, thereby improving the final RSV genotype classification accuracy. Figure 3 As shown, the hierarchical feature fusion steps are as follows:

[0105] First, the CNN network (the present invention uses ResNet50) is used to extract features from the two images. The input dimensions of the first image are all 26740×7×1, and the features are extracted through the ResNet50 network backbone layer composed of residual blocks. Each layer has a 3×3 convolution layer, Relu activation function and batch normalization processing to extract features. Since the image is based on onehot encoding, it is slightly reduced to 896×7×1 features through a 3-layer convolutional network.

[0106] The input dimension of the second image is 50×200×1. The features are also extracted through the ResNet50 network backbone layer composed of residual blocks. Each layer has a 3×3 convolution layer, Relu activation function and batch normalization processing to extract features. In order to facilitate fusion, the second image is also converted to 128×7×1 features through 3 layers of convolution. The obtained images are spliced ​​and fused to obtain the final feature map of 1024×7×1 for classification.

[0107] In the present invention, a feature weighted sum fusion method is used to fuse cross-modal features. The fusion layer weight matrix is ​​trainable, and the initial weight value of multi-layer feature fusion is set to 1. The feature fusion weight value will be back-propagated and updated along with the classification loss.

[0108] 4. Transformer-based multimodal RSV classification

[0109] In this paper, we explored the Transformer architecture for RSV. The network receives a feature map of size 1024×7×1 as input data. The feature encoding layer Transformer encoder consists of Multi-head Self-Attention (MSA) and MLP layers proposed by Vaswani. Layer Norm is used before executing MSA and MLP layers, and residual connections are used after the layers. MLP contains two fully connected layers with GELU nonlinearity.

[0110] We set up 23 categories in the classification layer. RSVA includes 14 categories, namely GA1, GA2, GA2.1, GA2.2, GA2.3, GA2.3.1, GA2.3.2, GA2.3.2a, GA2.3.2b, GA2.3.3, GA2.3.5, GA3, GA3.0.1, GA3.0.2; RSVB includes 9 categories, namely GB1, GB2, GB4, GB5, GB5.0.1, GB5.0.2, GB5.0.3, GB5.0.4a and GB5.0.5a. According to the probability value of typing, samples with the maximum probability value of typing less than 0.95 are defined as new RSV genotypes.

[0111] The Transformer architecture for RSV classification constructed in the present invention utilizes the self-attention mechanism of Transformer to improve the classification model's perception of the global information of RSV sequences, improve the ability to model complex relationships, and ultimately improve the accuracy of RSV classification.

[0112] Comparative Example 1

[0113] This comparative example is based on the construction of an evolutionary tree for RSV typing:

[0114] 1) Construct an evolutionary tree based on 10,000 sequences; 2) Use mafft software for multiple sequence alignment and iqtree software to construct an evolutionary tree with a bootstrap of 1000. Based on the evolutionary tree, use pplacer software to classify the 1,325 RSV sequences in the test set to see which branch these RSV sequences belong to.

[0115] Experimental Example 1

[0116] According to the typing method provided in Example 1 of the present invention, among 1325 RSV sequences, 1300 sequences were predicted correctly and 15 sequences were predicted incorrectly, with an accuracy rate of 98.87%;

[0117] According to the typing method provided in Comparative Example 1, 1279 sequences were predicted correctly, 36 sequences were predicted incorrectly, and the accuracy rate was 96.53%. In the evolutionary tree constructed from 10,000 sequences, there were 11 sequences with clustering errors.

[0118] Comprehensive comparison shows that the present invention is superior to the classification based on evolutionary tree.

[0119] Example 2

[0120] The present invention also provides a respiratory syncytial virus typing device, comprising: an input module, a control module and an output module;

[0121] The input module was configured to: input the respiratory syncytial virus genome;

[0122] The control modules include: reference genome construction module, hierarchical one hot encoding module, fusion feature module and genotyping module;

[0123] The output module includes: outputting respiratory syncytial virus typing results;

[0124] The reference genome construction module is configured as follows: based on the input respiratory syncytial virus genome, the sequence is divided into a training set, a validation set, and a test set; the genomes in part of the training set are sequenced to find the similar regions and variant regions between all the genomes in the training set, and a pan-genome map containing all genome information is constructed based on the alignment results; the variant regions are merged to output the reference genome, training set, validation set, and test set;

[0125] The hierarchical one hot encoding module is configured as follows: input the whole genome sequences of the reference genome, training set and validation set, align the whole genome sequences of the training set and validation set with the reference genome, and perform one hot encoding to obtain a picture with query vector * key vector * value vector of (X+1)*7*1, where X is the length of the reference genome; the key vector in the picture refers to the following seven types of information: whether there are A bases, T bases, C bases, G bases, missing or unknown bases, inserted mutations and whether the bases at the position are located in the coding gene region of the G protein at each position after alignment with the reference genome; output a picture with query vector * key vector * value vector of (X+1)*7*1; perform one hot encoding on the whole genome sequences of the training set and validation set again. Hot encoding, output query vector * key vector * value vector is 50*120~200*1 picture; the key vector in the 50*120~200*1 picture refers to the number of insertion mutations; if there is at least one insertion mutation, it is encoded in the front column of the 50*120~200*1 picture;

[0126] The fusion feature module is configured as follows: input (X+1)*7*1 pictures and 50*120~200*1 pictures, use CNN network to extract features, obtain the features of target dimensions respectively, splice and fuse the obtained images, and output fusion feature maps;

[0127] The genotyping module is configured to input the fused feature map, perform RSV genotyping using a neural network, and output the RSV genotyping results.

[0128] Example 3

[0129] The present invention also provides a device for establishing a respiratory syncytial virus typing model, comprising: an input module, a control module and an output module;

[0130] The input module was configured to: input the RSV genome in the database;

[0131] The control modules include: reference genome construction module, hierarchical one hot encoding module, fusion feature module and genotyping module;

[0132] The output module includes: outputting respiratory syncytial virus typing model;

[0133] The reference genome construction module is configured as follows: based on the input respiratory syncytial virus genome, the sequence is divided into a training set, a validation set, and a test set; the genomes in some training sets are sequenced, similar regions and variant regions between all genomes in some training sets are found, and a pan-genome map containing all genome information is constructed based on the alignment results; the variant regions are merged to output the reference genome, training set, validation set, and test set;

[0134] The hierarchical one hot encoding module is configured as follows: input the whole genome sequences of the reference genome, training set and validation set, align the whole genome sequences of the training set and validation set with the reference genome, and perform one hot encoding to obtain a picture with query vector * key vector * value vector of (X+1)*7*1, where X is the length of the reference genome; the key vector in the picture refers to the following seven types of information: whether there are A bases, T bases, C bases, G bases, missing or unknown bases, inserted mutations and whether the bases at the position are located in the coding gene region of the G protein at each position after alignment with the reference genome; output a picture with query vector * key vector * value vector of (X+1)*7*1; perform one hot encoding on the whole genome sequences of the training set and validation set again. Hot encoding, output query vector * key vector * value vector is 50*120~200*1 picture; the key vector in the 50*120~200*1 picture refers to the number of insertion mutations; if there is at least one insertion mutation, it is encoded in the front column of the 50*120~200*1 picture;

[0135] The fusion feature module is configured as follows: input (X+1)*7*1 pictures and 50*120~200*1 pictures, use CNN network to extract features, obtain the features of target dimensions respectively, splice and fuse the obtained images, and output fusion feature maps;

[0136] The genotyping module is configured to input the fused feature map, perform RSV genotyping using a neural network, and output a respiratory syncytial virus typing model.

[0137] Example 4

[0138] The present invention also provides a respiratory syncytial virus typing device, which includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, at least one program, a code set or an instruction set is loaded and executed by the processor to implement the above-mentioned information processing method for respiratory syncytial virus typing.

[0139] Example 5

[0140] The present invention also provides a device for establishing a respiratory syncytial virus typing model, the device comprising a processor and a memory, the memory storing at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, at least one program, a code set or an instruction set is loaded and executed by the processor to implement the above-mentioned method for constructing a respiratory syncytial virus typing model.

[0141] Specifically, the electronic device may include a memory, a processor, a bus, and a communication interface, and the memory, processor, and communication interface are electrically connected to each other directly or indirectly to achieve data transmission or interaction. For example, these elements can be electrically connected to each other via one or more buses or signal lines. The processor can process information and / or data related to target identification to perform one or more functions described in this application.

[0142] The memory can be, but is not limited to, random access memory (RAM), read only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable read-only memory (EEPROM), etc.

[0143] The processor can be an integrated circuit chip with signal processing capabilities. The processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0144] Example 6

[0145] The present invention also provides a computer-readable storage medium, which stores at least one instruction, at least one program, code set or instruction set, and the at least one instruction, at least one program, code set or instruction set is loaded and executed by a processor to implement the above-mentioned information processing method for respiratory syncytial virus typing or the method for constructing a respiratory syncytial virus typing model.

[0146] The present invention innovatively proposes an information processing method for respiratory syncytial virus typing. Based on the original onehot encoding and zero padding strategies, an innovative hierarchical one hot encoding strategy is proposed, which reduces the interference of the traditional zero padding method on the one hot encoding and enhances the robustness to indel (insertion) mutations.

[0147] The present invention obtains two pictures based on one hot encoding, a picture of (X+1)*7*1 and a picture of 50*120~200*1, based on proposing a CNN network to extract the features of the two pictures and align them, enhance the ability to capture mutation information in RSV genome images, make CNN easier to capture detailed information about RSV genome sequences, there is a stronger correspondence between the features extracted by CNN, and therefore it is easier to capture the supplementary information for insertion mutations in the second picture during the fusion process, and improve the final RSV genotype classification accuracy. By contrast, the classification accuracy of the RSV genotype classification method provided by the present invention is better than that of the classification method based on the evolutionary tree, and the clustering error rate is lower.

[0148] The development of respiratory syncytial virus typing strategies will help doctors and researchers to diagnose and evaluate more accurately. By optimizing the feature extraction of respiratory genome sequences and applying deep learning technology, more accurate diagnostic tools can be provided for the public safety field, helping doctors and researchers to better understand respiratory syncytial virus and provide more effective treatment options. Based on the hierarchical one hot encoding strategy and neural network image classification technology, it can help doctors and researchers make accurate diagnoses and treatment guidance, improve work efficiency, and help improve the efficiency and quality of diagnosis and treatment in the medical and public safety industries.

[0149] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for constructing a respiratory syncytial virus typing model, characterized in that: It includes the following steps: S1: Construction of reference genome: Selecting a respiratory syncytial virus genome for construction of a reference genome from the respiratory syncytial virus genome sequence database, dividing the sequence into a training set, a validation set, and a test set; performing sequence alignment on the genomes in part of the training set, finding similar regions and variant regions between all genomes in the said part of the training set, and constructing a pan-genome map containing all genome information based on the alignment results; merging variant regions to obtain a reference genome; S2: perform hierarchical one hot encoding: the length of the reference genome is X, the whole genome sequences of the training set and the validation set are aligned with the reference genome, and one hot encoding is performed to obtain a picture of query vector * key vector * value vector (X+1)*7*1; the key vector in the picture refers to the following seven types of information: whether there is an A base, a T base, a C base, a G base, a missing or unknown base, an insertion variation, and whether the base at the position is located in the coding gene region of the G protein after alignment with the reference genome; The whole genome sequences of the training set and the validation set are again ont hot encoded to obtain a query vector*key vector*value vector of a picture of 50*120~200*1; the key vector in the picture of 50*120~200*1 refers to the number of insertion mutations; if there is at least one insertion mutation, it is encoded in the front column of the picture of 50*120~200*1; S3: Fusing hierarchical features: using the CNN network to extract features of the (X+1)*7*1 image and the 50*120-200*1 image, respectively, to obtain features of the target dimensions, and splicing and fusing the obtained images to obtain a fused feature map; S4: RSV genotyping based on neural network: input fusion feature map, use neural network to perform RSV genotyping and output; S5: Input the full genome of the respiratory syncytial virus in the test set and perform steps S2-S4.

2. The method for constructing a respiratory syncytial virus typing model according to claim 1, characterized in that: The step S1 includes merging the variant regions according to the following rules: 1) if a certain position is only a substitution variant, only the base with the highest variant frequency is retained; 2) if a certain position has a deletion variant, the non-deleted base is retained; 3) if a certain position has an insertion variant, the insertion variant is retained.

3. The method for constructing a respiratory syncytial virus typing model according to claim 1, characterized in that: The rules for obtaining the ont hot encoding of a (X+1)*7*1 image are as follows: The 7 lines in the random picture are used as the coding lines in the coding gene region of A base, T base, C base, G base, insertion variation and G protein respectively; After comparison with the reference genome, whether there is an A base at each position, if there is an A base, the coding line of the A base is coded as 1 at this position; otherwise, it is coded as 0; whether there is a T base at each position after comparison with the reference genome, if there is a T base, the coding line of the T base is coded as 1 at this position, otherwise it is 0; whether there is a C base at each position after comparison with the reference genome, if there is a C base, the coding line of the C base is coded as 1 at this position, otherwise it is 0; whether there is a G base at each position after comparison with the reference genome, if there is a G base, the coding line of the G base is coded as 1 at this position, otherwise it is 0; whether there is a missing or unknown base at each position after comparison with the reference genome, if there is a missing or unknown base, the coding line of the missing or unknown base is coded as 1 at this position, otherwise it is 0; whether there is an insertion variation at each position after comparison with the reference genome, if there is an insertion variation, the coding line is coded as 1 at this position, otherwise it is 0; whether each position is located in the coding gene region of the G protein, if so, it is coded as 1, otherwise it is 0.

4. The method for constructing a respiratory syncytial virus typing model according to claim 1, characterized in that: The rules for obtaining ont hot encoding when obtaining images of 50*120 to 200*1 are as follows: If there is no insertion variation in the sequence, all the sequences are filled with 0; if there is at least one insertion variation in the sequence, it is encoded in the front column of the 50*120~200*1 image; if the length of an insertion variation is 1bp, the inserted base is one hot encoded in rows 1 to 4; If the length of an insertion variant is 2 bp, then the inserted bases are one-hot encoded in lines 1 to 8, and so on, until the length of an insertion variant is 50 bp, then the inserted bases are one-hot encoded in lines 1 to 200.

5. The method for constructing a respiratory syncytial virus typing model according to claim 4, characterized in that: When obtaining a 50*120-200*1 image, first perform sequence alignment on the whole genome sequence of the training set that does not participate in the construction of the reference genome and / or the RSV whole genome sequence in the verification set with the reference genome, determine the number of inserted mutations in the training set that does not participate in the construction of the reference genome and / or the verification set, and adjust the key vector according to the number of inserted mutations; Preferably, when step S3 fuses the hierarchical features, the ResNet50 network backbone layer is used to extract features, and each layer has a 3×3 convolution layer, a Relu activation function and a batch normalization process to extract features.

6. The method for constructing a respiratory syncytial virus typing model according to claim 1, characterized in that: The neural network is selected from BP neural network, RNN neural network, LSTM neural network, GRU neural network or Transformer architecture; Preferably, in step S4, the fused feature map is input, RSV genotyping is performed using the Transformer architecture and output; the feature encoding layer Tranformer encoder is composed of Multi-head Self-Attention (MSA) and MLP layers; Preferably, use Layer Norm normalization before performing MSA and MLP layers, and use residual connections after the layers; Preferably, if the maximum probability of typing after the whole genome sequence of the tested sample is compared with the reference genome is less than 0.95, the tested sample is judged to be a new RSV genotype; Preferably, when RSV genotyping, RSV type A is divided into GA1, GA2, GA2.1, GA2.2, GA2.3, GA2.3.1, GA2.3.2, GA2.3.2a, GA2.3.2b, GA2.3.3, GA2.3.5, GA3, GA3.0.1 and GA3.0.2; RSV type B is divided into GB1, GB2, GB4, GB5, GB5.0.1, GB5.0.2, GB5.0.3, GB5.0.4a and GB5.0.5a.

7. An information processing method for respiratory syncytial virus typing, characterized in that: It includes the following steps: A respiratory syncytial virus typing model is constructed according to the method for constructing a respiratory syncytial virus typing model according to any one of claims 1 to 6, and then the whole genome of the respiratory syncytial virus of the sample to be tested is input to perform steps S2 to S4.

8. A device for typing respiratory syncytial virus, characterized in that: include: Input module, control module and output module; The input module is configured to: input a respiratory syncytial virus genome; The control module includes: a reference genome construction module, a hierarchical one hot encoding module, a fusion feature module and a genotyping module; The output module includes: outputting respiratory syncytial virus typing results; The reference genome construction module is configured as follows: based on the input respiratory syncytial virus genome, the sequence is divided into a training set, a validation set and a test set; the genomes in part of the training set are sequenced to find similar regions and variant regions between all genomes in the part of the training set, and a pan-genome map containing all genome information is constructed based on the alignment results; the variant regions are merged to output the reference genome, the training set, the validation set and the test set; The hierarchical one hot encoding module is configured as follows: inputting the whole genome sequences of the reference genome, the training set and the validation set, aligning the whole genome sequences of the training set and the validation set with the reference genome, and performing one hot encoding to obtain a picture of query vector*key vector*value vector (X+1)*7*1, where X is the length of the reference genome; the key vector in the picture refers to the following seven types of information: whether there are A bases, T bases, C bases, G bases, missing or unknown bases, insertion mutations and whether the bases at the positions are located in the coding gene region of the G protein at each position after alignment with the reference genome; outputting a picture of query vector*key vector*value vector (X+1)*7*1; performing one hot encoding again on the whole genome sequences of the training set and the validation set. Hot encoding, output query vector * key vector * value vector is 50*120~200*1 picture; the key vector in the 50*120~200*1 picture refers to the number of insertion mutations; if there is at least one insertion mutation, it is encoded in the front column of the 50*120~200*1 picture; The fusion feature module is configured to: input the (X+1)*7*1 picture and the 50*120-200*1 picture, use the CNN network to perform feature extraction, obtain the features of the target dimension respectively, splice and fuse the obtained images, and output the fusion feature map; The genotyping module is configured to: input the fusion feature map, perform RSV genotyping using a neural network and output a respiratory syncytial virus typing result.

9. A respiratory syncytial virus typing device, characterized in that: The device includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the information processing method for respiratory syncytial virus typing as described in claim 7.

10. A computer-readable storage medium, characterized in that: The storage medium stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set are loaded and executed by the processor to implement the information processing method for respiratory syncytial virus typing as described in claim 7.

Citation Information

Patent Citations

  • Intelligent viral pneumonia diagnosis system based on multi-modal information fusion

    CN112530578A

  • Patient similarity classification method based on multi-source information

    CN115083550A

  • Variation pathogenicity annotation method, prediction variation effect atlas construction method and prediction variation effect atlas construction system

    CN117976040A

  • Method and system for predicting base editing efficiency

    WO2024164131A1

  • KR20190048926A