State prediction method and apparatus, device, storage medium and program product

By comparing gene sequence information with reference sequences to identify abnormal locations and types, and combining this with detection state information, a deep learning model is used for state prediction. This solves the problem of low accuracy in gene state prediction in traditional technologies and achieves more accurate state prediction.

WO2026113647A1PCT designated stage Publication Date: 2026-06-04TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
TENCENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2025-10-09
Publication Date
2026-06-04

Smart Images

  • Figure CN2025126418_04062026_PF_FP_ABST
    Figure CN2025126418_04062026_PF_FP_ABST
Patent Text Reader

Abstract

A state prediction method, which is executed by a computer device. The method comprises: acquiring gene sequence information of an object, and comparing the gene sequence information with a reference gene sequence to obtain abnormal sequence information in the gene sequence information, the abnormal sequence information being used for representing the abnormal position and the abnormal type of an abnormal gene unit in the gene sequence information (S202); acquiring detection state information of the object (S204); and inputting the abnormal sequence information and the detection state information into a state prediction model, so as to extract abnormal sequence features of the abnormal sequence information and detection state features of the detection state information, and, on the basis of the abnormal sequence features and the detection state features, obtaining prediction state information of the object (S206).
Need to check novelty before this filing date? Find Prior Art

Description

State prediction methods, apparatus, devices, storage media, and program products

[0001] Related applications

[0002] This application claims priority to Chinese patent application filed on November 29, 2024, with application number 202411735822.4, entitled "State Prediction Method, Apparatus, Device, Storage Medium and Program Product", the entire contents of which are incorporated herein by reference. Technical Field

[0003] This application relates to the field of artificial intelligence technology, and in particular to a state prediction method, apparatus, device, storage medium, and program product. Background Technology

[0004] In the medical field, the formation of gene states is influenced by a variety of factors, such as genetic, environmental, and emotional factors. Therefore, analyzing the process of gene state formation is crucial for predicting and treating gene states.

[0005] Traditional techniques typically employ statistical methods to collect extensive gene sample information and construct tables that characterize the relationship between macroscopic features such as regional and emotional traits and state labels. In practical applications, the state prediction result can be determined from the table based on the object's macroscopic characteristics. However, the accuracy of state prediction results using traditional techniques is relatively low. Summary of the Invention

[0006] This application provides a state prediction method, apparatus, device, storage medium, and program product.

[0007] This application provides a state prediction method. The method includes:

[0008] Obtain the gene sequence information of the object, compare the gene sequence information with the reference gene sequence, and obtain abnormal sequence information in the gene sequence information; the abnormal sequence information is used to characterize the abnormal location and abnormal type of the abnormal gene unit in the gene sequence information;

[0009] Obtain the detection status information of the object;

[0010] The abnormal sequence information and the detection state information are input into the state prediction model to extract the abnormal sequence features of the abnormal sequence information and the detection state features of the detection state information, and the predicted state information of the object is obtained based on the abnormal sequence features and the detection state features.

[0011] This application also provides a state prediction device. The device includes:

[0012] The sequence acquisition module is used to acquire the gene sequence information of the object, compare the gene sequence information with the reference gene sequence, and acquire abnormal sequence information in the gene sequence information; the abnormal sequence information includes the abnormal location and abnormal type of the gene unit;

[0013] A status acquisition module is used to acquire the detection status information of the object;

[0014] The state analysis module is used to input the abnormal sequence information and the detection state information into the state prediction model to extract the abnormal sequence features of the abnormal sequence information and the detection state features of the detection state information, and obtain the predicted state information of the object based on the abnormal sequence features and the detection state features.

[0015] A computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps described in the state prediction method above.

[0016] A computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, implements the steps described in the state prediction method above.

[0017] A computer program product. The computer program product includes a computer program that, when executed by a processor, implements the steps described in the state prediction method above.

[0018] Details of one or more embodiments of this application are set forth in the following drawings and description. Other features, objects, and advantages of this application will become apparent from the specification, drawings, and claims. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the published drawings without creative effort.

[0020] Figure 1 shows the application environment of the state prediction method in one embodiment;

[0021] Figure 2 is a schematic diagram of the first process of a state prediction method in one embodiment;

[0022] Figure 3 is a schematic diagram of the contents of abnormal sequence information in one embodiment;

[0023] Figure 4 is a schematic diagram of the first architecture of a state prediction model in one embodiment;

[0024] Figure 5 is a schematic diagram of the contents of the anomaly feature matrix in one embodiment;

[0025] Figure 6 is a schematic diagram of the second architecture of a state prediction model in one embodiment;

[0026] Figure 7 is a schematic diagram of the third architecture of the state prediction model in one embodiment;

[0027] Figure 8 is a schematic diagram of the second process of the state prediction method in one embodiment;

[0028] Figure 9 is a schematic diagram of the third process of the state prediction method in one embodiment;

[0029] Figure 10 is a flowchart of a gene disease type prediction model in one embodiment;

[0030] Figure 11 is a structural block diagram of a state prediction device in one embodiment;

[0031] Figure 12 is an internal structure diagram of a computer device in one embodiment;

[0032] Figure 13 is an internal structural diagram of a computer device in another embodiment. Detailed Implementation

[0033] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0034] The state prediction method provided in this application embodiment can be applied to the application environment shown in Figure 1. The terminal 102 communicates with the server 104 via a network. A data storage system can store the data that the server 104 needs to process. The data storage system can be set up independently, integrated into the server 104, or placed in the cloud or on other servers. The terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, projection devices, etc. Portable wearable devices can be smartwatches, smart bracelets, head-mounted devices, etc. Head-mounted devices can be virtual reality (VR) devices, augmented reality (AR) devices, smart glasses, etc. The server 104 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.

[0035] Both the terminal and the server can be used independently to execute the state prediction method provided in the embodiments of this application.

[0036] For example, the terminal acquires the object's gene sequence information, compares it with a reference gene sequence, and obtains abnormal sequence information from the gene sequence. This abnormal sequence information characterizes the abnormal location and type of abnormal gene units within the gene sequence. The terminal also acquires the object's detection status information, and uses a status prediction model to extract abnormal sequence features from the abnormal sequence information and detection status features from the detection status information. Based on these abnormal sequence features and detection status features, the terminal obtains the object's predicted status information.

[0037] Terminals and servers can also work together to execute the state prediction method provided in the embodiments of this application.

[0038] For example, the terminal acquires the gene sequence information of an object, compares it with a reference gene sequence, identifies abnormal sequence information within the gene sequence, and sends the object's detection status information and abnormal sequence information to the server. The server then uses a status prediction model to extract abnormal sequence features from the abnormal sequence information and detection status features from the detection status information. Based on these abnormal sequence features and detection status features, the server obtains the object's predicted status information and returns it to the terminal.

[0039] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data shall comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0040] In one embodiment, as shown in Figure 2, a state prediction method is provided. Taking the application of this method to a computer device as an example, the computer device can be a terminal or a server. The method can be executed independently by the terminal or server, or it can be implemented through interaction between the terminal and the server. Referring to Figure 2, the state prediction method includes the following steps:

[0041] Step S202: Obtain the gene sequence information of the object, compare the gene sequence information with the reference gene sequence, and obtain the abnormal sequence information in the gene sequence information; the abnormal sequence information is used to characterize the abnormal location and abnormal type of the abnormal gene unit in the gene sequence information.

[0042] The gene sequence information refers to the base sequence characterizing the genetic information of the object, specifically a sequence generated by arranging and combining adenine (A), thymine (T), cytosine (C), and guanine (G). The object's gene sequence information can be obtained by detecting samples of the object's tissues, such as cells, blood, or hair. In this embodiment, the gene sequence information can be a gene sequence matching a specific gene fragment of the object, or it can be the object's complete genome sequence.

[0043] In one scenario, the gene sequence information is the whole genome sequence information. The gene sequence information of an object can be pre-stored in a data storage system, and the data storage system stores the gene sequence information of multiple sample objects. In this scenario, in response to a state detection request for an object, the computer device queries and retrieves the object's gene sequence information from the data storage system based on the object's identifier.

[0044] In another scenario, the gene sequence information is the sequence of a specific gene fragment. The computer device can first obtain the whole gene sequence information of the object from the data storage system, and then extract the gene sequence information that matches the specific gene fragment from the whole gene sequence.

[0045] A reference gene sequence refers to the standard gene sequence of the species corresponding to the object, which also includes the four bases A, T, C, and G. It's important to note that although the base types of the gene sequence information and the reference gene sequence are the same, in practical applications, various uncertain factors such as environment, emotions, and lifestyle habits can influence the object's gene sequence information, potentially causing changes in the base positions, base arrangements, and gene sequence length—a phenomenon known as gene mutation. This leads to a certain degree of difference between the gene sequence information and the reference gene sequence. In other words, gene mutation refers to the phenomenon where various uncertain factors such as environment, emotions, and lifestyle habits affect the object's gene sequence information, causing changes in the base positions, base arrangements, and gene sequence length, resulting in differences between the gene sequence information and the reference gene sequence.

[0046] In this embodiment, abnormal gene sequences are identified by comparing gene sequence information with a reference gene sequence. Then, the abnormal positions and types of each abnormal gene unit within the abnormal gene sequence are analyzed to obtain abnormal sequence information, thus characterizing the microscopic gene features of the object at a fine-grained level. The abnormal position of the abnormal gene unit includes the chromosome where the abnormal gene unit is located and its absolute position in the reference gene sequence. The mutation type of the abnormal gene unit includes the current base type of the abnormal gene unit and the normal base type that matched the abnormal gene unit before mutation.

[0047] Abnormal gene sequences are determined by comparing gene sequence information with a reference gene sequence. They contain abnormal gene units that have undergone gene mutations, and in some cases, normal gene units at adjacent sites of each abnormal gene unit.

[0048] Optionally, after obtaining gene sequence information and a reference gene sequence, the gene sequence information and the reference gene sequence are compared at the same site using a site-by-site comparison method. Abnormal gene units that do not match the reference gene sequence, the corresponding reference gene units, and the locations of the abnormal gene units are screened from the gene sequence information. Then, base type identification is performed on the abnormal gene units and the reference gene units to determine the abnormal type of the abnormal gene units and to identify the locations of the abnormal gene units as abnormal positions. Finally, the abnormal type and abnormal position of each abnormal gene unit are summarized to obtain the abnormal sequence information in the gene sequence information.

[0049] In this process, given the gene sequence information and a reference gene sequence, two gene units at the same locus in the gene sequence information and the reference gene sequence are compared site by site. For two gene units at the same locus, if the gene unit in the gene sequence information is different from the gene unit in the reference gene sequence, the gene unit is determined to be an abnormal gene unit. If a gene unit exists at a certain locus in the gene sequence information, but the corresponding gene unit does not exist in the reference gene sequence, the gene unit at that locus in the reference gene sequence is a deleted gene unit. If both the gene sequence information and the reference gene sequence have gene units at a certain locus, but they are of different types, the gene unit at that locus in the gene sequence information is a replaced gene unit. After identifying abnormal gene units, their chromosome location and absolute position in the reference gene sequence are recorded as the abnormal position, and the current base type and the normal base type that matched before mutation are recorded as the abnormal type. Finally, the abnormal type and abnormal position of each abnormal gene unit are summarized to obtain the abnormal sequence information in the gene sequence information.

[0050] Optionally, after obtaining gene sequence information and a reference gene sequence, the abnormal gene segments in the gene sequence are identified by comparing the gene sequence information and the reference gene sequence according to the gene fragments. Then, the abnormal gene segments are compared with the reference gene segments site by site to determine the abnormal position and abnormal type of each abnormal gene unit in the abnormal gene segments, thereby obtaining abnormal sequence information.

[0051] In this process, after obtaining gene sequence information and a reference gene sequence, the gene sequence information and the reference gene sequence are compared according to gene fragments. The gene sequence information and the reference gene sequence are divided into gene fragments of equal length, and gene fragments at corresponding positions are compared. If a gene fragment in the gene sequence information differs from a gene fragment in the reference gene sequence, that gene fragment is an abnormal gene fragment. Next, the abnormal gene fragment is compared site-by-site with the reference gene fragment. For each gene unit in the abnormal gene fragment, if it differs from a gene unit at a corresponding site in the reference gene fragment, it is determined to be an abnormal gene unit. If a gene unit at a certain site exists in the abnormal gene fragment, but the corresponding site in the reference gene fragment does not exist, then that gene unit is a newly added gene unit; if a gene unit at a certain site does not exist in the abnormal gene fragment, but the corresponding site in the reference gene fragment does exist, then the gene unit at that site in the reference gene fragment is a deleted gene unit; if both the abnormal gene fragment and the reference gene fragment contain gene units at a certain site, but of different types, then the gene unit at that site in the abnormal gene fragment is a substituted gene unit. After identifying the abnormal gene unit, its location on the chromosome and its absolute position in the reference gene sequence are recorded as the abnormal location. The current base type of the abnormal gene unit and the normal base type it matched before mutation are recorded as the abnormal type. Finally, the abnormal type and abnormal location of each abnormal gene unit are summarized to obtain the abnormal sequence information.

[0052] In practical applications, abnormal gene sequences can include not only abnormal gene units that have undergone gene mutations, but also normal gene units at adjacent sites of each abnormal gene unit. In this case, the abnormal position of each abnormal gene unit in the abnormal sequence information obtained by organizing each gene unit of the abnormal gene sequence can be represented by the gene loci of adjacent gene units.

[0053] Step S204: Obtain the detection status information of the object.

[0054] The detection state information describes the state features of the object, including but not limited to geographic information, age, and region of interest images. Accordingly, the detection state information is represented in text, and / or numerical form, and / or image form.

[0055] Optionally, in response to a status detection request for an object, the computer device queries and retrieves the object's detection status information from a data storage system based on the object's identifier.

[0056] Optionally, the state features of the object can be collected from different omics systems. These state features are then converted into a unified format to obtain the object's detection state information, providing a coarse-grained representation of the object's macroscopic statistical characteristics. This can include collecting state features from different omics systems, such as geographic information, age information, and regions of interest (ROI) images. For geographic information, a mapping table between geographic names and numerical values ​​can be constructed to convert geographic information into corresponding numerical values. For age information, its numerical form is directly retained. For ROI images, convolutional neural networks (such as ResNet) are used to extract image features, and the extracted feature vectors are used as the numerical representation of the image. Then, the converted geographic values, age values, and image feature vectors are concatenated to obtain the object's detection state information, providing a coarse-grained representation of the object's macroscopic statistical characteristics.

[0057] Taking the object's state features, including geographic information, age information, and region of interest image, as an example, a numerical conversion model can be called to convert the geographic information, age information, and region of interest image information into state information in data format. Then, the converted state information is concatenated to obtain the object's detection state information.

[0058] Taking the object's state features, including geographic information, age information, and region of interest (ROI) image, as an example, a text conversion model can be called to convert the geographic information, age information, and ROI image information into text-format state information. Then, the converted state information is concatenated to obtain the object's detection state information.

[0059] Step S206: Input the abnormal sequence information and detection state information into the state prediction model to extract the abnormal sequence features of the abnormal sequence information and the detection state features of the detection state information, and obtain the predicted state information of the object based on the abnormal sequence features and detection state features.

[0060] The state prediction model is a deep learning model, such as the transformer model, which supports parallel computation on different types of input information and integrates the features of various input information to output predicted state information that matches the object. The number of predicted states can be one or multiple, and the content of the predicted state information includes each predicted state and its probability value.

[0061] Specifically, in this state prediction task, the transformer model mainly consists of an encoding layer and a decoding layer. The encoding layer includes a multi-head self-attention mechanism and a feed-forward network, used for feature extraction and encoding of the input anomalous sequence information and detection state information. The multi-head self-attention mechanism allows the model to focus on different parts of the input sequence in different representation subspaces, and its calculation process is as follows:

[0062] First, calculate the query, key, and value matrix: Q = XW Q K = XW K V = XW K

[0063] Where X is the input sequence, W Q W K and W V These are the weight matrices for the query, key, and value, respectively.

[0064] Then calculate the attention score:

[0065] Where d k denoted as the dimension of the key matrix.

[0066] The multi-head self-attention mechanism concatenates the outputs of multiple attention heads and performs a linear transformation: MultiHead(Q,K,V)=Concat(head1,…,head) h W O

[0067] Where h is the number of attention heads, W O This is the output weight matrix.

[0068] The feedforward neural network performs a nonlinear transformation on the output of the multi-head self-attention mechanism: FFN(x)=max(0,xW1+b1)W2+b2

[0069] Where W1, b1, W2, and b2 are the parameters of the feedforward neural network.

[0070] The decoding layer adds an encoder-decoder attention mechanism on top of the encoding layer. This mechanism integrates the output of the encoding layer with its own input to perform decoding, ultimately outputting predicted state information that matches the object.

[0071] In this embodiment, the state prediction model characterizes the complex correlation between multiple different types of state factors of an object and actual state information. The inputs to the state prediction model are: abnormal sequence information and detected state information of the object; the internal processing logic of the state prediction model is: by combining the individual analysis results of each type of input information and the correlation analysis results between different input information, the degree of matching between the object and multiple candidate state information is determined; the output of the state prediction model is: one or more predicted state information that match the object and the probability value of each predicted state information.

[0072] In one scenario, the state prediction model includes an encoding model and a decoding model. The encoding model is used for feature extraction, acquiring abnormal sequence features and detection state features. The decoding model is used to decode the extracted features to obtain the predicted state information. Specifically, the encoding model extracts abnormal sequence features from the abnormal sequence information and detection state features from the detection state information in parallel. The decoding model analyzes the abnormal sequence features and detection state features to obtain the predicted state information of the object.

[0073] In another scenario, the state prediction model includes an encoding model and a decoding model. The encoding model is used for feature extraction and feature fusion to obtain multimodal biometrics, while the decoding model is used to decode the fused multimodal biometrics to obtain the predicted state information. Specifically, the encoding model extracts anomalous sequence features from anomalous sequence information and detection state features from detection state information in parallel. The extracted anomalous sequence features and detection state features are then fused to obtain the object's multimodal biometrics. The decoding model analyzes these multimodal biometrics to obtain the object's predicted state information.

[0074] It is important to emphasize that abnormal sequence information includes abnormal location and abnormal type. Therefore, the abnormal sequence features extracted based on abnormal sequence information also include location features and type features. Furthermore, the multimodal biometrics obtained by fusing abnormal sequence features and detection state features include location features, type features, and detection state features. This allows the state prediction model to learn the influence of features on state information under each modality, the influence of the correlation of various features under different modalities on state information, and thus accurately predict the state information of the object.

[0075] In the above embodiments, the gene sequence information of the object is obtained, and the gene sequence information is compared with a reference gene sequence to obtain abnormal sequence information in the gene sequence information. The abnormal sequence information is used to characterize the abnormal location and abnormal type of abnormal gene units in the gene sequence information. The detection status information of the object is obtained. The abnormal sequence information and the detection status information are input into a status prediction model to extract the abnormal sequence features of the abnormal sequence information and the detection status features of the detection status information. Based on the abnormal sequence features and the detection status features, the predicted status information of the object is obtained. In this way, by using a preset status prediction model to extract, analyze, and predict features from the object's gene sequence information and detection status information, the interference of subjective factors is avoided to a certain extent, ensuring the objectivity of the status prediction information. Furthermore, the abnormal location and type of abnormal gene units in the gene sequence information are used to fine-grainedly characterize the microscopic genetic features of the object, while the detection state information is used to coarse-grainedly characterize the macroscopic statistical features of the object. During state prediction, the state prediction model analyzes this information, essentially taking into account information from different modalities and omics of the object. This allows the state prediction model to extract more comprehensive and richer multimodal biological features, ensuring the effectiveness of the feature analysis basis of the state prediction model and thus improving the authenticity and accuracy of the predicted state information obtained based on multimodal biological feature analysis. In addition, the information state prediction step is executed by the state prediction model deployed on the computer device. To avoid information recognition failures by the state prediction model, the object's gene sequence information is pre-processed before the state prediction model performs feature extraction and analysis steps. This ensures that the abnormal sequence information and detection state information meet the recognition requirements of the state prediction model, guaranteeing the normal operation of the state prediction model and fully utilizing the computing resources on the computer device.

[0076] In one embodiment, obtaining the gene sequence information of an object includes: obtaining the initial gene sequence of the object, the initial gene sequence including gene units at multiple sites; and performing quality control processing on the initial gene sequence to obtain gene sequence information.

[0077] The initial gene sequence can be a base sequence obtained through DNA sequencing or collected using sequencing equipment. The initial gene sequence comprises multiple gene units, each referring to a single base in the base sequence. A gene unit is a single base in a base sequence and is the basic unit that makes up a gene sequence; the initial gene sequence contains multiple such gene units.

[0078] After obtaining the initial gene sequence, a quality control process is performed to update the initial gene sequence and obtain the gene sequence information. Quality control of the initial gene sequence refers to the process of identifying and deleting low-quality gene data from the initial gene sequence, retaining high-quality gene units.

[0079] Optionally, quality control processing includes removing erroneous data. Erroneous data refers to data introduced during the testing of the initial gene sequence by the testing equipment due to hardware problems or testing methods. Specifically, according to a preset evaluation method, the quality of each gene unit in the initial gene sequence is quantitatively scored, and gene units with quantitative scores lower than the preset quality value are deleted from the initial gene sequence to obtain gene sequence information.

[0080] Specifically, the Phred quality score can be used as a preset evaluation method to quantify the quality of each gene unit in the initial gene sequence. The relationship between the base quality score Q and the base recognition error probability P is Q = -10log 10 P. Gene units with quantitative scores lower than a preset quality value (e.g., Q<20, indicating a base recognition error probability greater than 1%) are deleted from the initial gene sequence to obtain gene sequence information.

[0081] Optionally, the quality control process includes removing duplicate data, where duplicate data refers to the occurrence of the same or similar gene sequences multiple times in the initial gene sequence. Specifically, the initial gene sequence can be divided into multiple gene subsequences, the similarity between each pair of gene subsequences can be calculated, and either gene subsequence in the two gene subsequences with a similarity greater than a preset similarity threshold can be deleted to obtain the gene sequence information.

[0082] Specifically, the initial gene sequence can be divided into multiple gene subsequences, and the similarity between each pair of gene subsequences can be calculated using edit distance (Levenshtein distance). Edit distance refers to the minimum number of editing operations (insertion, deletion, replacement) required to transform one string into another. Similarity S can be defined as... Where d is the edit distance, and len(s1) and len(s2) are the lengths of the two gene subsequences, respectively. Then, in the two gene subsequences with a similarity greater than a preset similarity threshold (e.g., S>0.9), either gene subsequence is deleted to obtain the gene sequence information.

[0083] In the above embodiments, after obtaining the initial gene sequence of gene units including multiple sites, the initial gene sequence is subjected to quality control processing to obtain gene sequence information. This is equivalent to taking into account multiple factors, including test errors and redundancy of the initial gene sequence, in the process of obtaining the initial gene sequence, eliminating invalid data of the initial gene sequence, and screening out valid data. This improves the reliability and effectiveness of gene sequence information while reducing the computational burden of the subsequent state prediction model.

[0084] In one embodiment, comparing gene sequence information with a reference gene sequence to obtain abnormal sequence information in the gene sequence information includes:

[0085] The gene sequence information is aligned with the reference gene sequence to obtain candidate gene sequences that are located in the candidate site set. Each gene unit in the candidate gene sequence is compared with the corresponding gene unit in the reference gene sequence to identify normal and abnormal gene units in the candidate gene sequence. Based on the reference gene sequence, normal gene units, and abnormal gene units, the abnormal type sequence and abnormal position sequence of the object are constructed. According to the absolute gene position of the gene unit in the reference gene sequence, the abnormal type sequence and abnormal position sequence are aligned to obtain abnormal sequence information.

[0086] Sequence alignment refers to aligning two gene sequences according to the absolute gene position of a gene unit in a reference gene sequence. The candidate site set includes multiple sites, which can be preset sites for multiple gene units based on empirical values. Gene units that fall within the candidate site set are searched in the gene sequence information, and these gene units are sorted according to their absolute gene position to obtain the candidate gene sequence.

[0087] Each gene unit in the candidate gene sequence is compared with the gene unit at the corresponding site on the reference gene sequence. For any site on the candidate gene sequence, if the gene unit at that site is the same as the gene unit at that site in the reference gene sequence, then the gene unit at that site is determined to be a normal gene unit. If the gene unit at that site in the candidate gene sequence is of a different type than the gene unit at that site in the reference gene sequence, or if the candidate gene sequence has a gene unit at that site while the reference gene sequence does not, then a gene mutation has occurred at that site, and the gene unit at that site in the candidate gene sequence is an abnormal gene unit. If the candidate gene sequence does not have a gene unit at that site, then a gene mutation has occurred at that site, and the gene unit at that site on the reference gene sequence is an abnormal gene unit.

[0088] Based on the identification of normal and abnormal gene units, firstly, for each abnormal gene unit, the gene unit before and after the mutation of the abnormal gene unit is taken as a group of mutated gene combinations to obtain a two-dimensional vector corresponding to the abnormal gene unit; for each normal gene unit, the normal gene unit is copied twice to take as a group of unmutated gene combinations to obtain a two-dimensional vector corresponding to the normal gene unit.

[0089] Next, the absolute position of each abnormal gene unit is determined on the reference gene sequence. Based on the chromosomal position and absolute position of each abnormal gene unit, the two-dimensional vector corresponding to each abnormal gene unit is expanded to obtain the four-dimensional vector corresponding to each abnormal gene unit. The absolute position of each normal gene unit is determined on the reference gene sequence. Based on the chromosomal position and absolute position of each normal gene unit, the two-dimensional vector corresponding to each normal gene unit is expanded to obtain the four-dimensional vector corresponding to each normal gene unit.

[0090] Then, while keeping the vector length unchanged, the four-dimensional vectors corresponding to each abnormal gene unit and the four-dimensional vectors corresponding to each normal gene unit are stacked to construct a fourth-order matrix of the candidate gene sequence. Further, the two vectors representing each gene unit before and after mutation are used as the abnormality type sequence, and the vectors representing the chromosomal position and absolute position of each gene unit are used as the abnormal position sequence.

[0091] Taking a four-dimensional vector as an example, if the candidate gene sequence includes M mutated gene sequences and N non-mutated gene sequences, then the size of the fourth-order matrix of the candidate gene sequence is four rows and (M+N) columns. Each row represents one dimension of information about each gene unit in the candidate gene sequence, and each column represents information about a gene unit in four dimensions: gene type before mutation, gene type after mutation, mutated chromosome position, and absolute mutation position. In this case, the first and second rows of the fourth-order matrix are determined as abnormal type sequences, and the third and fourth rows are determined as abnormal position sequences.

[0092] Taking a four-dimensional vector as an example, if the candidate gene sequence includes M mutated gene sequences and N non-mutated gene sequences, then the size of the fourth-order matrix of the candidate gene sequence is four columns and (M+N) rows. Each column represents one dimension of information about each gene unit in the candidate gene sequence, and each row represents information about a gene unit in four dimensions: gene type before mutation, gene type after mutation, mutated chromosome position, and absolute mutation position. In this case, the first and second columns of the fourth-order matrix are determined as abnormal type sequences, and the third and fourth columns are determined as abnormal position sequences.

[0093] Finally, the abnormal type sequences and abnormal position sequences are sorted according to the absolute position of the gene units in the reference gene sequence from smallest to largest (or from largest to smallest) to obtain the aligned abnormal type sequences and abnormal position sequences, which serve as abnormal sequence information.

[0094] In the above embodiments, the gene sequence information and the reference gene sequence are aligned to facilitate the analysis of each gene unit in the gene sequence information. Then, candidate gene sequences are selected from the gene sequence information to form a candidate site set, which further narrows the analysis scope for subsequent gene unit variation analysis, improves the speed of gene abnormality analysis, and thus quickly and accurately identifies normal and abnormal gene units in the candidate gene sequences. Then, based on the two dimensions of abnormality type and abnormal location, the variation of each gene unit in the candidate gene sequence is characterized in a fine-grained manner to obtain abnormality type sequences and abnormal location sequences, enriching the abnormal sequence information content.

[0095] In one embodiment, the abnormal gene unit includes at least one of the following: a newly added gene unit, a deleted gene unit, and a replaced gene unit; the abnormal type sequence includes a gene reference sequence and a gene variant sequence; and the abnormal location sequence includes a chromosomal location sequence and a gene absolute location sequence.

[0096] Based on the reference gene sequence, normal gene units, and abnormal gene units, construct the abnormality type sequence and abnormality location sequence of the object, including:

[0097] Based on the absolute gene position of the gene unit in the reference gene sequence, the following steps are performed: The type identifier of each gene unit in the reference gene sequence corresponding to the candidate gene sequence and the gene addition identifier of each newly added gene unit are sorted and combined to obtain the gene reference sequence; The type identifiers of normal gene units, the type identifiers of replaced gene units, the type identifiers of newly added gene units, and the gene deletion identifiers of each deleted gene unit are sorted and combined to obtain the gene variant sequence; The chromosome position corresponding to each gene unit in the gene variant sequence and the chromosome position of each deleted gene unit are sorted and combined to obtain the chromosome position sequence; And, The absolute gene position of each gene unit in the gene variant sequence and the absolute gene position of each deleted gene unit are sorted and combined to obtain the absolute gene position sequence.

[0098] Specifically, "added gene unit" refers to the gene unit of the candidate gene sequence at a given locus that exists in the candidate gene sequence but not in the reference gene sequence. "Replaced gene unit" refers to the gene unit of the candidate gene sequence at a given locus that exists in both the candidate and reference gene sequences, but with different gene unit types. "Deleted gene unit" refers to the gene unit of the reference gene sequence at a given locus that does not exist in the candidate gene sequence but does exist in the reference gene sequence.

[0099] Taking abnormal gene units, including mutated abnormal gene units and adjacent normal gene units of abnormal gene units, as an example, please refer to Figure 3. Figure 3 is a schematic diagram of abnormal sequence information. The first two rows in Figure 3 are abnormal type sequences, and the first and second rows are gene reference sequences and gene mutation sequences, respectively. The last two rows in Figure 3 are abnormal position sequences, and the third and fourth rows are chromosome position sequences and gene absolute position sequences, respectively.

[0100] In Figure 3, the gene reference sequence includes five identifiers: "A", "T", "C", "G", and "+". "A", "T", "C", and "G" represent the base type of the gene unit, which is the type identifier of each gene unit in the reference gene sequence corresponding to the candidate gene sequence. "+" indicates the addition of a new gene unit. The gene variation sequence also includes five identifiers: "A", "T", "C", "G", and "-". "A", "T", "C", and "G" represent the base type of the gene unit, and "-" indicates gene deletion. The numerical value of the chromosome position indicates the chromosome number of the base. The numerical value of the gene absolute position sequence indicates the absolute position of the base on the reference genome.

[0101] Gene addition markers are used to indicate the addition of new gene units. In a gene reference sequence, a specific symbol (such as "+") is used to represent the addition of a gene unit. Gene deletion markers are used to indicate the deletion of gene units. In a gene mutation sequence, a specific symbol (such as "-") is used to represent the deletion of a gene unit.

[0102] As can be seen from the abnormal sequence information shown in Figure 3, at the gene position 20, the gene unit of the object mutated from G to C. The gene units at the gene positions 21-23 were deleted. The gene units at the gene positions 81-84 were added. The added gene units are “C”, “A”, “A”, and “A” bases respectively.

[0103] In the above embodiments, the addition of new gene units in the object is determined by identifying gene addition markers in the gene reference sequence; the deletion of gene units in the object is determined by identifying gene deletion markers in the gene mutation sequence; the base type of the object's gene units before and after mutation is determined by comparing the gene reference sequence and the gene mutation sequence; and the positions of added, deleted, and abnormal gene units in the reference genome sequence can be determined by using the chromosomal position sequence and the gene absolute position sequence.

[0104] In one embodiment, obtaining the detection status information of an object includes:

[0105] The system acquires the text state information, numerical state information, and image state information of the object; it performs data mapping on the text state information according to a preset text state transformation relationship to obtain the text feature information of the object; it performs vectorization processing on the image state information to obtain the image data information of the image state information, and summarizes the numerical state information and image data information to obtain the numerical feature information of the object; based on the text feature information and the numerical feature information, it obtains the detection state information.

[0106] For example, suppose the preset text state transition relationship is a dictionary. For gender information, the dictionary is {'male':1, 'female':0}; for the existence of a certain behavioral feature, the dictionary is {'exists':1, 'does not exist':0}. By mapping the text state information according to this dictionary, the text feature information of the object can be obtained.

[0107] When vectorizing image state information to obtain image state data, a pre-trained convolutional neural network (such as VGG16) can be used for feature extraction. The image is input into the VGG16 network, the last fully connected layer is removed, and the feature maps output from the intermediate layers are obtained. Then, global average pooling is performed on the feature maps to convert them into one-dimensional vectors, thus obtaining the image state data.

[0108] Textual status information includes object category information, such as the object's gender or whether a certain behavioral characteristic exists; numerical status information includes the object's life numerical information, such as the object's age or MSI value; image status information includes images of the object's region of interest, such as computed tomography (CT) images or magnetic resonance imaging (MRI) images.

[0109] For text state information, the text state transformation relationship can be a mapping relationship between text state and numerical value. According to the text state transformation relationship, the text state information is converted into matching numerical values ​​as text feature information. For example, using one-hot encoding, the gender male is converted to 1 and the gender female is converted to 0; using one-hot encoding, the presence of a certain behavioral feature is converted to 1 and the absence of a certain behavioral feature is converted to 0.

[0110] For numerical state information and image state information, the image features of the image state information are first extracted and then vectorized to obtain the feature values ​​of the image state information, which are used as image data information. Then, the numerical state information and image data information are concatenated in a preset order to obtain the numerical feature information of the object.

[0111] Following the order in which textual feature information precedes numerical feature information, the textual feature information and numerical feature information are summarized, that is, the textual feature information, numerical state information and image data information are concatenated to obtain the object detection state information.

[0112] It should be noted that both textual and numerical feature information are formatted uniformly and represented in numerical form. The purpose of this is to preserve the multi-dimensional feature information of the object while facilitating the identification and processing by the state prediction model, thereby improving the prediction speed of the state prediction model.

[0113] In the above embodiments, the text state information is mapped according to the preset text state transition relationship to obtain the text feature information of the object; the image state information is vectorized to obtain the image data information of the image state information; then the text feature information, numerical state information and image data information are summarized to obtain the detection state information. This is equivalent to unifying the form of text state information and image state information into numerical representation, and using objective data to comprehensively represent the state information of the object in multiple different dimensions, so as to improve the comprehensiveness and accuracy of the detection state information.

[0114] In one embodiment, as shown in Figure 4, the state prediction model includes an encoding network, a transformation network, and a decoding network. Abnormal sequence information and detection state information are input into the state prediction model to extract abnormal sequence features from the abnormal sequence information and detection state features from the detection state information. Based on the abnormal sequence features and detection state features, the predicted state information of the object is obtained, including:

[0115] The system extracts anomalous sequence features from the anomaly sequence information using an encoding network, and extracts detection state features from the detection state information. It then concatenates the anomalous sequence features and detection state features using a transformation network, and performs dimensionality upscaling on the concatenated features to obtain multimodal biometric features. Finally, it performs dimensionality downscaling on the multimodal biometric features using a decoding network to output the predicted state information of the object.

[0116] Dimensionality upscaling can be achieved using a Multilayer Perceptron (MLP). Assuming the concatenated feature vector is x, and the MLP contains two fully connected layers, the first fully connected layer has a weight matrix W1 and a bias vector b1, and the second fully connected layer has a weight matrix W2 and a bias vector b2. The dimensionality upscaling process is then: h1 = ReLU(xW1 + b1) h2 = h1W2 + b2

[0117] Where ReLU is the activation function, and h2 is the multimodal biological feature after dimensionality upgrade.

[0118] In the state prediction model, the state prediction process is as follows: Abnormal sequence information and detected state information are input into the encoding network to perform feature extraction steps for both types of information in parallel. Specifically, abnormal sequence features are extracted from the abnormal sequence information, and detected state features are extracted from the detected state information. Next, the extracted abnormal sequence features and detected state features are input into the transformation network to concatenate them, obtaining a fused feature. Further high-dimensional feature extraction is performed on the fused feature to obtain the object's multimodal biometric features. Finally, the multimodal biometric features are input into the decoding network to reduce the dimensionality of the high-dimensional multimodal biometric features, obtaining the probability of the object relative to multiple preset candidate state information. Then, based on the probability of the object relative to the multiple preset candidate state information, the predicted state information of the object is determined.

[0119] Optionally, multiple candidate state information and their probabilities can be aggregated to form the predicted state information of the object, or the candidate state information with the highest probability can be used as the predicted state information of the object.

[0120] In the above embodiments, anomaly sequence features of abnormal sequence information and detection state features of detection state information are extracted based on an encoding network. The abnormal sequence features and detection state features are concatenated based on a transformation network, and the concatenated features are then subjected to dimensionality upscaling to obtain multimodal biometric features. Finally, the multimodal biometric features are dimensionality reduced based on a decoding network to output the predicted state information of the object. This parallel encoding approach improves the model's feature extraction efficiency. Furthermore, considering the correlation between features of different modalities, the extracted features are concatenated, further enhancing the effectiveness of the multimodal biometric features. Finally, the concatenated features are subjected to dimensionality upscaling and downscaling to obtain accurate predicted state information.

[0121] In one embodiment, the coding network includes a sequence aberration coding layer; the aberration sequence information includes a gene reference sequence, a gene variant sequence, a chromosome position sequence, and a gene absolute position sequence; the aberration sequence features extracted based on the coding network include:

[0122] The sequence anomaly coding layer identifies gene reference sequences and gene variant sequences to determine the anomaly type of each gene unit in the anomalous sequence information. Based on the anomaly type of each gene unit, an anomaly type vector for each gene unit is generated, and these vectors are combined to obtain a type matrix for the anomalous sequence information. The sequence anomaly coding layer also identifies chromosome position sequences to determine the chromosome position of each gene unit in the anomalous sequence information. Based on the chromosome position of each gene unit, a position vector for each gene unit is generated, and these vectors are combined to obtain a chromosome position matrix for the anomalous sequence information. Finally, the type matrix, chromosome position matrix, and gene absolute position sequence are concatenated to obtain an anomaly feature matrix, which is then identified as the anomalous sequence features.

[0123] The input to the sequence anomaly coding layer is the gene reference sequence, gene variant sequence, chromosome position sequence, and gene absolute position sequence. The output is an anomaly feature matrix. The sequence anomaly coding layer can be a linear layer. The coding logic is to encode and transform multiple input sequences and use a unified representation form to describe the variation of different gene units.

[0124] The sequence anomaly coding layer identifies gene reference sequences and gene variant sequences, determining whether each gene unit in the anomalous sequence information has mutated. After identifying anomalous gene units, an anomalous type vector for each anomalous gene unit is determined from a pre-defined anomalous type vector according to the mutation type of the anomalous gene unit, and a normal type vector for each normal gene unit is determined according to the gene type of the normal gene unit. The anomalous type vector and the normal type vector have the same dimension. Next, the anomalous type vector of each anomalous gene unit and the normal type vector of each normal gene unit are combined to obtain the type matrix of the anomalous sequence information.

[0125] Chromosomal position sequences are identified by a sequence anomaly coding layer to determine the chromosomal position of each gene unit in the anomaly sequence information. The chromosomal position of each gene unit is one-hot encoded to obtain the position vector of each gene unit. The position vectors of each gene unit are combined to obtain the chromosomal position matrix of the anomaly sequence information.

[0126] The abnormal sequence information type matrix and chromosome position matrix are concatenated by the sequence abnormality coding layer. Then, the absolute gene position of each gene unit in the gene absolute position sequence is extracted and concatenated to the adjacent position of the chromosome position vector of the corresponding gene unit to obtain the abnormal feature matrix of the abnormal sequence information, that is, the abnormal sequence feature.

[0127] In this embodiment, a sequence anomaly coding layer is used to identify, process, and transform gene reference sequences and gene variant sequences to obtain a type matrix of anomalous sequence information. A sequence anomaly coding layer is also used to identify and transform chromosome position sequences to obtain a chromosome position matrix of anomalous sequence information. The type matrix, chromosome position matrix, and gene absolute position sequences are concatenated to obtain an anomaly feature matrix, which is then defined as the anomalous sequence features. Thus, the same format is used to represent the anomalous type and location of each gene unit in the anomalous sequence features, standardizing the expression of the variation of each gene unit. This facilitates the subsequent feature fusion of gene unit variation features and detection state features, improving the fusion speed of the transformation network in the state detection model and thereby shortening the state prediction time of the state detection model.

[0128] In one embodiment, based on the abnormality type of each gene unit, an abnormality type vector for each gene unit is generated, including:

[0129] For any gene unit, an initial type vector for the gene unit is constructed. The initial type vector includes newly added gene identifiers, deleted gene identifiers, reference gene identifiers of multiple types, and variant gene identifiers of multiple types. If the gene unit is a newly added gene unit, the values ​​of the newly added gene identifiers and the variant gene identifiers matching the type of the gene unit in the initial type vector are set to a first preset value, and the values ​​of other identifiers in the initial type vector are set to a second preset value to generate an abnormal type vector for the gene unit. If the gene unit is a deleted gene unit, the values ​​of the deleted gene identifiers and the reference gene identifiers matching the type of the gene unit in the initial type vector are set to a first preset value, and the values ​​of other identifiers in the initial type vector are set to a second preset value to generate an abnormal type vector for the gene unit. If the gene unit is a replaced gene unit, the historical gene type of the gene unit is determined, the values ​​of the reference gene identifiers matching the historical gene type and the variant gene identifiers matching the gene unit are set to a first preset value, and the values ​​of other identifiers in the initial type vector are set to a second preset value to generate an abnormal type vector for the gene unit.

[0130] The initial type vector of each gene unit has the same length and the same content, including one newly added gene identifier, one deleted gene identifier, four types of reference gene identifiers (A, G, C, T), and four types of variant gene identifiers (A, G, C, T). Furthermore, the length of the abnormality type vector of each gene unit is also the same; the difference between abnormality type vectors of different gene units lies in the specific numerical values ​​within the vectors.

[0131] When the gene unit is a newly added gene unit, the addition of the gene unit is characterized by setting the values ​​of the newly added gene identifier bits in the initial type vector and the values ​​of the identifier bits of the four types of variant gene identifier bits (A, G, C, T) that match the type of the newly added gene unit. In this embodiment, the values ​​of the newly added gene identifier bits and the values ​​of the variant gene identifier bits that match the type of the newly added gene unit are set to a first preset value, such as 1, and the values ​​of the other identifier bits in the initial type vector are set to a second preset value, such as 0, thereby achieving numerical encoding of the initial type vector and obtaining the abnormal type vector of the gene unit.

[0132] When the gene unit is a deleted gene unit, the deletion status of the gene unit is characterized by setting the value of the deleted gene identifier bit in the initial type vector and the values ​​of the identifier bits of the four types of reference gene identifier bits (A, G, C, T) that match the type of the deleted gene unit. In this embodiment, the values ​​of the deleted gene identifier bit and the values ​​of the reference gene identifier bits that match the type of the deleted gene unit are set to a first preset value, such as 1, and the values ​​of the other identifier bits in the initial type vector are set to a second preset value, such as 0, thereby realizing the numerical encoding of the initial type vector and obtaining the abnormal type vector of the gene unit.

[0133] When the gene unit is a replacement gene unit, the historical gene type of the gene unit is determined, that is, the gene unit type before replacement. The modification status of the gene unit is characterized by setting the values ​​of the reference gene identifiers (A, G, C, T) matching the historical gene type in the four types of reference gene identifiers in the initial type vector, and the values ​​of the variant gene identifiers (A, G, C, T) matching the replacement gene type in the four types of variant gene identifiers. In this embodiment, the values ​​of the reference gene identifiers matching the historical gene type and the variant gene identifiers matching the gene unit are set to a first preset value, such as 1. The values ​​of other identifiers in the initial type vector are set to a second preset value, such as 0, thereby achieving numerical encoding of the initial type vector and obtaining the abnormal type vector of the gene unit.

[0134] If the gene unit has not mutated, the values ​​of the newly added gene identifier bits in the initial type vector and the values ​​of the identifier bits in the four types of mutated gene identifier bits (A, G, C, T) that match the newly added gene unit type are all set to 0 to represent the case where the gene unit has not mutated.

[0135] Taking the abnormal sequence information shown in Figure 3, with the first preset value being 1 and the second preset value being 0 as an example, Figure 5 is a schematic diagram of the abnormal feature matrix output by the encoder. The abnormal feature matrix in Figure 5 includes 11 rows, each row representing a gene unit in Figure 3. The abnormal feature matrix in Figure 5 includes 35 columns. Columns 1-10 show the variation of the characteristic gene units, which are four types of reference gene markers (A, G, C, T), newly added gene unit markers, four types of variant gene markers (A, G, C, T), and deleted gene unit markers, respectively. Columns 11-34 show the chromosomal positions of the characteristic gene units, which are 22 autosomes and two sex chromosomes, respectively. Column 35 shows the absolute gene position of the characteristic gene unit.

[0136] The gene units represented in the first row of Figure 5 are the same as those represented in the first column of Figure 3, which are the gene units that have not undergone mutation. Therefore, in the first 10 columns of the first row, the value corresponding to each identifier is 0. The gene units in the sixth and eleventh rows of Figure 5 are the same as those represented in the sixth and eleventh columns of Figure 3, which are also the gene units that have not undergone mutation. The specific data settings are the same as above and will not be repeated here.

[0137] The gene units represented in the second row of Figure 5 are the same gene units represented in the second column of Figure 3, which are gene units that have undergone gene mutations, changing from G bases to C bases. In the first 10 columns of the second row, the reference gene marker corresponding to the G base (second column) and the variant gene marker corresponding to the C base (eighth column) have values ​​of 1, while the values ​​in the other columns are 0.

[0138] The gene units represented in the third row of Figure 5 are the same as those represented in the third column of Figure 3, which are the gene units where gene deletion has occurred. In the first 10 columns of the third row, the values ​​corresponding to the reference gene identifier (third column) and the deleted gene identifier (tenth column) of the deleted gene unit (C base) are 1, and the values ​​of the other columns are 0. The gene units in the fourth and fifth rows of Figure 5 are the same as those represented in the fourth and fifth columns of Figure 3, which are the gene units where gene deletion has occurred. The specific data settings are the same as above and will not be repeated here.

[0139] The gene units represented in the seventh row of Figure 5 are the same as those represented in the seventh column of Figure 3, which are the newly added gene units. In the first 10 columns of the seventh row, the value of the newly added gene identifier (the fifth column) and the value of the reference gene identifier (the eighth column) corresponding to the newly added gene unit (C base) are 1, and the values ​​of the other columns are 0. The gene units in the eighth to tenth rows of Figure 5 are the same as those represented in the eighth to tenth columns of Figure 3, which are the newly added gene units. The specific data settings are the same as above and will not be repeated here.

[0140] In the above embodiments, for any gene unit, an initial type vector of the gene unit is constructed. The unmutated state of the gene unit and the gene mutation state such as the addition, deletion, and replacement of gene units are represented by the values ​​of the newly added gene identifier, the deleted gene identifier, the reference gene identifier of multiple types, and the variant gene identifier of multiple types in the initial type vector, so as to standardize the mutation state of each gene unit.

[0141] In one embodiment, based on the chromosomal location of each gene unit, a position vector for each gene unit is generated, including:

[0142] For any gene unit, an initial position vector of the gene unit is constructed; the initial position vector includes multiple candidate chromosome identifiers; among the multiple candidate chromosome identifiers, the target chromosome identifier corresponding to the chromosome position of the gene unit is determined, and the value of the target chromosome identifier is set to a first preset value, and the values ​​of other chromosome identifiers in the initial position vector are set to a second preset value, so as to obtain the position vector of the gene unit.

[0143] It should be noted that regardless of whether the gene unit mutates, it corresponds to a chromosome location, and the number of chromosomes in the object is fixed at 24.

[0144] Continuing with the example of the abnormal sequence information shown in Figure 3, with a first preset value of 1 and a second preset value of 0, please refer to the abnormal feature matrix shown in Figure 5. The chromosome positions of the feature gene units in list 11-34 of Figure 5 are 22 autosomes and two sex chromosomes, respectively.

[0145] The gene units represented in the first row of Figure 5 are the same gene units represented in the first column of Figure 3, which are gene units with chromosome position 0. In the first row, columns 11-34, the candidate chromosome marker in column 11 is determined as the target chromosome marker, and the value of column 11 is set to 1, while the values ​​of columns 12-34 are set to 0.

[0146] The gene units represented in the second row of Figure 5 are the same as those represented in the second column of Figure 3, which are the gene units with chromosome position 1. Therefore, in columns 11-34 of the first row, the candidate chromosome marker in column 12 is identified as the target chromosome marker, and the value of column 12 is set to 1. The values ​​of columns 11 and 13-34 are set to 0. The gene units in the third to fifth rows of Figure 5 are the same as those represented in columns 3 to 5 of Figure 3. The specific data settings are the same as above and will not be repeated here.

[0147] The gene units represented in the sixth row of Figure 5 are the same as those represented in the sixth column of Figure 3, which are the gene units at chromosome position 2. Therefore, in columns 11-34 of the first row, the candidate chromosome marker in column 13 is identified as the target chromosome marker, and the value of column 13 is set to 1. The values ​​of columns 11-12 and 14-34 are set to 0. The gene units in rows 7 to 11 of Figure 5 are the same as those represented in columns 7 to 11 of Figure 3. The specific data settings are the same as above and will not be repeated here.

[0148] In this embodiment of the application, for any gene unit, an initial position vector of the gene unit is constructed, and the chromosome position of the gene unit is characterized by setting the values ​​of multiple candidate chromosome identifiers in the initial position vector, so as to standardize the characterization of the chromosome position of each gene unit.

[0149] In one embodiment, as shown in Figure 6, the encoding network includes a text encoding layer and a numerical encoding layer, extracting detection state features of the detection state information, including:

[0150] Text features are extracted from the detection state information based on the text encoding layer, and numerical features are extracted from the detection state information based on the numerical encoding layer; the detection state features are determined based on the text features and numerical features.

[0151] In Figure 6, the encoding network includes a text encoding layer, a numerical encoding layer, and a sequence anomaly encoding layer arranged in parallel. The text encoding layer includes a linear layer for extracting text features. In this embodiment, the input to the text encoding layer is the text feature information from the detection state information, and the output is the text features from the detection state information. The numerical encoding layer includes a linear layer for extracting numerical features. In this embodiment, the input to the numerical encoding layer is the numerical state information from the detection state information, and the output is the numerical features from the detection state information. The input to the sequence anomaly encoding layer is anomaly sequence information, and the output is an anomaly feature matrix. The one-hot encoding method for feature extraction is described in the aforementioned embodiment.

[0152] It should be noted that the text encoding layer, the numerical encoding layer, and the sequence anomaly encoding layer extract features in parallel. In this embodiment, the text features extracted by the text encoding layer and the numerical features extracted by the numerical encoding layer are combined to obtain the detection state features of the detection state information, and the anomaly feature matrix extracted by the sequence anomaly encoding layer is used as the anomaly sequence features.

[0153] In the above embodiments, a text encoding layer is used to extract text features from the detection state information, and a data encoding layer is used to extract numerical features from the detection state information. This is equivalent to extracting different types of features from the detection state information in parallel through different encoding layers, thereby improving the feature extraction speed of the detection state information and thus improving the efficiency of the state prediction model in predicting the state.

[0154] In one embodiment, as shown in Figure 7, the decoding network includes a pooling layer and a classification layer; based on the decoding network, dimensionality reduction processing is performed on multimodal biometrics to output the predicted state information of the object, including:

[0155] Based on the pooling layer, the multimodal biological features are dimensionality reduced to obtain the one-dimensional abnormal sequence features of the object. Based on the classification layer, the one-dimensional abnormal sequence features are further dimensionality reduced to predict the matching probability between the object and each candidate state information in the candidate state information set. The candidate state information whose matching probability is within the preset probability interval is determined as the predicted state information of the object.

[0156] The decoding network in Figure 7 includes a pooling layer and a classification layer connected in series. The pooling layer is connected to the transformation network and can be an average pooling layer. It is used to reduce the dimensionality of the multimodal biometric features output by the transformation network to obtain one-dimensional abnormal sequence features. The classification layer can include a linear layer, an activation layer, and a linear layer. It is used to further reduce the dimensionality of the one-dimensional abnormal sequence features, predict the matching probability between the object and each candidate state information based on the object's multimodal features, and use the candidate state information corresponding to one or more matching probabilities within a preset probability interval as the predicted state information.

[0157] Suppose the multimodal biometric features are represented by a two-dimensional matrix M with dimensions m×n. The average pooling layer divides matrix M into multiple non-overlapping regions. For each region, the average value of the elements within that region is calculated as the pooling result for that region. For example, using a k×k pooling window with a stride of s, each element v of the resulting one-dimensional anomaly sequence feature vector v is... i The calculation is as follows:

[0158] Where i is the index of the pooled vector. In this way, multimodal biological features are reduced to one-dimensional abnormal sequence features.

[0159] In one feasible scenario, the candidate state information corresponding to the maximum matching probability and the matching probability of the candidate state information can be used as the predicted state information of the object, or the candidate state information and the matching probability of each candidate state information can be directly summarized as the predicted state information.

[0160] In the above embodiments, the multimodal biological features are dimensionality reduced based on the pooling layer to fully integrate the multimodal biological features. Then, the one-dimensional abnormal sequence features obtained by the dimensionality reduction are further dimensionality reduced based on the classification layer to accurately predict the matching probability between the object and each candidate state information in the candidate state information set. The candidate state information whose matching probability is within the preset probability range is determined as the predicted state information of the object, providing a real and comprehensive basis for the object's state prediction results and improving the accuracy of the state prediction results.

[0161] In one embodiment, the training method for the state prediction model includes:

[0162] Obtain the state information dataset; the state information dataset includes historical abnormal sequence information, historical detection state information, and historical state information labels of multiple sample objects; predict the historical abnormal sequence information and historical detection state information using the state prediction model to be trained to obtain the sample predicted state information; train the state prediction model based on the loss value between the sample predicted state information and the historical state information labels to obtain the trained state prediction model.

[0163] The state information dataset is constructed based on historical data and includes actual anomaly sequence information, actual detection state information, and actual state information labels for multiple sample objects.

[0164] The state prediction model includes an initial encoding network, an initial transformation network, and an initial decoding network. The initial encoding network includes an initial sequence anomaly encoding layer, an initial text encoding layer, and an initial numerical encoding layer. The initial decoding network includes a pooling layer and an initial classification layer.

[0165] The state information dataset is input into the state prediction model. The state prediction model learns the logical relationship between the historical abnormal sequence information and historical detection state information of each sample object and the historical state information label of the sample object, thus obtaining the trained state prediction model.

[0166] Specifically, the state information dataset is input into the state prediction model, which is trained using the stochastic gradient descent algorithm. Assuming the parameters of the state prediction model are θ, for each sample object's historical anomaly sequence information X... abnormal and historical detection status information X detect The model outputs predicted state information. Where f is the model's mapping function. Define the loss function. To measure the predicted state information With historical status information label Y label Differences, such as those in the cross-entropy loss function, are observed by calculating the gradient of the loss function with respect to the parameter θ. And update parameters based on gradients Where α is the learning rate. This process is repeated until the loss function converges or the preset number of training rounds is reached, resulting in a trained state prediction model.

[0167] The training process of the state prediction model refers to the following: the state prediction model predicts the predicted state information of the sample based on historical abnormal sequence information and historical detection state information, inputs the predicted state information of the sample and the labels of historical state information into the loss function to obtain the model prediction accuracy, and optimizes the parameters of the state prediction model based on the model prediction accuracy until the loss value is less than a preset threshold, thus obtaining the trained state prediction model.

[0168] In the above embodiments, a state information dataset, constructed from historical anomaly sequence information, historical detection state information, and historical state information labels of multiple sample objects, is used to train the state prediction model, resulting in a trained state prediction model. This end-to-end training method reduces the difficulty of obtaining the state information dataset and provides a more diverse dataset for model training, thereby improving the generalization ability of the state prediction model. Furthermore, the trained state prediction model can be flexibly deployed on various mobile terminals according to actual application needs, further expanding the application scenarios of the state prediction model and thus improving the scenario applicability of the state prediction method.

[0169] In one embodiment, the state prediction model is a genetic disease prediction model; the method further includes:

[0170] The process involves obtaining the subject's gene reference sequence, gene variant sequence, chromosomal position sequence, and gene absolute position sequence; acquiring the subject's clinical and imaging information; and inputting the gene reference sequence, gene variant sequence, chromosomal position sequence, gene absolute position sequence, clinical and imaging information into a gene disease prediction model to extract the subject's multimodal biological characteristics. These multimodal biological characteristics are then analyzed to predict the subject's gene disease type.

[0171] Based on the biological tissue samples of the object, gene sequencing tools are used to sequence the biological tissue samples to obtain the initial gene sequence of the object. Then, the initial gene sequence is compared with the reference gene sequence to construct the object's gene reference sequence, gene variation sequence, chromosome position sequence and gene absolute position sequence.

[0172] Raw clinical omics information, such as gender, smoking status, and age, was collected from the subjects through questionnaires. The raw clinical omics information, including gender and smoking status, was digitized, and the digitized clinical omics information and age were used as the subjects' clinical information. Raw radiomics information, such as MRI images, was collected from the subjects through medical scanning. The raw radiomics information was then subjected to feature extraction and vectorization to obtain the imaging information.

[0173] Clinical information, imaging information, gene reference sequence, gene variant sequence, chromosome position sequence, and gene absolute position sequence are sequentially input into the gene disease prediction model. The gene disease prediction model extracts state features from the clinical and imaging information, and extracts sequence features from the gene reference sequence, gene variant sequence, chromosome position sequence, and gene absolute position sequence. Then, the state features and sequence features are fused to obtain the multimodal features of the object. Finally, the multimodal biofeedback is decoded and classified to obtain the probability of the object with each gene disease type, and the gene disease type corresponding to the highest probability is output as the gene disease type of the object.

[0174] For clinical information, a linear layer is used for feature extraction. Assume the clinical information is a vector x. clinical The weight matrix of the linear layer is W clinical The bias vector is b clinical The extracted clinical features are h. clinical =W clinical x clinical +b clinical For image information, convolutional neural networks (such as Inception) are used for feature extraction. The image is input into the network, and the feature vector h output by the network is obtained. image For gene reference sequences, gene variant sequences, chromosomal position sequences, and gene absolute position sequences, encoding transformation is performed through a sequence anomaly coding layer. The specific encoding method is as described above, resulting in an anomaly feature matrix as the sequence feature h. sequence Then, the clinical features h clinical Image features h image and sequence features h sequence By concatenating the features, we obtain the multimodal features of the object, hmulti-modal = [h...]. clinical h image h sequence Finally, a decoding network (such as fully connected layers and activation functions) is used to decode and classify multimodal biometric features, obtain the probability of an object being associated with each gene disease type, and output the gene disease type corresponding to the highest probability as the object's gene disease type.

[0175] In the above embodiments, the object's gene reference sequence, gene variant sequence, chromosome position sequence, gene absolute position sequence, clinical information, and imaging information are input into the gene disease prediction model to instruct the gene disease prediction model to predict the object's gene disease type based on the object's multimodal biological characteristics. This comprehensive consideration of multiple factors of imaging gene diseases results in a more accurate prediction of the gene disease type.

[0176] In one embodiment, the state prediction method can be applied to the scenario of gene disease prediction, as shown in Figure 8, and includes the following steps:

[0177] Step S801: Obtain the sequencing data of the object.

[0178] When the initial gene sequence obtained through second- or third-generation gene sequencing is acquired, the initial gene sequence is preprocessed, including but not limited to quality control and deduplication, to obtain the sequencing data of the target.

[0179] Step S802, sequence alignment.

[0180] The sequencing data sequence was aligned to the reference gene sequence.

[0181] Step S803: Perform mutation detection on the candidate gene sequences in the candidate site set to obtain abnormal sequence information.

[0182] Optionally, a gene mutation detection tool can be used to detect the two sets of gene sequences in the sequence alignment, output the mutation detection results, and then construct a gene reference sequence, gene mutation sequence, chromosome position sequence and gene absolute position sequence based on the mutation detection results as abnormal sequence information.

[0183] Step S804: Obtain the clinical omics information of the subjects.

[0184] Clinical omics information includes categorical and numerical features. Categorical features include gender and smoking status, while numerical features include age and microsatellite instability score (MSI). In practical applications, one-hot encoding is used to digitize the categorical features in the clinical omics information. The digitized categorical features then replace the categorical features in the clinical omics information to obtain the numerical form of the clinical omics information.

[0185] Step S805: Obtain the image omics information of the object.

[0186] Based on the obtained region of interest (ROI) image of the object, the feature values ​​of the ROI image are further obtained as the object's image omics information.

[0187] Step S806: Input the abnormal sequence information, clinical omics information and radiomics information into the gene disease prediction model, and output the probability of the gene disease type of the object.

[0188] Among them, the gene disease prediction model can be a transformer model, specifically a multimodal gene disease prediction model trained using a large number of historical datasets of sample objects before the gene disease prediction model is actually put into use.

[0189] In the above embodiments, the gene disease prediction model predicts the gene disease of the subject based on multiple information such as the subject's clinical information, gene mutations, and radiomics. This is equivalent to combining the subject's macroscopic tabular features with microscopic fine-grained gene features, enabling the multimodal gene disease discrimination model to learn the degree of influence between various complex factors and gene diseases, so as to make more accurate gene disease prediction results.

[0190] In one embodiment, as shown in Figure 9, the state prediction method includes the following steps:

[0191] Step S901: Obtain the text status information, numerical status information, and image status information of the object.

[0192] Step S902: According to the preset text state transition relationship, perform data mapping on the text state information to obtain the text feature information of the object.

[0193] Step S903: The image state information is vectorized to obtain the image data information of the image state information, and the numerical state information and the image data information are summarized to obtain the numerical feature information of the object.

[0194] Step S904: Obtain the initial gene sequence of the object, perform quality control processing on the initial gene sequence, and obtain gene sequence information.

[0195] Step S905: Compare the candidate gene sequences in the candidate site set with the reference gene sequences in the gene sequence information to construct the gene reference sequence, gene variation sequence, chromosome position sequence, and gene absolute position sequence.

[0196] Step S906: Extract text features based on the text encoding layer.

[0197] Step S907: Extract numerical features based on the numerical coding layer to obtain numerical feature information.

[0198] Step S908: Extract abnormal sequence features based on the sequence anomaly coding layer.

[0199] Step S909: Based on the transformation network, text features, numerical features and abnormal sequence features are concatenated, and the concatenated features are upgraded to obtain multimodal biological features.

[0200] Step S910: Dimensionality reduction of multimodal biological features based on pooling layer.

[0201] Step S911: Based on the classification layer, predict the dimensionality-reduced multimodal biological features and output the predicted state information of the object.

[0202] It should be noted that steps S901-S903 above are steps for obtaining detection state information, and steps S904-S905 above are steps for obtaining gene sequence information. In practical applications, the processes of obtaining detection state information and gene sequence information can be executed in parallel or sequentially. Steps S906-S911 above are flowcharts of the state prediction of the preset state prediction model.

[0203] In this embodiment, the method of extracting, analyzing, and predicting the gene sequence information and detection state information of the object through various network layers in the preset state prediction model can, to a certain extent, avoid the interference of subjective factors and ensure the objectivity of the prediction results of the state prediction model. Furthermore, when performing state prediction, both the macroscopic tabular features and the microscopic fine-grained gene features of the object are comprehensively considered to fully and accurately characterize the actual abnormal state of the object, thereby improving the authenticity and accuracy of the predicted state information obtained from multimodal biometric analysis.

[0204] In a specific embodiment, the state prediction method can be applied to a pan-genetic disease discrimination scenario. In this scenario, the state prediction model is a gene type prediction model for multiple types of diseases, and the predicted state information is a single gene disease type.

[0205] First, the computer device acquires multimodal information about the object, namely clinical omics information, radiomics information corresponding to multiple parts of the object, and gene sequences corresponding to multiple parts of the object. The acquired information is then processed to obtain various types of model input information, so that the disease type prediction model can identify and process it.

[0206] The clinical omics information includes gender, information on the presence or absence of smoking behavior, age, and microsatellite instability score. The clinical omics information is processed as follows: gender is numerically represented by 1 for male and 0 for female; information on the presence or absence of smoking behavior is also numerically represented by 1 for the presence of smoking behavior and 0 for the absence of smoking behavior.

[0207] The image-omics information includes, but is not limited to, CT scan images of the lungs and brain of the subject. The image-omics information is processed by extracting image features from the lung CT scan images and vectorizing them to obtain a first feature vector, and extracting image features from the brain CT scan images and vectorizing them to obtain a second feature vector.

[0208] The gene sequences corresponding to multiple parts of the object are processed. Specifically, candidate gene sequences corresponding to multiple parts are extracted from the object's whole genome sequence using a mutation detection tool, and mutation detection is performed on the candidate gene sequences to obtain mutation detection results. Then, the mutation detection results are parsed according to preset transformation conditions to construct gene reference sequences, gene variant sequences, chromosomal position sequences, and gene absolute position sequences. It should be noted that the gene reference sequences, gene variant sequences, chromosomal position sequences, and gene absolute position sequences have the same sequence length. In this embodiment, the above sequences can be summarized into a fourth-order gene sequence matrix.

[0209] Then, the numerically processed gender information, information indicating the presence of smoking behavior, age information, microsatellite instability score information, first feature vector, second feature vector, and gene sequence matrix are input into the gene type prediction model to obtain the probability of the object and multiple gene disease types output by the gene type prediction model. Finally, the gene disease type corresponding to the highest probability is taken as the gene disease type of the object.

[0210] It should be noted that clinical omics information and radiomics information can be selected selectively. For example, using clinical omics information and gene sequences for gene disease prediction, continuing with the example of male gender (mapping value 1), no smoking behavior (mapping value 0), age 60, and MSI score 0.5, please refer to Figure 10. Input the gene sequence matrix of "1", "0", "60", "0.5" and the fourth order into the preset gene disease prediction model. In the schematic diagram of the gene disease type prediction model architecture shown in Figure 10, the gene disease type prediction model performs the following encoding operations in parallel: extracting features corresponding to gender and the presence or absence of smoking behavior based on the text encoding layer, extracting features of age and MSI score based on the numerical encoding layer, and extracting features of the gene sequence matrix based on the sequence anomaly encoding layer. Then, the features extracted from the above three encoding layers are fused through the core network (i.e., the transformation network in the aforementioned embodiment), and the fused multimodal features are extracted. The core network includes multiple transformer layers. Then, the multimodal features are reduced in dimensionality by the average pooling layer in the decoding network. The matching probability between the reduced multimodal features and multiple gene disease types is predicted by the classification layer, and the gene disease type corresponding to the maximum matching probability is output as the predicted gene disease type.

[0211] In the aforementioned pan-genetic disease identification scenario, the state prediction method provided in this application embodiment comprehensively considers the clinical omics information of the object, the image omics information corresponding to multiple parts of the object, and multimodal biological characteristics including gene sequences, to accurately locate the potential genetic disease information corresponding to the object's current state, providing a reliable reference for subsequent disease prevention.

[0212] In a specific embodiment, the state prediction method can be applied to a specific biological tissue, such as the identification of gene diseases in the lungs. In this scenario, the state prediction model is a prediction model for the type of gene disease in the lungs, and the predicted state information is a specific lung gene disease. A computer device acquires the subject's gender, smoking behavior, age, symptom duration, lung images, and candidate gene sequences matching the lung disease. Next, gender is mapped, with 1 representing male and 0 representing female; smoking behavior is mapped, with 1 representing the presence of smoking behavior and 0 representing the absence of smoking behavior; the unit for symptom duration is unified to days; lung image feature vectors are extracted; and based on the candidate gene sequences, lung gene reference sequences, lung gene variant sequences, lung gene chromosomal position sequences, and lung gene absolute position sequences are constructed. Then, the mapped gender data, mapped smoking behavior feature data, age, symptom duration, lung image feature vector, gene reference sequence, lung gene mutation sequence, lung gene chromosomal position sequence, and lung gene absolute position sequence are input into the gene disease type prediction model. The model's category coding layer encodes the gender data and smoking behavior feature data to obtain category features, while the model's numerical coding layer encodes the age, symptom duration, and lung image feature vector to obtain numerical features. Finally, the model's abnormal gene coding layer encodes the gene reference sequence, lung gene mutation sequence, and lung... The chromosome position sequence of lung genes and the absolute position sequence of lung genes are encoded to obtain variation features. The core network of the model is used to concatenate the categorical features, numerical features and variation features to further extract multimodal features of the object related to the lungs. The multimodal features are reduced in dimensionality by the average pooling layer of the model. The classification module of the model is used to perform correlation matching between the dimensionality-reduced features and the features of multiple lung gene diseases to obtain the matching degree between the object and each lung gene disease. The lung gene disease with the highest matching degree is identified as the predicted lung gene disease of the object, providing a reliable reference for the subsequent formulation of disease treatment plans for the object.

[0213] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0214] Based on the same inventive concept, this application also provides a state prediction apparatus for implementing the state prediction method described above. The solution provided by this apparatus is similar to the implementation scheme described in the above method; therefore, the specific limitations in one or more state prediction apparatus embodiments provided below can be found in the limitations of the state prediction method described above, and will not be repeated here.

[0215] In one embodiment, as shown in FIG11, a state prediction device is provided, comprising: a sequence acquisition module 1101, a state acquisition module 1102, and a state analysis module 1103, wherein:

[0216] The sequence acquisition module 1101 is used to acquire the gene sequence information of the object, compare the gene sequence information with the reference gene sequence, and acquire abnormal sequence information in the gene sequence information; the abnormal sequence information includes the abnormal location and abnormal type of the gene unit;

[0217] The status acquisition module 1102 is used to acquire the detection status information of the object;

[0218] The state analysis module 1103 is used to input abnormal sequence information and detection state information into the state prediction model to extract abnormal sequence features of abnormal sequence information and detection state features of detection state information, and obtain the predicted state information of the object based on the abnormal sequence features and detection state features.

[0219] In one embodiment, the sequence acquisition module 1101 includes:

[0220] The initial sequence acquisition unit is used to acquire the initial gene sequence of the object. The initial gene sequence includes gene units at multiple sites.

[0221] The sequence quality control processing unit is used to perform quality control processing on the initial gene sequence to obtain gene sequence information.

[0222] In one embodiment, the sequence acquisition module 1101 includes:

[0223] The first sequence alignment unit is used to align the gene sequence information with the reference gene sequence to obtain the candidate gene sequence that is located in the candidate site set.

[0224] The gene unit comparison unit is used to compare each gene unit on the candidate gene sequence with the corresponding gene unit on the reference gene sequence to identify normal and abnormal gene units in the candidate gene sequence.

[0225] An abnormal sequence construction unit is used to construct the abnormal type sequence and abnormal location sequence of an object based on a reference gene sequence, a normal gene unit, and an abnormal gene unit.

[0226] The second sequence alignment unit is used to align the abnormal type sequence and the abnormal position sequence according to the absolute gene position of the gene unit in the reference gene sequence to obtain abnormal sequence information.

[0227] In one embodiment, the abnormal gene unit includes at least one of newly added gene units, deleted gene units, and replaced gene units; the abnormal type sequence includes a gene reference sequence and a gene mutation sequence; the abnormal position sequence includes a chromosome position sequence and a gene absolute position sequence; the abnormal sequence construction unit is used to perform the following steps according to the gene absolute position of the gene unit in the reference gene sequence: sorting and combining the type identifier of each gene unit in the reference gene sequence corresponding to the candidate gene sequence and the gene addition identifier of each newly added gene unit to obtain a gene reference sequence; sorting and combining the type identifier of the normal gene unit, the type identifier of the replaced gene unit, the type identifier of the newly added gene unit, and the gene deletion identifier of each deleted gene unit to obtain a gene mutation sequence; sorting and combining the chromosome position corresponding to each gene unit in the gene mutation sequence and the chromosome position of each deleted gene unit to obtain a chromosome position sequence; and sorting and combining the gene absolute position of each gene unit in the gene mutation sequence and the gene absolute position of each deleted gene unit to obtain a gene absolute position sequence.

[0228] In one embodiment, the status acquisition module 1102 includes:

[0229] The information acquisition unit is used to acquire the text status information, numerical status information and image status information of the object;

[0230] The information mapping unit is used to map text state information according to a preset text state transition relationship to obtain the text feature information of the object.

[0231] The vectorization processing unit is used to perform vectorization processing on the image state information to obtain the image data information of the image state information, and to summarize the numerical state information and the image data information to obtain the numerical feature information of the object.

[0232] The information aggregation unit is used to obtain detection status information based on textual feature information and numerical feature information.

[0233] In one embodiment, the state prediction model includes an encoding network, a transformation network, and a decoding network; the state analysis module 1103 includes:

[0234] The feature encoding unit is used to extract abnormal sequence features of abnormal sequence information based on the encoding network, and to extract detection state features of detection state information.

[0235] The feature splicing unit is used to splice abnormal sequence features and detection state features based on the transformation network, and to perform dimensionality upscaling on the spliced ​​features to obtain multimodal biological features.

[0236] The feature decoding unit is used to perform dimensionality reduction processing on multimodal biofeatures based on the decoding network and output the predicted state information of the object.

[0237] In one embodiment, the coding network includes a sequence anomaly coding layer; the anomalous sequence information includes a gene reference sequence, a gene variant sequence, a chromosome position sequence, and a gene absolute position sequence; a feature coding unit is further configured to identify the gene reference sequence and gene variant sequence through the sequence anomaly coding layer, determine the anomaly type of each gene unit in the anomalous sequence information, generate an anomaly type vector for each gene unit based on the anomaly type of each gene unit, and combine the anomaly type vectors of each gene unit to obtain a type matrix of the anomalous sequence information; identify the chromosome position sequence through the sequence anomaly coding layer, determine the chromosome position of each gene unit in the anomalous sequence information, generate a position vector for each gene unit based on the chromosome position of each gene unit, and combine the position vectors of each gene unit to obtain a chromosome position matrix of the anomalous sequence information; and concatenate the type matrix, chromosome position matrix, and gene absolute position sequence to obtain an anomalous feature matrix, which is then identified as the anomalous sequence feature.

[0238] In one embodiment, the feature encoding unit is further configured to construct an initial type vector for any gene unit; the initial type vector includes newly added gene identifiers, deleted gene identifiers, reference gene identifiers of multiple types, and variant gene identifiers of multiple types; if the gene unit is a newly added gene unit, the values ​​of the newly added gene identifiers and the values ​​of the variant gene identifiers matching the type of the gene unit in the initial type vector are set to a first preset value, and the values ​​of other identifiers in the initial type vector are set to a second preset value, thereby generating an abnormal type vector for the gene unit; if the gene unit is a deleted gene unit, the values ​​of the deleted gene identifiers and the values ​​of the reference gene identifiers matching the type of the gene unit in the initial type vector are set to a first preset value, and the values ​​of other identifiers in the initial type vector are set to a second preset value, thereby generating an abnormal type vector for the gene unit; if the gene unit is a replaced gene unit, the historical gene type of the gene unit is determined, the values ​​of the reference gene identifiers matching the historical gene type and the values ​​of the variant gene identifiers matching the gene unit are set to a first preset value, and the values ​​of other identifiers in the initial type vector are set to a second preset value, thereby generating an abnormal type vector for the gene unit.

[0239] In one embodiment, the feature encoding unit is further configured to construct an initial position vector for any gene unit; the initial position vector includes multiple candidate chromosome identifiers; among the multiple candidate chromosome identifiers, a target chromosome identifier corresponding to the chromosome position of the gene unit is determined, and the value of the target chromosome identifier is set to a first preset value, and the values ​​of other chromosome identifiers in the initial position vector are set to a second preset value, thereby obtaining the position vector of the gene unit.

[0240] In one embodiment, the encoding network includes a text encoding layer and a numerical encoding layer; the feature encoding unit is further configured to extract text features from the detection state information based on the text encoding layer, and extract numerical features from the detection state information based on the numerical encoding layer; and determine the detection state features based on the text features and the numerical features.

[0241] In one embodiment, the decoding network includes a pooling layer and a classification layer; a feature decoding unit is used to perform dimensionality reduction on multimodal biological features based on the pooling layer to obtain one-dimensional abnormal sequence features of the object; based on the classification layer, the one-dimensional abnormal sequence features are reduced in dimensionality to predict the matching probability between the object and each candidate state information in the candidate state information set, and the candidate state information whose matching probability is within a preset probability interval is determined as the predicted state information of the object.

[0242] In one embodiment, the state prediction device includes:

[0243] The dataset acquisition module is used to acquire the status information dataset; the status information dataset includes historical anomaly sequence information, historical detection status information, and historical status information labels for multiple sample objects;

[0244] The model training module is used to predict historical abnormal sequence information and historical detection state information through the state prediction model to obtain sample predicted state information. Based on the loss value of the sample predicted state information and the historical state information labels, the state prediction model is trained to obtain the trained state prediction model.

[0245] In one embodiment, the state prediction device includes:

[0246] The genomics information acquisition module is used to acquire the object's gene reference sequence, gene variation sequence, chromosome position sequence, and gene absolute position sequence.

[0247] The radiomics information acquisition module is used to acquire the clinical and imaging information of the subjects;

[0248] The disease type prediction module is used to input gene reference sequences, gene variant sequences, chromosome position sequences, gene absolute position sequences, clinical information, and imaging information into the gene disease prediction model to extract the multimodal biological characteristics of the object, analyze the multimodal biological characteristics, and predict the gene disease type of the object.

[0249] The aforementioned state prediction device, through a preset state prediction model, extracts, analyzes, and predicts features from the object's gene sequence information and detection state information. This approach can, to a certain extent, avoid interference from subjective factors and ensure the objectivity of the predicted state information. Furthermore, during state prediction, it comprehensively considers multimodal information including the object's detection state information, the abnormal location of abnormal gene units in the gene sequence, and the type of abnormality. This allows for a comprehensive and accurate characterization of the object's actual abnormal state characteristics, making the multimodal biofeedback extracted by the state prediction model more closely match the object's actual abnormal state. This, in turn, improves the authenticity and accuracy of the predicted state information obtained from the multimodal biofeedback analysis.

[0250] Each module in the aforementioned state prediction device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.

[0251] In one embodiment, a computer device, which may be a server, is provided, and its internal structure is shown in Figure 12. The computer device includes a processor, memory, input / output interfaces (I / O), and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is connected to the system bus via the I / O interfaces. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The database stores state prediction data. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. When the computer program is executed by the processor, it implements a state prediction method.

[0252] In one embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram is shown in Figure 13. The computer device includes a processor, memory, input / output interface, communication interface, display unit, and input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interface. The processor of the computer device provides computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The input / output interface of the computer device is used for exchanging information between the processor and external devices. The communication interface of the computer device is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a state prediction method. The display unit of the computer device is used to form a visually visible image. It can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.

[0253] Those skilled in the art will understand that the structures shown in Figures 12 and 13 are merely block diagrams of some structures related to the present application and do not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than shown in the figures, or combine certain components, or have different component arrangements.

[0254] In one embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.

[0255] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in the above method embodiments.

[0256] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0257] In summary, this application provides a state prediction method, apparatus, device, computer-readable storage medium, and computer program product. The computer device acquires the gene sequence information of an object, compares it with a reference gene sequence to obtain abnormal sequence information characterizing the abnormal location and type of abnormal gene units, and simultaneously acquires the object's detection state information. Both are input into a state prediction model to extract features, ultimately obtaining predicted state information. From a technical perspective, this method of combining the object's microscopic gene features with macroscopic detection state information greatly enriches the model's input information dimensions. Gene sequence information contains the object's most essential genetic characteristics, while detection state information reflects the object's actual state at the macroscopic level. By comprehensively analyzing these two different levels of information, the model can more comprehensively and deeply uncover the potential patterns of the object's state. During data processing, the potential bias and limitations of a single information source are avoided, enabling the model to accurately evaluate the object's state from multiple perspectives, thereby significantly improving the accuracy and reliability of the predicted state information. Furthermore, preprocessing the gene sequence information to meet the model's recognition requirements ensures the model can operate normally and efficiently, improving the utilization of computer equipment's computing resources and reducing unnecessary computational overhead.

[0258] Furthermore, when acquiring the target gene sequence information, an initial gene sequence containing gene units at multiple loci is first obtained, followed by quality control processing to remove low-quality and duplicate data. Low-quality gene data is often introduced due to factors such as hardware errors in the testing equipment, instability in the testing environment, or non-standard testing methods during the sequencing process. If this erroneous data enters the subsequent analysis process, it will interfere with the model's correct extraction of gene sequence features, reducing the model's accuracy. Removing duplicate data can effectively reduce data redundancy and complexity. In terms of data storage, this reduces the storage space occupied; in terms of model computation, it alleviates the burden of data processing on the model, improves the model's computational efficiency, and enables the model to converge to the optimal solution more quickly, thereby improving the overall analysis speed and efficiency.

[0259] Furthermore, when comparing gene sequence information with reference gene sequences to obtain abnormal sequence information, sequence alignment is first performed to obtain candidate gene sequences. Then, normal and abnormal gene units are identified, and abnormality type and location sequences are constructed. Finally, sequence alignment is performed again. Sequence alignment is a crucial step in gene analysis, providing a foundation for subsequent detailed analysis of gene units. By aligning gene sequences with reference gene sequences, potentially abnormal regions can be quickly located, and candidate gene sequences can be obtained, thereby narrowing the scope of analysis and reducing unnecessary computation. After identifying normal and abnormal gene units, abnormality type and location sequences are constructed. This refined information construction method accurately records the variation of gene units, providing rich and detailed information for subsequent model analysis. The second sequence alignment further verifies and refines the abnormal sequence information, ensuring the accuracy and reliability of the information, thereby improving the model's accuracy in analyzing gene abnormalities.

[0260] Furthermore, when the abnormal gene unit includes at least one of the added, deleted, and replaced gene units, the abnormal type sequence includes gene reference sequences and gene variant sequences, and the abnormal position sequence includes chromosomal position sequences and gene absolute position sequences, these sequences are constructed through specific sorting combinations. Specifically, identifying gene addition markers in the gene reference sequence and gene deletion markers in the gene variant sequence can accurately determine the addition and deletion of gene units; comparing the gene reference sequence and the gene variant sequence can clarify the base types before and after the gene unit mutation; and the chromosomal position sequence and gene absolute position sequence can accurately determine the position of the gene unit in the reference genome sequence. This detailed and accurate information construction method enables the model to conduct in-depth analysis of gene variation from multiple dimensions, providing a more accurate basis for state prediction. In the field of gene research, accurately grasping the type and location of gene variations is of great significance for understanding the pathogenesis of diseases and developing personalized treatment plans.

[0261] Furthermore, when acquiring the detection state information of an object, textual, numerical, and image state information are obtained. The textual information is data-mapped, and the image information is vectorized. Finally, the detection state information is summarized. Different forms of state information have different data formats and characteristics in their original state, which poses a significant challenge to the model's processing. By mapping the textual information to numerical form and vectorizing the image information, different forms of information are unified into numerical representations, ensuring a consistent data format and facilitating recognition and processing by the state prediction model. This unified data format not only improves the model's processing efficiency but also avoids information loss or inaccuracies caused by differences in data formats, enabling the model to acquire object state information more comprehensively and accurately, thereby improving the comprehensiveness and accuracy of the detection state information.

[0262] Furthermore, when the state prediction model includes an encoding network, a transformation network, and a decoding network, the encoding network extracts features from the abnormal sequence and the detected state in parallel. The transformation network concatenates and upscales these features to obtain multimodal biometrics, and the decoding network performs dimensionality reduction to output the predicted state information. The parallel processing of the encoding network fully utilizes the multi-core computing power of the computer, enabling simultaneous feature extraction from both the abnormal sequence and the detected state information, significantly improving the efficiency of feature extraction. When processing large-scale data, parallel computing can significantly shorten processing time and improve the system's response speed. The transformation network concatenates and upscales features from different sources, enabling the discovery of potential relationships between different modal features, making the multimodal biometrics more representative and effective. Dimensionality upscaling maps low-dimensional features to a high-dimensional space, increasing the expressive power of the features and allowing the model to better capture complex patterns in the data. The dimensionality reduction processing of the decoding network converts high-dimensional multimodal biometrics into low-dimensional predicted state information, removing redundant information, retaining key features, and improving the accuracy and interpretability of the prediction results.

[0263] Furthermore, when processing anomalous sequence information, the sequence anomaly coding layer in the coding network determines the anomalous type of gene unit by identifying gene references and variant sequences, generating anomaly type vectors and combining them into a type matrix. It also identifies chromosome position sequences, generates position vectors, and combines them into a chromosome position matrix. Finally, the type matrix, chromosome position matrix, and gene absolute position sequences are concatenated to obtain an anomaly feature matrix as the anomalous sequence features. This processing method structurally represents the anomalous information of gene units in matrix form, enabling the variation of different gene units to be stored and processed in a unified and standardized manner. The construction of the type matrix and position matrix makes the anomalous type and position information of gene units clearer and more explicit, facilitating model analysis and comparison. The concatenated anomaly feature matrix integrates information on the anomalous type and position of gene units, providing rich and comprehensive feature representations for subsequent model analysis, improving the model's ability to analyze and predict gene anomalies.

[0264] Furthermore, when generating anomaly type vectors based on the anomaly types of gene units, an initial type vector containing specific identifier bits is constructed. The values ​​of these identifier bits are set according to whether the gene unit is added, deleted, or replaced. This numerical encoding method accurately represents the anomaly types of gene units in digital form, enabling the storage and transmission of gene unit variation information in a standardized manner. The anomaly type vectors of different gene units have the same length and format, differing only in the values ​​of the identifier bits. This allows the model to easily compare and analyze the anomaly types of different gene units. During data processing, the digital representation facilitates rapid computation and processing by computer equipment, improving the efficiency and accuracy of data processing.

[0265] Furthermore, when generating position vectors based on the chromosome positions of gene units, an initial position vector containing multiple candidate chromosome markers is constructed, and the target chromosome marker is determined and its value is set. This method represents the chromosome positions of gene units in a standardized digital form, making chromosome position information clearer and more accurate. The position vectors of different gene units have a unified format, facilitating batch processing and analysis by the model. By setting the value of the target chromosome marker, the position of the gene unit on the chromosome can be accurately located, providing accurate positional information for subsequent gene analysis and research, and improving the efficiency and accuracy of processing gene unit position information.

[0266] Furthermore, when the encoding network includes text encoding layers and numerical encoding layers, it extracts textual and numerical features from the detection state information respectively, and then summarizes the extracted features to obtain the detection state features. Different encoding layers process different types of information specifically, enabling more effective extraction of key features. The text encoding layer performs semantic analysis on textual information, extracting important semantic features; the numerical encoding layer performs statistical analysis on numerical information, extracting numerical features. Parallel feature extraction fully utilizes the multi-core computing power of computer devices, improving the speed of feature extraction. Summarizing the extracted textual and numerical features to obtain the detection state features allows the model to consider different types of information simultaneously, improving the comprehensive utilization efficiency of detection state information and the model's predictive ability.

[0267] Furthermore, when the decoding network includes pooling and classification layers, the pooling layer performs dimensionality reduction on multimodal biometric features to obtain one-dimensional abnormal sequence features. The classification layer further reduces dimensionality and predicts the matching probability between the object and candidate state information. Candidate state information with matching probabilities within a preset probability range is used as the predicted state information. The dimensionality reduction process of the pooling layer removes redundant information and retains key features by locally aggregating multimodal biometric features, reducing the dimensionality and complexity of the data. In terms of data storage, it reduces storage space requirements; in terms of model computation, it reduces computational load and improves computational efficiency. The classification layer further reduces dimensionality and predicts matching probabilities, enabling a more accurate determination of the degree of matching between the object and different candidate state information. By setting a preset probability range, the most likely predicted state information can be selected, improving the accuracy and reliability of the prediction results.

[0268] Furthermore, the training method for the state prediction model involves acquiring a state information dataset containing historical information of multiple sample objects. The model then predicts the state information of the samples, and trains the model based on the loss value calculated by comparing the predicted state information with the historical state information labels. The diverse datasets contain information about different objects under different conditions, enabling the model to learn a wider range of features and patterns. By continuously adjusting the model's parameters, the prediction results gradually approach the real-world situation, improving the model's accuracy and generalization ability. The end-to-end training method reduces manual intervention, lowers the difficulty of acquiring datasets, and improves the efficiency and automation of model training.

[0269] Furthermore, when the state prediction model is a gene disease prediction model, it acquires the object's gene reference, variants, chromosomal positions, and absolute gene position sequences, as well as clinical and imaging information. This information is then input into the model to extract and analyze multimodal biomarkers to predict the type of gene disease. From a genetic perspective, abnormalities in gene sequences may be the root cause of disease; analysis of gene-related sequences can reveal potential disease risks. Clinical information reflects the patient's current physical condition and symptoms, while imaging information visually displays the body's internal structures and lesions. Integrating this multimodal information into the model allows it to assess gene diseases from multiple perspectives, more accurately capturing disease characteristics and patterns, thereby improving the accuracy of gene disease prediction.

[0270] Furthermore, in the context of gene disease prediction, sequencing data of the target population is acquired, sequence alignment and variant detection are performed to obtain abnormal sequence information, and clinical and radiomics information is obtained. This information is then input into the gene disease prediction model to output the probability of the gene disease type. The abnormal sequence information obtained from the processed sequencing data provides detailed abnormalities at the gene level, clinical omics information reflects the patient's macroscopic physical condition, and radiomics information provides intuitive imaging features. Through comprehensive analysis of this multimodal information, the model can gain a more comprehensive understanding of the patient's condition, learn the relationship between various complex factors and gene diseases, and thus make more accurate gene disease predictions. In the process of disease diagnosis and treatment, accurate gene disease prediction can provide important evidence for doctors to formulate personalized treatment plans, improving treatment effectiveness and success rates.

[0271] Furthermore, in another state prediction process, textual, numerical, and image state information of the object is acquired and processed to obtain detection state information. Gene sequence information is acquired and related sequences are constructed. Features are extracted through different coding layers, concatenated and upgraded using a transformation network, and then reduced in dimensionality using a decoding network to predict the predicted state information. The pre-defined state prediction model uses each network layer to extract, analyze, and predict features from gene sequence information and detection state information, avoiding interference from subjective factors and ensuring the objectivity of the prediction results. By comprehensively considering both the macroscopic tabular features of the object and the microscopic fine-grained gene features, the actual abnormal state of the object is comprehensively and accurately characterized. From a data processing perspective, this comprehensive utilization of multimodal information can fully mine the potential information in the data, improving the analytical capability and prediction accuracy of the object's state.

[0272] Furthermore, in the pan-genome disease identification scenario, clinical omics, radiomics, and gene sequence information of the subject are acquired, processed, and input into a gene type prediction model to obtain prediction results. By integrating multimodal biometrics, information about the subject can be obtained from different omics and levels, enabling the model to gain a more comprehensive understanding of the subject's condition. Clinical omics information provides basic physiological information and disease history of patients, radiomics information displays internal body structures and lesions, and gene sequence information reveals abnormalities at the gene level. Through comprehensive analysis of this information, the model can accurately locate potential genetic disease information, providing a more accurate basis for subsequent disease analysis and research. In terms of disease prevention and control, accurate genetic disease identification can help doctors promptly identify potential disease risks, take corresponding preventive measures, and reduce the incidence of diseases.

[0273] Furthermore, in the scenario of lung gene disease discrimination, relevant object information is acquired, processed, and input into the lung gene disease type prediction model. Different types of information are encoded through each coding layer of the model, multimodal features are extracted through core network concatenation, and the decoding network performs dimensionality reduction and matching to determine the predicted lung gene disease. Analyzing lung gene diseases from multiple dimensions, comprehensively considering clinical, genetic, and imaging information, enables the model to more accurately capture the characteristics and patterns of lung gene diseases. In disease diagnosis, this improves the accuracy and specificity of diagnosis; in treatment planning, it provides doctors with a more scientific basis, helping to develop more effective treatment plans and improve patients' treatment outcomes and quality of life.

[0274] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data shall comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0275] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0276] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0277] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. A state prediction method, executed by a computer device, the method comprising: Obtain the gene sequence information of the object, compare the gene sequence information with the reference gene sequence, and obtain abnormal sequence information in the gene sequence information; The abnormal sequence information is used to characterize the abnormal location and abnormal type of the abnormal gene unit in the gene sequence information; Obtain the detection status information of the object; and The abnormal sequence information and the detection state information are input into the state prediction model to extract the abnormal sequence features of the abnormal sequence information and the detection state features of the detection state information, and the predicted state information of the object is obtained based on the abnormal sequence features and the detection state features.

2. The method according to claim 1, wherein obtaining the gene sequence information of the object includes: Obtain the initial gene sequence of the object, the initial gene sequence comprising gene units at multiple sites; The initial gene sequence is subjected to quality control processing to obtain the gene sequence information.

3. The method according to claim 1 or 2, wherein comparing the gene sequence information with a reference gene sequence to obtain abnormal sequence information in the gene sequence information includes: The gene sequence information is aligned with the reference gene sequence to obtain candidate gene sequences that are located in the candidate site set. By comparing each gene unit on the candidate gene sequence with the corresponding gene unit on the reference gene sequence, normal gene units and abnormal gene units in the candidate gene sequence are determined. Based on the reference gene sequence, the normal gene unit, and the abnormal gene unit, construct the abnormality type sequence and abnormality location sequence of the object; Based on the absolute gene position of the gene unit in the reference gene sequence, the abnormal type sequence and the abnormal position sequence are sequence aligned to obtain the abnormal sequence information.

4. The method according to claim 3, wherein the abnormal gene unit includes at least one of a newly added gene unit, a deleted gene unit, and a replaced gene unit; the abnormal type sequence includes a gene reference sequence and a gene mutation sequence; the abnormal position sequence includes a chromosomal position sequence and a gene absolute position sequence; and constructing the abnormal type sequence and abnormal position sequence of the object based on the reference gene sequence, the normal gene unit, and the abnormal gene unit includes: According to the absolute gene position of the gene unit in the reference gene sequence, the following steps are performed: sort and combine the type identifier of each gene unit in the reference gene sequence corresponding to the candidate gene sequence and the gene addition identifier of each newly added gene unit to obtain the gene reference sequence; The type identifiers of the normal gene units, the type identifiers of the replaced gene units, the type identifiers of the newly added gene units, and the gene deletion identifiers of each deleted gene unit are sorted and combined to obtain the gene mutation sequence. The chromosome position sequence is obtained by sorting and combining the chromosome position corresponding to each gene unit in the gene mutation sequence and the chromosome position of each deleted gene unit; and the absolute gene position sequence is obtained by sorting and combining the absolute gene position of each gene unit in the gene mutation sequence and the absolute gene position of each deleted gene unit.

5. The method according to any one of claims 1 to 4, wherein obtaining the detection status information of the object includes: Obtain the text status information, numerical status information, and image status information of the object; According to the preset text state transition relationship, the text state information is data mapped to obtain the text feature information of the object; The image state information is vectorized to obtain image data information of the image state information, and the numerical state information and the image data information are summarized to obtain the numerical feature information of the object; The detection state information is obtained based on the text feature information and the numerical feature information.

6. The method according to any one of claims 1-5, wherein the state prediction model comprises an encoding network, a transformation network, and a decoding network; the step of inputting the abnormal sequence information and the detection state information into the state prediction model to extract the abnormal sequence features of the abnormal sequence information and the detection state features of the detection state information, and obtaining the predicted state information of the object based on the abnormal sequence features and the detection state features, comprises: Based on the coding network, abnormal sequence features of the abnormal sequence information are extracted, and detection state features of the detection state information are extracted. The abnormal sequence features and the detection state features are concatenated based on the transformation network, and the concatenated features are then subjected to dimensionality upscaling to obtain multimodal biological features. The multimodal biometrics are reduced in dimensionality using the decoding network to output the predicted state information of the object.

7. The method according to claim 6, wherein the coding network includes a sequence aberration coding layer; the aberration sequence information includes a gene reference sequence, a gene variant sequence, a chromosome position sequence, and a gene absolute position sequence; and the step of extracting aberration sequence features based on the coding network includes: The gene reference sequence and the gene variant sequence are identified by the sequence anomaly coding layer to determine the anomaly type of each gene unit in the anomaly sequence information. Based on the anomaly type of each gene unit, an anomaly type vector of each gene unit is generated, and the anomaly type vectors of each gene unit are combined to obtain the type matrix of the anomaly sequence information. The chromosome position sequence is identified by the sequence anomaly coding layer to determine the chromosome position of each gene unit in the anomaly sequence information. Based on the chromosome position of each gene unit, a position vector of each gene unit is generated, and the position vectors of each gene unit are combined to obtain the chromosome position matrix of the anomaly sequence information. The type matrix, the chromosome position matrix, and the gene absolute position sequence are concatenated to obtain an abnormal feature matrix, which is then identified as the abnormal sequence feature.

8. The method according to claim 7, wherein generating an abnormality type vector for each of the gene units based on the abnormality type of each gene unit comprises: For any gene unit, construct an initial type vector for that gene unit; The initial type vector includes newly added gene identifiers, deleted gene identifiers, multiple types of reference gene identifiers, and multiple types of variant gene identifiers; If the gene unit is a newly added gene unit, then the values ​​of the newly added gene identifier and the variant gene identifier that matches the type of the gene unit in the initial type vector are set to the first preset value, and the values ​​of other identifiers in the initial type vector are set to the second preset value, thereby generating the abnormal type vector of the gene unit. If the gene unit is a deleted gene unit, then the value of the deleted gene identifier in the initial type vector and the value of the reference gene identifier that matches the type of the gene unit are set to the first preset value, and the values ​​of other identifiers in the initial type vector are set to the second preset value, thereby generating an abnormal type vector of the gene unit. If the gene unit is a replacement gene unit, then the historical gene type of the gene unit is determined, and the values ​​of the reference gene identifier matching the historical gene type and the variant gene identifier matching the gene unit are set to the first preset value. The values ​​of other identifiers in the initial type vector are set to the second preset value, and the abnormal type vector of the gene unit is generated.

9. The method according to claim 7 or 8, wherein generating the position vector of each of the gene units based on their chromosomal positions comprises: For any gene unit, construct the initial position vector of the gene unit; The initial position vector includes multiple candidate chromosome identifier bits; Among the multiple candidate chromosome markers, the target chromosome marker corresponding to the chromosome position of the gene unit is determined, and the value of the target chromosome marker is set to a first preset value. The values ​​of other chromosome markers in the initial position vector are set to a second preset value to obtain the position vector of the gene unit.

10. The method according to any one of claims 6 to 9, wherein the encoding network comprises a text encoding layer and a numerical encoding layer; and the step of extracting the detection state features of the detection state information comprises: Text features are extracted from the detection state information based on the text encoding layer, and numerical features are extracted from the detection state information based on the numerical encoding layer. The detection state features are determined based on the text features and the numerical features.

11. The method according to any one of claims 6 to 10, wherein the decoding network comprises a pooling layer and a classification layer; the step of performing dimensionality reduction processing on the multimodal biometric features based on the decoding network and outputting the predicted state information of the object comprises: Based on the pooling layer, the multimodal biological features are dimensionality reduced to obtain one-dimensional abnormal sequence features of the object; Based on the classification layer, the one-dimensional abnormal sequence features are reduced in dimensionality to predict the matching probability between the object and each candidate state information in the candidate state information set. The candidate state information whose matching probability is within a preset probability interval is determined as the predicted state information of the object.

12. The method according to any one of claims 1 to 11, wherein the training method of the state prediction model comprises: Obtain a status information dataset; the status information dataset includes historical anomaly sequence information, historical detection status information, and historical status information labels for multiple sample objects; The state prediction model is used to predict the historical abnormal sequence information and the historical detection state information to obtain sample predicted state information. The state prediction model is then trained based on the loss value of the sample predicted state information and the historical state information label to obtain the state prediction model.

13. The method according to any one of claims 1 to 12, wherein the state prediction model is a gene disease prediction model; the method further comprises: Obtain the gene reference sequence, gene variant sequence, chromosome position sequence, and gene absolute position sequence of the object; Obtain the clinical and imaging information of the subject; The gene reference sequence, the gene variant sequence, the chromosome position sequence, the gene absolute position sequence, the clinical information, and the imaging information are input into the gene disease prediction model to extract the multimodal biological features of the object, and the multimodal biological features are analyzed to predict the gene disease type of the object.

14. A state prediction device, the device comprising: The sequence acquisition module is used to acquire the gene sequence information of the object, compare the gene sequence information with a reference gene sequence, and acquire abnormal sequence information in the gene sequence information. The abnormal sequence information includes the abnormal location and abnormal type of the gene unit; A status acquisition module is used to acquire the detection status information of the object; The state analysis module is used to input the abnormal sequence information and the detection state information into the state prediction model to extract the abnormal sequence features of the abnormal sequence information and the detection state features of the detection state information, and obtain the predicted state information of the object based on the abnormal sequence features and the detection state features.

15. A computer device comprising a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement the steps of the method according to any one of claims 1 to 13.

16. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 13.

17. A computer program product comprising a computer program that, when executed by a processor, implements the steps of the method according to any one of claims 1 to 13.