Text task processing method, system, device, medium and computer program product

Through text data-assisted methods, effective information is mined from gene expression data and accurate gene expression information text data are solved, which is a problem of low analysis efficiency and low accuracy caused by artificial participation in the prior art.

CN119495359BActive Publication Date: 2025-06-03INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510068250.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-06-03
Estimated Expiration
2045-01-16

AI Technical Summary

Technical Problem

The prior art requires manual participation in the selection and optimization of characteristic parameters when analyzing gene expression data, resulting in low accuracy, high cost and low efficiency of analysis results.

Method used

With text data assistance, more effective gene expression information is mined from gene expression data and text data containing more accurate gene expression information is generated. The specific method includes obtaining text data and gene expression data of the to-process text task, encoding the gene expression data into gene global vector data, and fusing it with the text data, and generating text data containing gene information through decoding processing.

Benefits of technology

It realizes efficient and accurate analysis of the gene expression matrix without manual participation in feature parameter selection and optimization, and generates text data containing accurate gene expression information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119495359B_ABST
    Figure CN119495359B_ABST
Patent Text Reader

Abstract

The present invention discloses a text task processing method, system, device, medium and computer program product, which is applied to the field of artificial intelligence technology. The method includes encoding the text data and gene expression data corresponding to the text task to be processed respectively to obtain gene global vector data, text vector data and global text vector data. Fusing the pre-order text sequence that already exists during the current decoding and the gene global vector data, and performing decoding processing on the text gene fusion data to obtain text data containing gene information corresponding to the current decoding step; determining the execution result of the text task to be processed according to the text data output by each decoding step. The present invention can solve the problem of manual intervention analysis based on expression data in the related art, and can mine more effective gene expression information from gene expression data with the assistance of text data, and can generate text data containing more accurate gene expression information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a text task processing method, system, device, medium and computer program product. Background Art

[0002] With the development of high-throughput sequencing technology, large-scale gene expression data can be obtained. The gene expression level analysis results of gene expression data are widely used in bioinformatics and molecular biology, such as constructing a gene regulatory network model for describing the regulatory relationships between genes through gene expression data.

[0003] When analyzing gene expression data in related technologies, it is necessary to rely on manual participation in the selection and optimization of feature parameters. For example, in the process of mining gene regulatory relationships from a gene expression matrix, it is necessary to rely on manual participation in the selection and optimization of feature parameters. The manual intervention method can neither guarantee the accuracy of the final gene analysis results, nor capture all the effective information, and it also leads to high costs and low efficiency in the entire gene analysis process.

[0004] In view of this, how to analyze a gene expression matrix more efficiently and accurately without relying on manual participation in the selection and optimization of feature parameters is a technical problem that needs to be solved by those skilled in the art.

[0005] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present application, and therefore may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0006] The present invention provides a text task processing method, system, electronic device, non-volatile storage medium and computer program product, which can mine more effective gene expression information from gene expression data through text data assistance, and generate text data containing more accurate gene expression information.

[0007] To solve the above technical problems, the present invention provides the following technical solutions:

[0008] On the one hand, the present invention provides a text task processing method, including:

[0009] Obtain the text data and gene expression data corresponding to the text task to be processed; the text task to be processed is a text task for generating text containing gene information; encode the gene expression data added with the position learning parameter and the global learning parameter into gene global vector data; encode the text data into text vector data according to the text task to be processed and / or encode the fusion data of the gene global vector data and the text data into global text vector data; fuse the pre-order text sequence that already exists during the current decoding and the gene global vector data, and perform decoding processing on the text gene fusion data to obtain the text data containing gene information corresponding to the current decoding step; determine the execution result of the text task to be processed according to the text data output by each decoding step; wherein, the position learning parameter is used to learn the sequential feature information of the gene expression data, and the global learning parameter is used to learn the global feature of the gene expression data.

[0010] In the first exemplary embodiment, the gene expression data is a gene expression matrix, and encoding the gene expression data added with the position learning parameter and the global learning parameter into gene global vector data includes: dividing the gene expression data into multiple sub-gene expression matrices, and adding the position learning parameter to each sub-gene expression matrix; adding the global learning parameter before the gene vector sequence data corresponding to each sub-gene expression matrix added with the position learning parameter as the gene expression matrix feature; inputting the gene expression matrix feature into the gene expression matrix encoding module to perform encoding processing on the gene expression matrix feature by using the gene expression matrix encoding module; the input position of the gene expression matrix encoding module includes a classification marker; taking the output corresponding to the classification marker of the gene expression matrix encoding module as the gene global vector data.

[0011] In the second exemplary embodiment, encoding the text data into text vector data includes: performing word segmentation on the text data to obtain multiple text words; inputting each text word into the text encoding module to perform encoding processing by using the text encoding module; the input position of the text encoding module includes a classification marker; taking the output corresponding to the classification marker of the text encoding module as the text vector data.

[0012] In a third exemplary embodiment, encoding the fusion data of the gene global vector data and the text data into global text vector data includes: performing word segmentation on the text data to obtain an initial text embedding sequence; concatenating the gene global vector data before the initial text embedding sequence to obtain fusion vector data; inputting the fusion vector data into a gene-aware text encoding module to perform encoding processing on the fusion vector data by using the gene-aware text encoding module; and using the output of the gene-aware text encoding module as the global text vector data.

[0013] In a fourth exemplary embodiment, encoding the fusion data of the gene global vector data and the text data into global text vector data includes: performing word segmentation on the text data to obtain an initial text embedding sequence; inputting the initial text embedding sequence and the gene global vector data into a gene-aware text encoding module to perform interactive calculation on the query matrix of each text word in the initial text embedding sequence and the key matrix and value matrix of the gene global vector data, and using the interactive calculation result as the fusion vector data for encoding processing; and using the output of the gene-aware text encoding module as the global text vector data.

[0014] In a fifth exemplary embodiment, fusing the pre-order text sequence that already exists during current decoding and the gene global vector data, and performing decoding processing on the text-gene fusion data to obtain text data containing gene information corresponding to the current decoding step includes: concatenating the gene global vector data before the pre-order text sequence to obtain text-gene fusion data; using a transformer network model as a gene-aware text decoding module, and inputting the text-gene fusion data into the gene-aware text decoding module to perform decoding processing through the gene-aware text decoding module.

[0015] In a sixth exemplary embodiment, the fusing of the pre-order text sequence that already exists during current decoding and the gene global vector data includes: creating an upper triangular matrix with the same length as the pre-order text sequence, and setting the values on the right side of the diagonal of the upper triangular matrix to infinity as mask information; performing a masking operation on each text word of the pre-order text sequence by using the mask information; determining attention weights according to the query matrix, key matrix, and value matrix of the pre-order text sequence, and determining the context vector of each text word of the pre-order text sequence according to the attention weights; generating a text segment of the pre-order text sequence according to the context vector of each text word, and fusing the text segment and the gene global vector data.

[0016] In a seventh exemplary embodiment, it further includes: pre-building a text task processing model; the text task processing model includes an input layer, an encoding layer, a decoding layer, and an output layer; obtaining a text training sample set matching the text task to be processed, the text training sample set includes multiple groups of text-gene training samples, and the text-gene training sample includes a text sample and a corresponding gene expression sample; using the text training sample set to train the text task processing model until the model training stop condition is reached, to obtain a text task processing model for executing the text task to be processed; wherein, the input layer inputs the text sample and the gene expression sample; the encoding layer includes a gene expression matrix encoder, a text encoder, and a gene-aware text encoder; the gene expression matrix encoder encodes the gene expression sample added with position learning parameters and global learning parameters into a gene global vector sample, the text encoder encodes the text sample into a text vector sample, and the gene-aware text encoder encodes the text sample spliced with the gene global vector sample into a global text vector sample; the decoding layer fuses the pre-order text sequence sample that already exists during the current decoding and the gene global vector sample, and decodes the text-gene fusion sample to obtain a text sample data containing gene information corresponding to the current decoding.

[0017] In an eighth exemplary embodiment, after obtaining the text data and gene expression data corresponding to the text task to be processed, it further includes: inputting the text data and the gene expression data into the text task processing model; the input layer of the text task processing model receives the text data and the gene expression data; the gene expression matrix encoder encodes the gene expression data added with position learning parameters and global learning parameters into a gene global vector data, the text encoder encodes the text data into a text vector data, and the gene-aware text encoder encodes the fusion data of the gene global vector data and the text data into a global text vector data; the decoding layer fuses the pre-order text sequence that already exists during the current decoding and the gene global vector data, and decodes the text-gene fusion data to obtain a text data containing gene information corresponding to the current decoding step; the output layer determines the execution result of the text task to be processed according to the text data output by each decoding step, and outputs it.

[0018] In a ninth exemplary embodiment, training the text task processing model using the text training sample set includes: pre-constructing an initial loss function for the text task processing model; the initial loss function includes a gene text retrieval loss, a gene-aware text encoding loss, and a gene-aware text generation loss; generating weight factors for the gene text retrieval loss, the gene-aware text encoding loss, and the gene-aware text generation loss according to the task type of the text task to be processed; adjusting the initial loss function according to the respective weight factors of the gene text retrieval loss, the gene-aware text encoding loss, and the gene-aware text generation loss to obtain a loss function; training the text task processing model based on the loss function using the text training sample set; wherein, the gene text retrieval loss is used to maximize the similarity between the text sample and its corresponding gene expression sample while minimizing the similarity between the text sample and other gene expression samples; the gene-aware text encoding loss is used to evaluate the matching degree of the text task processing model for the input text sample and gene expression sample; the gene-aware text generation loss is used to evaluate the ability of the text task processing model to generate text under the condition of gene expression context.

[0019] In a tenth exemplary embodiment, the text task to be processed is a gene-text matching task. Training the text task processing model using the text training sample set includes: initially training the text task processing model using the text training sample set to obtain text sample-gene expression sample pairs and corresponding similarity scores generated during the initial training process of the text task processing model, and using the similarity scores as confidence scores for the text sample-gene expression sample pairs; selecting target text sample-gene expression sample pairs that meet a preset similarity condition from each text sample-gene expression sample pair; obtaining variant text sample-gene expression sample pairs for each target text sample-gene expression sample pair by adding fluctuation conditions and / or replacing similar text descriptions; and retraining the text task processing model obtained from the initial training using each target text sample-gene expression sample pair and each variant text sample-gene expression sample pair.

[0020] The present invention also provides an electronic device including a processor, and the processor is used to implement the steps of the text task processing method as described in any one of the previous ones when executing a computer program stored in a memory.

[0021] Finally, the present invention also provides a non-volatile storage medium, on which a computer program is stored, and the computer program is used to implement the steps of the text task processing method as described in any one of the previous ones when executed by a processor.

[0022] The present invention also provides a computer program product, including computer programs / instructions, which, when executed by a processor, implement the steps of the text task processing method described in any one of the preceding paragraphs.

[0023] The present invention also provides a text task processing system, including a user interface and a text task processor; the user interface at least includes a task trigger area, a task processing result display area, a data input area, and / or a file import area, so that a user can describe a text task to be processed through the data input area, input text data and gene expression data through the data input area or the file import area, and send the text task to be processed, the text data, and the gene expression data to the text task processor through the task trigger area; the text task processor, based on the text task to be processed, processes the text data and the gene expression data by using the steps of the text task processing method described in any one of the preceding paragraphs, generates a text task processing result, and sends the text task processing result to the task processing result display area of the user interface.

[0024] The advantages of the technical solution provided by the present invention are as follows: the gene expression data and the text data are respectively encoded to construct a multi-modal latent space. By encoding and processing the gene expression data, non-linear features of gene expression can be extracted, and further, the deep non-linear relationship between the gene expression data and the text information can be captured, and the implicit biological information can be extracted. By using the combination of gene expression embedding and text embedding for encoding and processing, it is ensured that the content in the text can reflect the characteristics of gene expression and strengthen the understanding of gene expression by the text. During the text generation process, by introducing gene expression information, it is ensured that the generated text is highly relevant to the gene data, and the sparsity of the gene expression data and the context dependence of the text data can be effectively processed, realizing the effective combination of gene expression and text. Through the text, the deep expression relationship of genes can be efficiently and accurately mined from the gene expression data. The whole process does not require manual participation in the selection and optimization of feature parameters, and finally generates text data containing more accurate gene expression information. In addition, the present invention also provides a corresponding implementation system, electronic device, non-volatile storage medium, and computer program product for the text task processing method, further making the method more practical, and the system, electronic device, non-volatile storage medium, and computer program product have corresponding advantages.

[0025] The technical features mentioned above, the technical features to be mentioned below, and the technical features shown separately in the drawings can be combined with each other arbitrarily, as long as the combined technical features are not mutually contradictory. All feasible feature combinations are the technical contents clearly recorded in this article. Any one of the multiple sub-features included in the same statement can be applied independently without necessarily being applied together with other sub-features. It should be understood that the above general description and the following detailed description are only exemplary and do not limit the present invention. Description of the Drawings

[0026] In order to more clearly illustrate the technical solutions of the present invention or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0027] Figure 1 It is a schematic flow chart of a text task processing method provided by the present invention;

[0028] Figure 2 It is a schematic encoding flow chart of gene expression data provided by the present invention;

[0029] Figure 3 It is a schematic framework diagram of a text task processing model in an exemplary application scenario provided by the present invention;

[0030] Figure 4 It is a schematic structural framework diagram of an exemplary implementation manner of a text task processing device provided by the present invention;

[0031] Figure 5 It is a schematic structural framework diagram of an exemplary implementation manner of an electronic device provided by the present invention;

[0032] Figure 6 It is a schematic structural framework diagram of an exemplary implementation manner of a text task processing system provided by the present invention. Detailed Embodiments

[0033] To enable those skilled in the art to better understand the technical solutions of the present invention, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. Among them, the terms "first", "second", "third", "fourth", etc. in the specification and the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. The term "exemplary" means "serving as an example, embodiment or illustration". Any embodiment described herein as "exemplary" does not necessarily have to be construed as superior or better than other embodiments.

[0034] Gene expression data includes gene expression levels under different conditions. As the scale of legally obtained gene expression data becomes larger and larger, the analysis results of gene expression data are widely applied to biological-related fields, such as predicting gene regulatory networks based on gene expression matrices, and then constructing gene regulatory network models. Through this model, the regulatory relationships between genes can be obtained, the roles of genes in cell and tissue functions can be clarified, whether a certain type of gene has an important impact on the expression of other genes can be inferred, and key regulatory factors and regulatory pathways can be identified, which is beneficial to deeply understanding the complexity of biological systems and the occurrence mechanisms of biological processes, thereby supporting gene function annotation, gene expression pattern analysis, and biomarker discovery. Predicting gene regulatory networks usually uses gene expression matrices as input features and applies machine learning and statistical models, such as correlation analysis, linear models, clustering analysis, factor analysis, and Bayesian networks for analysis, and can discover potential regulatory relationships from large-scale gene expression data.

[0035] As described above, when analyzing gene expression data, related technologies rely on traditional machine learning and need to depend on feature parameters selected and optimized manually by professionals based on their knowledge and experience. Different selections of feature parameters will lead to different final gene analysis results, and the feature parameters selected manually cannot capture all information. This results in not only inaccurate results but also a large amount of human and time costs in the gene expression data analysis of related technologies. Taking the construction of a gene regulatory network model as an example, its model performance highly depends on the selection and optimization of feature parameters, which will lead to limitations in model performance, unable to capture deeper relationships and patterns, and also limit the expressive ability of shallow models. The ability to predict gene regulatory networks and the interpretability cannot meet user needs. Related technologies improve the analysis process of gene expression data through methods such as feature extraction modules, co-expression gene classification, and boundary recognition. However, this method usually requires a large number of manually labeled samples for training for each biological process category, and obtaining a large amount of high-quality labeled data is a time-consuming and costly process. Moreover, the quality and quantity of training samples will also affect the accuracy and reliability of the final results. In addition, there are related technologies that analyze gene expression data through the interaction between gene expression matrices and text, but they cannot finely distinguish functional co-expression modules from background noise, so they cannot identify the functions of co-expression modules at a finer granularity. To solve this problem, related technologies have also proposed an analysis method for optimizing gene expression matrices. Transfer learning learns semantic associations between meaningful biological process categories from noise, such as using the descriptions or attributes of biological processes to identify categories never seen in the training stage. However, the essence of these two methods is to align biological expression data and text data, and they do not fully utilize text data in the process of gene expression data analysis. In addition, there is a related technology that proposes to use discretization and autoregressive modeling to process gene expression data. This method maps high-dimensional and complex gene expression data to a low-dimensional and discretized latent space through a quantization and encoding process. On the basis of retaining the key biological information in the original data, it can reduce the representation complexity of gene expression data and improve the modeling efficiency of gene expression patterns. However, this method also does not fully utilize text data and cannot deeply explore the relationship between gene expression data and text data. In view of this, the present invention encodes gene expression data and text data separately, constructs a multi-modal latent space, and uses the combination of gene expression embedding and text embedding for encoding processing to ensure that the content in the text can reflect the characteristics of gene expression and strengthen the understanding of gene expression by the text. In the process of text generation, by introducing gene expression information, it is ensured that the generated text is highly relevant to gene data, and finally text data containing more accurate gene expression information is generated. After introducing the technical solution of the present invention, the various non-limiting embodiments of the present invention will be described in detail below.To better illustrate the present invention, numerous specific details are given in the following specific embodiments. Those skilled in the art should understand that the present invention can also be implemented without these specific details. In other instances, methods, means, elements, and circuits well-known to those skilled in the art are not described in detail to highlight the gist of the present invention.

[0036] First, please refer to Figure 1 , Figure 1 , which is a schematic flowchart of a text task processing method provided in this embodiment. This embodiment may include the following contents:

[0037] S101: Obtain text data and gene expression data corresponding to the text task to be processed.

[0038] The text task to be processed in this step is a text task for generating text containing gene information, including but not limited to text-gene matching tasks, gene annotation tasks, gene function prediction tasks, gene expression pattern analysis, text generation tasks. The text generation task may be, for example, the generation of the association analysis result between gene expression and disease symptoms. By analyzing gene expression data and relevant text data, text data containing the gene information required by the text task to be processed is finally output. The text data may be, for example, medical reports, symptom descriptions, or disease progress related to the gene expression matrix. The gene expression data may be a gene expression matrix, where the rows represent the expression of a gene under different environmental conditions or at different time points, and the columns represent the expression of all genes under different conditions or samples, such as tissues, experimental conditions, treatment factors, etc. The data in each cell represents the expression level of a specific gene in a specific sample.

[0039] S102: Encode the gene expression data added with position learning parameters and global learning parameters into gene global vector data, encode the text data into text vector data according to the text task to be processed and / or encode the fusion data of the gene global vector data and the text data into global text vector data.

[0040] It can be understood that gene expression data is usually represented by a gene expression matrix, which has high-dimensional sparsity, and text data has context dependence. It is usually difficult to directly model the deep association between the two. In order to establish an effective non-linear model between the gene expression matrix and text data and capture their complex interaction relationships, this step encodes the gene expression data and text data respectively to construct a multi-modal latent space and capture their deep interactions. It can be understood that for some text tasks to be processed, the demand for the deep relationship between text data and gene expression data is not high. In order to improve the execution efficiency of the text tasks to be processed, only the text data can be encoded into text vector data. Some text tasks to be processed need to capture the semantic associations within the text and generate context representations related to the gene expression matrix embedding at the same time, such as gene-text matching tasks. Therefore, it is also necessary to encode the fusion data of gene global vector data and text data.

[0041] Among them, since the order and position of the expression level of each gene in the gene expression matrix have biological significance under different conditions, in order to obtain a more accurate gene expression relationship, this step introduces position encoding when encoding the gene expression matrix, that is, learning the sequential feature information of the gene expression data through learnable position learning parameters during model training, so as to capture the local and global sequential information of the gene expression matrix. In order to obtain the global representation of the entire gene expression data, another learnable parameter, that is, the global learning parameter, can also be added to summarize the characteristics of the entire gene expression data.

[0042] S103: Fuse the pre-order text sequence that already exists during the current decoding and the gene global vector data, and decode the text-gene fusion data to obtain the text data containing gene information corresponding to the current decoding step; determine the execution result of the text task to be processed according to the text data output by each decoding step.

[0043] Based on the encoding of text data and gene expression data in the previous step, in order to effectively handle the sparsity of gene expression data and the context dependence of text data and achieve the effective combination of gene expression and text, this step combines gene expression data at each decoding step during the process of generating or translating text, so that the generated text data can contain important content related to the gene expression matrix. By introducing gene expression embedding, relevant words or sentence structures related to gene information can be adaptively selected during text generation, improving the accuracy and interpretability of gene-related text generation tasks.

[0044] In the technical solution provided in this embodiment, the gene expression data and text data are respectively encoded to construct a multi-modal latent space. By encoding the gene expression data, the non-linear features of gene expression can be extracted, and then the deep non-linear relationship between the gene expression data and text information can be captured, and the implicit biological information can be extracted. By using the combination of gene expression embedding and text embedding for encoding processing, it is ensured that the content in the text can reflect the characteristics of gene expression and strengthen the understanding of gene expression by the text. In the process of text generation, by introducing gene expression information, it is ensured that the generated text is highly relevant to the gene data, the sparsity of gene expression data and the context dependence of text data can be effectively processed, the effective combination of gene expression and text can be realized, and the deep expression relationship of genes can be efficiently and accurately mined from the gene expression data through the text. The whole process does not require manual participation in the selection and optimization of feature parameters, and finally generates text data containing more accurate gene expression information.

[0045] It should be noted that there is no strict order of execution between the steps in the present invention. As long as it conforms to the logical order, these steps can be executed simultaneously or in a certain preset order. Figure 1 This is just a schematic way and does not mean that it can only be such an execution order.

[0046] In the above embodiment, there is no limitation on how to encode the gene expression data to generate the gene global vector data. Based on the above embodiment, the present invention also gives an exemplary generation process of the gene global vector data, which may include the following content:

[0047] The gene expression data is a gene expression matrix. The gene expression data is divided into multiple sub-gene expression matrices, and position learning parameters are added to each sub-gene expression matrix; before the gene vector sequence data corresponding to each sub-gene expression matrix with added position learning parameters, a global learning parameter is added as the gene expression matrix feature; the gene expression matrix feature is input into the gene expression matrix encoding module to use the gene expression matrix encoding module to encode the gene expression matrix feature; the output corresponding to the classification label of the gene expression matrix encoding module is used as the gene global vector data.

[0048] In this embodiment, the input gene expression data can be first divided according to each sub-matrix, and then these divided parts are embedded into a vector space. Taking the gene expression data as a two-dimensional matrix For example, let G represent the number of genes, C represent the number of experimental conditions or samples, and R represent the matrix. It can be divided into N sub-matrices according to the expression data of g genes under c conditions for each sub-matrix. For the convenience of description, the divided sub-matrices are defined as sub-gene expression matrices. Correspondingly, the size of each sub-gene expression matrix is g×c. For the convenience of subsequent processing, each sub-gene expression matrix can be flattened into a one-dimensional vector. Correspondingly, the i-th sub-gene expression matrix can be represented as . For the convenience of subsequent data processing, the flattened sub-gene expression matrix can be projected into a higher-dimensional space through a learnable linear transformation matrix E. E represents the embedding dimension. Correspondingly, the i-th sub-gene expression matrix can represent an embedding vector . The set of vector representations of all sub-gene expression matrices is the gene vector sequence data. Similarly, considering that the order and position of the expression levels of each gene in the gene expression matrix have biological significance under different conditions, each sub-gene expression matrix of the gene expression matrix has position sensitivity. Therefore, in this embodiment, position encoding is introduced for each gene expression matrix to provide position information. The position encoding assigns a specific vector to the gene position and condition position of each sub-matrix, so that the model can capture the local and global order information of the gene expression matrix. Correspondingly, a learnable position information is added to the embedding vector of the i-th sub-gene expression matrix, that is: . is a learnable position encoding related to the relative position of the i-th sub-gene expression matrix, is the i-th sub-gene expression matrix with added position learning parameters. To obtain the global representation, after obtaining the gene vector sequence data with added position learning parameters, a [CLS] (classification) marker can also be added before it. For the convenience of description, the present invention defines it as the global learning parameter. The global learning parameter is used to represent the global representation of the entire gene expression matrix, and its output embedding after being processed by the subsequent gene expression matrix encoding module will be used to summarize the features of the entire gene expression matrix. The global learning parameter is a learnable parameter and can be represented as . At this time, the gene expression matrix features obtained can be represented as .

[0049] In this embodiment, the gene expression matrix encoding module can adopt any network model structure capable of encoding matrix data or image data. The input position of the gene expression matrix encoding module can include classification tags corresponding to global learning parameters. The output of the classification tags will serve as the global representation of the entire gene expression matrix and be used to capture the overall correlation information between gene expression and text during the execution of text processing tasks. For example, the gene expression matrix encoding module can adopt a Transformer encoder, such as Figure 2 shown. The gene expression matrix encoding module includes multiple Transformer blocks, each Transformer block being one layer. The Transformer block can include a multi-head attention layer based on the MHSA (Multi-Head Self-Attention) mechanism, a residual connection and a normalization layer, and an FFN (Feed-Forward Network) layer. The output of each layer is passed to the next layer through the residual connection and the normalization layer. Correspondingly, the gene expression matrix features are input into the gene expression matrix encoding module, and each layer of Transformer blocks in the gene expression matrix encoding module processes the gene expression matrix features. The output can be expressed as: . Among them, MHSA represents the multi-head attention layer, FFN represents the feed-forward network layer, L is the number of Transformer layers. After encoding by all L layers of Transformer, the final output of the classification tags represents the global embedding of the entire gene expression matrix, that is, the gene global vector data can be expressed as: , D represents the dimension, that is, the dimension after each sub-gene expression matrix is flattened and projected into a high-dimensional space.

[0050] Furthermore, in order to improve the processing efficiency of the gene expression matrix encoding module, a variational autoencoder (VAE) or a graph neural network (GNN) can be used to process the high-dimensional and sparse gene expression matrix first.

[0051] Further, in order to improve the efficient encoding of gene expression data, a gene expression matrix encoder may be pre-constructed in this embodiment. The input gene expression data is encoded by the gene expression matrix encoder, and gene global vector data is output. The gene expression matrix encoder includes an input layer, a gene expression matrix representation layer, a position information addition layer, a classification marker addition layer, a gene expression matrix encoding module, and an output layer. The input layer is used to receive gene expression data. The gene expression matrix representation layer divides the gene expression data into multiple sub-gene expression matrices. The position information addition layer adds position learning parameters to each sub-gene expression matrix. The classification marker addition layer adds global learning parameters before the gene vector sequence data corresponding to each sub-gene expression matrix with added position learning parameters as the gene expression matrix feature. The gene expression matrix feature is input into the gene expression matrix encoding module, and the gene expression matrix encoding module encodes the gene expression matrix feature. The output layer is used to use the output corresponding to the classification marker of the gene expression matrix encoding module as the gene global vector data.

[0052] As can be seen from the above, by extracting the non-linear features of gene expression data based on Transformer in this embodiment, the deep non-linear relationship between gene expression data and text information can be captured, and the implicit biological information can be extracted, thereby improving the analysis effect of gene expression data.

[0053] The above embodiments do not make any limitations on how to encode text data. Based on the above embodiments, the present invention also gives an exemplary generation method of text vector data, which may include the following content:

[0054] The text data is segmented to obtain multiple text words. Each text word is input into the text encoding module for encoding by the text encoding module. The output corresponding to the classification marker of the text encoding module is used as the text vector data.

[0055] In this embodiment, the text encoding module may use any network model structure capable of encoding text data, and the present invention does not make any limitations on this. Similarly, the input position of the text encoding module includes a classification marker. After encoding the input text data, the output of the [CLS] token (classification marker) is obtained as the global embedding of the entire text, that is, the text vector data is obtained. The text vector data can be expressed as , where D represents the dimension of matrix R. The text vector data can be used for alignment with the gene global vector data or other downstream tasks. The so-called alignment means that the literal description is consistent with the information contained in the gene expression matrix, that is, the text vector data can accurately describe the meaning of the gene expression matrix.

[0056] Further, in order to improve the efficient encoding of text data, a text encoder can be pre-constructed. For example, a text encoder can be constructed based on the architecture of BERT (Bidirectional Encoder Representations from Transformers, a pre-trained language model using bidirectional encoding). When the text data is input into the text encoder, the text data is first tokenized and embedded into the vector space. After being encoded by the L-layer BERT Transformer, the text vector data is obtained according to the [CLS] token output.

[0057] The above embodiments do not make any limitations on how to encode the fusion data of the gene global vector data and the text data. Based on the above embodiments, the present invention also provides two parallel global text vector data generation methods according to different fusion methods of the gene global vector data and the text data, which may include the following:

[0058] In an exemplary implementation, the text data is tokenized to obtain an initial text embedding sequence; the gene global vector data is concatenated before the initial text embedding sequence to obtain fusion vector data; the fusion vector data is input into the gene-aware text encoding module to encode the fusion vector data by using the gene-aware text encoding module; the output of the gene-aware text encoding module is used as the global text vector data. In another exemplary implementation, the text data is tokenized to obtain an initial text embedding sequence; the initial text embedding sequence and the gene global vector data are input into the gene-aware text encoding module to perform interactive calculations on the query matrix of each text word in the initial text embedding sequence and the key matrix and value matrix of the gene global vector data by using the gene-aware text encoding module, and the interactive calculation result is used as the fusion vector data for encoding; the output of the gene-aware text encoding module is used as the global text vector data.

[0059] In order to capture the content related to the characteristics of the gene expression matrix in the text in this embodiment, not only the text data needs to be encoded conventionally, but also the context information related to gene expression needs to be introduced. Exemplarily, on the basis of the standard text encoding, the global information embedding of the additional gene expression matrix is added. That is, in order to have the gene-aware ability, the initial text embedding sequence will be combined with the global embedding from the gene expression matrix. The global embedding of the gene expression matrix comes from the output of the aforementioned gene expression matrix encoder, that is, the gene global vector data is embedded in the text data. In order to realize the fusion of the gene global vector data and the text data, the text data can be first tokenized and embedded into the vector space to obtain an initial text embedding sequence, and the initial text embedding sequence can be expressed as , where denote the embedding vector of the i-th word in the text data, and n is the length of the text data. After obtaining the initial text embedding sequence, a simple data fusion method is to use vector concatenation, that is, concatenate the gene global vector data into the initial text embedding sequence to obtain fusion vector data that can provide additional information about the gene context. Exemplarily, the fusion vector data can be represented as . Any network model capable of encoding vector data can be used to build the gene-aware text encoding module. For example, the gene-aware text encoding module can be based on the multi-layer BERT network model structure. BERT includes multiple Transformer blocks. Input the fusion vector data into the gene-aware text encoding module. Each layer of the Transformer in the gene-aware text encoding module will process the fusion vector data through the multi-head self-attention mechanism and use the feed-forward network (FFN) to perform further calculations on the processed information. After L layers of Transformer calculations, a global representation of the text containing the influence of gene information is obtained. This process can be represented by the following relational expression , is the output of the gene-aware text encoding module. The [CLS] token after Transformer encoding will be used as the global embedding representation of the entire text, which not only contains the semantic information in the text but also combines the context information related to gene expression. The present invention defines it as global text vector data. The global text vector data can be represented, for example, as: , is the output corresponding to the [CLS] token. The global text vector data can be used for further alignment and matching with the embedding of the gene expression matrix. In downstream text processing tasks, such as gene-text matching tasks, gene function annotation prediction tasks, etc., implicit information related to genes in the text can be captured.

[0060] Compared with simply concatenating the gene expression global embedding in front of the text embedding sequence, in order to further improve the information fusion between the text and the gene expression matrix, the present invention can also achieve the information fusion between the text data and the gene expression matrix by introducing the cross-attention mechanism. In this embodiment, the gene-aware text encoding module further includes a cross-attention layer. During the encoding process of the gene global vector data and the text data by the gene-aware text encoding module, the cross-attention layer dynamically focuses on specific parts of the text and the relevant parts of the corresponding gene expression matrix. Through cross-attention, in each encoding step, not only the context information of the text can be captured, but also different text segments can be assigned different importance weights according to the characteristics of the gene expression matrix.

[0061] Furthermore, in order to improve the efficient encoding of the fused data of gene global vector data and text data, this embodiment may also pre-construct a gene-aware text encoder. The gene-aware text encoder fuses and encodes the text embeddings and gene expression embeddings obtained by the text encoder and the gene expression matrix encoder respectively, and finally outputs global text vector data. Exemplarily, the gene-aware text encoder may include a text sequence generation layer, a concatenation layer, a gene-aware text encoding module, and an output layer. The text sequence generation layer performs word segmentation on the text data to obtain an initial text embedding sequence; the concatenation layer concatenates the gene global vector data in front of the initial text embedding sequence to obtain fused vector data; the gene-aware text encoding module encodes the fused vector data; the output layer uses the output of the classification token of the gene-aware text encoding module as the global text vector data. Exemplarily, the gene-aware text encoder may also include a gene-aware text encoding module and an output layer. The gene-aware text encoding module obtains the initial text embedding sequence and the gene global vector data through the text encoder and the gene expression matrix encoder respectively. In the Transformer layer, the query matrix Q of each text word is interacted with the key matrix K and value matrix V of the gene expression matrix to generate an attention-weighted text embedding containing gene expression context, and then the interaction calculation result is used as the fused vector data for encoding processing. Finally, the output corresponding to the classification token of the gene-aware text encoding module is used as the global text vector data through the output layer. It can not only capture the potential relationship between gene information in the text, but also form a more accurate global text embedding representation by adaptively learning the part closely related to gene expression in the text data.

[0062] As can be seen from the above, through the gene-aware text encoder, this embodiment can obtain text data that perceives the global information of the gene expression matrix, encodes the text data into a high-dimensional vector representation, which is convenient for combining with the latent representation of the gene expression matrix, realizes the combination of gene expression embedding and text embedding, dynamically adjusts the representation of the text data through the multi-head self-attention mechanism, captures the correlation between words in the text, and the relationship between the text and the gene expression latent space, ensures that the content in the text can reflect the characteristics of gene expression, and strengthens the understanding of gene expression by the text. By fusing the global embedding information of the gene expression matrix into the text encoding process, the gene-aware text encoder fully combines the background information of gene expression while capturing the text semantics during the execution of text processing tasks, and enhances the performance in gene-related tasks.

[0063] The above embodiment does not make any limitation on how to perform decoding processing by combining the information of the gene expression matrix during the generation or translation of text. Based on the above embodiment, the present invention may further include the following content:

[0064] Before splicing the gene global vector data to the previous text sequence, text-gene fusion data is obtained; the Transformer network model is used as the gene-aware text decoding module, and the text-gene fusion data is input into the gene-aware text decoding module for decoding processing through the gene-aware text decoding module.

[0065] In this embodiment, the previous text sequence is the text data generated in the previous decoding step or the partial text data already generated during the previous process of the text processing task. This text data can be the text data input by the user, for example, it can be information prompting the task type to be executed by the gene-aware text decoding module. In each decoding step, the global embedding of the gene expression matrix, that is, the gene global vector data, is fused with the previous text sequence and input into the gene-aware text decoding module, so that the generated text depends not only on the previously generated content but also on the gene expression. The initial input of the gene-aware text decoding module, that is, before the gene global vector data is generated, only includes the existing text data, that is, the previous text sequence, which can be expressed as , represents the embedding of the m-th word in the previous text sequence, and m represents the text length of the previous text sequence. In each decoding step, the gene global vector data is dynamically introduced to ensure that the generated text data can refer to the gene expression information at each decoding stage, and the gene information is always involved in each step of the decoding process to ensure that the generated text can contain the context related to the gene expression. Exemplarily, the text-gene fusion data of the gene global vector data and the previous text sequence can be expressed as . The gene-aware text decoding module can be constructed based on the standard Transformer decoder structure in advance. Similarly, the gene-aware text decoding module includes multiple Transformer blocks, and each Transformer block includes a multi-head attention layer based on the multi-head self-attention mechanism, a cross-attention layer based on the cross-attention mechanism, and a feed-forward network layer. Among them, the multi-head attention layer is used for information exchange in the generated text sequence, and the cross-attention layer is used for interacting the gene global vector data with the previous text sequence during the decoding process, so that the gene-aware text decoding module can dynamically generate appropriate text content according to the characteristics of the gene expression. Exemplarily, when the gene-aware text decoding module decodes at the i-th step, the word embedding data generated through multiple layers of Transformer calculations can be expressed as . According to the further decoding of the word embedding data generated in each decoding step of the gene-aware text decoding module, text data that not only contains the semantic information of the text but also combines the context information in the gene expression matrix is generated, and this text data has more biological significance.

[0066] Further, in order to improve the sequentiality and context consistency of the finally generated text, the pre-order text sequence may be processed first to generate text fragments, which may include the following:

[0067] Create an upper triangular matrix with the same length as the pre-order text sequence, and set the values to the right of the diagonal of the upper triangular matrix to infinity as mask information; perform a masking operation on each text word in the pre-order text sequence using the mask information; determine the attention weights based on the query matrix, key matrix, and value matrix of the pre-order text sequence, and determine the context vectors of each text word in the pre-order text sequence according to the attention weights; generate text fragments of the pre-order text sequence based on the context vectors of each text word, and fuse the text fragments with the gene global vector data.

[0068] In this embodiment, if the pre-order text sequence is not a vector that maps words to a high-dimensional space, the pre-order text sequence may be first converted into an embedding, that is, each word in the sentence is mapped to a vector in a high-dimensional space. For example, for the sentence "The sun rises in the east", a dictionary can be created to map each word to a unique integer index, and then an embedding layer is used to convert these indexes into vector form. In order to only consider all the words before each word when generating it, without relying on future information, the future words of each processed word in the input text can be masked, so that only the previous context is considered when predicting each new word. To implement causal self-attention, this embodiment can use an upper triangular matrix to mask future information. Exemplarily, a matrix of the same size as the sequence length can be created, and the values to the right of the diagonal are set to infinity or a very large negative number, so that before applying the softmax (activation function name) function, the attention scores at these positions will become very small, almost zero. After the masking operation, the attention scores can be calculated first, and then the final attention weights are obtained by normalizing the attention scores through the softmax function. The influence of each word in the pre-order text sequence when generating the next word is determined by the attention weights. The context vector of each word is calculated using the attention weights. This vector is the weighted sum of all word vectors, and the weights are determined by the attention weights. This context vector will be used to predict the next word.

[0069] Finally, the context vectors are used to generate text fragments. At each step, only the information available before the current step is used to generate the next word, thus ensuring that the generated text is coherent and conforms to the context.

[0070] Furthermore, after obtaining the text data output from each decoding step, the global embedding information of the gene expression matrix can be further combined through a cross-attention mechanism. Exemplarily, each text data is dynamically adjusted again according to the gene global vector data, so that the finally generated text not only conforms to the semantic context, but also can effectively reflect the influence of gene-related information, and can adaptively process the relationship between the text and gene expression, thereby more precisely capturing the potential association between the text and gene expression during the generation process and enhancing the biological relevance of the generation result.

[0071] Furthermore, to improve the decoding efficiency, a gene-aware text decoder can be pre-constructed in this embodiment. During the text generation process, the gene-aware text decoder adaptively combines the gene expression matrix by introducing gene expression embedding information, and dynamically selects key information during the text generation process through a cross-attention mechanism, so that the generated text can contain important content related to the gene expression matrix. The gene-aware text decoder can include a causal attention layer, a text-gene fusion layer, a gene-aware text decoding module, a cross-attention layer, and an output layer; among them, the causal attention layer processes the previous text sequence to generate text fragments, the text-gene fusion layer concatenates the gene global vector data in front of the text fragments to obtain text-gene fusion data, the gene-aware text decoding module decodes the text-gene fusion data, the cross-attention layer adjusts the output of each decoding step of the gene-aware text decoding module again according to the gene global vector data, and the output layer is used to output the text data adjusted by the cross-attention layer as the execution result of the text task to be processed.

[0072] As can be seen from the above, through the gene-aware text decoder in this embodiment, relevant words or sentence structures related to gene information can be adaptively selected during the text generation process, ensuring that the generated text is highly relevant to the gene data and enhancing the accuracy, interpretability, and biological relevance of the gene-related text generation task. Further, based on the causal self-attention mechanism, the previous generated text fragments that conform to the context are effectively processed, ensuring that each newly generated word is inferred only based on the previously generated content, thereby maintaining the sequentiality and context consistency of the generated text.

[0073] When analyzing gene expression data in related technologies, linear methods are usually adopted. However, traditional linear methods cannot fully express the non-linear relationship between text data and gene expression matrix, resulting in inaccurate gene expression analysis results. In addition, such linear methods cannot build a robust model in the face of the differences between high-dimensional gene expression data and natural language text under the conditions of data scarcity and noise interference. Based on the above embodiments, the present invention also provides a text task processing model for processing gene expression data and text data, which can include the following content:

[0074] Such asFigure 3 As shown, a text task processing model is pre-constructed; a text training sample set matching the text task to be processed is obtained, and the text task processing model is trained using the text training sample set until the model training stop condition is reached, obtaining a text task processing model for executing the text task to be processed. Correspondingly, the process of using this text task processing model to execute the text task to be processed can be: inputting text data and gene expression data into the text task processing model; the input layer of the text task processing model receives the text data and gene expression data; the gene expression matrix encoder encodes the gene expression data with added position learning parameters and global learning parameters into gene global vector data, the text encoder encodes the text data into text vector data, and the gene-aware text encoder encodes the fusion data of the gene global vector data and the text data into global text vector data; the decoding layer fuses the pre-order text sequence and the gene global vector data that already exist during the current decoding, and performs decoding processing on the text-gene fusion data to obtain the text data containing gene information corresponding to the current decoding step; the output layer determines the execution result of the text task to be processed based on the text data output by each decoding step and outputs it.

[0075] Among them, the text training sample set includes multiple groups of text-gene training samples, and the text-gene training samples include text samples and corresponding gene expression samples; for example, if the text task to be processed is a gene-text matching task, then the text-gene training sample is a set of samples where the text and the gene expression matrix are already matched; if the text task to be processed is a gene annotation task, then the text data in the text-gene training sample is the annotation information of the gene expression matrix, and the gene expression sample is the gene expression matrix. The model training stop condition can be, for example, that the number of iterations reaches a preset number of iterations, or that the model accuracy exceeds a preset accuracy threshold, which does not affect the implementation of the present invention. The text task processing model includes an input layer, an encoding layer, a decoding layer, and an output layer; the input layer inputs text samples and gene expression samples; the encoding layer includes a gene expression matrix encoder, a text encoder, and a gene-aware text encoder; the gene expression matrix encoder encodes the gene expression sample with added position learning parameters and global learning parameters into a gene global vector sample, the text encoder encodes the text sample into a text vector sample, and the gene-aware text encoder encodes the text sample spliced with the gene global vector sample into a global text vector sample; the decoding layer fuses the pre-order text sequence sample and the gene global vector sample that already exist during the current decoding, and performs decoding processing on the text-gene fusion sample to obtain the text sample data containing gene information corresponding to the current decoding. Based on this, the training of the text task processing model can be:

[0076] Obtain a corresponding number of text-gene training samples from the text training sample set according to the preset batch size, and input the obtained text-gene training samples into the text task processing model. The gene expression samples and text samples are respectively encoded by their respective encoders to obtain gene global embeddings and text global embeddings. In the gene-aware text generation task, the text decoder receives the gene embeddings as context information. Calculate the gene-text retrieval loss, gene-aware text encoding loss, and gene-aware text generation loss respectively, and combine them according to weights to form a joint loss function. Calculate the gradients of the model parameters of the text task processing model through the backpropagation algorithm, and use the optimizer to update the model parameters of the text task processing model. Continuously obtain text-gene training samples of a corresponding scale from the text training sample set and input them into the text task processing model to execute the above process. During the training of the text task processing model, the training effect of the model can be evaluated by the change of the loss value on the validation set. By separately evaluating the performance of different tasks, it can be judged whether the performance of the model on each subtask meets the expectations, and accordingly adjust the model structure or training parameters of the text task processing model. Continuously optimize the performance of each task through multiple iterations until the loss function converges and the text task processing model reaches the best performance.

[0077] The above embodiments do not make any limitations on the loss function of the text task processing model. In order to improve the performance of the text task processing model and enhance the task execution accuracy, based on the above embodiments, the present invention also provides a method for determining the loss function, which may include the following contents:

[0078] Pre-construct an initial loss function of the text task processing model, generate weight factors for the gene-word retrieval loss, gene-aware word encoding loss, and gene-aware word generation loss according to the task type of the text task to be processed; adjust the initial loss function according to the weight factors of the gene-word retrieval loss, gene-aware word encoding loss, and gene-aware word generation loss respectively to obtain the loss function; use the text training sample set to train the text task processing model based on the loss function.

[0079] In this embodiment, the initial loss function includes the gene-word retrieval loss, gene-aware word encoding loss, and gene-aware word generation loss. The gene-word retrieval loss is used to maximize the similarity between the text sample and its corresponding gene expression sample, while minimizing the similarity between the text sample and other gene expression samples. Thus, during the training of the text task processing model based on the gene-word retrieval loss, the similarity of each correct pair is maximized, and the similarity of irrelevant pairs is made as low as possible, which is applicable to the gene-text retrieval task. Exemplarily, the gene-word retrieval loss can be represented by the following relational expression:

[0080] ;

[0081] In the formula, represents the gene-text retrieval loss, Ei is the i-th gene expression sample, Ti is the text sample corresponding to the i-th gene expression sample, Tj is the j-th text sample not corresponding to the i-th gene expression sample, N represents the total number of text-gene training samples, f represents the similarity calculation function, and e represents the natural logarithm. The gene-text retrieval loss can maximize the similarity between the gene expression matrix Ei and its corresponding text description Ti through contrastive learning, while minimizing the similarity with the irrelevant text Tj, thus being more suitable for the gene-text retrieval task. Exemplarily, the gene-text retrieval loss strengthens the correct gene-text pairing by calculating the similarity f(Ei, Ti) between the gene expression matrix Ei and its corresponding text description Ti. At the same time, it suppresses the similarity f(Ei, Tj) of the incorrect pairing Ei and Tj. Further, the gene-text retrieval loss uses the natural logarithm to weight the similarity of the positive samples while reducing the contribution of the negative samples, so that the text task processing model learns the significance of the correct pairing. By optimizing the sample pairs in the entire training batch, it ensures that the text task processing model can effectively distinguish the correct gene-text relationship.

[0082] Among them, the gene-aware text encoding loss is used to evaluate the matching degree of the text task processing model for the input text sample and the gene expression sample. During the process of training the text task processing model based on the gene-aware text encoding loss, the text task processing model is optimized by minimizing it, so that it can better predict the matching relationship between the gene expression matrix and the text, accurately predict whether the gene expression matrix and the text match, ensure that the model can accurately associate the gene expression data with the disease-related text description, and has a relatively high accuracy for disease prediction. For example, in the task of predicting brain diseases, it can improve the accuracy of understanding the disease mechanism and diagnostic prediction. Exemplarily, the gene-aware text encoding loss function adopts the form of binary cross-entropy to calculate the prediction accuracy of whether the model matches the gene expression matrix Ei and the text Ti. The gene-aware text encoding loss can be represented by the following relational expression:

[0083] ;

[0084] In the formula, Denote the gene-aware text encoding loss as $L$, $N$ as the total number of text-gene training samples, $y_i$ as the actual matching label, that is, the label indicating whether the gene expression matrix $E_i$ and the text $T_i$ match. For example, $y_i = 1$ indicates that $E_i$ and $T_i$ are correctly matched, while $y_i = 0$ indicates that $E_i$ and $T_i$ are not matched. $P_i$ is the matching probability of $E_i$ and $T_i$ predicted by the text task processing model, which can represent the confidence of the text task processing model in predicting a match. When $y_i = 1$, the loss function encourages $p_i$ to approach 1, indicating that it is correctly predicted that $E_i$ and $T_i$ are matched; when $y_i = 0$, the loss function encourages $p_i$ to approach 0, indicating that the model correctly predicts that $E_i$ and $T_i$ do not match.

[0085] Among them, the gene-aware text generation loss is used to evaluate the ability of the text task processing model to generate text under the condition of gene expression context. During the process of training the text task processing model based on the gene-aware text generation loss, it measures the probability of the text task processing model predicting the next word when given the previous generated word and the gene expression matrix, ensuring that the text task processing model can comprehensively utilize gene expression information during the process of generating text and provide a more accurate and targeted text output. Exemplarily, the gene-aware text generation loss can represent the log-likelihood loss of the text task processing model in generating words at each step. By summing up the words in the entire generated sequence and minimizing the prediction error of each word in the generation process, it can be expressed by the following relational expression:

[0086] ;

[0087] In the formula, denotes the gene-aware text generation loss, is the previous generated word, that is, each word in the previous text sequence, $E$ is the gene expression sample, $T$ represents the length of the previous text sequence, and $t$ is the $t$-th word in the previous text sequence. This loss function is based on the generative task and can be based on Predict the probability of the next word with the gene expression sample E. The formula represents the log-likelihood loss of the model for generating words at each step. The gene-aware text generation loss emphasizes that the text task processing model should not only generate the next word based on the previous text but also combine the global information of the gene expression matrix, so that the generated text content can accurately reflect the characteristics and context of gene expression. For text generation tasks related to diseases, since the text finally generated by the text task processing model can precisely capture the potential relationship between gene expression and disease progression, it is beneficial to improve the accuracy corresponding to such text generation tasks and can generate text data containing more accurate disease-related information. During the training process of the text task processing model, by minimizing the gene-aware text generation loss, the text task processing model will learn how to generate coherent and meaningful text related to the gene expression background. Especially in the fields of medicine and bioinformatics, such generation tasks can assist medical researchers in automatically generating reports or diagnostic descriptions related to gene expression.

[0088] To ensure that the text processing model can handle gene-text retrieval, gene-text matching, and text generation tasks under gene conditions simultaneously, the loss function of the text processing model can adopt a joint loss function that comprehensively considers the loss functions of each task. By sharing representations, the generalization ability of each task can be improved. For example, while the model is performing gene-text retrieval, it is also learning how to better encode the matching relationship between gene expression and text, which is beneficial to enhancing the overall performance of the text task processing model. During the actual training process of the text processing model, these loss functions will act together to optimize the overall performance of the model. Exemplarily, the loss function L of the text processing model can be expressed as:

[0089] ;

[0090] Among them, λ1 is the weight factor of the gene text retrieval loss, λ2 is the weight factor of the gene-aware text encoding loss, and λ3 is the weight factor of the gene-aware text generation loss. The weight coefficients in front of each loss function can weight the losses according to the importance of the tasks, ensuring the performance balance of the text task processing model on different tasks. The joint loss function converts the training process of the model into a multi-task learning problem, and the model simultaneously learns multiple tasks and shares the global embedding information of the gene expression matrix. In the initial stage of training the text task processing model, uniform weights can be used, that is, λ1 = λ2 = λ3 = 1. As the text task processing model is continuously trained, the weights can be dynamically adjusted according to the requirements and performance of different tasks. For example, if the text task processing model performs poorly in the generation task, the weight of λ3 can be appropriately increased to focus on optimizing the generation task. In the study of brain diseases, the gene-text matching task may be more practically significant than the text generation task. Correspondingly, the weight λ1 of the gene text retrieval loss can be increased. The weight allocation should reflect the priorities of the tasks to ensure that the text task processing model can meet the requirements of different tasks.

[0091] As can be seen from the above, this embodiment combines transfer learning with background noise processing, enabling the text task processing task to still have good performance in the face of data scarcity or unlabeled situations. Further, the gene-text retrieval loss, gene-aware text encoding loss, and gene-aware text generation loss are jointly optimized to ensure that the text task processing model can simultaneously achieve excellent performance in multiple gene-aware tasks such as gene-text matching, text generation, and gene-conditioned generation, improving the understanding, matching, and generation capabilities of the gene expression matrix and text, and enhancing the overall learning efficiency and generalization ability between tasks.

[0092] In order to further improve the execution accuracy of the text task processing model for the gene-text matching task and more accurately identify whether the input gene expression data and text data match, based on the above embodiments, the present invention can also perform dynamic data augmentation processing on the gene expression samples and text samples included in the text training sample set to ensure the diversity of training data and the effectiveness of model training, which can include the following contents:

[0093] The text task processing model is initially trained using a text training sample set, and text sample-gene expression sample pairs and corresponding similarity scores generated during the initial training of the text task processing model are obtained. The similarity score is used as the confidence score for the text sample-gene expression sample pair; target text sample-gene expression sample pairs that meet the preset similarity conditions are selected from each text sample-gene expression sample pair; variant text sample-gene expression sample pairs of each target text sample-gene expression sample pair are obtained by adding fluctuation conditions and / or replacing similar text descriptions to each target text sample-gene expression sample pair; the text task processing model obtained by the initial training is retrained using each target text sample-gene expression sample pair and each variant text sample-gene expression sample pair.

[0094] In this embodiment, the text task processing model is preliminarily trained using gene-text matching pairs in the text training sample set so that it learns the correspondence between the gene expression matrix and the text description. During this process, the text task processing model generates a large number of gene-text pairing samples and corresponding similarity scores, and the similarity score of each gene-text pairing can be used as the corresponding confidence score. After the preliminary training is completed, high-quality gene-text pairing samples can be screened out through the similarity score. The preset similarity condition can be a pre-set similarity screening condition. The higher the similarity, the more reliable the pairing of the sample. For each pair of gene expression matrix and text description, the following relational expression can be used to calculate the similarity value between the two, and a similarity threshold is set. The similarity threshold can take a value between 0.7 and 0.9, for example. If the gene expression matrix and the text description has a similarity value greater than the similarity threshold, it is considered to meet the preset similarity condition. Only gene-text pairings with a similarity higher than the threshold are retained, and low-similarity pairing samples are excluded to ensure that only high-quality training data is used in the data augmentation process. Among them, the calculation of the similarity value can be implemented by calling the following relational expression:

[0095] .

[0096] After screening out high-quality gene-text pairing samples, for the sake of convenience of description, these samples are defined as target text sample-gene expression sample pairs. For each target text sample-gene expression sample pair, new gene expression matrix samples can be generated by simulating gene expression under different experimental conditions, by randomly perturbing or adding noise. For example, for the target text sample-gene expression sample pair, a slight perturbation can be performed according to , and multiple variant samples are generated by randomly adding noise or other small-scale transformations such as scaling and translation, so as to obtain more similar but slightly different gene expression matrices. Among them, , is the standard deviation of the noise, which is usually set to be small to maintain the biological rationality of gene expression.

[0097] On this basis, the data diversity can also be increased for the target text sample - gene expression sample pair by replacing synonyms in the text. For the text descriptions matching the gene expression matrix, multiple different text versions can be generated by using methods such as synonym replacement, sentence structure change, and word order adjustment, ensuring that the enhanced text still accurately expresses gene - related content but has a certain degree of language diversity. For example, if the target text sample of the target text sample - gene expression sample is: This gene is highly expressed under condition X, it can be replaced with: Under condition X, this gene shows a high expression level to generate a new text sample. After generating new samples through the above - mentioned data enhancement, the enhanced gene - text paired samples are combined with the original training data to form a more abundant data set. These high - quality enhanced data are re - input into the text task processing model for training to further improve the performance of the model in the gene expression matrix and text pairing task, ensuring that it performs more robustly when facing complex gene - text matching tasks.

[0098] As can be seen from the above, in this embodiment, for the high - quality gene - text paired samples screened from the text training sample set, data enhancement processing is performed on these high - quality samples from the perturbation of the gene expression matrix and the diversification of text data to generate different gene expression matrices and diverse text samples, and the text task processing model is repeatedly trained to improve the robustness of the text task processing model. Through multiple iterative optimizations, it is ensured that the text task processing model can still show high robustness under data - scarce and high - noise conditions and achieve better generalization performance in complex bioinformatics tasks.

[0099] The present invention also provides a corresponding device for the text task processing method, further making the method more practical. Among them, the device can be described from the perspective of functional modules and the perspective of hardware respectively. The text task processing device provided by the present invention will be introduced below. This device is used to implement the text task processing method provided by the present invention. In this embodiment, the text task processing device can include or be divided into one or more program modules. These one or more program modules are stored in a storage medium and are executed by one or more processors to complete the text task processing method disclosed in Embodiment 1. The program modules referred to in this embodiment refer to a series of computer program instruction segments that can complete specific functions, and are more suitable for describing the execution process of the text task processing device in the storage medium than the program itself. The following description will specifically introduce the functions of each program module in this embodiment. The text task processing device described below can be correspondingly referred to the text task processing method described above.

[0100] From the perspective of functional modules, refer to Figure 4 , Figure 4 which is the structural diagram of the text task processing device provided in this embodiment under a specific implementation manner. The device may include:

[0101] A data acquisition module 401, configured to acquire text data and gene expression data corresponding to a text task to be processed; the text task to be processed is a text task for generating a text containing gene information.

[0102] An encoding module 402, configured to encode the gene expression data added with position learning parameters and global learning parameters into gene global vector data. Encode the text data into text vector data according to the text task to be processed and / or encode the fusion data of the gene global vector data and the text data into global text vector data. Among them, the position learning parameter is used to learn the sequential feature information of the gene expression data, and the global learning parameter is used to learn the global feature of the gene expression data.

[0103] A decoding module 403, configured to fuse the pre-order text sequence that already exists during the current decoding and the gene global vector data, and perform decoding processing on the text-gene fusion data to obtain the text data containing gene information corresponding to the current decoding step.

[0104] A result output module 404, configured to determine the execution result of the text task to be processed according to the text data output by each decoding step.

[0105] Exemplarily, in some implementation manners of this embodiment, the above encoding module 402 may further be configured to: divide the gene expression data into multiple sub-gene expression matrices, and add position learning parameters to each sub-gene expression matrix; add global learning parameters in front of the gene vector sequence data corresponding to each sub-gene expression matrix added with position learning parameters as the gene expression matrix feature; input the gene expression matrix feature into a gene expression matrix encoding module to perform encoding processing on the gene expression matrix feature by using the gene expression matrix encoding module; the input position of the gene expression matrix encoding module includes a classification tag; use the output corresponding to the classification tag of the gene expression matrix encoding module as the gene global vector data.

[0106] Exemplarily, in some other implementation manners of this embodiment, the above encoding module 402 may further be configured to: perform word segmentation processing on the text data to obtain multiple text words; input each text word into a text encoding module to perform encoding processing by using the text encoding module; the input position of the text encoding module includes a classification tag; use the output corresponding to the classification tag of the text encoding module as the text vector data.

[0107] Exemplarily, in some other embodiments of this embodiment, the above encoding module 402 can also be used to: perform word segmentation on the text data to obtain an initial text embedding sequence; splice the gene global vector data in front of the initial text embedding sequence to obtain fused vector data; input the fused vector data into the gene-aware text encoding module to use the gene-aware text encoding module to perform encoding processing on the fused vector data; use the output of the gene-aware text encoding module as the global text vector data.

[0108] Exemplarily, in some other embodiments of this embodiment, the above encoding module 402 can also be used to: encode the fusion data of the gene global vector data and the text data into global text vector data, including: performing word segmentation on the text data to obtain an initial text embedding sequence; inputting the initial text embedding sequence and the gene global vector data into the gene-aware text encoding module to use the gene-aware text encoding module to perform interactive calculations on the query matrix of each text word in the initial text embedding sequence and the key matrix and value matrix of the gene global vector data, and using the interactive calculation result as the fused vector data for encoding processing; using the output of the gene-aware text encoding module as the global text vector data.

[0109] Exemplarily, in some other embodiments of this embodiment, the above decoding module 403 can also be used to: splice the gene global vector data in front of the previous text sequence to obtain text-gene fusion data; use the Transformer network model as the gene-aware text decoding module, and input the text-gene fusion data into the gene-aware text decoding module to perform decoding processing through the gene-aware text decoding module.

[0110] As an exemplary embodiment of the above embodiment, the above decoding module 403 can further be used to: create an upper triangular matrix with the same length as the previous text sequence, and set the values on the right side of the diagonal of the upper triangular matrix to infinity to serve as mask information; use the mask information to perform a masking operation on each text word in the previous text sequence; determine the attention weights according to the query matrix, key matrix, and value matrix of the previous text sequence, and determine the context vectors of each text word in the previous text sequence according to the attention weights; generate a text segment of the previous text sequence according to the context vectors of each text word, and fuse the text segment and the gene global vector data.

[0111] Exemplarily, in some other embodiments of this embodiment, the above device may further include, for example, a model training module, which is used to pre-construct a text task processing model; the text task processing model includes an input layer, an encoding layer, a decoding layer, and an output layer; obtain a text training sample set that matches the text task to be processed, the text training sample set includes multiple groups of text-gene training samples, and the text-gene training sample includes a text sample and a corresponding gene expression sample; use the text training sample set to train the text task processing model until the model training stop condition is reached, and obtain a text task processing model for executing the text task to be processed; wherein, the input layer inputs the text sample and the gene expression sample, and the encoding layer includes a gene expression matrix encoder, a text encoder, and a gene-aware text encoder; the gene expression matrix encoder encodes the gene expression sample with added position learning parameters and global learning parameters into a gene global vector sample, the text encoder encodes the text sample into a text vector sample, and the gene-aware text encoder encodes the text sample spliced with the gene global vector sample into a global text vector sample; the decoding layer fuses the pre-order text sequence sample and the gene global vector sample that already exist during the current decoding, and decodes the text-gene fusion sample to obtain the text sample data containing gene information corresponding to the current decoding.

[0112] As an exemplary implementation of the above embodiment, the above device may further include, for example, a task execution module, which can be used to input text data and gene expression data into the text task processing model; the input layer of the text task processing model receives the text data and the gene expression data; the gene expression matrix encoder encodes the gene expression data with added position learning parameters and global learning parameters into a gene global vector data, the text encoder encodes the text data into a text vector data, and the gene-aware text encoder encodes the fusion data of the gene global vector data and the text data into a global text vector data; the decoding layer fuses the pre-order text sequence and the gene global vector data that already exist during the current decoding, and decodes the text-gene fusion data to obtain the text data containing gene information corresponding to the current decoding step; the output layer determines the execution result of the text task to be processed according to the text data output by each decoding step, and outputs it.

[0113] As an exemplary implementation of the above embodiment, the above model training module may further be configured to: pre-construct an initial loss function of the text task processing model; the initial loss function includes a gene text retrieval loss, a gene-aware text encoding loss, and a gene-aware text generation loss; generate weight factors for the gene text retrieval loss, the gene-aware text encoding loss, and the gene-aware text generation loss according to the task type of the text task to be processed; adjust the initial loss function according to the respective weight factors of the gene text retrieval loss, the gene-aware text encoding loss, and the gene-aware text generation loss to obtain a loss function; use the text training sample set to train the text task processing model based on the loss function; wherein, the gene text retrieval loss is used to maximize the similarity between a text sample and its corresponding gene expression sample while minimizing the similarity between the text sample and other gene expression samples; the gene-aware text encoding loss is used to evaluate the matching degree of the text task processing model to the input text sample and gene expression sample; the gene-aware text generation loss is used to evaluate the ability of the text task processing model to generate text under the condition of the gene expression context.

[0114] As another exemplary implementation of the above embodiment, the above model training module may further be configured to: perform initial training on the text task processing model using the text training sample set, obtain text sample-gene expression sample pairs and corresponding similarity scores generated during the initial training of the text task processing model, and use the similarity scores as confidence scores for the text sample-gene expression sample pairs; select target text sample-gene expression sample pairs that meet a preset similarity condition from the text sample-gene expression sample pairs; obtain variant text sample-gene expression sample pairs of the target text sample-gene expression sample pairs by adding fluctuation conditions and / or replacing similar text descriptions to the target text sample-gene expression sample pairs; use the target text sample-gene expression sample pairs and the variant text sample-gene expression sample pairs to retrain the text task processing model obtained from the initial training.

[0115] The text task processing device mentioned above is described from the perspective of functional modules. Further, the present invention also provides an electronic device, which is described from the perspective of hardware. Figure 5 FIG. is a schematic structural diagram of the electronic device provided by the embodiment of the present invention in one implementation manner. As Figure 5 shown, the electronic device includes a memory 50 for storing a computer program; a processor 51 for implementing the steps of the text task processing method as mentioned in any of the above embodiments when executing the computer program.

[0116] Among them, the processor 51 may include one or more processing cores, such as a 4-core processor or an 8-core processor. The processor 51 may also be a controller, a microcontroller, a microprocessor, or other data processing chips, etc. The processor 51 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor 51 may also include a main processor and a coprocessor. The main processor is a processor used to process data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the processor 51 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for the rendering and drawing of the content to be displayed on the display screen. In some embodiments, the processor 51 may also include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.

[0117] The memory 50 may include one or more computer non-volatile storage media, which may be non-transitory. The memory 50 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. The memory 50 may be an internal storage unit of an electronic device in some embodiments, such as the hard disk of a server. The memory 50 may also be an external storage device of an electronic device in other embodiments, such as a plug-in hard disk equipped on a server, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 50 may also include both an internal storage unit and an external storage device of an electronic device. The memory 50 can be used not only to store application software installed on the electronic device and various types of data, such as the code of a program during the execution of a text task processing method, etc., but also to temporarily store data that has been output or will be output. In this embodiment, the memory 50 is at least used to store the following computer program 501. After being loaded and executed by the processor 51, the computer program can implement the relevant steps of the text task processing method disclosed in any of the foregoing embodiments. Additionally, the resources stored in the memory 50 may also include an operating system 502 and data 503, etc., and the storage method can be temporary storage or permanent storage. Among them, the operating system 502 may include Windows, Unix, Linux, etc. The data 503 may include, but is not limited to, data corresponding to the text task processing results, etc.

[0118] In some embodiments, the electronic device may further include a display screen 52, an input / output interface 53, a communication interface 54 or a network interface, a power supply 55 and a communication bus 56. Among them, the display screen 52 and the input / output interface 53, such as a keyboard, belong to the user interface, and exemplary user interfaces may also include standard wired interfaces, wireless interfaces, etc. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, and an OLED (Organic Light-Emitting Diode) touch device, etc. The display may also be appropriately referred to as a display screen or a display unit, which is used to display information processed in the electronic device and to display a visual user interface. The communication interface 54 may exemplarily include a wired interface and / or a wireless interface, such as a WI-FI interface, a Bluetooth interface, etc., which are generally used to establish a communication connection between the electronic device and other electronic devices. The communication bus 56 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 5 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.

[0119] Those skilled in the art will understand that Figure 5 The structure shown in the figure does not constitute a limitation on the electronic device, and may include more or fewer components than those shown in the figure, for example, it may also include a sensor 57 for realizing various functions.

[0120] It can be understood that if the text task processing method in the above embodiments is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a non-volatile storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the related technology, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and executes all or part of the steps of the methods in the various embodiments of the present invention. The foregoing storage medium includes, but is not limited to: USB flash drive, mobile hard disk, read-only memory (ROM), random access memory (RAM), electrically erasable programmable ROM, register, hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), magnetic memory, removable disk, CD-ROM, magnetic disk or optical disk, etc., various media that can store program codes. Based on this, the present invention also provides a non-volatile storage medium, on which a computer program is stored. When the computer program is executed by a processor, it performs the steps of the text task processing method described in any one of the above embodiments.

[0121] It can be understood that if the text task processing method in the above embodiments is implemented in the form of a software functional unit and sold or used as an independent product, this computer software product may not need to be stored in a physical storage medium. For example, it can be directly transmitted to a device with information processing capabilities, such as a computer, through a wired network or a wireless network to execute all or part of the steps of the methods in the various embodiments of the present invention. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the related technology, or all or part of this technical solution, can be embodied in the form of a software product. Based on this, the present invention also provides a computer program product, storing a computer program. When the computer program is executed by a processor, it performs the steps of the text task processing method described in any one of the above embodiments.

[0122] Finally, the present invention also provides a text task processing system. Please refer to Figure 6, the text task processing system may include a user interface 601 at the front end and a text task processor 602 at the back end. The user interface 601 at least includes a task trigger area, a task processing result display area, a data input area, and / or a file import area to enable the user to describe the text task to be processed through the data input area, input text data and gene expression data through the data input area or the file import area, and send the text task to be processed, the text data, and the gene expression data to the text task processor through the task trigger area; the text task processor 602 processes the text data and the gene expression data based on the text task to be processed, using the steps of the text task processing method described in any of the above embodiments, generates a text task processing result, and sends the text task processing result to the task processing result display area of the user interface. The text task processing result can be directly displayed in text, or in image form, or it can also be a file, and at the same time, identification information for the user to download is displayed. The file import area can import the corresponding data by uploading a file. The data input area may include, for example, a text input area and a voice input area, and the user can choose to input the corresponding data in text or voice form.

[0123] In this embodiment, the text task processing system can be, for example, a web page or an application package. The user side can pre-deploy the text task processing system and display the user page of the text task processing system through the human-computer interaction page of the user side. The user operates on this user page and feeds back the finally generated text data on the user page.

[0124] The above has introduced in detail a text task processing method, apparatus, electronic device, non-volatile storage medium, and computer program product provided by the present invention. Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. Whether the units and algorithm steps of each example described in the disclosed embodiments are executed in the form of electronic hardware or computer software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, and such implementation should not be considered to exceed the scope of the present invention. Without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the present invention.

Claims

1. A text task processing method, characterized in that: include: Pre-constructing a text task processing model, obtaining text data and gene expression data corresponding to the text task to be processed, and inputting the text data and the gene expression data into the text task processing model; the text task to be processed is a task of generating a text containing gene information; Encoding the gene expression data with added positional learning parameters and global learning parameters into gene global vector data; Encoding the text data into text vector data according to the text task to be processed, and / or encoding the fusion data of the gene global vector data and the text data into global text vector data; The preamble text sequence already existing in the current decoding is merged with the gene global vector data, and the text gene fusion data is decoded to obtain the text data containing gene information corresponding to the current decoding step; Determining the execution result of the text task to be processed according to the text data outputted by each decoding step; Among them, the position learning parameters are used to learn the sequential feature information of the gene expression data, and the global learning parameters are used to learn the global features of the gene expression data; the generation process of the gene global vector data includes: dividing the gene expression data into multiple sub-gene expression matrices, and adding position learning parameters to each sub-gene expression matrix; adding global learning parameters before the gene vector sequence data corresponding to each sub-gene expression matrix to which the position learning parameters are added, as gene expression matrix features; inputting the gene expression matrix features into a gene expression matrix encoding module, so as to encode the gene expression matrix features using the gene expression matrix encoding module; the input position of the gene expression matrix encoding module includes a classification mark; and using the output corresponding to the classification mark of the gene expression matrix encoding module as the gene global vector data.

2. The text task processing method according to claim 1, characterized in that: Encoding the text data into text vector data includes: Performing word segmentation processing on the text data to obtain a plurality of text words; Inputting each text word into a text encoding module to perform encoding processing using the text encoding module; the input position of the text encoding module includes a classification mark; The output corresponding to the classification mark of the text encoding module is used as text vector data.

3. The text task processing method according to claim 1, characterized in that: Encoding the fusion data of the gene global vector data and the text data into global text vector data includes: Performing word segmentation processing on the text data to obtain an initial text embedding sequence; splicing the gene global vector data to before the initial text embedding sequence to obtain fused vector data; Inputting the fused vector data into a gene-aware text encoding module, so as to encode the fused vector data using the gene-aware text encoding module; The output of the gene-aware text encoding module is used as global text vector data.

4. The text task processing method according to claim 1, characterized in that: Encoding the fusion data of the gene global vector data and the text data into global text vector data includes: Performing word segmentation processing on the text data to obtain an initial text embedding sequence; Inputting the initial text embedding sequence and the gene global vector data into a gene-aware text encoding module, using the gene-aware text encoding module to interactively calculate the query matrix of each text word of the initial text embedding sequence with the key matrix and the value matrix of the gene global vector data, and encoding the interactive calculation result as fused vector data; The output of the gene-aware text encoding module is used as global text vector data.

5. The text task processing method according to claim 1, characterized in that: The preamble text sequence that already exists during the current decoding is merged with the gene global vector data, and the text gene fusion data is decoded to obtain text data containing gene information corresponding to the current decoding step, including: Splicing the gene global vector data to the front of the preceding text sequence to obtain text gene fusion data; The converter network model is used as a gene-aware text decoding module, and the text gene fusion data is input into the gene-aware text decoding module for decoding processing by the gene-aware text decoding module.

6. The text task processing method according to claim 5, characterized in that: The preamble text sequence that already exists during the current decoding is merged with the gene global vector data, including: Create an upper triangular matrix with the same length as the preceding text sequence, and set the value on the right side of the diagonal of the upper triangular matrix to infinity as mask information; Using the mask information, a mask operation is performed on each text word of the preceding text sequence; Determining attention weights according to a query matrix, a key matrix, and a value matrix of the preceding text sequence, and determining a context vector of each text word of the preceding text sequence according to the attention weights; A text segment of the preceding text sequence is generated according to the context vector of each text word, and the text segment is fused with the gene global vector data.

7. The text task processing method according to any one of claims 1 to 6, characterized in that: Also includes: The text task processing model includes an input layer, an encoding layer, a decoding layer and an output layer; Acquire a text training sample set that matches the text task to be processed, wherein the text training sample set includes multiple groups of text-gene training samples, and the text-gene training samples include text samples and corresponding gene expression samples; Using the text training sample set, training the text task processing model until a model training stop condition is reached, thereby obtaining a text task processing model for executing the text task to be processed; Wherein, the input layer inputs the text sample and the gene expression sample; The encoding layer includes a gene expression matrix encoder, a text encoder and a gene-aware text encoder; the gene expression matrix encoder encodes the gene expression sample with added position learning parameters and global learning parameters into a gene global vector sample, the text encoder encodes the text sample into a text vector sample, and the gene-aware text encoder encodes the text sample spliced ​​with the gene global vector sample into a global text vector sample; The decoding layer fuses the preamble text sequence sample and the gene global vector sample that already exist during the current decoding, and performs decoding processing on the text gene fusion sample to obtain text sample data containing gene information corresponding to the current decoding.

8. The text task processing method according to claim 7, characterized in that: After obtaining the text data and gene expression data corresponding to the text task to be processed, it also includes: The input layer of the text task processing model receives the text data and the gene expression data; the gene expression matrix encoder encodes the gene expression data with added position learning parameters and global learning parameters into gene global vector data, the text encoder encodes the text data into text vector data, and the gene-aware text encoder encodes the fusion data of the gene global vector data and the text data into global text vector data; the decoding layer fuses the preamble text sequence that already exists during the current decoding with the gene global vector data, and decodes the text-gene fusion data to obtain the text data containing gene information corresponding to the current decoding step; the output layer determines the execution result of the text task to be processed according to the text data output by each decoding step, and outputs it.

9. The text task processing method according to claim 7, characterized in that: Using the text training sample set to train the text task processing model includes: Pre-constructing an initial loss function of the text task processing model; the initial loss function includes gene text retrieval loss, gene-aware text encoding loss and gene-aware text generation loss; According to the task type to which the to-be-processed text task belongs, generating weight factors of the gene-aware text retrieval loss, the gene-aware text encoding loss, and the gene-aware text generation loss; According to the respective weight factors of the gene-aware word retrieval loss, the gene-aware word encoding loss and the gene-aware word generation loss, the initial loss function is adjusted to obtain a loss function; Using the text training sample set, training the text task processing model based on the loss function; Among them, the gene-text retrieval loss is used to maximize the similarity between the text sample and its corresponding gene expression sample while minimizing the similarity between the text sample and other gene expression samples; the gene-aware text encoding loss is used to evaluate the degree of matching between the input text sample and the gene expression sample by the text task processing model; the gene-aware text generation loss is used to evaluate the ability of the text task processing model to generate text under the condition of gene expression context.

10. The text task processing method according to claim 7, characterized in that: The text task to be processed is a gene-text matching task, and the text task processing model is trained using the text training sample set, including: Performing initial training on the text task processing model using the text training sample set, obtaining text sample-gene expression sample pairs and corresponding similarity scores generated by the text task processing model during the initial training process, and using the similarity scores as confidence scores for the text sample-gene expression sample pairs; Selecting a target text sample-gene expression sample pair that meets a preset similarity condition from each text sample-gene expression sample pair; By adding fluctuation conditions and / or replacing similar text descriptions for each target text sample-gene expression sample pair, a variant text sample-gene expression sample pair of each target text sample-gene expression sample pair is obtained; The text task processing model obtained by the initial training is trained again using each target text sample-gene expression sample pair and each variant text sample-gene expression sample pair.

11. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the processor is used to implement the steps of the text task processing method according to any one of claims 1 to 10 when executing a computer program stored in the memory.

12. A non-volatile storage medium, characterized in that: The non-volatile storage medium stores a computer program, and when the computer program is executed by the processor, the steps of the text task processing method according to any one of claims 1 to 10 are implemented.

13. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, the steps of the text task processing method described in any one of claims 1 to 10 are implemented.

14. A text task processing system, characterized in that: Includes user interface and text task processor; The user interface at least comprises a task triggering area, a task processing result displaying area, a data input area and / or a file importing area, so that a user can describe a text task to be processed through the data input area, input text data and gene expression data through the data input area or the file importing area, and send the text task to be processed, the text data and the gene expression data to the text task processor through the task triggering area; The text task processor, based on the text task to be processed, processes the text data and the gene expression data using the steps of the text task processing method as described in any one of claims 1 to 10, generates a text task processing result, and sends the text task processing result to the task processing result display area of ​​the user interface.

Citation Information

Patent Citations

  • Text processing method and device

    CN110210032A

  • Text processing method and device, electronic equipment and storage medium

    CN118069814A