Metabolism-related fatty liver disease prediction method based on self-attention mechanism, control server and medium
By adopting a prediction method based on self-attention mechanism in MAFLD screening, combining blood tests and traditional Chinese medicine data, the problem of ignoring traditional Chinese medicine indicators and difficulty in dealing with massive heterogeneous data in the existing technology is solved, and a more accurate and lower-cost MAFLD screening effect is achieved.
Patent Information
- Application Number
- CN202510148910.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art is difficult to effectively process massive heterogeneous medical data in metabolic-related fatty liver disease (MAFLD) screening, ignoring the complex dependence between patient sequence data, and traditional models ignore the important role of traditional Chinese medicine indicators, resulting in poor screening results.
A fatty liver disease prediction method based on self-attention mechanism is adopted, combined with blood test data and traditional Chinese medicine data (such as tongue and pulse), and by increasing timing information encoding and using multi-scale convolution and stratified attention mechanisms, a fatty liver detection model is constructed for training to obtain more accurate prediction results.
It improves the accuracy and prediction accuracy of MAFLD screening, effectively integrates traditional Chinese medicine data, reduces screening costs, and can more comprehensively analyze patient stage indicators to achieve a more accurate diagnosis of MAFLD.
Smart Images

Figure CN120072303A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of medicine and computer science, and particularly relates to a method for predicting metabolic associated fatty liver disease based on self-attention mechanism, a control server and a medium. Background Art
[0002] Metabolic associated fatty liver disease (MAFLD) is a common chronic liver disease, which is closely related to metabolic diseases such as obesity, insulin resistance, and type 2 diabetes. Due to the complex pathogenesis of MAFLD and the non-linear relationship between its disease progression and multiple physiological indicators, early screening and risk assessment are of great significance for the prognosis of patients. However, traditional screening methods usually rely on doctors' clinical experience and rule-based risk assessment models, which often ignore the complex dependence relationships between patient sequence data and are difficult to effectively process massive heterogeneous medical data, thus limiting their application effects in MAFLD screening.
[0003] With the rapid development of artificial intelligence technology, especially deep learning, risk prediction and disease screening in the medical field have gradually shifted from traditional feature engineering-based models to data-driven methods. For example, Pei-Yuan Su et al. used machine learning models to predict the incidence probability of fatty liver under normal weight; Aylin Tahmasebi et al. used supervised machine learning models to distinguish between alcoholic fatty liver and non-alcoholic fatty liver; Hong-Ye Peng et al. used a method combining LASSO regression and random forest to screen important predictors of non-alcoholic fatty liver risk.
[0004] However, current disease screening models mainly rely on convolutional neural networks or structures based on BP neural networks. Although they have achieved certain success in some fields, these models face multiple challenges when dealing with screening tasks for complex diseases such as MAFLD.
[0005] First, most existing models tend to model single physiological indicators and ignore the temporal correlation between multiple physiological signals of patients, which may lead to misjudgment of the metabolic changes of MAFLD patients.
[0006] Second, most existing models ignore the important role of traditional Chinese medicine indicators in MAFLD screening, especially indicators such as tongue image and pulse condition, which have a significant impact on the diagnosis of MAFLD.
[0007] In addition, most existing models prefer to use blood test data or imaging data such as elastography ultrasound and CT. Such methods not only result in a single model inference environment, but also increase the medical costs for patients, and more expensive examinations are required to use such models for inference. Summary of the Invention
[0008] To solve the above problems, the present invention provides a method for predicting metabolic associated fatty liver disease based on self-attention mechanism, which can obtain more objective, lower-cost and more accurate prediction results of metabolic associated fatty liver disease based on the integrated traditional Chinese and Western medicine data.
[0009] To achieve the above object, the technical solution adopted by the present invention is as follows: A method for predicting metabolic associated fatty liver disease based on self-attention mechanism, comprising the following steps:
[0010] Obtain the patient's blood test and four diagnostic instruments data, and construct a fatty liver disease data set;
[0011] For the same patient's same index data in different periods, add temporal information encoding to form an optimized fatty liver disease data set with sequence features.
[0012] Construct a fatty liver detection model and train it. The fatty liver detection model includes a convolutional neural network, a hierarchical attention module, and a feed-forward neural network layer;
[0013] Input the patient data into the trained fatty liver detection model to obtain the prediction grade of the patient's metabolic associated fatty liver disease.
[0014] As can be seen from the above technical solution of the present invention, compared with the existing screening technologies, the significant advantages of the analysis method provided by the present invention are: (1) Based on the blood test data, the model deeply integrates traditional Chinese medicine data such as tongue image and pulse condition, effectively improving the accuracy of the model; (2) By adding learnable position encoding, the sequence features of the same patient's same index in different periods are obtained, realizing a comprehensive analysis of the patient's stage indicators; (3) Through multi-scale convolution and hierarchical attention, fine-grained reasoning of the patient's time series indicators is realized, improving the prediction accuracy of the model. Brief Description of the Drawings
[0015] Figure 1 is a flow chart of the method for screening metabolic associated fatty liver disease in an embodiment of the present invention.
[0016] Figure 2 is a schematic diagram of multi-scale convolution in an embodiment of the present invention.
[0017] Figure 3 is a schematic diagram of hierarchical attention in an embodiment of the present invention. Detailed Description of the Invention
[0018] Example 1, as Figure 1 shown, according to a preferred embodiment of the present invention, a method for predicting metabolic associated fatty liver disease based on self-attention mechanism is used to predict the patient's index data, and the implementation of the above solution will be described in detail in combination with Figure 1 shown.
[0019] Step 1: Obtain the patient's blood test and four diagnostic instrument data, and construct a relevant fatty liver disease dataset for training the fatty liver detection model, as follows:
[0020] The dataset includes two parts: blood test data and four diagnostic instrument data. The blood test indicators include: gender, age, basophil count, hemoglobin, platelet count, neutrophil count, red blood cell count, lymphocyte percentage, lymphocyte absolute value, white blood cell count, monocyte count, hematocrit, eosinophil count, total bilirubin, total protein, albumin, direct bilirubin, glutamyl transpeptidase, alkaline phosphatase, alanine aminotransferase, aspartate aminotransferase, creatinine, uric acid, urea, total cholesterol, triglyceride, high-density lipoprotein cholesterol, low-density lipoprotein cholesterol, fasting blood glucose, alpha-fetoprotein, carcinoembryonic antigen, BMI value. All indicators are stored in digital form.
[0021] The four diagnostic instrument data include: tongue color, redness at the edge and tip of the tongue, ecchymosis and petechiae, tongue coating color, thickness of the tongue coating, greasiness of the tongue coating, curdiness of the tongue coating, exfoliation of the tongue coating, fatness and thinness of the tongue, tooth marks, prickles, cracks, facial color, flushing of the zygomatic regions, dark circles under the eyes, luster, lip color, average heart rate, pericardium meridian, liver meridian, kidney meridian, spleen meridian, lung meridian, stomach meridian, gallbladder meridian, bladder meridian. All indicators are stored in digital form.
[0022] Finally, the dataset is first segmented using the patient ID as the index to form the index vectors of different patients, expressed as follows:
[0023] P(n) = [XHDB(n), XXBJS(n), SJLXBJS(n), ……, PGJ(n)]
[0024] In the formula, P(n) represents the index matrix of the nth patient, XHDB(n) represents the hemoglobin index, XXBJS(n) represents the platelet count index, SJLXBJS(n) represents the specific index of basophil count, and PGJ(n) represents the bladder meridian index.
[0025] Among them, each index is represented separately according to different collection times to form an index vector, expressed as follows:
[0026] XHDB(n) = [xhdb(1), xhdb2), xhdb(3) …… xhdb(m)]
[0027] In the formula, XHDB(n) represents the hemoglobin index of the nth patient, xhdb(1) represents the hemoglobin value collected for the first time by the nth patient, xhdb(2) represents the hemoglobin value collected for the second time by the nth patient, and xhdb(m) represents the hemoglobin value collected for the mth time by the nth patient.
[0028] Step 2: Add temporal information encoding to the dataset, as follows:
[0029] Encode the collected data with temporal information to form an optimized fatty liver disease dataset with sequence features. For the same indicator of the same patient at different times, add sequence information through trigonometric functions. The expression of positional encoding is as follows:
[0030]
[0031]
[0032] In the formula, w represents the learnable parameter, pos represents the position in the sequence (from 0 to the sequence length minus 1), and i represents the dimension index of the vector. The sine function is used for even dimensions and the cosine function is used for odd dimensions.
[0033] By adding the positional encoding to the original indicator values, a dataset with time series information is obtained, and the expression is as follows:
[0034] XHDBP(n) = [xhdb(1) + PE(1), Xhdb(2) + PE(2), Xhdb(3) + PE(3)…Xhdb(n) + PE(n)]
[0035] XXBJS(n) = [xxbjs(1) + PE(1), xxbjs(2) + PE(2), xxbjs(3) + PE(3)…xxbjs(n) + PE(n)]
[0036] PGJ(n) = [pgj(1) + PE(1), pgj(2) + PE(2), pgj)3) + PE(3)…pgj(n) + PE(n)]
[0037] In the formula, XHDBP(n) represents the hemoglobin indicator of the nth patient with time series information, XXBJS(n) represents the platelet count indicator of the nth patient with time series information, PGJ(n) represents the bladder meridian indicator of the nth patient with time series information, and so on for other indicators.
[0038] Step 3: Construct and train a fatty liver detection model. The fatty liver detection model includes an improved convolutional neural network, a hierarchical attention module, and a feedforward neural network layer. The training process includes:
[0039] (1) Use a convolutional neural network to extract features from the time series data in the fatty liver disease dataset. After downsampling the extracted features, unify the scales and splice them to obtain a multi-scale feature matrix. Taking the hemoglobin indicator XHDBP(n) as an example;
[0040] 1) Convolve the indicator with time series information through a multi-scale convolutional kernel, and the expression is as follows:
[0041] K belongs to {K 1 , K 2 , …, K m}
[0042]
[0043] In the formula, K represents the set of convolution kernels, and K m represents the m-th convolution kernel, and XHDBP e represents the e-th hemoglobin index with time series information, n is the number of hemoglobin indexes with time series information, m represents the number of convolution kernels, and K c represents the c-th convolution kernel, and Y(n) represents the convolution features obtained after convolution calculation;
[0044] 2) For multi-scale features, continuously downsample from the calculation result of the first convolution kernel to the calculation result of the last convolution kernel, and splice them after dimension unification to obtain a multi-scale feature matrix, which is expressed as follows:
[0045]
[0046] D = concat(Downsample_Y 0 , Downsample_Y 1 , …, Downsample_Y n )
[0047] In the formula, Y(n) represents the features obtained after parallel convolution, window(3×1) represents the (3×1) sliding pooling window, the concat function represents feature splicing, and Downsample_Y n represents the maximum value within the pooling window, and D represents the multi-scale feature matrix obtained after splicing.
[0048] (2) Use a hierarchical attention module to process the dependency relationships of different scales in the multi-scale feature matrix to obtain hierarchical attention features y_norm;
[0049] 1) Perform non-linear calculation on the multi-scale feature matrix through the tanh function, which is expressed as follows:
[0050] T n = tanh(WD n + b)
[0051]
[0052] In the formula, D nwhere $D_n$ represents the $n$-th downsampling value in the multi-scale feature matrix obtained after splicing, $W$ represents the learnable weight, $b$ represents the learnable bias, $\sinh(x)$ represents the hyperbolic sine function, $\cosh(x)$ is the hyperbolic cosine function, and $\tanh$ represents the hyperbolic tangent function.
[0053] 2) Calculate the proportion of the convolution result of each layer in the total convolution result and perform weighted summation, as expressed below:
[0054]
[0055] In the formula, $T$ n represents the result of performing $\tanh$ non-linear calculation on the $n$-th downsampling value $D$ n in the multi-scale feature matrix obtained after splicing, $\alpha$ n represents the proportion of the $n$-th result in all results, and $N$ represents the range of $n$.
[0056] 3) Weight $T$ n and perform residual connection and normalization calculations, as expressed below:
[0057]
[0058] where $T$ n represents the result of performing $\tanh$ non-linear calculation on the $n$-th downsampling value $D$ n in the multi-scale feature matrix obtained after splicing, $\alpha$ n represents the proportion of the $n$-th result in all results, $\mu$ and $\sigma$ 2 are the mean and variance of $T$ respectively, $\epsilon$ takes the value of $1\times10$ -4 , represents the normalized value of $T$ n , $\lambda$ and $\beta$ are learnable parameters, and $y_{norm}$ is the result after residual connection and normalization calculations.
[0059] (3) Use the feedforward neural network module for training and optimize it through the residual connection and normalization module;
[0060] 1) Train through the feedforward network module. The feedforward network has a total of $L$ layers, as expressed below:
[0061] $y_{FFN}=f(y_{norm})$
[0062] $=\sigma$ (L) $(W$ (L) $\sigma$ (L-1 $(W$ (L-1) $\cdots\sigma$ (1) $(W$ (1) $y_{norm}+b$ (1) )$+b$ (L-1) )$+b$ (L) )
[0063] σ(x) = wing(x) + β
[0064]
[0065] Wherein, w represents the weight matrix, b is the bias vector, σ (L) is the activation function of the L-th layer feedforward network, L is the total number of layers of the feedforward network, β is the learnable parameter of the bias term, w and ε are both hyperparameters, C is a constant, and y_FFN is the output vector of the feedforward network.
[0066] 2) Perform residual connection and normalization calculations, and the calculation process is the same as that of the hierarchical attention module.
[0067] 3) Obtain the severity probability of each relevant fatty liver disease through the softmax function, expressed as follows:
[0068]
[0069] Wherein, n is the total number of categories, P(y = j|y_FFN) is the probability of the j-th category, is the final predicted category, and y_FFN j represents the value of the j-th y_FFN.
[0070] (4) Obtain the weights of the fatty liver detection model through backpropagation, expressed as follows:
[0071]
[0072] Wherein, η represents the hyperparameter learning rate, with a value of 0.01, W represents the learnable weight matrix, b represents the learnable bias term, λ and β represent learnable parameters, y i represents the actual value of the i-th data, is the predicted value of the i-th data, and L is the loss function.
[0073] When the training loss in this round is less than 10 compared with the previous round -4 , stop training and use the weights obtained currently as the final inference weights.
[0074] Fourth step: Input the patient data into the trained fatty liver detection model to obtain the MAFLD risk rating of the patient, including four categories: healthy, low, medium, and high.
[0075] Result comparison
[0076] Taking the data of a hospital physical examination center as the dataset, compare this model with the support vector machine, random forest, and multi-layer perceptron, and the results are as follows:
[0077]
[0078] Taking five indicators of accuracy, precision, recall, F1-score, and the number of parameters as a reference, it can be seen that this model has achieved good performance in all five results. The accuracy reached 88%, higher than that of support vector machines, random forests, and multi-layer perceptrons; the precision reached 92%, indicating that 98% of the samples predicted as positive classes by this model are truly positive class samples; the recall rate was 87%, indicating that this model can identify 92% of the truly positive class samples; the F1-score was 90%, which comprehensively considered precision and recall and is an indicator for the comprehensive evaluation of the model's performance; the number of parameters of this method was 370,000, slightly higher than that of support vector machines and random forests.
[0079] Example 2
[0080] A control server, presented in the form of a general-purpose computing device, includes but is not limited to: one or more processors or processing units, a system memory, and a bus connecting the system memory and the processing unit.
[0081] The bus represents one or more of several types of bus structures, including a memory bus or a memory controller peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the multiple bus structures. For example, these architectures include but are not limited to Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MAC) bus, Enhanced ISA bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus.
[0082] The control server includes a variety of computer system-readable media. These media can be any available media that can be accessed by the control server, including volatile and non-volatile media, removable and non-removable media.
[0083] The system memory can include computer system-readable media in the form of volatile memory, such as random access memory (RAM) and / or cache memory. The control server can further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, a storage system can be used for reading and writing non-removable, non-volatile magnetic media (commonly referred to as a "hard disk drive"). A disk drive can be provided for reading and writing removable non-volatile disks (such as "floppy disks"), and an optical disk drive for reading and writing removable non-volatile optical disks (such as CD-ROM, DVD-ROM, or other optical media). In these cases, each drive can be connected to the bus through one or more data media interfaces. The system memory can include at least one program product having a set (such as at least one) of program modules configured to perform the functions of the embodiments of the present invention.
[0084] A program / util utility having a set (at least one) of program modules can be stored, for example, in a system memory. Such program modules include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment. The program modules generally execute the functions and / or methods in the embodiments described in the present invention.
[0085] The control server can also communicate with one or more external devices, such as a keyboard, a pointing device, a display, etc., and can also communicate with one or more devices that enable a user to interact with the device, and / or communicate with any device that enables the control server to communicate with one or more other computing devices (such as a network card, a modem, etc.). Such communication can be carried out through an input / output (I / O) interface. Moreover, the server can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter. The network adapter communicates with other modules of the control server through a bus. It should be understood that although not shown in the figure, other hardware and / or software modules can be used in combination with the control server, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc. The processing unit executes various functional applications and data processing by running programs stored in the system memory, such as implementing the method for predicting metabolic associated fatty liver disease based on the self-attention mechanism provided in the embodiments of the present invention.
[0086] Embodiment 3
[0087] A computer-readable storage medium stores a computer program, and when the program is executed by a processor, it can implement any of the methods for predicting metabolic associated fatty liver disease based on the self-attention mechanism in the above embodiments.
[0088] The computer storage medium of the embodiments of the present invention may adopt any combination of one or more computer-readable media. The computer-readable media may be computer-readable signal media or computer-readable storage media. The computer-readable storage media may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage media include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage media may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0089] The computer-readable signal media may include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal media may also be any computer-readable media other than the computer-readable storage media, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0090] The program code contained on the computer-readable media may be transmitted by any appropriate medium, including but not limited to:
[0091] wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0092] The computer program code for performing the operations of the present invention may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (for example, by connecting through the Internet service provider via the Internet).
Claims
1. A method for predicting metabolic-related fatty liver disease based on a self-attention mechanism, characterized in that: The following steps are involved: Obtain patient blood test and four diagnostic instrument data to build a fatty liver disease dataset; For the same indicator data of the same patient in different periods, time series information coding is added to form an optimized fatty liver disease data set with sequence characteristics; Constructing and training a fatty liver detection model, wherein the fatty liver detection model includes a convolutional neural network, a hierarchical attention module, and a feedforward neural network layer; The patient data is input into the trained fatty liver detection model to obtain the patient's metabolic-related fatty liver disease prediction level.
2. The method for predicting metabolic-related fatty liver disease based on self-attention mechanism according to claim 1, characterized in that: Blood test indicators include: gender, age, basophil count, hemoglobin, platelet count, neutrophil count, red blood cell count, lymphocyte percentage, lymphocyte absolute value, white blood cell count, monocyte count, hematocrit, eosinophil count, total bilirubin, total protein, albumin, direct bilirubin, glutamyl transpeptidase, alkaline phosphatase, alanine aminotransferase, aspartate aminotransferase, creatinine, uric acid, urea, total cholesterol, triglycerides, high-density lipoprotein cholesterol, low-density lipoprotein cholesterol, fasting blood glucose, α-fetoprotein, carcinoembryonic antigen, BMI value, all indicators are stored in digital form; The data of the four diagnostic instruments include: tongue color, redness of the edges and tip, ecchymosis and petechiae, tongue coating color, thickness of tongue coating, greasy tongue coating, rotten tongue coating, peeled tongue coating, tongue fatness or thinness, tooth marks, punctures, cracks, complexion, redness of the cheekbones, dark eye sockets, luster, lip color, average heart rate, pericardium meridian, liver meridian, kidney meridian, spleen meridian, lung meridian, stomach meridian, gallbladder meridian, and bladder meridian. All indicators are stored in digital form.
3. The method for predicting metabolic-related fatty liver disease based on self-attention mechanism according to claim 1, characterized in that: The constructed fatty liver disease dataset is segmented using the patient ID as the index to form indicator vectors for different patients, which are expressed as follows: P(n)=[XHDB(n), XXBJS(n), SJLXBJS(n),..., PGJ(n)] Where P(n) represents the index matrix of patient n, XHDB(n) represents the hemoglobin index, XXBKS(n) represents the platelet count index, SJLXBJS(n) represents the specific index basophil count, and PGJ(n) represents the bladder meridian index; Among them, each indicator is represented according to different collection times to form an indicator vector, which is expressed as follows: XHDB(n)=[xhdb(1), xhdb(2), xhdb(3)...xhdb(m)] Where XHDB(n) represents the hemoglobin index of patient n, xhdb(1) represents the first hemoglobin value collected by patient n, xhdb(2) represents the second hemoglobin value collected by patient n, and xhdb(m) represents the mth hemoglobin value collected by patient n.
4. The method for predicting metabolic-related fatty liver disease based on self-attention mechanism according to claim 1, characterized in that: For the same indicator data of the same patient in different periods, the specific method of adding time series information coding to form an optimized fatty liver disease data set with sequence characteristics is as follows: For the same indicator of the same patient at different times, the sequence information is added through trigonometric functions, and the position coding is expressed as follows: In the formula, w represents the learnable parameter, pos represents the position in the sequence, and i represents the dimension index of the vector, where the sine function is used for even dimensions and the cosine function is used for odd dimensions; By adding the position code to the original indicator value, a data set with time series information is obtained.
5. The method for predicting metabolic-related fatty liver disease based on self-attention mechanism according to claim 1, characterized in that: The specific method of building a fatty liver detection model and training it is as follows: The convolutional neural network used was used to extract features from the time series data in the fatty liver disease dataset. The extracted features were downsampled, scaled and concatenated to obtain a multi-scale feature matrix. Use a hierarchical attention module to process dependencies of different scales in a multi-scale feature matrix and obtain hierarchical attention features; The feedforward neural network module is used for training and optimized through the residual connection and normalization modules to obtain a trained fatty liver detection model.
6. The method for predicting metabolic-related fatty liver disease based on self-attention mechanism according to claim 5, characterized in that: The convolutional neural network used was used to extract features from the time series data in the fatty liver disease dataset. The extracted features were downsampled, scaled and concatenated to obtain a multi-scale feature matrix. The specific method is as follows: The indicators with time series information are convolved through multi-scale convolution kernels, which can be expressed as follows: K∈{K1,K2,…,K m } In the formula, K represents the convolution kernel set, K m represents the mth convolution kernel, Z e represents the eth indicator with time series information, n is the number of indicators with time series information, m represents the number of convolution kernels, K c represents the cth convolution kernel, and Y(n) represents the convolution feature obtained by convolution calculation; For multi-scale features, we continuously downsample from the calculation result of the first convolution kernel to the calculation result of the last convolution kernel, unify the dimensions and then splice them to obtain the multi-scale feature matrix, which is expressed as follows: D=concat(Downsample_Y0,Downsample_Y1,...,Downsample_Y n ) In the formula, Y(n) represents the features obtained after parallel convolution, window(3×1) represents the (3×1) sliding pooling window, concat function represents feature concatenation, Downsample_Y n represents the maximum value in the pooling window, and D represents the multi-scale feature matrix obtained after splicing.
7. The method for predicting metabolic-related fatty liver disease based on self-attention mechanism according to claim 5, characterized in that: A hierarchical attention module is used to process dependencies of different scales in a multi-scale feature matrix. The specific method for obtaining hierarchical attention features is as follows: 1) Perform nonlinear calculation on the multi-scale feature matrix through the tanh function, which is expressed as follows: T n =tanh(WD n +b) Where D n represents the nth downsampled value in the multi-scale feature matrix obtained after splicing, W represents the learnable weight, b represents the learnable bias, sinh(x) represents the hyperbolic sine function, cosh(x) is the hyperbolic cosine function, and tanh represents the hyperbolic tangent function. 2) Calculate the proportion of each layer of convolution results to the total convolution results and make a weighted sum, expressed as follows: Where, T n Represents the nth downsampled value D in the multi-scale feature matrix obtained after concatenation n The result of tanh nonlinear calculation, α n It indicates the proportion of the nth result in all the results, and N indicates the range of n. 3) T n Weighted, residual connection and normalized calculation are performed, which can be expressed as follows: Among them, T n Represents the nth downsampled value D in the multi-scale feature matrix obtained after concatenation n The result of tanh nonlinear calculation, α n Indicates the proportion of the nth result in all results, μ and σ 2 are the mean and variance of T respectively, and ε is 1×10 -4 , Indicates T n The normalized value of , λ and β are learnable parameters, and y_norm is the result of residual connection and normalization calculation.
8. The method for predicting metabolic-related fatty liver disease based on self-attention mechanism according to claim 7, characterized in that: The specific method of using the feedforward neural network module for training and optimizing through the residual connection and normalization modules to obtain the trained fatty liver detection model is as follows: 1) Training is performed through a feedforward network module. The feedforward network has a total of L layers, which can be expressed as follows: y_FFN=f(y_norm)=σ (L) (w (L) s (L-1) (w (L-1) …s (1) (w (1) y_norm+b (1) )+b (L-1) )+b (L) ) σ(x)=wing(x)+β In the formula, w represents the weight matrix, b is the bias vector, σ( L ) is the activation function of the L-th layer feedforward network, L is the total number of layers of the feedforward network, β is the learnable parameter of the bias term, w and ε are both hyperparameters, C is a constant, and y_FFN is the output vector of the feedforward network; 2) Perform residual connection and normalization calculations. The calculation process is the same as that of the hierarchical attention module. 3) The probability of severity of each related fatty liver disease is obtained by normalized exponential function, expressed as follows: Where n is the total number of categories, P(y=j|y_FFN) is the probability of category j, is the final predicted category, y_FFN j Indicates the value of the jth y_FFN; (4) The weight of the fatty liver detection model is obtained through back propagation, which is expressed as follows: In the formula, η represents the hyperparameter learning rate, which is 0.01, W represents the learnable weight matrix, b represents the learnable bias term, λ and β represent the learnable parameters, and y i Indicates the actual value of the i-th data. is the predicted value of the i-th data, and L is the loss function; When the loss of this round of training is less than the set value compared with the previous round, the training is stopped and the current weight is used as the final inference weight.
9. A control server, characterized in that: include: one or more processors; A storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the metabolic-related fatty liver disease prediction method based on the self-attention mechanism as described in any one of claims 1-8.
10. A storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method for predicting metabolic-related fatty liver disease based on a self-attention mechanism as described in any one of claims 1 to 8 is implemented.
Citation Information
Cited By
Power robot control method and system based on visual voice action model
CN121043156A