Method for risk scoring of methylation detection results of cervical cancer based on machine learning model

CN122842934APending Publication Date: 2026-09-29QUANYI INVESTMENT (GUANGZHOU) MEDICAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610999741.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-07
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0003]然而,现有甲基化数据的处理方法多采用静态分析策略,通常仅基于单个时间点的甲基化水平进行特征提取,未能充分利用甲基化水平在时间维度上的动态变化信息

Benefits of technology

[0015]上述基于机器学习模型的宫颈癌甲基化检测结果风险评分方法、系统、计算机设备及存储介质,通过将甲基化数据与细胞形态学图像融合构建为三维时空张量,并在时间维度上计算各基因的甲基化变化率,能够有效捕捉甲基化水平在时间维度上的动态变化特征,克服了现有静态分析方法无法反映甲基化动态演变过程的缺陷。通过非平稳因果推断算法分析基因甲基化变化与细胞形态学特征之间的动态因果关系并生成动态因果权重参数,能够从因果层面揭示甲基化异常与细胞形态改变之间的内在关联,为特征选择提供具有生物学意义的指导。通过构建强化学习模型并定义状态空间、动作空间和奖励函数,利用强化学习模型动态调节可变形卷积核的采样位置和因果权重衰减系数,使特征提取过程能够根据数据特性自适应地调整,并在闭环迭代中持续优化,从而提升了跨模态时空特征提取的适应性和准确性,增强了模型对数据分布变化的鲁棒性,进而在后续由医师结合临床诊断标准及其他辅助检查结果使用得到的特征重要性评分时,能够为宫颈癌风险评估提供更准确的参考依据。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122842934A_ABST
    Figure CN122842934A_ABST
Patent Text Reader

Abstract

The application relates to a cervical cancer methylation detection result risk scoring method and system based on a machine learning model, a computer device and a storage medium. The method comprises the following steps: acquiring methylation data and cell morphology images, and constructing a three-dimensional space-time tensor; calculating the methylation change rate of each gene to generate dynamic causal weight parameters; constructing a reinforcement learning model, taking the dynamic causal weight parameters and the tensor statistics as a state space, taking the deformable convolution kernel offset adjustment amount and the causal weight decay coefficient adjustment amount as an action space, and taking the feature extraction time and the cross-modal feature correlation error as a reward function; extracting cross-modal space-time interaction features, inputting the features into the model, outputting the adjustment amount, and updating the sampling position and the decay coefficient; updating the model parameters and repeating iterations until a termination condition is met; and calculating a feature importance score. The application can adaptively optimize the feature extraction process, improve the feature expression and cross-modal information fusion effect, and further improve the accuracy of subsequent auxiliary analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of methylation data processing technology, and in particular to a risk scoring method, system, computer device and storage medium for cervical cancer methylation detection results based on a machine learning model. Background Technology

[0002] Cervical cancer is one of the leading malignant tumors threatening women's health. Currently, high-risk human papillomavirus (HPV) testing and thin-layer liquid-based cytology (TCT) are the most widely used clinical detection methods. In recent years, DNA methylation testing has attracted attention due to its specific change patterns during cervical lesions. Existing studies have found that the methylation levels of multiple genes, such as PAX1, SOX1, and DAPK1, are associated with the progression of cervical lesions.

[0003] However, existing methods for processing methylation data mostly employ static analysis strategies, typically extracting features based only on methylation levels at a single time point, failing to fully utilize the dynamic changes in methylation levels over time. Furthermore, existing methods often treat methylation data and cell morphology information as independent modalities during feature extraction, ignoring the inherent relationships between different modalities and resulting in the ineffective utilization of cross-modal complementary information. This static and isolated data processing approach leads to insufficient feature representation capabilities, failing to fully reflect the complex relationship between methylation changes and cell morphological alterations, thus limiting the effectiveness of subsequent data analysis and applications. Summary of the Invention

[0004] Therefore, it is necessary to provide a risk scoring method, system, computer device, and storage medium for cervical cancer methylation detection results based on a machine learning model that can integrate dynamic changes in methylation with cross-modal causal features and adaptively optimize the feature extraction process.

[0005] Firstly, this application provides a risk scoring method for cervical cancer methylation detection results based on a machine learning model. The method includes the following steps: S1: Acquire methylation data and cell morphology images, and construct a three-dimensional spatiotemporal tensor by combining the methylation data and the cell morphology images; S2: Calculate the methylation change rate of each gene from the time dimension of the three-dimensional spatiotemporal tensor, and apply a non-stationary causal inference algorithm to analyze the dynamic causal relationship between the methylation change rate of each gene and cell morphological characteristics, and generate dynamic causal weight parameters. S3: Using the dynamic causal weight parameters and the statistics of the three-dimensional spatiotemporal tensor as the state space, the offset adjustment of the deformable convolution kernel and the causal weight decay coefficient adjustment as the action space, and the feature extraction time and cross-modal feature correlation error as the reward function, a reinforcement learning model is obtained. S4: Using the three-dimensional spatiotemporal tensor as input, extract cross-modal spatiotemporal interaction features using the deformable convolutional kernel; S5: Input the dynamic causal weight parameters, the statistics of the three-dimensional spatiotemporal tensor, and the cross-modal spatiotemporal interaction features as the current state into the reinforcement learning model. The reinforcement learning model outputs the offset adjustment amount and the decay coefficient adjustment amount. The sampling position of the deformable convolution kernel is updated according to the offset adjustment amount, and the decay coefficient of the dynamic causal weight parameters is updated according to the decay coefficient adjustment amount. S6: Update the parameters of the reinforcement learning model based on the reward function; S7: Repeat steps S4 to S6 until the preset termination condition is met; S8: Calculate the feature importance score of each gene marker in the methylation data based on the parameters of the trained reinforcement learning model.

[0006] In one embodiment, the first dimension of the three-dimensional spatiotemporal tensor corresponds to gene information, the second dimension corresponds to time point information, and the third dimension corresponds to spatial location information.

[0007] In one embodiment, S2 includes: A sliding time window is set on the time dimension of the three-dimensional spatiotemporal tensor; Within each of the aforementioned sliding time windows, the conditional mutual information value between the methylation change rate of each gene and the change rate of each morphological feature in the cell morphology image is calculated; The causal influence direction and influence weight of each gene on cell morphological characteristics are determined based on the conditional mutual information value, and the average value of the influence weight of each gene in each time window is used as the dynamic causal weight parameter.

[0008] In one embodiment, the state space includes: the dynamic causal weight parameters, and the mean and variance of the three-dimensional spatiotemporal tensor.

[0009] In one embodiment, the expression for the reward function is: in, The feature extraction time; The cross-modal feature correlation error, , This represents the methylation data feature portion in the cross-modal spatiotemporal interaction features. With cell morphology characteristics Mutual information values ​​between them; , These are preset positive weighting coefficients.

[0010] In one embodiment, the expression for the feature importance score is: in, Indicates gene markers, Indicates the current time point, This indicates the current iteration number of the reinforcement learning model; and These are the weight coefficients related to the number of iteration steps, and and The value of varies The increase is monotonically increasing; Genes identified based on gene function databases The functional significance coefficient; For genes At the point of time The absolute value of the methylation change rate; These are preset spatial heterogeneity weighting coefficients; For genes based on the three-dimensional spatiotemporal tensor At the point of time The spatial heterogeneity index determined by the spatial distribution.

[0011] In one embodiment, the preset termination condition is: the value of the reward function is less than a preset rate of change threshold in N consecutive iterations, where N is a preset positive integer and the rate of change threshold is a preset positive number.

[0012] Secondly, this application also provides a risk scoring system for cervical cancer methylation detection results based on a machine learning model. The system includes: The tensor construction module is used to acquire methylation data and cell morphology images, and construct a three-dimensional spatiotemporal tensor from the methylation data and the cell morphology images; The causal inference module is used to calculate the methylation change rate of each gene from the time dimension of the three-dimensional spatiotemporal tensor, and to apply a non-stationary causal inference algorithm to analyze the dynamic causal relationship between the methylation change rate of each gene and cell morphological characteristics, and generate dynamic causal weight parameters. The model building module is used to obtain a reinforcement learning model by using the dynamic causal weight parameters and the statistics of the three-dimensional spatiotemporal tensor as the state space, the offset adjustment of the deformable convolution kernel and the adjustment of the causal weight decay coefficient as the action space, and the feature extraction time and cross-modal feature correlation error as the reward function. The feature extraction module is used to extract cross-modal spatiotemporal interaction features by taking the three-dimensional spatiotemporal tensor as input and applying the deformable convolution kernel. The reinforcement learning processing module is used to input the dynamic causal weight parameters, the statistics of the three-dimensional spatiotemporal tensor, and the cross-modal spatiotemporal interaction features as the current state into the reinforcement learning model. The reinforcement learning model outputs the offset adjustment amount and the decay coefficient adjustment amount. The module updates the sampling position of the deformable convolution kernel according to the offset adjustment amount and updates the decay coefficient of the dynamic causal weight parameters according to the decay coefficient adjustment amount. The parameter update module is used to update the parameters of the reinforcement learning model based on the reward function; The loop control module is used to trigger the feature extraction module, the reinforcement learning processing module and the parameter update module to execute repeatedly in sequence until a preset termination condition is met. The scoring calculation module is used to calculate the feature importance score of each gene marker in the methylation data based on the parameters of the trained reinforcement learning model.

[0013] Thirdly, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the aforementioned risk scoring method for cervical cancer methylation detection results based on a machine learning model.

[0014] Fourthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, implements the aforementioned risk scoring method for cervical cancer methylation detection results based on a machine learning model.

[0015] The aforementioned risk scoring method, system, computer equipment, and storage medium for cervical cancer methylation detection results based on a machine learning model, constructs a three-dimensional spatiotemporal tensor by fusing methylation data with cell morphology images, and calculates the methylation change rate of each gene in the time dimension. This effectively captures the dynamic changes in methylation levels over time, overcoming the limitation of existing static analysis methods that cannot reflect the dynamic evolution of methylation. By analyzing the dynamic causal relationship between gene methylation changes and cell morphology features using a non-stationary causal inference algorithm and generating dynamic causal weight parameters, the intrinsic association between methylation abnormalities and cell morphological changes can be revealed at the causal level, providing biologically significant guidance for feature selection. By constructing a reinforcement learning model and defining the state space, action space, and reward function, the model dynamically adjusts the sampling position of deformable convolutional kernels and the causal weight decay coefficient. This enables the feature extraction process to adaptively adjust according to data characteristics and continuously optimize in closed-loop iterations, thereby improving the adaptability and accuracy of cross-modal spatiotemporal feature extraction and enhancing the model's robustness to changes in data distribution. Consequently, when physicians use the feature importance scores obtained in conjunction with clinical diagnostic criteria and other auxiliary examination results, it can provide a more accurate reference for cervical cancer risk assessment. Attached Figure Description

[0016] Figure 1 This is a diagram illustrating the application environment of a risk scoring method for cervical cancer methylation detection results based on a machine learning model in one embodiment. Figure 2 This is a flowchart illustrating a risk scoring method for cervical cancer methylation detection results based on a machine learning model in one embodiment. Figure 3 This is a schematic diagram of the process for generating dynamic causal weight parameters in one embodiment; Figure 4 This is a block diagram of a risk scoring system for cervical cancer methylation detection results based on a machine learning model, as shown in one embodiment. Figure 5 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application. First, to facilitate understanding of the technical solutions provided by the embodiments of this application, the background technology involved in the embodiments of this application will be described below.

[0018] Cervical cancer is one of the leading malignant tumors threatening women's health. Currently, high-risk human papillomavirus (HPV) testing and thin-layer liquid-based cytology (TCT) are the most widely used clinical detection methods. However, TCT results have poor reproducibility and low sensitivity, while HPV DNA testing lacks specificity, easily leading to overdiagnosis and overtreatment. Therefore, finding more stable and objective auxiliary detection indicators has become a research hotspot in this field.

[0019] In recent years, DNA methylation detection has attracted attention due to its specific patterns of change during cervical lesions. As an important epigenetic modification, DNA methylation levels of specific genes change significantly during the development and progression of cervical cancer. Existing studies have found that the methylation levels of several genes, including PAX1, SOX1, and DAPK1, are associated with the progression of cervical lesions. Therefore, detecting and analyzing the methylation levels of these genes is expected to provide supplementary information for cervical cancer screening.

[0020] However, existing methods for processing methylation data mostly employ static analysis strategies. Specifically, current methods typically perform feature extraction and analysis based only on methylation data collected at a single time point, failing to fully utilize the dynamic changes in methylation levels over time. In fact, during the progression of cervical lesions, gene methylation levels are not constant but exhibit a gradual change over time. This dynamic change itself contains rich biological significance, and existing static analysis methods, by not taking this into account, fail to fully reflect the occurrence and development of methylation abnormalities, thus limiting the accuracy and reliability of subsequent data analysis.

[0021] Furthermore, existing methods often treat methylation data and cell morphology information as independent modalities during feature extraction. Specifically, current methods typically extract numerical features from methylation data and visual features from cell morphology images separately, then simply concatenate or fuse them. Because they ignore the intrinsic relationships between different modalities, cross-modal complementary information is not effectively utilized, hindering further improvements in feature expression capabilities. In reality, there is a close biological link between gene methylation status and cell morphology; abnormal methylation of specific genes can lead to changes in the expression of related proteins, thereby affecting cell morphology. Existing methods fail to effectively model and utilize this cross-modal intrinsic relationship, resulting in extracted features that cannot fully reflect the complex relationship between methylation changes and cell morphology alterations.

[0022] The aforementioned static and isolated data processing methods limit the effectiveness of feature expression and cross-modal information fusion, making it difficult for the final feature importance score to accurately reflect the actual contribution of each gene marker to dynamic changes in methylation and changes in cell morphology, thereby affecting the effectiveness of subsequent data analysis and applications.

[0023] Therefore, embodiments of this application provide a method, system, computer device, and storage medium for risk scoring of cervical cancer methylation detection results based on a machine learning model.

[0024] The risk scoring method for cervical cancer methylation detection results based on a machine learning model provided in this application can be applied to, for example... Figure 1 In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104 or placed on a cloud or other network server. Terminal 102 acquires methylation data and cell morphology images and sends them to server 104 via the network. After receiving the methylation data and cell morphology images, server 104 constructs a three-dimensional spatiotemporal tensor, calculates the methylation change rate of genes, and applies a non-stationary causal inference algorithm to generate dynamic causal weight parameters. Based on this, a reinforcement learning model is constructed, using state space, action space, and reward function to drive deformable convolutional kernels for adaptive feature extraction. After closed-loop iterative optimization, the model outputs feature importance scores for each gene marker, providing intermediate results information for subsequent data analysis. Terminal 102 can be, but is not limited to, various computer devices with data acquisition and communication functions, such as desktop computers, laptops, and medical data terminals deployed in medical institutions. Server 104 can be a standalone server or a server cluster composed of multiple servers.

[0025] It should be noted that methylation data and cell morphology images belong to medical and health data. According to relevant regulations, medical and health data is considered sensitive personal information, and processing sensitive personal information requires the individual's separate consent. The processing of personal information should adhere to the principles of legality, legitimacy, necessity, and good faith. In this application embodiment, the acquisition and use of the aforementioned data are based on legality and compliance, i.e., explicit authorization from the data subject or other legal basis stipulated by laws and regulations, and are used solely for the purpose of feature extraction and scoring calculation as described in this application method. Necessary security protection measures are taken during data storage, transmission, and processing to ensure the confidentiality, integrity, and availability of the data. Simultaneously, according to relevant regulations, data processing activities should strengthen risk monitoring and fulfill data security protection obligations. The methylation data involved in this application scheme belongs to human genetic resource information, and the processing of related data also complies with the relevant laws and regulations on human genetic resource management. Furthermore, the method of this application directly produces a feature importance score, which serves as intermediate result information for auxiliary analysis, rather than being directly used for determining diagnostic conclusions, and does not involve disease diagnosis methods.

[0026] Firstly, as mentioned above, existing methylation data processing typically extracts and analyzes features based only on methylation data collected at a single time point, failing to fully utilize the dynamic changes in methylation levels over time. Furthermore, it often treats methylation data and cell morphology images as independent modalities, ignoring the inherent correlation between different modalities, resulting in extracted features that cannot comprehensively reflect the occurrence and development of methylation anomalies. Therefore, in one embodiment, such as... Figure 2 As shown, a risk scoring method for cervical cancer methylation detection results based on a machine learning model is provided, and this method is applied to... Figure 1 Taking terminal 102 as an example, the explanation includes the following steps: S1: Acquire methylation data and cell morphology images, and construct a three-dimensional spatiotemporal tensor from the methylation data and cell morphology images.

[0027] In this step, methylation data refers to the methylation level information of each gene locus obtained from the sample using methylation detection technology. The methylation data can be obtained as follows: First, DNA is extracted from cervical cell samples. The extracted DNA is then subjected to bisulfite conversion, converting unmethylated cytosine to uracil, while methylated cytosine remains unchanged. Next, the target gene region is amplified by PCR. The amplified product is hybridized with a methylation detection chip or sequenced to obtain the methylation signal intensity of each gene locus. Finally, the methylation signal intensity is converted into a methylation level value. The formula for calculating the methylation level value is: The value ranges from 0 to 1. In practice, methylation data can be represented as a matrix, where rows correspond to gene markers, columns correspond to time points or samples, and matrix elements represent the methylation level of the corresponding gene marker at the given time point or sample. Gene markers include, but are not limited to, genes known to be associated with the progression of cervical lesions, such as PAX1, SOX1, and DAPK1.

[0028] Cell morphology images refer to cervical cell morphology images acquired through microscopic imaging techniques. The acquisition method for cell morphology images involves: first, preparing cervical cell samples into cell smears, fixing and staining the smears; then, scanning and imaging the stained cell smears using a microscope and a digital image acquisition system to obtain digital images containing morphological information such as cell nuclear morphology, cytoplasmic characteristics, and nucleocytoplasmic ratio. In specific implementations, cell morphology images are two-dimensional pixel matrices, where the pixel value of each pixel represents the optical density or color intensity information at that location.

[0029] A three-dimensional spatiotemporal tensor is a three-dimensional data structure formed by fusing methylation data with cell morphology images. Specifically, the first dimension of the three-dimensional spatiotemporal tensor corresponds to gene information, the second dimension corresponds to time point information, and the third dimension corresponds to spatial location information. The three-dimensional spatiotemporal tensor is constructed as follows: for each gene marker, the methylation level value at each time point is mapped to the corresponding gene and time dimension positions in the three-dimensional spatiotemporal tensor, and simultaneously associated with the corresponding spatial location information in the cell morphology image. Spatial alignment is required during the specific construction of the three-dimensional spatiotemporal tensor. Spatial alignment refers to mapping gene loci in the methylation data to their spatial locations in the cell morphology image. Specifically, for each gene marker, its corresponding cell type or cell region in the cell sample is determined, and then, based on the spatial location range occupied by that cell type or cell region in the cell morphology image, the gene marker is associated with its corresponding spatial location. Spatial alignment can be achieved using image registration algorithms and cell segmentation algorithms. Any of the existing technologies can be used, such as image registration algorithms based on mutual information or cell segmentation algorithms based on deep learning; this embodiment does not limit this approach.

[0030] S2: Calculate the methylation change rate of each gene from the time dimension of the three-dimensional spatiotemporal tensor, and apply a non-stationary causal inference algorithm to analyze the dynamic causal relationship between the methylation change rate of each gene and cell morphological characteristics, and generate dynamic causal weight parameters.

[0031] In this step, the methylation change rate refers to the rate at which the methylation level of each gene marker changes over time. The specific implementation of calculating the methylation change rate of each gene from the time dimension of the three-dimensional spatiotemporal tensor is as follows: For each gene marker and each spatial location, the methylation level value of the gene marker at adjacent time points is obtained in the time dimension. The change in methylation level between adjacent time points is calculated, and then the change in methylation level is divided by the time interval to obtain the methylation change rate of the gene marker at that spatial location and time interval. For each gene marker, the methylation change rate at all spatial locations is averaged along the spatial dimension to obtain the aggregate methylation change rate of the gene marker. Through the above method, a methylation change rate sequence for each gene marker at different time intervals is obtained.

[0032] Non-stationary causal inference algorithms are used to analyze causal relationships in non-stationary time series data. The "non-stationarity" in non-stationary causal inference algorithms refers to the fact that the statistical properties (such as mean and variance) of the time series change over time, as opposed to stationary time series whose statistical properties do not change over time. In this application, a non-stationary causal inference algorithm is used to analyze the dynamic causal relationship between the methylation change rate of each gene and cell morphological characteristics. Specifically, a non-stationary causal model is established using the methylation change rate sequence of each gene as the causal variable and the cell morphological characteristic sequence as the outcome variable. This model indicates that the cell morphological characteristics at the current moment are determined by the methylation change rate at past moments through a time-varying function. By estimating the parameters of the time-varying function, the causal influence strength of the methylation change rate of each gene on cell morphological characteristics can be obtained.

[0033] The dynamic causal weight parameter refers to a numerical parameter that quantifies the degree of influence of methylation changes of each gene on cell morphological characteristics. In this application, each gene marker corresponds to a dynamic causal weight parameter, and the value of the dynamic causal weight parameter reflects the strength of the influence of the gene's methylation changes on cell morphological characteristics. The value range of the dynamic causal weight parameter is 0 to 1, and the larger the value, the stronger the influence of the gene's methylation changes on cell morphological characteristics. Specifically, the dynamic causal weight parameter is calculated as follows: for each gene, the causal influence strength between the gene's methylation change rate sequence and each morphological characteristic sequence is calculated, and then the weighted average of the causal influence strength of the gene on all morphological characteristics is taken as the dynamic causal weight parameter of the gene. The non-stationary causal inference algorithm can adopt any non-stationary time series causal inference algorithm in the prior art, such as the PCMCI+ algorithm or the time-varying VAR model, and this embodiment does not limit it.

[0034] S3: The reinforcement learning model is obtained by using the dynamic causal weight parameters and the statistics of the three-dimensional spatiotemporal tensor as the state space, the offset adjustment of the deformable convolution kernel and the causal weight decay coefficient adjustment as the action space, and the feature extraction time and cross-modal feature correlation error as the reward function.

[0035] In this step, the state space is the set of all observable states of the reinforcement learning model. In this application, the state space is composed of dynamic causal weight parameters and statistics of a three-dimensional spatiotemporal tensor. The dynamic causal weight parameters are a vector composed of the dynamic causal weight parameters of each gene, and the statistics of the three-dimensional spatiotemporal tensor include the mean and variance of the three-dimensional spatiotemporal tensor along each dimension. The state space contains complete data distribution information at the gene, temporal, and spatial levels.

[0036] The action space is the set of all actions that a reinforcement learning model can execute. In this application, the action space consists of the offset adjustment of the deformable convolutional kernel and the adjustment of the causal weight decay coefficient. The offset adjustment of the deformable convolutional kernel is a vector whose dimension is equal to the number of sampling points in the deformable convolutional kernel, and each element represents the offset adjustment value of the corresponding sampling point. The adjustment of the causal weight decay coefficient is a scalar representing the adjustment magnitude of the current decay coefficient.

[0037] A deformable convolutional kernel is a type of convolutional kernel that adaptively adjusts its sampling shape based on local features of the input data by adding a learnable offset at each sampling point location of a standard convolutional kernel. Standard convolutional kernels have fixed sampling positions within a regular grid on the feature map, lacking the ability to model geometric deformations. Deformable convolutional kernels, by adding offsets to the regular grid sampling positions, allow the kernel to adaptively adjust its sampling range at different locations on the feature map, thereby better capturing local deformation features and cross-modal interaction information in the input data.

[0038] The reward function is a function used in reinforcement learning models to evaluate the reward gained by an agent after performing an action. The reward function is determined based on feature extraction time and cross-modal feature correlation error. Feature extraction time refers to the time required to extract cross-modal spatiotemporal interaction features from a three-dimensional spatiotemporal tensor. Cross-modal feature correlation error refers to the error in the correlation between methylated data features and cell morphology features.

[0039] Reinforcement learning is a machine learning model that learns optimal decision-making strategies by interacting with the environment and receiving reward signals from the environment. In this application, the reinforcement learning model can be a deep Q-network model, or other reinforcement learning models (such as deep deterministic policy gradient models, proximal policy optimization models, etc.), and this embodiment does not limit this. This embodiment uses a deep Q-network, which is a reinforcement learning model combining deep neural networks and the Q-learning algorithm. Its structure consists of an input layer, hidden layers, and an output layer. The size of the input layer is equal to the dimension of the state space, i.e., the sum of the number of gene markers G and the statistical dimensions (mean and variance, two in total) of the three-dimensional spatiotemporal tensor, plus 1 (the flattened dimension of the cross-modal spatiotemporal interaction features, the specific size of which depends on the feature map size). In practical applications, the input layer dimension can be adjusted according to the actual data scale, and this embodiment does not limit this. The hidden layer includes multiple fully connected layers, each followed by an activation function for nonlinear transformation of the input features. The first fully connected layer has an input dimension equal to the input layer size, and its output dimension can be set to 256. The second fully connected layer has an input dimension of 256 and an output dimension of 128. The third fully connected layer has an input dimension of 128 and an output dimension of 64. A ReLU activation function is used to introduce non-linearity between the fully connected layers. The output layer is a fully connected layer with an input dimension of 64. Its output dimension is equal to the dimension of the action space, which is the sum of the number of sampling points N of the deformable convolutional kernel and 1 (the adjustment amount of the causal weight decay coefficient). Each neuron in the output layer corresponds to the Q-value of an action.

[0040] The training process of a deep Q-network includes the following steps: First, a batch of experiences is randomly sampled from the experience replay pool. Each experience includes the current state, the action performed, the reward obtained, and the next state. Then, for each experience, a target Q-value is calculated using a target network. The target network has the same structure as the main network, and its parameters are periodically copied from the main network. The target Q-value is calculated as the current reward plus a discount factor multiplied by the maximum Q-value of the next state. The discount factor ranges from 0 to 1 and controls the current value of future rewards. Next, the mean squared error between the Q-value output by the main network and the target Q-value is calculated as the loss function. Finally, the gradient of the loss function with respect to the main network parameters is calculated using the backpropagation algorithm, and the parameters of the main network are updated using an optimizer. The optimizer can be the Adam optimizer, with a learning rate ranging from 0.0001 to 0.01. After a preset number of steps (e.g., 100 steps), the parameters of the main network are copied to the target network. In practical applications, specific parameters such as the number of hidden layers, the number of neurons in each layer, the type of activation function, the type of optimizer, the learning rate, and the discount factor can be adjusted according to the data scale and computing resources. This embodiment does not limit these parameters.

[0041] S4: Using a three-dimensional spatiotemporal tensor as input, deformable convolution kernels are applied to extract cross-modal spatiotemporal interaction features.

[0042] In this step, a 3D spatiotemporal tensor is used as the input for deformable convolution operations. A deformable convolution kernel is used to perform convolution operations on the 3D spatiotemporal tensor to extract cross-modal spatiotemporal interaction features. The deformable convolution operation is as follows: for each position on the 3D spatiotemporal tensor, the output of the deformable convolution is the weighted sum of all sampling points around that position, where the position of each sampling point is the regular grid position plus a learnable offset. When the sampling point position is a non-integer position, the value at that position is obtained through interpolation. After performing the above convolution operation on all positions in the 3D spatiotemporal tensor, a cross-modal spatiotemporal interaction feature map is obtained. This map contains the interaction information between the methylation data modality and the cell morphology image modality. The total number of channels in the cross-modal spatiotemporal interaction feature map is equal to the sum of the number of methylation data-related feature channels and the number of cell morphology-related feature channels; the former contains the methylation data feature portion, and the latter contains the cell morphology feature portion. When initializing the weights of deformable convolutional kernels, the weights of the convolutional kernels for the methylated data feature part and the cell morphology feature part are initialized separately, and then adaptively adjusted through network training.

[0043] S5: Input the dynamic causal weight parameters, the statistics of the three-dimensional spatiotemporal tensor, and the cross-modal spatiotemporal interaction features as the current state into the reinforcement learning model. The reinforcement learning model outputs the offset adjustment and the decay coefficient adjustment. The sampling position of the deformable convolution kernel is updated according to the offset adjustment, and the decay coefficient of the dynamic causal weight parameters is updated according to the decay coefficient adjustment.

[0044] In this step, the current state is the state information observed by the reinforcement learning model at time step t, which consists of the dynamic causal weight parameter vector, the statistics of the three-dimensional spatiotemporal tensor, and the cross-modal spatiotemporal interaction features. Figure 3 The process consists of several parts. After the current state is input into the deep Q-network, the network outputs the Q-value for each action in the action space, and then selects the action with the largest Q-value for output. The action includes offset adjustment and attenuation coefficient adjustment.

[0045] The offset adjustment is a vector containing the offset adjustment value for each sampling point in the deformable convolution kernel. When updating the sampling position of the deformable convolution kernel, the offset adjustment is added to the current offset of each sampling point in the deformable convolution kernel to obtain the updated offset, and then the new sampling position is calculated based on the updated offset. The decay coefficient controls the update magnitude of the dynamic causal weight parameter during the iteration process and is a scalar with a value range of 0 to 1. When updating the decay coefficient, the decay coefficient adjustment is added to the current decay coefficient to obtain the updated decay coefficient, and then the updated decay coefficient is limited to the range of [0,1] to ensure that the value of the decay coefficient is always within the effective range.

[0046] S6: Update the parameters of the reinforcement learning model based on the reward function.

[0047] In this step, the reinforcement learning model is the deep Q-network constructed in step S3. After executing the action and updating the sampling position of the deformable convolutional kernel and the decay coefficient of the dynamic causal weight parameters, step S4 is re-executed under the updated deformable convolutional kernel sampling position and decay coefficient to extract cross-modal spatiotemporal interaction features, and then the reward value after executing the action is calculated according to the reward function.

[0048] During the training of a deep Q-network, the current state, the action performed, the reward value obtained, and the next state are stored as an experience in the experience replay pool. The experience replay pool is a fixed-size buffer; when the number of stored experiences reaches a preset threshold, a batch of experiences is randomly sampled from the pool. For each sampled experience, the target Q-value is first calculated using the target network. The discount factor controls the current value of future rewards, ranging from 0 to 1 (typically 0.9). The average of the squared differences between the predicted Q-value and the target Q-value is then calculated as the loss function, where the predicted Q-value is output by the main network. Finally, the parameters of the main network are updated using gradient descent, as follows: Through the training process described above, the deep Q-network gradually optimizes its parameters, enabling the model to select actions that yield higher reward values ​​in the future.

[0049] S7: Repeat steps S4 to S6 until the preset termination condition is met.

[0050] In this step, steps S4 to S6 constitute a complete iterative cycle of reinforcement learning training. In each iteration, step S4 is first executed to extract cross-modal spatiotemporal interaction features, then step S5 is executed to input the current state into the reinforcement learning model and output the action, update the sampling position and decay coefficient, and finally step S6 is executed to update the parameters of the reinforcement learning model based on the reward function. By continuously repeating the above iterative process, the reinforcement learning model gradually optimizes its parameters, and the sampling position of the deformable convolution kernel and the decay coefficient of the dynamic causal weight parameter are also gradually adjusted to a better state. When the preset termination condition is met, the iteration stops. The number of times S4 to S6 are repeated can be determined by those skilled in the art based on experience or conventional experiments, and this embodiment does not limit this.

[0051] S8: Calculate the feature importance score of each gene marker in the methylation data based on the parameters of the trained reinforcement learning model.

[0052] In this step, after the reinforcement learning model has been trained, the feature importance score of each gene marker is calculated using the parameters of the trained deep Q-network. The feature importance score reflects the contribution of each gene marker to the association between methylation changes and cell morphological features, and this score can be used as intermediate result information for subsequent data analysis.

[0053] Through the above steps, methylation data feature extraction based on a machine learning model was achieved. This method constructs a three-dimensional spatiotemporal tensor from methylation data and cell morphology images and calculates the methylation change rate, capturing the dynamic changes in methylation levels over time. Dynamic causal weight parameters are generated using a non-stationary causal inference algorithm, revealing the intrinsic relationship between gene methylation changes and cell morphological features from a causal perspective. Adaptive optimization of the feature extraction process is achieved by dynamically adjusting the sampling position of deformable convolution kernels and the causal weight decay coefficient using a reinforcement learning model. The quality of feature extraction is continuously improved through closed-loop iterative training, ultimately outputting feature importance scores for each gene marker, providing more accurate intermediate results for subsequent data analysis.

[0054] In the construction of a three-dimensional spatiotemporal tensor, the specific information type corresponding to each dimension after data fusion directly determines the correctness of data extraction and calculation along each dimension in subsequent steps. If the dimension definition is unclear, directional errors may occur when extracting and calculating data along each dimension in subsequent steps, causing the calculation results to deviate from expectations. Traditional methods do not uniformly define the semantics of each dimension after data fusion, making it difficult for the data stored in the tensor to be correctly parsed and used in subsequent steps, affecting the accuracy of the entire feature extraction process. Therefore, in one embodiment, the first dimension of the three-dimensional spatiotemporal tensor corresponds to gene information, the second dimension corresponds to time point information, and the third dimension corresponds to spatial location information.

[0055] Specifically, the three-dimensional spacetime tensor is a three-dimensional data structure with dimensions of . ,in The total number of gene markers, The total number of time points. This represents the total number of spatial locations. The value depends on the number of gene markers being detected, and may include multiple gene markers such as PAX1, SOX1, and DAPK1. The value depends on the number of longitudinal follow-up time points, which may include multiple time points such as baseline, 6th month, 12th month, etc. The value of depends on the number of spatial locations divided in the cell morphology image; for example, it can be divided into multiple spatial locations based on the cell distribution area on a cell smear. The first dimension of the three-dimensional spatiotemporal tensor. The corresponding gene information, specifically the methylation level data and related feature data of different gene markers stored along the first dimension, corresponds to the second dimension of the three-dimensional spatiotemporal tensor. The corresponding time point information, that is, data collected at different time points, is stored along the second dimension. The third dimension of the three-dimensional spatiotemporal tensor. The corresponding spatial location information, that is, the data stored along the third dimension, are data from different spatial locations in the cell morphology image.

[0056] By mapping the first, second, and third dimensions of the three-dimensional spatiotemporal tensor to gene information, time point information, and spatial location information, respectively, the data at each location in the three-dimensional spatiotemporal tensor has a clear semantic meaning. Subsequent steps of data extraction and calculation along each dimension can accurately locate and acquire target data, avoiding calculation errors caused by unclear dimension meanings, and providing a unified and reliable data foundation for subsequent cross-modal spatiotemporal feature extraction.

[0057] When analyzing the causal relationship between gene methylation changes and cell morphological features, simply describing the application of non-stationary causal inference algorithms without providing specific implementation details will leave those skilled in the art unclear on how to extract the methylation change rate of each gene from a three-dimensional spatiotemporal tensor and analyze its dynamic causal relationship with morphological features. Traditional methods often perform causal inference within a single time interval, ignoring the non-stationar nature of causal relationships over time, resulting in causal inference results that fail to reflect the dynamic evolution of the relationship between gene methylation changes and cell morphological features. Therefore, in one embodiment, such as... Figure 3 As shown, step S2 includes: S21: Set a sliding time window in the time dimension of the three-dimensional spacetime tensor; S22: Within each sliding time window, calculate the conditional mutual information value between the methylation change rate of each gene and the change rate of each morphological feature in the cell morphology image; S23: Determine the causal influence direction and influence weight of each gene on cell morphological characteristics based on the conditional mutual information value, and use the average influence weight of each gene in each time window as the dynamic causal weight parameter.

[0058] In this application, the sliding time window is a fixed-length time interval defined on the time dimension of a three-dimensional spacetime tensor. Let the sequence of time points on the time dimension be... The window length is L (i.e., the window contains L consecutive time points, such as 3 time points), and the window sliding step size is D (i.e., the difference in the number of time points between the starting positions of adjacent windows, such as 1 time point). The time points covered by the j-th sliding time window are... ,in , This represents the total number of sliding time windows. In practical applications, the window length L and window sliding step size D can be set according to the temporal resolution of the data and the analysis requirements; this embodiment does not impose any limitations on these.

[0059] For each sliding time window, the methylation change rate of each gene and the change rate of each morphological feature within that window are first calculated. Morphological features may include, but are not limited to, features such as nuclear area, nuclear perimeter, nucleocytoplasmic ratio, cell morphology index, cytoplasmic area, and nuclear roundness. Then, for each gene and each morphological feature, the conditional mutual information value between the methylation change rate of that gene and the change rate of that morphological feature is calculated. The conditional mutual information value is used to quantify the statistical dependence between the methylation change rate of that gene and the change rate of that morphological feature, given other variables (including the methylation change rates of other genes and the change rates of other morphological features). The larger the conditional mutual information value, the stronger the association between the methylation change of that gene and the change of that morphological feature. The conditional mutual information can be calculated using any existing calculation method, such as the conditional mutual information calculation method based on nuclear density estimation or the conditional mutual information calculation method based on k-nearest neighbor estimation; this embodiment does not limit this.

[0060] Based on the calculated conditional mutual information values, the causal influence direction and influence weight of each gene on cell morphological features are determined. Specifically, for each gene and each morphological feature, the conditional mutual information value reflects the strength of the association between the methylation change rate of the gene and the change rate of the morphological feature. If the conditional mutual information value is greater than a preset significance threshold, the gene is determined to have a causal influence on the morphological feature. The direction of causal influence is determined by comparing the conditional mutual information values ​​in different time lag directions: if the conditional mutual information value when the gene's methylation change precedes the morphological feature change is greater than the reverse conditional mutual information value, the gene is determined to have a causal influence on the morphological feature. The influence weight represents the strength of the causal influence of the gene's methylation change on the morphological feature, and is obtained by normalizing the conditional mutual information values. For each gene, the causal influence weight of the gene on all morphological features is averaged to obtain the causal influence weight of the gene within a single time window. For each gene, the influence weights within all time windows are averaged to obtain the dynamic causal weight parameter of the gene.

[0061] By setting a sliding time window in the time dimension of the three-dimensional spatiotemporal tensor and calculating the conditional mutual information value between the rate of change of gene methylation and the rate of change of morphological characteristics in each window, the non-stationary characteristics of causal relationship changes over time can be captured. This allows the dynamic causal weight parameter to integrate the causal influence information of different time stages, thus reflecting the dynamic causal relationship between gene methylation changes and cell morphological characteristics more comprehensively.

[0062] In constructing reinforcement learning models, the definition of the state space directly impacts the model's decision quality. If the state space contains only partial information or is incompletely defined, the reinforcement learning model will lack sufficient data features to support optimal decisions, causing the model's output actions to deviate from the optimal policy, thus affecting the performance of the entire feature extraction process. Traditional methods often define the state space only including currently extracted features, failing to incorporate biological prior knowledge and data distribution characteristics, resulting in a lack of multi-dimensional information support for the model's decisions. Therefore, in one embodiment, the state space includes: dynamic causal weight parameters, and the mean and variance of a three-dimensional spatiotemporal tensor.

[0063] Specifically, the state space is composed of dynamic causal weight parameters and statistics of the three-dimensional spatiotemporal tensor. The dynamic causal weight parameters are a vector composed of the dynamic causal weight parameters of each gene, reflecting the degree of influence of gene methylation changes on cell morphological characteristics. The statistics of the three-dimensional spatiotemporal tensor include the mean and variance. The mean of the three-dimensional spatiotemporal tensor is calculated by summing all elements and dividing by the total number of elements; the variance of the three-dimensional spatiotemporal tensor is calculated by averaging the squares of the differences between each element and the mean. The mean reflects the overall level of the tensor data, while the variance reflects the degree of dispersion of the tensor data.

[0064] By incorporating dynamic causal weight parameters into the state space, reinforcement learning models can perceive the biological prior knowledge—the degree to which changes in gene methylation affect cell morphological characteristics—when making decisions. By incorporating the mean and variance of the three-dimensional spatiotemporal tensor into the state space, the model can perceive the statistical distribution characteristics—the overall level and dispersion of the current tensor data. The state space integrates biological prior information and statistical data information, enabling the model to make more rational action choices, thereby improving the adaptability and accuracy of feature extraction.

[0065] In reinforcement learning model training, the reward function is the core signal driving model parameter optimization, and its specific expression directly affects the model's learning direction and final performance. If the reward function cannot accurately reflect feature extraction efficiency and cross-modal fusion quality, the model will be unable to learn the optimal feature extraction strategy. Traditional methods often design reward functions based on only a single metric (such as classification accuracy), failing to simultaneously consider feature extraction time and cross-modal fusion quality, leading to a trade-off in the model's optimization process. Therefore, in one embodiment, the expression for the reward function is: in, For feature extraction time; For cross-modal feature correlation error, , This represents the methylation data feature portion in cross-modal spatiotemporal interaction features. With cell morphology characteristics Mutual information values ​​between them; , These are preset positive weighting coefficients.

[0066] Specifically, in this embodiment, the reward function It consists of two parts. The first part is the feature extraction efficiency term, which takes a value of . This represents the processing time, in milliseconds, required to extract cross-modal spatiotemporal interaction features from the 3D spatiotemporal tensor in step S4. This encourages the model to complete feature extraction in the shortest possible time. The shorter the length, the larger the value of this item, and the higher the reward value.

[0067] The second term is the cross-modal feature correlation term, with a value of [value missing]. Mutual information value This is used to measure the degree of interdependence between methylation data features and cell morphology features. The mutual information value can be calculated using any existing mutual information calculation method, such as histogram-based or kernel density-based methods; this embodiment is not limited to any particular method. A larger mutual information value indicates a stronger correlation between methylation data features and cell morphology features, resulting in better cross-modal fusion performance. Because... It is the opposite of the mutual information value, therefore The smaller the value (i.e., the larger the mutual information value), the better the second term. The larger the value of , the higher the reward. Overall, the second term indicates that the reward increases with the increase of cross-modal feature correlation.

[0068] In practical applications, and The value can be adjusted according to the different requirements for efficiency and accuracy in specific application scenarios. For example, in scenarios with high real-time requirements, it can be set to... The value is greater than The value of is chosen to emphasize feature extraction efficiency; in scenarios where high accuracy is required, it can be set to . The value is greater than The value of is chosen to emphasize the correlation of cross-modal features, such as the value of . , ,when and When the values ​​of are equal, feature extraction efficiency and cross-modal feature correlation have the same weight in the reward function, which is suitable for scenarios where a balance between efficiency and accuracy is required. It should be noted that the specific values ​​of the above weight coefficients are only examples and can be adjusted according to specific needs in actual applications. This embodiment does not limit this.

[0069] By incorporating feature extraction time into the first term of the reward function, the model aims to shorten feature extraction time during training, thus improving computational efficiency. Conversely, by incorporating cross-modal feature correlation error into the second term of the reward function, the model aims to enhance the correlation between methylated data features and cell morphology features during training, thereby improving the quality of cross-modal fusion. This weighted combination allows the reinforcement learning model to simultaneously optimize both feature extraction efficiency and cross-modal feature correlation, resulting in features that are both computationally efficient and exhibit good cross-modal fusion performance.

[0070] Feature importance scoring is the final output of this method, and its calculation method directly affects the reliability and usability of the scoring results. If the definitions of the weight coefficients and parameters in the scoring formula are not clear enough, the scoring results will lack technical basis and cannot provide a reliable reference for subsequent data analysis. Traditional methods often assess gene importance based on only a single dimension (such as only based on methylation level or only based on gene function), without simultaneously considering multiple dimensions such as the biological significance of gene function, the degree of dynamic changes in methylation, and spatial heterogeneity, leading to one-sided assessment results. Therefore, in one embodiment, the expression for feature importance scoring is: in, Indicates gene markers, Indicates the current time point, This indicates the current iteration number of the reinforcement learning model; and These are the weight coefficients related to the number of iteration steps, and and The value of varies The increase is monotonically increasing; Genes identified based on gene function databases The functional significance coefficient; For genes At the point of time The absolute value of the methylation change rate; These are preset spatial heterogeneity weighting coefficients; For genes based on three-dimensional spatiotemporal tensors At the point of time The spatial heterogeneity index determined by the spatial distribution.

[0071] Specifically, in this embodiment, feature importance scoring It consists of three items. The first item is the gene function significance item, which takes a value of Functional significance coefficients are based on genes in gene function databases (such as the GeneOntology database). Biological function annotation information can be used to determine this, for example, based on genes. Using annotation information from the GeneOntology database, the enrichment degree of the gene in relevant functional categories such as "cellular processes," "biological regulation," and "metabolic processes" was calculated, and the enrichment degree value was used as the functional significance coefficient. The first term, whose value ranges from 0 to 1, represents the dynamic change in methylation. The second term is the dynamic change term in methylation, with a value of [value missing]. . The value is calculated in step S2. The third term is the spatial heterogeneity term, and its value is... . Genes based on three-dimensional spatiotemporal tensors At the point of time The distribution of methylation levels along spatial dimensions is determined. In one implementation, The calculation method can be: for genes At the point of time The spatial distribution sequence is calculated, and the ratio of the standard deviation to the mean of the sequence, i.e., the coefficient of variation, is used as the spatial heterogeneity index. The larger the coefficient of variation, the more uneven the spatial distribution of methylation levels, and the higher the spatial heterogeneity index. In practical applications, the spatial heterogeneity index can also be calculated using any existing spatial statistic (such as the Moran index, Gini coefficient, etc.), and this embodiment does not limit this.

[0072] for and In the early stages of training (when k is relatively small). and The value of k is relatively small, and the contributions of the gene function significance term and the dynamic change term of methylation are relatively low; as training progresses (k increases), the contribution of k increases. and As the value of gradually increases, the contributions of these two items gradually increase as well. This is a preset spatial heterogeneity weighting coefficient, ranging from 0 to 1, used to adjust the weight of the spatial heterogeneity index in the feature importance score. For example, , , ,in The preset maximum number of iterations can be set to 1000. At that time, α , , These values ​​result in an initial ratio of approximately 3:4:3 for the gene function significance, methylation dynamics, and spatial heterogeneity terms in the feature importance score. In practical applications, and The increasing function form and The specific value can be adjusted according to specific needs, and this embodiment does not limit it.

[0073] By incorporating gene function significance coefficients into the feature importance score, the score results reflect the known importance of each gene in its biological function. By including the absolute value of methylation change rate in the score, the score results reflect the dynamic changes of each gene over time. By incorporating the spatial heterogeneity index into the score, the score results reflect the uneven distribution of each gene in space. This three-pronged approach allows the feature importance score to comprehensively assess the importance of each gene marker across functional, temporal, and spatial dimensions, integrating prior biological knowledge and data-driven features, thus providing the score results with more robust technical evidence.

[0074] In reinforcement learning training, the criteria for determining when to stop training directly affect whether the model can converge to a better state and the computational resource consumption during training. If the termination condition is unclear or unreasonable, it may lead to premature (model not converging) or late (wasted computational resources), affecting training efficiency and model quality. Traditional methods often use a fixed number of iterations as the termination condition, failing to dynamically determine the stopping time based on the model's actual convergence state, resulting in wasted computational resources or insufficient training. Therefore, in one embodiment, the preset termination condition is: the value of the reward function is less than a preset rate of change threshold for N consecutive iterations, where N is a preset positive integer and the rate of change threshold is a preset positive number.

[0075] Specifically, the termination condition is used to determine when the reinforcement learning training process should stop. After each iteration, the rate of change between the reward function value of the current iteration and the reward function value of the previous iteration is calculated. The rate of change is calculated by dividing the absolute value of the difference between the current value and the previous value by the absolute value of the previous value. If the current iteration number is greater than or equal to N, and the rate of change of the most recent N iterations is less than a preset rate of change threshold, then the reward function is considered to have stabilized, and the performance of the reinforcement learning model has reached convergence, at which point training stops. Here, N is a preset positive integer representing the number of iterations that need to be observed continuously, for example, N=10; the rate of change threshold is a preset positive number representing the critical value for determining whether the change in the reward function value is acceptable, for example, the rate of change threshold can be set to 0.01 (i.e., the rate of change is less than 1%). In practical applications, the specific values ​​of N and the rate of change threshold can be set according to the training efficiency and accuracy requirements, for example, N=5 to 20, and the rate of change threshold=0.001 to 0.05. This embodiment does not limit this.

[0076] By using the rate of change of the reward function value as the convergence criterion, the training process can dynamically determine the stopping point based on the actual convergence state of the model. By requiring the rate of change to be less than a preset threshold for N consecutive iterations, premature stopping due to accidental fluctuations in a single iteration is avoided, ensuring that the model has been sufficiently validated before convergence is determined. This termination condition allows the reinforcement learning model to automatically stop training after performance convergence, avoiding unnecessary computational resource consumption and improving training efficiency.

[0077] In medical data applications, patient privacy protection is a crucial factor that must be considered. Methylation data and cell morphology images from different medical institutions involve sensitive patient information; directly sharing raw data poses compliance risks and limits the collaborative use of multi-center data. Traditional methods require centralizing data from various centers for model training, which not only carries the risk of privacy breaches but may also violate data security laws and regulations. Therefore, in one embodiment, this method is applied to a federated learning architecture, and the reward function is determined based on the federated learning communication bandwidth.

[0078] Specifically, federated learning architecture is a distributed machine learning framework that allows participants to collaboratively train a global model by exchanging model parameters or gradient information, while storing the original data locally. In federated learning architecture, participants include a central server and multiple local terminals. Local terminals can be, but are not limited to, computing devices in various medical institutions, and the central server can be, but is not limited to, a standalone server or a cloud server.

[0079] In this application, the specific process of applying the method to the federated learning architecture is as follows: First, the central server initializes the parameters of the global reinforcement learning model and distributes the global parameters to each local terminal. Each local terminal executes steps S1 to S8 locally, training the reinforcement learning model based on local data to obtain local model parameters. Then, each local terminal encrypts its local model parameters and sends the encrypted local model parameters to the central server. The central server aggregates the encrypted local model parameters to obtain the global model parameters. The aggregation operation can use a federated averaging algorithm, that is, a weighted average of the local model parameters. The weights of each local terminal can be set according to the amount of local data, which is not limited in this embodiment. The central server distributes the global model parameters back to each local terminal. Each local terminal updates its local reinforcement learning model based on the global model parameters and repeats the above training and aggregation process until the global model converges.

[0080] Under the aforementioned federated learning architecture, encryption processing can be implemented using a homomorphic encryption scheme (such as CKKS homomorphic encryption), enabling the central server to perform parameter aggregation in ciphertext. CKKS homomorphic encryption is a homomorphic encryption scheme that supports approximate addition and multiplication operations on encrypted data. It allows direct aggregation operations on encrypted model parameters without decryption, thereby protecting the privacy of model parameters on each local terminal. In practical applications, encryption processing can also employ any existing privacy protection technology, such as secure multi-party computation or differential privacy; this embodiment does not limit this approach.

[0081] In a federated learning architecture, the reward function is also determined based on the federated learning communication bandwidth; that is, the expression for the reward function is: ,in For federal learning communication bandwidth, These are preset positive weighting coefficients. Federated learning communication bandwidth. This refers to the communication bandwidth consumed when transmitting model parameters between the central server and each local terminal. In one implementation, The calculation method can be as follows: the amount of encrypted parameters transmitted in a single communication (in bits) is divided by the communication time (in seconds) to obtain the actual communication bandwidth consumption value. By incorporating the communication bandwidth of federated learning into the reward function, the reinforcement learning model will tend to choose actions with lower communication overhead during training (such as choosing a more compact model parameter representation or a more efficient encryption scheme), thereby balancing communication efficiency with privacy protection. The value range is from 0 to 1, such as the value... , , ,when When the value is small, the impact of communication bandwidth on the reward function is relatively weak; when When the value is large, the communication bandwidth has a stronger impact on the reward function. In practical applications, , , The specific value can be adjusted according to specific needs, and this embodiment does not limit it.

[0082] By applying this method to a federated learning architecture, each local terminal trains the model independently while storing the original data locally, sharing only the encrypted model parameters. The central server performs aggregation in encrypted form, enabling collaborative training of the global model without sharing the original methylation data and cell morphology images. This protects patient privacy while fully utilizing multi-center data to improve model performance. Furthermore, by incorporating the federated learning communication bandwidth into the reward function, the model balances communication efficiency during training, achieving a balance between privacy protection and communication overhead.

[0083] Based on the above, the risk scoring method for cervical cancer methylation detection results based on a machine learning model provided in this application can be applied to various application scenarios, such as: (1) In medical institutions, methylation detection and cell morphology image acquisition are performed on the collected cervical cell samples. The characteristic importance score of each gene marker is calculated using the method of this application to provide auxiliary analysis information for physicians.

[0084] (2) In public health screening projects, methylation detection and cell morphology image acquisition are performed on cervical cell samples from large populations. The method of this application is used for efficient methylation data processing and feature extraction to provide data support for subsequent epidemiological analysis and public health decision-making.

[0085] (3) In the field of biomedical research, it is used to discover and analyze key gene markers related to the progression of cervical lesions, providing technical means for the study of the pathogenesis of cervical cancer and the discovery of new biomarkers.

[0086] In the above application scenarios, the feature importance score output by the method provided in this application embodiment is used as intermediate result information for subsequent analysis and decision-making, rather than directly outputting diagnostic conclusions. In practical application scenarios, there are various implementation methods, such as execution on terminal 102 or server 104, or deployment locally or in the cloud, and this application does not limit this.

[0087] It should be understood that although the steps in the flowcharts of the above embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps.

[0088] Secondly, based on the same inventive concept, this application also provides a system for implementing the aforementioned risk scoring method for cervical cancer methylation detection results based on a machine learning model. The solution provided by this system is similar to the implementation described in the above method; therefore, the specific limitations in one or more system embodiments provided below can be found in the above-described limitations regarding the risk scoring method for cervical cancer methylation detection results based on a machine learning model, and will not be repeated here.

[0089] In one embodiment, such as Figure 4 As shown, the system includes: a tensor construction module, a causal inference module, a model construction module, a feature extraction module, a reinforcement learning processing module, a parameter update module, a loop control module, and a score calculation module. Among them: The tensor construction module is used to acquire methylation data and cell morphology images, and construct a three-dimensional spatiotemporal tensor from the methylation data and cell morphology images; The causal inference module is used to calculate the methylation change rate of each gene from the time dimension of the three-dimensional spatiotemporal tensor, and to apply a non-stationary causal inference algorithm to analyze the dynamic causal relationship between the methylation change rate of each gene and cell morphological characteristics, and generate dynamic causal weight parameters. The model building module is used to obtain a reinforcement learning model with the dynamic causal weight parameters and the statistics of the three-dimensional spatiotemporal tensor as the state space, the offset adjustment of the deformable convolution kernel and the adjustment of the causal weight decay coefficient as the action space, and the feature extraction time and cross-modal feature correlation error as the reward function. The feature extraction module is used to extract cross-modal spatiotemporal interaction features by taking a three-dimensional spatiotemporal tensor as input and applying deformable convolution kernels. The reinforcement learning processing module is used to input the dynamic causal weight parameters, the statistics of the three-dimensional spatiotemporal tensor, and the cross-modal spatiotemporal interaction features as the current state into the reinforcement learning model. The reinforcement learning model outputs the offset adjustment and the decay coefficient adjustment. The module updates the sampling position of the deformable convolution kernel according to the offset adjustment and updates the decay coefficient of the dynamic causal weight parameters according to the decay coefficient adjustment. The parameter update module is used to update the parameters of the reinforcement learning model based on the reward function; The loop control module is used to trigger the feature extraction module, reinforcement learning processing module and parameter update module to execute repeatedly in sequence until the preset termination condition is met. The scoring calculation module is used to calculate the feature importance score of each gene marker in the methylation data based on the parameters of the trained reinforcement learning model.

[0090] In one embodiment, the tensor construction module includes tensor construction units. These tensor construction units are configured to define a three-dimensional spatiotemporal tensor with a first dimension corresponding to gene information, a second dimension corresponding to time point information, and a third dimension corresponding to spatial location information.

[0091] In one embodiment, the causal inference module includes a causal inference unit. The causal inference unit is configured to: set a sliding time window in the time dimension of a three-dimensional spatiotemporal tensor; within each sliding time window, calculate the conditional mutual information value between the methylation change rate of each gene and the change rate of each morphological feature in the cell morphology image; determine the causal influence direction and influence weight of each gene on the cell morphology features based on the conditional mutual information value, and use the average influence weight of each gene within each time window as a dynamic causal weight parameter.

[0092] In one embodiment, the model building module includes an action space configuration unit. The action space configuration unit is configured to configure the state space using dynamic causal weight parameters and the mean and variance of a three-dimensional spatiotemporal tensor.

[0093] In one embodiment, the model building module includes a reward function setting unit. This reward function setting unit is configured to calculate the reward function using the following expression: in, For feature extraction time; For cross-modal feature correlation error, , This represents the methylation data feature portion in cross-modal spatiotemporal interaction features. With cell morphology characteristics Mutual information values ​​between them; , These are preset positive weighting coefficients.

[0094] In one embodiment, the scoring calculation module includes a scoring calculation unit. The scoring calculation unit is configured to calculate a feature importance score using the following expression: in, Indicates gene markers, Indicates the current time point, This indicates the current iteration number of the reinforcement learning model; and These are the weight coefficients related to the number of iteration steps, and and The value of varies The increase is monotonically increasing; Genes identified based on gene function databases The functional significance coefficient; For genes At the point of time The absolute value of the methylation change rate; These are preset spatial heterogeneity weighting coefficients; For genes based on three-dimensional spatiotemporal tensors At the point of time The spatial heterogeneity index determined by the spatial distribution.

[0095] In one embodiment, the loop control module includes a termination setting unit. The termination setting unit is configured to use the fact that the value of the reward function is less than a preset rate of change threshold for N consecutive iterations as the termination condition, where N is a preset positive integer and the rate of change threshold is a preset positive number.

[0096] The modules in the aforementioned risk scoring system for cervical cancer methylation detection results based on machine learning models can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.

[0097] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor, memory, and a network interface connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operation of the operating system and computer programs in the non-volatile storage media. The database stores all data required for system operation. The network interface communicates with external terminals via a network connection. When executed by the processor, the computer program implements a risk scoring method for cervical cancer methylation detection results based on a machine learning model.

[0098] Those skilled in the art will understand that Figure 5 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0099] Thirdly, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of any of the methods described in the above method embodiments.

[0100] Fourthly, a computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the steps of any of the methods described in the above method embodiments.

[0101] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0102] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0103] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A risk scoring method for cervical cancer methylation detection results based on a machine learning model, characterized in that, The method includes the following steps: S1: Acquire methylation data and cell morphology images, and construct a three-dimensional spatiotemporal tensor by combining the methylation data and the cell morphology images; S2: Calculate the methylation change rate of each gene from the time dimension of the three-dimensional spatiotemporal tensor, and apply a non-stationary causal inference algorithm to analyze the dynamic causal relationship between the methylation change rate of each gene and cell morphological characteristics, and generate dynamic causal weight parameters. S3: Using the dynamic causal weight parameters and the statistics of the three-dimensional spatiotemporal tensor as the state space, the offset adjustment of the deformable convolution kernel and the causal weight decay coefficient adjustment as the action space, and the feature extraction time and cross-modal feature correlation error as the reward function, a reinforcement learning model is obtained. S4: Using the three-dimensional spatiotemporal tensor as input, extract cross-modal spatiotemporal interaction features using the deformable convolutional kernel; S5: Input the dynamic causal weight parameters, the statistics of the three-dimensional spatiotemporal tensor, and the cross-modal spatiotemporal interaction features as the current state into the reinforcement learning model. The reinforcement learning model outputs the offset adjustment amount and the decay coefficient adjustment amount. The sampling position of the deformable convolution kernel is updated according to the offset adjustment amount, and the decay coefficient of the dynamic causal weight parameters is updated according to the decay coefficient adjustment amount. S6: Update the parameters of the reinforcement learning model based on the reward function; S7: Repeat steps S4 to S6 until the preset termination condition is met; S8: Calculate the feature importance score of each gene marker in the methylation data based on the parameters of the trained reinforcement learning model.

2. The method according to claim 1, characterized in that, The first dimension of the three-dimensional spatiotemporal tensor corresponds to gene information, the second dimension corresponds to time point information, and the third dimension corresponds to spatial location information.

3. The method according to claim 1, characterized in that, S2 includes: S21: Set a sliding time window on the time dimension of the three-dimensional spatiotemporal tensor; S22: Within each of the sliding time windows, calculate the conditional mutual information value between the methylation change rate of each gene and the change rate of each morphological feature in the cell morphology image; S23: Determine the causal influence direction and influence weight of each gene on cell morphological characteristics based on the conditional mutual information value, and use the average value of the influence weight of each gene in each time window as the dynamic causal weight parameter.

4. The method according to claim 1, characterized in that, The state space includes: the dynamic causal weight parameters, and the mean and variance of the three-dimensional spatiotemporal tensor.

5. The method according to claim 1, characterized in that, The expression for the reward function is: in, The feature extraction time; The cross-modal feature correlation error, , This represents the methylation data feature portion in the cross-modal spatiotemporal interaction features. With cell morphology characteristics Mutual information values ​​between them; , These are preset positive weighting coefficients.

6. The method according to claim 1, characterized in that, The expression for the feature importance score is: in, Indicates gene markers, Indicates the current time point, This indicates the current iteration number of the reinforcement learning model; and These are the weight coefficients related to the number of iteration steps, and and The value of varies The increase is monotonically increasing; Genes identified based on gene function databases The functional significance coefficient; For genes At the point of time The absolute value of the methylation change rate; These are preset spatial heterogeneity weighting coefficients; For genes based on the three-dimensional spatiotemporal tensor At the point of time The spatial heterogeneity index determined by the spatial distribution.

7. The method according to claim 1 or 5, characterized in that, The preset termination condition is: the value of the reward function is less than the preset rate of change threshold in N consecutive iterations, where N is a preset positive integer and the rate of change threshold is a preset positive number.

8. A risk scoring system for cervical cancer methylation detection results based on a machine learning model, characterized in that, The system includes: The tensor construction module is used to acquire methylation data and cell morphology images, and construct a three-dimensional spatiotemporal tensor from the methylation data and the cell morphology images; The causal inference module is used to calculate the methylation change rate of each gene from the time dimension of the three-dimensional spatiotemporal tensor, and to apply a non-stationary causal inference algorithm to analyze the dynamic causal relationship between the methylation change rate of each gene and cell morphological characteristics, and generate dynamic causal weight parameters. The model building module is used to obtain a reinforcement learning model by using the dynamic causal weight parameters and the statistics of the three-dimensional spatiotemporal tensor as the state space, the offset adjustment of the deformable convolution kernel and the adjustment of the causal weight decay coefficient as the action space, and the feature extraction time and cross-modal feature correlation error as the reward function. The feature extraction module is used to extract cross-modal spatiotemporal interaction features by taking the three-dimensional spatiotemporal tensor as input and applying the deformable convolution kernel. The reinforcement learning processing module is used to input the dynamic causal weight parameters, the statistics of the three-dimensional spatiotemporal tensor, and the cross-modal spatiotemporal interaction features as the current state into the reinforcement learning model. The reinforcement learning model outputs the offset adjustment amount and the decay coefficient adjustment amount. The module updates the sampling position of the deformable convolution kernel according to the offset adjustment amount and updates the decay coefficient of the dynamic causal weight parameters according to the decay coefficient adjustment amount. The parameter update module is used to update the parameters of the reinforcement learning model based on the reward function; The loop control module is used to trigger the feature extraction module, the reinforcement learning processing module and the parameter update module to execute repeatedly in sequence until a preset termination condition is met. The scoring calculation module is used to calculate the feature importance score of each gene marker in the methylation data based on the parameters of the trained reinforcement learning model.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the risk scoring method for cervical cancer methylation detection results based on a machine learning model, as described in any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the risk scoring method for cervical cancer methylation detection results based on a machine learning model, as described in any one of claims 1 to 7.