Breast cancer absolute risk prediction method based on generative representation and relative risk decomposition
By employing generative characterization and relative risk decomposition methods, the instability caused by imaging style and parameter fluctuations in breast cancer risk prediction was addressed. This enabled the generation of absolute risk sequences based on baseline incidence rates and personalized screening interval recommendations, thereby improving the accuracy and consistency of breast cancer screening management.
Patent Information
- Application Number
- CN202610140581.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-02-02
- Publication Date
- 2026-05-19
- Estimated Expiration
- 2046-02-02
AI Technical Summary
Existing methods for predicting breast cancer risk are unstable in terms of output when faced with different imaging styles and fluctuations in imaging parameters, and the relative risk results are difficult to reconcile with the baseline incidence rate in the population, making it impossible to accurately schedule screening intervals.
By using generative characterization and relative risk decomposition methods, an image generation characterization model is established to unify mammography results with different imaging styles. By combining basic clinical factors and baseline incidence rates in the population, a stable absolute risk sequence is calculated, and an individualized screening plan is output based on the risk level and screening interval mapping table.
It enables the stable generation of absolute breast cancer risk sequences on an annual timescale, provides personalized screening interval recommendations, and improves the accuracy and consistency of breast cancer screening management.
Smart Images

Figure CN121641458B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of breast imaging risk assessment technology, and in particular to a method for predicting the absolute risk of breast cancer based on generative representation and relative risk decomposition. Background Technology
[0002] Breast cancer is one of the leading malignant tumors threatening women's health. Clinically, mammography combined with basic clinical data is commonly used for screening and risk assessment. Current prevention and control practices rely on two main approaches: firstly, epidemiological models based on population statistics, which provide phased incidence rates according to factors such as age and region, guiding follow-up strategies at the average level; secondly, at the imaging diagnostic level, radiologists' experience or computer-aided detection systems identify focal lesions and morphological features such as breast density to predict an individual's short-term or medium-term likelihood of developing cancer. With the development of artificial intelligence, deep learning methods are beginning to be used to automatically extract breast structural features and tissue texture information, attempting to provide quantitative risk indicators in routine screening processes and assist in developing more individualized screening intervals and management plans.
[0003] In existing technologies, a common approach is to construct convolutional neural networks or other feature extraction models based on a single mammogram to extract characteristics such as overall breast density, focal suspicious areas, and texture distribution. These characteristics are then input into a risk prediction model along with clinical factors such as age, past medical history, and reproductive history. The model then outputs an individual breast cancer risk score or risk level through classification or regression. Some approaches, after obtaining the risk score, combine it with publicly available population incidence statistics to estimate the probability of disease within a specific time window, providing follow-up recommendations within a certain range.
[0004] However, such schemes often use fixed imaging results as input and lack systematic modeling of different imaging styles and imaging parameter fluctuations in the same examination. The model output is easily affected by changes in image style and is unstable. At the same time, the relative risk results lack a unified annual time axis alignment with the baseline incidence rate of the population divided by region and age, making it difficult to directly generate absolute risk sequences that can be used to accurately arrange screening intervals.
[0005] Therefore, a method for predicting the absolute risk of breast cancer that can overcome the shortcomings of the existing technology is a problem that needs to be solved by those skilled in the art. Summary of the Invention
[0006] One objective of this invention is to propose a technical solution that, while taking into account the baseline incidence rate of regional and age groups, can uniformly characterize and decompose the imaging information obtained from the same mammogram under different imaging styles and basic clinical factors, stably synthesize the absolute risk sequence of breast cancer on an annual timescale, and thereby derive a matching individualized screening interval.
[0007] A method for predicting absolute risk of breast cancer based on generative representation and relative risk decomposition according to an embodiment of the present invention includes:
[0008] S1. Receive mammogram images, basic clinical factors, regional identifiers, and age, and determine the baseline incidence rate and risk time point list for the population based on regional identifiers and age;
[0009] S2. Based on mammograms, establish an image generation and characterization model. Following a generation process from blurry to clear, form an imaging style variant set for the same examination. Under continuous constraints of the generation process, merge the images of the imaging style variant set to the same feature line. Extract long-term tissue feature vectors reflecting long-term tissue status from the end of the feature line and fix the model parameters.
[0010] S3. Construct a staged risk calculation framework based on long-term tissue feature vectors and basic clinical factors, calculate the relative risk coefficient for the population, perform a consistency test on the relative risk coefficient based on the imaging style variant set, obtain the coefficient difference and compare it with the consistency threshold, adjust the relative risk coefficient under the constraint of the consistency threshold, and output the corrected relative risk coefficient.
[0011] S4. Based on the corrected relative risk coefficient, the baseline incidence rate of the population, and the list of risk time points, calculate the absolute risk sequence for each time point and output the absolute risk sequence;
[0012] S5. Based on the absolute risk sequence and the classification threshold table, determine the risk level at each time point and output the risk level sequence.
[0013] S6. Calculate the screening interval and output the screening interval based on the risk level sequence and interval mapping table.
[0014] Optionally, S1 is as follows:
[0015] Using mammograms, basic clinical factors, regional identifiers, and age as inputs, we determined a risk calculation time reference starting from age, and used age as the sole starting point for the risk time point list to establish a starting time point corresponding to age.
[0016] The baseline incidence rate of the population is determined based on the regional identifier. The average level of the population is used as a reference. The regional identifier is matched one-to-one with the baseline incidence rate of the population to align the baseline incidence rate of the population with the starting time point, thus obtaining the baseline incidence rate of the population that matches the starting time point.
[0017] Based on the starting point, a list of risk points is generated on an annual basis. Multiple points are formed by recursively using an integer year sequence, maintaining a monotonically increasing chronological order, so that the risk point list covers all points of time required for subsequent absolute risk sequence calculations.
[0018] The baseline incidence rate of the population matched with the starting time point is combined with the list of risk time points as the output, which is used to calculate the absolute risk series at each time point in conjunction with the corrected relative risk coefficient, ensuring that the absolute risk series is synthesized with the baseline incidence rate of the population as a reference.
[0019] Optionally, S2 is as follows:
[0020] Mammograms were stitched together into a four-channel grayscale image from both sides and two perspectives. The image was then input into the image generation and representation model, and an extraction path was established with a convolutional neuron grid as the input layer.
[0021] Three sets of downsampling convolutional layers were used on the four-channel grayscale image, with the number of channels being 64, 128, and 256 respectively. Each set contained two layers of convolutional neurons and one layer of nonlinear transformation to extract breast contour and tissue structure features, forming shared downsampling features.
[0022] Based on shared downsampling features, three groups of generation process stages are established, and an imaging style variant set is generated in order from blurry to sharp. Continuous constraints of the generation process are calculated at each stage.
[0023] The generation process continuously constrains adjacent stages to limit the alignment error between adjacent stages, aligns the positional relationship between the breast contour and the tissue region, and allows the imaging style variant set to gradually merge into the same feature line in the generation process.
[0024] The resolution is restored step by step along the same feature line using three sets of upsampling convolutional layers with the number of channels being 128, 64 and 32 respectively. The aligned image is output and the upsampled features are fed into the feature output layer.
[0025] In the feature output layer, the upsampled features are compressed into long-term tissue feature vectors by fully connected neurons. The long-term tissue feature vectors have a dimension of 128, which correspond to the long-term state of the overall density of the breast, the strip texture and the background structure.
[0026] After obtaining the long-term tissue feature vector, the model parameters of the image generation representation model are fixed to form a fixed extraction path. The long-term tissue feature vector is used as the only image input source for calculating the relative risk coefficient, and the imaging style variant set is output for consistency verification.
[0027] Optionally, S3 specifically refers to:
[0028] Using long-term organizational feature vectors and basic clinical factors as inputs, an input layer for a staged risk calculation framework is established. The two types of vectors are concatenated to obtain a fusion vector with a dimension of 148.
[0029] The fusion vector is fed into three layers of fully connected neurons, with the number of neurons being 128, 64, and 32 respectively. Each layer is followed by a nonlinear transformation to form a feature for calculating the relative risk coefficient of the target population.
[0030] In the output layer, the relative risk coefficient for the target population is calculated with reference to the average level of the population, and the single-value coefficient is output as the initial relative risk coefficient for the same examination.
[0031] Multiple sets of long-term tissue feature vectors are generated based on the imaging style variant set. Each set of long-term tissue feature vectors is concatenated with basic clinical factors to form a fusion vector, which is then fed into a three-layer fully connected neuron to calculate the relative risk coefficient for the population and obtain a coefficient list.
[0032] Calculate the difference between the maximum and minimum values in the coefficient list to obtain the coefficient difference, and compare it with the consistency threshold to form the consistency test result;
[0033] When the coefficient difference exceeds the consistency threshold, the parameters of the staged risk calculation framework are updated parametrically under the constraint of the consistency threshold. The model parameters of the image generation characterization model are not updated, and the long-term tissue feature vector remains unchanged, so that the relative risk coefficient of the same examination meets the consistency threshold requirement.
[0034] After the parameterization update is completed, the coefficients that meet the consistency threshold are used as the corrected relative risk coefficients. The corrected relative risk coefficients are then output and used in conjunction with the baseline incidence rate of the population to calculate the absolute risk sequence.
[0035] Optionally, S4 specifically refers to:
[0036] Using the corrected relative risk coefficient, the baseline incidence rate of the population, and the list of risk time points as input, the baseline incidence rate of the population is matched with time points on an annual basis, starting from the first time point in the risk time point list, to form a time point baseline incidence rate covering all time points in the risk time point list.
[0037] The corrected relative risk coefficient is applied to the baseline incidence rate at each time point, and the risk amount at each time point is calculated item by item according to the risk time point list. The single-value expression of the corrected relative risk coefficient remains unchanged, resulting in a time point risk amount sequence with the same length as the risk time point list.
[0038] The risk quantity sequence at each time point is recursively calculated based on the order of the risk time point list. The cumulative result of the next time point is updated with the calculation result of the previous time point to obtain the absolute risk value of each time point, and then the absolute risk sequence is formed by combining them in order.
[0039] The absolute risk sequence is used as the output to determine the risk level sequence with the graded threshold table, and is used in conjunction with the interval mapping table to calculate the screening interval.
[0040] Optional, S5 specifically includes:
[0041] Using the absolute risk sequence and the graded threshold table as input, the absolute risk values are read item by item in the time order of the risk time point list. The thresholds of each level in the graded threshold table are called for matching, and a comparison object for judgment is established for each time point.
[0042] The absolute risk value at each point in time is compared with the threshold values for each level. Non-overlapping threshold ranges for each level are used. When the absolute risk value falls into a certain threshold range, it is determined to be of that level.
[0043] When the absolute risk value is equal to the boundary threshold of a certain level, it is determined to be that level. The boundary value is taken directly according to the boundary threshold, without extrapolation to adjacent levels.
[0044] The risk levels at each point in time are combined into a risk level sequence according to the time order of the risk time point list, maintaining a one-to-one correspondence with the absolute risk sequence, which is used to calculate the screening interval with the interval mapping table.
[0045] Optional, S6 specifically includes:
[0046] Using the risk level sequence and interval mapping table as input, the corresponding screening interval is searched in the interval mapping table according to the time order of the risk time point list, resulting in a list of interval values with the same length as the risk time point list.
[0047] Based on the interval value list, the minimum value is taken as the synthesis rule, and the items in the interval value list are compared in order to form a single-value screening interval, so that the screening interval maintains a definite correspondence with the risk level sequence.
[0048] The single-value screening interval is used as the output, serving as a time reference consistent with subsequent use of the absolute risk sequence.
[0049] The beneficial effects of this invention are:
[0050] 1. This proposal presents an improved method for calculating the relative risk coefficient of breast cancer. It employs generative representation models, continuous constraints in the generation process, and feature line extraction techniques to uniformly map multiple imaging styles from the same mammogram into long-term tissue feature vectors. Based on this, a population-oriented single-value relative risk coefficient is obtained through a staged risk calculation framework. By applying continuous constraints to the breast contours, tissue region positional relationships, and feature line projections of different imaging styles during the generation process, imaging style variants converge to the same feature line in the feature space, reducing the interference of imaging parameter fluctuations on the representation results. Furthermore, the relative risk coefficient is consistent with the set of imaging style variants. Only the risk calculation framework parameters are updated while the image generation representation model remains fixed, ensuring that the risk output under different imaging styles in the same examination remains within a consistent threshold range. This improvement allows the relative risk coefficient to more stably reflect the long-term state of breast density and tissue structure, rather than short-term imaging style differences. This distinguishes it from existing algorithms that directly regress risk based on single-image features, making it more suitable for repeated use in real-world screening scenarios.
[0051] 2. This proposal presents a novel method for synthesizing absolute risk of breast cancer. It employs techniques such as a population baseline incidence rate function based on regional identifiers and age, a risk point list aligned to whole ages, and amplification of relative risk values to construct individual absolute risk sequences on a unified annual timescale. By establishing a time axis with an age-rounded starting point, the corresponding population baseline incidence rate is retrieved from a pre-built database based on regional identifiers. Each year, a point-in-time baseline incidence rate corresponding to the risk point list is generated. Then, a single corrected relative risk coefficient is used to amplify and truncate the annual baseline incidence rate, constructing the annual point-in-time risk quantity. An explicit recursive formula is used to accumulate the annual absolute risk values. This method, without altering the regional and age-dependent relationship, strictly aligns individual relative risk with population baseline incidence rates, outputting an absolute risk sequence with clear annual meaning. This differs from existing schemes that only provide dimensionless risk scores or are not closely linked to population statistics, and facilitates the establishment of a directly usable quantitative bridge between public health strategies and individual follow-up plans.
[0052] 3. This proposal puts forward an integrated approach to breast cancer risk stratification and screening management. It employs techniques such as absolute risk sequence stratification, risk level-screening interval mapping, and minimum value interval synthesis to transform annual absolute risk into single-value screening intervals that match clinical follow-up procedures. In this method, the absolute risk sequence is first determined year-by-year based on a non-overlapping stratification threshold table. A left-closed, right-open interval design and boundary value direct extraction rule ensure the uniqueness of the risk level for each year and the consistency of boundary processing. Subsequently, through a pre-set interval mapping table, the risk levels of each year are converted into screening interval values for the same time unit, and the minimum value is used to synthesize a single screening interval for actual scheduling over the entire time range. Through this process, this proposal, while maintaining consistency with the baseline incidence rate and annual timeline, integrates stable relative risk representation and absolute risk calculation results into specific screening interval decisions. This differs from existing algorithms that only output risk scores or rough suggestions, and is more conducive to forming executable individualized follow-up plans in screening management scenarios. Attached Figure Description
[0053] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:
[0054] Figure 1 This is a flowchart of a method for predicting absolute risk of breast cancer based on generative characterization and relative risk decomposition proposed in this invention.
[0055] Figure 2 This is a flowchart of data acceptance and processing for a breast cancer absolute risk prediction method based on generative representation and relative risk decomposition proposed in this invention.
[0056] Figure 3 This is a flowchart of the feature extraction process for a breast cancer absolute risk prediction method based on generative representation and relative risk decomposition proposed in this invention.
[0057] Figure 4 This is a flowchart illustrating the construction of multiple sets of relative risk coefficients for a breast cancer absolute risk prediction method based on generative characterization and relative risk decomposition proposed in this invention.
[0058] Figure 5 This is a flowchart illustrating the risk output of a breast cancer absolute risk prediction method based on generative representation and relative risk decomposition proposed in this invention.
[0059] Figure 6 This is an image generation and representation model diagram of a breast cancer absolute risk prediction method based on generative representation and relative risk decomposition proposed in this invention.
[0060] Figure 7 Multi-style inspection of stability risk plots using traditional methods;
[0061] Figure 8 This invention presents a multistyle examination stable risk map of a breast cancer absolute risk prediction method based on generative representation and relative risk decomposition.
[0062] Figure 9 This is a schematic diagram of the generation process of a method for predicting absolute risk of breast cancer based on generative representation and relative risk decomposition proposed in this invention.
[0063] Figure 10 This is a schematic diagram illustrating the three-layer concept of the breast cancer absolute risk prediction method based on generative representation and relative risk decomposition proposed in this invention, from the image style layer to the tissue feature layer and then to the risk / strategy layer. Detailed Implementation
[0064] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.
[0065] refer to Figures 1 to 10 A method for predicting absolute risk of breast cancer based on generative representation and relative risk decomposition, characterized by comprising:
[0066] S1. Receive mammogram images, basic clinical factors, regional identifiers, and age, and determine the baseline incidence rate and risk time point list for the population based on regional identifiers and age;
[0067] S2. Based on mammograms, establish an image generation and characterization model. Following a generation process from blurry to clear, form an imaging style variant set for the same examination. Under continuous constraints of the generation process, merge the images of the imaging style variant set to the same feature line. Extract long-term tissue feature vectors reflecting long-term tissue status from the end of the feature line and fix the model parameters.
[0068] S3. Construct a staged risk calculation framework based on long-term tissue feature vectors and basic clinical factors, calculate the relative risk coefficient for the population, perform a consistency test on the relative risk coefficient based on the imaging style variant set, obtain the coefficient difference and compare it with the consistency threshold, adjust the relative risk coefficient under the constraint of the consistency threshold, and output the corrected relative risk coefficient.
[0069] S4. Based on the corrected relative risk coefficient, the baseline incidence rate of the population, and the list of risk time points, calculate the absolute risk sequence for each time point and output the absolute risk sequence;
[0070] S5. Based on the absolute risk sequence and the classification threshold table, determine the risk level at each time point and output the risk level sequence.
[0071] S6. Calculate the screening interval and output the screening interval based on the risk level sequence and interval mapping table.
[0072] In this specific embodiment, step S1 is as follows:
[0073] Using mammograms as image tensors (Used as input for subsequent image generation representation models, not processed in this step), basic clinical factors are used as vectors. (This is used as input for the subsequent phased risk calculation framework and is not processed in this step.) The region identifier is used as an enumeration quantity. Using age as a real-value scalar (Unit: year) As a reference for risk calculation time, whole age alignment is used, and the starting point is defined as... ,in For the floor operator, when Take when it is an integer Thus, a foundation was established. A timeline starting with [a specific timeframe] is used to generate a list of risk points and provide a unified time base for subsequent absolute risk series calculations. and It is directly passed to the image generation and characterization model and the staged risk calculation framework, maintaining a fixed extraction path and an adjustable risk mapping in terms of responsibility;
[0074] According to the region identifier Retrieve the corresponding data table from the pre-set population baseline incidence rate database, and establish a population baseline incidence rate function with the population average level as a reference. ,in For whole age, Indicates regional identifier The target population is the whole age group The one-year baseline incidence rate in the population, Perform initial alignment, take The baseline incidence rate of the population corresponding to the starting time point is consistent with the starting time point, without introducing intermediate features related to the image generation and representation model, and without changing the dependence on region and age.
[0075] Starting time point Based on this, set the annual coverage length. ,in Represents a set of positive integers, using a fixed configuration in the implementation. Define the number of time points as Generate a list of risk points in time. Recursively calculated in integer year sequence, using Forming a monotonically increasing time series, for Each item Increasing only by whole numbers in age, without inserting non-integer time points, ensures that the risk time point list maintains consistency in its annual granularity with the population-oriented relative risk coefficients output by the phased risk calculation framework. When the integer is not an integer, through The calculation starts from the whole age and aligns the first calculation year with the calendar year boundary to avoid non-annual time spans in the risk time point list;
[0076] Constructing the baseline incidence vector of the population ,use exist Take values one by one, so that and Correspondingly, with As a risk point list output, As the baseline incidence rate output for the population, it is used to calculate the corrected relative risk coefficient at each time point. In subsequent calculations, it is used as... The provided annual index will locate the year. The provided baseline incidence rates are combined with the corrected relative risk coefficients to form the point-in-time inputs required for the absolute risk sequence. By aligning the starting point, fixing the annual coverage length, and retrieving the baseline incidence rate values by region identifier, the baseline incidence rates are kept consistent with the risk point-in-time list at the annual granularity. This provides a directly usable input alignment for absolute risk synthesis based on the population-oriented relative risk coefficients output by the phased risk calculation framework.
[0077] In this specific embodiment, step S2 is as follows:
[0078] Bilateral mammograms from two different perspectives were stitched together to form a four-channel grayscale image, which was then represented as an image tensor. where pixel width pixel height Number of channels ,right Intensity normalization is performed, linearly mapping pixel intensity to... ,Will The input layer of the image generation representation model consists of a grid of convolutional neurons, using a kernel size of [size missing]. Fixed convolution kernel, for Feature extraction is performed using three groups of downsampling convolutional layers. Let the number of downsampling convolutional layers be... The number of output channels in the three groups are as follows: , , Each group contains two convolutional neurons and one nonlinear transformation layer. The outputs of the three downsampled convolutional layers are concatenated and merged along the channel dimension to obtain a shared downsampled feature map. It is used to uniformly represent the contour and tissue structure characteristics of the breast;
[0079] exist Based on this, establish three groups of generation process stages, and set the number of generation process stages. Stage Index The set of imaging style variants for the same examination is generated in order from blurry to sharp, and the number of imaging styles for the same examination is defined as follows: , constituting a set of imaging style variants Each Indicates the stage The next To achieve a controllable transition from blurry to sharp, the generation results of various imaging styles are achieved by setting the detail recovery intensity within each stage according to a fixed incremental strategy. The sequence of stage one is weak detail, stage two is medium detail, and stage three is strong detail, so that the output of each stage maintains a progressive level of detail.
[0080] During the generation process phase, the continuity constraints of the generation process are calculated, using a fixed response threshold. exist The breast tissue mask is extracted, and the largest connected region is selected as the breast tissue region through connected component screening. A fixed number of sampling points are then used on the boundary of this region. Sampling boundary pixel coordinates sequentially to form a contour coordinate set and with the geometric center of the region pixels Using this as the base point, calculate from radial distance to each boundary sampling point , forming a positional relationship vector The characteristic line is defined as a parametric curve along the path of maximum response inside the mammary gland mask. The path is The maximum response value is taken as the standard, along the Based on a fixed sampling step size Pixels are sampled at equal intervals, and the number of sampling points is set to... The intensity values at the sampling points are recorded on the aligned image to form the feature line projection vector. Each of them Normalized intensity;
[0081] To standardize the calculation and dimensions of constraints, the following continuous constraint function of the generation process is used as the only constraint term in the generation process for iterative minimization, and a threshold is used to determine whether the constraint is satisfied:
[0082] ;
[0083] in, To generate process continuity constraint functions, To generate a process stage index, For imaging style indexing, To generate the number of process stages, The number of imaging styles in the same examination. , , The weighting coefficients are non-negative and are set using spatial scale normalization. , Intensity normalization settings are adopted. , This is an absolute value summation operator that sums the absolute values of the components of a vector or coordinate sequence. This is a sum-of-squares operator that sums the squares of the components of a vector. In the stage Next A set of contour coordinates for each imaging style, whose components are pixel coordinate values. In the stage Next The tissue region location relationship vector for each imaging style, whose components are pixel coordinates and pixel distances. In the stage Next The feature line projection vectors of a certain imaging style have components that are normalized intensity values. In the stage The mean vector of the feature line projection vectors of all imaging styles is used to measure the degree of convergence to the same feature line.
[0084] An iterative minimization strategy is adopted, in which the transformation parameters of the generation process stage are updated in each iteration, so that... Monotonically decreasing, set continuous constraint threshold As a criterion for cessation, when The generation process continuously satisfies constraints, aligns the positional relationship between the breast contour and tissue region in adjacent stages, and gradually merges the projections of the imaging style variant set onto the feature line at each stage, forming a... Alignment results achieved by advancing along a single path at the top;
[0085] After completing the continuous constraints of the generation process, three sets of upsampling convolutional layers are used to restore the resolution step by step along the feature line. The number of output channels for the three sets are respectively , , The upsampled output is used to generate an aligned image according to the feature line order, and the upsampled features are fed into the feature output layer. The feature output layer uses fully connected neurons to compress the upsampled features into a long-term tissue feature vector. ,in This is a real-valued vector of dimension 128, corresponding to the long-term state of the overall breast density, stripe texture, and background structure. Through the continuous constraints and feature line advancements in the above generation process, it is made... The formation of [the image] depends solely on the alignment results within the same inspection and is unaffected by fluctuations in imaging style.
[0086] After obtaining the long-term organizational feature vector Then, the model parameters of the image generation representation model are fixed, and the model parameter set is set as follows. ,fixed This indicates that during the subsequent parameterization update process of the phased risk calculation framework, no... Update the feature vector while maintaining a fixed extraction path to ensure long-term organization. It will become the sole image input source for calculating the relative risk coefficient. With imaging style variant set As the output of this step, Used to perform consistency checks within the phased risk calculation framework, generating process continuity constraints and characteristic lines as... The formation of this model is based on the unified input of a four-channel grayscale image, the shared feature carrying of three sets of downsampling convolutional layers, the continuous constraints of three sets of generation process stages, and the resolution restoration of three sets of upsampling convolutional layers. This completes the generation process from blurry to sharp, achieving imaging style alignment within the same inspection and stably extracting long-term tissue feature vectors at the ends of feature lines. This provides a consistent input basis for the phased risk calculation framework, which is consistent with the scenario.
[0087] In this specific embodiment, step S3 is as follows:
[0088] The long-term organizational feature vector is represented as ,in Let the basic clinical factors be represented as a real-valued vector of dimension 128. ,in The vectors are real-valued vectors of dimension 20, arranged in a fixed order and normalized before input. They are then used in the input layer of the phased risk calculation framework. and The vector is obtained by concatenating the two vectors. ,in For a real-valued vector of dimension 148, The input is fed into three fully connected neurons, and intermediate feature vectors are obtained respectively. , , ,in The dimension is 128. The dimension is sixty-four. The dimension is thirty-two, with each layer followed by an element-wise nonlinear transformation, using a fixed linear rectified function. The relative risk coefficient for the target population is calculated by feeding it into the output layer. ,in Let be a positive real number, referencing the average level of the population and employing a single-node positive value constraint. The long-term tissue feature vector corresponding to the aligned image in stage three of the generation process is denoted as . ,by and In the current parameter set The forward calculation result is used as the initial relative risk coefficient, denoted as . This serves as a benchmark for subsequent consistency checks;
[0089] Based on the output of the imaging style variant set during the generation process, the imaging style variant set for the same examination is constructed. ,Will The number of samples is expressed as A fixed extraction path and a set of model parameters for the image generation and representation model are used. right Extract long-term organizational feature vectors one by one to obtain a long-term organizational feature vector set. Each of them For a real-valued vector of dimension 128, each and Concatenate to generate a fused vector Each of them A real-valued vector of dimension 148, representing the current parameter set of the phased risk calculation framework. right By performing forward calculations one by one, a list of relative risk coefficients for each population group is obtained. Each of them It is a positive real number;
[0090] right To perform a consistency check, calculate the maximum and minimum values and take the difference to obtain the coefficient difference. ,in For non-negative real numbers, set a consistency threshold. ,in For positive real numbers, when The consistency check result is recorded as passed, and the result is directly entered into the system. As the corrected relative risk coefficient output, when Then, the parameter update process is initiated, which fixes the set of model parameters for the image generation representation model. Only update the parameter set of the phased risk calculation framework. Set the learning rate With maximum number of iterations ,in It is a positive real number. For positive integers, and A single objective variable is constructed as the input, and a consistency term, a baseline preservation term, and a weight decay term are synthesized. In each iteration, the gradient signal is updated based on this objective variable. Refresh after each update And recalculate ,when or reach The iteration stops when the constraint is applied directly to the relative risk coefficient level for the population. The following unique objective function is minimized in the parameterized update process:
[0091] ;
[0092] in, The single-valued output of the objective function. , , These are non-negative weighting coefficients used to control the contributions of the consistency term, the benchmark preservation term, and the weight decay term, respectively. This represents the number of samples in the imaging style variant set from the same examination. For the summation operator, For absolute value operators, For the current parameter set The following is a fusion vector The relative risk coefficient for the target population, obtained through forward calculation, takes the value of a positive real number. For the current parameter set Below and The relative risk coefficient for the target population, obtained through forward calculation, takes the value of a positive real number. For the initial parameter set In the same input and The relative risk coefficient obtained from the above-mentioned population group takes the value of a positive real number. For parameter set The result of the sum of squares operator, The scaling constant of the parameter set is taken as the value of the L2 norm of the initial parameter set, which is used to make the weight decay term dimensionless.
[0093] After the iteration stops, use the updated parameter set. right and Perform forward computation to obtain the coefficients that satisfy the consistency threshold constraint, and denote these coefficients as... ,in For positive real numbers, The corrected relative risk coefficient output is used to calculate the absolute risk sequence in conjunction with the baseline incidence rate at each time point. Throughout the process, the staging risk calculation framework only accepts long-term tissue feature vectors and basic clinical factors as input, does not accept the original images, and does not generate a set of model parameters for the image representation model. Updates are performed to ensure that the fixed extraction path and the adjustable risk mapping are separated in terms of responsibility. The consistency check directly applies to the relative risk coefficient for the population, ensuring that input variations brought about by the imaging style variant set are constrained to a unified decision quantity expression within the same inspection, ultimately achieving... The coefficient output is used exclusively for synthesizing absolute risk sequences.
[0094] In this specific embodiment, step S4 is as follows:
[0095] The corrected relative risk coefficient is expressed as ,in The positive real number originates from the output of a single inspection within the phased risk calculation framework, representing the risk point-in-time list as follows: ,in For the first At each point in time in a year, Given a list of risk time points with a length that is a positive integer, represent the baseline incidence rate of the population as a vector. Each of them To and The corresponding one-year baseline incidence rate is within the range. ,by Starting from the first point in time, and dividing by year. Perform time point correspondence to ensure and The one-to-one correspondence remains unchanged, forming a cover. The set of baseline incidence rates at all points in time;
[0096] Will Acting on Construct a time-point risk sequence Each of them To and To maintain the validity of the probability values, a truncation threshold is set for the corresponding point-in-time risk. Calculated by term-by-term multiplication and truncation The possible values of:
[0097] ;
[0098] in, For the first The point-in-time risk for each year is a dimensionless positive real number. The corrected relative risk coefficient is a dimensionless positive real number. For the first The baseline incidence rate in the population for each year is a dimensionless positive real number. The threshold value is a dimensionless positive real number. As a minimum value operator, it returns the smaller of the two inputs. The above definition ensures that the point-in-time risk is strictly less than one and keeps the single-valued expression of the corrected relative risk coefficient unchanged in each year, so that the point-in-time risk only changes with the change of the baseline incidence rate at the point-in-time.
[0099] Based on the order of the risk time point list By performing recursion, the absolute risk value at each time point is obtained. Let the absolute risk value sequence be... ,in To and The corresponding absolute risk value adopts a cumulative update strategy starting from zero, initializing the first item and accumulating the incidence probability of the same examination over the annual dimension to form a value corresponding to the risk level. For strictly aligned annual cumulative sequences, to clarify the recursive relationship and ensure the reproducibility of the calculation, the following formula is used for year-by-year updates:
[0100] ;
[0101] in, For the first The absolute risk value for each year is a dimensionless positive real number. This represents the absolute risk value for the previous year, and is a dimensionless positive real number. For the first The point-in-time risk for each year is a dimensionless positive real number, and the product term is... The probability of no disease in the current year under the condition that no disease occurred in the previous year is a dimensionless positive real number. The above recursion ensures that the annual cumulative progresses in monotonic time order and uniquely depends on the result of the previous time point and the risk quantity at the current time point in each update, thus avoiding information leakage across years.
[0102] After the recursion is completed, according to The time sequence is combined into an absolute risk sequence, maintaining the same... The one-to-one correspondence remains unchanged. The absolute risk sequence is used as the output to determine the risk level sequence with the tiered threshold table, and is used in conjunction with the interval mapping table to calculate the screening interval. To ensure the usability of the output in subsequent determinations and mappings, it is retained by time point index. The correspondence between these parameters allows the calculation of risk level sequences and screening intervals to be directly mapped at the annual granularity.
[0103] In the implementation, the input is... As an annual index, As an annual reference incidence rate, As the only relative risk amplification factor in the same inspection, for Without interpolation or extrapolation, only by Each value is retrieved and truncated before being multiplied to avoid bias caused by inconsistencies in time scales. Calculation and The recursion is processed in a fixed order using a single channel, without parallel updates across years. This ensures that the annual calculation is completed in a fixed memory access order in the hardware implementation. Through the above process of year-to-year correspondence, item-by-item amplification and recursive accumulation, the corrected relative risk coefficient and the baseline incidence rate of the population are combined on a unified annual scale. The output absolute risk sequence directly connects to the subsequent steps of risk level sequence determination and screening interval calculation.
[0104] In this specific embodiment, step S5 is as follows:
[0105] The absolute risk sequence is represented as Each of them These are dimensionless positive real numbers derived from absolute risk values calculated recursively over the years. The risk point-in-time list is represented as follows: ,in For the first At each point in time in a year, Given a list of risk points with a length that is a positive integer, represent the tiered threshold table as follows: ,in As a risk level identifier, For the number of levels, The lower threshold of this level, This is the upper threshold for this level, with a value range of [value range missing]. Using non-overlapping level threshold intervals, Arrange intervals from low to high according to the threshold and create an interval set. ,in For , Used for the highest level to include the upper bound, for and Alignment, by Read item by item in chronological order And construct comparison objects for each point in time. The input set used to drive a single decision;
[0106] The absolute risk value at each time point is compared with the threshold values for each level, using a single-channel sequential search, starting from... arrive Determine in sequence, if Fall into a Then it is determined to be a level. Immediately stop the search at that point in time to ensure uniqueness of the determination. Since the intervals were arranged according to the non-overlapping principle during construction, any... The matching result is a unique level identifier and no duplicate matching occurs. To avoid double assignment at the boundary, the internal interval is defined as left-closed and right-open, and only the highest level interval is defined as right-closed to ensure that the upper bound is contained.
[0107] when When the threshold value equals that of a certain level, the decision is made according to the direct threshold determination rule. or Then according to Assignment based on the endpoints of the closed interval: if it falls on the left endpoint, it is assigned to the current level. When it falls on the right closed endpoint of the highest level, it belongs to the highest level. Under this rule, no extrapolation to neighboring levels is performed, and the numerical expression of the threshold and interval boundary is not adjusted, ensuring the consistency of boundary value processing.
[0108] The risk levels at each point in time are grouped into a risk level sequence according to the chronological order of the risk time point list. Each of them For at time point The obtained level identifier is derived from... ,Keep and One-to-one correspondence, record Triples are used to support subsequent mappings. The screening interval is calculated using the interval mapping table as input. The interval mapping table is represented as follows: ,in For level The corresponding screening interval value, in months or years, is set as a fixed unit in the system configuration and reads items sequentially by year. ,exist Search and Consistent Rank Items Output As the screening interval at that point in time, For at time point The screening interval value, in units of Consistent, used to form screening interval sequences and maintain consistency with The indexes are consistent;
[0109] In implementation, The provided annual index drives the traversal, The provided absolute risk value is compared with a threshold value to determine the risk level. The provided interval set is used for level determination, in order to Perform screening interval mapping and prohibit modification during the comparison process. The values should not be interpolated or extrapolated, and adjustments are prohibited after the interval is constructed. and The boundary attribution is ensured through the above process of reading at specific times, interval matching, direct boundary retrieval, and sequence combination, so as to ensure that the risk level sequence and the absolute risk sequence are in one-to-one correspondence, and the schedule for follow-up and re-examination is output in fixed unit screening intervals.
[0110] In this specific embodiment, step S6 is as follows:
[0111] The risk level sequence is represented as Each of them For at time point The risk level identifier represents the list of risk points as follows: ,in For the first At each point in time in a year, Given a list of risk points with a length that is a positive integer, represent the interval mapping table as follows: ,in As a risk level identifier, For level The corresponding screening interval value is set to a fixed unit in the system configuration and remains unchanged. The number of levels is a positive integer. and Perform index alignment during traversal. Read item by item in chronological order This ensures that a unique level identifier exists for each point in time for lookup. Perform a key-value lookup, and Perform a match and retrieve the matching pairs. Removed later And defined as the screening interval value at that point in time. All time points Arranged into a list of interval values in chronological order. Each of them For time point The corresponding screening interval value, unit and Consistency remains unchanged; to ensure the determinism of the search, it is required that... cover All level identifiers in the code will not be temporarily rewritten or extrapolated;
[0112] List of interval values Based on this, a minimum value synthesis rule is adopted to form a single-value screening interval. Time sequence Perform a single-channel traversal, initializing the current minimum value as the first item. In the At this point in time Compare with the current minimum value, if If it is less than the current minimum value, then use Cover the current minimum value, otherwise leave it unchanged. After completing the traversal, obtain the single-value screening interval. ,in To and For positive real numbers with consistent units, to avoid comparison errors caused by inconsistent units, the units must be checked before entering the composition rule. The system performs a consistency check on the units; if the system is configured as "month", then all units will be checked. Compare using "months" as the unit; if the system is configured to "years", then all... Comparing in "years", the composition rule of taking the minimum value is used to make... and To maintain a definite correspondence within the same inspection and avoid ambiguity due to changes over time, the minimum value operator is denoted as [insert operator here]. Its function is to return the minimum value in the input set, which is achieved in this embodiment through sequential comparison. It does not introduce parallel reduction or randomization processes;
[0113] Single-value screening interval As an output, a time reference consistent with subsequent use of the absolute risk sequence is provided to ensure... It can be directly entered into the timeline of the follow-up plan and recorded during output. The cooperative relationship, among which Used to generate the time interval for the next screening. Provide the year index corresponding to the current check; the output stage is incorrect. Interpolation, extrapolation, or rounding are performed, and the expression is based solely on the fixed units configured by the system. When the system configuration requires integer units, the expression is performed before entering the composition rule. Unit consistency is achieved and a fixed integer constraint with rounding up is adopted to ensure execution in a defined storage and scheduling manner in the hardware implementation;
[0114] In the implementation, a linear process of index-driven and key-value lookup is adopted. The provided annual index drives the traversal, Use the provided level identifier to search, The provided mapping between levels and intervals is used to retrieve values. Perform minimum value synthesis; the entire process remains unchanged. and The content does not dynamically rescale or switch units for interval values, ensuring that interval calculations within the same check rely solely on established mapping relationships and defined synthesis rules. Through the linear processing of item-by-item lookup, unit consistency, and minimum value synthesis, the single-value screening interval is obtained. It outputs in fixed units, providing a unified time reference and clear scheduling interval for subsequent use with the absolute risk sequence. Further explanation of the character meanings is as follows: As a risk level sequence, List of risk points, For the first At each point in time in a year, The length of the risk point list, For interval mapping table, As a risk level identifier, This represents the screening interval value. This is a list of interval values. For at time point The screening interval value, For single-value screening intervals, It is the minimum value operator.
[0115] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A method for predicting absolute risk of breast cancer based on generative representation and relative risk decomposition, characterized in that, include: S1. Receive mammogram images, basic clinical factors, regional identifiers, and age, and determine the baseline incidence rate and risk time point list for the population based on regional identifiers and age; S2. Based on mammograms, establish an image generation and characterization model. Following a generation process from blurry to clear, form an imaging style variant set for the same examination. Under continuous constraints of the generation process, merge the images of the imaging style variant set to the same feature line. Extract long-term tissue feature vectors reflecting long-term tissue status from the end of the feature line and fix the model parameters. S3. Construct a staged risk calculation framework based on long-term tissue feature vectors and basic clinical factors, calculate the relative risk coefficient for the population, perform a consistency test on the relative risk coefficient based on the imaging style variant set, obtain the coefficient difference and compare it with the consistency threshold, adjust the relative risk coefficient under the constraint of the consistency threshold, and output the corrected relative risk coefficient. S4. Based on the corrected relative risk coefficient, the baseline incidence rate of the population, and the list of risk time points, calculate the absolute risk sequence for each time point and output the absolute risk sequence; S5. Based on the absolute risk sequence and the classification threshold table, determine the risk level at each time point and output the risk level sequence. S6. Calculate the screening interval and output the screening interval based on the risk level sequence and interval mapping table; S2 includes: Mammograms were stitched together into a four-channel grayscale image from both sides and two perspectives. The image was then input into the image generation and representation model, and an extraction path was established with a convolutional neuron grid as the input layer. Three groups of downsampling convolutional layers were used on the four-channel grayscale image, with the number of channels being 64, 128, and 256 respectively. Each group contained two layers of convolutional neurons and one layer of nonlinear transformation to extract breast contour and tissue structure features, forming shared downsampling features. Based on shared downsampling features, three groups of generation process stages are established, and an imaging style variant set is generated in order from blurry to sharp. Continuous constraints of the generation process are calculated at each stage. The generation process continuously constrains adjacent stages to limit the alignment error between adjacent stages, aligns the positional relationship between the breast contour and the tissue region, and allows the imaging style variant set to gradually merge into the same feature line in the generation process. The resolution is restored step by step along the same feature line using three sets of upsampling convolutional layers with the number of channels being 128, 64 and 32 respectively. The aligned image is output and the upsampled features are fed into the feature output layer. In the feature output layer, the upsampled features are compressed into long-term tissue feature vectors by fully connected neurons. The long-term tissue feature vectors have a dimension of 128, which correspond to the long-term state of the overall density of the breast, the strip texture and the background structure. After obtaining the long-term tissue feature vector, the model parameters of the image generation and characterization model are fixed to form a fixed extraction path. The long-term tissue feature vector is used as the only image input source for calculating the relative risk coefficient, and the imaging style variant set is output for consistency verification.
2. The method for predicting absolute risk of breast cancer based on generative representation and relative risk decomposition according to claim 1, characterized in that, S1 specifically refers to: Using mammograms, basic clinical factors, regional identifiers, and age as inputs, we determined a risk calculation time reference starting from age, and used age as the sole starting point for the risk time point list to establish a starting time point corresponding to age. The baseline incidence rate of the population is determined based on the regional identifier. The average level of the population is used as a reference. The regional identifier is matched one-to-one with the baseline incidence rate of the population to align the baseline incidence rate of the population with the starting time point, thus obtaining the baseline incidence rate of the population that matches the starting time point. Based on the starting point, a list of risk points is generated on an annual basis. Multiple points are formed by recursion in an integer year sequence, maintaining a monotonically increasing time order, so that the risk point list covers all points of time required for subsequent absolute risk sequence calculation. The baseline incidence rate of the population matched with the starting time point is combined with the list of risk time points as the output, which is used to calculate the absolute risk series at each time point in conjunction with the corrected relative risk coefficient, ensuring that the absolute risk series is synthesized with the baseline incidence rate of the population as a reference.
3. The method for predicting absolute risk of breast cancer based on generative representation and relative risk decomposition according to claim 1, characterized in that, S3 specifically refers to: Using long-term organizational feature vectors and basic clinical factors as inputs, an input layer for a staged risk calculation framework is established. The two types of vectors are concatenated to obtain a fusion vector with a dimension of 148. The fusion vector is fed into three layers of fully connected neurons, with the number of neurons being 128, 64, and 32 respectively. Each layer is followed by a nonlinear transformation to form a feature for calculating the relative risk coefficient of the target population. In the output layer, the relative risk coefficient for the target population is calculated with reference to the average level of the population, and the single-value coefficient is output as the initial relative risk coefficient for the same examination. Multiple sets of long-term tissue feature vectors are generated based on the imaging style variant set. Each set of long-term tissue feature vectors is concatenated with basic clinical factors to form a fusion vector, which is then fed into a three-layer fully connected neuron to calculate the relative risk coefficient for the population and obtain a coefficient list. Calculate the difference between the maximum and minimum values in the coefficient list to obtain the coefficient difference, and compare it with the consistency threshold to form the consistency test result; When the coefficient difference exceeds the consistency threshold, the parameters of the staged risk calculation framework are updated parametrically under the constraint of the consistency threshold. The model parameters of the image generation characterization model are not updated, and the long-term tissue feature vector remains unchanged, so that the relative risk coefficient of the same examination meets the consistency threshold requirement. After the parameterization update is completed, the coefficients that meet the consistency threshold are used as the corrected relative risk coefficients. The corrected relative risk coefficients are then output and used in conjunction with the baseline incidence rate of the population to calculate the absolute risk sequence.
4. The method for predicting absolute risk of breast cancer based on generative representation and relative risk decomposition according to claim 1, characterized in that, S4 specifically refers to: Using the corrected relative risk coefficient, the baseline incidence rate of the population, and the list of risk time points as input, the baseline incidence rate of the population is matched with time points on an annual basis, starting from the first time point in the risk time point list, to form a time point baseline incidence rate covering all time points in the risk time point list. The corrected relative risk coefficient is applied to the baseline incidence rate at each time point, and the risk amount at each time point is calculated item by item according to the risk time point list. The single-value expression of the corrected relative risk coefficient remains unchanged, resulting in a time point risk amount sequence with the same length as the risk time point list. The risk quantity sequence at each time point is recursively calculated based on the order of the risk time point list. The cumulative result of the next time point is updated with the calculation result of the previous time point to obtain the absolute risk value of each time point, and then the absolute risk sequence is formed by combining them in order. The absolute risk sequence is used as the output to determine the risk level sequence with the graded threshold table, and is used in conjunction with the interval mapping table to calculate the screening interval.
5. The method for predicting absolute risk of breast cancer based on generative representation and relative risk decomposition according to claim 1, characterized in that, S5 specifically refers to: Using the absolute risk sequence and the graded threshold table as input, the absolute risk values are read item by item in the time order of the risk time point list. The thresholds of each level in the graded threshold table are called for matching, and a comparison object for judgment is established for each time point. The absolute risk value at each point in time is compared with the threshold values for each level. Non-overlapping threshold ranges for each level are used. When the absolute risk value falls into a certain threshold range, it is determined to be of that level. When the absolute risk value is equal to the boundary threshold of a certain level, it is determined to be that level. The boundary value is taken directly according to the boundary threshold, without extrapolation to adjacent levels. The risk levels at each point in time are combined into a risk level sequence according to the time order of the risk time point list, maintaining a one-to-one correspondence with the absolute risk sequence, which is used to calculate the screening interval with the interval mapping table.
6. The method for predicting absolute risk of breast cancer based on generative representation and relative risk decomposition according to claim 1, characterized in that, S6 specifically refers to: Using the risk level sequence and interval mapping table as input, the corresponding screening interval is searched in the interval mapping table according to the time order of the risk time point list, resulting in a list of interval values with the same length as the risk time point list. Based on the interval value list, the minimum value is taken as the synthesis rule, and the items in the interval value list are compared in order to form a single-value screening interval, so that the screening interval maintains a definite correspondence with the risk level sequence. The single-value screening interval is used as the output, serving as a time reference consistent with subsequent use of the absolute risk sequence.