Cascading type multi-factor bone age evaluation device, method and equipment based on space-frequency conversion network and medium

Through the cascading multi-factor bone age assessment method based on the space-frequency conversion network, the problems of cumbersome and strong subjectivity of traditional bone age assessment methods are solved, and more efficient and accurate bone age assessment is achieved. Combined with a variety of bone development impact factors, the development of intelligent medicine is promoted.

CN120125567APending Publication Date: 2025-06-10ZHENJIANG NO 1 PEOPLES HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510291377.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

Traditional bone age assessment methods are cumbersome, rely on doctors’ professional knowledge, are susceptible to subjective factors, and have failed to fully consider a variety of factors affecting bone development.

Method used

Using a cascaded multi-factor bone age evaluation method based on the space-frequency conversion network, bone age evaluation is divided into three cascade functional steps: bone marker detection, fine-grained multi-factor bone age score and standardized control prediction. Combined with the patient's gender, genetic factors, hormone levels and other factors, images and factor characteristics are extracted through the space-frequency conversion network, multi-factor joint characteristics are calculated, and bone age scores and development trend prediction are finally carried out.

Benefits of technology

It improves the accuracy and efficiency of bone age assessment, reduces the impact of subjective factors on the results, and can more accurately consider a variety of impact factors on bone development, which promotes the development of intelligent medicine.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125567A_ABST
    Figure CN120125567A_ABST
Patent Text Reader

Abstract

The invention discloses a cascaded multi-factor bone age evaluation device, method and equipment based on a space-frequency conversion network and a medium, and the method comprises the steps: extracting a detection region pool of a skeleton marker composed of a plurality of predefined detection frames from a processed X-ray image, and for each predefined detection frame, screening out the skeleton marker through a space-frequency conversion network A; inputting a certain skeleton marker region into a space-frequency conversion network B, extracting region features, inputting various skeleton development influence factors into a factor feature extraction network to obtain factor features, and further determining multi-factor joint features; inputting the multi-factor joint features into a bone age scoring network to obtain a scoring result of the nth bone mark area, and adding the scoring results of all the bone mark areas to obtain a final bone age score; determining the bone physiological age according to the bone age score, and predicting the development trend of the bone based on the bone physiological age and the actual age. According to the method, the bone age evaluation precision and efficiency are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the application of artificial intelligence technology in the medical field, and specifically to a cascaded multi-factor bone age assessment device, method, equipment and medium based on a space-frequency conversion network. Background Art

[0002] Bone age estimation is the process of determining a person's biological age by analyzing bone development, especially for children and adolescents. This assessment is usually performed by observing X-rays of the wrist, hand or other parts of the body to assess the maturity of the bones. Bone age may not be completely consistent with actual age (i.e. physiological age). Bone age estimation can help determine whether there are growth and development abnormalities or assess a child's growth potential. In medicine, bone age estimation is often used in the following situations: (1) Growth and development assessment: Doctors use bone age estimation to assess children's growth and development to help determine whether there is developmental delay or advancement, especially in the diagnosis of abnormal height or puberty. (2) Diagnosis and treatment: Bone age estimation helps diagnose endocrine diseases related to growth and development, such as growth hormone deficiency, precocious puberty, thyroid dysfunction, etc., and assists in specifying treatment plans; (3) Bone health assessment: Some chronic diseases or hormonal problems may affect the development of bone age. Estimating bone age helps to detect these problems early; (4) Forensic medicine: Bone age estimation is used to determine the age of an individual, especially in the absence of identity documents, which has important legal significance.

[0003] The Greulich-Pyle (GP) method and the Tanner-Whitehouse (TW) method are two common traditional methods for assessing bone age. Their core goal is to estimate the biological age of children and adolescents by analyzing their skeletal development. The GP method is one of the earliest and most commonly used methods for assessing bone age, especially in radiology. This method assesses bone age by comparing the skeletal development in X-rays with a set of standard control atlases. The GP method relies on a set of standard bone age atlases constructed from a large number of children's X-rays. These atlases show the typical development of the wrist and hand bones of children of different ages. Doctors will take X-rays of children's wrists or hands and compare them with images of different age groups in the standard atlas to find the most similar images. The TW rule is more complex and is a finely quantified bone age assessment method. It not only relies on X-ray images, but also combines different signs of bone development to assess bone age. This method focuses on multiple bone structures in the hand, assesses the maturity of each bone, and scores each bone based on the development of the bones in the X-ray. Each bone development stage corresponds to a specific score, and the scores of all bones are added together to get a total score, which represents the bone age. The TW method usually sets different scoring standards for children of different age groups, so it can provide a more accurate bone age assessment.

[0004] The above traditional evaluation method has a rather cumbersome process, highly depends on the professional knowledge and experience level of doctors, and is easily affected by subjective factors. With the rapid development of current deep learning and computer-aided diagnosis technologies, intelligent bone age estimation has become a research content with important clinical significance. Deep learning models can capture complex patterns in images through multi-level feature learning. Especially when dealing with complex or irregular bone development situations, deep learning models can often better identify and predict bone age, reducing human errors. Once the deep learning model is trained, it can automatically perform bone age evaluation, eliminating the process of manually comparing atlases and scoring. This greatly reduces the workload of clinicians, especially when performing bone age evaluation among a large number of patients, and can significantly improve work efficiency. Traditional bone age evaluation methods usually rely on doctors' experience and judgment, with certain subjectivity. Especially for cases with ambiguous development stages, the evaluation results may be biased. While deep learning models can automatically learn based on a large amount of labeled data and evaluate with a consistent standard, reducing the influence of subjective factors on the results. Summary of the Invention

[0005] In view of this, the present invention provides a cascaded multi-factor bone age evaluation device, method, equipment and medium based on a spatio-frequency conversion network. The present invention divides bone age evaluation into three cascaded functional steps: skeletal landmark detection, fine-grained multi-factor bone age scoring, and standardized control prediction, effectively improving the accuracy of the bone age evaluation task and approaching the actual clinical needs. At the same time, in order to further improve the calculation efficiency and personalized evaluation performance, the present invention also proposes a spatio-frequency conversion network to realize the conversion of the neural network in the spatial domain and the frequency domain, and combines various additional influencing factors of bone development such as the patient's gender, genetic factors, and hormone levels.

[0006] The present invention achieves the above technical objectives through the following technical means.

[0007] A cascaded multi-factor bone age evaluation method based on a spatio-frequency conversion network:

[0008] Take a skeletal X-ray image and process it;

[0009] Extract a detection region pool of skeletal landmarks from the processed X-ray image. The detection region pool of skeletal landmarks consists of multiple predefined detection boxes. Specifically, 9 predefined detection boxes are generated at each pixel position of the X-ray image; for each predefined detection box, use the spatio-frequency conversion network A to screen and locate the skeletal landmarks from it;

[0010] Input a certain skeletal landmark region into the spatio-frequency conversion network B to extract regional features Meanwhile, input multiple bone development influencing factors into the factor feature extraction network to obtain factor features. Calculate the shared information between the region features and the factor features. Utilize the shared information, the remaining information of the region features, and the remaining information of the factor features to calculate the image weight and the factor weight; from the image weight, the region features, the factor weight, and the factor features, through element-wise summation operations, calculate the multi-factor joint features.

[0011] Input the multi-factor joint features into the bone age scoring network for inference to obtain the scoring results of the nth bone landmark region, and add up the scoring results of all bone landmark regions to obtain the final bone age score.

[0012] Map the final bone age score to the industry standard quantification index to obtain the skeletal physiological age assessment result, and place the coordinates composed of the skeletal physiological age and the actual age in the standard skeletal development curve for comparison to predict the development trend of the skeleton.

[0013] Furthermore, the structures of the spatio-frequency conversion network A and the spatio-frequency conversion network B are the same, and both are composed of an input layer, five spatio-frequency conversion layers, two max pooling layers, and an output layer; after the input layer receives the input image, high-dimensional image features are extracted through five spatio-frequency conversion layers; a max pooling layer is connected respectively after the third spatio-frequency conversion layer and the fourth spatio-frequency conversion layer for downsampling operations; the output layer outputs the results.

[0014] Furthermore, in the spatio-frequency conversion layer, for the input image I(x, y) and the conversion kernel K(x, y), first, through Fourier transform, map the information in the spatial domain to the frequency domain:

[0015]

[0016] Among them, is the image in the frequency domain, is the conversion kernel in the frequency domain, and the function represents the Fourier transform;

[0017] After that, perform frequency domain feature extraction operations, and perform pixel-wise multiplication on the input image and the conversion kernel in the frequency domain:

[0018]

[0019] Among them, is the feature of the input image in the frequency domain;

[0020] Then, convert the feature back to the spatial domain through inverse Fourier transform:

[0021]

[0022] Among them, C(x, y) is the spatial domain image feature after spatio-frequency conversion, and the function represents the inverse Fourier transform operation.

[0023] Furthermore, the method for screening and locating bone markers by using the spatio-frequency conversion network A is specifically as follows: the spatio-frequency conversion network A classifies each predefined detection box according to the marker type, and the sum of all class probabilities is 1. If the maximum value of the class probability is higher than the threshold, the detection box is retained, and the offset error between the retained detection box and the true bounding box is calculated to accurately locate the bone markers.

[0024] Furthermore, the shared information between the region feature and the factor feature

[0025]

[0026] is: is the information entropy of the image feature, is the information entropy of the factor feature, is the joint information entropy of the image feature and the factor feature, is the probability distribution of the image feature, is the probability distribution of the factor feature, is the joint distribution of the image feature and the factor feature.

[0027] Furthermore, the remaining information of the region feature and the remaining information

[0028]

[0029] of the factor feature image weight

[0030] A cascade multi-factor bone age assessment device based on a spatio-frequency conversion network, comprising:

[0031] A data processing module that standardizes, denoises, and normalizes the skeletal X-ray image;

[0032] A bone marker detection module that detects bone markers from the processed X-ray image by using the spatio-frequency conversion network A;

[0033] A fine-grained multi-factor bone age scoring module that designs a multi-factor joint algorithm and comprehensively scores multiple bone marker regions by combining various factors affecting bone development;

[0034] The standardized control prediction module maps the bone age score to a standardized evaluation index to obtain the skeletal physiological age, and predicts the development trajectory of the skeleton from the skeletal physiological age and the actual age.

[0035] An electronic device includes a memory and a processor;

[0036] The memory is used to store a computer program;

[0037] The processor is used to execute the computer program and implement the above-mentioned cascade multi-factor bone age evaluation method based on the spatio-frequency conversion network when executing the computer program.

[0038] A storage medium stores a computer program, and when the computer program is executed by a processor, the processor is caused to execute the cascade multi-factor bone age evaluation method based on the spatio-frequency conversion network as described above.

[0039] The beneficial effects of the present invention are as follows:

[0040] (1) The present invention simulates the real situation of the clinical bone age assessment task, systematically divides the bone age assessment into three stages (i.e., detecting markers, multi-factor bone age scoring, and standardized control prediction), and combines advanced technologies of artificial intelligence, greatly improving the accuracy and efficiency of bone age assessment and promoting the further development of intelligent medicine.

[0041] (2) The present invention designs a spatio-frequency conversion network based on the convolution theorem, converts the complex convolution operation of medical images in the traditional spatial domain into a simple multiplication operation in the frequency domain, reduces the computational complexity of the network from O(N 2 ) to O(NlogN), thereby improving the overall inference efficiency of the model to ensure the real-time requirements of bone age assessment in clinical practice.

[0042] (3) Existing bone age intelligent assessment systems often only input the left hand X-ray image of the patient into the neural network for analysis. However, in real medical practice, bone age assessment also needs to consider factors affecting skeletal development such as the patient's gender, genetic factors, and hormone levels. The present invention designs a multi-factor combined bone age scoring method based on this medical reality, fully considering various skeletal development factors with medical significance, and thus giving a more accurate bone age assessment result.

[0043] (4) The present invention has important clinical application significance, can be deployed in various medical scenarios, reduce the workload of doctors, optimize the allocation of medical resources, and improve the overall efficiency of the hospital. At the same time, it also saves a large amount of labor costs for the hospital, expands the business scope of the hospital, and has great commercial value. Description of the Drawings

[0044] Figure 1Flow chart of cascade multi-factor bone age assessment based on spatio-frequency conversion network according to the present invention;

[0045] Figure 2 Schematic diagram of the structure of the spatio-frequency conversion network according to the present invention;

[0046] Figure 3 Flow chart of skeletal marker detection according to the present invention;

[0047] Figure 4 Flow chart of fine-grained multi-factor bone age scoring according to the present invention;

[0048] Figure 5 Schematic diagram of the multi-factor joint inference algorithm according to the present invention;

[0049] Figure 6 Flow chart of standardized control prediction according to the present invention. Detailed implementation manners

[0050] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, but the protection scope of the present invention is not limited thereto.

[0051] As Figure 1 shown, a cascade multi-factor bone age assessment method based on a spatio-frequency conversion network according to the present invention specifically includes the following steps:

[0052] Step (1), taking and processing skeletal X-ray images, including standardization, denoising, and normalization operations.

[0053] In order to effectively analyze the images, it is necessary to perform standardization processing on the input X-ray images so that the image size is unified to 512×512. Assuming that the size of the original input X-ray image is H×W, the scaling ratio of the image can be calculated by the following method:

[0054]

[0055] where r h and r w are the scaling ratios of the X-ray image in the vertical and horizontal directions respectively.

[0056] After that, the bilinear interpolation method is used to calculate the position of the pixel point (x ′ , y ′ ) in the scaled image corresponding to the pixel point (x, y) in the original input X-ray image:

[0057]

[0058] Due to the influence of X-ray imaging conditions, there may be different degrees of noise in the image. The GaussianBlur() function in the OpenCV toolkit is called to perform Gaussian filtering on the image, thereby achieving the denoising operation.

[0059] Due to the differences in X-ray imaging equipment, there may be differences between data from different sources, which affect the learning efficiency of the model. Therefore, a normalization operation is adopted to unify the pixel values of all X-ray images between 0 and 1. The implementation method is as follows:

[0060]

[0061] Among them, min(X) and max(X) represent the minimum pixel value and the maximum pixel value in the X-ray image, respectively.

[0062] Step (2), construct a spatial-frequency conversion network.

[0063] In order to reduce the computational overhead of the network and improve the bone age assessment efficiency, the present invention constructs a spatial-frequency conversion network. According to the convolution theorem, the convolution calculation in the spatial domain is equivalent to the product operation in the frequency domain, and its formulaic expression is as follows:

[0064] F{f(t)*g(t)}=F{f(t)}·F{g(t)}

[0065] Among them, F{f(t)} and F{g(t)} are the frequency domain transforms of the functions f(t) and g(t) respectively, * and · represent convolution and product operations respectively, and f(t), g(t) represent the function representations of the spatial domain data.

[0066] Since the convolution in the spatial domain is converted into the product operation in the frequency domain, constructing a neural network in the frequency domain can reduce the computational complexity.

[0067] The architecture of the spatial-frequency conversion network is as Figure 2 shown. The network consists of an input layer, five spatial-frequency conversion layers, two max-pooling layers, and an output layer.

[0068] The conversion kernel size of each spatial-frequency conversion layer is 3×3, the stride is 1, and the activation function is the ReLU function. In the spatial-frequency conversion layer, for the input image I(x, y) and the conversion kernel K(x, y), first, the information in the spatial domain is mapped to the frequency domain through Fourier transform, and its calculation is as follows:

[0069]

[0070] Among them, and are the images and conversion kernels in the frequency domain respectively, and the function Denote the Fourier transform, which is implemented by calling the fft module in the numpy toolkit. Then, perform frequency-domain feature extraction operations, and perform pixel-by-pixel multiplication on the input image and the transformation kernel in the frequency domain:

[0071]

[0072] Among them, is the feature of the input image in the frequency domain, and then the feature is transformed to the spatial domain through the inverse Fourier transform:

[0073]

[0074] Among them, C(x, y) is the spatial-domain image feature after the spatial-frequency conversion, and the function represents the inverse Fourier transform operation, which is implemented by calling the fft.iff function in the numpy toolkit.

[0075] After the input layer receives the input image, high-dimensional image features are extracted through five spatial-frequency conversion layers. The output channel numbers of the five spatial-frequency conversion layers are 64, 128, 256, 512, and 1024 respectively. After the third and fourth spatial-frequency conversion layers, a max pooling layer is connected respectively for downsampling operations to further reduce the computational overhead; the pooling size of the max pooling layer is 2×2, the stride is 2, and the output sizes are 256×256×256 and 128×128×1024 respectively.

[0076] The output layer consists of a Flatten layer, two fully connected layers, and a classification layer. The Flatten layer first flattens the features extracted by the spatial-frequency conversion layer into a one-dimensional tensor, and then inputs them into two fully connected layers. The two fully connected layers have 512 and 256 neurons respectively, and the activation function is the ReLU function. Finally, it enters the classification layer, and the softmax function is used to output the prediction result.

[0077] Step (3), detect bone markers.

[0078] Next, perform bone marker detection to discover the markers used for bone age assessment in the image, including carpal bones, articular cartilage, and growth plates of various bones, etc. The process is as Figure 3 shown. Specifically:

[0079] 1) Extract the detection region pool of bone markers

[0080] Extract a large number of predefined detection boxes from the processed X-ray image to form the detection region pool of bone markers. Specifically, 9 predefined detection boxes are generated at each pixel position (x, y) of the X-ray image corresponding to 3 different scales and 3 different aspect ratio regions, which can be expressed as:

[0081]

[0082] Among them, w k and h k respectively represent the width and height of the k-th predefined detection box.

[0083] 2) Screening target skeletal markers

[0084] For each predefined detection box in the detection region pool Use the spatio-temporal conversion network A to screen and locate the skeletal markers from it. The spatio-temporal conversion network A first classifies each predefined detection box according to the marker type. The sum of all class probabilities is 1. If the maximum class probability is higher than the threshold σ (empirical value), then retain this detection box. After that, calculate the offset error L i , y i , w i , h i ) between the retained detection box and the true bounding box (x loc (i.e., the localization loss) to accurately locate the skeletal marker. The calculation method is as follows:

[0085]

[0086] Among them, are the parameters of the retained detection box, that is, the position, width, and height.

[0087] 3) Definition of loss function

[0088] In the training process of the spatio-temporal conversion network A during skeletal marker detection, the loss function is the sum of the classification loss and the localization loss of the predefined detection box:

[0089] L = L cls + L loc

[0090] Among them, the classification loss L cls adopts the standard cross-entropy loss.

[0091] Step (4), fine-grained multi-factor bone age scoring.

[0092] After extracting the skeletal markers in step (3), a processed X-ray image is divided into a set of multiple target regions (i.e., the screened skeletal marker regions):

[0093]

[0094] Next, a fine-grained analysis and quantitative scoring will be performed on each skeletal marker to obtain the final bone age prediction result. As Figure 4As shown, the present invention proposes a bone age scoring method combining multiple factors, which not only considers obtaining valuable medical information from images, but also comprehensively scores by combining multiple factors affecting skeletal development such as gender, genetic factors, and hormone levels. For ease of description, in this embodiment, the symbol H is used to represent one or more of the multiple factors affecting skeletal development, and any number of factors can be combined according to this method in actual applications.

[0095] 1) Feature extraction

[0096] For a certain skeletal marker region r n , input it into the spatio-temporal conversion network B to extract the region features At the same time, input the factor H into the factor feature extraction network ω to obtain the factor features The factor feature extraction network consists of an input layer, a fully connected layer, and an output layer. And After passing through their respective feature extraction networks, they are both mapped to a feature tensor of 128×128×1024.

[0097] 2) Multi-factor combination

[0098] As Figure 5 shown, the present invention proposes a multi-factor joint inference algorithm to combine multiple factors H affecting skeletal development and the images of multiple skeletal marker regions.

[0099] Calculate the shared information between the region features and the factor features in the following way

[0100]

[0101] Where is the information entropy of the image features, is the information entropy of the factor features, is the joint information entropy of the image features and the factor features, And are respectively 's probability distributions, is And 's joint distribution. The shared information reflects the degree of information overlap between the image and the additional factor. More importantly, after removing the shared information, the remaining information of the image features and the remaining information of the factor features The more the remaining information, the greater the information contribution of the image or a certain factor to the bone age score and the greater the proportion. The two remaining information are calculated in the following way:

[0102]

[0103] Calculate the image and factor weights using the shared information and the remaining information:

[0104]

[0105] Therefore, the final multi-factor joint feature is:

[0106]

[0107] where, represents the element-wise summation operation.

[0108] 3) Joint feature inference

[0109] Input the obtained multi-factor joint feature Z into the bone age scoring network for inference. The bone age scoring network consists of two fully connected layers, with 512 and 256 neurons respectively, and the activation function is the ReLU function for both; finally, obtain the scoring result s of the nth bone landmark region through the softmax function n , and sum up the scoring results of all bone landmark regions to obtain the final bone age score:

[0110]

[0111] Constrain the training of the bone age scoring network through the following loss function:

[0112]

[0113] where, N represents the total number of training samples, and S n is the true label of the bone age score of the nth sample.

[0114] Step (5), standardized control prediction.

[0115] Map the final bone age score obtained in step (4) to the industry standard quantization index to obtain a medically significant skeletal physiological age assessment result; place the coordinates composed of the skeletal physiological age and the actual age in the standard skeletal development curve for comparison to predict the development trend of the skeleton. See Figure 6 .

[0116] A cascade multi-factor bone age assessment device based on a spatio-temporal conversion network, including a data processing module, a bone landmark detection module, a fine-grained multi-factor bone age scoring module, and a standardized control prediction module;

[0117] The data processing module performs standardization, denoising, and normalization processing on the input image to obtain appropriate network input and a unified analysis standard;

[0118] The bone landmark detection module extracts medically valuable bone landmarks from the processed image using the spatio-temporal conversion network;

[0119] The fine-grained multi-factor bone age scoring module scores multiple skeletal marker regions. This module designs a multi-factor joint algorithm to comprehensively score by combining multiple factors affecting skeletal development.

[0120] The standardized control prediction module maps the bone age score to a standardized evaluation index for comparison to evaluate the patient's skeletal development and predict the skeletal development trajectory.

[0121] The present invention relates to a cascaded multi-factor bone age assessment device based on a spatio-frequency conversion network, which can run in the form of a computer program on various computing devices. These devices include but are not limited to terminal devices (such as tablet computers, desktop computers, laptops, mobile phones, wearable devices or personal digital assistants) and servers, where the latter can be independent or a cluster composed of multiple servers.

[0122] The computing device consists of a network interface, a memory, and a processor connected by a system bus. The memory includes a memory and a non-volatile storage medium, and the latter is used to store an operating system and a computer program. The computer program contains program instructions, and when the processor executes these instructions, it can achieve multi-modal classification for missing modality data. The processor is responsible for providing control and computing capabilities to ensure the normal operation of the device. The memory provides a running environment for the computer program in the non-volatile storage medium, enabling classification operations when the processor executes instructions.

[0123] In addition, the network interface supports network communication functions such as task allocation. It should be noted that the processor can be a digital signal processor (DSP), a central processing unit (CPU), a general-purpose processor, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), or other programmable logic devices, discrete hardware components, etc. A general-purpose processor usually refers to a microprocessor or other conventional processors.

[0124] The present invention also provides a computer-readable and writable storage medium storing a computer program, the program instructions contained therein being executable by a processor to implement a multi-factor bone age assessment method based on a spatio-frequency conversion network. The storage medium can be a device memory hard disk or an external storage device, such as a smart memory card, a plug-in hard disk, a flash memory card, a secure digital card (SD), etc.

[0125] The described embodiments are the preferred embodiments of the present invention, but the present invention is not limited to the above embodiments. Without departing from the essence of the present invention, any obvious improvements, substitutions, or variations that those skilled in the art can make fall within the protection scope of the present invention.

Claims

1. A cascaded multi-factor bone age assessment method based on a space-frequency conversion network, characterized in that: Take and process bone X-ray images; Extracting a detection region pool of bone markers from the processed X-ray image, wherein the detection region pool of bone markers is composed of a plurality of predefined detection frames, specifically generating 9 predefined detection frames at each pixel position of the X-ray image; for each predefined detection frame, using a space-frequency conversion network A to screen and locate the bone markers therein; Input a certain bone landmark region into the space-frequency conversion network B to extract regional features At the same time, multiple factors affecting bone development are input into the factor feature extraction network to obtain factor features. Calculate the shared information between regional features and factor features The image weight and factor weight are calculated by using the shared information, the residual information of the regional features, and the residual information of the factor features; the multi-factor joint feature is calculated by element-by-element summation operation of the image weight, the regional features, the factor weight, and the factor features; The multi-factor joint features are input into the bone age scoring network for inference to obtain the scoring result of the nth bone landmark area, and the scoring results of all bone landmark areas are added together to obtain the final bone age score; The final bone age score is mapped to industry standard quantitative indicators to obtain the bone physiological age assessment result. The coordinates of the bone physiological age and actual age are placed on the standard bone development curve chart for comparison to predict the bone development trend.

2. The cascaded multi-factor bone age assessment method based on space-frequency conversion network according to claim 1 is characterized in that: The space-frequency conversion network A and the space-frequency conversion network B have the same structure, both consisting of an input layer, five space-frequency conversion layers, two maximum pooling layers and an output layer; after the input layer receives the input image, the high-dimensional image features are extracted through the five space-frequency conversion layers; The third space-frequency conversion layer and the fourth space-frequency conversion layer are respectively connected to a maximum pooling layer for downsampling operation; the output layer outputs the result.

3. The cascaded multi-factor bone age assessment method based on space-frequency conversion network according to claim 2 is characterized in that: In the space-frequency conversion layer, for the input image I(x, y) and the conversion kernel K(x, y), the spatial domain information is first mapped to the frequency domain through Fourier transform: in, is the image in the frequency domain, is the conversion kernel in the frequency domain, function represents Fourier transform; Then the frequency domain feature extraction operation is performed, and the input image and the conversion kernel are multiplied pixel by pixel in the frequency domain: in, is the feature of the input image in the frequency domain; Then the features are converted to the spatial domain through inverse Fourier transform: Among them, C(x,y) is the spatial domain image feature after space-frequency conversion, and the function Represents an inverse Fourier transform operation.

4. The cascaded multi-factor bone age assessment method based on space-frequency conversion network according to claim 3 is characterized in that: The method of using the space-frequency conversion network A to screen and locate the bone markers is as follows: the space-frequency conversion network A classifies each predefined detection frame according to the marker type, and the sum of all category probabilities is 1. If the maximum category probability is higher than a threshold, the detection frame is retained, and the offset error between the retained detection frame and the true boundary frame is calculated to accurately locate the bone markers.

5. The cascaded multi-factor bone age assessment method based on space-frequency conversion network according to claim 3 is characterized in that: The shared information between the regional features and the factor features for: in, is the information entropy of the image features, is the information entropy of factor features, is the joint information entropy of image features and factor features, is the probability distribution of image features, is the probability distribution of factor characteristics, is the joint distribution of image features and factor features.

6. The cascaded multi-factor bone age assessment method based on space-frequency conversion network according to claim 5 is characterized in that: Residual information of regional characteristics Residual information of factor features They are:

7. The cascaded multi-factor bone age assessment method based on space-frequency conversion network according to claim 6 is characterized in that: Image weight Factor Weights 8. A device for implementing the cascade multi-factor bone age assessment method based on space-frequency conversion network as described in any one of claims 1 to 7, characterized in that: include: The data processing module performs standardization, denoising and normalization on bone X-ray images; The bone landmark detection module detects bone landmarks from the processed X-ray images using the space-frequency conversion network A; The fine-grained multi-factor bone age scoring module designs a multi-factor joint algorithm that combines multiple factors affecting bone development to conduct a comprehensive scoring of multiple bone marker areas. The standardized control prediction module maps the bone age score to standardized evaluation indicators to obtain the bone physiological age, and predicts the bone development trajectory based on the bone physiological age and actual age.

9. An electronic device, characterized in that: including memory and processor; The memory is used to store computer programs; The processor is used to execute the computer program and implement the cascaded multi-factor bone age assessment method based on space-frequency conversion network as described in any one of claims 1-7 when executing the computer program.

10. A storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by the processor, the processor executes the cascaded multi-factor bone age assessment method based on the space-frequency conversion network as described in any one of claims 1-7.