Breast tumor benignity and malignancy discrimination method based on multi-modal images

CN122531701APending Publication Date: 2026-08-07THE FIRST AFFILIATED HOSPITAL OF ARMY MEDICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
THE FIRST AFFILIATED HOSPITAL OF ARMY MEDICAL UNIV
Filing Date
2026-06-03
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0005]为了解决传统乳腺肿瘤判别方式依赖单一影像来源提取局部区域轮廓与基础像素分布特征,面对复杂病灶组织内部多样化结构时难以获取多维空间深层纹理映射关系,单一影像在成像过程易受干扰导致边界提取精度受限,仅凭基础形状参数与灰度统计数值构建分类映射难以捕捉深层结构频段能量分布差异,在面临不同病变表现时特征提取稳定性较差,导致病变性质分类边界划分不够精准,整体判别准确度与泛化适用性受极大制约的技术问题,本发明实施例提供了基于多模态影像的乳腺肿瘤良恶性判别方法

Benefits of technology

本发明中,基于多源数据时间对齐机制整合多维度特征梯度阵列以构建深层空间映射,通过频域变换解构空间纹理并量化不同频段能量分布占比以指导卷积提取过程,针对高低频属性差异自适应调整采样范围与步长分离出目标显著与背景特征,依据能量比重加权融合空间结构与像素灰度分布参数确立综合评估指标,克服同源数据信息匮乏与单一提取维度限制所致的边界模糊弊端,精准剥离病灶组织固有纹理特性与背景干扰因素,大幅增强恶性倾向特征识别灵敏度与全局判别精确性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122531701A_ABST
    Figure CN122531701A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of deep learning, in particular to a breast tumor benign and malignant discrimination method based on multi-modal images, comprising the following steps: correcting image acquisition timing to generate synchronous alignment data, calculating lesion area gradient to construct a basic feature mapping matrix, performing spatial frequency domain transformation to calculate energy proportion to generate a frequency energy distribution matrix, splitting high and low frequency regions and adjusting convolution kernel parameters to extract modal decoupling feature maps, and generating a breast tumor discrimination result based on decoupled feature weighted fusion. In the present application, deep mapping is constructed by integrating feature gradient based on multi-source data time alignment mechanism, different frequency band energy proportions are guided for convolution extraction by deconstructing texture and quantifying energy proportion through frequency domain transformation, target and background features are separated by adjusting sampling range and step length according to high and low frequency attribute differences, lesion tissue inherent texture and background interference factors are accurately stripped, and sensitivity of malignant tendency recognition and global discrimination accuracy are comprehensively enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of deep learning technology, and in particular to a method for distinguishing between benign and malignant breast tumors based on multimodal imaging. Background Technology

[0002] Deep learning technology refers to a class of computational methods that are based on multi-layer neural network structures. By constructing a network model containing an input layer, multiple hidden layers, and an output layer, and using large-scale data for iterative parameter training, it achieves automatic feature extraction and pattern mapping. Its core aspects include convolution operations for local feature extraction, activation functions for introducing non-linear expressions, loss functions for defining error metrics, and backpropagation based on gradient descent for weight updates. This technology field is widely involved in image analysis, speech recognition, and medical image processing. In particular, in medical image processing, pixel-level or region-level learning is used to complete classification and discrimination tasks from image data.

[0003] Traditional methods for distinguishing between benign and malignant breast tumors refer to a type of processing approach that differentiates the nature of tumors in breast imaging data. This approach involves processing a single modality of image, such as ultrasound or X-ray images. First, the boundaries of the tumor region are obtained manually or through threshold-based segmentation. Then, features are extracted based on the shape and contour parameters of the tumor region, such as perimeter-to-area ratio and edge smoothness, as well as gray-level distribution characteristics, such as average gray value and gray-level variance. These features are then input into a classifier such as a support vector machine or k-nearest neighbor classifier. By constructing a mapping relationship between the sample feature vectors and known class labels, benign or malignant tumor classification is completed. In addition, some methods also perform preprocessing steps such as image filtering and denoising, contrast enhancement, and gray-level normalization to improve the stability of feature extraction.

[0004] Traditional methods for identifying breast tumors rely on extracting local contours and basic pixel distribution features from a single image source. When faced with complex lesions and diverse internal structures, it is difficult to obtain multi-dimensional spatial deep texture mapping relationships. A single image is easily interfered with during the imaging process, which limits the accuracy of boundary extraction. It is difficult to capture the differences in frequency band energy distribution of deep structures by constructing classification mappings based solely on basic shape parameters and gray-level statistical values. When faced with different lesion manifestations, the stability of feature extraction is poor, resulting in insufficient precision in the delineation of lesion nature classification boundaries. The overall accuracy and generalization applicability are greatly limited. Summary of the Invention

[0005] To address the challenges of traditional breast tumor discrimination methods that rely on extracting local contours and basic pixel distribution features from a single image source, making it difficult to obtain multi-dimensional deep texture mapping relationships when faced with complex lesions and diverse internal structures, the limitations of single images in boundary extraction accuracy due to susceptibility to interference during imaging, the inability to capture differences in frequency band energy distribution in deep structures by constructing classification maps solely based on basic shape parameters and grayscale statistics, and the poor stability of feature extraction when facing different lesion manifestations, resulting in inaccurate lesion classification boundary delineation and severely restricting overall discrimination accuracy and generalization applicability, this invention provides a method for distinguishing benign and malignant breast tumors based on multimodal images.

[0006] To achieve the above objectives, this invention employs a method for distinguishing benign and malignant breast tumors based on multimodal imaging, comprising the following steps: S1: Acquire raw image data and time records, calculate the time difference between ultrasound images, X-ray images and magnetic resonance images, correct the image acquisition sequence, and generate synchronized image data. S2: Based on the synchronized image data, calculate the gray-level gradient of each modality of breast image and the texture direction gradient of the lesion area to construct a feature difference map, and input it into the structural feature layering algorithm to perform spatial domain feature extraction and construct a basic feature mapping matrix; S3: Perform a Fourier transform on the basic feature mapping matrix to convert the spatial texture features into frequency components, calculate the proportion of high-frequency and low-frequency energy in the lesion area, and generate a frequency domain energy distribution matrix. S4: Based on the frequency domain energy distribution matrix, the basic feature mapping matrix is ​​split into a high-frequency mapping region and a low-frequency mapping region. The convolution kernel radius is adjusted for the high-frequency mapping region, and the convolution kernel stride is adjusted for the low-frequency mapping region to extract the modal decoupling feature map. S5: Perform weighted fusion based on the modal decoupling feature map and the frequency domain energy distribution matrix, calculate the fused texture continuity parameter and gray-level balance parameter, compare them with the preset benign and malignant boundary threshold, and generate breast tumor discrimination result.

[0007] As a further aspect of the present invention, the synchronized image data includes a voxel matrix, spatial registration parameters, and an interpolation resampling grid; the basic feature mapping matrix includes a gradient magnitude map, local binary mode features, and an directional gradient histogram; the frequency domain energy distribution matrix includes an amplitude spectrum, a phase spectrum, and a power spectral density; the modal decoupling feature map includes a saliency map, an orthogonal feature matrix, and a spatial attention mask; and the breast tumor discrimination result includes benign / malignant classification labels, a malignancy probability prediction value, and an image report data system rating.

[0008] As a further aspect of the present invention, the specific steps of S1 are as follows: S101: Acquire raw image data produced by multimodal image acquisition equipment, retrieve corresponding metadata fields of raw image data, parse ultrasound occurrence time, extract X-ray occurrence time and magnetic resonance occurrence time, aggregate the three types of occurrence time, and establish a multimodal time cluster; S102: Select the magnetic resonance occurrence time within the multimodal time cluster as the reference zero point, perform difference calculations between the ultrasound occurrence time and the X-ray occurrence time and the reference zero point respectively to obtain the ultrasound time-biased scalar and the radiation time-biased scalar, map the two types of time-biased scalars to a one-dimensional array, and generate a multimodal time-biased matrix. S103: Call the multimodal time-biased matrix to extract the original image data belonging to the ultrasound image frame sequence and the X-ray image frame sequence. Perform frame sequence shifting according to the ultrasound time-biased scalar and the radiographic time-biased scalar. Associate the shifted frame sequence and the magnetic resonance image frame sequence with a fixed time axis to obtain synchronized image data.

[0009] As a further aspect of the present invention, the specific steps of S2 are as follows: S201: Based on the synchronized image data, extract the pixel array of each modality image, apply the edge detection operator to calculate the bidirectional gray-level derivative and synthesize the gray-level gradient magnitude, identify lesion areas that exceed the preset judgment threshold, use the co-occurrence matrix to obtain the texture direction gradient of the lesion area, and establish a modal multidimensional gradient set. S202: For the modal multidimensional gradient set, separate the gray-level gradient magnitude and texture direction gradient of different modes, perform subtraction on the gray-level gradient magnitude of adjacent modes according to the coordinate correspondence to obtain the gray-level difference, perform subtraction on the texture direction gradient at the same position to obtain the texture difference, and map the two types of differences to the grid nodes to obtain the feature difference map. S203: The feature difference map is called to divide the spatial domain into grid blocks. The grid blocks are scanned by the convolution operator to extract the pixel distribution frequency to form local spatial features. The feature values ​​are accumulated by the operator moving along the step size, and the feature values ​​are arranged in the original order to establish a basic feature mapping matrix.

[0010] As a further aspect of the present invention, the preset judgment threshold is determined by an adaptive value selection method based on the statistical distribution of bidirectional gray-level derivatives. Specifically, the bidirectional gray-level derivatives are statistically analyzed using a histogram in the lesion area, and the gray-level gradient amplitude corresponding to the cumulative distribution located in the preset high-level proportion interval is selected as the preset judgment threshold.

[0011] As a further aspect of the present invention, the specific steps of S3 are as follows: S301: Call the basic feature mapping matrix to extract the local spatial feature sequence, transform the local spatial feature sequence to the frequency domain through Fourier transform, analyze the signal intensity features and signal offset state to obtain complex frequency components, extract the corresponding amplitude spectrum parameters and phase spectrum parameters, and establish a set of spatial frequency components. S302: Extract frequency domain data from the spatial frequency component set, decouple high-frequency texture components and low-frequency background components according to a preset cutoff frequency threshold, analyze the corresponding amplitude spectrum parameters, quantify the high-frequency energy attributes of the lesion edge and the low-frequency energy attributes of the main structure, and analyze the distribution ratio of the two to obtain the frequency band energy ratio. S303: Call the frequency band energy ratio item to retrieve the original row and column index identifier, map the frequency band energy ratio item to the grid node according to the original row and column index identifier and perform numerical filling, perform pure zero filling operation for the grid coordinates of non-lesion areas to complete the background features, and generate a frequency domain energy distribution matrix.

[0012] As a further aspect of the present invention, the preset cutoff frequency threshold is adaptively determined based on the amplitude spectrum parameter distribution characteristics of the spatial frequency component set. After the amplitude spectrum parameters are sorted from low to high frequency, the corresponding frequency at which the cumulative amplitude spectrum energy reaches a predetermined proportion of the total amplitude spectrum energy is selected as the preset cutoff frequency threshold.

[0013] As a further aspect of the present invention, the specific steps of S4 are as follows: S401: Extract the node values ​​of the frequency domain energy distribution matrix and perform a difference operation with a preset baseline. Construct a high-frequency distribution mask and a low-frequency distribution mask based on the positive and negative attributes of the difference. Apply the high-frequency mask and the low-frequency mask to perform dot product separation on the basic feature mapping matrix to establish a multi-band mapping combination. S402: Extract the energy density parameters of the high-frequency mapping region in the multi-band mapping combination, convert the energy density parameters into scaling weights through preset dimensionless coefficients and multiply them with the initial radius to obtain the target receptive field boundary, apply the convolution kernel corresponding to the target receptive field boundary along the high-frequency mapping region to perform dot product traversal, extract dense texture vectors, and obtain a variable radius texture feature matrix. S403: Extract the energy sparse parameters of the low-frequency mapping region based on the multi-band mapping combination, extract the dimensionless step size adjustment coefficient based on the energy sparse parameters and multiply it with the initial step size to obtain the target sliding step size, apply the convolution kernel corresponding to the target sliding step size to perform jump dot product along the low-frequency mapping region, extract the sparse background vector, and perform channel concatenation between the variable radius texture feature matrix and the sparse background vector to generate a modal decoupling feature map.

[0014] As a further aspect of the present invention, the specific steps of S5 are as follows: S501: The modal decoupling feature map and frequency domain energy distribution matrix are called to extract local feature pixels and frequency domain distribution values. The node fusion weights are set according to the frequency domain distribution values. The local feature pixels are multiplied by the node fusion weights to perform numerical mapping. The matrix channel accumulation is performed along the spatial dimension to establish a cross-modal feature matrix. S502: Extract the internal texture pixel array based on the cross-modal feature matrix, determine the texture continuity parameter by calculating the gradient difference between adjacent grid nodes, determine the gray balance parameter based on the global pixel gray level distribution variance, and concatenate the texture continuity parameter and the gray level balance parameter to obtain the morphological evaluation parameter set. S503: For the set of morphological evaluation parameters, the texture continuity parameter and the grayscale balance parameter are multiplied by their corresponding weight coefficients and summed to obtain the evaluation morphological index. This index is then compared with a preset benign / malignant boundary threshold. If the index is greater than the benign / malignant boundary threshold, a malignant warning label is assigned; otherwise, a benign / normal label is assigned, thus generating a breast tumor discrimination result.

[0015] As a further aspect of the present invention, the preset benign-malignant boundary threshold is limited to a fixed value within the quantile range obtained based on historical sample statistics, and the quantile range is determined by the range between the median and the upper quartile of the evaluation morphological indicators corresponding to all training samples.

[0016] Compared with the prior art, the advantages and positive effects of the present invention are as follows: In this invention, a multi-dimensional feature gradient array is integrated based on a multi-source data time alignment mechanism to construct a deep spatial mapping. Spatial texture is deconstructed through frequency domain transformation and the energy distribution ratio of different frequency bands is quantified to guide the convolution extraction process. The sampling range and stride are adaptively adjusted to separate the salient features of the target from the background based on the differences in high and low frequency attributes. A comprehensive evaluation index is established by fusing spatial structure and pixel grayscale distribution parameters based on energy weight. This overcomes the drawbacks of boundary blurring caused by the lack of information from the same source data and the limitation of a single extraction dimension. It accurately removes the inherent texture characteristics of lesion tissue and background interference factors, and greatly enhances the sensitivity of malignant tendency feature recognition and the accuracy of global discrimination. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the accompanying drawings without creative effort.

[0018] Figure 1 This is a schematic diagram of the steps of the present invention; Figure 2 This is a detailed schematic diagram of S1 of the present invention; Figure 3 This is a detailed schematic diagram of S2 of the present invention; Figure 4 This is a detailed schematic diagram of S3 of the present invention; Figure 5 This is a detailed schematic diagram of S4 of the present invention; Figure 6 This is a detailed schematic diagram of S5 of the present invention. Detailed Implementation

[0019] The technical solution of the present invention will now be described with reference to the accompanying drawings.

[0020] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.

[0021] Please see Figure 1 This invention provides a method for distinguishing between benign and malignant breast tumors based on multimodal imaging, comprising the following steps: S1: Acquire raw image data and time records, calculate the time difference between ultrasound images, X-ray images and magnetic resonance images, correct the image acquisition sequence, and generate synchronized image data. S2: Based on synchronized image data, calculate the gray-level gradient of each modality of breast image and the texture direction gradient of the lesion area to construct a feature difference map, and input it into the structural feature layering algorithm to perform spatial domain feature extraction and construct a basic feature mapping matrix; S3: Perform a Fourier transform on the basic feature mapping matrix to convert spatial texture features into frequency components, calculate the proportion of high-frequency and low-frequency energy in the lesion area, and generate a frequency domain energy distribution matrix. S4: Based on the frequency domain energy distribution matrix, the basic feature mapping matrix is ​​split into high-frequency mapping region and low-frequency mapping region. The convolution kernel radius is adjusted for the high-frequency mapping region, and the convolution kernel stride is adjusted for the low-frequency mapping region to extract the modality decoupling feature map. S5: Perform weighted fusion based on the modal decoupling feature map and the frequency domain energy distribution matrix, calculate the fused texture continuity parameter and gray-level balance parameter, compare them with the preset benign and malignant boundary threshold, and generate breast tumor discrimination results.

[0022] Synchronous alignment of image data includes voxel matrix, spatial registration parameters, and interpolation resampling grid; basic feature mapping matrix includes gradient magnitude map, local binary mode features, and directional gradient histogram; frequency domain energy distribution matrix includes amplitude spectrum, phase spectrum, and power spectral density; modal decoupling feature map includes saliency map, orthogonal feature matrix, and spatial attention mask; breast tumor discrimination results include benign / malignant classification labels, malignancy probability prediction value, and image report data system rating.

[0023] Please see Figure 2 The specific steps of S1 are as follows: S101: Acquire raw image data produced by multimodal image acquisition equipment, retrieve corresponding metadata fields of raw image data, parse ultrasound occurrence time, extract X-ray occurrence time and magnetic resonance occurrence time, aggregate the three types of occurrence time, and establish a multimodal time cluster; The system retrieves a collection of raw examination files from a medical imaging database, containing multimodal physical attributes such as ultrasound scan data, X-ray scan data, and MRI scan data. A dirty data cleaning operation is performed on the retrieved raw examination files, examining the header structure of each data packet within each file, filtering out invalid files with missing header identifiers or containing garbled characters, and retaining the original image data with intact formatting. The system retrieves the corresponding data attribute record dictionary within the cleaned raw image data, locating the metadata field recording the specific scanning action time within each record. The system analyzes the timing of the ultrasound probe contacting the patient's body surface, extracts the timing of the X-ray generator releasing the detection dose, and simultaneously extracts the timing of the MRI radio frequency coil emitting electromagnetic pulses. The string content in hour, minute, second, and millisecond formats contained in these three types of timings is uniformly converted into a total value of consecutive milliseconds calculated based on Coordinated Universal Time (UTC). For example, the parsed first ultrasound time record is 10:10:20:50 AM, which, after conversion, yields the first ultrasound millisecond count as 1680343820500. The extracted X-ray time record is 10:10:21:000 AM, which, after conversion, yields the first X-ray millisecond count as 1680343821000. The extracted MRI time record is 10:10:19:000 AM, which, after conversion, yields the first MRI millisecond count as 1680343819000. Based on the identification number of the same examination subject, the first ultrasound millisecond count, first X-ray millisecond count, and first MRI millisecond count belonging to the same subject are grouped and stored in a shared sequence set, establishing a multimodal time cluster. Table 1 shows the original data table of the multimodal time cluster.

[0024] Table 1: Original Data Table of Multimodal Time Clusters 2001 1680343820500 1680343821000 1680343819000 2002 1680343850100 1680343850500 1680343849100 Table 1 lists the aggregated millisecond-level timestamp information of multiple subjects. The advantage of this operation logic is that by uniformly converting heterogeneous raw time strings into continuous absolute millisecond values ​​and merging them for storage, it eliminates the misalignment interference caused by differences in the underlying clock recording formats of multiple devices.

[0025] S102: Select the time of magnetic resonance occurrence within the multimodal time cluster as the reference zero point, perform difference calculations between the ultrasound occurrence time and the X-ray occurrence time and the reference zero point respectively to obtain the ultrasound time-biased scalar and the radiation time-biased scalar, map the two types of time-biased scalars to a one-dimensional array, and generate a multimodal time-biased matrix. Data records within the multimodal time cluster are extracted, and the value corresponding to the MRI occurrence milliseconds in the record entry is selected as the global alignment reference zero point. The corresponding ultrasound occurrence milliseconds within the same record are retrieved, and the reference zero point value is subtracted from this ultrasound occurrence milliseconds to obtain the ultrasound scan delay relative to the MRI scan, which is used as the ultrasound time offset scale. Similarly, the X-ray occurrence milliseconds within the same record are retrieved, and the reference zero point value is subtracted from this X-ray occurrence milliseconds to obtain the X-ray scan delay relative to the MRI scan, which is used as the radiographic time offset scale. A single-row, multi-column data container structure is created, and the previously calculated ultrasound and radiographic time offset scales are sequentially filled into the corresponding column positions of this container, resulting in a one-dimensional array for a single subject. Following the arrangement order of different subjects within the multimodal time cluster, the multiple one-dimensional arrays generated for each subject are stacked and concatenated one by one according to the row and column order of a matrix, ultimately generating a matrix containing all time offset data. In the actual numerical calculation, the data corresponding to the first subject number 2001 in Table 1 was retrieved. The MRI occurrence time in milliseconds (1680343819000) was selected as the reference zero point. The ultrasound occurrence time in milliseconds (1680343820500) was extracted. Subtracting 1680343819000 from 1680343820500 yielded a difference of 1500, which was used to obtain the ultrasound time-scaling value of 1500 milliseconds. Similarly, the X-ray occurrence time in milliseconds (1680343821000) was extracted. Subtracting 1680343819000 from 1680343821000 yielded a difference of 2000, which was used to obtain the radiation time-scaling value of 2000 milliseconds. The values ​​1500 and 2000 were then filled into a single-row structure, resulting in a one-dimensional array of values ​​consisting of 1500 and 2000. Following this rule, data number 2002 is processed. The MRI value 1680343849100 is used as the baseline. The ultrasound value 1680343850100 is subtracted from the baseline to obtain the ultrasound time-partial scalar value of 1000. The X-ray value 1680343850500 is subtracted from the baseline to obtain the radiation time-partial scalar value of 1400. This results in a second one-dimensional array of values, consisting of 1000 and 1400. These two arrays are then concatenated to generate a multimodal time-partial matrix.

[0026] S103: Call the multimodal time-biased matrix, extract the original image data belonging to the ultrasound image frame sequence and X-ray image frame sequence, perform frame sequence shifting according to the ultrasound time-biased scalar and the radiographic time-biased scalar, associate the shifted frame sequence and the magnetic resonance image frame sequence to a fixed time axis, and obtain synchronously aligned image data. The offset values ​​of each row in the multimodal time-biased matrix are read. Based on the subject number recorded in the matrix, the ultrasound image frame sequence set and the X-ray image frame sequence set belonging to the subject are extracted from the cleaned raw image data set. Based on the ultrasound time-biased scalar recorded in the multimodal time-biased matrix, combined with the preset frame rate per second parameter of the ultrasound equipment, the ultrasound time-biased scalar is multiplied by the frame rate per second and divided by 1000 to obtain the number of ultrasound image frames that need to be shifted. The ultrasound image frame sequence is then shifted backward in the time dimension according to this number of shifted frames. Similarly, the radiographic time-biased scalar is multiplied by the preset frame rate per second of the X-ray equipment and divided by 1000 to obtain the number of X-ray image frames that need to be shifted. The X-ray image frame sequence is then shifted accordingly according to this number. A fixed time axis with a reference zero point as the starting scale is established. The ultrasound image frame sequence, X-ray image frame sequence after the aforementioned sliding shift processing, and the original MRI image frame sequence without shifting are sequentially mounted and pasted onto the corresponding scale nodes of this fixed time axis according to the same frame correspondence position, resulting in synchronized image data with strictly corresponding time scales. In the specific numerical calculation, assuming the ultrasound time-scale deviation for subject 2001 is 1500 milliseconds, and the preset ultrasound acquisition frame rate is 30 frames per second, multiplying 1500 by 30 yields 45000, and dividing 45000 by 1000 gives the number of shifted frames as 45 frames. The entire ultrasound image frame sequence of this subject is then shifted backward by a distance of 45 frames. Assuming the X-ray time-scale deviation is 2000 milliseconds, and the X-ray acquisition frame rate is set to 15 frames per second, multiplying 2000 by 15 yields 30000, and dividing 30000 by 1000 gives the number of shifted frames as 30 frames. The entire sequence of X-ray images of the subject was shifted backward by 30 frames. Subsequently, the ultrasound sequence shifted by 45 frames, the X-ray sequence shifted by 30 frames, and the MRI sequence were correlated to a fixed time axis starting from 0.

[0027] Please see Figure 3 The specific steps of S2 are as follows: S201: Based on synchronized image data, extract the pixel array of each modality image, apply the edge detection operator to calculate the bidirectional gray-level derivative and synthesize the gray-level gradient magnitude, identify lesion areas exceeding the preset judgment threshold, use the co-occurrence matrix to obtain the texture direction gradient of the lesion area, and establish a modal multidimensional gradient set. The independent modal data packets recorded within the synchronized image data are read. The extracted magnetic resonance imaging (MRI) frame sequence and ultrasound image frame sequence are decomposed into independent-dimensional two-dimensional matrix structures, and the full-image pixel array corresponding to each independent modality is extracted. An edge detection operator containing horizontal and vertical detection masks is invoked, and the detection mask is slid across the MRI full-image pixel array pixel by pixel. For the center pixel at the 10th row and 20th column, the weighted difference of the brightness grayscale values ​​of the center pixel and its left and right adjacent pixels is used to obtain the horizontal grayscale derivative value of 96, and the weighted difference of the brightness grayscale values ​​of the vertically adjacent pixels is used to obtain the vertical grayscale derivative value of 72. The horizontal gray-level derivative 96 is multiplied by itself to obtain the horizontal square value 9216, and the vertical gray-level derivative 72 is multiplied by itself to obtain the vertical square value 5184. Adding 9216 and 5184 gives 14400. Taking the square root of 14400 yields a magnetic resonance gray-level gradient amplitude of 120 for the 10th row and 20th column. Similarly, for the pixel at this coordinate in the full-image ultrasound pixel array, the horizontal gray-level derivative 72 and the vertical gray-level derivative 54 are calculated, squared, added together to obtain 8100, and the square root gives an ultrasound gray-level gradient amplitude of 90. The gray-level gradient amplitudes of 1000 pixels in the full-image ultrasound are statistically analyzed to construct a distribution histogram. When the cumulative amplitude reaches 80, there are 860 pixels. Dividing 860 by 1000 gives a cumulative distribution ratio of 0.86, which falls within the range of 0.85 to 0.95. 80 is set as the preset threshold. Since the amplitude of 120 is greater than 80, this point is retained as the lesion region. Each pixel within the lesion region is extracted, and the joint probability distribution of adjacent pixel pairs is statistically analyzed to form a co-occurrence matrix. Based on this matrix, the magnetic resonance texture direction gradient at this point is calculated to be 45. Similarly, the ultrasound texture direction gradient is calculated to be 20. A modal multidimensional gradient set is established using the amplitudes of 120 and 90, and the gradients of 45 and 20.

[0028] S202: For the modal multidimensional gradient set, separate the gray-level gradient magnitude and texture direction gradient of different modes. According to the coordinate correspondence, perform subtraction on the gray-level gradient magnitude of adjacent modes to obtain the gray-level difference, and perform subtraction on the texture direction gradient at the same position to obtain the texture difference. Map the two types of differences to the grid nodes to obtain the feature difference map. The data entries within the established modal multidimensional gradient set are read. Based on the type label of the scanning instrument, the mixed data within the set are split into independent modal grayscale gradient amplitude arrays and corresponding independent modal texture orientation gradient arrays. The magnetic resonance (MRI) grayscale gradient amplitude and ultrasound grayscale gradient amplitude are extracted for the same subject at the same spatial coordinate point. The MRI grayscale gradient amplitude is subtracted from the ultrasound grayscale gradient amplitude to obtain the pure numerical difference between the two, thus obtaining the grayscale difference at that coordinate point. Similarly, the MRI texture orientation gradient and ultrasound texture orientation gradient are extracted at the same spatial coordinate point. The MRI texture orientation gradient is subtracted from the ultrasound texture orientation gradient to obtain the texture difference at that spatial location. A two-dimensional blank grid structure with dimensions of 512 units in both length and width is constructed. Based on the original spatial horizontal and vertical coordinate parameters, the calculated grayscale difference data and texture difference data are sequentially filled into the matching grid nodes within this two-dimensional blank grid structure. After filling, a feature difference map full of numerical differences is generated. In the specific calculation, for the pixel node with a spatial coordinate of 10th row and 20th column, the split dataset is called to extract the magnetic resonance grayscale gradient amplitude at this node as 120, and the corresponding ultrasonic grayscale gradient amplitude as 90. A subtraction operation is performed to subtract 90 from 120, resulting in a difference of 30, i.e., the grayscale difference is 30. Simultaneously, the magnetic resonance texture direction gradient at this node is extracted as 45, and the corresponding ultrasonic texture direction gradient is 20. A subtraction operation is performed to subtract 20 from 45, resulting in a difference of 25, i.e., the texture difference is 25. Then, this pair of values, grayscale difference 30 and texture difference 25, are directly filled into the corresponding node in a 512-unit 2D grid, based on the original coordinate index of the 10th row and 20th column. This subtraction and mapping operation is repeated until all coordinate positions are processed.

[0029] S203: Call the feature difference map to divide the spatial domain grid blocks, use the convolution operator to scan the grid blocks, extract the pixel distribution frequency to form local spatial features, move the operator along the step size to accumulate feature values, arrange the feature values ​​in the original order, and establish the basic feature mapping matrix. The generated feature difference map, after mapping, is called and its total dimensions are divided into multiple square spatial domain grid blocks, each with a side length of 16 pixels, according to a fixed stride size. A convolutional scanning network model is introduced, consisting of one input layer, one spatial hidden layer, and one feature output layer. The input layer receives the difference data within each square spatial domain grid block. The spatial hidden layer uses a 3x3 convolution operator mask and performs a sliding scan within the grid block in a left-to-right, top-to-bottom order. At each stop, the absolute frequency of each gray-level difference value falling within the 3x3 range is counted, and the pixel distribution frequency is directly used as the local spatial feature. Data transmission between layers within the network is performed using local connection logic. After the spatial hidden layer outputs the scan results, a modified linear unit activation function is applied to non-linearly filter the frequency values, forcibly replacing all values ​​less than 0 with 0. As the convolution operator mask continues to slide at a set stride distance, the extracted and filtered feature values ​​are continuously accumulated. The output layer rearranges the accumulated feature values ​​according to the original row and column distribution order of the spatial domain grid blocks in the entire image, ultimately establishing the basic feature mapping matrix. For example, a 16x16 square grid block located in the upper left corner is extracted and fed into the input layer of the convolutional network. The 3x3 convolution operator in the spatial hidden layer covers the first 9 pixels in the upper left corner, counting the occurrences of pixels with a gray-level difference of 30 as 4 times, and outputting this as the first local spatial feature. The operator slides to the right by a distance of 2 pixels, counting the occurrences of pixels with a gray-level difference of 25 as 3 times in the newly covered area. The value remains unchanged after applying the activation function. The occurrence count of 4 in the first step is added to the occurrence count of 3 in the second step to obtain the accumulated feature value of 7. After repeating the traversal of the entire grid block, the final accumulated value is obtained, for example, 56. According to the original order of the grid block in the first row and first column of the entire image, 56 is filled into the starting position of the output layer result array.

[0030] Please see Figure 4 The specific steps of S3 are as follows: S301: Call the basic feature mapping matrix, extract the local spatial feature sequence, transform the local spatial feature sequence to the frequency domain through Fourier transform, analyze the signal intensity features and signal offset state to obtain complex frequency components, extract the corresponding amplitude spectrum parameters and phase spectrum parameters, and establish a set of spatial frequency components. Read all the neatly arranged feature values ​​within the basic feature mapping matrix, and peel off the values ​​row by row according to the horizontal row index of the matrix, reorganizing them into a one-dimensional continuous local spatial feature sequence. Perform a Fast Fourier Transform (FFT) algorithm on the extracted local spatial feature sequence to transform the feature values ​​dominated by spatial location coordinates into a frequency domain coordinate system dominated by frequency changes. Analyze the transformed frequency domain sequence, extracting the corresponding complex representation for each frequency node, and decompose it into a real part representing the signal intensity feature and an imaginary part representing the signal offset state. Combine these two parts to obtain the complex frequency components. Extract the real part value and multiply it by itself to obtain the square root; extract the imaginary part value and multiply it by itself to obtain the square root; add these two squared results and perform a square root operation to obtain the amplitude spectrum parameter of that frequency node; simultaneously, extract the imaginary part value and divide it by the real part value to obtain the ratio, and then perform an arctangent angle calculation on this ratio to obtain the phase spectrum parameter of that frequency node. Classify and summarize the amplitude spectrum parameters and phase spectrum parameters corresponding to all frequency nodes to establish a spatial frequency component set. In a specific computational exercise, assuming that the feature sequence extracted from the matrix, after transformation, yields a real part of 8 representing signal strength and an imaginary part of 6 representing signal offset at a node with a frequency index of 10 Hz. These two parts together constitute the complex frequency component of that node. Multiplying the real part 8 by itself yields a squared result of 64, and multiplying the imaginary part 6 by itself yields a squared result of 36. Adding 64 and 36 gives a sum of 100. Taking the square root of 100 yields an amplitude spectrum parameter of 10. Next, dividing the imaginary part 6 by the real part 8 yields a ratio of 0.75. Calculating the arctangent angle of 0.75 yields a phase spectrum parameter of 36.8 degrees. The same computational process is applied to the node with a frequency index of 20 Hz, yielding a real part of 12 and an imaginary part of 5. Squaring and adding these squares (144 + 25) gives 169. Taking the square root yields an amplitude spectrum parameter of 13, and the arctangent angle yields a phase spectrum parameter of 22.6 degrees. A set is created by summarizing the amplitude and phase data.

[0031] S302: Extract frequency domain data from the set of spatial frequency components, decouple high-frequency texture components and low-frequency background components according to the preset cutoff frequency threshold, analyze the corresponding amplitude spectrum parameters, quantify the high-frequency energy attributes of the lesion edge and the low-frequency energy attributes of the main structure, and analyze the distribution ratio of the two to obtain the frequency band energy ratio. Read all data nodes within the established spatial frequency component set and extract the amplitude spectrum parameters contained in each node individually. For these amplitude spectrum parameters, sort them sequentially from low to high frequency according to their corresponding frequency values. Add all sorted amplitude spectrum parameters to calculate the total amplitude spectrum energy. Then, sequentially extract the corresponding amplitude spectrum parameters in ascending frequency order and accumulate them one by one. For each added amplitude spectrum parameter, divide the accumulated amplitude spectrum energy by the previously calculated total amplitude spectrum energy to obtain a dynamically updated energy distribution ratio. When this distribution ratio first reaches or exceeds a preset ratio threshold (e.g., 0.80), accumulation immediately stops, and the frequency value of the last added amplitude spectrum parameter is directly determined as the preset cutoff frequency threshold. All frequencies greater than the preset cutoff frequency threshold are classified as high-frequency texture components, and their edge high-frequency energy attributes are quantified; frequencies less than or equal to the preset cutoff frequency threshold are classified as low-frequency background components, and their main structure low-frequency energy attributes are quantified. The sum of the high-frequency energy attributes at the edge is divided by the sum of the low-frequency energy attributes of the main structure to obtain the frequency band energy ratio. In the actual numerical calculation, there are four frequency levels, ordered from low to high. The corresponding amplitude spectrum parameters are 40, 30, 20, and 10. Adding them all together (40 + 30 + 20 + 10) gives a total amplitude spectrum energy of 100. The addition proceeds sequentially: first, the lowest frequency amplitude spectrum parameter 40 is extracted, and 40 is divided by the total energy of 100 to obtain a ratio of 0.40; then, the second lowest frequency amplitude spectrum parameter 30 is added, resulting in a cumulative energy of 70, which is divided by 100 to obtain a ratio of 0.70; finally, the third-level frequency amplitude spectrum parameter 20 is added, resulting in a cumulative energy of 90, which is divided by 100 to obtain a ratio of 0.90. Since 0.90 exceeds the threshold value of 0.80 for the first time, the frequency value corresponding to the third-level amplitude spectrum parameter 20, for example, 60 Hz, is determined as the preset cutoff frequency threshold. The high-frequency energy greater than 60 Hz is 10, and the low-frequency energy less than or equal to 60 Hz is 90. Dividing the high-frequency energy 10 by the low-frequency energy 90 yields a frequency band energy ratio of 0.11.

[0032] Table 2: Frequency and Energy Distribution Table 1 20 40 0.40 2 40 30 0.70 3 60 20 0.90 Table 2 shows the energy ratio accumulation process for each frequency node. The advantage of this operational logic is that it uses an adaptive optimization method based on energy ratio to determine the cutoff frequency, dynamically adapting to the abundance of boundary spiculations in different breast lesion types.

[0033] S303: Call the frequency band energy ratio item to retrieve the original row and column index identifier, map the frequency band energy ratio item to the grid node according to the original row and column index identifier, perform numerical filling, perform pure zero filling operation for the grid coordinates of non-lesion areas to complete the background features, and generate a frequency domain energy distribution matrix; Extract the specific values ​​of the frequency band energy ratios obtained from the aforementioned operations. Retrieve the original spatial horizontal and vertical coordinate indices of the basic feature mapping matrix retained during the spatial domain transformation. Construct a new two-dimensional blank grid structure with the same number of rows and columns as the original basic feature mapping matrix. Based on the retrieved original horizontal and vertical coordinate indices, directly fill the obtained frequency band energy ratio values ​​one by one into the corresponding nodes in the two-dimensional blank grid structure, performing a value filling operation. Subsequently, retrieve the lesion area location marker records generated during the aforementioned edge detection stage and check and compare each coordinate node in the two-dimensional blank grid structure. For non-lesion area grid coordinate points that are not covered by the lesion area location markers, forcibly enter pure numbers 0 into their nodes for value coverage, performing a pure zero filling operation to erase stray frequency band features in irrelevant areas, thereby completing the overall clean background feature appearance. After the above complete value mapping and background zeroing operations, the entire two-dimensional grid structure is transformed into a frequency domain energy distribution matrix. For example, in the previous steps, the frequency band energy ratio of a specific spatial segment was found to be 0.11. A search revealed that the segment corresponding to this value was originally located in the 5th row and 8th column of the basic feature mapping matrix. At this point, the number 0.11 was directly written into the node cell corresponding to the 5th row and 8th column of the newly created equal-sized two-dimensional blank mesh structure to complete the numerical filling. Next, a mask comparison of the lesion area revealed that the 5th row and 8th column belonged to the lesion area, so the number 0.11 was retained; however, a search revealed that the node coordinates in the 1st row and 1st column were not within the lesion area location marker range, so the number 0 was immediately written directly into the node cell in the 1st row and 1st column to complete the pure zero-filling operation for the non-lesion area. This process was repeated to complete the full image scan.

[0034] Please see Figure 5 The specific steps of S4 are as follows: S401: Extract the node values ​​of the frequency domain energy distribution matrix and perform a difference operation with the preset baseline. Construct a high-frequency distribution mask and a low-frequency distribution mask based on the positive and negative attributes of the difference. Apply the high-frequency mask and the low-frequency mask to perform dot product separation on the basic feature mapping matrix to establish a multi-band mapping combination. Extract the numerical content of all corresponding nodes in the frequency domain energy distribution matrix, sort all node values ​​from largest to smallest, and find the median value in the exact middle position. Set this median value as the preset baseline. For each node in the frequency domain energy distribution matrix, subtract the preset baseline from the extracted value, performing a difference operation. Check the sign attribute of each difference operation result. If the difference result is positive or equal to zero, generate the number 1 at that node position as a marker, and combine all nodes marked with 1 to form a high-frequency distribution mask; if the difference result is negative, generate the number 1 at that node position as a marker, and set the other positions to 0, combining them to form a low-frequency distribution mask. Retrieve the previously generated basic feature mapping matrix, multiply the matrix region containing the number 1 in the high-frequency distribution mask with the corresponding node values ​​in the basic feature mapping matrix, separating the high-frequency feature layer; similarly, multiply the low-frequency distribution mask with the corresponding node values ​​in the basic feature mapping matrix, separating the low-frequency feature layer. The high-frequency and low-frequency layers are merged to create a multi-band mapping combination. In the specific numerical extrapolation, assume that the extracted values ​​of the five nodes in the matrix are 0.11, 0.09, 0.05, 0.03, and 0.00. The value in the middle after sorting is 0.05, which is set as the preset baseline. Extract the first node value 0.11, subtract the baseline 0.05 from 0.11, and the difference is 0.06. Since 0.06 is a positive number, record the number 1 in the first position of the mask, belonging to the high-frequency distribution mask coverage area. Extract the node value 0.05, subtract the baseline 0.05 from it, and the difference is 0. Since 0 equals zero, record the number 1 in the corresponding position of the mask, belonging to the high-frequency distribution mask coverage area. Extract another node value 0.03, subtract the baseline 0.05 from 0.03, and the difference is negative 0.02. Since 0.02 is a negative number, record the number 1 in the corresponding position of the low-frequency mask array. Subsequently, if the feature value at the first position in the basic feature mapping matrix is ​​56, multiply 56 by 1 in the high-frequency distribution mask to obtain 56, which is then separated into the high-frequency layer. If the feature value at another position is 40, multiply it by the corresponding position in the low-frequency distribution mask to separate the layers.

[0035] S402: Extract the energy density parameters of the high-frequency mapping region in the multi-band mapping combination, convert the energy density parameters into scaling weights through preset dimensionless coefficients and multiply them with the initial radius to obtain the target receptive field boundary, apply the corresponding convolution kernel of the target receptive field boundary along the high-frequency mapping region to perform dot product traversal, extract dense texture vectors, and obtain the variable radius texture feature matrix. Feature content of high-frequency mapping regions contained in multi-band mapping combinations is extracted. The number of all pixels falling within the high-frequency mapping region is counted, and the feature values ​​corresponding to each pixel are summed to obtain the total energy of the region features. The calculated total energy of the region features is divided by the total number of pixels, and the ratio is directly defined as the energy density parameter. An initial fixed scanning radius is set, and the energy density parameter is converted into scaling weights using a preset dimensionless coefficient. The scaling weights are multiplied by the initial scanning radius to calculate the product that needs to be expanded. The product is rounded to the nearest integer and used as the target receptive field boundary size. A square two-dimensional convolutional kernel array with the same size as the target receptive field boundary is constructed, and the two-dimensional convolutional kernel is continuously moved row by row and column by column along the spatial coordinate axis of the high-frequency mapping region. At each stop, the region value covered by the convolutional kernel is multiplied and added one-to-one with the weight value of the convolutional kernel itself, and the results are reassembled and stitched into a dense texture vector. All results are combined to output a variable radius texture feature matrix. During the calculation of specific parameters, it was found that the high-frequency mapping region contains a total of 400 feature pixels. The feature values ​​of these 400 pixels were summed one by one, resulting in a total of 600. Dividing the total energy of 600 by the number of pixels (400) yields a ratio of 1.5, indicating an energy density parameter of 1.5. A preset dimensionless coefficient (e.g., set to 1.0) is introduced to convert the energy-dimensional density parameter 1.5 into a dimensionless scaling weight of 1.5. The default initial scan radius is set to 2 pixels. Multiplying the obtained scaling weight 1.5 by the initial scan radius of 2 yields a product of 3. Therefore, the target receptive field boundary is established as 3 pixels in length and width. Subsequently, a 3x3 two-dimensional convolution kernel is generated and slid along the high-frequency mapping region pixel by pixel, performing dot product accumulation on the currently covered 9 pixels to extract a dense texture vector. The advantage of this operational logic is that it transforms the feature density of the lesion site into a dynamic product amplification coefficient after dimensionless processing, which adjusts the size of the receptive field, so that the capture window for the edge of the fine burr can expand and contract automatically according to the intensity of the lesion.

[0036] S403: Extract energy sparse parameters of the low-frequency mapping region based on the multi-band mapping combination, extract the dimensionless step size adjustment coefficient based on the energy sparse parameters and multiply it with the initial step size to obtain the target sliding step size, apply the convolution kernel corresponding to the target sliding step size to perform jump dot product along the low-frequency mapping region, extract the sparse background vector, and perform channel concatenation between the variable radius texture feature matrix and the sparse background vector to generate a modal decoupling feature map. Read the data structure of the low-frequency mapping region obtained from the aforementioned multi-band mapping combination decomposition, and count the total number of non-zero feature pixels scattered within this low-frequency mapping region. Plan the fixed total number of grid pixels occupied by the total region, divide the total number of non-zero feature pixels by the total number of grid pixels, and calculate its sparsity ratio. Use the reciprocal of this ratio or directly specify the derived value of the reciprocal as the energy sparsity parameter. Pre-set an initial step size value for the basic displacement, extract a dimensionless step size adjustment coefficient based on the obtained energy sparsity parameter, and multiply this dimensionless step size adjustment coefficient with the initial step size value to obtain the final target sliding step size used for the actual scanning span. Construct a low-frequency detection convolution kernel of the corresponding size, and along the spatial coordinate network of the low-frequency mapping region, force discontinuous jump movement and dot product feature extraction operations according to the target sliding step size interval obtained above, arranging the smooth feature data gathered at each jump dwell point to construct a sparse background vector. The variable radius texture feature matrix generated in the previous stage is called. The dimensions of all data arrays contained in this feature matrix are then combined with the channel dimensions of the currently extracted sparse background vector to perform a side-by-side concatenation operation, thus expanding the channel dimension of the feature map and generating a modally decoupled feature map. In the digital inference calculation, it is assumed that the number of non-zero feature pixels found in the low-frequency mapping region is 100, while the total number of grid pixels in this region is 1000. Dividing 100 by 1000 yields a sparsity ratio of 0.1. For ease of calculation, the energy sparsity parameter corresponding to this sparse feature is directly set to the number 1 (derived reciprocal value). Based on the energy sparsity parameter 1, a dimensionless step size adjustment coefficient of 1.5 is extracted through a preset mapping relationship. The preset initial step size value is 2. Multiplication logic is executed, multiplying the dimensionless step size adjustment coefficient 1.5 by the initial step size value 2 to obtain a final target sliding step size of 3 pixel intervals. After each dot product extraction operation, the convolutional kernel jumps directly to the right or down, spanning a distance of 3 pixels, to perform the next scan and obtain the sparse background vector. The original variable radius matrix has 64 channels. The 64 channels formed by the extracted sparse background vector are concatenated side by side to the back of the matrix, doubling the total number of channels to 128.

[0037] Please see Figure 6 The specific steps of S5 are as follows: S501: Call the modal decoupling feature map and frequency domain energy distribution matrix to extract local feature pixels and frequency domain distribution values, set node fusion weights based on frequency domain distribution values, multiply local feature pixels with node fusion weights to perform numerical mapping, perform matrix channel accumulation along spatial dimensions, and establish a cross-modal feature matrix; Extract the modal decoupling feature map generated by the above stages and the frequency domain energy distribution matrix established in the previous process. From each channel array of the modal decoupling feature map, read the local feature pixel values ​​contained at specific spatial coordinate points; simultaneously, extract the corresponding frequency domain distribution values ​​from the frequency domain energy distribution matrix based on the same spatial coordinate point index. Use the extracted frequency domain distribution value at that point directly as the node fusion weight to measure the importance of the feature at that point. Execute weighted multiplication logic, multiplying the previously extracted local feature pixel values ​​by the node fusion weight to calculate the weighted enhanced mapping value, and perform the numerical mapping operation. For each spatial coordinate point, after completing the weighted multiplication mapping, perform continuous accumulation operations along the depth dimension direction composed of 128 stitched channels, summing all weighted mapping values ​​of that point in each channel to compress the channel thickness and generate a comprehensive feature value. Combine and arrange the comprehensive feature values ​​generated for all spatial coordinate points in the entire image to construct a new two-dimensional cross-modal feature matrix. In the specific assignment and calculation process, for a fixed point at the 30th row and 40th column of the two-dimensional coordinate system, the local feature pixel value of 80 is read from the first channel of the modal decoupling feature map. Using the same coordinate position, the frequency domain distribution value of this point is extracted from the frequency domain energy distribution matrix, which is 0.8. This frequency domain distribution value of 0.8 is directly set as the node fusion weight. A multiplication operation is performed, multiplying the pixel value of 80 by the weight value of 0.8, to calculate the weighted mapping value of 64 for this channel. The feature value of this coordinate in the second channel, for example, 50, is read and multiplied by the weight of 0.8 to obtain a mapping value of 40. This operation is continued to process all 128 channels. Then, addition is performed along the depth channel dimension, adding the 64 of the first channel to the 40 of the second channel and the mapping values ​​of all other channels, assuming that the final summation yields a comprehensive feature value of 4500. 4500 is filled into the 30th row and 40th column position of the newly created two-dimensional matrix. The matrix is ​​constructed by traversing all coordinate nodes.

[0038] S502: Extract the internal texture pixel array based on the cross-modal feature matrix, determine the texture continuity parameter by calculating the gradient difference between adjacent grid nodes, determine the gray balance parameter based on the global pixel gray level distribution variance, and concatenate the texture continuity parameter and the gray level balance parameter to obtain the morphological evaluation parameter set. The cross-modal feature matrix, after dimensionality reduction and merging, is read. A pre-planned lesion boundary mask is used to cover the entire image, filtering out invalid peripheral data and specifically extracting the texture pixel array combination within the lesion coverage area. Within the extracted internal texture pixel array, following the network order from left to right, the feature values ​​of the current grid node and its next rightmost grid node are obtained one by one. The feature value of the current node is subtracted from the feature value of its rightmost neighbor, and the absolute value of the difference is taken as the gradient difference between adjacent grid nodes. All the obtained gradient differences between adjacent grid nodes within the array are summed, and then divided by the total number of logarithmic subtractions to obtain the average difference, which is defined as the texture continuity parameter. The feature values ​​of all pixel grids within the lesion coverage area are collected, summed, and divided by the total number of pixels to obtain the global average feature value. Then, the difference between each individual feature value and the global average feature value is calculated. This difference is multiplied by itself and squared. All squared values ​​are summed and divided by the total number of pixels to calculate the global pixel grayscale distribution variance. This variance is directly set as the grayscale balance parameter. A container is created to hold the feature description vector set of the data. The texture continuity parameter and grayscale balance parameter calculated above are stored side-by-side in the container in order and then concatenated to obtain the complete morphological evaluation parameter set. In the numerical calculation example, the array contains three consecutive node values: 80, 85, and 82. First, the absolute value of 80 minus 85 is taken as 5, and then the absolute value of 85 minus 82 is taken as 3. Adding 5 and 3 gives 8, which is divided by the comparison logarithm 2 to obtain the average value of 4, i.e., the texture continuity parameter is 4. Additionally, the feature values ​​of the region covering 5 pixels are extracted as 10, 12, 14, 16, and 18. These are summed to 70, which is divided by 5 to obtain the average value of 14. Square the differences one by one: 10 minus 14 equals -4 squared, which is 16; 12 minus 14 equals -2 squared, which is 4; 14 minus 14 squared is 0; 16 minus 14 squared is 4; 18 minus 14 squared is 16. Add 16, 4, 0, 4, and 16 to get 40. Divide 40 by 5 to get the variance value, which is 8. Set 8 directly as the grayscale balance parameter. Then, fill the feature vector queue with the texture continuity parameter 4 and the grayscale balance parameter 8 and concatenate them.

[0039] S503: For the morphological evaluation parameter set, the texture continuity parameter and grayscale balance parameter are multiplied by their corresponding weight coefficients and summed to obtain the evaluation morphological index. This index is then compared with the preset benign / malignant boundary threshold. If the index is greater than the benign / malignant boundary threshold, a malignant warning label is assigned; otherwise, a benign / normal label is assigned, generating a breast tumor discrimination result. Extract the values ​​of each parameter from the previously constructed morphological evaluation parameter set. Retrieve the pre-set first weighting coefficient for texture continuity and the second weighting coefficient for grayscale balance. Perform a mixed weighted calculation, directly multiplying the extracted texture continuity parameter value by the first weighting coefficient value to obtain the texture product; and directly multiplying the grayscale balance parameter value by the second weighting coefficient value to obtain the grayscale product. Summate the texture product and the grayscale product to obtain the final morphological evaluation index value used for overall judgment. Retrieve all morphological evaluation index values ​​generated in previous diagnoses from all historical training samples retained in the database, and sort the historical index values ​​from smallest to largest. Select the median parameter, which is located in the middle of the entire sorted queue, and select the upper quartile parameter, which is located at the 75th percentile of the sorted queue. Add the median parameter and the upper quartile parameter and calculate the arithmetic mean. Set this mean as a fixed value and use it as the preset benign / malignant boundary threshold. The values ​​of the evaluation morphological indicators obtained by summing the aforementioned values ​​are compared and verified with the preset benign / malignant threshold. If the evaluation morphological indicator value is greater than the preset benign / malignant threshold, a malignant warning label is directly assigned to it; if it is less than or equal to the preset benign / malignant threshold, a benign / normal label is assigned to it, thus generating the final breast tumor discrimination result. In the specific calculation process, the results calculated in the previous steps are used to extract the texture continuity parameter as 4 and the grayscale balance parameter as 8. The first weight coefficient is set to 0.6 and the second weight coefficient is set to 0.4. Multiplying 4 by 0.6 yields a texture product result of 2.4; multiplying 8 by 0.4 yields a grayscale product result of 3.2. Adding 2.4 and 3.2, the final evaluation morphological indicator value is 5.6. In the historical sample sorting, the median parameter of the historical indicator is found to be 4.0 and the upper quartile parameter is 6.0. Add 4.0 and 6.0 to get 10.0, divide by 2 to get the average value of 5.0, and set 5.0 as the preset threshold for distinguishing between benign and malignant. At this time, compare the value of the evaluation morphology index to be tested, 5.6, with the threshold of 5.0. Since 5.6 is greater than 5.0, the warning logic is triggered, and a judgment conclusion with a malignant warning label is output.

[0040] Table 3: Sample Morphological Indicators and Classification Table Test batch one 5.6 5.0 Malicious warning label Test Batch Two 3.8 5.0 benign routine label Table 3 lists the implementation records of binary classification of benign and malignant diseases using threshold comparison. The advantage of this operation logic is that it uses a combination of the median and upper quartile values ​​extracted from a large sample of historical data to construct a rigid dividing line, reducing misjudgments caused by local overfitting.

[0041] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of protection of the described technical solutions.

Claims

1. A method for discriminating between benign and malignant breast tumors based on multi-modal images, characterized by, Includes the following steps: S1: Acquire raw image data and time records, calculate the time difference between ultrasound images, X-ray images and magnetic resonance images, correct the image acquisition sequence, and generate synchronized image data. S2: Based on the synchronized image data, calculate the gray-level gradient of each modality of breast image and the texture direction gradient of the lesion area to construct a feature difference map, and input it into the structural feature layering algorithm to perform spatial domain feature extraction and construct a basic feature mapping matrix; S3: Perform a Fourier transform on the basic feature mapping matrix to convert the spatial texture features into frequency components, calculate the proportion of high-frequency and low-frequency energy in the lesion area, and generate a frequency domain energy distribution matrix. S4: Based on the frequency domain energy distribution matrix, the basic feature mapping matrix is ​​split into a high-frequency mapping region and a low-frequency mapping region. The convolution kernel radius is adjusted for the high-frequency mapping region, and the convolution kernel stride is adjusted for the low-frequency mapping region to extract the modal decoupling feature map. S5: Perform weighted fusion based on the modal decoupling feature map and the frequency domain energy distribution matrix, calculate the fused texture continuity parameter and gray-level balance parameter, compare them with the preset benign and malignant boundary threshold, and generate breast tumor discrimination result. 2.The method of claim 1, wherein, The synchronized image data includes a voxel matrix, spatial registration parameters, and interpolation resampling grid. The basic feature mapping matrix includes a gradient magnitude map, local binary mode features, and directional gradient histogram. The frequency domain energy distribution matrix includes amplitude spectrum, phase spectrum, and power spectral density. The modal decoupling feature map includes a saliency map, orthogonal feature matrix, and spatial attention mask. The breast tumor discrimination result includes benign / malignant classification labels, malignancy probability prediction value, and image report data system rating. 3.The method of claim 1, wherein, The specific steps of S1 are as follows: S101: Acquire raw image data produced by multimodal image acquisition equipment, retrieve corresponding metadata fields of raw image data, parse ultrasound occurrence time, extract X-ray occurrence time and magnetic resonance occurrence time, aggregate the three types of occurrence time, and establish a multimodal time cluster; S102: Select the magnetic resonance occurrence time within the multimodal time cluster as the reference zero point, perform difference calculations between the ultrasound occurrence time and the X-ray occurrence time and the reference zero point respectively to obtain the ultrasound time-biased scalar and the radiation time-biased scalar, map the two types of time-biased scalars to a one-dimensional array, and generate a multimodal time-biased matrix. S103: Call the multimodal time-biased matrix to extract the original image data belonging to the ultrasound image frame sequence and the X-ray image frame sequence. Perform frame sequence shifting according to the ultrasound time-biased scalar and the radiographic time-biased scalar. Associate the shifted frame sequence and the magnetic resonance image frame sequence with a fixed time axis to obtain synchronized image data.

4. The method for determining the benign or malignant nature of breast tumors based on multimodal imaging according to claim 3, characterized in that, The specific steps of S2 are as follows: S201: Based on the synchronized image data, extract the pixel array of each modality image, apply the edge detection operator to calculate the bidirectional gray-level derivative and synthesize the gray-level gradient magnitude, identify lesion areas that exceed the preset judgment threshold, use the co-occurrence matrix to obtain the texture direction gradient of the lesion area, and establish a modal multidimensional gradient set. S202: For the modal multidimensional gradient set, separate the gray-level gradient magnitude and texture direction gradient of different modes, perform subtraction on the gray-level gradient magnitude of adjacent modes according to the coordinate correspondence to obtain the gray-level difference, perform subtraction on the texture direction gradient at the same position to obtain the texture difference, and map the two types of differences to the grid nodes to obtain the feature difference map. S203: The feature difference map is called to divide the spatial domain into grid blocks. The grid blocks are scanned by the convolution operator to extract the pixel distribution frequency to form local spatial features. The feature values ​​are accumulated by the operator moving along the step size, and the feature values ​​are arranged in the original order to establish a basic feature mapping matrix.

5. The method for determining the benign or malignant nature of breast tumors based on multimodal imaging according to claim 4, characterized in that, The preset judgment threshold is determined by an adaptive value selection method based on the statistical distribution of bidirectional gray-level derivatives. Specifically, the bidirectional gray-level derivatives are statistically analyzed using a histogram in the lesion area, and the gray-level gradient amplitude corresponding to the cumulative distribution located in the preset high-level proportion interval is selected as the preset judgment threshold.

6. The method for determining the benign or malignant nature of breast tumors based on multimodal imaging according to claim 4, characterized in that, The specific steps for S3 are as follows: S301: Call the basic feature mapping matrix to extract the local spatial feature sequence, transform the local spatial feature sequence to the frequency domain through Fourier transform, analyze the signal intensity features and signal offset state to obtain complex frequency components, extract the corresponding amplitude spectrum parameters and phase spectrum parameters, and establish a set of spatial frequency components. S302: Extract frequency domain data from the spatial frequency component set, decouple high-frequency texture components and low-frequency background components according to a preset cutoff frequency threshold, analyze the corresponding amplitude spectrum parameters, quantify the high-frequency energy attributes of the lesion edge and the low-frequency energy attributes of the main structure, and analyze the distribution ratio of the two to obtain the frequency band energy ratio. S303: Call the frequency band energy ratio item to retrieve the original row and column index identifier, map the frequency band energy ratio item to the grid node according to the original row and column index identifier and perform numerical filling, perform pure zero filling operation for the grid coordinates of non-lesion areas to complete the background features, and generate a frequency domain energy distribution matrix.

7. The method for determining the benign or malignant nature of breast tumors based on multimodal imaging according to claim 6, characterized in that, The preset cutoff frequency threshold is adaptively determined based on the amplitude spectrum parameter distribution characteristics of the spatial frequency component set. After sorting the amplitude spectrum parameters from low to high frequency, the corresponding frequency at which the cumulative amplitude spectrum energy reaches a predetermined proportion of the total amplitude spectrum energy is selected as the preset cutoff frequency threshold.

8. The method for distinguishing benign and malignant breast tumors based on multimodal imaging according to claim 6, characterized in that, The specific steps of S4 are as follows: S401: Extract the node values ​​of the frequency domain energy distribution matrix and perform a difference operation with a preset baseline. Construct a high-frequency distribution mask and a low-frequency distribution mask based on the positive and negative attributes of the difference. Apply the high-frequency mask and the low-frequency mask to perform dot product separation on the basic feature mapping matrix to establish a multi-band mapping combination. S402: Extract the energy density parameters of the high-frequency mapping region in the multi-band mapping combination, convert the energy density parameters into scaling weights through preset dimensionless coefficients and multiply them with the initial radius to obtain the target receptive field boundary, apply the convolution kernel corresponding to the target receptive field boundary along the high-frequency mapping region to perform dot product traversal, extract dense texture vectors, and obtain a variable radius texture feature matrix. S403: Extract the energy sparse parameters of the low-frequency mapping region based on the multi-band mapping combination, extract the dimensionless step size adjustment coefficient based on the energy sparse parameters and multiply it with the initial step size to obtain the target sliding step size, apply the convolution kernel corresponding to the target sliding step size to perform jump dot product along the low-frequency mapping region, extract the sparse background vector, and perform channel concatenation between the variable radius texture feature matrix and the sparse background vector to generate a modal decoupling feature map.

9. The method for determining the benign or malignant nature of breast tumors based on multimodal imaging according to claim 8, characterized in that, The specific steps of S5 are as follows: S501: The modal decoupling feature map and frequency domain energy distribution matrix are called to extract local feature pixels and frequency domain distribution values. The node fusion weights are set according to the frequency domain distribution values. The local feature pixels are multiplied by the node fusion weights to perform numerical mapping. The matrix channel accumulation is performed along the spatial dimension to establish a cross-modal feature matrix. S502: Extract the internal texture pixel array based on the cross-modal feature matrix, determine the texture continuity parameter by calculating the gradient difference between adjacent grid nodes, determine the gray balance parameter based on the global pixel gray level distribution variance, and concatenate the texture continuity parameter and the gray level balance parameter to obtain the morphological evaluation parameter set. S503: For the set of morphological evaluation parameters, the texture continuity parameter and the grayscale balance parameter are multiplied by their corresponding weight coefficients and summed to obtain the evaluation morphological index. This index is then compared with a preset benign / malignant boundary threshold. If the index is greater than the benign / malignant boundary threshold, a malignant warning label is assigned; otherwise, a benign / normal label is assigned, thus generating a breast tumor discrimination result.

10. The method for determining the benign or malignant nature of breast tumors based on multimodal imaging according to claim 9, characterized in that, The preset benign / malignant boundary threshold is limited to a fixed value within the quantile range obtained from historical sample statistics, and the quantile range is determined by the range between the median and the upper quartile of the evaluation morphological indicators corresponding to all training samples.