Post-structured decomposition method based on image deep learning
By using a dual-path residual convolutional neural network and an adaptive attention mechanism to fuse local and global features in deep learning, and performing hierarchical feature analysis and dynamic weighting strategies, the shortcomings of feature extraction and structured decomposition in the existing technology are solved, and efficient and accurate image data processing and cross-domain data fusion are achieved.
Patent Information
- Application Number
- CN202510240651.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-06-10
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the process of deep learning feature extraction, post-structured decomposition and data structure, the lack of hierarchical structure of feature expression, post-processing methods fail to effectively reflect the intrinsic correlation of image features, and data structure fails to meet the needs of cross-domain fusion.
The post-structured decomposition method based on deep learning of images is adopted, and local and global features are fused through a dual-path residual convolutional neural network and an adaptive attention mechanism. Hierarchical feature analysis and dynamic weighting strategies are adopted to carry out post-structured decomposition and feature subset generation, and finally the feature subset is converted into a standardized structured data format.
It effectively improves the accuracy and efficiency of image data processing, avoids information redundancy and loss, realizes high-quality feature representation conversion, has efficient, accurate and automated data analysis capabilities, and supports cross-domain data fusion and comprehensive analysis.
Smart Images

Figure CN120125841A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of structured data, and particularly to a post-structured decomposition method based on image deep learning. Background Art
[0002] With the progress of digital image acquisition technology and the wide application of deep learning in various fields, the acquisition, processing, and analysis of image data have become increasingly important. Especially in the fields of medicine, security, industrial inspection, remote sensing monitoring, etc., the generation of image data has grown explosively. With the diversification and complexity of image data, traditional image processing methods have gradually been unable to meet the needs of modern data analysis. In this context, image analysis technology based on deep learning has emerged and achieved remarkable results in tasks such as image recognition, object detection, segmentation, and classification.
[0003] Existing image analysis technologies mostly rely on deep learning networks (such as convolutional neural networks, CNNs) to automatically extract features from image data. These deep learning networks can gradually extract low-level features to high-level features from images through multiple layers of convolution and pooling operations, and finally obtain a high-dimensional feature representation that can represent the content of the image. This feature extraction method has significant advantages over traditional methods (such as manually designed feature extraction algorithms) because it can automatically learn features suitable for the task from the data without manual intervention. However, although these deep learning methods can extract rich features from image data and have strong generalization ability, the high-dimensional features they extract are often still unstructured. This means that although the features contain a large amount of useful information, this information does not have a clear hierarchical structure and is not necessarily easily utilized directly by subsequent analysis and decision support systems.
[0004] Existing deep learning methods usually focus on extracting local features and global features, but often separate the two in the processing process and lack effective fusion. Traditional convolutional neural networks (CNNs) perform well in extracting local features, but for capturing global features and overall understanding of image content, they often rely on complex multi-layer networks and attention mechanisms, which often fail to fully integrate local details and global context, resulting in some detail information being ignored in global analysis. At the same time, although some improved convolutional neural networks (such as residual networks) and attention mechanisms can capture more global information, these methods still face the problem of how to effectively fuse local features and global features. Especially in the application of complex image data, how to balance and fuse local and global information to obtain an accurate high-dimensional feature representation remains a challenge.
[0005] In addition, although existing deep learning methods have been able to extract various meaningful features from images, these features are usually high-dimensional unstructured data. Although deep learning networks can optimize these high-dimensional features, they lack hierarchical and structured expressions of feature information. This leads to the difficulty of directly applying high-dimensional features to actual decision support systems during subsequent analysis. For example, in medical image analysis, doctors need to extract certain key structural information from image data, such as the size, shape, and location of tumors. These information are not only the "features" of the images, but also require hierarchical representation and structured output for effective utilization in subsequent clinical decisions. However, traditional feature extraction and representation methods have not fully met this requirement and cannot transform the complexity of image data into structured data that is easy to understand and utilize, thus affecting the effectiveness and efficiency of subsequent analysis.
[0006] Another major problem in the prior art is the post-processing and structuring of features. When performing image analysis, it is usually necessary to post-structure the extracted high-dimensional features in order to transform them into a data format suitable for further analysis, decision-making, or docking with other systems. However, traditional feature processing methods mainly rely on simple grouping and clustering algorithms, which classify or cluster features based on static similarity metrics. Although this method may be effective in some simple tasks, when dealing with complex image data, due to the complex and hierarchical relationships between features, traditional methods often cannot effectively distinguish different levels of feature information, resulting in feature redundancy and information loss. This method also often fails to fully utilize the multi-level information in image data and cannot effectively extract semantic information at each level from images, thereby affecting the effectiveness of subsequent tasks.
[0007] In addition, in the process of converting image data into structured data, the prior art often relies on simple statistical methods and rule processing. These methods fail to fully explore the deep semantics in image data, resulting in the generated structured data lacking sufficient semantic information and being difficult to meet the requirements of complex decision support systems. In many practical applications, image data needs to be fused with other types of data, such as patient medical record information, time series of image data, environmental data, etc. There is a lack of effective mechanisms in the prior art to achieve cross-domain data fusion and comprehensive analysis, resulting in low data utilization efficiency and the phenomenon of information silos.
[0008] In summary, there are still many deficiencies in the prior art during the process of image deep learning feature extraction, post-structural decomposition, and data structuring. The main problems are concentrated in: insufficient fusion of local and global information, making it difficult to comprehensively extract multi-level information of images; the representation of high-dimensional unstructured features lacks a hierarchical structure and is difficult to directly apply to practical applications; the grouping and clustering methods in the post-processing stage fail to effectively reflect the internal correlation of image features, resulting in information loss and redundancy; the data structuring process fails to fully meet the requirements of cross-domain fusion. Summary of the Invention
[0009] An object of the present invention is to propose a post-structural decomposition method based on image deep learning. The present invention combines deep learning with the post-structural decomposition method, and through fusing local and global features, adopting a hierarchical feature parsing and dynamic weighting strategy, innovatively extracts and structures image data, effectively improving the accuracy and efficiency of image data processing. At the same time, it avoids the problems of information redundancy and loss in traditional technologies, realizes high-quality feature representation conversion, and has efficient, accurate, and automated data analysis capabilities.
[0010] A post-structural decomposition method based on image deep learning according to an embodiment of the present invention includes the following steps:
[0011] S1. Collect original image data through an image acquisition device, and preprocess the original image data to generate preprocessed image data;
[0012] S2. Based on the preprocessed image data, use a dual-path residual convolutional neural network and an adaptive attention mechanism to simultaneously extract low-level and high-level local features through two parallel convolutional branches;
[0013] S3. Perform a linear mapping operation on the extracted low-level and high-level local features to generate query, key, and value matrices, use an improved attention mechanism, introduce a dynamic adjustment term to calculate the attention score, and perform weighted summation on the value matrix to generate a global feature;
[0014] S4. Input the local feature and the global feature into a feature fusion network for feature fusion, and adopt an adaptive weighting strategy to generate a fused high-dimensional unstructured feature;
[0015] S5. Input the high-dimensional unstructured feature representation into a post-structural decomposition module, perform an overall analysis and decomposition on the high-dimensional unstructured feature representation to form a feature subset describing each level of information in the image data;
[0016] S6. Convert the feature subset into a standardized structured data format.
[0017] Optionally, the original image data is a digital image containing multiple color channels.
[0018] Optionally, the preprocessing includes noise filtering, image enhancement, resolution correction, and color normalization of the original image data.
[0019] Optionally, step S2 specifically includes:
[0020] S21. Input the preprocessed image data into a dual-path residual convolutional neural network, which includes a first convolutional branch and a second convolutional branch;
[0021] S22. Use the first convolutional branch to extract local features from the preprocessed image data:
[0022]
[0023] where F local1 (x, y) represents the local features output by the first convolutional branch, ReLU represents the rectified linear activation function, k represents the convolutional kernel radius, W 1 (i, j) represents the weight of the first convolutional kernel at position (i, j), b 1 represents the corresponding bias parameter, I preproc (x + i, y + j) represents the input preprocessed image data, and (x, y) represents the pixel coordinates;
[0024] S23. Use the second convolutional branch to extract local features from the preprocessed image data:
[0025]
[0026] where F local2 (x, y) represents the local features output by the second convolutional branch, ReLU represents the rectified linear activation function, k represents the convolutional kernel radius, W 2 (i, j) represents the weight of the second convolutional kernel at position (i, j), b 2 represents the corresponding bias parameter;
[0027] S24. Perform weighted fusion on the local features F local1 (x, y) and F local2 (x, y) output by the first and second convolutional branches to obtain fused local features:
[0028] F local (x, y) = β · F local1 (x, y) + (1 - β) · F local2 (x, y);
[0029] where F local (x, y) represents the fused local features, and β represents the weighting coefficient;
[0030] S25. Apply an adaptive attention mechanism to the fused local feature F local (x, y) to calculate the attention coefficient for color channel c:
[0031]
[0032] where A(c) represents the attention coefficient for color channel c, σ represents the Sigmoid activation function, M and N respectively represent the number of pixels in the vertical and horizontal directions of the image, and c represents the color channel index;
[0033] S26. Use the attention coefficient A(c) to weight the fused local feature F local (x, y, c) to obtain the local feature representation:
[0034] F′ local (x, y, c) = A(c) · F local (x, y, c);
[0035] where F′ local (x, y, c) represents the local feature representation, and F local (x, y, c) represents the fused local feature.
[0036] Optionally, the specific steps of S3 include:
[0037] S31. Receive the local feature representation F′ output by step S31 local , where the local feature representation F′ local represents the local feature matrix processed by the dual-path convolutional neural network and the adaptive attention mechanism. The dimension of the local feature matrix is H×W×C, where H and W in the dimension respectively represent the height and width of the feature map, and C represents the number of feature channels;
[0038] S32. Perform a linear mapping operation on the local feature representation F′ local to generate a query matrix, a key matrix, and a value matrix:
[0039] Q = F′ local ·W Q , K = F′ local ·W K , V = F′ local ·W V ;
[0040] where Q represents the query matrix, K represents the key matrix, V represents the value matrix, and W Q represents the weight matrix of the query matrix, W K represents the weight matrix of the key matrix, and W V represents the weight matrix of the value matrix;
[0041] S33. Perform a dynamic convolution operation on the local feature representation F′ local to generate a dynamic adjustment term. By performing a convolution operation on F′ local with the dynamic convolution kernel weights and adding the bias, then using a non-linear activation function for mapping:
[0042]
[0043] where, Δ represents the dynamic adjustment term, represents the non-linear activation function, Conv represents the two-dimensional convolution operation, W d represents the dynamic convolution kernel weights, b d represents the bias parameter;
[0044] S34. Calculate the attention score matrix:
[0045]
[0046] where, A represents the attention score matrix, softmax represents row-wise normalization, Q represents the query matrix, K represents the key matrix, T represents the matrix transpose operation, represents the scaling factor to ensure numerical stability;
[0047] S35. Use the attention score matrix to perform weighted summation on the value matrix to generate the global feature representation:
[0048] F global = A · V;
[0049] where, F global represents the global feature representation.
[0050] Optionally, the feature fusion network adopts an adaptive weighting strategy to generate the fused feature:
[0051] F fusion = αF local + (1 - α)F global ;
[0052] where, F fusion represents the fused feature, F local represents the local feature, F global represents the global feature, α represents the adaptive weight:
[0053] α = σ(W α [F local , F global + b α );
[0054] where, σ represents the Sigmoid activation function, which is used to ensure that the value range of the adaptive weight α is between 0 and 1, Wα represents the fusion weight parameter matrix, b α represents the bias parameter.
[0055] Optionally, the S5 specifically includes:
[0056] S51. Calculate the cosine similarity between any two feature vectors F fusion in the fused high-dimensional unstructured feature representation F fusion (i) and F fusion (j):
[0057]
[0058] where S(i, j) represents the cosine similarity between any two feature vectors, F fusion (i) and F fusion (j) represent any two feature vectors in the high-dimensional unstructured feature table, and ∥·∥ represents the Euclidean norm;
[0059] S52. Construct an adjacency matrix according to a preset similarity threshold:
[0060]
[0061] where A(i, j) represents the adjacency matrix, and θ represents the preset similarity threshold;
[0062] S53. Analyze the adjacency matrix A(i, j) to obtain a preliminary grouping set {G 1 , G 2 , …, G K}, and each grouping G K in the preliminary grouping set includes all feature vectors whose cosine similarity to each other is not less than θ;
[0063] S54. Calculate the centroid vector for each grouping in the preliminary grouping set:
[0064]
[0065] where μ k represents the centroid vector, |G k | represents the number of feature vectors included in the grouping G K , and F fusion (i) represents the feature vector of the high-dimensional unstructured feature representation;
[0066] S55. Calculate the cosine similarity between any two groupings of the centroid vectors of all groupings:
[0067]
[0068] where S(G p, G q ) represents the cosine similarity between any two groups G p and G q , μ p and μ q respectively represent the centroid vectors of groups G p and G q , ∥μ p ∥ represents the Euclidean norm of group G p ;
[0069] S56. Set the initial hierarchical threshold θ (1) = θ, and in the l-th level, for any two groups G p and G q , if the following conditions are met:
[0070] S(G p , G q ) ≥ θ (l) ;
[0071] then merge groups G p and G q into a new group G r = G p ∪G q , and update the centroid vector:
[0072]
[0073] where μ r represents the centroid vector of the new group G r , |G r | represents the number of feature vectors contained in group G r ;
[0074] S57. Update the threshold, and repeat the merging process of steps S55 and S56 until there are no groups that meet the merging conditions, forming the intermediate clustering result of iterative hierarchical clustering:
[0075] θ (l+1) = θ (l) - δ;
[0076]
[0077] where θ (l+1) represents the updated threshold, δ > 0 represents a predetermined constant, F (l) represents the intermediate clustering result of iterative hierarchical clustering, represents the K (l) -th clustering group formed during the l-th level clustering process, K (l) represents the total number of clustering groups formed during the l-th level clustering process;
[0078] S58. Integrate the intermediate clustering results F at each level, merge the groups that have intersections or inheritance relationships in different clustering levels, and define the set of final feature subsets. Each feature subset F in the set of final feature subsets (l) satisfies: k
[0079]
[0080] where L k represents the index set of all clustering levels that contain the same or compatible feature information, and ensures that any feature vector appears in only one unique feature subset. F k represents the feature subset in the set of final feature subsets, represents the k-th clustering group formed during the clustering process at the l-th level, represents the merging of all clustering groups k belonging to the index set L to obtain the final feature subset F k , and ∪ represents the union operation of sets;
[0081] S59. Verify the set of integrated feature subsets and output the set of final feature subsets For any two feature vectors F k (i) and F fusion (j) in any F fusion , there is:
[0082]
[0083] where |F k | represents the number of feature vectors contained in the feature subset F k , θ min represents the predetermined consistency threshold, and K final represents the total number of subsets in the set of final feature subsets obtained after integrating the intermediate clustering results.
[0084] Optionally, the S6 specifically includes:
[0085] S61. Perform statistical summary processing on each feature subset F in the set of final feature subsets to generate a summary feature vector: k
[0086]
[0087] where V k represents the summary feature vector, |F k | represents the number of feature vectors contained in the feature subset F k , and f represents the feature vector belonging to the feature subset F k Any eigenvector among them;
[0088] S62. Perform linear normalization on the aggregated eigenvector V k to generate a normalized eigenvector:
[0089]
[0090] where R k represents the normalized eigenvector, and μ and σ represent the mean vector and the standard deviation vector respectively:
[0091]
[0092]
[0093] where K final represents the total number of subsets in the set of final feature subsets obtained after integrating the intermediate clustering results;
[0094] S63. Construct the normalized eigenvector R k into a data record D in a fixed format k , and the data record D k is defined as the field {ID k , R k}, and the ID in the field {ID k , R k} k represents the unique identifier of the feature subset F k , and R k represents the corresponding predetermined fields {F 1 , F 2 , …, F d} in sequence for each component;
[0095] S64. Arrange all the data records D k in a fixed order to form a structured data set:
[0096] D = {D k ∣k = 1, 2, …, K final};
[0097] where D represents the structured data set, and the structured data set is stored in tabular form, with each row corresponding to a data record D k , and each column corresponding to a predetermined field.
[0098] The beneficial effects of the present invention are:
[0099] First, in view of the deficiency of the traditional image data processing method in separately extracting local features and global features, the present invention effectively combines the advantages of both by fusing local and global information, and comprehensively captures the detailed and semantic information in the image data by using a dual-path residual convolutional neural network and an adaptive attention mechanism. Through this innovative design, the network can simultaneously extract low-level and high-level local features, thereby improving the integrity and accuracy of feature representation, overcoming the limitation of the traditional method that can only extract features locally, and making image analysis more refined and comprehensive.
[0100] Second, after feature extraction, the present invention further introduces a post-structured decomposition module, which adopts a hierarchical feature parsing method. Different from the traditional simple clustering method, the present invention can effectively decompose and organize the multi-level information in the image data by calculating the similarity of high-dimensional unstructured features, constructing an adjacency matrix, and applying a dynamic threshold and a hierarchical clustering strategy. Through this innovative decomposition method, the features at each level in the image data can be better distinguished and expressed, making the subsequent feature subsets more hierarchical and structured, thus better supporting data analysis and decision-making support. Compared with the single grouping and clustering methods in the prior art, the post-structured decomposition method of the present invention can more accurately capture the subtle differences in the image data, avoiding information redundancy and loss.
[0101] In addition, the present invention combines local features and global features through an adaptive weighted fusion strategy, and introduces a dynamic adjustment term in the feature fusion process, so that the weights of features can be automatically adjusted at each stage of the extraction process. This adaptive mechanism enables the present invention to automatically adjust the relationship between features according to specific image data, further enhancing the flexibility and accuracy of feature representation. Through this mechanism, the present invention can better adapt to different types of image data and provide high-quality feature representation in different application scenarios.
[0102] Finally, in the structured data conversion stage, the present invention ensures the standardization and consistency of the data by statistically summarizing, normalizing the generated feature subsets, and converting them into data records in a fixed format. This processing not only enables the image data to be directly applied to data analysis and decision-making support systems, but also enables seamless docking and fusion with other types of data systems. By converting high-dimensional unstructured features into structured data that is easy to process and analyze, the present invention improves the utilization efficiency of the image data and avoids the defects of complex and inefficient data conversion in the traditional image data processing method. BRIEF DESCRIPTION OF THE DRAWINGS
[0103] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification, and are used to explain the present invention together with the embodiments of the present invention, and do not constitute a limitation to the present invention. In the drawings:
[0104] Figure 1 Flow chart of a post-structured decomposition method based on image deep learning proposed by the present invention;
[0105] Figure 2 Detailed structure diagram of the post-structured decomposition module of a post-structured decomposition method based on image deep learning proposed by the present invention. Specific embodiments
[0106] Now, the present invention will be further described in detail with reference to the accompanying drawings. These drawings are all simplified schematic diagrams, only showing the basic structure of the present invention in a schematic way, so they only show the components related to the present invention.
[0107] Reference Figure 1 and Figure 2 , a post-structured decomposition method based on image deep learning, includes the following steps:
[0108] S1. Collect the original image data through an image acquisition device, and preprocess the original image data to generate preprocessed image data;
[0109] S2. Based on the preprocessed image data, use a dual-path residual convolutional neural network and an adaptive attention mechanism to simultaneously extract low-level and high-level local features through two parallel convolutional branches;
[0110] S3. Perform a linear mapping operation on the extracted low-level and high-level local features to generate query, key, and value matrices, use an improved attention mechanism, introduce a dynamic adjustment term to calculate the attention score, and perform weighted summation on the value matrix to generate a global feature;
[0111] S4. Input the local features and the global feature into a feature fusion network for feature fusion, and adopt an adaptive weighting strategy to generate the fused high-dimensional unstructured feature;
[0112] S5. Input the high-dimensional unstructured feature representation into the post-structured decomposition module, perform overall analysis and decomposition on the high-dimensional unstructured feature representation, and form a feature subset describing the information at each level in the image data;
[0113] S6. Convert the feature subset into a standardized structured data format.
[0114] In this embodiment, the original image data is a digital image including multiple color channels.
[0115] In this embodiment, the preprocessing includes noise filtering, image enhancement, resolution correction, and color normalization of the original image data.
[0116] In this embodiment, the specific content of S2 includes:
[0117] S21. Input the preprocessed image data into a dual-path residual convolutional neural network, where the dual-path residual convolutional neural network includes a first convolutional branch and a second convolutional branch;
[0118] S22. Use the first convolutional branch to extract local features from the preprocessed image data:
[0119]
[0120] where F local1 (x, y) represents the local features output by the first convolutional branch, ReLU represents the rectified linear unit activation function, k represents the convolutional kernel radius, W 1 (i, j) represents the weight of the first convolutional kernel at position (i, j), b 1 represents the corresponding bias parameter, I preproc (x + i, y + j) represents the input preprocessed image data, and (x, y) represents the pixel coordinates;
[0121] S23. Use the second convolutional branch to extract local features from the preprocessed image data:
[0122]
[0123] where F local2 (x, y) represents the local features output by the second convolutional branch, ReLU represents the rectified linear unit activation function, k represents the convolutional kernel radius, W 2 (i, j) represents the weight of the second convolutional kernel at position (i, j), b 2 represents the corresponding bias parameter;
[0124] S24. Perform weighted fusion on the local features F local1 9x, y) and F local2 (x, y) output by the first and second convolutional branches to obtain fused local features:
[0125] F local (x, y) = β · F local1 (x, y) + (1 - β) · F local2 (x, y);
[0126] where F local (x, y) represents the fused local features, and β represents the weighting coefficient;
[0127] S25. Apply an adaptive attention mechanism to the fused local features F local (x, y) to calculate the attention coefficient for color channel c:
[0128]
[0129] Among them, A(c) represents the attention coefficient of color channel c, σ represents the Sigmoid activation function, M and N respectively represent the number of pixels in the vertical and horizontal directions of the image, and c represents the color channel index;
[0130] S26. Use the attention coefficient A(c) to weight the fused local feature F local (x, y, c) to obtain a local feature representation:
[0131] F′ local (x, y, c) = A(c) · F local (x, y, c);
[0132] Among them, F′ local (x, y, c) represents the local feature representation, and F local (x, y, c) represents the fused local feature.
[0133] In this embodiment, the S3 specifically includes:
[0134] S31. Receive the local feature representation F′ output by step S31 local , and the local feature representation F′ local represents the local feature matrix processed by the dual-path convolutional neural network and the adaptive attention mechanism. The dimension of the local feature matrix is H×W×C, where H and W in the dimension respectively represent the height and width of the feature map, and C represents the number of feature channels;
[0135] S32. Perform a linear mapping operation on the local feature representation F′ local to generate a query matrix, a key matrix, and a value matrix:
[0136] Q = F′ local ·W Q , K = F′ local ·W K , V = F′ local ·W V ;
[0137] Among them, Q represents the query matrix, K represents the key matrix, V represents the value matrix, and W Q represents the weight matrix of the query matrix, W K represents the weight matrix of the key matrix, and W V represents the weight matrix of the value matrix;
[0138] S33. Perform a dynamic convolution operation on the local feature representation F′ local to generate a dynamic adjustment term. Through convolution operation on F′ local and the dynamic convolution kernel weight, and adding a bias, then use a non-linear activation function for mapping:
[0139]
[0140] Among them, Δ represents the dynamic adjustment term, represents the non-linear activation function, Conv represents the two-dimensional convolution operation, W d represents the dynamic convolution kernel weight, b d represents the bias parameter;
[0141] S34. Calculate the attention score matrix:
[0142]
[0143] Among them, A represents the attention score matrix, softmax represents row-wise normalization processing, Q represents the query matrix, K represents the key matrix, T represents the matrix transpose operation, represents the scaling factor to ensure numerical stability;
[0144] S35. Use the attention score matrix to perform weighted summation on the value matrix to generate the global feature representation:
[0145] F global = A · V;
[0146] Among them, F global represents the global feature representation.
[0147] In this embodiment, the feature fusion network adopts an adaptive weighting strategy to generate the fusion feature:
[0148] F fusion = αF local + (1 - α)F global ;
[0149] Among them, F fusion represents the fused feature, F local represents the local feature, F global represents the global feature, α represents the adaptive weight:
[0150] α = σ(W α [F local , F global + b α );
[0151] Among them, σ represents the Sigmoid activation function, which is used to ensure that the value range of the adaptive weight α is between 0 and 1, W α represents the fusion weight parameter matrix, b α represents the bias parameter.
[0152] In this embodiment, the specific content of S5 includes:
[0153] S51. Calculate the cosine similarity for any two feature vectors F fusion in the fused high-dimensional unstructured feature representation F fusion (i) and F fusion (j):
[0154]
[0155] where S(i, j) represents the cosine similarity of any two feature vectors, F fusion (i) and F fusion (j) represent any two feature vectors in the high-dimensional unstructured feature table, and ∥·∥ represents the Euclidean norm;
[0156] S52. Construct an adjacency matrix based on a preset similarity threshold:
[0157]
[0158] where A(i, j) represents the adjacency matrix and θ represents the preset similarity threshold;
[0159] S53. Analyze the adjacency matrix A(i, j) to obtain a preliminary grouping set {G 1 , G 2 , …, G K}, where each grouping G K in the preliminary grouping set includes all feature vectors whose cosine similarity to each other is not less than θ;
[0160] S54. Calculate the centroid vector for each grouping in the preliminary grouping set:
[0161]
[0162] where μ k represents the centroid vector, |G k | represents the number of feature vectors included in the grouping G K , and F fusion (i) represents the feature vector of the high-dimensional unstructured feature representation;
[0163] S55. Calculate the cosine similarity between any two groupings for the centroid vectors of all groupings:
[0164]
[0165] where S(G p , G q ) represents the cosine similarity between any two groupings G p and G q , and μ p and μ q respectively represent the centroid vectors of the groupings Gp and G q 's centroid vector, ∥μ p ∥ represents the Euclidean norm of the grouping G p ;
[0166] S56. Set the initial hierarchical threshold θ (1) = θ, and in the l-th level, for any two groupings G p and G q , if the following condition is satisfied:
[0167] S(G p , G q ) ≥ θ (l) ;
[0168] Then merge the groupings G p and G q into a new grouping G r = G p ∪ G q , and update the centroid vector:
[0169]
[0170] where μ r represents the centroid vector of the new grouping G r , |G r | represents the number of feature vectors contained in the grouping G r ;
[0171] S57. Update the threshold, and repeat the merging process of steps S55 and S56 until there are no groupings that meet the merging conditions, forming the intermediate clustering result of iterative hierarchical clustering:
[0172] θ (l+1) = θ (l) - δ;
[0173]
[0174] where θ (l+1) represents the updated threshold, δ > 0 represents a predetermined constant, F (l) represents the intermediate clustering result of iterative hierarchical clustering, represents the K (l) -th clustering group formed during the l-th level clustering process, K (l) represents the total number of clustering groups formed during the l-th level clustering process;
[0175] S58. Integrate the intermediate clustering results F (l) at each level, unify and merge the groupings that have intersections or inheritance relationships in different clustering levels, and define the set of final feature subsets. Each feature subset F in the set of final feature subsetsk Satisfy:
[0176]
[0177] Wherein, L k represents the index set of all clustering levels containing the same or compatible feature information, and ensures that any feature vector appears in only a unique feature subset, F k represents the feature subset in the set of final feature subsets, represents the k-th clustering group formed in the l-th level clustering process, represents combining all the clustering groups belonging to the index set L k to obtain the final feature subset F , ∪ represents the union operation of sets; k , ∪ represents the union operation of sets;
[0178] S59. Verify the integrated set of feature subsets and output the set of final feature subsets For any two feature vectors F k in any F fusion (i) and F fusion (j), there is:
[0179]
[0180] Wherein, |F k | represents the number of feature vectors contained in the feature subset F k , θ min represents a predetermined consistency threshold, and K final represents the total number of subsets in the set of final feature subsets obtained after integrating the intermediate clustering results.
[0181] In this embodiment, the S6 specifically includes:
[0182] S61. Perform statistical summarization processing on each feature subset F k in the set of final feature subsets to generate a summary feature vector:
[0183]
[0184] Wherein, V k represents the summary feature vector, |F k | represents the number of feature vectors contained in the feature subset F k , and f represents any feature vector belonging to the feature subset F k ;
[0185] S62. Perform linear normalization processing on the summary feature vector V k to generate a normalized feature vector:
[0186]
[0187] Among them, R k represents a normalized eigenvector, and μ and σ represent the mean vector and the standard deviation vector respectively:
[0188]
[0189] Among them, K final represents the total number of subsets in the set of final feature subsets obtained after integrating the intermediate clustering results;
[0190] S63. Construct the normalized eigenvector R k into a data record D in a fixed format k , and the data record D k is defined as the field {ID k , R k}, and the ID in the field {ID k , R k} k represents the unique identifier of the feature subset F k , and R k represents the corresponding predetermined fields {F 1 , F 2 , …, F d} in sequence for each component;
[0191] S64. Arrange all the data records D k in a fixed order to form a structured data set:
[0192] D = {D k ∣k = 1, 2, …, K final};
[0193] Among them, D represents the structured data set, and the structured data set is stored in tabular form, with each row corresponding to a data record D k , and each column corresponding to a predetermined field.
[0194] Example 1:
[0195] To verify the feasibility of the present invention in implementation, the present invention is applied to a medical imaging diagnosis system, especially for experimentation in the automatic diagnosis of lung CT images. The automatic diagnosis of lung CT images is an important field in medical image analysis. Especially in the early screening of lung cancer, rapid diagnosis and monitoring of lung diseases, it can effectively improve the diagnosis efficiency and accuracy. However, a common problem in current technologies is that the high-dimensional features extracted by existing deep learning methods are usually unstructured and lack fine hierarchical information, making it difficult to directly apply the image data to a medical decision support system. At the same time, traditional image analysis methods usually rely on manual analysis or rule-based segmentation methods, which are time-consuming and error-prone. Therefore, solving how to effectively convert complex image data into structured data and provide support for medical decisions has become a key technical problem in solving the automatic diagnosis of medical images.
[0196] In this embodiment, the image data used comes from the lung CT image library of a certain tertiary hospital. This library contains the lung CT scan images of 2000 patients, and the images include different types of lung diseases, such as lung cancer, pulmonary tuberculosis, pneumonia, etc. All image data has been manually annotated by the hospital's medical imaging department, including the annotation of the lung lesion areas for each case. Through the method of the present invention, we use a deep learning network to automatically analyze these CT images, extract the detailed features in the images, and then convert these features into structured data for more efficient diagnostic support.
[0197] In this scenario, first, the original lung CT image data is obtained through an image acquisition device. Since most of these CT image data are two-dimensional images and contain a large amount of detailed information, traditional image analysis methods usually have difficulty in efficiently extracting the key information in the images. According to the method of the present invention, first, preprocessing steps such as noise filtering, image enhancement, resolution correction, and color normalization are performed on the original image data. Through these steps, the quality of the images has been significantly improved. Especially in some low-quality CT scans, removing noise and enhancing image details make the images clearer and enhance the effect of subsequent feature extraction.
[0198] Next, the preprocessed image data is processed using a dual-path residual convolutional neural network and an attention mechanism, which can simultaneously extract the local details and global semantic information in the images. The extraction of local features helps to accurately locate the lung lesion areas, while the capture of global features can provide the overall health status of the lungs. The fusion of this local and global information not only improves the positioning accuracy of the lesion areas but also helps the model better understand the overall condition of the lungs, solving the problem that traditional methods only rely on local features.
[0199] After that, the extracted high-dimensional unstructured features are input into the post-structural decomposition module. Through calculating the similarity between features, constructing an adjacency matrix, and performing hierarchical clustering, this module finally converts the extracted high-dimensional features into multiple hierarchical feature subsets. Each feature subset represents a certain type of important information in the image, such as the shape, size, boundary, etc. of the lesion area. Through this process, the unstructured data in the image is transformed into a standardized structured data format, and these structured data can be directly used in subsequent analysis and decision support systems, greatly improving the efficiency and accuracy of image analysis.
[0200] Through these steps, the finally output structured data can provide real-time support for the diagnosis of lung CT images. Doctors can make more accurate judgments based on these structured data, improving the accuracy and efficiency of diagnosis.
[0201] To verify the effectiveness of the method of the present invention, a comparative experiment was conducted in a certain top-three hospital. In the experiment, the lung CT image data of 2000 patients were used. These data were manually annotated by professional doctors and were subjected to image segmentation and annotation through the hospital's radiology department. Two sets of data processing methods were set in the experiment: one set used the method of the present invention, and the other set used the traditional image analysis method (a segmentation algorithm based on handcrafted features and rules). By comparing indicators such as the accuracy rate, diagnosis time, and error rate of the two methods in the diagnosis of lung diseases, the effectiveness of the method of the present invention was verified.
[0202] Table 1 Comparative table of experimental data
[0203]
[0204] According to the above comparative table of experimental data, it can be clearly seen that the method of the present invention is superior to the traditional method in multiple key performance indicators. First of all, in terms of the accuracy rate, the method of the present invention reached 94.5%, significantly higher than 85.3% of the traditional method. This indicates that the present invention has a higher correct classification ability in the automatic diagnosis of lung CT images, can more accurately identify the lesion areas in the images, thus improving the reliability and effectiveness of diagnosis.
[0205] In terms of the diagnosis time, the diagnosis time of the method of the present invention is 25.4 seconds. Compared with 52.1 seconds of the traditional method, it reduces by nearly half of the time. This means that adopting the method of the present invention can not only provide more accurate diagnosis results, but also greatly improve the diagnosis efficiency, save time for doctors and speed up the diagnosis process, especially suitable for the processing of large-scale image data and real-time medical scenarios.
[0206] In terms of the error rate, the error rate of the method of the present invention is 5.5%, which is significantly lower than 14.7% of the traditional method. This indicates that the method of the present invention performs excellently in reducing misdiagnosis and missed diagnosis, and has higher stability and reliability. A low error rate is particularly important for medical image analysis because any misdiagnosis or missed diagnosis will have a significant impact on the health of patients, and the method of the present invention can effectively avoid this problem.
[0207] In addition, the method of the present invention also performs excellently in two indicators of detection sensitivity and detection specificity. The sensitivity is 92.3% and the specificity is 96.1%, both of which are higher than 80.2% and 87.5% of the traditional method. This shows that the method of the present invention can not only efficiently identify diseased samples, but also accurately exclude non-diseased samples. High sensitivity means that potential lesions can be detected as early as possible, while high specificity can reduce false alarms and ensure accurate diagnostic results.
[0208] In terms of data processing efficiency, the method of the present invention can process 144 images per minute, which is much higher than 69 of the traditional method. This advantage enables the method of the present invention to have an obvious speed advantage in processing large-scale image data, especially in an efficient medical system, it can provide faster diagnostic support.
[0209] In summary, the method of the present invention performs excellently in terms of accuracy, diagnostic time, error rate, detection ability and data processing efficiency, and is significantly superior to the traditional method. Through these advantages, the present invention not only improves the diagnostic accuracy of lung CT images, but also enhances the diagnostic efficiency and stability, greatly promoting the intelligent process of image diagnosis, and has broad application prospects and practical value.
[0210] As described above, it is only the preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and should be covered within the protection scope of the present invention.
Claims
1. A post-structured decomposition method based on image deep learning, characterized in that: The steps include: S1. Collecting original image data through an image acquisition device, and preprocessing the original image data to generate preprocessed image data; S2, based on the preprocessed image data, using the dual-path residual convolutional neural network and adaptive attention mechanism, extracting low-level and high-level local features simultaneously through two parallel convolution branches; S3, linear mapping operations are performed on the extracted low-level and high-level local features to generate query, key and value matrices, and the attention score is calculated by introducing dynamic adjustment items using the improved attention mechanism, and the value matrix is weighted and summed to generate global features; S4, inputting local features and global features into the feature fusion network for feature fusion, and using an adaptive weighting strategy to generate fused high-dimensional unstructured features; S5, inputting the high-dimensional unstructured feature representation into a post-structured decomposition module, performing overall analysis and decomposition on the high-dimensional unstructured feature representation, and forming feature subsets that describe information at each level in the image data; S6. Convert the feature subset into a standardized structured data format.
2. The post-structured decomposition method based on image deep learning according to claim 1, characterized in that: The original image data is a digital image containing multiple color channels.
3. The post-structured decomposition method based on image deep learning according to claim 1, characterized in that: The preprocessing includes performing noise filtering, image enhancement, resolution correction and color normalization on the original image data.
4. The post-structured decomposition method based on image deep learning according to claim 1, characterized in that: The S2 specifically includes: S21, inputting the preprocessed image data into a dual-path residual convolutional neural network, wherein the dual-path residual convolutional neural network comprises a first convolution branch and a second convolution branch; S22, using the first convolution branch to extract local features of the preprocessed image data: Among them, F local1 (x, y) represents the local feature of the output of the first convolution branch, ReLU represents the rectification activation function, k represents the radius of the convolution kernel, W1(i, j) represents the weight of the first convolution kernel at position (i, j), b1 represents the corresponding bias parameter, I preproc (x+i, y+j) represents the input preprocessed image data, and (x, y) represents the pixel coordinates; S23, using the second convolution branch to extract local features from the preprocessed image data: Among them, F local2 (x, y) represents the local feature of the second convolution branch output, ReLU represents the rectification activation function, k represents the convolution kernel radius, W2(i, j) represents the weight of the second convolution kernel at position (i, j), and b2 represents the corresponding bias parameter; S24, local features F output by the first and second convolution branches local1 (x,y) and F local2 (x, y) is weighted fused to obtain the fused local features: F local (x,y)=β·F local1 (x,y)+(1-β)·F local2 (x,y); Among them, F local (x, y) represents the local features after fusion, and β represents the weighting coefficient; S25, the fused local feature F local (x,y) applies an adaptive attention mechanism and calculates the attention coefficient of color channel c: Where A(c) represents the attention coefficient of color channel c, σ represents the Sigmoid activation function, M and N represent the number of pixels in the vertical and horizontal directions of the image, respectively, and c represents the color channel index; S26. Use the attention coefficient A(c) to fusion local feature F local (x, y, c) is weighted to obtain the local feature representation: F′ local (x,y,c)=A(c)·F local (x,y,c); Among them, F′ local (x, y, c) represents the local feature representation, F local (x, y, c) represents the fusion of local features.
5. The post-structured decomposition method based on image deep learning according to claim 1, characterized in that: The S3 specifically includes: S31, receiving the local feature representation F' output from step S31 local , the local feature representation F′ local represents the local feature matrix after being processed by the dual-path convolutional neural network and the adaptive attention mechanism, and the dimension of the local feature matrix is H×W×C, where H and W in the dimension represent the height and width of the feature map respectively, and C represents the number of feature channels; S32, local feature representation F′ local Perform a linear mapping operation to generate the query matrix, key matrix, and value matrix: Q=F′ local ·W Q ,K=F′ local ·W K ,V=F′ local ·W V ; Among them, Q represents the query matrix, K represents the key matrix, V represents the value matrix, and W Q represents the weight matrix of the query matrix, W K represents the weight matrix of the bond matrix, W V The weight matrix representing the value matrix; S33, local feature representation F' local Perform dynamic convolution operation to generate dynamic adjustment terms, by adjusting F′ local After convolution operation with dynamic convolution kernel weights and adding bias, nonlinear activation function is used for mapping: Among them, Δ represents the dynamic adjustment term, represents a nonlinear activation function, Conv represents a two-dimensional convolution operation, and W d represents the dynamic convolution kernel weight, b d represents the bias parameter; S34. Calculate the attention score matrix: Among them, A represents the attention score matrix, softmax represents row normalization, Q represents the query matrix, K represents the key matrix, and T represents the matrix transpose operation. represents the scaling factor to ensure numerical stability; S35. Use the attention score matrix to perform weighted summation on the value matrix to generate a global feature representation: F global =A·V; Among them, F global Represents the global feature representation.
6. The post-structured decomposition method based on image deep learning according to claim 1, characterized in that: The feature fusion network uses an adaptive weighting strategy to generate fusion features: F fusion =αF local +(1-α)F global ; Among them, F fusion represents the fused features, F local represents the local features, F global represents the global feature, and α represents the adaptive weight: α=σ(W α [F local ,F global ]+b α ); Among them, σ represents the Sigmoid activation function, which is used to ensure that the value range of the adaptive weight α is between 0 and 1, W α represents the fusion weight parameter matrix, b α Represents the bias parameter.
7. The post-structured decomposition method based on image deep learning according to claim 1, characterized in that: The S5 specifically includes: S51, the fused high-dimensional unstructured feature representation F fusion Any two eigenvectors F fusion (i) and F fusion (j) Calculate cosine similarity: Among them, S(i,j) represents the cosine similarity of any two feature vectors, F fusion (i) and F fusion (j) represents any two feature vectors in a high-dimensional unstructured feature table, and ∥·∥ represents the Euclidean norm; S52, constructing an adjacency matrix according to a preset similarity threshold: Where A(i,j) represents the adjacency matrix, and θ represents the preset similarity threshold; S53, analyze the adjacency matrix A(i,j) to obtain a preliminary grouping set {G1,G2,…,G K }, each group G in the preliminary group set K Include all feature vectors whose cosine similarity is not less than θ; S54, calculating the centroid vector for each group in the preliminary grouping set: Among them, μ k represents the centroid vector, |G k | indicates group G K The number of eigenvectors contained in F fusion (i) Feature vector representing high-dimensional unstructured feature representation; S55. Calculate the cosine similarity between any two groups for the centroid vectors of all groups: Among them, S(G p ,G q ) represents any two groups G p With G q The cosine similarity of p and μ q Respectively represent group G p and G q The centroid vector of p ∥ indicates group G p The Euclidean norm of ; S56, setting the initial level threshold θ (1) =θ, and at level l, for any two groups G p With G q , if it satisfies: S(G p ,G q )≥θ (l) ; Then group G p With G q Merge into new group G r =G p ∪G q , and update the centroid vector: Among them, μ r Represents a new group G r The centroid vector of |G r | indicates group G r The number of eigenvectors contained in ; S57, update the threshold, repeat the merging process of step S55 and step S56 until there is no group that meets the merging condition, and form an intermediate clustering result of iterative hierarchical clustering: i (l+1) =θ (l) -d; Among them, θ (l+1) represents the updated threshold, δ>0 represents a predetermined constant, F (l) represents the intermediate clustering result of iterative hierarchical clustering, Indicates the Kth cluster formed in the lth level clustering process (l) cluster groups, K (l) represents the total number of cluster groups formed during the l-th level clustering process; S58, for each level of intermediate clustering results F (l) Integrate, unify and merge the groups with intersection or inheritance relationship in different clustering levels, and define the final feature subset set, each feature subset F in the final feature subset set k satisfy: Among them, L k Represents the index set of all clustering levels that contain the same or compatible feature information, and ensures that any feature vector only appears in a unique feature subset, F k represents the feature subset in the final feature subset set, represents the kth cluster group formed in the lth level clustering process, Indicates that all the index sets L k Clustering of Merge to get the final feature subset F k , ∪ represents the union operation of sets; S59: Verify the integrated feature subset set and output the final feature subset set For any F k Any two eigenvectors F fusion (i) and F fusion (j) with: Among them, |F k | represents the feature subset F k The number of eigenvectors contained in θ min represents the predetermined consistency threshold, K final Represents the total number of subsets in the final feature subset set obtained after integrating the intermediate clustering results.
8. The post-structured decomposition method based on image deep learning according to claim 1, characterized in that: The S6 specifically includes: S61, for each feature subset F in the final feature subset set k Perform statistical aggregation processing to generate a summary feature vector: Among them, V k represents the summary feature vector, |F k | represents the feature subset F k The number of feature vectors contained in, f represents the feature subset F k Any eigenvector in ; S62, summarize the feature vector V k Perform linear normalization to generate normalized feature vectors: Among them, R k represents the normalized eigenvector, μ and σ represent the mean vector and standard deviation vector respectively: Among them, K final Represents the total number of subsets in the final feature subset set obtained after integrating the intermediate clustering results; S63, the normalized feature vector R k Structured as a fixed format data record D k , the data record D k Defined as field {ID k ,R k }, the field {ID k ,R k } k Represents the feature subset F k The unique identifier, R k Indicates the predetermined fields {F1, F2, ..., F d }; S64. Record all data D k Arrange in a fixed order to form a structured data set: D={D k ∣k=1,2,…,K final }; Where D represents a structured data set, which is stored in a table format, with each row corresponding to a data record D k , each column corresponds to a predetermined field.