Image encoder learning method and device based on topology enhancement, equipment and medium
By quantifying the topological complexity index and transforming the topological constraints, the problem of topological feature destruction in the learning of self-supervised image encoders is solved, and the image encoder achieves stable preservation of topological features and improved enhancement effect.
Patent Information
- Application Number
- CN202511059906.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-10-31
AI Technical Summary
Existing self-supervised image encoder learning methods lack explicit modeling of the underlying structure of images, resulting in the destruction of key topological features during the enhancement process.
By quantifying topological complexity, the underlying topological structure is transformed into a computable metric, providing a clear basis for topological protection in subsequent enhancement operations. This includes multi-threshold topological feature tracking, definition of topological complexity metrics, sampling parameters for topological constraint transformation, topological preservation enhancement, refined feature extraction, and dual loss analysis.
It improves the image encoder's ability to preserve topological features and the stability of the enhancement effect, ensuring that key topological structures are not destroyed and improving the quality of model output.
Smart Images

Figure CN120876629A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of machine learning technology, and in particular to an image encoder optimization method, apparatus, device, and medium based on topology enhancement. Background Technology
[0002] Self-supervised image encoder learning is an image feature learning method that does not require manually labeled data. Its core is to design "self-supervised tasks" (such as image reconstruction, contrastive learning, and topological alignment) to allow the encoder to automatically learn effective features from massive amounts of unlabeled images. In recent years, self-supervised image encoder learning has also found important applications in various fields. For example, in the healthcare field, encoders learn the texture and structural topological features of normal / abnormal tissues. After pre-training on unlabeled images, only a small amount of labeled data is needed for fine-tuning, which can assist in the detection of early lesions. In the fintech field, encoders are pre-trained on massive amounts of unlabeled scanned images of invoices and financial statements to learn features such as text layout, seal position, and digital structure. Key information can be extracted without manual annotation for automated reconciliation and anti-counterfeiting purposes.
[0003] Currently, traditional self-supervised image encoder learning methods often rely on heuristic data augmentation strategies such as random geometric transformations (e.g., cropping, rotation) or color perturbations. However, this approach lacks explicit modeling of the underlying structure of the image, which may lead to the destruction of key topological features during the augmentation process. Summary of the Invention
[0004] This invention provides an image encoder optimization method, apparatus, device, and medium based on topology enhancement. By quantifying topology complexity, the underlying topology structure is transformed into a computable index, providing a clear basis for topology protection in subsequent enhancement operations.
[0005] Firstly, a topology-enhanced image encoder optimization method is provided, including: The original image to be processed is obtained, and multi-threshold topological feature tracking is performed on the original image to obtain a continuous homology map; The topology complexity index is determined based on the continuous homology graph, and the sampling parameters of the preset topology constraint transformation are defined through the topology complexity index. The original image is topologically preserved and enhanced according to the sampling parameters to obtain an enhanced view, and the enhanced view is refined and feature extracted to obtain enhanced features; Perform dual loss analysis on the enhanced features, and then weight and fuse the analyzed losses to generate a comprehensive loss value; The parameters of the preset image encoder are adjusted based on the comprehensive loss value to obtain the adjusted image encoder.
[0006] Secondly, a topology-enhanced image encoder optimization device is provided, comprising: The acquisition and tracking module is used to acquire the original image to be processed, perform multi-threshold topological feature tracking on the original image, and obtain a continuous homology map; The definition module is used to determine the topology complexity index based on the continuous homology graph, and to define the sampling parameters of the preset topology constraint transformation through the topology complexity index; An enhancement module is used to perform topology-preserving enhancement on the original image based on the sampling parameters to obtain an enhanced view; The extraction module is used to extract refined features from the enhanced view to obtain enhanced features; The analysis and fusion module is used to perform dual loss analysis on the enhanced features and weightedly fuse the analyzed losses to generate a comprehensive loss value. The adjustment module is used to adjust the parameters of the preset image encoder according to the comprehensive loss value to obtain the adjusted image encoder.
[0007] Thirdly, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described topology-enhanced image encoder optimization method.
[0008] Fourthly, a computer-readable storage medium is provided, which stores a computer program that, when executed by a processor, implements the steps of the above-described topology-enhanced image encoder optimization method.
[0009] In the aforementioned scheme implemented by the image encoder optimization method, apparatus, computer equipment, and storage medium based on topology enhancement, the original image is acquired and multi-threshold topological feature tracking is performed. This captures the topological structure of the original image at different scales. The generated continuous homology map can quantify the evolution process of topological features, providing a comprehensive and objective visual data foundation for subsequent analysis of the image's topological complexity, avoiding feature omissions caused by a single threshold. The topological complexity index quantifies the complexity of the image's topological structure, thereby defining the sampling parameters for topological constraint transformations. This ensures that subsequent enhancement operations focus on key topological features, avoiding indiscriminate transformations that damage the core structure of the image, and improving parameter settings. The enhancement is targeted; topology-preserving enhancement enhances image details while ensuring that the core topological structure is not destroyed through sampling parameter constraints; refined feature extraction focuses on key effective features after enhancement, reducing redundant information and providing high-quality input for subsequent loss analysis; dual loss analysis evaluates the enhancement effect from multiple dimensions, and the weighted fusion of the comprehensive loss value can balance the importance of different losses and comprehensively reflect the deviation in the enhancement process; the comprehensive loss value is used as a feedback signal to adjust the encoder parameters, which can drive the encoder to learn a mapping relationship that better conforms to the topological constraints and feature enhancement needs, improve the encoder's ability to preserve image topological features and the stability of the enhancement effect, and ultimately optimize the output quality of the model. Attached Figure Description
[0010] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 This is a schematic diagram of an application environment for an image encoder optimization method based on topology enhancement according to an embodiment of the present invention; Figure 2 This is a flowchart illustrating an image encoder optimization method based on topology enhancement according to an embodiment of the present invention; Figure 3 This is a schematic diagram of a topology-enhanced image encoder optimization device according to an embodiment of the present invention; Figure 4 This is a schematic diagram of the structure of a computer device according to one embodiment of the present invention; Figure 5 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation
[0012] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0013] This invention provides an image encoder optimization method based on topology enhancement, which can be applied to applications such as... Figure 1 In this application environment, the client communicates with the server via a network. The server can acquire the original image to be processed, perform multi-threshold topological feature tracking on the original image to obtain a continuous homology map; determine the topological complexity index based on the continuous homology map, and define sampling parameters for a preset topological constraint transformation through the topological complexity index; perform topological preservation enhancement on the original image based on the sampling parameters to obtain an enhanced view, and extract refined features from the enhanced view to obtain enhanced features; perform dual loss analysis on the enhanced features, and weight and fuse the analyzed losses to generate a comprehensive loss value; adjust the parameters of a preset image encoder based on the comprehensive loss value to obtain an adjusted image encoder, and feed the adjusted image encoder back to the client. This invention provides an image encoder optimization device based on topological enhancement. For the adjusted image encoder service, by quantifying the topological complexity, the underlying topological structure is transformed into a computable index, providing a clear topological protection basis for subsequent enhancement operations. The client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster composed of multiple servers. The invention will be described in detail below through specific embodiments.
[0014] Please see Figure 2 As shown, Figure 2 A flowchart illustrating an image encoder optimization method based on topology enhancement, provided in this embodiment of the invention, includes the following steps: S1. Obtain the original image to be processed, and perform multi-threshold topological feature tracking on the original image to obtain a continuous homology map.
[0015] In this embodiment of the invention, the acquisition refers to reading or collecting raw image data from a data source, and the multi-threshold topological feature tracking refers to extracting the topological features of the image under different thresholds and tracking the evolution of these features as the threshold changes.
[0016] Specifically, the original image to be analyzed is first acquired. Then, by adjusting multiple threshold parameters, the "birth-existence-death" process of topological features (such as connected regions, holes, structural correlations, etc.) in the image at different scales is dynamically tracked. Finally, these evolutionary patterns are presented in the form of a continuous homology map, which can quantitatively reflect the stability and complexity of the core topological structure in the image and provide structured topological guidance information for subsequent analysis.
[0017] In detail, for the input image First, construct a cubic complex of its grayscale values, expressed by the following formula:
[0018] Where ∈ is an increasing threshold parameter, K ∈ This represents the set of pixels whose grayscale values do not exceed ∈. As ∈ increases from 0 to 1, the generation and disappearance of topological features such as connected components (β0) and holes (β1) are recorded, forming a persistent homology map, which can be expressed by the following formula: Here b i ,d i These represent the birth and death thresholds for the i-th topological feature, respectively. For example, in a CT image of a lung nodule, the persistence point of the nodule region corresponds to (b 结节 ,d 结节 Its durability d 结节 -b 结节 The significance of the nodules was quantified.
[0019] In specific healthcare scenarios, by tracking topological features such as connectivity of lesion areas and evolution of cavity structures inside tumors through multi-threshold tracking, continuous homology maps can accurately capture microstructural changes in lesions and assist doctors in determining the type of lesion.
[0020] In fintech scenarios, by tracking the connectivity of transaction nodes under different thresholds and the topological evolution of risk propagation paths, continuous coherence graphs can reveal the potential interconnected structures of financial networks, identify key nodes in risk transmission, provide deep data support based on topology for risk warning and anti-fraud analysis, and enhance the stability monitoring capabilities of financial systems.
[0021] In this embodiment of the invention, the step of performing multi-threshold topological feature tracking on the original image to obtain a continuous homology map includes: Construct the spatial connection relationship between pixels in the original image, and transform the original image into a cubic complex based on the spatial connection relationship; The topological features of the cubic complex are dynamically tracked using preset multi-threshold parameters to obtain the tracking results; Based on the tracking results, the topological feature evolution law of the original image is extracted, and the continuous homology map is determined through the topological feature evolution law.
[0022] In this embodiment of the invention, the construction refers to defining the neighborhood relationship between pixels in the image and establishing topological connection rules between pixels; the transformation refers to mapping the pixels / voxels of the image and their connection relationships into a discrete geometric structure; the dynamic tracking of topological features refers to gradually constructing a cubic complex and tracking the evolution process of topological features under multiple thresholds; the extraction refers to the process of extracting the life cycle information of topological features from the dynamic tracking data; and the determination refers to the process of visualizing the evolution law of topological features and generating a continuous homology map.
[0023] Specifically, firstly, each pixel of the original image is traversed, and "connection edges" between pixels are defined according to the spatial adjacency relationship between pixels (such as 8-neighborhood or 4-neighborhood connection rules). Adjacent pixels are regarded as "association units" in a topological sense. Then, based on these connection relationships, the two-dimensional image is transformed into a cubic complex. Here, a cubic complex refers to an abstract mathematical model composed of multiple levels of topological structures such as pixels, edges formed by adjacent pixels, and faces formed by adjacent edges, realizing the transformation of the image from a pixel grid to a topological structure.
[0024] Furthermore, for the constructed cubic complex, a preset multi-threshold parameter is introduced, such as a threshold sequence based on pixel grayscale value, gradient value, or texture features, to dynamically monitor changes in topological features under different threshold conditions: when the threshold increases, the "death" of low grayscale regions is tracked (such as the disappearance of small connected regions); when the threshold decreases, the "birth" of high grayscale regions is tracked (such as the merging of large connected regions), and the occurrence time (threshold), duration, and evolution path of each topological feature (such as connected components and holes) are recorded to form a complete tracking result.
[0025] Furthermore, based on the dynamic tracking results of topological features, core evolutionary patterns are extracted: including the appearance and disappearance thresholds of different topological features, structural stability within the survival interval, such as whether they merge or split, and topological mutation events at key threshold points, such as the formation or disappearance of holes; then these patterns are mapped onto a persistent homology graph—with the horizontal axis as the threshold parameter and the vertical axis as the topological feature dimension, such as 0-dimensional connected components and 1-dimensional holes, and the survival interval of each topological feature is marked with points or line segments, ultimately forming a visualized persistent homology graph.
[0026] In this embodiment of the invention, discrete pixel information is transformed into a structured set of topological units, providing a mathematical basis for subsequent topological analysis and overcoming the locality limitations of traditional pixel-level analysis; a multi-threshold strategy covers topological features at different scales, avoiding feature omissions caused by a single threshold and ensuring the capture of subtle topological changes; the extraction of evolutionary laws transforms complex topological changes into quantifiable rules, providing a logical framework for deep image structure analysis.
[0027] In this embodiment of the invention, the limitations of traditional image analysis relying on pixel-level features are overcome. By topological feature tracking and continuous homology graph quantization, the essential features of the image can be captured from the structural level. Moreover, the multi-threshold strategy ensures the integrity of topological information at different scales, providing a reliable structural basis for subsequent topological constraint-based tasks and improving the robustness and deep insight capabilities of image analysis.
[0028] S2. Determine the topology complexity index based on the continuous homology graph, and define the sampling parameters of the preset topology constraint transformation through the topology complexity index.
[0029] In this embodiment of the invention, "determining" refers to extracting numerical indicators of the quantized topological structure from the continuous homology graph, and "defining" refers to setting parameters for constraint transformation based on the topological complexity index to ensure that the transformed data retains key topological features.
[0030] Specifically, based on the continuous homology graph, indicators of the quantifiable topological complexity, such as the number of connected components and the complexity of holes, are first extracted from it. Then, based on the numerical characteristics of these indicators, such as high complexity or specific topological association types, sampling parameters for topological constraint transformations are defined in a targeted manner. These are the rules that control the magnitude and range of image enhancement operations, ensuring that subsequent enhancement operations retain the core topological structure while meeting the preset topological similarity requirements.
[0031] In detail, define the preset topological constraint transformation. The sampling parameters are expressed by the following formula:
[0032] in The normalization constant is Controlling topology fidelity strength To address the bottleneck distance between continuous homology graphs, This indicates a continuous coherence plot. Indicates the enhanced image The continuous coherence plot, This represents the probability of choosing enhancement transformation T based on the original image I; The bottleneck distance is expressed by the following formula:
[0033] in, This refers to the optimal matching method. This represents a point in a continuous homology graph. The formula can be understood as: finding an optimal matching method. ,Bundle The dots Mapped to The dots To minimize the maximum point-to-point distance.
[0034] In healthcare settings, continuous coherence maps are used to determine the topological complexity of lesion regions, such as the number of cavities within a tumor and the connectivity complexity of the vascular network. If the indicators show a complex lesion topology (such as multiple micropores), strict sampling parameters are defined to limit the enhancement amplitude to avoid damaging the subtle topological features of the lesion. If the indicators show a simple connected organ outline, the parameters can be relaxed to allow for a greater enhancement amplitude, which improves image clarity while ensuring that the key diagnostic topology remains unchanged, helping doctors to more accurately identify lesion features.
[0035] In financial scenarios, topological complexity indicators are determined based on the continuous coherence graph of the transaction network or risk distribution map, such as the connectivity density of abnormal transaction clusters and the branch complexity of risk transmission paths. If the indicators show a high-complexity risk association network, the sampling parameters are used to limit the magnitude of topological transformation to prevent distortion of the association relationship of key risk nodes. If the indicators show a structured multi-regional transaction network, flexible sampling parameters are defined to allow for appropriate structural adjustments while retaining the topological characteristics of the core transaction links, providing more reliable augmented data for risk modeling and fraud detection.
[0036] In this embodiment of the invention, determining the topological complexity index based on the persistent homology graph includes: Identify the core topological feature trajectories in the persistent homology graph; Extract the Betti number sequence from the original image based on the core topological feature trajectory; The topological complexity index of the original image is determined based on the Betti number sequence.
[0037] In this embodiment of the invention, the identification refers to selecting the most representative features of the data topology (such as connected components, holes, voids, etc.) from the continuous homology map; the Betti number sequence extraction refers to converting the dynamic changes of the core feature trajectory into a Betti number sequence that changes with the filtering parameters; and the determination refers to quantifying the topological complexity of the original image based on the dynamic changes of the Betti number sequence.
[0038] Specifically, the study focuses on identifying the "birth-existence-death" trajectories of various topological structures (such as connected components, holes, and high-dimensional topological features) in persistent homology graphs. It clarifies the existence status of core topological features at different scales, such as which topological structures are stable in the long term and which are minor features that appear only briefly. Based on the obtained topological feature trajectories, the study further quantifies and extracts the Betti number sequence of the original image. This sequence numerically represents the core topological attributes of the image in different dimensions (such as zero-dimensional Betti number reflecting the number of connected components, one-dimensional Betti number reflecting the number of holes, etc.). This transforms the abstract topological evolution information in persistent homology graphs into measurable quantitative indicators.
[0039] Furthermore, by combining the magnitude, variation, and stability of the values in the sequence, the topological complexity of the image is comprehensively determined: for example, if the high-dimensional Betti numbers in the sequence (such as the values corresponding to holes) are large and fluctuate frequently, it indicates the existence of complex hole structures, corresponding to high topological complexity; if the zero-dimensional Betti numbers (such as the number of connected components) are numerous but the associations are stable, it indicates that multi-connected regions are the main feature, corresponding to a specific type of topological complexity. Finally, these determination results are integrated into a topological complexity index that reflects the complexity of the deep topological structure of the image.
[0040] In this embodiment of the invention, defining the sampling parameters for the preset topological constraint transformation through the topological complexity index includes: The topological complexity index is used to analyze the topological structure type of the original image; The basic constraints are set for the topology type to obtain the basic constraints. The parameter range of the preset topology constraint transformation is adjusted according to the preset Gibbs distribution weighting strategy and the basic constraints. The sampling parameters are determined based on the parameter range and the preset topological similarity threshold.
[0041] In this embodiment of the invention, the topology type analysis refers to classifying the topology of the original image into predefined types based on topology complexity indicators, such as the volatility of the Betti number and the proportion of high-dimensional Betti numbers. The basic constraint setting refers to defining a set of hard or soft constraints according to the topology type to limit the feasible space of subsequent topology transformations. The adjustment refers to dynamically scaling the feasible range of topology transformation parameters under the basic constraints, combining the probability weighting mechanism of the Gibbs distribution, to balance exploration and development. The determination refers to selecting sampling points that meet the topology similarity requirements within the adjusted parameter range.
[0042] Specifically, using the topological complexity index as input, and combining the characteristics of the Betti number sequence in the index (such as numerical size and variation pattern), the topological structure of the original image is classified into different types. For example, if the index shows that the high-dimensional Betti number (corresponding to holes) is large and fluctuates frequently, it is judged as "complex hole structure type"; if the zero-dimensional Betti number (corresponding to connected components) is large but the correlation is stable, it is judged as "multi-connected region type".
[0043] Furthermore, corresponding topology preservation rules are formulated for each type: for the "complex void structure type", the core constraint is clearly defined as "strictly limiting the range of operations that may damage the integrity of the void"; for the "multi-connected region type", the core constraint is clearly defined as "allowing moderate spatial deformation but maintaining the connectivity between regions".
[0044] Furthermore, a Gibbs distribution weighting strategy is adopted to quantify and adjust the parameter range of the preset topological constraint transformation. For "complex hollow structure type", the value range of enhancement amplitude parameters is narrowed by weight allocation (such as the selectable range of rotation angle and tensile strength). For "multi-connected region type", the restrictions on spatial deformation parameters are relaxed by weight tilting, but the constraints on connectivity-related parameters are strengthened. Finally, the preset topological similarity threshold is used for verification to simulate the enhancement effect under different parameter combinations, ensuring that the transformation operation corresponding to the parameters can make the topological similarity between the enhanced image and the original image meet the threshold requirements, and finally determining the sampling parameters that meet the topological constraints.
[0045] Furthermore, a weighted sampling strategy can also be used, as shown below: Calculate the Betti number sequence of the original image. ,in, Indicates connected components, Indicates emptiness; like (The presence of complex void structures) limits the magnitude of enhancement. ;like (No complex void structure exists), no limit on the magnitude of enhancement. .
[0046] like (Multi-connected region), allowing for large spatial deformation while maintaining... ,like Large spatial deformation is not allowed and it does not maintain .
[0047] In this embodiment of the invention, topological complexity index is used to analyze topological structure types. Based on quantitative indicators, image topological types are accurately classified, avoiding bias from subjective human judgment and providing an objective basis for subsequent constraint setting. This ensures differentiated processing for topological structures of different complexity. Basic constraints are set for topological structure types, and basic rules are formulated for specific topological types to ensure the safety of bottom-level constraint transformations, prevent enhancement operations from damaging core topological features, and guarantee structural integrity. Through a probability weighting strategy based on Gibbs distribution, the parameter range is dynamically adjusted within the basic constraint framework, making parameter adjustments more consistent with the statistical laws of topological features, balancing enhancement diversity and structural stability, and reducing ineffective or risky transformations.
[0048] In this embodiment of the invention, a bridge is built between the continuous homology graph and the sampling parameters by using the topology complexity index, thereby realizing the accurate conversion of topology analysis results into actual operation rules.
[0049] S3. Perform topological preservation enhancement on the original image according to the sampling parameters to obtain an enhanced view, and perform refined feature extraction on the enhanced view to obtain enhanced features.
[0050] In this embodiment of the invention, the topology-preserving enhancement refers to enhancing the original image through geometric transformation or noise injection, while strictly constraining the enhanced image to maintain the same topological structure as the original image. The refined feature extraction refers to filtering out features that are strongly correlated with the topological structure through topological constraints or attention mechanisms, while suppressing irrelevant or redundant information.
[0051] Specifically, based on the determined sampling parameters, a topology-preserving enhancement operation is applied to the original image to generate an enhanced view, allowing the image to obtain richer features while maintaining its basic topological structure. Then, the enhanced view is further extracted and refined to obtain enhanced features that can more accurately reflect the key information of the image. This is a coherent process from image preprocessing to feature mining.
[0052] In specific healthcare scenarios, original images such as CT and MRI images can be enhanced with topology preservation to highlight subtle lesion features without damaging the topological structure of lesions such as tumors and vascular malformations, making them easier for doctors to observe.
[0053] In financial scenarios, the original image can be compared to a transaction network graph or a risk heat map. Topology preservation enhancement can maintain the topological structure of the financial network, such as the connection relationship between transaction nodes and the risk transmission path, and avoid the loss of key related information. Refined feature extraction can uncover the characteristics of transaction clusters and key risk nodes, which can be used for anti-fraud and risk warning, and help financial risk management and decision-making.
[0054] In this embodiment of the invention, the step of performing topology-preserving enhancement on the original image based on the sampling parameters to obtain an enhanced view includes: Candidate transformations that meet the topological constraints are selected based on the sampling parameters; The original image is enhanced based on the candidate transformation to obtain a preliminary enhanced view; The enhanced view is obtained by comparing and verifying the topological features of the initial enhanced view with those of the original image.
[0055] In this embodiment of the invention, the filtering refers to selecting a set of transformations that meet the conditions from all possible image transformations according to preset topological constraints. The enhancement process refers to applying the selected candidate transformations to the original image to generate a set of topologically consistent enhanced views. The comparison verification refers to comparing the preliminary enhanced views with the topological features of the original image and selecting the enhanced view with the highest topological consistency as the final output.
[0056] Specifically, the enhancement types included in the sampling parameters are clearly defined, such as rotation, clipping, deformation, etc., and their corresponding operation ranges, such as rotation angle ranges and clipping ratio limits. Candidate transformations that meet the topological constraints are screened from the parameterized enhancement distribution. Combined with the Gibbs distribution weighting strategy, transformation operations that have less impact on the core topological structure (such as connectivity and holes) are given priority, while transformation types that may destroy the preset topological similarity threshold are excluded.
[0057] Furthermore, for images with complex void structures, slight rotation or local cropping is performed according to the limited amplitude; for images with multiple connected regions, moderate spatial deformation is performed within the allowable range, but the connectivity between regions is ensured to remain unchanged through real-time topology verification, ultimately resulting in an enhanced view with a topology consistent with the original image.
[0058] Furthermore, by comparing the topological features of the preliminary enhanced view with the original image, it is verified whether the enhanced image meets the preset topological similarity threshold; if it does not meet the threshold, the process returns to the second step to re-select candidate transformations; if it meets the threshold, the enhanced view is determined as the final output.
[0059] In this embodiment of the invention, the step of refining feature extraction from the enhanced view to obtain enhanced features includes: Deep feature extraction is performed on the enhanced view to obtain a high-dimensional original feature vector; The high-dimensional original feature vector is optimized by performing feature dimension representation to obtain intermediate features; The correlation between the intermediate features and the original image is enhanced by using a preset topological alignment mechanism to obtain enhanced features.
[0060] In this embodiment of the invention, the deep feature extraction refers to extracting features layer by layer from the enhanced view using a pre-trained deep learning model; the feature dimension expression optimization refers to optimizing the high-dimensional original feature vector through dimensionality reduction or feature compression techniques to obtain low-dimensional intermediate features; and the original relevant feature enhancement refers to enhancing the intermediate features according to a preset topological alignment mechanism to make them closer to the topological structure of the original image, while suppressing irrelevant features.
[0061] Specifically, the enhanced view is fed into a preset image encoder. The encoder performs deep feature extraction on the enhanced view through multi-layer convolution, pooling and other operations to capture visual semantic information in the image, such as texture, contour and structural features preserved by topological constraints, such as connected region morphology and key topological relationships. The high-dimensional original feature vector is fed into the projection head network for processing. The projection head adjusts the dimensionality and refines the semantics of the original features through linear transformation, non-linear activation and other operations, filters redundant information and enhances the discriminativeness of core features, making the features more suitable for subsequent loss calculation needs, and outputs the initially optimized intermediate features.
[0062] Furthermore, a topological alignment mechanism is introduced to further refine the features. By comparing the topological features of the enhanced view with those of the original image, such as topological labels based on persistent homology graphs, the feature components related to the core topological structure are strengthened in the feature space, and non-topological noise that may be introduced by the enhancement is suppressed, ultimately resulting in enhanced features that have both semantic integrity and topological stability.
[0063] In this embodiment of the invention, transformation methods that may damage the image topology are pre-filtered by sampling parameters, reducing invalid enhancement operations from the source and ensuring that candidate transformations are within the "topologically safe" range, laying the foundation for the effectiveness of subsequent enhancements. Enhancement is performed based on the filtered safe candidate transformations, which can enrich the diversity of the image and improve the generalization ability of features, while avoiding structural distortion caused by improper transformations, initially achieving the effect of "enhancing without distortion". Through secondary verification of topological features, possible residual topological deviations in the initial enhancement are further eliminated.
[0064] In this embodiment of the invention, topology preservation enhancement ensures that the core structural information of the image is not destroyed, laying a reliable foundation for subsequent feature extraction; refined feature extraction further mines key information, making the features more representative and distinguishable. Whether it is medical diagnosis or financial analysis, it can improve the accuracy and effectiveness of task processing, making image- or graph-based analysis and decision-making more scientific and reasonable.
[0065] S4. Perform dual loss analysis on the enhanced features, and then weight and fuse the analyzed losses to generate a comprehensive loss value.
[0066] In this embodiment of the invention, the dual loss analysis refers to the process of simultaneously using two different loss functions or evaluation criteria to independently calculate the same set of enhanced features output by the model, so as to comprehensively measure the performance of the features in different dimensions, and finally generate a comprehensive loss value through weighted fusion.
[0067] Specifically, we first calculate the loss generated by the enhanced features during model training or task execution from two different dimensions, such as one focusing on feature semantic matching and the other on topological structure fit. Then, we assign different weights to these two losses according to task requirements and combine them into a comprehensive loss value to guide model optimization or result evaluation.
[0068] In this embodiment of the invention, the step of performing dual loss analysis on the enhanced features and weightedly fusing the analyzed losses to generate a comprehensive loss value includes: The enhanced features are subjected to feature contrast loss analysis to obtain feature contrast loss values; Perform topology consistency loss analysis on the enhanced features to obtain the topology consistency loss value; The feature comparison loss value and the topology consistency loss value are weighted and summed according to a preset weight allocation strategy to generate a comprehensive loss value.
[0069] In this embodiment of the invention, the feature contrast loss analysis refers to the process of forcing the model to learn discriminative feature representations by measuring the distance between sample pairs in the feature space. The topology consistency loss analysis refers to the process of ensuring that the feature representations learned by the model retain the topology of the original data by constraining the neighborhood relationships of samples in the feature space. The weighted summation refers to the process of obtaining the final comprehensive loss value by assigning different weights to the feature contrast loss and the topology consistency loss and combining the optimization objectives of the two.
[0070] Specifically, enhanced features are selected along with other reference features, such as different enhanced view features of the same image and negative sample image features. Through a contrastive learning framework, the similarity between enhanced features and positive sample features and the difference between enhanced features and negative sample features are calculated. This contradiction between "expected similarity and actual difference" is quantified into a loss value, namely the feature contrast loss value.
[0071] Furthermore, the topological structure corresponding to the enhanced features is extracted, such as the connectivity and holes reflected in the continuous homology map, and then compared with the original image or target topological structure to quantify the difference between the enhanced feature topology and the reference topology. For example, the bottleneck distance is used to calculate the structural deviation between the topological maps, and this difference is converted into a loss value. The more inconsistent the topology, the higher the loss.
[0072] Furthermore, the weights of feature contrast loss and topology consistency loss are predefined, such as by setting them according to the task's requirements for semantic discriminativity and topology preservation. , The weights are calculated by multiplying each of the two loss values by their respective weights and then summing the results to generate the overall loss value, expressed by the formula: Total loss value = Feature contrast loss + Topology consistency loss Among them, given enhanced view pairs Its projection characteristics The optimization objective is:
[0073] The first term is the standard contrast loss, the second term is the topology consistency penalty term, and the indicator function is... Ensure that the bottleneck distance between views is less than the threshold only. Apply constraints at the time, This is the balance coefficient.
[0074] In this embodiment of the invention, feature contrast loss analysis is performed to force features to learn the expression of "distinguishing between similar and dissimilar types", so that the enhanced features have strong semantic discriminative power and downstream tasks can make more accurate judgments using features; topological consistency loss analysis is performed on the enhanced features to avoid the enhancement operation from destroying key structures; the two optimization objectives are flexibly balanced with weights to adapt to the priority requirements of different businesses for semantic discrimination and topological preservation.
[0075] In this embodiment of the invention, the limitation of focusing only on one-sided optimization objectives with a single loss is overcome. By balancing multi-dimensional losses, the model or analysis process can take into account different key information, such as semantics and topology, thereby improving the comprehensiveness and accuracy of task results.
[0076] S5. Adjust the parameters of the preset image encoder according to the comprehensive loss value to obtain the adjusted image encoder.
[0077] In this embodiment of the invention, the parameter adjustment refers to the process of iteratively updating the network parameters of the encoder through an optimization algorithm to minimize the overall loss value.
[0078] Specifically, there is a preset image encoder with initial network parameters, such as convolutional layer weights and Transformer attention parameters, as well as a batch of labeled image data or image data used for contrastive learning, and a pre-calculated comprehensive loss value calculation logic. The original image is input into the preset image encoder, and after the operation of each layer of the encoder, the feature representation of the image is obtained. Then, based on the calculation of feature contrast loss and topological consistency loss, and then the fusion according to weights, the comprehensive loss value is calculated.
[0079] Furthermore, based on the comprehensive loss value, optimization algorithms such as stochastic gradient descent and Adam are used to calculate the gradient of the loss value with respect to the parameters of each layer of the encoder. Along the reverse direction of the gradient, the network parameters of the encoder are adjusted, such as fine-tuning the convolutional kernel weights and updating the attention mechanism parameters, so that the comprehensive loss value is reduced as much as possible, thus completing one parameter update. The process of calculating the loss through forward propagation and adjusting the parameters through backpropagation is repeated. After multiple iterations (traversing multiple batches of data), when the comprehensive loss value stabilizes at a low level or reaches the preset number of training rounds, the adjusted image encoder is obtained.
[0080] In this embodiment of the invention, by adjusting the parameters in conjunction with the comprehensive loss value, the features learned by the encoder are not only semantically distinguishable but also retain the image topological structure information, making the features output by the encoder more in line with business needs. At the same time, by continuously iterating and optimizing the parameters, the encoder can better extract effective features, which can improve the accuracy and robustness of the model in downstream tasks, and make the model more adaptable to complex scenarios.
[0081] As can be seen, in the above scheme, for the adjusted image encoder service, the original image to be processed is acquired, and multi-threshold topological feature tracking is performed on the original image to obtain a continuous homology map; a topological complexity index is determined based on the continuous homology map, and sampling parameters for preset topological constraint transformation are defined through the topological complexity index; topological preservation enhancement is performed on the original image based on the sampling parameters to obtain an enhanced view, and refined feature extraction is performed on the enhanced view to obtain enhanced features; dual loss analysis is performed on the enhanced features, and the analyzed losses are weighted and fused to generate a comprehensive loss value; the parameters of the preset image encoder are adjusted based on the comprehensive loss value to obtain the adjusted image encoder. By quantifying the topological complexity, the underlying topological structure is transformed into a computable index, providing a clear basis for topological protection for subsequent enhancement operations.
[0082] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0083] In one embodiment, a topology-enhanced image encoder optimization apparatus is provided, which corresponds one-to-one with the topology-enhanced image encoder optimization method described in the above embodiments. For example... Figure 3 As shown, the image encoder optimization device based on topology enhancement includes an acquisition and tracking module 101, a definition determination module 102, an enhancement module 103, an extraction module 104, an analysis and fusion module 105, and an adjustment module 106. Detailed descriptions of each functional module are as follows: The acquisition and tracking module 101 is used to acquire the original image to be processed, perform multi-threshold topological feature tracking on the original image, and obtain a continuous homology map. The definition module 102 is used to determine the topology complexity index based on the continuous homology graph, and to define the sampling parameters of the preset topology constraint transformation through the topology complexity index; Enhancement module 103 is used to perform topology-preserving enhancement on the original image according to the sampling parameters to obtain an enhanced view; Extraction module 104 is used to extract refined features from the enhanced view to obtain enhanced features; The analysis and fusion module 105 is used to perform dual loss analysis on the enhanced features and weightedly fuse the analyzed losses to generate a comprehensive loss value. The adjustment module 106 is used to adjust the parameters of the preset image encoder according to the comprehensive loss value to obtain the adjusted image encoder.
[0084] In one embodiment, the acquisition tracking module 101, when acquiring the original image to be processed and performing multi-threshold topological feature tracking on the original image to obtain a continuous homology map, is used for: Construct the spatial connection relationship between pixels in the original image, and transform the original image into a cubic complex based on the spatial connection relationship; The topological features of the cubic complex are dynamically tracked using preset multi-threshold parameters to obtain the tracking results; Based on the tracking results, the topological feature evolution law of the original image is extracted, and the continuous homology map is determined through the topological feature evolution law.
[0085] In one embodiment, when determining the topological complexity index based on the persistent homology graph, the definition module 102 is used to: Identify the core topological feature trajectories in the persistent homology graph; Extract the Betti number sequence from the original image based on the core topological feature trajectory; The topological complexity index of the original image is determined based on the Betti number sequence.
[0086] When defining the sampling parameters for a preset topological constraint transformation using the aforementioned topological complexity index, it is used for: The topological complexity index is used to analyze the topological structure type of the original image; The basic constraints are set for the topology type to obtain the basic constraints. The parameter range of the preset topology constraint transformation is adjusted according to the preset Gibbs distribution weighting strategy and the basic constraints. The sampling parameters are determined based on the parameter range and the preset topological similarity threshold.
[0087] In one embodiment, when the enhancement module 103 performs topology-preserving enhancement on the original image according to the sampling parameters to obtain an enhanced view, it is used to: Candidate transformations that meet the topological constraints are selected based on the sampling parameters; The original image is enhanced based on the candidate transformation to obtain a preliminary enhanced view; The enhanced view is obtained by comparing and verifying the topological features of the initial enhanced view with those of the original image.
[0088] In one embodiment, when the extraction module 104 extracts refined features from the enhanced view to obtain enhanced features, it is used to: Deep feature extraction is performed on the enhanced view to obtain a high-dimensional original feature vector; The high-dimensional original feature vector is optimized by performing feature dimension representation to obtain intermediate features; The correlation between the intermediate features and the original image is enhanced by using a preset topological alignment mechanism to obtain enhanced features.
[0089] In one embodiment, when the analysis and fusion module 105 performs dual loss analysis on the enhanced features and weights and fuses the analyzed losses to generate a comprehensive loss value, it is used to: The enhanced features are subjected to feature contrast loss analysis to obtain feature contrast loss values; Perform topology consistency loss analysis on the enhanced features to obtain the topology consistency loss value; The feature comparison loss value and the topology consistency loss value are weighted and summed according to a preset weight allocation strategy to generate a comprehensive loss value.
[0090] This invention provides an image encoder optimization device based on topology enhancement. For an adjusted image encoder service, the device acquires the original image to be processed, performs multi-threshold topological feature tracking on the original image to obtain a continuous homology map, determines a topological complexity index based on the continuous homology map, and defines sampling parameters for a preset topological constraint transformation using the topological complexity index, performs topological preservation enhancement on the original image based on the sampling parameters to obtain an enhanced view, and extracts refined features from the enhanced view to obtain enhanced features, performs dual loss analysis on the enhanced features, and weights and fuses the analyzed losses to generate a comprehensive loss value, and adjusts the preset image encoder parameters based on the comprehensive loss value to obtain an adjusted image encoder. By quantifying the topological complexity, the underlying topological structure is transformed into a computable index, providing a clear basis for topological protection in subsequent enhancement operations.
[0091] For specific limitations regarding the topology-enhanced image encoder optimization device, please refer to the limitations of the topology-enhanced image encoder optimization method described above, which will not be repeated here. Each module in the aforementioned topology-enhanced image encoder optimization device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0092] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When executed by the processor, the computer program implements the functions or steps of a topology-enhanced image encoder optimization method on the server side.
[0093] In one embodiment, a computer device is provided, which may be a client, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When executed by the processor, the computer program implements client-side functions or steps of a topology-enhanced image encoder optimization method.
[0094] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps: The original image to be processed is obtained, and multi-threshold topological feature tracking is performed on the original image to obtain a continuous homology map; The topology complexity index is determined based on the continuous homology graph, and the sampling parameters of the preset topology constraint transformation are defined through the topology complexity index. The original image is topologically preserved and enhanced according to the sampling parameters to obtain an enhanced view, and the enhanced view is refined and feature extracted to obtain enhanced features; Perform dual loss analysis on the enhanced features, and then weight and fuse the analyzed losses to generate a comprehensive loss value; The parameters of the preset image encoder are adjusted based on the comprehensive loss value to obtain the adjusted image encoder.
[0095] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor: The original image to be processed is obtained, and multi-threshold topological feature tracking is performed on the original image to obtain a continuous homology map; The topology complexity index is determined based on the continuous homology graph, and the sampling parameters of the preset topology constraint transformation are defined through the topology complexity index. The original image is topologically preserved and enhanced according to the sampling parameters to obtain an enhanced view, and the enhanced view is refined and feature extracted to obtain enhanced features; Perform dual loss analysis on the enhanced features, and then weight and fuse the analyzed losses to generate a comprehensive loss value; The parameters of the preset image encoder are adjusted based on the comprehensive loss value to obtain the adjusted image encoder.
[0096] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0097] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0098] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0099] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them. If any software tools or components other than those of our company appear in the embodiments, they are merely illustrative examples and do not represent actual use. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. An image encoder optimization method based on topology enhancement, characterized in that, include: The original image to be processed is obtained, and multi-threshold topological feature tracking is performed on the original image to obtain a continuous homology map; The topology complexity index is determined based on the continuous homology graph, and the sampling parameters of the preset topology constraint transformation are defined by the topology complexity index. The original image is topologically preserved and enhanced according to the sampling parameters to obtain an enhanced view, and the enhanced view is then refined to extract enhanced features. Perform dual loss analysis on the enhanced features, and then weight and fuse the analyzed losses to generate a comprehensive loss value; The parameters of the preset image encoder are adjusted based on the comprehensive loss value to obtain the adjusted image encoder.
2. The image encoder optimization method based on topology enhancement as described in claim 1, characterized in that, The step of performing multi-threshold topological feature tracking on the original image to obtain a continuous homology map includes: Construct the spatial connection relationship between pixels in the original image, and transform the original image into a cubic complex based on the spatial connection relationship; The topological features of the cubic complex are dynamically tracked using preset multi-threshold parameters to obtain the tracking results; Based on the tracking results, the topological feature evolution law of the original image is extracted, and the continuous homology map is determined through the topological feature evolution law.
3. The image encoder optimization method based on topology enhancement as described in claim 1, characterized in that, The step of determining the topological complexity index based on the persistent homology graph includes: Identify the core topological feature trajectories in the persistent homology graph; Extract the Betti number sequence from the original image based on the core topological feature trajectory; The topological complexity index of the original image is determined based on the Betti number sequence.
4. The image encoder optimization method based on topology enhancement as described in claim 1, characterized in that, The sampling parameters for defining the preset topological constraint transformation through the topological complexity index include: The topological complexity index is used to analyze the topological structure type of the original image; The basic constraints are set for the topology type to obtain the basic constraints. The parameter range of the preset topology constraint transformation is adjusted according to the preset Gibbs distribution weighting strategy and the basic constraints. The sampling parameters are determined based on the parameter range and the preset topological similarity threshold.
5. The image encoder optimization method based on topology enhancement as described in claim 1, characterized in that, The step of performing topology-preserving enhancement on the original image based on the sampling parameters to obtain an enhanced view includes: Candidate transformations that meet the topological constraints are selected based on the sampling parameters; The original image is enhanced based on the candidate transformation to obtain a preliminary enhanced view; The enhanced view is obtained by comparing and verifying the topological features of the initial enhanced view with those of the original image.
6. The image encoder optimization method based on topology enhancement as described in claim 1, characterized in that, The process of refining and extracting enhanced features from the enhanced view to obtain enhanced features includes: Deep feature extraction is performed on the enhanced view to obtain a high-dimensional original feature vector; The high-dimensional original feature vector is optimized by performing feature dimension representation to obtain intermediate features; The correlation between the intermediate features and the original image is enhanced by using a preset topological alignment mechanism to obtain enhanced features.
7. The image encoder optimization method based on topology enhancement as described in claim 1, characterized in that, The step of performing dual loss analysis on the enhanced features and weighting and fusing the analyzed losses to generate a comprehensive loss value includes: The enhanced features are subjected to feature contrast loss analysis to obtain feature contrast loss values; Perform topology consistency loss analysis on the enhanced features to obtain the topology consistency loss value; The feature comparison loss value and the topology consistency loss value are weighted and summed according to a preset weight allocation strategy to generate a comprehensive loss value.
8. An image encoder optimization device based on topology enhancement, characterized in that, include: The acquisition and tracking module is used to acquire the original image to be processed, perform multi-threshold topological feature tracking on the original image, and obtain a continuous homology map; The definition module is used to determine the topology complexity index based on the continuous homology graph, and to define the sampling parameters of the preset topology constraint transformation through the topology complexity index; An enhancement module is used to perform topology-preserving enhancement on the original image based on the sampling parameters to obtain an enhanced view; The extraction module is used to extract refined features from the enhanced view to obtain enhanced features; The analysis and fusion module is used to perform dual loss analysis on the enhanced features and weightedly fuse the analyzed losses to generate a comprehensive loss value. The adjustment module is used to adjust the parameters of the preset image encoder according to the comprehensive loss value to obtain the adjusted image encoder.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the topology-enhanced image encoder optimization method as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the topology-enhanced image encoder optimization method as described in any one of claims 1 to 7.