Weak supervision target positioning method and system of similarity alignment distillation network based on CLIP

By combining CLIP model and distillation network technology, the problems of incomplete positioning and high labeling cost in traditional weak supervision target positioning methods are solved, and high-precision and low-cost target positioning are achieved, which is suitable for fields such as intelligent security and autonomous driving.

CN120599221APending Publication Date: 2025-09-05WUHAN UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510701712.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Traditional weakly supervised target positioning methods often ignore certain key areas of the target in the image, resulting in incomplete positioning and high cost of relying on pixel-level labeling, which is difficult to meet the actual application needs.

Method used

The similarity-aligned distillation network based on CLIP is adopted, combined with class activation map (CAM) and foreground prediction map (FPM), visual and semantic features are extracted through the CLIP model, and features are optimized using decoder and distillation technology. The exponential attenuation technology is used to distinguish the prospect from the background, and a high-precision target positioning map is generated.

Benefits of technology

It significantly improves the accuracy and clarity of target positioning, reduces the annotation cost, improves the adaptability and flexibility of the model, and is suitable for complex and changeable practical application environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120599221A_ABST
    Figure CN120599221A_ABST
Patent Text Reader

Abstract

The invention belongs to but not limited to the technical field of computer vision, and discloses a weak supervision target positioning method of a similarity alignment distillation network based on CLIP, which comprises the following steps: inputting image and text data into a pre-trained CLIP model for processing; the CLIP model extracts advanced visual and semantic features from the image and the text by using the deep learning ability of the CLIP model, and generates a self-attention map; image features are transmitted to a decoder, and the decoder carries out detailed analysis and fine adjustment on the features so as to better adapt to specific positioning requirements; the decoded image features and the text features are jointly used for calculating similarity, and a foreground prediction map is generated; the class activation graph and the foreground prediction graph are further optimized under the guidance of the CGDM module, meanwhile, the foreground prediction graph is processed by the EDFE module, and the EDFE module strengthens the foreground and inhibits the background through an exponential decay technology, so that the definition of the positioning graph is improved; and combining the class activation graph and the foreground prediction graph to generate a final positioning graph.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to but is not limited to the field of computer vision technology, and in particular relates to a weakly supervised target localization method and system based on a CLIP similarity alignment distillation network. Background Art

[0002] Object localization is a fundamental and critical technology in computer vision, involving the identification and location of specific objects in images. Traditional object localization methods typically rely on fully supervised learning, which requires each training example to be accurately labeled, such as with bounding boxes or pixel-level annotations. While this approach can achieve high accuracy, it also faces challenges such as high labeling costs, a time-consuming, and error-prone process.

[0003] With the growth of large-scale datasets and the advancement of deep learning technology, the limitations of fully supervised methods have become increasingly apparent, especially in resource-constrained applications. Therefore, reducing the reliance on detailed annotations has become a research hotspot in this field. Weakly supervised object localization (WSOL) technology has emerged to alleviate the annotation burden, allowing models to locate objects in images using only image-level labels. This approach does not require expensive annotations such as bounding boxes, but instead infers the location of objects by learning the overall classification label of the image.

[0004] Class Activation Map (CAM) and its derivatives are the mainstream methods to achieve this goal. CAM can visualize image areas related to specific categories through a global average pooling layer, but often only covers part of the object. To improve this limitation, researchers have proposed a variety of strategies, such as the self-attention mechanism of visual Transformer (TS-CAM), feature direction alignment, semantic similarity, and spatial relationship enhancement methods. In addition, a variety of techniques have been developed to expand the activation area, reduce background interference, and improve target coverage, such as B-CAM and E-CAM. 2 Net.

[0005] As an innovative technology in weakly supervised object localization (WSOL), foreground prediction map (FPM) focuses on leveraging the underlying features of the image to generate more detailed activation maps, thereby comprehensively covering all parts of the target. Compared with the traditional class activation map (CAM) method, FPM has significant advantages in accuracy and detail capture. In the BAS architecture, a weakly supervised localization paradigm based on FPM is proposed. By suppressing background activation values, the problem of premature convergence of cross entropy loss is overcome, thereby generating a more complete foreground prediction map and significantly improving the performance of weakly supervised object localization. Although FPM-based methods have achieved remarkable results, they still suffer from the problem of incomplete object localization.

[0006] In view of the above analysis, the technical problems that need to be solved urgently in the existing technology are:

[0007] Some traditional CAM-based methods can only highlight certain significant parts of the target in the image, while ignoring other key areas, resulting in incomplete positioning; some FPM-based methods have achieved significant results, but there is still the problem of incomplete target positioning. Summary of the Invention

[0008] In response to the problems existing in the prior art, the present invention provides a weakly supervised target localization method and system based on a CLIP-based similarity alignment distillation network.

[0009] The present invention is implemented as follows: a weakly supervised target localization method based on CLIP similarity alignment distillation network, comprising the following steps:

[0010] Step 1: Input image and text data into the pre-trained CLIP model for processing;

[0011] In step 2, the CLIP model uses its deep learning capabilities to extract high-level visual and semantic features from images and text, and generates self-attention maps to reveal the influence of different regions in the image on the prediction;

[0012] Step 3: The image features are transmitted to the decoder, which analyzes and fine-tunes the features in detail to better suit specific positioning requirements.

[0013] In step 4, the decoded image features and text features are used together to calculate similarity and generate a foreground prediction map (FPM). At the same time, these features are also used to generate a class activation map (CAM);

[0014] Step 5: The foreground prediction map (FPM) is processed by the EDFE module, which uses exponential decay technology to enhance the foreground and suppress the background, thereby improving the clarity of the positioning map.

[0015] In step six, the class activation map (CAM) and foreground prediction map (FPM) are combined and further optimized by the CGDM module to generate the final localization map.

[0016] Furthermore, the CLIP model is an efficient pre-trained deep learning model for extracting powerful features from images and related text; these features include but are not limited to the visual content of the image and its semantic correspondence with the text;

[0017] The CLIP model uses its feature extraction and object localization capabilities in CSDN to obtain preliminary features of images. In this way, the CLIP model provides a strong foundation for CSDN.

[0018] Furthermore, the decoder further processes and fine-tunes the features extracted from the CLIP model. This step mainly involves parsing, optimizing, and adjusting the features to ensure that these features are more suitable for specific data and application scenarios. The decoder is used to enhance and refine the original features provided by the CLIP model to make them more accurately meet the needs of target positioning.

[0019] Furthermore, the CGDM module uses distillation technology to optimize and correct the object recognition and localization process of CSDN by leveraging the existing advanced localization capabilities of the pre-trained CLIP model; CGDM provides a performance benchmark by analyzing and applying the output features of the CLIP model and its self-attention map to generate improved class activation maps (CAMs) and foreground prediction maps (FPMs).

[0020] Furthermore, the EDFE module uses exponential decay technology to precisely adjust the activation intensity of the foreground and background in the image, with a particular focus on enhancing the saliency of the foreground area. By exponentially strengthening the activation values ​​of foreground features and attenuating the activation values ​​of background features, it effectively distinguishes the foreground and background layers of the image.

[0021] Another object of the present invention is to provide a CLIP-based weakly supervised target localization similarity alignment distillation network for implementing the weakly supervised target localization method, comprising:

[0022] The CLIP model is an efficient pre-trained deep learning model for extracting powerful features from images and associated text; these features include but are not limited to the visual content of the image and its semantic correspondence with the text;

[0023] The decoder module further processes and fine-tunes the features extracted from the CLIP model. This step mainly involves parsing, optimizing, and adjusting the features to ensure that they are more suitable for specific data and application scenarios.

[0024] The CGDM module uses distillation technology to optimize and correct the object recognition and localization process of CSDN by leveraging the existing advanced localization capabilities of the pre-trained CLIP model;

[0025] The EDFE module exponentially enhances the activation value of the foreground feature and attenuates the activation value of the background feature to effectively distinguish the foreground and background layers of the image.

[0026] Another object of the present invention is to provide a weakly supervised target localization system based on a weakly supervised target localization method of a CLIP similarity alignment distillation network, comprising:

[0027] Data input module, inputs image and text data into the pre-trained CLIP model for processing;

[0028] Feature extraction module: The CLIP model uses its deep learning capabilities to extract high-level visual and semantic features from images and text, and generates self-attention maps to reveal the impact of different regions in the image on prediction;

[0029] Feature parsing and fine-tuning module: Image features are transmitted to the decoder, which performs detailed parsing and fine-tuning of the features to better meet specific positioning requirements;

[0030] In the FPM and CAM generation module, the decoded image features and text features are used together to calculate the similarity and generate the foreground prediction map (FPM). At the same time, these features are also used to generate the class activation map (CAM);

[0031] FPM processing module,The foreground prediction map (FPM) is processed by the EDFE module, which uses the exponential decay technology to enhance the foreground and suppress the background, thereby improving the clarity of the positioning map;

[0032] CAM and FPM combination module, the class activation map (CAM) and foreground prediction map (FPM) are combined and further optimized by the CGDM module to generate the final localization map.

[0033] Another object of the present invention is to provide a computer device, which includes a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the weakly supervised target localization method based on the CLIP similarity alignment distillation network.

[0034] Another object of the present invention is to provide a computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to perform the steps of the weakly supervised target localization method based on the CLIP-based similarity alignment distillation network.

[0035] Another object of the present invention is to provide an information data processing terminal, which includes the weakly supervised target positioning system.

[0036] In combination with the above technical solutions and the technical problems solved, the advantages and positive effects of the technical solutions to be protected by the present invention are as follows:

[0037] First, the CLIP-driven Similarity-aligned Distillation Network (CSDN) framework proposed in this paper successfully addresses many limitations of traditional weakly supervised object localization (WSOL) methods. By integrating CAM and FPM technologies and adding unique modules, it significantly improves the accuracy of object localization and the practicality of the model.

[0038] 1. Accurate target positioning

[0039] By integrating CAM and FPM, the proposed method leverages CAM's class activation capabilities and FPM's fine-grained feature utilization, enabling the system to comprehensively cover and accurately locate the entire target. This fusion strategy significantly improves target location accuracy, especially in complex and changing real-world environments.

[0040] 2. Increased attention to prospects

[0041] The CLIP Guidance Distillation Module (CGDM) is a key component of our CSDN architecture. The CLIP model, thanks to its cross-modal training, has demonstrated remarkable capabilities in understanding and localizing objects in images. The CGDM module leverages this capability to transfer high-level visual features from CLIP to CSDN through distillation, inheriting CLIP's precise calibration of key image regions and improving focus on foreground areas.

[0042] 3. Background suppression and foreground highlighting

[0043] The Exponential Decay Foreground Emphasis (EDFE) module of this invention employs an innovative exponential decay strategy to enhance foreground signals and suppress background noise. By adjusting the decay rate of foreground and background activation values, this module effectively separates foreground and background, thereby improving the clarity of the localization map and the recognition of targets. This approach not only optimizes foreground visibility but also significantly reduces background interference, thereby enhancing the model's adaptability to diverse scenarios while maintaining high-quality output.

[0044] 4. Cost-effectiveness and simplified training process

[0045] The CSDN framework of the present invention significantly reduces the amount of data and computing resources required for model training by using pre-trained CLIP models. This strategy of reusing pre-trained models not only reduces costs but also accelerates the model development cycle, providing a highly cost-effective solution. Furthermore, the high transferability and adaptability of pre-trained models enables CSDN to adapt to new tasks and environments in a short period of time, greatly improving the flexibility and efficiency of R&D.

[0046] Second, traditional fully supervised learning relies on pixel-level manual annotation (such as bounding box and mask annotation), which requires a lot of professional manpower and time costs. The present invention uses image-level weakly supervised learning technology, which only requires the binary label of "whether the target object is contained" to achieve highly robust target positioning. By associating image-level labels with local feature responses, the target heat map is automatically generated and converted into a positioning frame, completely avoiding the manpower consumption of pixel-by-pixel annotation. In some fields of massive image data processing, this technology can significantly reduce the cost of annotation, while lowering the training threshold for annotation personnel from professional labelers to ordinary operation and maintenance personnel, significantly accelerating the iteration and deployment process of AI models.

[0047] For a long time, weakly supervised target positioning tasks have faced problems such as incomplete target area activation and severe background interference when relying solely on image-level annotation, resulting in target positioning accuracy and reliability that are difficult to meet practical application requirements. The present invention effectively integrates the two mechanisms of CAM and FPM, and introduces a multimodal alignment strategy based on text-image similarity to accurately guide the foreground activation of image features under weak supervision conditions, effectively eliminating the technical bottlenecks of traditional methods in positioning integrity and background suppression. Furthermore, by designing an exponential decay module, the ability to distinguish between foreground and background is enhanced, and dynamic regulation of target area activation intensity and background response is achieved, providing a new technical means to improve positioning quality.

[0048] Traditional weakly supervised target localization methods generally believe that unimodal visual features are sufficient to achieve satisfactory positioning results, while less attention is paid to the potential of text-image multimodal information fusion. In addition, the industry generally believes that the class activation map (CAM) and foreground prediction map (FPM) methods are each independent research routes, and there are few attempts to organically combine the two. To this end, the present invention attempts to organically combine the CAM and FPM mechanisms, and introduces a CLIP-based multimodal pre-training model to achieve visual-language feature alignment and knowledge distillation, thereby achieving a more significant improvement in positioning performance. This method verifies the feasibility and effectiveness of the multimodal fusion strategy in weakly supervised scenarios, and is expected to provide new research ideas and practical paths for subsequent weakly supervised visual tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 This is a flow chart of a weakly supervised target localization method based on a CLIP similarity alignment distillation network provided by an embodiment of the present invention;

[0050] Figure 2 is a diagram showing the structure of a CLIP-based weakly supervised target localization similarity alignment distillation network provided by an embodiment of the present invention;

[0051] Figure 3 2 is a structural diagram of a weakly supervised target localization system based on a CLIP similarity alignment distillation network provided by an embodiment of the present invention;

[0052] Figure 4 It is a schematic diagram of the implementation effect of the weakly supervised target localization method of the CLIP-based similarity alignment distillation network provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0053] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0054] like Figure 1 As shown, an embodiment of the present invention provides a weakly supervised target localization method based on a CLIP similarity alignment distillation network, comprising the following steps:

[0055] Step 1: Input image and text data into the pre-trained CLIP model for processing;

[0056] In step 2, the CLIP model uses its deep learning capabilities to extract high-level visual and semantic features from images and text, and generates self-attention maps to reveal the influence of different regions in the image on the prediction;

[0057] Step 3: The image features are transmitted to the decoder, which analyzes and fine-tunes the features in detail to better suit specific positioning requirements.

[0058] In step 4, the decoded image features and text features are used together to calculate similarity and generate a foreground prediction map (FPM). At the same time, these features are also used to generate a class activation map (CAM);

[0059] Step 5: The foreground prediction map (FPM) is processed by the EDFE module, which uses exponential decay technology to enhance the foreground and suppress the background, thereby improving the clarity of the positioning map.

[0060] In step six, the class activation map (CAM) and foreground prediction map (FPM) are combined and further optimized by the CGDM module to generate the final localization map.

[0061] CLIP Image and Text Encoder:

[0062] The CLIP model is an efficient pre-trained deep learning model that is primarily used to extract powerful features from images and associated text. These features include, but are not limited to, the visual content of the image and its semantic correspondence with the text.

[0063] The CLIP model's primary role in CSDN is to leverage its superior feature extraction and object localization capabilities to obtain preliminary image features. This provides a powerful foundation for CSDN, enabling the entire system to improve efficiency and achieve high-precision object localization in complex visual scenes.

[0064] Decoder module:

[0065] The decoder module is used to further process and fine-tune the features extracted from the CLIP model. This step mainly involves parsing, optimizing, and adjusting the features to ensure that they are more suitable for specific data and application scenarios.

[0066] The decoder's core role is to enhance and refine the raw features provided by the CLIP model, making them more precisely suited to target localization. By further processing the features, the decoder ensures that the output features not only retain their original rich information but also increase sensitivity to detail and adaptability to specific tasks. This fine-tuning process is key to improving overall system performance, particularly when dealing with challenging visual environments, significantly enhancing the accuracy of target recognition and localization.

[0067] CGDM module:

[0068] The CGDM module mainly uses distillation technology to optimize and correct the target recognition and localization process of CSDN by leveraging the existing advanced localization capabilities of the pre-trained CLIP model.

[0069] The core purpose of CGDM is to provide a performance benchmark by analyzing and applying the output features of the CLIP model and its self-attention map to generate an improved Class Activation Map (CAM). This process enables CSDN to draw on CLIP's efficient localization strategy, significantly improving the accuracy of object localization. In this way, CGDM not only improves CSDN's ability to identify and locate objects in images, but also reduces the model's misidentification rate when dealing with complex or interfering backgrounds.

[0070] EDFE module:

[0071] The EDFE module uses exponential decay technology to precisely adjust the activation strength of the foreground and background in an image, with a particular emphasis on enhancing the saliency of the foreground region. This method effectively distinguishes the foreground and background layers of an image by exponentially increasing the activation values ​​of foreground features while simultaneously attenuating the activation values ​​of background features.

[0072] The core function of the EDFE module is to significantly enhance foreground highlighting and significantly reduce the impact of background noise through meticulous adjustments to the activation map. This method of highlighting the foreground and suppressing the background not only clearly defines the target area but also increases the contrast between the target and the background, making the positioning map more visually clear and easier to distinguish, thereby significantly improving the accuracy and reliability of the entire system's positioning.

[0073] like Figure 2 As shown, an embodiment of the present invention provides a CLIP-based weakly supervised target positioning similarity alignment distillation network for implementing the weakly supervised target positioning method, including:

[0074] The CLIP model is an efficient pre-trained deep learning model for extracting powerful features from images and associated text; these features include but are not limited to the visual content of the image and its semantic correspondence with the text;

[0075] The decoder module further processes and fine-tunes the features extracted from the CLIP model. This step mainly involves parsing, optimizing, and adjusting the features to ensure that they are more suitable for specific data and application scenarios.

[0076] The CGDM module uses distillation technology to optimize and correct the object recognition and localization process of CSDN by leveraging the existing advanced localization capabilities of the pre-trained CLIP model;

[0077] The EDFE module exponentially enhances the activation value of the foreground feature and attenuates the activation value of the background feature to effectively distinguish the foreground and background layers of the image.

[0078] like Figure 3 As shown, the embodiment of the present invention provides a weakly supervised target positioning system of a weakly supervised target positioning method based on a CLIP similarity alignment distillation network, including:

[0079] Data input module, inputs image and text data into the pre-trained CLIP model for processing;

[0080] Feature extraction module: The CLIP model uses its deep learning capabilities to extract high-level visual and semantic features from images and text, and generates self-attention maps to reveal the impact of different regions in the image on prediction;

[0081] Feature parsing and fine-tuning module: Image features are transmitted to the decoder, which performs detailed parsing and fine-tuning of the features to better meet specific positioning requirements;

[0082] In the FPM and CAM generation module, the decoded image features and text features are used together to calculate the similarity and generate the foreground prediction map (FPM). At the same time, these features are also used to generate the class activation map (CAM);

[0083] FPM processing module,The foreground prediction map (FPM) is processed by the EDFE module, which uses the exponential decay technology to enhance the foreground and suppress the background, thereby improving the clarity of the positioning map;

[0084] CAM and FPM combination module, the class activation map (CAM) and foreground prediction map (FPM) are combined and further optimized by the CGDM module to generate the final localization map.

[0085] The specific application fields or related products of the present invention.

[0086] 1. Intelligent security monitoring system

[0087] In the field of public safety, weakly supervised target positioning technology can significantly improve the automation level of monitoring systems. Taking crowded places such as airports and subway stations as an example, traditional monitoring systems rely on manually labeled large amounts of bounding box data to train target detection models, which is costly and difficult to cover dynamic scenes. Through a weakly supervised learning framework, the present invention only needs to label the monitoring video frames with image-level labels (such as "suspicious objects exist" or "normal scenes") to automatically generate target positioning heat maps. For example, when an unattended suitcase appears in the monitoring screen, the system analyzes the correlation between the image-level label and the feature activation area, frames the position of the suspicious object in real time, triggers an audible and visual alarm, and superimposes the positioning results on the control interface of the security personnel. The response time can be shortened to less than 200 milliseconds.

[0088] For perimeter security scenarios (such as industrial parks and nuclear power plants), traditional intrusion detection requires pixel-level annotation of individuals climbing over fences or illegally entering. However, this invention can directly train models using weakly annotated video data (such as the "intrusion event" label). During deployment, the system automatically identifies abnormal behavior by analyzing heat maps of moving targets in the video stream and activates defense devices (such as automatic searchlights and drone tracking).

[0089] 2. Autonomous Driving Perception System

[0090] In the field of autonomous driving, weakly supervised target localization technology addresses the core pain point of a lack of high-precision annotated data. Traditional methods require bounding box annotation of vehicles and pedestrians in every frame of lidar or camera data, which is costly for single-frame annotation. However, a multi-camera fusion framework allows training images to be labeled "vehicle / pedestrian present" to generate target heat maps and extract coarse localization areas. For example, in urban road scenarios, the system can simultaneously process input from both front-view and side-view cameras, outputting real-time heat maps of traffic participants within a 50-meter radius of the vehicle.

[0091] Traditional models are prone to missing detections in long-tail scenarios (such as fallen trees and wild animals on mountain roads) due to a lack of labeled data. This invention, through a weakly supervised learning strategy, requires only a few hundred images labeled "obstacles present" to train a localization model adapted to these specific scenarios. Technically, an attention mechanism is employed to enhance the model's sensitivity to rare objects, and transfer learning is used to transfer general scene knowledge to long-tail tasks.

[0092] Relevant evidence of the technical effects achieved by the embodiments of the present invention.

[0093] The implementation effect is as follows Figure 4 As shown in the figure: the original image is on the far left, the localization map generated through a series of operations is in the middle, and the localization box generated based on the localization map and the intersection over union (IoU) ratio between the calculated localization box and the true target area are on the far right. This IoU value is a key indicator for measuring localization accuracy. A larger value indicates a higher degree of overlap between the predicted localization box and the true target area, which means that our target localization method has a higher accuracy.

[0094] In the evaluation process of weakly supervised object localization (WSOL), three indicators, Top-1 localization accuracy, Top-5 localization accuracy and GT-Known localization accuracy, are usually used to measure the performance of the model. Top-1 localization accuracy measures whether the model can correctly mark the location of the target in its most confident category prediction, that is, whether the class activation map (CAM) of the most likely category predicted by the model correctly covers at least half of the true target area. Top-5 localization accuracy checks whether the top five most likely categories predicted by the model contain the correct target location, allowing the model to be considered correct if it correctly marks the target location in any of the top five predictions. GT-Known localization accuracy focuses on the performance of the model when the correct category is known, and evaluates the model's positioning ability when the target category is clear. Table 1 shows the comparison of the method of the present invention and some other methods in terms of these three indicators:

[0095] Table 1 Experimental comparison

[0096]

[0097] It should be noted that the embodiments of the present invention can be implemented by hardware, software, or a combination of software and hardware. The hardware portion can be implemented using dedicated logic; the software portion can be stored in a memory and executed by an appropriate instruction execution system, such as a microprocessor or dedicated design hardware. Those skilled in the art will appreciate that the above-mentioned devices and methods can be implemented using computer-executable instructions and / or contained in processor control code, for example, such as a carrier medium such as a disk, CD or DVD-ROM, a programmable memory such as a read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The devices and modules of the present invention can be implemented by hardware circuits such as very large-scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, or programmable hardware devices such as field programmable gate arrays, programmable logic devices, etc., can also be implemented by software executed by various types of processors, or can be implemented by a combination of the above-mentioned hardware circuits and software, such as firmware.

[0098] The above description is only a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications, equivalent substitutions and improvements made by any technician familiar with this technical field within the technical scope disclosed by the present invention and within the spirit and principles of the present invention should be covered by the scope of protection of the present invention.

Claims

1. A weakly supervised target localization method based on CLIP similarity alignment distillation network, characterized by: The following steps are involved: Step 1: Input image and text data into the pre-trained CLIP model for processing; In step 2, the CLIP model uses its deep learning capabilities to extract high-level visual and semantic features from images and text, and generates self-attention maps to reveal the influence of different regions in the image on the prediction; Step 3: The image features are transmitted to the decoder, which analyzes and fine-tunes the features in detail to better suit specific positioning requirements. In step 4, the decoded image features and text features are used together to calculate similarity and generate a foreground prediction map (FPM). At the same time, these features are also used to generate a class activation map (CAM); Step 5: The foreground prediction map (FPM) is processed by the EDFE module, which uses exponential decay technology to enhance the foreground and suppress the background, thereby improving the clarity of the positioning map. In step six, the class activation map (CAM) and foreground prediction map (FPM) are combined and further optimized by the CGDM module to generate the final localization map.

2. The weakly supervised target localization method based on CLIP similarity alignment distillation network according to claim 1, characterized in that: The CLIP model is an efficient pre-trained deep learning model for extracting powerful features from images and associated text; these features include but are not limited to the visual content of the image and its semantic correspondence with the text; The CLIP model uses its feature extraction and object localization capabilities in CSDN to obtain preliminary features of images. In this way, the CLIP model provides a strong foundation for CSDN.

3. The weakly supervised target localization method based on CLIP similarity alignment distillation network according to claim 1, characterized in that: The decoder further processes and fine-tunes the features extracted from the CLIP model. This step mainly involves parsing, optimizing, and adjusting the features to ensure that they are more suitable for specific data and application scenarios. The decoder is used to enhance and refine the original features provided by the CLIP model to make them more accurately meet the needs of target positioning.

4. The weakly supervised target localization method based on CLIP similarity alignment distillation network according to claim 1, characterized in that: The CGDM module uses distillation technology to optimize and correct the object recognition and localization process of CSDN by leveraging the existing advanced localization capabilities of the pre-trained CLIP model; CGDM provides a performance benchmark by analyzing and applying the output features of the CLIP model and its self-attention map to generate an improved class activation map (CAM).

5. The weakly supervised target localization method based on CLIP similarity alignment distillation network according to claim 1, characterized in that: The EDFE module uses exponential decay technology to precisely adjust the activation intensity of the foreground and background in the image, with a particular focus on enhancing the saliency of the foreground area. By exponentially strengthening the activation values ​​of foreground features and attenuating the activation values ​​of background features, it effectively distinguishes the foreground and background layers of the image.

6. A CLIP-based weakly supervised target localization similarity alignment distillation network implementing the weakly supervised target localization method according to any one of claims 1 to 5, characterized in that: include: CLIP model, an efficient pre-trained deep learning model for extracting powerful features from images and associated text; These features include, but are not limited to, the visual content of the image and its semantic correspondence with the text; The decoder module further processes and fine-tunes the features extracted from the CLIP model. This step mainly involves parsing, optimizing, and adjusting the features to ensure that they are more suitable for specific data and application scenarios. The CGDM module uses distillation technology to optimize and correct the object recognition and localization process of CSDN by leveraging the existing advanced localization capabilities of the pre-trained CLIP model; The EDFE module exponentially enhances the activation value of the foreground feature and attenuates the activation value of the background feature to effectively distinguish the foreground and background layers of the image.

7. A weakly supervised target localization system of the weakly supervised target localization method based on CLIP similarity alignment distillation network according to any one of claims 1 to 5, characterized in that: include: Data input module, inputs image and text data into the pre-trained CLIP model for processing; Feature extraction module: The CLIP model uses its deep learning capabilities to extract high-level visual and semantic features from images and text, and generates self-attention maps to reveal the impact of different regions in the image on prediction; Feature parsing and fine-tuning module: Image features are transmitted to the decoder, which performs detailed parsing and fine-tuning of the features to better meet specific positioning requirements; In the FPM and CAM generation module, the decoded image features and text features are used together to calculate the similarity and generate the foreground prediction map (FPM). At the same time, these features are also used to generate the class activation map (CAM); FPM processing module,The foreground prediction map (FPM) is processed by the EDFE module, which uses the exponential decay technology to enhance the foreground and suppress the background, thereby improving the clarity of the positioning map; CAM and FPM combination module, the class activation map (CAM) and foreground prediction map (FPM) are combined and further optimized by the CGDM module to generate the final localization map.

8. A computer device, characterized in that: The computer device includes a memory and a processor, the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of the weakly supervised target localization method based on the CLIP similarity alignment distillation network as described in any one of claims 1 to 5.

9. A computer-readable storage medium, characterized in that A computer program is stored, and when the computer program is executed by a processor, the processor performs the steps of the weakly supervised target localization method based on the CLIP similarity alignment distillation network according to any one of claims 1 to 5.

10. An information data processing terminal, characterized in that: The information data processing terminal includes the weakly supervised target positioning system as described in claim 7.

Citation Information

Cited By

  • Training method, detection method and system of cerebral hemorrhage focus detection model

    CN121746385A