Method, device and storage medium for multi-source image forged region segmentation
Through the pre-training and point prompt mechanism of the image recognition model and combined with the loss function, the precise segmentation of multi-source areas in the image is achieved, solving the problem of not being able to identify different source areas in the prior art, and improving the accuracy and efficiency of image recognition.
Patent Information
- Application Number
- CN202510141472.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2045-02-08
AI Technical Summary
The prior art cannot effectively identify and segment areas of different sources in images, limiting their application in complex situations.
By pre-training the image recognition model, using area comparison learning and point prompt mechanisms, combining the InfoNCE loss function and area adaptive source segmentation loss function, precise segmentation of multi-source areas in the image is achieved.
It improves the accuracy and efficiency of image recognition, can fully present the tampered information of the image, and enhances the understanding and segmentation ability of multi-source composition.
Smart Images

Figure CN119579626B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of digital image processing, and in particular to a method, device and storage medium for multi-source image forgery region segmentation. Background Art
[0002] With the rapid development of digital technology, image editing and generation tools are becoming increasingly popular, greatly enriching people's visual expression. However, this also facilitates the malicious manipulation and forgery of images, posing a serious challenge to image authenticity. In news reporting, false images can mislead the public and influence public opinion. In the legal field, forged image evidence can interfere with judicial justice and undermine legal order. In commercial advertising, false product images can mislead consumers and undermine market integrity.
[0003] Traditional image forgery localization (IFL) methods mostly use binary segmentation, marking the original area as 0 and the tampered area as 1. However, this method can only distinguish between true and false, and cannot further analyze multiple source regions in the image, limiting its application in complex situations. Summary of the Invention
[0004] The present invention provides a method, device and storage medium for multi-source image forgery region segmentation, which are used to solve the problem that the prior art cannot accurately identify and segment regions from different sources in an image.
[0005] In order to achieve the above-mentioned purpose, an embodiment of the present invention provides a method for multi-source image forgery area segmentation, the method comprising the following steps: obtaining an image to be detected and corresponding one or more point prompts arranged in a grid, and inputting them into a trained image recognition model, each of the one or more point prompts is used to indicate the source area of the image to be detected; using the encoder of the image recognition model to operate on the image to be detected to obtain a feature representation of the image to be detected; for each point prompt, using the mask decoder of the image recognition model to perform segmentation prediction on the feature representation of the image to be detected, to obtain the predicted area and prediction result corresponding to each point prompt; using a preset clustering algorithm to integrate the prediction results to form a multi-source partition map.
[0006] Optionally, before obtaining the image to be detected and the corresponding one or more point prompts arranged in a grid, the method further includes: pre-training the image recognition model using regional contrast learning.
[0007] Optionally, the pre-training of the image recognition model using regional contrast learning includes: converting the input image into an image embedding through the image encoder of the image recognition model; calculating the similarity and difference between the image embeddings using the InfoNCE loss function, and applying the result calculated by the InfoNCE loss function to all source regions of the input image to obtain the contrast loss between regions. , according to the contrast loss between the regions , adjust the parameters of the image encoder to complete the pre-training of the image recognition model.
[0008] Optionally, the contrast loss between the regions Expressed as:
[0009]
[0010] in, (·) represents the loss value, represents the number of source regions in the input image, Indicates the The feature set corresponding to the source region, represents the number of elements in the first source region, Indicates that from source regions to remove the current query vector The remaining feature set after is taken as the positive sample. Indicates that except for All other features outside the source region are regarded as negative samples.
[0011] Optionally, after pre-training the image recognition model using regional contrast learning, the method further includes training the image recognition model, including: fixing the pre-trained image encoder and keeping the prompt encoder of the image recognition model in its original state; converting the binary label dataset of the input image into a point mask through connected component analysis; configuring a loss function based on the point mask, the loss function including an area-adaptive source segmentation loss function and a confidence loss function; and optimizing the parameters of the mask decoder based on the loss function.
[0012] Optionally, the feature representation includes image embedding and prompt embedding, and the encoder of the image recognition model is used to operate on the image to be detected to obtain the feature representation of the image to be detected, including: encoding the image to be detected through a trained image encoder to obtain image embedding; encoding the point prompt through a trained prompt encoder to obtain prompt embedding.
[0013] Optionally, for each point prompt, the mask decoder of the image recognition model is used to perform segmentation prediction on the feature representation of the image to be detected to obtain the prediction area and prediction result corresponding to each point prompt, including: using the mask decoder to perform segmentation prediction on the image embedding and the prompt embedding to generate a corresponding segmentation mask and confidence score for each point prompt; determining the corresponding prediction area based on the segmentation mask, and calculating the average image embedding corresponding to the area based on the determined prediction area as a representative feature; obtaining the prediction result based on the confidence score and the representative feature.
[0014] Optionally, the prediction results are integrated using a preset clustering algorithm to form a multi-source partition map, including: in the clustering process, the segmentation mask with the highest confidence score is screened out from each cluster as the final representative feature of the cluster, and the multi-source partition map is formed through integration.
[0015] On the other hand, the present invention further provides a control device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement any one of the above methods.
[0016] On the other hand, the present invention further provides a machine-readable storage medium having instructions stored thereon, wherein the instructions enable a machine to execute any one of the above methods.
[0017] The present invention provides a method, device, and storage medium for segmenting forged regions in multi-source images. By pre-training the image recognition model, the model's ability to understand and capture subtle differences is improved, thereby enhancing the accuracy of image recognition. Through a point prompt mechanism, refined segmentation of regions from different sources in the image is achieved, which can fully present the image's tampering information or multi-source composition. Through the above method, the present invention establishes a method for effectively segmenting forged regions in multi-source images. With an efficient reasoning process, efficiency is improved, accuracy is enhanced, and the problem of locating image forgeries is comprehensively solved, with important application value in multiple fields. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions of the present invention or the prior art, the following briefly introduces the drawings required for the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work. In the drawings:
[0019] Figure 11 is a flow chart of a method for segmenting forged regions in multi-source images provided by an embodiment of the present invention;
[0020] Figure 2 is another flowchart of a method for segmenting forged regions in multi-source images provided by an embodiment of the present invention;
[0021] Figure 3 1 is a schematic structural diagram of a method for multi-source image forgery region segmentation provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0022] The following describes the specific implementation of the embodiment of the present invention in detail with reference to the accompanying drawings. It should be understood that the specific implementation described herein is only used to illustrate and explain the embodiment of the present invention and is not used to limit the embodiment of the present invention.
[0023] It should be noted that the acquisition, transmission, storage, use, and processing of data in the technical solution of this application are in compliance with the relevant provisions of national laws and regulations. In the embodiments of this application, certain software, components, models, and other existing solutions in the industry may be mentioned. These should be considered as exemplary. Their purpose is only to illustrate the feasibility of implementing the technical solution of this application, but it does not mean that the applicant has or will necessarily use such solutions.
[0024] As mentioned above, with the rapid development of digital technology, image editing and generation tools are becoming increasingly popular, which has facilitated the malicious tampering and forgery of images. The authenticity of images faces severe challenges. Existing technologies are unable to analyze the specific sources of each area and it is difficult to present complete tampering information. Therefore, a method that can accurately identify and segment different source areas in an image is needed to realize the identification of forged images.
[0025] To address this problem, the present invention provides a method, device and storage medium for multi-source image forgery region segmentation. Through a point prompt mechanism, accurate segmentation of multi-source regions in an image is achieved. Through inter-region contrast learning and pre-training, the image recognition model can effectively capture subtle feature differences between regions and improve feature resolution. Through the loss function, it ensures that both large and small regions can be accurately segmented, thereby improving the reliability of the output results. Moreover, through a dual-region paired point sampling strategy, the image recognition model can learn different regions more balancedly, greatly improving the accuracy of image recognition.
[0026] The following combination Figure 1-Figure 3 The present invention will be described in detail.
[0027] Figure 1 FIG. 1 is a flow chart of a method for segmenting forged regions of multi-source images provided by an embodiment of the present invention. Figure 1As shown, the method for multi-source image forgery region segmentation provided by the embodiment of the present invention may be executed by a controller, and the method includes the following steps:
[0028] Step 101: Obtain an image to be detected and corresponding one or more point prompts arranged in a grid, and input them into a trained image recognition model, where each of the one or more point prompts is used to indicate the source area of the image to be detected;
[0029] Step 102: using the encoder of the image recognition model to perform operations on the image to be detected to obtain a feature representation of the image to be detected;
[0030] Step 103: For each point hint, use the mask decoder of the image recognition model to perform segmentation prediction on the feature representation of the image to be detected, and obtain the prediction area and prediction result corresponding to each point hint;
[0031] Step 104: Utilize a preset clustering algorithm to integrate the prediction results to form a multi-source partition map.
[0032] The method for multi-source image forgery region segmentation provided by the present invention utilizes a point prompt mechanism, which not only achieves accurate segmentation of multi-source regions in an image, breaking the limitations of traditional binary segmentation, but also greatly enhances the interactivity between users and the system, meeting diverse application needs. A trained image recognition model can accurately locate image forgery regions. At the same time, when detecting a large number of images, there is no need to recalculate image embeddings for each point prompt, saving a large amount of computing resources and time.
[0033] Please refer to Figure 2 and Figure 3 For example, preferably, before step 101, the method further includes: pre-training the image recognition model using region-to-region contrastive learning. Region-to-region contrastive learning can enhance the image encoder's understanding of source region features.
[0034] Further preferably, the pre-training of the image recognition model using regional contrast learning may include: converting the input image into an image embedding by using the image encoder of the image recognition model; calculating the similarity and difference between the image embeddings by using the InfoNCE loss function, and applying the result calculated by the InfoNCE loss function to all source regions of the input image to obtain the contrast loss between regions. , according to the contrast loss between the regions , adjust the parameters of the image encoder to complete the pre-training of the image recognition model.
[0035] In a preferred embodiment of the present invention, the InfoNCE loss function can measure the degree of distinction between positive samples and negative samples, and by mining the relationship between internal regions of the data, it can effectively improve the performance of the model in various tasks. In the pre-training stage, the image embedding E=E(I) generated by the image encoder is used for subsequent inter-region contrast learning. For example, the InfoNCE loss function is used to calculate the similarities and differences between these embeddings. For feature points in the same source region, their distance in the feature space should be as close as possible; for feature points from different source regions, they should be as far as possible. Subsequently, the InfoNCE loss is applied to all source regions of the entire image to obtain the inter-region contrast loss , inter-region contrast loss It aims to minimize the distance between feature points in the same source region and maximize the distance between feature points in different source regions.
[0036] Further preferably, the loss function of the encoder can be expressed by the following formula:
[0037]
[0038] in, Represents the query vector, which is the feature extracted from a specific location. Represents a positive sample vector, another feature point that belongs to the same source region as the query vector. represents the set of negative samples, which are feature points from other source regions. τ represents the temperature parameter, which controls the scaling of the similarity score. A smaller τ value results in higher scores for correct matches and lower scores for incorrect matches, thereby increasing the probability of selecting the correct match. · Represents the dot product operation of two vectors, measuring the similarity between them. (·) represents the exponential function, which is used to convert the dot product result into a positive value and strengthen the difference between high and low scores.
[0039] Further preferably, the contrast loss between the regions Expressed as:
[0040]
[0041] in, (·) represents the loss value, represents the number of source regions in the input image, Indicates the The feature set corresponding to the source region, represents the number of elements in the first source region, Indicates that from source regions to remove the current query vector The remaining feature set after is taken as the positive sample. Indicates that except for All other features outside the source region are regarded as negative samples.
[0042] Preferably, after pre-training, the method further comprises: training the image recognition model to optimize the decoder.
[0043] Further preferably, the training of the image recognition model includes: fixing the pre-trained image encoder and keeping the prompt encoder of the image recognition model in its original state; converting the binary label data set of the input image into a point mask through connected component analysis; configuring a loss function based on the point mask, the loss function including an area-adaptive source segmentation loss function and a confidence loss function; and optimizing the parameters of the mask decoder based on the loss function.
[0044] In a preferred embodiment of the present invention, the image encoder remains in its pre-trained state, while the hint encoder uses the original SAM state. For binary labeled datasets, connected components are analyzed to convert them into point masks suitable for multi-source prediction. The point mask creation process involves analyzing the connected components in the binary labeled dataset and converting them into a point mask suitable for multi-source prediction. This process is performed automatically by the algorithm. For a given point, if it is within a connected region, that region is labeled as 1, while adjacent regions that do not belong to the region are labeled as 0, and non-adjacent regions are assigned the ignore label -1. Subsequently, an area-adaptive source segmentation loss function assigns higher weights to smaller regions based on the areas represented by the point mask. For example, if an image contains both large original background regions and small tampered object regions, the area-adaptive source segmentation loss function guides the mask decoder to accurately segment source regions of different sizes through reasonable weight distribution (determined based on factors such as region area ratio, image feature distribution, and task requirements), avoiding segmentation bias caused by differences in region area and ensuring that both large and small regions are accurately predicted by the image recognition model. Traditional image forgery detection technology often only outputs segmentation results, making it difficult for users to intuitively understand how credible the results are.
[0045] This method uses confidence loss to measure the difference between the predicted confidence and the ground truth, quantifying it using metrics such as Euclidean distance and cosine similarity. During training, the mask decoder aims to minimize this confidence score, generating highly reliable segmentation masks. This results in segmentation masks with a confidence score assigned to each mask when outputting the segmentation results. This allows users to assess the reliability of the results, avoiding misjudgments caused by unreliable segmentation results and enhancing the practical value of the entire technology.
[0046] In addition, the present invention adopts dual-region paired point sampling. This sampling strategy ensures that the image recognition model pays equal attention to different source regions during the training process, and evenly learns the differences in their respective features, textures, colors, etc., as well as the relationships between them, so that the image recognition model has a better understanding and discrimination ability for various types of regions, thereby improving the overall segmentation performance.
[0047] Preferably, the feature representation includes image embedding and prompt embedding, and the method of using the encoder of the image recognition model to operate on the image to be detected to obtain the feature representation of the image to be detected includes: encoding the image to be detected by a trained image encoder to obtain an image embedding; and encoding the point prompt by a trained prompt encoder to obtain a prompt embedding. Image embedding achieves efficient processing, cross-modal fusion, and feature sharing by extracting and converting image features into low-dimensional vectors, and has the advantages of high efficiency, semantic preservation, and transferability; prompt embedding provides additional guidance information for the model, helping it focus on key points, enhance semantic understanding, and customize tasks, and has the advantages of flexibility, strong interpretability, and improved performance.
[0048] Preferably, in step 103, for each point prompt, the mask decoder of the image recognition model is used to perform segmentation prediction on the feature representation of the image to be detected to obtain the prediction area and prediction result corresponding to each point prompt, including: using the mask decoder to perform segmentation prediction on the image embedding and the prompt embedding to generate a corresponding segmentation mask and confidence score for each point prompt; determining the corresponding prediction area according to the segmentation mask, and calculating the average image embedding corresponding to the area according to the determined prediction area as a representative feature; obtaining the prediction result according to the confidence score and the representative feature.
[0049] In a preferred embodiment of the present invention, the calculation of image embedding is independent of the point hint, and only one encoder operation is required for each image to obtain a complete feature representation. This one-time calculation method reduces repeated operations and greatly improves efficiency compared to traditional methods. The mask decoder generates a segmentation mask and its confidence score for each point hint based on the image embedding and hint embedding, combines the image features with the hint information to predict the possible source area and provide a reliability measure, paving the way for subsequent screening of accurate results. After determining the area according to the segmentation mask, the average image embedding is calculated as a representative feature, which provides a basis for cluster analysis. By extracting the representative features of the region, it facilitates the subsequent integrated analysis of the relationship between different regions.
[0050] Preferably, in step 104, the prediction results are integrated using a preset clustering algorithm to form a multi-source partition map, including: in the clustering process, the segmentation mask with the highest confidence score is screened out from each cluster as the final representative feature of the cluster, and the multi-source partition map is formed through integration.
[0051] In a preferred embodiment of the present invention, a clustering algorithm (such as k-means) is used to integrate the preliminary results to form a final multi-source partition map. The number of clusters can be set according to the number of sources that are known or automatically inferred, and the mask with the highest confidence level is selected from each cluster as the final representative feature to ensure that the final output is both accurate and reliable, realizing the process from scattered preliminary predictions to the overall accurate multi-source region segmentation map presentation.
[0052] As an example, the process of the method provided in the embodiment of the present invention is as follows:
[0053] (1) Input preparation. The user provides the image to be detected and a set of point prompts arranged in a grid. These point prompts are used to guide the model to identify the various source areas in the image. The user is required to provide point prompts during the inference phase, but it does not rely entirely on the user's expert level. The design of a set of point prompts arranged in a grid ensures that even if a single point prompt is not accurate or reasonable, the image can be fully covered as a whole, thereby ensuring the integrity and accuracy of the segmentation results. The grid-shaped point prompts ensure full coverage of the entire image, thereby improving the integrity of the segmentation results.
[0054] (2) Feature extraction. The encoder obtains the feature representation of the image. Since the calculation of image embedding is independent of the point prompt, only one encoder operation is required for each image to obtain the complete feature representation. This one-time calculation method reduces repeated operations, improves efficiency, and reduces resource consumption.
[0055] (3) Segmentation prediction. For each point cue, the mask decoder generates a preliminary segmentation mask and its confidence score based on the image embedding and the cue embedding. Each segmentation mask represents a possible source region, and the confidence score reflects the reliability of the prediction.
[0056] (4) Representative feature calculation. Based on the region determined by the preliminary mask, the average image embedding corresponding to the region is calculated as the representative feature. This step provides the basis for subsequent clustering analysis.
[0057] (5) Aggregation and clustering. Use a clustering algorithm (e.g., k-means) to integrate all preliminary results and form the final multi-source partition map. When the number of sources is known, the number of clusters is pre-set; otherwise, the algorithm is allowed to automatically infer the appropriate number. During the clustering process, the mask with the highest confidence score is selected from each cluster as the final representative feature of the cluster, ensuring that the final output is both accurate and reliable.
[0058] (6) Final output. After the above steps, one or more segmentation maps representing different source regions will be generated, which intuitively shows the true source of each part of the image.
[0059] The method for multi-source image forgery region segmentation provided by the present invention guides the model to identify the various source regions in the image by introducing a point prompt mechanism. This method allows the user or system to obtain accurate segmentation results by simply providing a few key points, greatly enhancing the interactivity and flexibility of the system. Through inter-region contrast learning pre-training, the model can better separate regions of different sources in the feature space, improving the model's understanding and capture of subtle differences. Through the area-adaptive source segmentation loss function, smaller regions are given higher weights to ensure that all regions receive sufficient attention and accurate predictions. Through confidence loss, it helps to improve the reliability of the final output and enables the model to self-evaluate the accuracy of its predictions. In addition, this method effectively solves the label agnostic problem in the forgery positioning task (i.e., it is impossible to clearly distinguish between tampered and non-tampered areas), and achieves a stable learning effect by considering the interchangeability between tampered areas and original areas in the real world.
[0060] An embodiment of the present invention further provides a control device, which may include: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above method.
[0061] An embodiment of the present invention further provides a machine-readable storage medium, on which instructions are stored, and the instructions enable a machine to execute the above method.
[0062] It should be understood that in various embodiments of the present invention, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0063] Additionally, the terms "system" and "network" are often used interchangeably. The term "and / or" is simply used to describe a relationship between related objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the character " / " generally indicates an "or" relationship between the related objects.
[0064] It should be understood that in the embodiments of the present invention, "B corresponding to A" means that B is associated with A and B can be determined based on A. However, it should also be understood that determining B based on A does not mean determining B based solely on A; B can also be determined based on A and / or other information.
[0065] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the composition and steps of each example according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0066] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0067] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, or can be electrical, mechanical or other forms of connection.
[0068] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the objectives of the embodiments of the present invention.
[0069] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0070] From the above description of the embodiments, it will be apparent to those skilled in the art that the present invention can be implemented using hardware, firmware, or a combination thereof. When implemented using software, the aforementioned functionality may be stored in a computer-readable medium or transmitted as one or more instructions or codes on the computer-readable medium. Computer-readable media include computer storage media and communication media, wherein communication media includes any medium that facilitates the transfer of computer programs from one location to another. Storage media can be any available medium that can be accessed by a computer. By way of example and not limitation, computer-readable media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer. Furthermore, any suitable connection may constitute a computer-readable medium. For example, if the software is transmitted from a website, server, or other remote source using a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of the medium. As used herein, the terms "disk" and "disc" include compact discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), floppy disks, and Blu-ray discs. Disks typically reproduce data magnetically, while discs use lasers to reproduce data optically. Combinations of the above should also be included within the scope of protection for computer-readable media.
[0071] In short, the above description is only a preferred embodiment of the technical solution of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the scope of protection of the present invention.
Claims
1. A method for multi-source image forgery region segmentation, characterized in that: The method comprises the following steps: Obtain an image to be detected and one or more corresponding point prompts arranged in a grid, and input them into a trained image recognition model, where each of the one or more point prompts is used to indicate the source area of the image to be detected. Among them, regional contrast learning is used to pre-train the image recognition model. During pre-training, a dual-region paired point sampling strategy is adopted to select points from the source region of the input image and points from other source regions. Based on the contrast loss between regions, Adjusting parameters of the image recognition model; After the pre-training, the image recognition model is trained, including: Fixing the pre-trained image encoder and keeping the prompt encoder of the image recognition model in an original state; Convert the binary labeled dataset of input images into point masks through connected component analysis; According to the point mask, a loss function is configured, wherein the loss function includes an area-adaptive source segmentation loss function and a confidence loss function; Optimizing parameters of a mask decoder of the image recognition model according to the loss function; Using the encoder of the image recognition model, operating on the image to be detected to obtain a feature representation of the image to be detected; For each point prompt, using the mask decoder, perform segmentation prediction on the feature representation of the image to be detected to obtain a prediction area and a prediction result corresponding to each point prompt; The prediction results are integrated using a preset clustering algorithm to form a multi-source partition map.
2. The method according to claim 1, characterized in that The contrast loss between the regions Expressed as: in, (·) represents the loss value, represents the number of source regions in the input image, Indicates the The feature set corresponding to the source region, represents the number of elements in the first source region, Indicates that from source regions to remove the current query vector The remaining feature set after is taken as the positive sample. Indicates that except for All other features outside the source region are regarded as negative samples.
3. The method according to claim 1, characterized in that The feature representation includes image embedding and prompt embedding, and the encoder of the image recognition model is used to operate on the image to be detected to obtain the feature representation of the image to be detected, which includes: Encoding the image to be detected by using a trained image encoder to obtain image embedding; The point prompt is encoded by the trained prompt encoder to obtain the prompt embedding.
4. The method according to claim 3, characterized in that For each of the point prompts, using the mask decoder of the image recognition model to perform segmentation prediction on the feature representation of the image to be detected, and obtaining the prediction area and prediction result corresponding to each of the point prompts, including: performing segmentation prediction on the image embedding and the cue embedding using the mask decoder to generate a corresponding segmentation mask and a confidence score for each point cue; Determine a corresponding prediction region according to the segmentation mask, and calculate an average image embedding corresponding to the region according to the determined prediction region as a representative feature; The prediction result is obtained according to the confidence score and the representative features.
5. The method according to claim 4, characterized in that The method of utilizing a preset clustering algorithm to integrate the prediction results to form a multi-source partition map includes: During the clustering process, the segmentation mask with the highest confidence score is selected from each cluster as the final representative feature of the cluster, and is integrated to form the multi-source partition map.
6. A control device, characterized in that: The control device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor executes the computer program to implement the method according to any one of claims 1 to 5.
7. A machine-readable storage medium, characterized in that The machine-readable storage medium stores instructions, which enable the machine to execute the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Image forgery detection method and device based on adaptive clustering strategy, and medium
CN118351426A
Model processing method and device, equipment, medium and product
CN118470466A