Image feature indexing method and device, equipment and medium
By performing data enhancement and masking on the original image data set, combining the encoding network and the comparison learning module, an image feature index library is built, and the error detection and missed detection problems in the recognition of highly similar targets in the prior art are solved, achieving higher recognition accuracy and efficiency.
Patent Information
- Application Number
- CN202510141898.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art is difficult to effectively distinguish when dealing with highly similar target recognition, especially in the case of partial occlusion of the foreground or partially modified image, resulting in high false detection rates and missed detection rates.
By acquiring the original image data set, multiple enhancement images are generated, and masked, inputting the encoding network for reconstruction, extracting image feature representations, and optimizing feature differences through comparing learning modules to build an image feature index library.
Accurate recognition of repeated target images is achieved, error detection rates and missed detection rates are reduced, and the accuracy and real-timeness of the search system are improved.
Smart Images

Figure CN119992257A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of artificial intelligence technology and medical health, and in particular to an image feature indexing method, device, equipment and storage medium. Background Art
[0002] In the fields of healthcare and finance, the application scenarios of image retrieval technology are gradually increasing, but the existing technology still has many shortcomings in dealing with the recognition of highly similar targets. In the field of healthcare, the retrieval and analysis of medical images need to deal with a large amount of complex and diverse image data. These images often show differences due to different shooting angles, lighting conditions, and imaging equipment, making it difficult for existing retrieval technologies to accurately extract the core features of the lesion area. In addition, medical images often contain a large amount of irrelevant background information, and the existing global feature extraction methods are difficult to focus on the lesion area, resulting in retrieval results that are easily affected by background interference and are inaccurate. At the same time, insufficient local feature extraction makes it difficult to identify subtle lesion features, affecting the accurate diagnosis of the disease.
[0003] In the financial field, image retrieval technology is mainly used for identity authentication and anti-fraud. For example, in the scenario of animal insurance, insurance companies need to verify the identity of insured animals through images to prevent duplicate insurance and insurance fraud. However, existing image retrieval technology is difficult to adapt to the uncontrollable image acquisition conditions in financial scenarios. For example, animal images uploaded by different policyholders may have differences in angle, resolution, and lighting conditions, which makes the image retrieval system have a low recognition accuracy in practical applications. In addition, animals of the same species may have highly similar appearance features, and traditional retrieval methods are difficult to effectively distinguish these individual differences, resulting in frequent false detections or missed detections.
[0004] Existing image retrieval technologies mainly include global feature retrieval methods and local feature retrieval methods. The global feature retrieval method extracts the overall features of the image for similarity comparison, but the recognition effect is significantly reduced when dealing with images with foreground occlusion and complex background. At the same time, this method ignores the key local details in the image and it is difficult to accurately distinguish highly similar target individuals. Although the local feature retrieval method can identify the detailed features in the image, it lacks the ability to understand the semantics of the entire image. When there are multiple similar local areas in the image, recognition errors are prone to occur. Summary of the invention
[0005] The main purpose of the present invention is to provide an image feature indexing method, device, equipment and storage medium, aiming to solve the technical problem that the prior art is difficult to effectively distinguish highly similar image targets in repeated target recognition, especially when the foreground is partially occluded or the image is partially modified, resulting in high false detection rate and missed detection rate.
[0006] To achieve the above object, the present invention provides an image feature indexing method, comprising:
[0007] Acquire an original image data set, perform a data enhancement operation on each original image in the original image data set, and generate a plurality of enhanced images corresponding to each original image;
[0008] Mark all enhanced images generated from the same original image as similar images, and mark enhanced images generated from different original images as dissimilar images;
[0009] Perform masking on a partial image area of each enhanced image, and input the masked image into the encoding network;
[0010] Reconstructing the masked image region through the encoding network to generate a corresponding reconstructed image, wherein the reconstructed image retains similarity marks or dissimilarity marks of the enhanced image;
[0011] extracting image feature representation from the reconstructed image through an encoding network;
[0012] Through the contrast learning module, the difference between the image feature representations is reduced for the reconstructed images marked as similar images, and the difference between the image feature representations is increased for the reconstructed images marked as dissimilar images;
[0013] According to the processed image feature representation, an image feature index library is constructed.
[0014] Furthermore, to achieve the above object, the present invention provides an image feature indexing device, comprising:
[0015] An image marking module, used for marking all enhanced images generated from the same original image as similar images, and marking enhanced images generated from different original images as dissimilar images;
[0016] A mask processing module, used for performing mask shielding processing on a partial image area of each enhanced image, and inputting the mask shielded image into the encoding network;
[0017] An image reconstruction module, used to reconstruct the masked image area through the encoding network to generate a corresponding reconstructed image, wherein the reconstructed image retains similarity marks or dissimilarity marks of the enhanced image;
[0018] A feature extraction module, used for extracting image feature representation from the reconstructed image through an encoding network;
[0019] A contrastive learning module, used for reducing the difference between image feature representations for reconstructed images marked as similar images and increasing the difference between image feature representations for reconstructed images marked as dissimilar images through the contrastive learning module;
[0020] The feature index module is used to construct an image feature index library based on the processed image feature representation.
[0021] Furthermore, to achieve the above-mentioned purpose, the present invention also provides a computer device, which includes a memory, a processor, and an image feature indexing program stored in the memory and executable on the processor, wherein the image feature indexing program implements the steps of the image feature indexing method described above when executed by the processor.
[0022] Furthermore, to achieve the above-mentioned purpose, the present invention also provides a computer-readable storage medium, on which an image feature indexing program is stored, and when the image feature indexing program is executed by a processor, the steps of the image feature indexing method described above are implemented.
[0023] Beneficial effects: The present invention relates to the fields of artificial intelligence technology and medical health, and discloses an image feature indexing method, including: obtaining an original image data set and performing a data enhancement operation, marking the enhanced image as a similar image or a dissimilar image; performing a mask shielding process on the enhanced image and inputting it into a coding network to generate a reconstructed image; extracting image feature representation from the reconstructed image, optimizing the feature differences between similar images and dissimilar images through contrastive learning; constructing an image feature index library based on the processed image feature representation to achieve retrieval and recognition of repeated target images. The present invention can accurately identify repeated target images by comprehensively extracting the global features and local features of the image. The self-supervised learning method reduces the dependence on manually labeled data and improves the robustness and generalization ability of the image feature representation. The contrastive learning module optimizes the difference in image features and improves the system's recognition accuracy for similar and dissimilar images. It can effectively solve the problem of identifying repeated target images, significantly reduce the false detection rate and missed detection rate, and improve the accuracy and real-time performance of the retrieval system. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The present invention will be further described below with reference to the accompanying drawings and embodiments, in which:
[0025] Figure 1 A schematic diagram of an application environment of an image feature indexing method according to an embodiment of the present invention;
[0026] Figure 2 A schematic diagram of a flow chart of an embodiment of an image feature indexing method of the present invention;
[0027] Figure 3 A schematic diagram of functional modules of a preferred embodiment of the image feature indexing device of the present invention;
[0028] Figure 4 A schematic diagram of the structure of a computer device in one embodiment of the present invention;
[0029] Figure 5 FIG. 4 is another schematic diagram of the structure of a computer device in one embodiment of the present invention. DETAILED DESCRIPTION
[0030] It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.
[0031] The image feature indexing method provided by the embodiment of the present invention can be applied in the following aspects: Figure 1 In the application environment, the user end communicates with the server end through the network. The server end can obtain the original image data set through the user end and perform data enhancement operation, mark the enhanced image as a similar image or a dissimilar image; perform mask shielding processing on the enhanced image and input it into the encoding network to generate a reconstructed image; extract image feature representation from the reconstructed image, optimize the feature difference between similar images and dissimilar images through contrast learning; build an image feature index library based on the processed image feature representation to achieve retrieval and recognition of repeated target images. The present invention can accurately identify repeated target images by comprehensively extracting global features and local features of the image. The self-supervised learning method reduces the dependence on manually labeled data and improves the robustness and generalization ability of image feature representation. The contrast learning module optimizes the difference of image features and improves the recognition accuracy of the system for similar images and dissimilar images. It can effectively solve the recognition problem of repeated target images, significantly reduce the false detection rate and missed detection rate, and improve the accuracy and real-time performance of the retrieval system. Among them, the user end can be but not limited to various personal computers, laptops, smart phones, tablet computers and portable wearable devices. The server end can be implemented by an independent server or a server cluster composed of multiple servers. The present invention is described in detail below through specific embodiments.
[0032] See also Figure 2 , Figure 2 This is a flow chart of an embodiment of an image feature indexing method provided by the present invention. It should be noted that although a logical sequence is shown in the flow chart, in some cases, the steps shown or described may be performed in a different order than that shown here.
[0033] like Figure 2 As shown, the image feature indexing method proposed by the present invention comprises the following steps:
[0034] S10, obtaining an original image data set, performing a data enhancement operation on each original image in the original image data set, and generating a plurality of enhanced images corresponding to each original image;
[0035] In this embodiment, the raw image dataset refers to a collection of multiple unprocessed images, which is used for subsequent data enhancement and model training. The dataset can contain images from different sources, such as photos taken by different devices, image files in different formats, or historical image data extracted from a database. The content of these image data may include face images, animal images, medical images, etc. depending on the specific application scenario. In the scenario of repeated target recognition, the raw image dataset is the basic data, which contains the images of the objects to be recognized.
[0036] Ways to obtain raw image data sets include extracting image data from external databases through API interfaces, importing image data by uploading files, or directly loading image data from existing image resource libraries. In practical applications, when obtaining data, attention should be paid to the compatibility of image formats, image clarity requirements, and compliance of data sources. For example, in the scenario of animal insurance, animal images taken by farmers can be uploaded through mobile phone applications and used as part of the raw image data set.
[0037] Data augmentation refers to the process of generating multiple versions of images by performing multiple transformations or perturbations on the original image to expand the size of the data set and improve the generalization ability of the model. Data augmentation can simulate the effects of different shooting conditions or image transformations without changing the semantics of the image, helping the model to better adapt to complex scenes. Common data augmentation operations include image cropping, rotation, flipping, brightness adjustment, and noise addition.
[0038] When applying data augmentation to each original image, two strategies can be used: random augmentation and fixed augmentation. Random augmentation refers to randomly selecting multiple augmentation methods and randomly setting augmentation parameters to generate multiple different versions of images. For example, randomly cropping a part of the original image, or randomly adjusting the brightness and contrast of the image. Fixed augmentation is to perform consistent transformations on each image according to a preset augmentation strategy, such as rotating each image 90 degrees or mirror-flipping it. In the implementation process, image processing libraries (such as OpenCV, Pillow, or TensorFlow's image processing module) are usually used to complete data augmentation operations.
[0039] Through data augmentation operations, each original image generates multiple augmented images, which are versions of the original image under different transformation conditions. The augmented images retain the semantic information of the original image, but show differences at the pixel level. These differences can include changes in image brightness, contrast, rotation angle, cropping position, etc. The purpose of this augmentation is to enable the model to recognize images that still belong to the same object under different shooting conditions, thereby improving the recognition robustness of the model.
[0040] In order to generate multiple enhanced images, batch processing can be used to input each original image into the data enhancement module and output multiple enhanced images. For example, after rotating the original image A, adding noise and adjusting the brightness, enhanced images A1, A2 and A3 are generated. These enhanced images can be stored in the file system or database, or directly passed as input data to the subsequent model training process. In the specific implementation, loop processing or parallel processing is usually used to improve the efficiency of data enhancement to ensure that a large number of enhanced images are generated in a short time.
[0041] Example description: In the pathological image recognition scenario in the medical and health field, doctors often need to use medical images to determine the nature of the lesion area. However, the lesion areas of different patients may show large differences due to different shooting equipment and shooting angles. By performing data enhancement operations on the images of the lesion area, multiple images with different feature changes can be generated, such as images of lesion areas with different brightness and different angles. These enhanced images can help the model better identify the lesion area, improve the accuracy of automatic diagnosis, and reduce the workload of doctors when manually comparing images. For example, in the diagnosis of skin diseases, the original image may be a erythema on the patient's arm. Through data enhancement operations, images of erythema with lower brightness and partially occluded erythema can be generated. These enhanced images can help the model better identify the characteristics of erythema, and even in the actual scene, when there is insufficient light or occlusion, the lesion area can be effectively identified. This method reduces the dependence on manual annotation and improves the recognition robustness of the model under different acquisition conditions.
[0042] Similarly, in the insurance scenario of livestock insurance, insurance companies need to verify the identity of the insured livestock to prevent the insured from repeatedly taking photos of the same livestock for repeated insurance or insurance fraud. Due to the large differences in image acquisition conditions among different farmers, livestock pictures may have different angles, lighting, and resolutions, and the insured may even manually modify or block the features of the livestock to take images. For example, the insured may take photos of the same cow from different angles, or mark or paint on the cow's body, attempting to deceive the insurance system through these image differences.
[0043] By performing data enhancement operations on the original image dataset, the recognition system's ability to distinguish highly similar images can be effectively improved. For example, the system performs enhancement operations such as cropping, rotating, and adjusting the brightness of an original photo of a cow to generate multiple enhanced images with different feature changes. These enhanced images can help the model learn the changes in the appearance features of livestock under different shooting conditions, thereby improving the generalization ability of the recognition system.
[0044] In a specific application, suppose that an original photo of a cow generates multiple enhanced images from different angles after data enhancement. The system can mark these images as similar images through the comparative learning module, and automatically identify whether the images submitted by the insured are duplicate photos of the same livestock during the actual insurance verification process. In this way, even if the insured tries to take photos of the same livestock from different angles or in different environments, the system can recognize the similarities of these images and effectively prevent duplicate insurance and insurance fraud. Through data enhancement and image feature extraction technology, insurance companies can greatly improve the accuracy and timeliness of image verification without increasing labor costs, thereby realizing automated anti-fraud and reducing the insurance company's claims risk.
[0045] By enhancing the original image dataset, a large amount of image data with different transformation effects can be generated without changing the semantic information of the original image, thereby effectively expanding the scale of the training dataset and improving the generalization ability of the model. Compared with the training method that relies only on the original image, it can effectively deal with image recognition problems under different shooting conditions, significantly improve the recognition accuracy of highly similar image targets, and reduce the false detection rate and missed detection rate.
[0046] S20, marking all enhanced images generated from the same original image as similar images, and marking enhanced images generated from different original images as dissimilar images;
[0047] In this embodiment, multiple images generated from the same original image after data augmentation may visually show different feature changes, such as differences in brightness, rotation angle, and cropping position, but these enhanced images still come from the same object and are therefore marked as similar images. The marking process of similar images is to help the model learn the feature consistency of the same object under different transformation conditions, thereby improving the model's ability to recognize the same target and enhancing the robustness of images under different shooting conditions.
[0048] In the implementation process, a unique identifier (ID) can be automatically generated for each original image through the data enhancement module, and all enhanced images generated from the original image can be uniformly marked as similar images. For example, the enhanced images A1, A2, and A3 generated for the original image A are all given the same identifier ID_A and stored in the similar image set. The labeling process can be managed through the label field of the database or implemented through the input label during model training.
[0049] Different original images may represent different objects, and the images generated after data enhancement may also show similar visual features, but these images essentially belong to different objects. Therefore, it is necessary to mark the enhanced images generated from different original images as dissimilar images to help the model learn how to distinguish the feature differences between different targets. The labeling process of dissimilar images can effectively improve the accuracy of the model when dealing with differential recognition between similar objects.
[0050] In practical applications, the enhanced images of different original images can be marked as dissimilar images by generating a unique identifier for each original image and ensuring that the identifiers of different original images are different. For example, the enhanced images A1, A2, B1, and B2 generated by the original image A and the original image B, respectively, where A1 and B1 are marked as dissimilar images. This labeling process is usually implemented by constructing positive and negative sample pairs, and the positive sample pairs (similar images) and negative sample pairs (dissimilar images) are input during model training to guide the model to optimize the accuracy of feature extraction.
[0051] By marking similar images and dissimilar images, the model is effectively guided to learn how to distinguish different images of the same object and images of different objects, improving the model's recognition ability when facing highly similar targets. It can automatically generate positive and negative labels for training samples, reducing dependence on manually labeled data, while improving the robustness and generalization ability of the model under different acquisition conditions and transformation conditions.
[0052] S30, performing masking processing on a partial image area of each enhanced image, and inputting the masked image into the encoding network;
[0053] In this embodiment, masking refers to covering some areas of the image so that the pixel values of these areas are replaced with fixed values or random values, thereby simulating the information missing scene of the image. The purpose of masking is to let the model learn how to accurately extract effective features when the image is incomplete or some information is missing. This method can improve the robustness of the model, so that it can still maintain a high recognition accuracy when facing image occlusion, damage or incompleteness in real scenes.
[0054] The mask area can be determined by a randomly generated mask template, which identifies the area in the image that needs to be masked. The mask ratio is usually a fixed preset ratio or a randomly selected dynamic ratio, such as masking 20% to 60% of the image. The masked area can be a continuous block (such as a randomly cropped area) or a non-continuous block (such as multiple small masked blocks).
[0055] Generate a random mask template to determine the area in the image that needs to be masked; apply the mask template to mask the specified area of each enhanced image, and replace the pixel values of these areas with zero values, random noise values, or other fixed values; mask processing can be implemented through image processing tools or frameworks, such as OpenCV, Pillow, TensorFlow, etc. In practical applications, the size and position of the mask area can be adjusted according to different image types and task requirements. For example, in animal recognition tasks, mask processing can be performed on the head or body parts to allow the model to learn to recognize targets when different parts are blocked.
[0056] The masked image contains masked areas and unmasked areas. These images need to be input into the encoding network to extract the feature representation of the image. The encoding network is usually a neural network based on deep learning, such as a convolutional neural network (CNN) or a Transformer network, which is used to convert the image into a feature vector. These feature vectors can represent the overall semantics and local detail information of the image, providing basic data for subsequent reconstruction learning and feature comparison.
[0057] The masked image is input to the input layer of the encoding network; the encoding network processes the image in blocks and divides the image into multiple small blocks (such as a sequence of image blocks); the encoding network extracts the global and local features of the image layer by layer and generates a feature vector representation of the image. In practical applications, you can choose an adaptive encoding network structure, such as ResNet, EfficientNet, ViT (Vision Transformer), etc., and adjust the depth and width of the network according to the complexity of the task and the requirements of feature extraction.
[0058] Example description: In the medical image analysis in the field of healthcare, doctors often need to identify lesions in CT images or MRI images. However, there may be partially missing or occluded areas in medical images. For example, due to the shooting angle, changes in patient position or limitations of image acquisition equipment, some information of the lesion area is missing. By masking the enhanced image of each medical image, the missing information can be simulated, allowing the model to learn how to infer the features of the occluded area from the unoccluded area. For example, in the analysis of lung CT images, the small lung nodule area of the image can be randomly masked, and the training model can automatically identify the complete features of the missing area in actual detection. This processing method is particularly important for early lesion detection. For example, in a patient's CT image, the lesion area may be partially occluded or missing. Through masking, the model can learn the potential features of the lesion, thereby improving the recognition rate of early lesions and providing effective auxiliary support for doctors' diagnosis and treatment. Through this mask shielding processing, the medical image recognition system can adapt to different image defect situations in practical applications, improve the recognition accuracy of incomplete images, effectively reduce missed detections and false detections, reduce the workload of doctors, and improve the efficiency and accuracy of medical diagnosis.
[0059] Similarly, in the insurance business in the financial sector, especially in the underwriting and claims process of livestock insurance and vehicle insurance, insurance companies usually need to verify pictures of the insured, such as livestock photos, vehicle photos, etc. However, due to different shooting environments, angles, and lighting conditions, insurance companies often encounter images with missing or blocked information, which affects the accuracy of image verification. In addition, the insured may create duplicate insurance or fraudulent insurance images by partially blocking or manually modifying them, such as applying marks on the same livestock, covering ear tags, etc., in an attempt to create different insured objects.
[0060] By masking some areas of each enhanced image, the image verification system of insurance companies can adapt to scenarios with missing or blocked information and improve the ability to identify insured objects. For example, when verifying livestock insurance, the head marks, ear tags or body features in the image are randomly masked, and the training model can still recognize repeated photos of the same livestock even when some information is missing, effectively preventing duplicate insurance and insurance fraud. Specific application scenarios may include the following insurance scenarios:
[0061] Livestock insurance: The insured may take photos of the same cow from different angles, or even mark or cover the cow to create different insurance photos. Through masking, the model can learn how to identify the same livestock when some information is missing, preventing duplicate insurance.
[0062] Vehicle insurance: During the claims process, the insured may submit multiple claims for the same vehicle by covering the license plate or vehicle body details. Through masking, the system can still identify the photo of the same vehicle even when the license plate is covered or part of the vehicle body is missing, reducing the risk of duplicate claims.
[0063] Crop insurance: When applying for crop insurance, farmers may submit partially obscured or modified crop photos in an attempt to duplicate insurance or defraud compensation. Through masking, the model can identify missing or obscured crop features, effectively improving the insurance company's underwriting and claims efficiency.
[0064] By masking some areas of the enhanced image, the information loss and occlusion scenarios in the actual image acquisition process are effectively simulated, and the model's ability to recognize incomplete images is improved. At the same time, by inputting the masked image into the encoding network to extract features, the model can focus more on the unobstructed areas of the image, thereby improving the model's ability to extract effective features. The model's dependence on irrelevant background information is reduced, and the recognition accuracy and robustness of the model in complex environments are improved.
[0065] S40, reconstructing the masked image area through the encoding network to generate a corresponding reconstructed image, wherein the reconstructed image retains similarity marks or dissimilarity marks of the enhanced image;
[0066] In this embodiment, reconstruction processing refers to the model completing the information of the masked area in the image, so that the masked part can regenerate close to real content through the learning process of the model. During the image reconstruction process, the encoding network will infer the features of the masked area based on the information of the uncovered image area, and generate a reconstructed feature to fill the missing part.
[0067] The core purpose of reconstruction processing is to let the model learn how to infer missing information when the image information is incomplete, so as to improve the recognition ability and robustness of the model when facing partially occluded or defective images. This reconstruction technology is crucial to solving the recognition problem of incomplete images in real scenes.
[0068] The input masked image is fed into the encoding network; the encoding network analyzes the uncovered areas in the image by extracting and mapping features layer by layer; the encoding network completes the masked areas based on the extracted global and local features to generate reconstructed features; the reconstructed features are filled into the masked areas to generate a reconstructed image.
[0069] The reconstructed image refers to the image after the masked area is completed through the reconstruction process of the encoding network. The reconstructed image contains not only the original uncovered part, but also the reconstructed part generated by the model. These reconstructed images can be used for subsequent contrast learning and image feature indexing to help the model further optimize the feature representation of the image.
[0070] The reconstructed features generated by the encoding network are fused with the uncovered image parts of the original image. The generated reconstructed image is subject to consistency verification to ensure that the features of the reconstructed area are consistent with those of the uncovered part of the original image. The reconstructed image can be used as input for model training to help the model learn how to complete partially occluded images in different scenarios.
[0071] During the data augmentation and masking process, each image is assigned a similar or dissimilar label to distinguish different images of the same object and images of different objects. The reconstructed image needs to inherit the label information of the enhanced image to ensure that the model can correctly identify similar and dissimilar images in the contrastive learning module, thereby optimizing the accuracy of feature extraction and recognition.
[0072] In the process of reconstructing the image generation, the label information of the input image is inherited; in the comparative learning stage of the model, the label information is used to distinguish between similar image sets and dissimilar image sets; the retention of label information can be managed through the label field of the database or the input label of the model.
[0073] Example description: In the medical image analysis in the field of healthcare, doctors usually need to identify and diagnose lesions on patients' medical images (such as CT, MRI, X-rays, etc.). However, actual medical image data may be partially missing or image occluded due to acquisition conditions, changes in patient position, equipment noise and other factors. In this case, traditional image recognition models are difficult to effectively analyze incomplete image data, and are prone to problems such as missed detection and false detection. In order to improve the robustness of the medical image analysis system, masking is performed on some areas to simulate image defects or occlusion scenes, and reconstruction learning is performed through the encoding network to generate a reconstructed image with complete features. In this way, even in the case of incomplete image data, the model can still accurately identify the lesion area and improve the accuracy and reliability of medical image recognition.
[0074] In medical image analysis, masking can be used to simulate information loss or occlusion scenarios, such as masking small nodule areas in the lungs in CT images or masking the edge of tumors in brain MRI. This can help the model learn how to extract effective lesion features when some information is missing.
[0075] Lung CT analysis: Small nodules in lung images are masked, and the training model can still recognize the complete lesion morphology when the nodule is partially covered.
[0076] Brain MRI analysis: Randomly mask tumor edges or hemorrhage areas in brain images and let the model learn how to infer the complete information of the missing areas.
[0077] After the masking process, the model needs to reconstruct the masked area through the encoding network. This step can help the model infer the missing information, thereby generating a reconstructed image with complete features, enabling the model to accurately identify lesion features even when the image is partially missing.
[0078] Early screening for lung cancer: Reconstruct partially obscured small nodule areas, allowing the model to infer the size and shape of the obscured nodule areas based on the characteristics of the uncovered lung tissue, thereby achieving automatic screening for early lung cancer.
[0079] Stroke detection: Reconstruct some hemorrhage areas in brain CT images, so that the model can still accurately determine the location and area of the stroke when there is noise or defects in the image.
[0080] After the encoding network completes the reconstruction of the masked area, the generated reconstructed image contains the overall medical imaging information and local lesion details, which enables the model to further optimize the accuracy of feature extraction and lesion identification during the contrastive learning process.
[0081] Skin Lesion Detection: For partially covered skin lesion images, the model can reconstruct the covered lesion area based on the color and texture features of the surrounding uncovered areas, thereby accurately identifying the lesion type and severity.
[0082] Breast cancer screening: In breast X-ray images, certain areas are masked to simulate image defects, and the model is trained to identify potential tumors or calcifications after reconstructing the image, helping doctors make early screening decisions.
[0083] By reconstructing and learning the masked image areas, the model can accurately infer the features of the missing areas when some information is missing or the image is incomplete, thereby improving the accuracy of image recognition. The reconstructed image retains the similar or dissimilar marks of the original image, ensuring that the model correctly optimizes the differences in image features during the comparative learning process, thereby effectively improving the recognition accuracy and robustness of the model in repeated target recognition tasks.
[0084] S50, extracting image feature representation from the reconstructed image through an encoding network;
[0085] In this embodiment, image feature representation refers to converting the visual information of the image into a feature vector that the model can understand and process. The reconstructed image is input into the encoding network, which converts the global information (overall structure) and local information (detail features) in the image into feature representations. These feature representations can help the model quickly and accurately match similar images or distinguish different images in tasks such as image retrieval, target recognition, and repeated target determination.
[0086] Input the reconstructed image to the encoding network; the encoding network divides the reconstructed image into a sequence of image blocks and converts each image block into a feature vector; extracts the global feature representation to represent the overall structural information of the image; extracts the local feature representation to capture the detailed features of the image; fuses the global features and local features to generate an image feature representation that contains both overall information and detailed information.
[0087] Example description: In medical image analysis, extracting global and local feature representations of images is crucial for identifying lesions and abnormalities. For example, in brain MRI images, global features can help identify the overall structure of the brain, and local features can capture the edges and shape details of the tumor.
[0088] In lung CT images, by extracting features from reconstructed images, the model can simultaneously capture the structural information of the entire lung and the detailed features of local nodules. This helps the model to accurately identify the size, shape and location of small nodules even when there are partial defects or noise interference, thus improving the accuracy of early lung cancer screening.
[0089] In skin lesion images, global feature representation can capture the overall distribution of the lesion area, while local feature representation can capture details such as the texture and edge of the lesion. Through the feature extraction process, the model can identify different types of skin lesions, such as melanoma, psoriasis, eczema, etc., to assist doctors in early diagnosis.
[0090] In the livestock insurance business, the insured may submit photos of the same livestock taken from multiple angles and in different environments in an attempt to create duplicate insurance or insurance fraud. By extracting features from the reconstructed image, the system can extract the overall features (such as body shape and color) and local features (such as ear tags and fur texture) of the livestock. These feature representations can help the system identify duplicate photos of the same livestock even when occluded or missing information, effectively reducing the risk of duplicate claims for insurance companies.
[0091] By extracting features from reconstructed images, the model can obtain complete feature representations from incomplete or partially occluded images, improving the accuracy and robustness of the model in tasks such as image retrieval and repeated target recognition. It can capture both the overall features and local details of the image, effectively solving the recognition problem when the image information is incomplete in complex environments or acquisition conditions.
[0092] S60, reducing the difference between the image feature representations for the reconstructed images marked as similar images and increasing the difference between the image feature representations for the reconstructed images marked as dissimilar images through a comparative learning module;
[0093] In this embodiment, during the process of image data enhancement, mask shielding processing, and encoding network reconstruction, the labeling information of each image (similar labels or dissimilar labels) will run through the entire data stream and be transmitted to the reconstructed image.
[0094] Similar images: Multiple enhanced images generated from the same original image are labeled as similar images.
[0095] Dissimilar images: The enhanced images generated from different original images are labeled as dissimilar images.
[0096] During the reconstruction process, the model completes the masked areas of the image, but the unmasked areas still retain the main visual features of the enhanced image. Therefore, the reconstructed image can directly inherit the tags of the enhanced image without re-labeling.
[0097] In the image data enhancement stage, corresponding similar or dissimilar labels are generated for each image; in the reconstruction stage, the reconstructed image completes the masked areas through the encoding network while retaining the unmasked area information; the generated reconstructed image inherits the labels of the original enhanced image and is used as input data in the subsequent contrastive learning module.
[0098] Contrastive Learning is an unsupervised learning method based on the comparison of feature similarities and differences. Its core goals are:
[0099] For similar images, the model needs to minimize the difference in feature representation between them, making them closer in the feature space;
[0100] For dissimilar images, the model needs to maximize the difference in feature representation between them, making them farther away in the feature space.
[0101] Feature extraction is performed on each reconstructed image to generate global feature vectors and local feature vectors; for the reconstructed images marked as similar images, the cosine similarity between them is calculated; for the reconstructed images marked as dissimilar images, the cosine similarity between them is also calculated.
[0102] For similar images, the model reduces the difference in feature representation by minimizing the cosine distance loss function; for dissimilar images, the model increases the difference in feature representation by maximizing the cosine distance loss function.
[0103] For similar images (i.e. enhanced images generated from the same original image), the goal of the model is to make the feature representations of these images as similar as possible, thereby improving the clustering ability of the model and making it easier for the model to better recognize different images of the same object.
[0104] For each group of similar images, calculate the cosine similarity between them; use a loss function (such as InfoNCELoss) to minimize the difference in feature representations of these images; during training, the model automatically adjusts parameters so that the feature representations of similar images are closer in the feature space.
[0105] For dissimilar images (i.e., enhanced images generated from different original images), the goal of the model is to make the feature representations of these images as different as possible, thereby improving the model's ability to distinguish and facilitating the model to better recognize images of different objects.
[0106] For each group of dissimilar images, the cosine similarity between them is calculated; a loss function is used to maximize the difference in the feature representations of these images; during training, the model automatically adjusts parameters so that the feature representations of dissimilar images are further apart in the feature space.
[0107] By performing comparative learning optimization on similar and dissimilar images, the model's feature differentiation ability is effectively improved. By reducing the feature differences of similar images and increasing the feature differences of dissimilar images, the model can better cluster similar objects and accurately distinguish different objects. It can more accurately process complex image scenes and improve the accuracy and robustness of the model in repeated object recognition tasks.
[0108] S70, constructing an image feature index library according to the processed image feature representation.
[0109] In this embodiment, the image feature index library is a data structure for storing and managing image feature representations. By constructing the index library, efficient image retrieval, repeated target recognition, and similarity comparison operations can be achieved. The image feature representation extracted by the encoding network will be stored in the index library. The construction process of the index library includes not only the storage of feature data, but also the optimization of the index structure to support subsequent rapid retrieval and comparison.
[0110] The core goal of building an image feature index library is to improve the efficiency and accuracy of image retrieval, especially in the scenario of large-scale data sets. The index library can quickly locate image collections similar to the target image.
[0111] The image feature representations extracted and processed by the encoding network are stored, and the feature vector of each image is associated with its image identifier (such as ID or file name) to ensure that the original image can be quickly located during retrieval.
[0112] Generate a unique image identifier for each reconstructed image; extract the global feature vector and local feature vector of the image; associate each feature vector with the corresponding image identifier and store them in an index library.
[0113] When storing image feature representations in an index library, the data structure of the index library needs to be optimized to ensure fast positioning and similarity comparison during subsequent retrieval.
[0114] Select a suitable index structure (such as inverted index, KD tree, ANN index, etc.) to improve retrieval efficiency; bucket, sort and cluster the feature data in the index library to reduce the computational overhead during retrieval; dynamically update the index library, and perform real-time maintenance and optimization of the index library as new image feature representations are added.
[0115] The key goal of building an index library is to achieve fast retrieval and similarity comparison. By choosing a suitable index structure, you can quickly find the image that is closest to the target image in a large-scale dataset.
[0116] Use inverted index to store feature vectors in blocks to improve retrieval speed; use algorithms such as Locality Sensitive Hashing (LSH) to bucket the index library to support fast similarity comparison; deduplicate and denoise the feature representations in the index library to remove redundant or abnormal feature vectors.
[0117] Example description: In the field of medical health, the analysis and diagnosis of medical images (such as CT, MRI, X-ray, etc.) usually requires a large amount of image data to be used for lesion comparison, similar case retrieval, and lesion trend analysis. In order to improve the retrieval efficiency of image data and the accuracy of diagnosis, an image feature index library can be constructed to organize the image feature representation of medical images into an efficient retrieval structure to support fast similar case search and lesion feature analysis.
[0118] Specifically, in the medical image processing process, the system extracts global features (such as the overall morphology of organs) and local features (such as the shape, edge, texture, etc. of lesions) from each medical image, and stores these image feature representations in the index library. The construction process of the index library includes the storage of feature data, optimization of the index structure, and dynamic updating to ensure that historical cases with similar features to the target image can be quickly retrieved later.
[0119] In actual applications, when doctors upload new patient images, the system will quickly locate cases similar to the target images from the image feature index library and refer to the diagnosis and treatment plans of these cases. This effectively solves the problem of low efficiency of one-by-one comparison and difficulty in quickly locating similar cases in the traditional image comparison process, improves the retrieval and analysis capabilities of medical images, and provides technical support for precision medicine.
[0120] In the insurance business in the financial field, especially in the claims process of livestock insurance, the problem of duplicate insurance and duplicate claims is often encountered. For example, the same livestock may be submitted multiple insurance photos by the insured at different times and locations with different shooting angles or environments, attempting to duplicate insurance or defraud insurance. The traditional manual review method mainly relies on human eye recognition, which is not only time-consuming and labor-intensive, but also prone to missed inspections due to fatigue, and cannot meet the business needs of large-scale image review.
[0121] In order to improve the efficiency and accuracy of repeated target recognition, an image feature index library can be constructed to store the feature representations of livestock images submitted by the insured in the index library. The index library construction process includes extracting global and local features of the image, storing feature representations, and optimizing the index structure to achieve fast retrieval and comparison.
[0122] Specifically, each time a new insurance application is submitted, the system extracts the feature representation of the target image and compares it with the existing feature representations in the image feature index library. Based on the comparison results, the system can quickly determine whether there are identical or highly similar images and mark them as possible duplicate insurance behaviors. In addition, the index library will be dynamically updated to ensure that the system can promptly incorporate new image features into the index library when processing new insurance images to maintain the latest feature data.
[0123] For photos of the same livestock taken from different angles, the system can quickly identify multiple insurance records of the same livestock through index library comparison and issue duplicate insurance warnings to avoid losses caused by duplicate claims.
[0124] In addition, by using the historical image feature representation in the index library, the system can quickly compare newly submitted claim photos to determine whether the livestock has had similar claims records. If a high-similarity historical claim image is detected, the system will automatically mark it as a high-risk claim application, prompting the insurance company to conduct further manual review.
[0125] By building an image feature index library, we can achieve fast retrieval and similarity comparison in large-scale image data sets, effectively improving the model's ability to identify repeated targets and image retrieval efficiency. It can significantly reduce computing overhead and improve the system's response speed in real-time scenarios. In addition, the dynamic update mechanism of the index library can ensure that the system always maintains the latest feature data, thereby improving recognition accuracy and robustness.
[0126] The present invention relates to the fields of artificial intelligence technology and medical health, and discloses an image feature indexing method, including: obtaining an original image data set and performing a data enhancement operation, marking the enhanced image as a similar image or a dissimilar image; performing a mask shielding process on the enhanced image and inputting it into a coding network to generate a reconstructed image; extracting image feature representation from the reconstructed image, optimizing the feature differences between similar images and dissimilar images through contrastive learning; constructing an image feature index library based on the processed image feature representation to achieve retrieval and recognition of repeated target images. The present invention can accurately identify repeated target images by comprehensively extracting global features and local features of the image. The self-supervised learning method reduces the dependence on manually labeled data and improves the robustness and generalization ability of the image feature representation. The contrastive learning module optimizes the difference in image features and improves the system's recognition accuracy for similar and dissimilar images.
[0127] In one embodiment, the above S10 includes:
[0128] S101, obtaining an original image data set from an image data source, where the original image data set includes a plurality of images;
[0129] S102, performing image size adjustment, format conversion and / or color space standardization processing on each original image in the original image data set;
[0130] S103, based on a preset enhancement strategy, randomly selecting a plurality of data enhancement methods for each original image, wherein the data enhancement methods include cropping, rotating, flipping, adjusting brightness, and adding noise;
[0131] S104, applying the selected multiple data enhancement methods to each original image to generate multiple enhanced images with different feature changes.
[0132] In this embodiment, the image data source is the basis for constructing the image dataset, which usually comes from a local database, cloud storage, real-time acquisition equipment or a third-party image service. The original image dataset needs to contain multiple images, which are the original inputs for the system to perform image enhancement and subsequent processing.
[0133] The image data source can be the image repository within the insurance company, the medical imaging system of the hospital, or images collected by real-time photography equipment. The system obtains images in batches from the data source and generates a unique image identifier for each image for subsequent data association and management. The transmission process of image data must ensure data integrity and security, and encrypted transmission protocols can be used to prevent data leakage.
[0134] The purpose of preprocessing is to improve the quality and consistency of image data and provide standardized input for subsequent data enhancement and feature extraction. Preprocessing operations include adjusting image size, format conversion, and color space standardization.
[0135] Adjust image size: Unify images of different resolutions into a fixed size to reduce model performance fluctuations caused by image size differences.
[0136] Format conversion: Convert image formats to standardized image formats (such as JPEG, PNG, etc.) to ensure image compatibility between different systems.
[0137] Color space standardization: Convert the image's color space to RGB or grayscale to eliminate color differences caused by different acquisition devices.
[0138] Data augmentation is to enrich the training samples of the model and improve the generalization and robustness of the model. According to the preset augmentation strategy, the system will randomly select multiple augmentation methods for each original image, including cropping, rotation, flipping, adjusting brightness, and adding noise.
[0139] Cropping: Randomly crop an area of the image to simulate different shooting angles or fields of view.
[0140] Rotation: Randomly rotate the image by a certain angle to simulate the change of image orientation.
[0141] Flip: Flip the image horizontally or vertically to enhance the model's ability to recognize symmetrical images.
[0142] Adjust Brightness: Randomly adjust the brightness of the image to simulate different lighting conditions.
[0143] Add noise: Add Gaussian noise or salt and pepper noise to the image to enhance the model's noise resistance.
[0144] By applying a variety of data augmentation methods, multiple enhanced images can be generated from each original image. These enhanced images have different feature changes, thereby improving the model's adaptability to different scenes and different acquisition conditions.
[0145] Combination of multiple enhancement methods: The system applies a combination of multiple enhancement methods to each original image based on a random enhancement strategy.
[0146] Storage and management of enhanced images: The generated enhanced images are stored in association with the identifiers of the original images to ensure that the source of each enhanced image can be traced.
[0147] Verification of enhanced images: Perform quality verification on the generated enhanced images to ensure that they meet the expected feature change effects.
[0148] This embodiment generates multiple enhanced images with different feature changes through data enhancement operations, which effectively improves the adaptability of the model in different scenarios and acquisition conditions, and enhances the generalization ability and robustness of the model.
[0149] In one embodiment, the above S30 includes:
[0150] S301, randomly generating a mask template based on a preset mask area ratio, where the mask template is used to identify a mask area in an enhanced image;
[0151] S302, applying the mask template to each enhanced image to perform mask processing on the mask area identified by the mask template in each enhanced image, and generating an image after mask processing;
[0152] S303: Input the masked image into a coding network.
[0153] In this embodiment, the input enhanced image is masked and shielded. By covering part of the image area, the encoding network is forced to infer the masked area during the learning process, thereby improving the model's image reconstruction and feature extraction capabilities.
[0154] The mask template is a binary matrix used to identify the areas that need to be masked in the image. In the mask template, 1 represents the area that needs to be masked, and 0 represents the uncovered area. The process of generating the mask template needs to consider the proportion of the mask area, that is, how much of the total pixels of the image will be randomly masked. By randomly generating mask templates, different degrees of image missing can be simulated, so that the model can learn under different masking conditions.
[0155] A mask ratio range is preset, such as 40% to 70%. A random generation algorithm is used to generate a mask template that meets the preset ratio. The generation of the mask template can be based on the image block method, randomly selecting several blocks to cover, or using a random dot matrix method to cover several pixels of the image.
[0156] After the mask template is generated, it needs to be applied to each enhanced image one by one. The masking of the masked area is achieved by setting the corresponding pixel values of the image to zero or replacing them with random noise values. The image after the masking process will have some information missing, which is used to train the reconstruction ability of the model.
[0157] Read the pixel matrix of each enhanced image, and perform pixel-by-pixel operations on the generated mask template and the pixel matrix of the image. For the areas marked as 1 in the mask template, replace the corresponding pixel values of the image with zero values or random noise values. For the areas marked as 0 in the mask template, keep the original pixel values of the image unchanged. Generate a masked image containing the mask information to ensure that the pixel information of some areas is effectively covered.
[0158] After the masking process is completed, the generated image will be input into the encoding network for feature extraction and reconstruction learning. The goal of the encoding network is to infer the content of the masked area by learning the features of the unmasked image area, thereby improving the reconstruction ability.
[0159] The masked image is formatted according to the model input requirements, including image resizing, batch normalization, and other operations. It is input to the input layer of the encoding network to start the feature extraction and reconstruction learning tasks. The feature vector is extracted in the middle layer of the network, and the reconstructed image is generated through the output layer of the network, completing the image reconstruction and feature representation learning process.
[0160] This embodiment can effectively improve the reconstruction and feature extraction capabilities of the model by masking some areas of the enhanced image and inputting the masked image into the encoding network. By simulating information loss, the model can infer the content of the masked area during the learning process, thereby improving the recognition accuracy and robustness of the model in the case of image loss or occlusion.
[0161] In one embodiment, the above S40 includes:
[0162] S401, the masked image includes a masked image area and an unmasked image area, and a feature vector containing information of the masked area is extracted from the masked image through the encoding network;
[0163] S402, reconstructing and learning the feature vector through the encoding network to generate a reconstruction feature for filling the masked image area;
[0164] S403: Fusing the reconstructed features with the unshielded image region to generate a reconstructed image that inherits the similarity mark or dissimilarity mark of the corresponding enhanced image.
[0165] In this embodiment, the encoding network is used to extract features and reconstruct the input masked image to generate a reconstructed image. The reconstructed image not only contains the information of the unmasked area of the original image, but also infers the content of the masked area through the reconstructed features. In addition, the reconstructed image needs to inherit the similarity mark or dissimilar mark of the original enhanced image to ensure that the model can accurately identify the relationship between similar and dissimilar images in the subsequent comparison process.
[0166] The masked image is an image with some areas covered. After receiving the image, the encoding network needs to extract the features of the uncovered area and the feature vector containing the information of the covered area from the image. These feature vectors can help the model infer the content of the covered area, thereby improving the reconstruction ability of the model.
[0167] The encoding network decomposes the input masked image into multiple image blocks and extracts the feature vector of each image block. The feature vector of the unmasked image block is mainly used to learn the overall characteristics of the image; the feature vector of the masked image block is used to infer the missing information. When extracting the feature vector, the encoding network can use the self-attention mechanism to better capture the contextual information in the image and improve the accuracy of reconstruction.
[0168] The encoding network reconstructs and learns the extracted feature vectors to restore the covered image area. The reconstruction learning process is similar to the decoding stage of the autoencoder. The network learns the features of the uncovered area and makes reasonable inferences and reconstructions about the content of the covered area.
[0169] The extracted feature vectors are nonlinearly mapped to generate reconstructed features. The reconstructed features need to be able to match the texture, color, and shape of the original image to ensure that the generated reconstructed area is seamlessly connected with other areas of the original image. The encoding network can use adversarial training or loss function optimization to improve the authenticity and accuracy of the reconstructed features.
[0170] After generating the reconstructed features, these reconstructed features need to be fused with the uncovered image areas to generate the final reconstructed image. The reconstructed image retains the similar or dissimilar markers of the enhanced image to ensure that subsequent contrastive learning can accurately identify the similarity relationship between images.
[0171] The image blocks in the uncovered area are directly copied, and the reconstructed features are filled into the covered image blocks. Image stitching and fusion algorithms are used to ensure smooth transition of the edges of the reconstructed image and avoid obvious stitching marks. After the reconstructed image is generated, the similar or dissimilar marks of the original enhanced image are inherited to the reconstructed image to ensure the consistency of the marks.
[0172] This embodiment can effectively improve the reconstruction capability of the model by reconstructing the covered image area, so that the model can still generate a complete image feature representation when processing images with partial information missing or damaged. It improves the model's ability to recover missing information and enhances the robustness of the model; when the image is partially damaged or missing, it can still achieve high-accuracy feature extraction; in scenarios such as medical imaging and insurance claims, it helps identify blurred or missing image features, thereby improving the effect of image retrieval and comparison.
[0173] In one embodiment, the above S50 includes:
[0174] S501, inputting the reconstructed image into a coding network, performing image block processing on the reconstructed image through the coding network to generate an image block sequence;
[0175] S502, in a coding network, converting the image block sequence into a feature vector sequence, and extracting a local feature vector and a global feature vector from the feature vector sequence;
[0176] S503: Fusing the local feature vector and the global feature vector to generate an image feature representation containing overall information and local details.
[0177] In this embodiment, the reconstructed image is input into the encoding network, and the network performs image segmentation, feature extraction and feature fusion on the image, and finally generates an image feature representation containing overall information and local details. The extracted image feature representation is used for image retrieval, comparison and repeated object recognition.
[0178] After the reconstructed image is input into the encoding network, the network will perform image segmentation on the image. The purpose of image segmentation is to split the large image into multiple small blocks so that the encoding network can extract image features block by block to capture detail information and local changes.
[0179] The encoding network receives the input reconstructed image and divides it into image blocks of fixed size. For example, the image is divided into 16×16 image blocks, each of which contains a certain number of pixels. Image blocks can be divided into non-overlapping ways, that is, there is no overlapping area between each image block, or into sliding windows, that is, there is a partial overlapping area between adjacent image blocks. The image block sequence after image block division is used for subsequent feature extraction to ensure that the encoding network can fully capture the local detail information of the image.
[0180] The image block sequence is the input data form of the encoding network. The network converts each image block into a feature vector to construct a feature vector sequence. The local feature vector extracted by the network is used to represent the local details of the image, while the global feature vector is used to represent the overall information of the image.
[0181] The middle layer of the encoding network converts each image block into a fixed-length feature vector to form a feature vector sequence. The local feature vector is used to capture the detailed features of the image block, such as texture, edge, and color distribution. The global feature vector is generated by the network's CLS (classification) tagging or pooling operation to represent the semantic information of the entire image. During the feature extraction process, the encoding network models the contextual information of each image block to ensure that the extracted feature vector can accurately reflect the content of the image.
[0182] In order to improve the integrity and expressiveness of image feature representation, it is necessary to fuse local feature vectors and global feature vectors. The fused image feature representation can contain both the overall information and local details of the image, thereby improving the accuracy of image retrieval and comparison.
[0183] The local feature vector and the global feature vector are concatenated or weighted fused to generate the final image feature representation. The fusion process can use an attention mechanism or a weighted average mechanism to ensure that the fused feature representation can balance the overall information and local details. The fused image feature representation is used for image retrieval and comparison to ensure that the system can quickly identify similar or duplicate images.
[0184] This embodiment can effectively improve the model's feature extraction capability by extracting image feature representation from the reconstructed image, ensuring that the model can capture both overall information and local details. It improves the integrity and expressiveness of image feature representation; improves the system's image recognition and comparison capabilities under different shooting conditions and scenes; effectively supports large-scale image retrieval and repeated image recognition, and is applied to the fields of finance, insurance, medical care, and health.
[0185] In one embodiment, the above S60 includes:
[0186] S601, dividing the reconstructed image into a similar image set and a dissimilar image set based on the similarity mark or dissimilarity mark of the reconstructed image;
[0187] S602, performing cosine similarity analysis on image feature representations in the similar image set and the dissimilar image set through a comparative learning module;
[0188] S603, for the image feature representations in the similar image set, reducing the differences between the image feature representations by minimizing the cosine similarity analysis result;
[0189] S604: for the image feature representations in the dissimilar image set, increase the difference between the image feature representations by maximizing the cosine similarity analysis result.
[0190] In this embodiment, the feature representation of the reconstructed image is analyzed and optimized by the contrast learning module, thereby reducing the feature differences between similar images and increasing the feature differences between dissimilar images. The core method is to use cosine similarity analysis to compare the directional consistency of the image feature representation and optimize the classification and discrimination capabilities of the model.
[0191] The reconstructed images have inherited the corresponding similar or dissimilar labels when they were generated, so the contrastive learning module can divide the reconstructed images into similar image sets and dissimilar image sets according to these labels. The purpose of the division is to perform different optimization strategies for the image feature representations of different sets.
[0192] Similar images collection: Contains the reconstructed images of all enhanced images generated from the same original image.
[0193] Dissimilar images collection: contains reconstructed images generated from different original images.
[0194] By traversing the similar or dissimilar labels of the reconstructed images, the reconstructed images are automatically classified into corresponding sets. In order to improve processing efficiency, a batch partitioning strategy can be adopted to divide multiple reconstructed images into sets at one time.
[0195] Cosine similarity is a common indicator for measuring the directional consistency between two feature vectors. The calculation formula is the cosine value of the angle between the feature vectors. In the contrastive learning module, the feature representations of the similar image set and the dissimilar image set are compared through cosine similarity analysis to evaluate the similarity between the images.
[0196] Extract the feature vector of each image and calculate the cosine similarity between the two feature vectors. For each pair of image feature representations in the similar image set, calculate the positive cosine similarity; for each pair of image feature representations in the dissimilar image set, calculate the negative cosine similarity.
[0197] The analysis results are optimized using the loss function in the contrastive learning module to ensure that the model can accurately distinguish between similar and dissimilar images.
[0198] In a collection of similar images, the goal is to make the reconstructed images generated by the same original image have closer feature representations. By minimizing the cosine similarity analysis results, the model can reduce the feature differences between images of the same category and improve the clustering ability of the model.
[0199] Calculate the cosine similarity for each pair of image feature representations in the similar image set, and make the feature representations of similar images closer by minimizing the loss function. Use the Infonce loss function or the contrast loss function to ensure that the angle between the feature vectors of images of the same category is close to zero. Use the gradient descent optimization algorithm to continuously update the model parameters to minimize the feature difference.
[0200] In a set of dissimilar images, the goal is to make the reconstructed images generated by different original images have greater feature differences. By maximizing the cosine similarity analysis results, the model can distinguish between images of different categories and improve the classification ability of the model.
[0201] Calculate the cosine similarity for each pair of image feature representations in the dissimilar image set, and maximize the loss function to make the feature representations of dissimilar images more different. Use contrast loss function or triple loss function to ensure that the angle between feature vectors of different categories of images is as close to 90 degrees as possible. Further improve the model's ability to distinguish through adversarial training or data enhancement.
[0202] This embodiment optimizes the feature representation of the reconstructed image through the contrastive learning module, so that the model can effectively distinguish similar images from dissimilar images. It improves the classification and differentiation capabilities of the model, effectively reduces the misclassification of similar images, enhances the robustness of the model when processing large-scale image data sets, and improves the accuracy of image retrieval and comparison.
[0203] In one embodiment, after the above S70, the method further includes:
[0204] S801, performing a similarity comparison between the feature representation of the target image and each image feature representation in the image feature index library to generate a similarity comparison result set;
[0205] S802, screening out candidate image feature representations having similarity values higher than a preset threshold from the image feature index library according to the similarity values in the similarity comparison result set, and determining candidate images corresponding to the candidate image feature representations;
[0206] S803: Perform duplicate target judgment on the candidate image and the target image to determine whether the candidate image is an image that is duplicated with the target image.
[0207] In this embodiment, after the image feature extraction and processing are completed, these image feature representations need to be organized into an index library for fast retrieval and comparison. The core goal of building an image feature index library is to improve the efficiency and accuracy of image retrieval. The index library can quickly locate image feature representations similar to the target image.
[0208] The image feature index library uses structured storage to associate the feature vector of each image with its corresponding image identifier. The index library can use KD tree, LSH (local sensitive hashing) or vector database to accelerate retrieval and improve the comparison speed. The construction process of the index library includes feature normalization, denoising and feature clustering to ensure the consistency and robustness of the feature data in the library.
[0209] When a new target image needs to be retrieved, the system extracts the feature representation of the target image and compares it with each feature representation in the image feature index library. The result of the similarity comparison is used to measure the similarity between the target image and the images in the library and generate a similarity comparison result set.
[0210] Extract the feature representation of the target image and normalize it. Use cosine similarity, Euclidean distance, or Manhattan distance as the similarity calculation method to compare the feature representation of the target image with each feature representation in the index library one by one. Generate a similarity comparison result set, record the similarity value between each index library image and the target image, and sort them according to the similarity value.
[0211] In order to improve the retrieval efficiency, the system will filter out image feature representations with similarity values higher than the preset threshold from the similarity comparison result set, and use the images corresponding to these feature representations as candidate images. The candidate image set is the basis for further judging duplicate targets, which can effectively reduce the computational complexity of subsequent comparisons.
[0212] Set a similarity threshold and select a reasonable threshold range (for example, 0.8 or 0.9) according to the actual application scenario. Traverse the similarity comparison result set and filter out image feature representations with similarity values higher than the threshold. Determine the image identifier corresponding to the candidate image feature representation and generate a candidate image set for subsequent repeated target judgment.
[0213] After generating a set of candidate images, the system will conduct a more in-depth comparison of the candidate images with the target images to determine whether they belong to the same repeated object. This process usually uses more sophisticated similarity calculation methods, such as feature matching or local feature analysis, to ensure that the recognition results of repeated objects are accurate and reliable.
[0214] Perform fine feature comparison between each candidate image and the target image, such as using SIFT (Scale Invariant Feature Transform) and ORB (Rapid Feature Point Matching) to match local features of the image. Calculate the coverage or overlap of local feature matching to determine whether the two images belong to the same target. Record the associated information of the repeated target image, including the repeated image identification, similarity value, and number of matching features, for subsequent analysis and processing.
[0215] Through the above steps, the system of this embodiment can efficiently build an image feature index library and use the index library to achieve fast image retrieval and comparison. It improves the efficiency and accuracy of image retrieval and quickly screens out candidate image sets; reduces the amount of calculation and storage costs, and optimizes the retrieval performance of the system by setting a similarity threshold.
[0216] In one embodiment, an image feature indexing device is provided, and the image feature indexing device corresponds one-to-one to the image feature indexing method in the above embodiment. Figure 3 , Figure 3 The figure is a functional module diagram of a preferred embodiment of the image feature indexing device of the present invention. The data enhancement module 10, the image marking module 20, the mask processing module 30, the image reconstruction module 40, the feature extraction module 50, the contrast learning module 60 and the feature indexing module 70. The functional modules are described in detail as follows:
[0217] The data enhancement module 10 is used to obtain an original image data set, perform a data enhancement operation on each original image in the original image data set, and generate a plurality of enhanced images corresponding to each original image;
[0218] An image marking module 20, used to mark all enhanced images generated from the same original image as similar images, and to mark enhanced images generated from different original images as dissimilar images;
[0219] The mask processing module 30 is used to perform mask processing on a partial image area of each enhanced image, and input the masked image into the encoding network;
[0220] An image reconstruction module 40, configured to reconstruct the masked image region through the encoding network to generate a corresponding reconstructed image, wherein the reconstructed image retains similarity marks or dissimilarity marks of the enhanced image;
[0221] A feature extraction module 50, configured to extract image feature representation from the reconstructed image through an encoding network;
[0222] A contrastive learning module 60, configured to reduce the difference between image feature representations for reconstructed images marked as similar images and increase the difference between image feature representations for reconstructed images marked as dissimilar images through the contrastive learning module;
[0223] The feature index module 70 is used to construct an image feature index library according to the processed image feature representation.
[0224] In one embodiment, the data enhancement module 10 is specifically used for:
[0225] Acquire an original image dataset from an image data source, where the original image dataset includes multiple images;
[0226] Performing image size adjustment, format conversion and / or color space standardization processing on each original image in the original image data set;
[0227] Based on a preset enhancement strategy, a plurality of data enhancement methods are randomly selected for each original image, wherein the data enhancement methods include cropping, rotating, flipping, adjusting brightness, and adding noise;
[0228] Apply the selected multiple data augmentation methods to each original image to generate multiple enhanced images with different feature changes.
[0229] In one embodiment, the mask processing module 30 is specifically configured to:
[0230] Based on a preset mask area ratio, a mask template is randomly generated, where the mask template is used to identify the mask area in the enhanced image;
[0231] Applying the mask template to each enhanced image to perform mask processing on the mask area identified by the mask template in each enhanced image to generate an image after mask processing;
[0232] The masked image is input into the encoding network.
[0233] In one embodiment, the image reconstruction module 40 is specifically configured to:
[0234] The masked image includes a masked image area and an unmasked image area, and a feature vector including information of the masked area is extracted from the masked image through the encoding network;
[0235] Reconstructing and learning the feature vector through the encoding network to generate reconstruction features for filling the masked image area;
[0236] The reconstructed features are fused with the unshielded image region to generate a reconstructed image that inherits similarity marks or dissimilarity marks of the corresponding enhanced image.
[0237] In one embodiment, the feature extraction module 50 is specifically used for:
[0238] Inputting the reconstructed image into a coding network, performing image block processing on the reconstructed image through the coding network to generate an image block sequence;
[0239] In the encoding network, the image block sequence is converted into a feature vector sequence, and a local feature vector and a global feature vector are extracted from the feature vector sequence;
[0240] The local feature vector and the global feature vector are fused to generate an image feature representation containing overall information and local details.
[0241] In one embodiment, the comparative learning module 60 is specifically used for:
[0242] Based on the similarity marks or dissimilarity marks of the reconstructed images, the reconstructed images are divided into a similar image set and a dissimilar image set;
[0243] Performing cosine similarity analysis on the image feature representations in the similar image set and the dissimilar image set through a comparative learning module;
[0244] For the image feature representations in a set of similar images, the differences between the image feature representations are reduced by minimizing the cosine similarity analysis results;
[0245] For the image feature representations in a set of dissimilar images, the differences between the image feature representations are increased by maximizing the cosine similarity analysis results.
[0246] In one embodiment, the feature index module 70 is specifically configured to:
[0247] Performing a similarity comparison between the feature representation of the target image and each image feature representation in the image feature index library to generate a similarity comparison result set;
[0248] According to the similarity values in the similarity comparison result set, candidate image feature representations having similarity values higher than a preset threshold are screened out from the image feature index library, and candidate images corresponding to the candidate image feature representations are determined;
[0249] A duplicate target judgment is performed on the candidate image and the target image to determine whether the candidate image is an image that is duplicated with the target image.
[0250] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 4 As shown. The computer device includes a processor, a memory, a network interface and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external user terminal through a network connection. When the computer program is executed by the processor, it realizes the functions or steps of the service side of an image feature indexing method.
[0251] In one embodiment, a computer device is provided. The computer device may be a user terminal, and its internal structure diagram may be as follows: Figure 5 As shown. The computer device includes a processor, a memory, a network interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server through a network connection. When the computer program is executed by the processor, it realizes the functions or steps of a user side of an image feature indexing method
[0252] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the following steps are implemented:
[0253] Acquire an original image data set, perform a data enhancement operation on each original image in the original image data set, and generate a plurality of enhanced images corresponding to each original image;
[0254] Mark all enhanced images generated from the same original image as similar images, and mark enhanced images generated from different original images as dissimilar images;
[0255] Perform masking on a partial image area of each enhanced image, and input the masked image into the encoding network;
[0256] Reconstructing the masked image region through the encoding network to generate a corresponding reconstructed image, wherein the reconstructed image retains similarity marks or dissimilarity marks of the enhanced image;
[0257] extracting image feature representation from the reconstructed image through an encoding network;
[0258] Through the contrast learning module, the difference between the image feature representations is reduced for the reconstructed images marked as similar images, and the difference between the image feature representations is increased for the reconstructed images marked as dissimilar images;
[0259] According to the processed image feature representation, an image feature index library is constructed.
[0260] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:
[0261] Acquire an original image data set, perform a data enhancement operation on each original image in the original image data set, and generate a plurality of enhanced images corresponding to each original image;
[0262] Mark all enhanced images generated from the same original image as similar images, and mark enhanced images generated from different original images as dissimilar images;
[0263] Perform masking on a partial image area of each enhanced image, and input the masked image into the encoding network;
[0264] Reconstructing the masked image region through the encoding network to generate a corresponding reconstructed image, wherein the reconstructed image retains similarity marks or dissimilarity marks of the enhanced image;
[0265] extracting image feature representation from the reconstructed image through an encoding network;
[0266] Through the contrast learning module, the difference between the image feature representations is reduced for the reconstructed images marked as similar images, and the difference between the image feature representations is increased for the reconstructed images marked as dissimilar images;
[0267] According to the processed image feature representation, an image feature index library is constructed.
[0268] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can refer to the relevant descriptions on the server side and the user side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.
[0269] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0270] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0271] It should be noted that if software tools or components other than those of the Company appear in the embodiments of the present application, they are only used for illustration and do not represent actual use. The above-described embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the above-mentioned embodiments, a person of ordinary skill in the art should understand that the technical solutions described in the above-mentioned embodiments can still be modified, or some of the technical features thereof can be replaced by equivalents; and these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.
Claims
1. An image feature indexing method, characterized in that: The following steps are involved: Acquire an original image data set, perform a data enhancement operation on each original image in the original image data set, and generate a plurality of enhanced images corresponding to each original image; Mark all enhanced images generated from the same original image as similar images, and mark enhanced images generated from different original images as dissimilar images; Perform masking on a partial image area of each enhanced image, and input the masked image into the encoding network; Reconstructing the masked image region through the encoding network to generate a corresponding reconstructed image, wherein the reconstructed image retains similarity marks or dissimilarity marks of the enhanced image; extracting image feature representation from the reconstructed image through an encoding network; Through the contrast learning module, the difference between the image feature representations is reduced for the reconstructed images marked as similar images, and the difference between the image feature representations is increased for the reconstructed images marked as dissimilar images; According to the processed image feature representation, an image feature index library is constructed.
2. The image feature indexing method according to claim 1, characterized in that: Acquire an original image data set, perform a data enhancement operation on each original image in the original image data set, and generate multiple enhanced images corresponding to each original image, including: Acquire an original image dataset from an image data source, where the original image dataset includes multiple images; Performing image size adjustment, format conversion and / or color space standardization processing on each original image in the original image data set; Based on a preset enhancement strategy, a plurality of data enhancement methods are randomly selected for each original image, wherein the data enhancement methods include cropping, rotating, flipping, adjusting brightness, and adding noise; Apply the selected multiple data augmentation methods to each original image to generate multiple enhanced images with different feature changes.
3. The image feature indexing method according to claim 1, characterized in that: Masking is performed on a portion of the image area of each enhanced image, and the masked image is input into the encoding network, including: Based on a preset mask area ratio, a mask template is randomly generated, where the mask template is used to identify the mask area in the enhanced image; Applying the mask template to each enhanced image to perform mask processing on the mask area identified by the mask template in each enhanced image to generate an image after mask processing; The masked image is input into the encoding network.
4. The image feature indexing method according to claim 1, characterized in that: Reconstructing the masked image region through the encoding network to generate a corresponding reconstructed image, wherein the reconstructed image retains similarity marks or dissimilarity marks of the enhanced image, including: The masked image includes a masked image area and an unmasked image area, and a feature vector including information of the masked area is extracted from the masked image through the encoding network; Reconstructing and learning the feature vector through the encoding network to generate reconstruction features for filling the masked image area; The reconstructed features are fused with the unshielded image region to generate a reconstructed image that inherits similarity marks or dissimilarity marks of the corresponding enhanced image.
5. The image feature indexing method according to claim 1, characterized in that: Extracting image feature representation from the reconstructed image through an encoding network, including: Inputting the reconstructed image into a coding network, performing image block processing on the reconstructed image through the coding network to generate an image block sequence; In the encoding network, the image block sequence is converted into a feature vector sequence, and a local feature vector and a global feature vector are extracted from the feature vector sequence; The local feature vector and the global feature vector are fused to generate an image feature representation containing overall information and local details.
6. The image feature indexing method according to claim 1, characterized in that: Through the contrast learning module, the differences between the image feature representations of the reconstructed images marked as similar images are reduced, and the differences between the image feature representations of the reconstructed images marked as dissimilar images are increased, including: Based on the similarity marks or dissimilarity marks of the reconstructed images, the reconstructed images are divided into a similar image set and a dissimilar image set; Performing cosine similarity analysis on the image feature representations in the similar image set and the dissimilar image set through a comparative learning module; For the image feature representations in a set of similar images, the differences between the image feature representations are reduced by minimizing the cosine similarity analysis results; For the image feature representations in a set of dissimilar images, the differences between the image feature representations are increased by maximizing the cosine similarity analysis results.
7. The image feature indexing method according to claim 1, characterized in that: According to the processed image feature representation, after building the image feature index library, it also includes: Performing a similarity comparison between the feature representation of the target image and each image feature representation in the image feature index library to generate a similarity comparison result set; According to the similarity values in the similarity comparison result set, candidate image feature representations having similarity values higher than a preset threshold are screened out from the image feature index library, and candidate images corresponding to the candidate image feature representations are determined; A duplicate target judgment is performed on the candidate image and the target image to determine whether the candidate image is an image that is duplicated with the target image.
8. An image feature indexing device, characterized in that: The image feature indexing device comprises: A data enhancement module is used to obtain an original image data set, perform a data enhancement operation on each original image in the original image data set, and generate a plurality of enhanced images corresponding to each original image; An image marking module, used for marking all enhanced images generated from the same original image as similar images, and marking enhanced images generated from different original images as dissimilar images; A mask processing module, used for performing mask shielding processing on a partial image area of each enhanced image, and inputting the mask shielded image into the encoding network; An image reconstruction module, used to reconstruct the masked image area through the encoding network to generate a corresponding reconstructed image, wherein the reconstructed image retains similarity marks or dissimilarity marks of the enhanced image; A feature extraction module, used for extracting image feature representation from the reconstructed image through an encoding network; A contrastive learning module, used for reducing the difference between image feature representations for reconstructed images marked as similar images and increasing the difference between image feature representations for reconstructed images marked as dissimilar images through the contrastive learning module; The feature index module is used to construct an image feature index library based on the processed image feature representation.
9. A computer device, characterized in that: The computer device includes a memory, a processor, and an image feature indexing program stored in the memory and executable on the processor. When the image feature indexing program is executed by the processor, the steps of the image feature indexing method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that: The storage medium stores an image feature indexing program, which, when executed by a processor, implements the steps of the image feature indexing method according to any one of claims 1 to 7.
Citation Information
Cited By
Image feature indexing method and apparatus, device, and medium
WO2026166147A1