Image retrieval method and device, computer equipment and storage medium

By segmenting and encoding images and generating image indexes based on location information, the problems of high computational overhead and low performance in existing technologies are solved, and efficient and accurate image retrieval is achieved, which is suitable for large-scale image library management and fast retrieval.

CN120832428AActive Publication Date: 2025-10-24PING AN TECH (SHENZHEN) CO LTD

Patent Information

Application Number
CN202510698854.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-10-24
Estimated Expiration
2045-05-27

AI Technical Summary

Technical Problem

Existing image retrieval methods have high computational overhead and low performance, making it difficult to efficiently manage and retrieve large-scale image libraries.

Method used

Image segmentation and encoding technology is used to divide the image into sub-images, and the image index is generated by combining the location information. The improved Vision Transformer model is used for encoding to generate a compact hash code index, which is stored in the image retrieval database.

Benefits of technology

It achieves efficient and accurate image retrieval, improves the efficiency of image library management and retrieval speed, and is suitable for large-scale image library management and fast retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120832428A_ABST
    Figure CN120832428A_ABST
Patent Text Reader

Abstract

The invention discloses an image retrieval method and device, computer equipment and a storage medium, belongs to the technical field of artificial intelligence, and is provided with a retrieval system applied to an insurance marketing scene image. According to the method and the device, the to-be-processed image is segmented and coded, the content features of the image are combined with the position information to form high-quality image representation, and the corresponding image index information is generated by utilizing the preset image index generator, so that efficient indexing and retrieval of the image are realized. In the encoding process, local features of the image are reserved, spatial position information is fused, discrimination and uniqueness of image indexing are improved, generated image indexing information and an original image are stored in an image retrieval database in an associated mode, it is ensured that the image retrieval process has high matching precision and response speed, and the image retrieval efficiency is improved. The image retrieval accuracy and processing efficiency are effectively improved, and the method is suitable for large-scale image library management and rapid retrieval.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of artificial intelligence, and particularly relates to an image retrieval method and device, a computer device and a storage medium. BACKGROUND

[0002] In marketing activities, image materials play a key role. However, in the face of various businesses, different focuses and diverse material sources, the image material library is increasingly large, which brings challenges to efficient retrieval and classification of materials, and business personnel need to reprocess the retrieved information, which reduces work efficiency. Taking the insurance marketing scene as an example, for example, health insurance needs to prepare various disease comparison pictures before and after claims, and professional claims team service pictures; car insurance needs to show accident scene investigation, fast claims to account pictures. The business personnel need to reorganize after retrieval, which affects the work progress.

[0003] Traditional image retrieval methods mainly rely on manual labels or descriptions, which not only consumes a large amount of human resources, but also is prone to subjective bias, affecting the accuracy of retrieval. In recent years, with the development of artificial intelligence technology, some solutions using deep learning networks (such as CNN convolutional models) for deep extraction of image features have appeared, but there are problems of large computational overhead and low performance on large-scale training sets. SUMMARY

[0004] The purpose of the embodiments of the present application is to provide an image retrieval method, device, computer device and storage medium to solve the technical problems of large computational overhead and low performance of existing image retrieval methods.

[0005] To solve the above technical problems, the embodiments of the present application provide an image retrieval method, which adopts the technical solutions as follows:

[0006] An image retrieval method comprises:

[0007] Obtaining pre-collected to-be-processed images, and performing image segmentation on the to-be-processed images to obtain a plurality of to-be-processed sub-images;

[0008] Performing image encoding on the to-be-processed sub-images to obtain a first encoding vector;

[0009] Adding image position information to each to-be-processed sub-image, and performing position encoding on the image position information of the to-be-processed sub-images to obtain a second encoding vector;

[0010] Inputting the first encoding vector and the second encoding vector into a preset image index generator to obtain image index information of the to-be-processed images;

[0011] Correlate the image index information with the image to be processed to obtain associated image data, and store the associated image data in the image retrieval database.

[0012] In response to an image retrieval instruction, obtain target image index information matched with the image to be retrieved, and search the image retrieval database for a target image matched with the target image index information.

[0013] To solve the above technical problems, the embodiment of the application further provides an image retrieval device, which adopts the technical scheme as follows:

[0014] An image retrieval device comprises:

[0015] An image segmentation module is configured to obtain pre-collected images to be processed, and perform image segmentation on the images to be processed to obtain a plurality of sub-images to be processed.

[0016] An image encoding module is configured to perform image encoding on the sub-images to be processed to obtain first encoding vectors.

[0017] A position encoding module is configured to add image position information to each of the sub-images to be processed, and perform position encoding on the image position information of the sub-images to be processed to obtain second encoding vectors.

[0018] An index generation module is configured to input the first encoding vectors and the second encoding vectors into a preset image index generator to obtain image index information of the image to be processed.

[0019] A data correlation module is configured to correlate the image index information with the image to be processed to obtain associated image data, and store the associated image data in the image retrieval database.

[0020] An image retrieval module is configured to, in response to an image retrieval instruction, obtain target image index information matched with the image to be retrieved, and search the image retrieval database for a target image matched with the target image index information.

[0021] To solve the above technical problems, the embodiment of the application further provides a computer device, which adopts the technical scheme as follows:

[0022] A computer device comprises a memory and a processor, the memory stores computer readable instructions, and the processor executes the computer readable instructions to realize the steps of the image retrieval method according to any one of the above.

[0023] To solve the above technical problems, the embodiment of the application further provides a computer readable storage medium, which adopts the technical scheme as follows:

[0024] A computer readable storage medium having stored thereon computer readable instructions which, when executed by a processor, implement the steps of the image retrieval method of any one of the above.

[0025] Compared with the prior art, the embodiments of the present application have the following beneficial effects:

[0026] The present application discloses an image retrieval method and device, computer equipment and a storage medium, belonging to the field of artificial intelligence, and has an image retrieval system applied to an insurance marketing scene. The present application combines the content features and position information of an image by segmenting and encoding a to-be-processed image, forms a high-quality image representation, and generates corresponding image index information using a preset image index generator, thereby realizing efficient indexing and retrieval of images. In the encoding process, not only the local features of the image are retained, but also spatial position information is fused, improving the discriminability and uniqueness of the image index. By associating and storing the generated image index information with the original image in an image retrieval database, the image retrieval process is ensured to have high matching accuracy and response speed. In actual application, when an image retrieval instruction is received, the system can quickly obtain images matching the target image index information from the database, realizing efficient and accurate image retrieval. The present application effectively improves the accuracy and processing efficiency of image retrieval, and is suitable for large-scale image library management and fast retrieval. BRIEF DESCRIPTION OF DRAWINGS

[0027] In order to more clearly illustrate the schemes in the present application, the drawings needed in the description of the embodiments of the present application will be briefly introduced as follows. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor on the basis of these drawings.

[0028] Figure 1 An exemplary system architecture diagram to which the present application can be applied is shown;

[0029] Figure 2 A flowchart of one embodiment of the image retrieval method according to the present application is shown;

[0030] Figure 3 A structural schematic diagram of one embodiment of the image retrieval device according to the present application is shown;

[0031] Figure 4 A structural schematic diagram of one embodiment of the computer equipment according to the present application is shown. DETAILED DESCRIPTION

[0032] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs; the terms used in the specification are intended to describe the particular embodiments and are not intended to limit the application; the terms "include" and "have" and their any variations used in the specification and the claims and the above description of drawings are intended to cover the non-exclusive inclusion; the terms "first", "second" and the like used in the specification and the claims and the above description of drawings are intended to distinguish different objects, not to describe a particular order.

[0033] Reference herein to "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment can be included in at least one embodiment of the application. The appearances of the phrase in various places in the specification are not necessarily all referring to the same embodiment, nor are they necessarily mutually exclusive of one another. It is expressly understood that the embodiments described herein are merely examples and are not intended to limit the scope of the application.

[0034] In order to make the person skilled in the art better understand the scheme of the present application, the technical solutions in the embodiments of the present application will be described clearly and completely in conjunction with the drawings below.

[0035] As shown in Figure 1 The system architecture 100 can include a terminal device 101, a network 102 and a server 103, and the terminal device 101 can be a notebook computer 1011, a tablet computer 1012 or a mobile phone 1013. The network 102 is a medium for providing a communication link between the terminal device 101 and the server 103. The network 102 can include various connection types, such as wired, wireless communication links or optical fiber cables, etc.

[0036] The user can use the terminal device 101 to interact with the server 103 through the network 102 to receive or send messages, etc. Various communication client applications can be installed on the terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.

[0037] The terminal device 101 can be various electronic devices with a display screen and supporting web browsing, in addition to the notebook computer 1011, the tablet computer 1012 or the mobile phone 1013, the terminal device 101 can also be an electronic book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 player (Moving Picture Experts Group Audio Layer IV), a laptop computer and a desktop computer, etc.

[0038] The server 103 can be a server providing various services, for example, a background server providing support for a page displayed on the terminal device 101.

[0039] It should be noted that the image retrieval method provided by the embodiments of the present application is generally executed by a server / terminal device, and accordingly, the image retrieval apparatus is generally arranged in a server / terminal device.

[0040] It should be understood that, Figure 1 The number of terminal devices, networks and servers in the above system is only illustrative, and the above system can have any number of terminal devices, networks and servers according to the implementation needs.

[0041] With reference to Figure 2 , a flow chart of one embodiment of the image retrieval method according to the present application is shown. The image retrieval method includes the following steps:

[0042] S201, acquiring a pre-collected image to be processed, and performing image segmentation on the image to be processed to obtain a plurality of sub-images to be processed;

[0043] Specifically, first, the original image to be processed is acquired from a pre-set image acquisition module or a business system, and the image source can include a local material library, an image acquisition device or an external database. In order to facilitate processing by the Transformer network, the system divides the original image, adopts an image segmentation strategy consistent with the Vision Transformer (ViT), divides the whole image into a plurality of fixed-size image patches (sub-images), and each sub-image represents a small area of the image. This image segmentation method can convert a two-dimensional image into a one-dimensional patch sequence, which is convenient for feeding into the encoder of the Transformer for processing.

[0044] In the technical solution, ViT is used as the core image feature extraction model. ViT is a deep learning model based on Transformer architecture, unlike traditional convolutional neural networks (CNN), which do not rely on convolution operations to extract local features. Instead, the input image is divided into a number of fixed-size image blocks (patches), each of which is then flattened and linearly mapped to a vector sequence, with position information added, and then input into the Transformer encoder as a whole. ViT models the relationship between regions of an image in a global range through self-attention mechanisms, has excellent global perception capabilities, can capture dependencies between distant pixels, and extracts more complete and accurate image semantic information.

[0045] In the present solution, two key optimizations are made to address the high computational load and Token redundancy of the ViT native model in image retrieval tasks: first, the number of input Tokens is reduced by merging image patches to reduce redundant calculations; second, the standard multi-head attention mechanism is simplified to a linear attention mechanism, significantly reducing computational complexity and improving processing efficiency while ensuring the need for feature expression in image retrieval.

[0046] Compared to the sliding window method of traditional CNN, this division method preserves local information while providing a basis for subsequent global modeling. In addition, to further improve model efficiency, the present technology also introduces a patch merging strategy, which merges patches in adjacent or similar regions during segmentation, thereby reducing the number of tokens without significantly losing feature information, effectively alleviating the computational burden in subsequent processing.

[0047] S202, image encoding is performed on the to-be-processed sub-image to obtain a first encoding vector;

[0048] Specifically, the divided image patches will be input into the image encoding module, which is built based on the improved Vision Transformer (ViT) architecture. Each sub-image is first expanded into a vector form, and then projected to a unified feature dimension through a linear mapping layer to form a basic feature vector.

[0049] Different from the traditional ViT, in order to improve the computational efficiency and adapt to the image retrieval task, the linear attention mechanism is used in the encoder part instead of the standard multi-head self-attention (Multi-Head Self-Attention) to significantly reduce the complexity. Linear attention replaces the attention matrix of full connection by inner product, avoiding the problem of token explosion and large computational overhead of the original ViT in high-resolution image processing. Under this design, the global relationship between image patches can still be modeled, while the encoding efficiency is greatly improved. The first encoding vector output by this step represents the semantic content information of each sub-image in the image space.

[0050] S203, adding image position information to each to-be-processed sub-image, and position encoding the image position information of the to-be-processed sub-image to obtain a second encoding vector;

[0051] Specifically, the ViT model itself does not have the spatial induction bias inherent in CNN, so in order to let the model understand the spatial position of each sub-image, the two-dimensional position coordinate information of each patch in the original image needs to be added. The position information is first represented in the form of absolute position (such as row and column index or normalized coordinates), and then converted by a position encoding function to obtain a second encoding vector with the same dimension as the first encoding vector. The position encoding can use fixed sine function encoding or training obtained position embedding vector, and the present technology preferentially uses trainable position encoding in combination with the task characteristics to enhance the model's learning ability of spatial structure. These position vectors will be fused with the first encoding vector in the subsequent encoding to provide spatial position information support for the model. Through this step, the model can not only understand the semantic features of the image content when processing features, but also master the spatial layout of each sub-image in the whole image, so that the finally generated image index has stronger distinguishability and stability, thereby improving the matching accuracy and robustness in the image retrieval task.

[0052] S204, inputting the first encoding vector and the second encoding vector into a preset image index generator to obtain image index information of the to-be-processed image;

[0053] Specifically, the content encoding (first encoding vector) of the image is fused with the position encoding (second encoding vector), which can be fused in the form of vector addition, splicing or attention mechanism weighting, to form a complete set of image representation features. The fused features are input into a custom-designed image index generator, which is composed of a dropout layer and two dense layers. The dropout layer is used to prevent overfitting and enhance the generalization ability of the model. Then the high-dimensional features are mapped to a lower-dimensional latent space, and finally a fixed-length binary hash code is output as the index information of the image. This hash code can significantly compress the storage space and support fast similarity calculation (such as Hamming distance matching), which is particularly suitable for use in large-scale image retrieval databases. The entire process generates a unique, compact and efficiently matchable index identifier for the image, balancing accuracy and computational efficiency.

[0054] S205, associate the image index information with the image to be processed to obtain associated image data, and store the associated image data in the image retrieval database;

[0055] Specifically, the system establishes a one-to-one association between the image index information generated in the previous step and the original image to form image-index pairs. These associated image data are uniformly formatted and written into the image retrieval database. The image retrieval database supports high-performance hash code index management structures, such as inverted index, local sensitive hashing (LSH) or product quantization (PQ) index schemes, to ensure fast search in a large number of images. The completion of this step indicates that the image has the ability to be quickly searched, and the image library is also structured and managed.

[0056] In addition, to support multi-scene retrieval requirements, the database can also store metadata related to the image, such as business tags, image categories, upload time, usage frequency, etc., to facilitate multi-dimensional filtering and sorting during retrieval.

[0057] S206, in response to the image retrieval instruction, obtaining target image index information matched with the image to be retrieved, and searching for target images matched with the target image index information in the image retrieval database.

[0058] Specifically, when the system receives a retrieval request, it re-encodes the image to be retrieved according to the process from S201 to S204 to generate its index information. The index information is input into the image retrieval database as a query keyword, and the target image index closest to it is quickly located by calculating the similarity between hash codes (for example, calculating the Hamming distance or approximate nearest neighbor matching). After a successful match, the system further extracts the corresponding target image and its related information from the database and returns the result to the user. Due to the use of a compact and efficient hash index, the retrieval speed is greatly improved, supporting a response in seconds; at the same time, because the index integrates image content and spatial location information, the matching accuracy is higher. This retrieval process adopts an end-to-end single-stage architecture, avoiding the problems of coarse sorting error accumulation and repeated calculations in traditional two-stage retrieval schemes, improving the practicality and user experience of the overall system, and is particularly suitable for industry scenarios with large image volumes and frequent retrieval, such as insurance marketing, customer service, image content auditing, etc.

[0059] In the above embodiment, the present application segments and encodes the image to be processed, combines the content features of the image with the position information to form a high-quality image representation, and uses a preset image index generator to generate the corresponding image index information, thereby realizing efficient indexing and retrieval of the image. During the encoding process, not only the local features of the image are retained, but also the spatial position information is integrated, thereby improving the discriminability and uniqueness of the image index. By associating the generated image index information with the original image and storing it in the image retrieval database, it is ensured that the image retrieval process has high matching accuracy and response speed. In actual applications, after receiving the image retrieval instruction, the system can quickly obtain the image that matches the target image index information from the database, thereby realizing efficient and accurate image retrieval. The present application effectively improves the accuracy and processing efficiency of image retrieval, and is suitable for large-scale image library management and rapid retrieval.

[0060] For example, in insurance marketing scenarios, sales personnel often need to retrieve and display a large number of images related to different types of insurance. For example, health insurance requires images of physical examinations, hospitals, and service processes; auto insurance requires images of accident scenes, vehicle damage, and claims services. In practice, due to the large number of images, manual organization and search are inefficient, which can easily delay customer communication and business progress.

[0061] Taking a car insurance accident claim scene as an example, when an accident vehicle image (such as a rear-end accident scene) is uploaded into the system, the image will first be subjected to preprocessing by the image segmentation module. For example, the system automatically identifies key areas in the image, such as the damaged vehicle head, ground scratches, surrounding environment, etc., and divides them into multiple sub-images. Each sub-image is flattened, encoded, and combined with its spatial location information in the image to extract features through the ViT improved model. The system then inputs the extracted content features and location information into the index generator, which is processed through the Dropout layer and the fully connected layer to generate structured first and second target vectors.

[0062] Next, the content and location information are fused by introducing a linear attention mechanism to obtain a weighted feature vector, which is mapped to a hash space to ultimately generate a string of binary hash codes such as "011010101011…", which is the unique index information of the image. The index is stored in the image retrieval database and corresponds one-to-one with the image sample.

[0063] Further, the method further includes the steps of obtaining a pre-collected image to be processed and performing image segmentation on the image to be processed to obtain a plurality of sub-images to be processed, which specifically includes:

[0064] Determining any pixel point in the image to be processed as a pixel seed point;

[0065] Using a region growing method to identify whether the adjacent pixel points of the pixel seed point match the pixel seed point;

[0066] Determining the sub-image region according to the pixel point matching result;

[0067] Performing image segmentation on the image to be processed according to the sub-image region to obtain a plurality of sub-images to be processed.

[0068] In this embodiment, the image segmentation no longer uses the traditional fixed size division method, but introduces a region growing method based on content perception to generate sub-image regions, thereby obtaining image segments with more semantic consistency. Specifically, first, a pixel point in the image is selected as an initial pixel seed point, which can be initialized according to the image brightness, color histogram, or texture features. Then, the region growing algorithm is used to gradually expand the neighborhood region of the seed point, and determine whether the adjacent pixel points are similar to the seed point in a specific feature space (for example, the color difference is lower than a threshold, the texture direction is consistent, etc.). If they match, they are included in the same region and continue to expand outward until they cannot meet the similarity condition, forming a complete sub-image region. This process can be repeated until the entire image is divided into a plurality of semantically related and clearly bounded sub-images.

[0069] Compared with the fixed-size division method, this method is more in line with the actual structural information in the image, can improve the effectiveness and expressiveness of subsequent feature extraction, and provide more representative input data for the ViT encoder.

[0070] Through the above steps, a fine division based on the semantics of image content is achieved, which improves the structural consistency and feature expression ability of sub-images.

[0071] Furthermore, the step of determining the sub-image area according to the pixel matching result specifically includes:

[0072] When adjacent pixel points match the pixel seed point, the area where the adjacent pixel points and the pixel seed point are located is determined as a sub-image area;

[0073] When the adjacent pixel points do not match the pixel seed point, the unmatched pixel points are obtained;

[0074] Use the unmatched pixels as new pixel seed points;

[0075] Continue to identify whether the adjacent pixel points of the new pixel seed point match the new pixel seed point until all pixel points in the image to be processed are matched, and obtain all sub-image areas.

[0076] In this embodiment, the system adaptively segments the image based on the region growing method. The key is to use pixel seed points as the core and form multiple semantically consistent sub-image areas through continuous iterative expansion. First, an initial pixel point is selected from the image as the seed point, and its adjacent pixels are identified to see whether they are similar to it in dimensions such as color, brightness, and texture direction. Once the match is successful, the adjacent pixel will be incorporated into the current sub-image area. Each newly added pixel point in the area will continue to be matched outward as an extension source point, thereby gradually expanding the area until no more qualified adjacent pixel points can be added. At the same time, for pixel points that do not match the current area, they are independently extracted as new pixel seed points, and the above-mentioned area expansion process is repeated until all pixels in the entire image are attributed to a sub-image area.

[0077] Through the above steps, full coverage content division of each pixel in the image is achieved, and the semantic consistency and structural accuracy of the sub-image area are improved.

[0078] Furthermore, the step of performing image encoding on the sub-image to be processed to obtain a first encoding vector specifically includes:

[0079] Flatten the sub-image to be processed into a one-dimensional vector to obtain the initial features of the sub-image;

[0080] Perform dimensionality increase and dimensionality reduction operations on the initial features of the sub-image to obtain the multi-dimensional features of the sub-image;

[0081] The preset encoder is used to encode the multi-dimensional features of the sub-image to obtain a first encoding vector.

[0082] In this embodiment, in order to fully exploit the feature information in the sub-image to be processed, the system first flattens each sub-image into a one-dimensional vector to obtain its initial feature representation. The flattening operation can preserve the basic structural information of the original pixel arrangement. Then, the system performs dimension increasing and dimension decreasing processing on the one-dimensional vector. The dimension increasing operation projects the feature vector into a higher-dimensional feature space through the introduction of a linear transformation (such as a fully connected layer), thereby increasing its expression capacity, enabling the model to learn more complex and detailed feature relationships. The dimension decreasing operation is used to suppress redundant information, improve subsequent computational efficiency by compressing the feature dimension, and avoid overfitting and resource waste caused by high-dimensional features. This dimension increasing and decreasing process balances the expressiveness and computational cost of the sub-image features. Subsequently, the system inputs the processed multi-dimensional features into a preset encoder. The encoder is designed based on the structure of Vision Transformer (ViT) and uses its powerful self-attention mechanism to extract deep feature relationships of the sub-image in a global range, thereby outputting a high-quality first encoding vector.

[0083] Through the above steps, the richness and abstractness of the sub-image feature representation are improved, providing a solid foundation for the accuracy of image indexing.

[0084] Further, the step of adding image position information to each sub-image to be processed and position encoding the image position information of the sub-image to be processed to obtain a second encoding vector includes:

[0085] Constructing a plane coordinate system based on the plane in which the image to be processed is located;

[0086] Obtaining the image position information corresponding to each sub-image to be processed based on the plane coordinate system;

[0087] Associating each image position information with the corresponding sub-image to be processed;

[0088] Position encoding the image position information of the sub-image to be processed to obtain a second encoding vector.

[0089] In this embodiment, in order to preserve the spatial relationship between each sub-image in the image, the system introduces a position information processing step after image segmentation. First, a unified two-dimensional plane coordinate system is established on the plane where the entire image to be processed is located, and each sub-image has a unique position identifier in the coordinate system, which is usually represented by the upper left corner coordinates or the center point coordinates of the sub-image in the original image. The system then generates a position identifier for each sub-image based on these coordinate information, and associates the position with the corresponding sub-image features. Then, the image position information is processed by position encoding to adapt to the sequence input requirement of the Transformer structure. The position encoding can use sine and cosine function encoding, learnable position embedding or hybrid method to map the spatial position information into a vector form with the same dimension as the image feature vector. The position encoding result is the second encoding vector, which plays a role in guiding the attention mechanism to focus on different image regions in the model, so that the model not only pays attention to the content features, but also perceives the spatial layout, thereby more effectively capturing the relationship between the global structure and the local dependency in the image.

[0090] Through the above steps, the explicit expression of image spatial structure information is realized, and the understanding ability of the model for the global relationship of the image is enhanced.

[0091] Further, the image index generator includes a random dropout layer and a fully connected layer, and the step of inputting the first encoding vector and the second encoding vector into the preset image index generator to obtain the image index information of the image to be processed includes:

[0092] using the random dropout layer to randomly deactivate the first encoding vector to obtain a first intermediate vector;

[0093] using the fully connected layer to perform linear transformation processing on the first intermediate vector to obtain a first target vector;

[0094] using the random dropout layer to randomly deactivate the second encoding vector to obtain a second intermediate vector;

[0095] using the fully connected layer to perform linear transformation processing on the second intermediate vector to obtain a second target vector;

[0096] hashing the first target vector and the second target vector to obtain a binary hash code;

[0097] using the binary hash code as the image index information of the image to be processed.

[0098] In this embodiment, in order to construct efficient and discriminative image index information, the system designs an image index generator based on a deep neural network, mainly including a dropout layer and a dense layer. This structure not only helps to improve the generalization ability of the model, but also compresses the feature dimension and enhances the compactness of the index. Specifically, first, the first encoding vector output by the ViT encoder is input to the Dropout layer. Random inactivation can effectively reduce the dependence between neurons and prevent the model from overfitting during training, thereby improving the overall robustness. Subsequently, the first intermediate vector after inactivation is sent to the Dense Layer layer for linear transformation to extract higher-level abstract features and form the first target vector. At the same time, the second encoding vector after position encoding also undergoes the same processing procedure, and the second target vector is obtained after random inactivation and linear transformation. Next, the system performs feature fusion on the first and second target vectors and performs binary hash mapping operation through a specific hash function to generate a fixed-length binary hash code. This hash code not only preserves the content features of the image but also embeds its spatial position information, which is a compact representation of the global semantic and structural information of the image. Finally, this hash code serves as the unique index information of the image, which can be used for fast matching and image retrieval operations, and also has good storage efficiency and search performance.

[0099] Through the above steps, a compact hash index with content recognition and spatial structure perception ability is constructed, which improves the accuracy of image retrieval and the response speed of the system.

[0100] Further, the step of performing hash value mapping on the first target vector and the second target vector to obtain a binary hash code includes:

[0101] The linear attention mechanism is used to calculate the linear attention weight value of the first target vector and the second target vector.

[0102] Based on the linear attention weight value, the first target vector and the second target vector are weighted and summed to obtain a weighted feature vector.

[0103] The weighted feature vector is introduced into a preset hash mapping space, and the hash value mapping of the weighted feature vector is completed in the hash mapping space to obtain a binary hash code.

[0104] In this embodiment, the system adopts a linear attention mechanism to efficiently fuse image content features and location information, and then generates discriminative binary hash codes. First, for the first target vector (representing image content features) and the second target vector (representing image location information), the system calculates the attention weights between the two by introducing a linear attention mechanism. Unlike traditional self-attention, linear attention reduces the complexity from quadratic to linear, significantly reducing the consumption of computing resources, while retaining the key ability to model contextual dependencies. The system performs weighted sum processing on the two target vectors based on these weight values, fusing their core information in terms of image semantics and spatial location to obtain a unified weighted feature vector. Subsequently, the weighted feature vector is mapped to a pre-set hash space, which can be constructed by a trainable hash layer, or by a set of projection matrices and threshold functions to convert floating-point vectors into discrete binary representations. Finally, the system generates a compact and semantically rich binary hash code based on the distribution of the weighted feature in the hash mapping space, which is used as a unique index identifier for image retrieval.

[0105] Through the above steps, the depth fusion of image semantics and spatial information is achieved, and is converted into an efficient and matchable hash index, improving the response efficiency and accuracy of the system in large-scale image retrieval scenarios.

[0106] In the above embodiment, the present application discloses an image retrieval method, which belongs to the field of artificial intelligence and has an image retrieval system applied to insurance marketing scenarios. The present application combines the content features and location information of the image by segmenting and encoding the image to be processed, forms a high-quality image representation, and generates corresponding image index information using a pre-set image index generator, thereby realizing efficient indexing and retrieval of images. In the encoding process, not only the local features of the image are preserved, but also the spatial location information is fused, improving the discriminability and uniqueness of the image index. By associating the generated image index information with the original image and storing them in the image retrieval database, the image retrieval process is ensured to have high matching accuracy and response speed. In actual application, when receiving an image retrieval instruction, the system can quickly obtain images matching the target image index information from the database, realizing efficient and accurate image retrieval. The present application effectively improves the accuracy and processing efficiency of image retrieval, and is suitable for large-scale image library management and fast retrieval.

[0107] In this embodiment, the image retrieval method is run on an electronic device (e.g. Figure 1The server shown) can receive instructions or obtain data through wired or wireless connections. It should be noted that the wireless connection can include, but is not limited to, 3G / 4G connection, WiFi connection, Bluetooth connection, Wi MAX connection, Zigbee connection, UWB (ultra wideband) connection, and other now known or future developed wireless connection.

[0108] It should be emphasized that, in order to further ensure the privacy and security of the above-mentioned image information, the above-mentioned image information can also be stored in a node of a block chain.

[0109] The blockchain referred to in the present application is a new application mode of distributed data storage, peer-to-peer transmission, consensus mechanism, encryption algorithm and other computer technologies. Blockchain, in essence, is a decentralized database, a series of data blocks associated using cryptographic methods, each containing a batch of network transaction information for verifying the validity (anti-fake) of the information and generating the next block. Blockchain can include blockchain underlying platform, platform product service layer, and application service layer, etc.

[0110] The embodiments of the present application can acquire and process related data based on artificial intelligence technology. Among them, artificial intelligence (Artificial Intelligence, AI) is to use digital computers or digital computer controlled machines to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. Theory, method, technology and application system.

[0111] The basic technology of artificial intelligence generally includes technologies such as sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction system, mechatronics, etc. Artificial intelligence software technology mainly includes computer vision technology, robot technology, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning, etc. Several major directions.

[0112] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be completed by computer readable instructions instructing related hardware, which can be stored in a computer readable storage medium. The program can include the processes of the above-mentioned embodiments when executed, wherein the storage medium can be a magnetic disc, an optical disc, a read-only memory (ROM) or other non-volatile storage medium, or a random access memory (RAM) etc.

[0113] It should be understood that although each step in the flowchart of the accompanying drawings is shown in sequence according to the direction of the arrow, these steps are not necessarily executed in sequence according to the direction of the arrow. Unless otherwise specified herein, the execution of these steps is not strictly limited in sequence, and they can be executed in other sequences. Moreover, at least part of the steps in the flowchart of the accompanying drawings can include multiple sub-steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution sequence is not necessarily sequential, but can be alternately executed with at least part of other steps or sub-steps or stages of other steps.

[0114] Further referring to Figure 3 , as an implementation of the method shown in the above Figure 2 , the present application provides an embodiment of an image retrieval device, which corresponds to the method embodiment shown in Figure 2 , and the device can be specifically applied to various electronic devices.

[0115] As shown in Figure 3 , the image retrieval device 300 described in the embodiment includes:

[0116] An image segmentation module 301 is configured to acquire a pre-collected to-be-processed image, and perform image segmentation on the to-be-processed image to obtain a plurality of to-be-processed sub-images.

[0117] An image encoding module 302 is configured to perform image encoding on the to-be-processed sub-images to obtain a first encoding vector.

[0118] A position encoding module 303 is configured to add image position information to each to-be-processed sub-image, and perform position encoding on the image position information of the to-be-processed sub-images to obtain a second encoding vector.

[0119] An index generation module 304 is configured to input the first encoding vector and the second encoding vector into a pre-set image index generator to obtain image index information of the to-be-processed image.

[0120] A data association module 305 is configured to associate the image index information with the to-be-processed image to obtain associated image data, and store the associated image data into an image retrieval database.

[0121] An image retrieval module 306 is configured to acquire target image index information matched with a to-be-retrieved image in response to an image retrieval instruction, and find a target image matched with the target image index information in the image retrieval database.

[0122] Further, the image segmentation module 301 specifically includes:

[0123] A pixel seed point unit is configured to determine an arbitrary pixel point in the image to be processed as a pixel seed point;

[0124] A pixel matching unit is configured to identify whether adjacent pixel points of the pixel seed point match the pixel seed point using a region growing method;

[0125] An image region unit is configured to determine a sub-image region according to a pixel point matching result;

[0126] An image segmentation unit is configured to perform image segmentation on the image to be processed according to the sub-image region, to obtain a plurality of sub-images to be processed.

[0127] Further, the image region unit specifically includes:

[0128] A first matching sub-unit is configured to, when the adjacent pixel points match the pixel seed point, determine a region in which the adjacent pixel points and the pixel seed point are located as a sub-image region;

[0129] A second matching sub-unit is configured to, when the adjacent pixel points do not match the pixel seed point, acquire non-matching pixel points;

[0130] A new pixel seed point sub-unit is configured to take the non-matching pixel points as new pixel seed points;

[0131] A continuous matching sub-unit is configured to continue to identify whether adjacent pixel points of the new pixel seed points match the new pixel seed points, until all pixel points in the image to be processed participate in the matching, to obtain all sub-image regions.

[0132] Further, the image encoding module 302 specifically includes:

[0133] An image flattening unit is configured to flatten the sub-image to be processed into a one-dimensional vector, to obtain a sub-image initial feature;

[0134] A dimension transformation unit is configured to perform dimension increasing operation and dimension decreasing operation on the sub-image initial feature, to obtain a sub-image multi-dimensional feature;

[0135] An image encoding unit is configured to perform encoding operation on the sub-image multi-dimensional feature using a preset encoder, to obtain a first encoding vector.

[0136] Further, the position encoding module 303 specifically includes:

[0137] A plane coordinate system unit is configured to construct a plane coordinate system on a plane in which the image to be processed is located;

[0138] A position acquisition unit is configured to acquire image position information corresponding to each sub-image to be processed based on the plane coordinate system;

[0139] The image association unit is configured to associate each image position information with a corresponding to-be-processed sub-image.

[0140] The position encoding unit is configured to perform position encoding on the image position information of the to-be-processed sub-image to obtain a second encoding vector.

[0141] Further, the image index generator comprises a random inactivation layer and a full connection layer, and the index generation module 304 specifically comprises:

[0142] The first inactivation processing unit is configured to perform random inactivation processing on the first encoding vector by using the random inactivation layer to obtain a first intermediate vector.

[0143] The first linear transformation unit is configured to perform linear transformation processing on the first intermediate vector by using the full connection layer to obtain a first target vector.

[0144] The second inactivation processing unit is configured to perform random inactivation processing on the second encoding vector by using the random inactivation layer to obtain a second intermediate vector.

[0145] The second linear transformation unit is configured to perform linear transformation processing on the second intermediate vector by using the full connection layer to obtain a second target vector.

[0146] The hash mapping unit is configured to perform hash value mapping on the first target vector and the second target vector to obtain a binary hash code.

[0147] The image index unit is configured to take the binary hash code as image index information of the to-be-processed image.

[0148] Further, the hash mapping unit specifically comprises:

[0149] The linear attention sub-unit is configured to calculate linear attention weight values of the first target vector and the second target vector by using a linear attention mechanism.

[0150] The weighted summation sub-unit is configured to perform weighted summation on the first target vector and the second target vector based on the linear attention weight values to obtain a weighted feature vector.

[0151] The hash mapping sub-unit is configured to import the weighted feature vector into a preset hash mapping space, and perform hash value mapping on the weighted feature vector in the hash mapping space to obtain the binary hash code.

[0152] In the above embodiment, the application discloses an image retrieval device, which belongs to the technical field of artificial intelligence and has an image retrieval system applied to an insurance marketing scene. The application combines the content features and position information of an image by segmenting and encoding a to-be-processed image, forms a high-quality image representation, and generates corresponding image index information by using a preset image index generator, so that efficient indexing and retrieval of the image are realized. In the encoding process, not only the local features of the image are retained, but also the spatial position information is fused, so that the discriminability and uniqueness of the image index are improved. The generated image index information is stored in association with the original image in an image retrieval database, so that the image retrieval process has high matching accuracy and response speed. In actual application, after receiving an image retrieval instruction, the system can quickly acquire an image matched with the target image index information from the database, and efficient and accurate image retrieval is realized. The application effectively improves the accuracy and processing efficiency of image retrieval, and is suitable for large-scale image library management and fast retrieval.

[0153] To solve the above technical problems, the application embodiment further provides a computer device. For details, please refer to Figure 4 , Figure 4 The basic structure block diagram of the computer device of the embodiment is shown in the figure.

[0154] The computer device 4 includes a memory 41, a processor 42 and a network interface 43 which are connected to each other through a system bus. It should be pointed out that only the computer device 4 with the memory 41, the processor 42 and the network interface 43 is shown in the figure, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented. Among them, those skilled in the art can understand that the computer device here is a device capable of automatically performing numerical calculation and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0155] The computer device can be a desktop computer, a notebook computer, a palm computer and a cloud server, etc. The computer device can interact with the user through a keyboard, a mouse, a remote controller, a touchpad or a voice control device, etc.

[0156] The memory 41 includes at least one type of readable storage medium, such as a flash memory, a hard disk, a multimedia card, a card-type memory (e.g., an SD or DX memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the memory 41 can be an internal storage unit of the computer device 4, such as a hard disk or a memory of the computer device 4. In other embodiments, the memory 41 can also be an external storage device of the computer device 4, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device 4. Of course, the memory 41 can also include both the internal storage unit and the external storage device of the computer device 4. In this embodiment, the memory 41 is generally used to store an operating system and various application software installed on the computer device 4, such as computer readable instructions of the image retrieval method, etc. In addition, the memory 41 can also be used to temporarily store various data that have been output or will be output.

[0157] The processor 42 can be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip in some embodiments. The processor 42 is generally used to control the overall operation of the computer device 4. In this embodiment, the processor 42 is used to run computer readable instructions or process data stored in the memory 41, such as computer readable instructions of the image retrieval method.

[0158] The network interface 43 can include a wireless network interface or a wired network interface, and is generally used to establish a communication connection between the computer device 4 and other electronic devices.

[0159] The present application also provides an embodiment, i.e., to provide a computer device, which includes a memory and a processor, the memory stores computer readable instructions, and the processor executes the computer readable instructions to implement the steps of the user insurance demand assessment method as described above, i.e., to implement:

[0160] An image retrieval method includes:

[0161] Obtaining a pre-collected image to be processed, and performing image segmentation on the image to be processed to obtain a plurality of image sub-images to be processed;

[0162] encoding the to-be-processed sub-image to obtain a first encoding vector;

[0163] adding image position information to each to-be-processed sub-image and encoding the image position information of the to-be-processed sub-image to obtain a second encoding vector;

[0164] inputting the first encoding vector and the second encoding vector into a preset image index generator to obtain image index information of the to-be-processed image;

[0165] associating the image index information with the to-be-processed image to obtain associated image data, and storing the associated image data into an image retrieval database;

[0166] in response to an image retrieval instruction, acquiring target image index information matched with a to-be-retrieved image, and searching the image retrieval database for a target image matched with the target image index information.

[0167] The application further provides another implementation, namely providing a computer readable storage medium storing computer readable instructions executable by at least one processor to enable the at least one processor to perform the steps of the image retrieval method as described above, namely to implement:

[0168] An image retrieval method comprises:

[0169] acquiring pre-collected to-be-processed images and performing image segmentation on the to-be-processed images to obtain a plurality of to-be-processed sub-images;

[0170] encoding the to-be-processed sub-image to obtain a first encoding vector;

[0171] adding image position information to each to-be-processed sub-image and encoding the image position information of the to-be-processed sub-image to obtain a second encoding vector;

[0172] inputting the first encoding vector and the second encoding vector into a preset image index generator to obtain image index information of the to-be-processed image;

[0173] associating the image index information with the to-be-processed image to obtain associated image data, and storing the associated image data into an image retrieval database;

[0174] in response to an image retrieval instruction, acquiring target image index information matched with a to-be-retrieved image, and searching the image retrieval database for a target image matched with the target image index information.

[0175] Those skilled in the art can clearly understand that the above-mentioned embodiment method can be realized by means of software and the necessary general hardware platform, of course, it can also be realized by hardware, but in many cases, the former is a better embodiment. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes a plurality of instructions for making a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device) execute the methods described in various embodiments of the present application.

[0176] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld devices or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in a distributed computing environment in which tasks are performed by remote processing devices connected by a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.

[0177] It should be noted that the non-company software tools or components appearing in various embodiments of the present application are only illustrative and do not represent actual use.

[0178] Obviously, the above-described embodiments are only part of the embodiments of the present application, and are not all the embodiments. The preferred embodiments of the present application are given in the drawings, but do not limit the patent scope of the present application. The present application can be realized in many different forms, and on the contrary, the purpose of providing these embodiments is to make the disclosure of the present application more thorough and comprehensive. Although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions recorded in the foregoing specific embodiments, or make equivalent replacements to some technical features. Any equivalent structure made by using the contents of the specification and drawings, directly or indirectly applied to other related technical fields, is also within the scope of the patent protection of the present application.

Claims

1. An image retrieval method characterized by, The method comprises the following steps: acquiring a pre-collected image to be processed, and performing image segmentation on the image to be processed to obtain a plurality of sub-images to be processed; performing image coding on the sub-images to be processed to obtain a first coding vector; adding image position information to each of the sub-images to be processed, and performing position coding on the image position information of the sub-images to be processed to obtain a second coding vector; inputting the first coding vector and the second coding vector into a preset image index generator to obtain image index information of the image to be processed; associating the image index information with the image to be processed to obtain associated image data, and storing the associated image data into an image retrieval database; in response to an image retrieval instruction, acquiring target image index information matched with an image to be retrieved, and searching the image retrieval database for a target image matched with the target image index information.

2. The image retrieval method of claim 1, wherein, The step of acquiring a pre-collected image to be processed, and performing image segmentation on the image to be processed to obtain a plurality of sub-images to be processed specifically comprises the following steps: determining any pixel point in the image to be processed as a pixel seed point; using a region growing method to identify whether adjacent pixel points of the pixel seed point match the pixel seed point; determining a sub-image region according to the pixel point matching result; performing image segmentation on the image to be processed according to the sub-image region to obtain a plurality of sub-images to be processed.

3. The image retrieval method of claim 2, wherein, The step of determining a sub-image region according to the pixel point matching result specifically comprises the following steps: when the adjacent pixel points match the pixel seed point, determining a region where the adjacent pixel points and the pixel seed point are located as a sub-image region; when the adjacent pixel points do not match the pixel seed point, acquiring non-matching pixel points; taking the non-matching pixel points as new pixel seed points; continuing to identify whether adjacent pixel points of the new pixel seed points match the new pixel seed points until all pixel points in the image to be processed participate in matching, thereby obtaining all sub-image regions.

4. The image retrieval method of claim 1, wherein, The step of performing image coding on the sub-images to be processed to obtain a first coding vector specifically comprises the following steps: flattening the sub-images to be processed into one-dimensional vectors to obtain sub-image initial features; performing dimension increasing and dimension decreasing operations on the sub-image initial features respectively to obtain sub-image multi-dimensional features; using a preset encoder to perform coding operation on the sub-image multi-dimensional features to obtain the first coding vector.

5. The image retrieval method of claim 1, wherein, The step of adding image position information to each of the sub-images to be processed, and performing position coding on the image position information of the sub-images to be processed to obtain a second coding vector specifically comprises the following steps: constructing a plane coordinate system on a plane where the image to be processed is located; acquiring image position information corresponding to each of the sub-images to be processed based on the plane coordinate system; associating each of the image position information with the corresponding sub-image to be processed; performing position coding on the image position information of the sub-images to be processed to obtain the second coding vector.

6. The image retrieval method of claim 1, wherein, The image index generator comprises a random inactivation layer and a full connection layer, the step of inputting the first encoding vector and the second encoding vector into a preset image index generator to obtain image index information of the to-be-processed image specifically comprises: randomly inactivating the first encoding vector using the random inactivation layer to obtain a first intermediate vector; linearly transforming the first intermediate vector using the full connection layer to obtain a first target vector; randomly inactivating the second encoding vector using the random inactivation layer to obtain a second intermediate vector; linearly transforming the second intermediate vector using the full connection layer to obtain a second target vector; performing hash value mapping on the first target vector and the second target vector to obtain a binary hash code; taking the binary hash code as the image index information of the to-be-processed image.

7. The image retrieval method of claim 6, wherein, The step of performing hash value mapping on the first target vector and the second target vector to obtain a binary hash code specifically comprises: calculating linear attention weight values of the first target vector and the second target vector by adopting a linear attention mechanism; performing weighted summation on the first target vector and the second target vector based on the linear attention weight values to obtain a weighted feature vector; importing the weighted feature vector into a preset hash mapping space, and completing hash value mapping of the weighted feature vector in the hash mapping space to obtain the binary hash code.

8. An image retrieval apparatus characterized by comprising: comprise: an image segmentation module configured to acquire a to-be-processed image collected in advance, and perform image segmentation on the to-be-processed image to obtain a plurality of to-be-processed sub-images; an image encoding module configured to perform image encoding on the to-be-processed sub-images to obtain a first encoding vector; a position encoding module configured to add image position information to each to-be-processed sub-image, and perform position encoding on the image position information of the to-be-processed sub-images to obtain a second encoding vector; an index generation module configured to input the first encoding vector and the second encoding vector into a preset image index generator to obtain image index information of the to-be-processed image; a data association module configured to associate the image index information with the to-be-processed image to obtain associated image data, and store the associated image data into an image retrieval database; an image retrieval module configured to acquire target image index information matched with a to-be-retrieved image in response to an image retrieval instruction, and find a target image matched with the target image index information in the image retrieval database.

9. A computer device, comprising: comprise a memory and a processor, the memory stores computer readable instructions, and the processor implements the steps of the image retrieval method according to any one of claims 1 to 7 when executing the computer readable instructions.

10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer readable instructions, and the computer readable instructions are executed by the processor to implement the steps of the image retrieval method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Binary code image retrieval method and system based on comparative learning

    CN117573915A

  • Method of training image-text retrieval model, method of multimodal image retrieval, electronic device and medium

    US20220391587A1

Cited By

  • Quick retrieval method and equipment for three-dimensional model and medium

    CN121808097A

  • A method, device and medium for fast retrieval of a three-dimensional model

    CN121808097B