An image retrieval method, apparatus, computer device, and storage medium

CN120832428BActive Publication Date: 2026-09-01PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510698854.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2026-09-01
Estimated Expiration
2045-05-27

AI Technical Summary

Technical Problem

[0004]本申请实施例的目的在于提出一种图像检索方法、装置、计算机设备及存储介质,以解决现有图像检索方法存在的计算开销大,性能较低的技术问题

Benefits of technology

[0026] This application discloses an image retrieval method, apparatus, computer equipment, and storage medium, belonging to the field of artificial intelligence technology, and features an image retrieval system applicable to insurance marketing scenarios. This application segments and encodes the image to be processed, combining the image's content features with its location information to form a high-quality image representation. A pre-defined image index generator then generates corresponding image index information, thereby achieving efficient image indexing and retrieval. During the encoding process, not only are local image features preserved, but spatial location information is also integrated, improving the discriminativeness and uniqueness of the image index. By associating the generated image index information with the original image and storing it in an image retrieval database, high matching accuracy and response speed are ensured during the image retrieval process. In practical applications, upon receiving an image retrieval command, the system can quickly retrieve an image from the database that matches the target image index information, achieving efficient and accurate image retrieval. This application effectively improves the accuracy and processing efficiency of image retrieval, and is suitable for large-scale image library management and rapid retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120832428B_ABST
    Figure CN120832428B_ABST
Patent Text Reader

Abstract

This application discloses an image retrieval method, apparatus, computer equipment, and storage medium, belonging to the field of artificial intelligence technology, and possessing an image retrieval system applicable to insurance marketing scenarios. This application segments and encodes the image to be processed, combining the image's content features with location information to form a high-quality image representation, and uses a preset image index generator to generate corresponding image index information, thereby achieving efficient image indexing and retrieval. During the encoding process, not only are local features of the image preserved, but spatial location information is also integrated, improving the discriminativeness and uniqueness of the image index. By associating the generated image index information with the original image and storing it in an image retrieval database, the image retrieval process ensures high matching accuracy and response speed, effectively improving the accuracy and processing efficiency of image retrieval, and is suitable for large-scale image library management and rapid retrieval.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of artificial intelligence technology, specifically relating to an image retrieval method, apparatus, computer equipment, and storage medium. Background Technology

[0002] Image materials play a crucial role in marketing campaigns. However, the diverse range of business types, varying focuses, and diverse sources of materials have led to an ever-growing image library, posing challenges to efficient retrieval and classification. Business personnel need to reprocess the retrieved information, reducing work efficiency. For example, in insurance marketing scenarios, health insurance requires preparing before-and-after comparison images of various illnesses and images of professional claims teams; car insurance requires showcasing images of accident scene investigations and rapid claims processing. The sheer volume of materials necessitates re-integration after retrieval, impacting work progress.

[0003] Traditional image retrieval methods primarily rely on manual labeling or description, which not only consumes significant human resources but is also prone to subjective bias, affecting retrieval accuracy. In recent years, with the development of artificial intelligence technology, some schemes have emerged that use deep learning networks (such as CNN convolutional models) for deep feature extraction from images. However, these schemes suffer from high computational overhead and relatively low performance on large-scale training datasets. Summary of the Invention

[0004] The purpose of this application is to provide an image retrieval method, apparatus, computer device, and storage medium to solve the technical problems of high computational overhead and low performance in existing image retrieval methods.

[0005] To address the aforementioned technical problems, this application provides an image retrieval method, employing the following technical solution:

[0006] An image retrieval method, comprising:

[0007] Acquire pre-collected images to be processed, and perform image segmentation on the images to be processed to obtain several sub-images to be processed;

[0008] Image encoding is performed on the sub-image to be processed to obtain the first encoding vector;

[0009] Image position information is added to each sub-image to be processed, and the image position information of the sub-image to be processed is positionally encoded to obtain a second encoding vector;

[0010] The first and second encoding vectors are input into a preset image index generator to obtain the image index information of the image to be processed;

[0011] The image index information is associated with the image to be processed to obtain associated image data, and the associated image data is stored in the image retrieval database;

[0012] In response to an image retrieval command, the system obtains the target image index information that matches the image to be retrieved, and searches the image retrieval database for the target image that matches the target image index information.

[0013] To address the aforementioned technical problems, this application also provides an image retrieval device, which employs the following technical solution:

[0014] An image retrieval device, comprising:

[0015] The image segmentation module is used to acquire pre-collected images to be processed and to segment the images to be processed into several sub-images to be processed.

[0016] The image encoding module is used to encode the sub-image to be processed to obtain the first encoding vector;

[0017] The position encoding module is used to add image position information to each sub-image to be processed, and to perform position encoding on the image position information of the sub-image to be processed to obtain a second encoding vector;

[0018] The index generation module is used to input the first encoding vector and the second encoding vector into a preset image index generator to obtain the image index information of the image to be processed;

[0019] The data association module is used to associate image index information with the image to be processed, obtain associated image data, and store the associated image data in the image retrieval database;

[0020] The image retrieval module is used to respond to image retrieval commands, obtain the target image index information that matches the image to be retrieved, and search for the target image that matches the target image index information in the image retrieval database.

[0021] To address the aforementioned technical problems, this application also provides a computer device that employs the following technical solution:

[0022] A computer device includes a memory and a processor, the memory storing computer-readable instructions, the processor executing the computer-readable instructions to implement the steps of the image retrieval method as described in any of the preceding claims.

[0023] To address the aforementioned technical problems, this application also provides a computer-readable storage medium, employing the technical solution described below:

[0024] A computer-readable storage medium storing computer-readable instructions, which, when executed by a processor, implement the steps of the image retrieval method as described in any one of the above descriptions.

[0025] Compared with the prior art, the embodiments of this application have the following main advantages:

[0026] This application discloses an image retrieval method, apparatus, computer equipment, and storage medium, belonging to the field of artificial intelligence technology, and features an image retrieval system applicable to insurance marketing scenarios. This application segments and encodes the image to be processed, combining the image's content features with its location information to form a high-quality image representation. A pre-defined image index generator then generates corresponding image index information, thereby achieving efficient image indexing and retrieval. During the encoding process, not only are local image features preserved, but spatial location information is also integrated, improving the discriminativeness and uniqueness of the image index. By associating the generated image index information with the original image and storing it in an image retrieval database, high matching accuracy and response speed are ensured during the image retrieval process. In practical applications, upon receiving an image retrieval command, the system can quickly retrieve an image from the database that matches the target image index information, achieving efficient and accurate image retrieval. This application effectively improves the accuracy and processing efficiency of image retrieval, and is suitable for large-scale image library management and rapid retrieval. Attached Figure Description

[0027] To more clearly illustrate the solutions in this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0028] Figure 1 An exemplary system architecture diagram is shown, in which this application can be applied;

[0029] Figure 2 A flowchart of one embodiment of the image retrieval method according to this application is shown;

[0030] Figure 3 A schematic diagram of one embodiment of the image retrieval apparatus according to this application is shown;

[0031] Figure 4 A schematic diagram of the structure of one embodiment of a computer device according to this application is shown. Detailed Implementation

[0032] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein in the specification of the application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application; the terms "comprising" and "having," and any variations thereof, in the specification, claims, and foregoing drawings of this application, are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the specification, claims, or foregoing drawings of this application are used to distinguish different objects, not to describe a particular order.

[0033] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0034] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.

[0035] like Figure 1 As shown, system architecture 100 may include terminal device 101, network 102, and server 103. Terminal device 101 may be a laptop 1011, tablet 1012, or mobile phone 1013. Network 102 is used as a medium to provide a communication link between terminal device 101 and server 103. Network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables.

[0036] Users can use terminal device 101 to interact with server 103 via network 102 to receive or send messages, etc. Various communication client applications can be installed on terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social media platform software, etc.

[0037] Terminal device 101 can be various electronic devices with a display screen and support web browsing. In addition to laptops 1011, tablets 1012, or mobile phones 1013, terminal device 101 can also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 player (Moving Picture Experts Group Audio Layer IV), a laptop computer, and a desktop computer, etc.

[0038] Server 103 can be a server that provides various services, such as a backend server that provides support for the pages displayed on terminal device 101.

[0039] It should be noted that the image retrieval method provided in this application embodiment is generally executed by a server / terminal device, and correspondingly, the image retrieval device is generally set in the server / terminal device.

[0040] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative; the system can have any number of terminal devices, networks, and servers depending on implementation needs.

[0041] Continue to refer to Figure 2 A flowchart of an embodiment of an image retrieval method according to this application is shown. The image retrieval method includes the following steps:

[0042] S201, acquire the pre-collected images to be processed, and perform image segmentation on the images to be processed to obtain several sub-images to be processed;

[0043] Specifically, the system first acquires the raw image to be processed from a pre-defined image acquisition module or business system. The image source can include a local media library, image acquisition equipment, or an external database. To facilitate processing via the Transformer network, the system divides the raw image using an image segmentation strategy consistent with Vision Transformer (ViT), dividing the entire image into several fixed-size image patches (sub-images). Each sub-image represents a small region of the image. This image segmentation method transforms a two-dimensional image into a one-dimensional patch sequence, making it easier to feed into the Transformer encoder for processing.

[0044] In this technical solution, ViT is used as the core image feature extraction model. ViT is a deep learning model based on the Transformer architecture. Unlike traditional convolutional neural networks (CNNs), it does not rely on convolution operations to extract local features. Instead, it divides the input image into several fixed-size image patches, flattens each patch, linearly maps it into a vector sequence, adds positional information, and then inputs the whole into the Transformer encoder. ViT models the relationships between different regions of the image globally through a self-attention mechanism, possessing excellent global perception capabilities. It can capture the dependencies between distant pixels and extract more complete and accurate image semantic information.

[0045] In this solution, two key optimizations were made to address the high computational cost and token redundancy issues of the native ViT model in image retrieval tasks: First, the number of input tokens was reduced by merging image patches, thereby reducing redundant computation; second, the standard multi-head attention mechanism was simplified to a linear attention mechanism, which significantly reduced computational complexity, improved processing efficiency, and ensured the requirement for feature expressiveness in image retrieval.

[0046] Compared to the sliding window approach of traditional CNNs, this segmentation method preserves local information while providing a foundation for subsequent global modeling. Furthermore, to further improve model efficiency, this technique introduces a patch merging strategy, merging patches from adjacent or similar regions during segmentation. This reduces the number of tokens without significantly losing feature information, effectively alleviating the computational burden in subsequent processing.

[0047] S202, Image encoding is performed on the sub-image to be processed to obtain the first encoding vector;

[0048] Specifically, the segmented image patches are fed as input into the image encoding module, which is built on an improved Vision Transformer (ViT) architecture. Each sub-image is first unfolded into a vector form, and then projected onto a uniform feature dimension through a linear mapping layer to form a basic feature vector.

[0049] Unlike traditional ViT, this technique employs a linear attention mechanism in the encoder section instead of the standard multi-head self-attention to improve computational efficiency and adapt to image retrieval tasks, significantly reducing complexity. Linear attention replaces the fully connected attention matrix with an inner product, avoiding the token inflation and high computational cost issues inherent in the original ViT when processing high-resolution images. This design allows for the modeling of global relationships between image patches while greatly improving encoding efficiency. The first encoded vector output from this step represents the semantic content information of each sub-image in the image space.

[0050] S203, add image position information to each sub-image to be processed, and perform position encoding on the image position information of the sub-image to be processed to obtain the second encoding vector;

[0051] Specifically, the ViT model itself lacks the spatial inductive bias inherent in CNNs. Therefore, to enable the model to understand the spatial location of each sub-image, it is necessary to append its two-dimensional positional coordinates in the original image to each patch. This positional information is first represented in absolute position (such as row and column indices or normalized coordinates), and then transformed through a positional encoding function to obtain a second encoding vector with the same dimension as the first encoding vector. Positional encoding can use a fixed sine function encoding or a trained positional embedding vector. This technique prioritizes trainable positional encoding methods based on task characteristics to enhance the model's ability to learn spatial structure. These positional vectors will be fused with the first encoding vector in subsequent encodings to provide spatial positional information support for the model. Through this step, the model can not only understand the semantic features of image content when processing features, but also grasp the spatial layout of each sub-image in the entire image, making the final generated image index more discriminative and stable, thereby improving the matching accuracy and robustness in image retrieval tasks.

[0052] S204, input the first encoding vector and the second encoding vector into the preset image index generator to obtain the image index information of the image to be processed;

[0053] Specifically, the content encoding (first encoding vector) and position encoding (second encoding vector) of an image are fused. This can be achieved through vector addition, concatenation, or weighted fusion using attention mechanisms, forming a complete set of image representation features. The fused features are then input into a custom-designed image index generator. This generator consists of a random deactivation layer (Dropout layer) and two fully connected (Dense Layers). The Dropout layer prevents overfitting and enhances the model's generalization ability. The high-dimensional features are then mapped to a lower-dimensional latent space, and a fixed-length binary hash code is output. This hash code serves as the image's index information, significantly compressing storage space and supporting fast similarity calculations (such as Hamming distance matching), making it particularly suitable for use in large-scale image retrieval databases. The entire process generates unique, compact, and efficiently matching index identifiers for images, balancing accuracy and computational efficiency.

[0054] S205, associate the image index information with the image to be processed to obtain associated image data, and store the associated image data in the image retrieval database;

[0055] Specifically, the system establishes a one-to-one correspondence between the image index information generated in the previous step and the original images, forming image-index pairs. These associated image data are then uniformly formatted and written into the image retrieval database. The image retrieval database supports high-performance hash code index management structures, such as indexing schemes based on inverted indexes, Locality Sensitive Hash (LSH), or Product Quantization (PQ), ensuring fast searching even among a large number of images. The completion of this step signifies that the images have the capability for rapid retrieval, and the image library has achieved structured management.

[0056] In addition, to support multi-scenario retrieval needs, the database can also synchronously store image-related metadata, such as business tags, image categories, upload time, and usage frequency, which facilitates multi-dimensional filtering and sorting during retrieval.

[0057] S206, responding to the image retrieval command, obtain the target image index information that matches the image to be retrieved, and search for the target image that matches the target image index information in the image retrieval database.

[0058] Specifically, upon receiving a retrieval request, the system re-encodes the image to be retrieved according to the process from S201 to S204 to generate its index information. This index information is input into the image retrieval database as the query keyword. Through similarity calculation between hash codes (e.g., calculating Hamming distance or approximate nearest neighbor matching), the system quickly locates the index of the closest target image. After a successful match, the system further extracts the corresponding target image and its related information from the database and returns the result to the user. Due to the use of a compact and efficient hash index, the retrieval speed is significantly improved, supporting second-level response times. Simultaneously, because the index integrates image content and spatial location information, the matching accuracy is higher. This retrieval process adopts an end-to-end single-stage architecture, avoiding the problems of coarse-sorting error accumulation and redundant calculations in traditional two-stage retrieval schemes. This improves the overall system's practicality and user experience, making it particularly suitable for industry scenarios with large image volumes and frequent retrievals, such as insurance marketing, customer service, and image content auditing.

[0059] In the above embodiments, this application segments and encodes the image to be processed, combining the image's content features with its location information to form a high-quality image representation. A pre-defined image index generator then generates corresponding image index information, thereby achieving efficient image indexing and retrieval. During the encoding process, not only are local image features preserved, but spatial location information is also integrated, improving the discriminativeness and uniqueness of the image index. By associating the generated image index information with the original image and storing it in the image retrieval database, high matching accuracy and response speed are ensured during the image retrieval process. In practical applications, upon receiving an image retrieval command, the system can quickly retrieve an image from the database that matches the target image index information, achieving efficient and accurate image retrieval. This application effectively improves the accuracy and processing efficiency of image retrieval and is suitable for large-scale image library management and rapid retrieval.

[0060] For example, in insurance marketing scenarios, sales personnel often need to retrieve and display a large amount of image materials related to different types of insurance. For instance, health insurance requires images of physical examinations, hospitals, and service processes; while auto insurance requires images of accident scenes, vehicle damage, and claims services. In practice, due to the sheer volume of image materials, manual sorting and searching are inefficient and can easily disrupt customer communication and business progress.

[0061] Taking a car insurance accident claim scenario as an example, when an image of an accident vehicle (such as a rear-end collision scene) is uploaded to the system, the image first undergoes image segmentation via a preprocessing module. For instance, the system automatically identifies key regions in the image, such as the damaged front of the vehicle, ground scratches, and the surrounding environment, using a region growing algorithm, and segments them into multiple sub-images. Each sub-image is flattened, encoded, and combined with its spatial location information within the image, features are extracted using an improved ViT model. The system then inputs the extracted content features and location information into an index generator, which, after processing through a Dr opout layer and a fully connected layer, generates structured first and second target vectors.

[0062] Next, a linear attention mechanism is introduced to fuse content and location information to obtain a weighted feature vector, which is then mapped to a hash space to generate a binary hash code such as "011010101011...". This hash code is the unique index information of the image. The index is stored in the image retrieval database and corresponds one-to-one with the image sample.

[0063] Further, the steps of acquiring pre-collected images to be processed and performing image segmentation on these images to obtain several sub-images to be processed specifically include:

[0064] Determine any pixel in the image to be processed as a pixel seed point;

[0065] The region growing method is used to identify whether the adjacent pixels of a pixel seed point match the pixel seed point.

[0066] Sub-image regions are determined based on pixel matching results;

[0067] The image to be processed is segmented according to the sub-image regions to obtain several sub-images to be processed.

[0068] In this embodiment, image segmentation no longer employs the traditional fixed-size partitioning method. Instead, a content-aware region growing algorithm is introduced to generate sub-image regions, thereby obtaining image fragments with greater semantic consistency. Specifically, firstly, a pixel in the image is randomly selected as the initial seed pixel. This seed pixel can be initialized based on image brightness, color histogram, or texture features. Subsequently, the region growing algorithm is used to progressively expand the neighborhood region of the seed pixel. It is determined whether adjacent pixels are similar to the seed pixel in a specific feature space (e.g., color difference below a threshold, consistent texture direction, etc.). If a match is found, the pixels are included in the same region, and the expansion continues until the similarity condition can no longer be met, forming a complete sub-image region. This process can be repeated until the entire image is divided into several semantically related sub-images with clear boundaries.

[0069] Compared to fixed-size partitioning, this method better reflects the actual structural information in the image, which can improve the effectiveness and expressive power of subsequent feature extraction and provide more representative input data for the ViT encoder.

[0070] Through the above steps, fine segmentation based on image content semantics is achieved, improving the structural consistency and feature representation ability of sub-images.

[0071] Furthermore, the step of determining the sub-image region based on the pixel matching results specifically includes:

[0072] When an adjacent pixel matches a pixel seed point, the region containing the adjacent pixel and the pixel seed point is defined as a sub-image region.

[0073] When an adjacent pixel does not match the pixel seed point, the mismatched pixel is obtained.

[0074] Use the mismatched pixels as new pixel seed points;

[0075] Continue identifying whether the neighboring pixels of the new pixel seed point match the new pixel seed point until all pixels in the image to be processed participate in the matching, thus obtaining all sub-image regions.

[0076] In this embodiment, the system performs adaptive image segmentation based on the region growing method. The key lies in using a seed pixel as the core and iteratively expanding to form multiple semantically consistent sub-image regions. First, an initial pixel is selected from the image as the seed pixel. Its neighboring pixels are identified for similarity in dimensions such as color, brightness, and texture direction. Once a match is found, the neighboring pixel is incorporated into the current sub-image region. Each newly added pixel within the region continues to serve as an expansion source point, gradually expanding the region until no more matching neighboring pixels can be added. Simultaneously, pixels that do not match the current region are independently extracted as new seed pixels, and the region expansion process is repeated until all pixels in the entire image are assigned to a specific sub-image region.

[0077] Through the above steps, full-coverage content segmentation of each pixel in the image is achieved, improving the semantic consistency and structural accuracy of sub-image regions.

[0078] Further, the step of image encoding the sub-image to be processed to obtain the first encoded vector specifically includes:

[0079] Flatten the sub-image to be processed into a one-dimensional vector to obtain the initial features of the sub-image;

[0080] The initial features of the sub-image are subjected to dimensionality-up operations and dimensionality-down operations respectively to obtain the multi-dimensional features of the sub-image;

[0081] The multidimensional features of the sub-image are encoded using a preset encoder to obtain the first encoded vector.

[0082] In this embodiment, to fully extract the feature information from the sub-images to be processed, the system first flattens each sub-image into a one-dimensional vector to obtain its initial feature representation. The flattening operation preserves the basic structural information of the original pixel arrangement. Next, the system performs dimensionality increase and decrease operations on this one-dimensional vector. The dimensionality increase operation introduces a linear transformation (such as a fully connected layer) to project the feature vector into a higher-dimensional feature space, thereby increasing its expressive power and enabling the model to learn more complex and detailed feature relationships. The dimensionality decrease operation is used to suppress redundant information, improve subsequent computational efficiency by compressing the feature dimension, and avoid overfitting and resource waste caused by high-dimensional features. This dimensionality increase and decrease process achieves a balance between the expressive power and computational cost of the sub-image features. Subsequently, the system inputs the processed multi-dimensional features into a preset encoder. This encoder is based on the Vision Transformer (ViT) architecture and utilizes its powerful self-attention mechanism to extract deep feature relationships of the sub-images globally, thereby outputting a high-quality first encoded vector.

[0083] The above steps enhance the richness and abstractness of sub-image feature representation, providing a solid foundation for the accuracy of image indexing.

[0084] Further, the steps of adding image position information to each sub-image to be processed and performing position encoding on the image position information of the sub-image to obtain the second encoding vector specifically include:

[0085] Construct a planar coordinate system in the plane containing the image to be processed;

[0086] Obtain the image position information corresponding to each sub-image to be processed based on the planar coordinate system;

[0087] Associate the location information of each image with the corresponding sub-image to be processed;

[0088] The image location information of the sub-image to be processed is positionally encoded to obtain the second encoding vector.

[0089] In this embodiment, to preserve the spatial relationships between sub-images, the system introduces a location information processing step after image segmentation. First, a unified two-dimensional coordinate system is established on the plane containing the entire image to be processed. Each sub-image has a unique location identifier within this coordinate system, typically represented by the coordinates of its top-left corner or center point in the original image. The system then generates a location identifier for each sub-image based on these coordinates and associates this location with the corresponding sub-image features. Next, these image location information undergo location encoding to adapt to the Transformer structure's requirements for sequential input. Location encoding can employ sine / cosine function encoding, learnable location embedding, or a hybrid approach, mapping the spatial location information into a vector form with the same dimension as the image feature vector. This location encoding result is the second encoding vector, which guides the attention mechanism to focus on different image regions, enabling the model to not only focus on content features but also perceive their spatial layout, thereby more effectively capturing the relationships between global structure and local dependencies in the image.

[0090] Through the above steps, the spatial structure information of the image is explicitly expressed, enhancing the model's ability to understand the global relationships of the image.

[0091] Furthermore, the image index generator includes a random deactivation layer and a fully connected layer. The step of inputting the first encoding vector and the second encoding vector into the preset image index generator to obtain the image index information of the image to be processed specifically includes:

[0092] The first encoded vector is randomly deactivated using a random deactivation layer to obtain the first intermediate vector.

[0093] The first intermediate vector is linearly transformed using a fully connected layer to obtain the first target vector.

[0094] The second encoded vector is randomly deactivated using a random deactivation layer to obtain the second intermediate vector.

[0095] The second intermediate vector is linearly transformed using a fully connected layer to obtain the second target vector.

[0096] The first and second target vectors are mapped using hash values ​​to obtain binary hash codes.

[0097] Use the binary hash code as the image index information of the image to be processed.

[0098] In this embodiment, to construct efficient and discriminative image indexing information, the system designs an image index generator based on a deep neural network, mainly including a random deactivation layer (Dropout) and a fully connected layer (DenseLayer). This structure not only helps improve the generalization ability of the model but also compresses the feature dimension and enhances the compactness of the index. Specifically, the first encoded vector output by the ViT encoder is first input to the Dropout layer. Random deactivation can effectively reduce the dependencies between neurons, prevent the model from overfitting during training, and thus improve the overall robustness. Subsequently, the deactivated first intermediate vector is fed into the Dense Layer for linear transformation to extract higher-level abstract features, forming the first target vector. At the same time, the same processing flow is performed on the position-encoded second encoded vector, obtaining the second target vector through random deactivation and linear transformation. Next, the system fuses the features of the first and second target vectors and performs a binary hash mapping operation through a specific hash function to generate a fixed-length binary hash code. This hash code retains the content features of the image and embeds its spatial location information, representing a compact expression of the global semantic and structural information of the image. Ultimately, this hash code serves as the unique index information for the image, enabling fast matching and image retrieval operations, while also possessing good storage efficiency and search performance.

[0099] Through the above steps, a compact hash index that combines content recognition and spatial structure awareness is constructed, improving the accuracy of image retrieval and the system response speed.

[0100] Further, the step of mapping the first target vector and the second target vector to obtain the binary hash code specifically includes:

[0101] The linear attention mechanism is used to calculate the linear attention weights of the first and second target vectors;

[0102] The first target vector and the second target vector are weighted and summed based on the linear attention weight values ​​to obtain a weighted feature vector.

[0103] The weighted feature vector is imported into a preset hash mapping space, and the hash value mapping of the weighted feature vector is completed in the hash mapping space to obtain the binary hash code.

[0104] In this embodiment, the system employs a linear attention mechanism to efficiently fuse image content features and location information, thereby generating a discriminative binary hash code. First, for the first target vector (representing image content features) and the second target vector (representing image location information), the system calculates the attention weights between them using a linear attention mechanism. Unlike traditional self-attention, linear attention reduces complexity from quadratic to linear, significantly reducing computational resource consumption while retaining crucial context-dependent modeling capabilities. Based on these weight values, the system performs a weighted summation of the two target vectors, fusing their core information in image semantics and spatial location to obtain a unified weighted feature vector. Subsequently, this weighted feature vector is mapped to a preset hash space, which can be constructed using a trainable hash layer or by converting floating-point vectors into discrete binary representations using a set projection matrix and threshold function. Finally, based on the distribution of the weighted features in the hash mapping space, the system generates a compact and semantically rich binary hash code, which serves as a unique index identifier for the image and is used for image retrieval.

[0105] Through the above steps, deep fusion of image semantics and spatial information is achieved and transformed into an efficient and matchable hash index, thereby improving the system's response efficiency and accuracy in large-scale image retrieval scenarios.

[0106] In the above embodiments, this application discloses an image retrieval method, belonging to the field of artificial intelligence technology, and has an image retrieval system applied to insurance marketing scenarios. This application segments and encodes the image to be processed, combining the image's content features with location information to form a high-quality image representation, and uses a preset image index generator to generate corresponding image index information, thereby achieving efficient image indexing and retrieval. During the encoding process, not only are local image features preserved, but spatial location information is also integrated, improving the discriminativeness and uniqueness of the image index. By associating the generated image index information with the original image and storing it in an image retrieval database, the image retrieval process ensures high matching accuracy and response speed. In practical applications, upon receiving an image retrieval command, the system can quickly retrieve an image from the database that matches the target image index information, achieving efficient and accurate image retrieval. This application effectively improves the accuracy and processing efficiency of image retrieval, and is suitable for large-scale image library management and rapid retrieval.

[0107] In this embodiment, the image retrieval method runs on an electronic device (e.g., Figure 1The server shown can receive instructions or acquire data via wired or wireless connection. It should be noted that the aforementioned wireless connection methods may include, but are not limited to, 3G / 4G connections, WiFi connections, Bluetooth connections, WiMAX connections, Zigbee connections, UWB (ultra-wideband) connections, and other currently known or future wireless connection methods.

[0108] It should be emphasized that, to further ensure the privacy and security of the aforementioned image information, the image information can also be stored in a blockchain node.

[0109] The blockchain referred to in this application is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Essentially, a blockchain is a decentralized database, a chain of data blocks linked together using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity of the information (anti-counterfeiting) and generate the next block. A blockchain can include an underlying blockchain platform, a platform product service layer, and an application service layer.

[0110] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.

[0111] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0112] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by instructing related hardware through computer-readable instructions. These computer-readable instructions can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the methods described above. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, optical disk, or read-only memory (ROM), or random access memory (RAM).

[0113] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.

[0114] Further reference Figure 3 As a response to the above Figure 2 To implement the method shown, this application provides an embodiment of an image retrieval device, which is similar to... Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.

[0115] like Figure 3 As shown, the image retrieval device 300 described in this embodiment includes:

[0116] The image segmentation module 301 is used to acquire pre-collected images to be processed and to segment the images to be processed to obtain several sub-images to be processed.

[0117] Image encoding module 302 is used to encode the sub-image to be processed to obtain a first encoding vector;

[0118] The position encoding module 303 is used to add image position information to each sub-image to be processed, and to perform position encoding on the image position information of the sub-image to be processed to obtain a second encoding vector;

[0119] The index generation module 304 is used to input the first encoding vector and the second encoding vector into a preset image index generator to obtain the image index information of the image to be processed;

[0120] The data association module 305 is used to associate image index information with the image to be processed, obtain associated image data, and store the associated image data in the image retrieval database;

[0121] The image retrieval module 306 is used to respond to image retrieval commands, obtain target image index information that matches the image to be retrieved, and search for target images that match the target image index information in the image retrieval database.

[0122] Furthermore, the image segmentation module 301 specifically includes:

[0123] The pixel seed point unit is used to determine any pixel in the image to be processed as a pixel seed point;

[0124] The pixel matching unit is used to identify whether the neighboring pixels of the pixel seed point match the pixel seed point using the region growing method.

[0125] Image region unit, used to determine sub-image regions based on pixel matching results;

[0126] The image segmentation unit is used to segment the image to be processed according to the sub-image regions, and obtain several sub-images to be processed.

[0127] Furthermore, the image region unit specifically includes:

[0128] The first matching subunit is used to determine the area where the adjacent pixel and the pixel seed point are located as a sub-image region when the adjacent pixel is matched with the pixel seed point.

[0129] The second matching subunit is used to obtain the mismatched pixel when the adjacent pixel does not match the pixel seed point;

[0130] A new pixel seed point subunit is used to take mismatched pixels as new pixel seed points;

[0131] The continuous matching sub-unit is used to continue identifying whether the neighboring pixels of the new pixel seed point match the new pixel seed point, until all pixels in the image to be processed participate in the matching, thus obtaining all sub-image regions.

[0132] Furthermore, the image encoding module 302 specifically includes:

[0133] The image flattening unit is used to flatten the sub-image to be processed into a one-dimensional vector to obtain the initial features of the sub-image.

[0134] The dimension transformation unit is used to perform dimension up and dimension down operations on the initial features of the sub-image to obtain multi-dimensional features of the sub-image.

[0135] The image coding unit is used to encode the multidimensional features of the sub-image using a preset encoder to obtain the first coding vector.

[0136] Furthermore, the location encoding module 303 specifically includes:

[0137] A planar coordinate system unit is used to construct a planar coordinate system in the plane where the image to be processed is located.

[0138] The position acquisition unit is used to acquire the image position information corresponding to each sub-image to be processed based on the planar coordinate system.

[0139] The image association unit is used to associate the location information of each image with the corresponding sub-image to be processed;

[0140] The position encoding unit is used to encode the image position information of the sub-image to be processed, and obtain the second encoding vector.

[0141] Furthermore, the image index generator includes a random deactivation layer and a fully connected layer, and the index generation module 304 specifically includes:

[0142] The first deactivation processing unit is used to perform random deactivation processing on the first encoded vector using a random deactivation layer to obtain the first intermediate vector;

[0143] The first linear transformation unit is used to perform a linear transformation on the first intermediate vector using a fully connected layer to obtain the first target vector;

[0144] The second deactivation processing unit is used to perform random deactivation processing on the second encoded vector using a random deactivation layer to obtain the second intermediate vector.

[0145] The second linear transformation unit is used to perform a linear transformation on the second intermediate vector using a fully connected layer to obtain the second target vector.

[0146] The hash mapping unit is used to map the hash values ​​of the first target vector and the second target vector to obtain a binary hash code.

[0147] The image indexing unit is used to use binary hash codes as image index information for the image to be processed.

[0148] Furthermore, the hash mapping unit specifically includes:

[0149] The linear attention subunit is used to calculate the linear attention weights of the first target vector and the second target vector using a linear attention mechanism.

[0150] The weighted summation subunit is used to sum the first target vector and the second target vector based on the linear attention weight values ​​to obtain a weighted feature vector;

[0151] The hash mapping subunit is used to import the weighted feature vector into the preset hash mapping space, and complete the hash value mapping of the weighted feature vector within the hash mapping space to obtain the binary hash code.

[0152] In the above embodiments, this application discloses an image retrieval device, belonging to the field of artificial intelligence technology, which has an image retrieval system applied to insurance marketing scenarios. This application segments and encodes the image to be processed, combining the image's content features with location information to form a high-quality image representation, and uses a preset image index generator to generate corresponding image index information, thereby achieving efficient image indexing and retrieval. During the encoding process, not only are local features of the image preserved, but spatial location information is also integrated, improving the discriminativeness and uniqueness of the image index. By associating the generated image index information with the original image and storing it in an image retrieval database, the image retrieval process ensures high matching accuracy and response speed. In practical applications, upon receiving an image retrieval command, the system can quickly retrieve an image from the database that matches the target image index information, achieving efficient and accurate image retrieval. This application effectively improves the accuracy and processing efficiency of image retrieval, and is suitable for large-scale image library management and rapid retrieval.

[0153] To address the aforementioned technical problems, embodiments of this application also provide a computer device. Please refer to [link / reference] for details. Figure 4 , Figure 4 This is a basic structural block diagram of the computer device in this embodiment.

[0154] The computer device 4 includes a memory 41, a processor 42, and a network interface 43 that are interconnected via a system bus. It should be noted that only the computer device 4 with memory 41, processor 42, and network interface 43 is shown in the figure; however, it should be understood that it is not required to implement all the components shown, and more or fewer components can be implemented alternatively. Those skilled in the art will understand that the computer device described here is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0155] The computer device can be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device can interact with the user via a keyboard, mouse, remote control, touchpad, or voice control.

[0156] The memory 41 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 41 may be an internal storage unit of the computer device 4, such as the hard disk or memory of the computer device 4. In other embodiments, the memory 41 may also be an external storage device of the computer device 4, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device 4. Of course, the memory 41 may also include both the internal storage unit and its external storage device of the computer device 4. In this embodiment, the memory 41 is typically used to store the operating system and various application software installed on the computer device 4, such as computer-readable instructions for image retrieval methods. In addition, the memory 41 can also be used to temporarily store various types of data that have been output or will be output.

[0157] In some embodiments, the processor 42 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor 42 is typically used to control the overall operation of the computer device 4. In this embodiment, the processor 42 is used to execute computer-readable instructions stored in the memory 41 or to process data, for example, to execute computer-readable instructions for the image retrieval method.

[0158] The network interface 43 may include a wireless network interface or a wired network interface, which is typically used to establish communication connections between the computer device 4 and other electronic devices.

[0159] This application also provides an implementation method, namely, a computer device including a memory and a processor. The memory stores computer-readable instructions, and the processor, when executing the computer-readable instructions, implements the steps of the user insurance demand assessment method described above, that is, it implements:

[0160] An image retrieval method, comprising:

[0161] Acquire pre-collected images to be processed, and perform image segmentation on the images to be processed to obtain several sub-images to be processed;

[0162] Image encoding is performed on the sub-image to be processed to obtain the first encoding vector;

[0163] Image position information is added to each sub-image to be processed, and the image position information of the sub-image to be processed is positionally encoded to obtain a second encoding vector;

[0164] The first and second encoding vectors are input into a preset image index generator to obtain the image index information of the image to be processed;

[0165] The image index information is associated with the image to be processed to obtain associated image data, and the associated image data is stored in the image retrieval database;

[0166] In response to an image retrieval command, the system obtains the target image index information that matches the image to be retrieved, and searches the image retrieval database for the target image that matches the target image index information.

[0167] This application also provides another embodiment, namely, a computer-readable storage medium storing computer-readable instructions that can be executed by at least one processor to cause the at least one processor to perform the steps of the image retrieval method described above, i.e., to implement:

[0168] An image retrieval method, comprising:

[0169] Acquire pre-collected images to be processed, and perform image segmentation on the images to be processed to obtain several sub-images to be processed;

[0170] Image encoding is performed on the sub-image to be processed to obtain the first encoding vector;

[0171] Image position information is added to each sub-image to be processed, and the image position information of the sub-image to be processed is positionally encoded to obtain a second encoding vector;

[0172] The first and second encoding vectors are input into a preset image index generator to obtain the image index information of the image to be processed;

[0173] The image index information is associated with the image to be processed to obtain associated image data, and the associated image data is stored in the image retrieval database;

[0174] In response to an image retrieval command, the system obtains the target image index information that matches the image to be retrieved, and searches the image retrieval database for the target image that matches the target image index information.

[0175] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0176] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics devices, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0177] It should be noted that the software tools or components not belonging to this company that appear in the various embodiments of this application are merely illustrative examples and do not represent actual use.

[0178] Obviously, the embodiments described above are only some embodiments of this application, not all embodiments. The accompanying drawings show preferred embodiments of this application, but do not limit the patent scope of this application. This application can be implemented in many different forms; rather, the purpose of providing these embodiments is to provide a more thorough and comprehensive understanding of the disclosure of this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments, or make equivalent substitutions for some of the technical features. Any equivalent structures made using the content of this application's specification and drawings, directly or indirectly applied to other related technical fields, are similarly within the scope of patent protection of this application.

Claims

1. An image retrieval method, characterized in that, include: Acquire pre-collected images to be processed, and perform image segmentation on the images to be processed to obtain several sub-images to be processed; The sub-image to be processed is image encoded to obtain a first encoding vector; Image position information is added to each of the sub-images to be processed, and the image position information of the sub-images to be processed is positionally encoded to obtain a second encoding vector; The first encoding vector and the second encoding vector are input into a preset image index generator to obtain the image index information of the image to be processed; The image index information is associated with the image to be processed to obtain associated image data, and the associated image data is stored in the image retrieval database; In response to an image retrieval command, the system obtains the target image index information that matches the image to be retrieved, and searches for the target image that matches the target image index information in the image retrieval database. The image index generator includes a random deactivation layer and a fully connected layer. The step of inputting the first encoding vector and the second encoding vector into the preset image index generator to obtain the image index information of the image to be processed specifically includes: The first encoded vector is randomly deactivated using the random deactivation layer to obtain the first intermediate vector. The first intermediate vector is linearly transformed using the fully connected layer to obtain the first target vector. The second encoded vector is randomly deactivated using the random deactivation layer to obtain the second intermediate vector; The second intermediate vector is linearly transformed using the fully connected layer to obtain the second target vector. The first target vector and the second target vector are mapped using hash values ​​to obtain binary hash codes; The binary hash code is used as the image index information of the image to be processed. The step of mapping the first target vector and the second target vector to obtain a binary hash code specifically includes: The linear attention mechanism is used to calculate the linear attention weight values ​​of the first target vector and the second target vector; Based on the linear attention weight values, the first target vector and the second target vector are weighted and summed to obtain a weighted feature vector; The weighted feature vector is imported into a preset hash mapping space, and the hash value mapping of the weighted feature vector is completed in the hash mapping space to obtain the binary hash code.

2. The image retrieval method as described in claim 1, characterized in that, The steps of acquiring pre-collected images to be processed and performing image segmentation on the images to be processed to obtain several sub-images to be processed specifically include: Any pixel in the image to be processed is determined as a pixel seed point; The region growing method is used to identify whether the neighboring pixels of the pixel seed point match the pixel seed point. Sub-image regions are determined based on pixel matching results; The image to be processed is segmented according to the sub-image regions to obtain several sub-images to be processed.

3. The image retrieval method as described in claim 2, characterized in that, The step of determining the sub-image region based on the pixel matching result specifically includes: When the adjacent pixel matches the pixel seed point, the area where the adjacent pixel and the pixel seed point are located is determined as a sub-image region; When the adjacent pixel does not match the pixel seed point, the mismatched pixel is obtained; Use the mismatched pixels as new pixel seed points; Continue identifying whether the neighboring pixels of the new pixel seed point match the new pixel seed point until all pixels in the image to be processed participate in the matching, thus obtaining all sub-image regions.

4. The image retrieval method as described in claim 1, characterized in that, The step of encoding the sub-image to be processed to obtain the first encoding vector specifically includes: Flatten the sub-image to be processed into a one-dimensional vector to obtain the initial features of the sub-image; Perform dimensionality increase and dimensionality decrease operations on the initial features of the sub-image to obtain multi-dimensional features of the sub-image; The first encoded vector is obtained by encoding the multidimensional features of the sub-image using a preset encoder.

5. The image retrieval method as described in claim 1, characterized in that, The step of adding image position information to each of the sub-images to be processed and performing position encoding on the image position information of the sub-images to be processed to obtain a second encoding vector specifically includes: Construct a planar coordinate system on the plane where the image to be processed is located; Based on the planar coordinate system, obtain the image position information corresponding to each of the sub-images to be processed; Associate each of the image location information with the corresponding sub-image to be processed; The image position information of the sub-image to be processed is positionally encoded to obtain the second encoding vector.

6. An image retrieval device, characterized in that, The image retrieval device implements the steps of the image retrieval method as described in any one of claims 1 to 5, and the image retrieval device comprises: The image segmentation module is used to acquire pre-collected images to be processed and to segment the images to be processed to obtain several sub-images to be processed. The image encoding module is used to encode the sub-image to be processed to obtain a first encoding vector; The position encoding module is used to add image position information to each of the sub-images to be processed, and to perform position encoding on the image position information of the sub-images to be processed to obtain a second encoding vector; An index generation module is used to input the first encoding vector and the second encoding vector into a preset image index generator to obtain the image index information of the image to be processed; The data association module is used to associate the image index information with the image to be processed to obtain associated image data, and store the associated image data in the image retrieval database; The image retrieval module is used to respond to image retrieval commands, obtain target image index information that matches the image to be retrieved, and search for target images that match the target image index information in the image retrieval database.

7. A computer device, characterized in that, The device includes a memory and a processor, wherein the memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the steps of the image retrieval method as described in any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the image retrieval method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Binary code image retrieval method and system based on comparative learning

    CN117573915A

  • Method of training image-text retrieval model, method of multimodal image retrieval, electronic device and medium

    US20220391587A1