Image retrieval method, electronic device, computer readable medium and computer program product

By aligning feature vectors in the same semantic space through the lightweight MobileCLIP model, the problems of large model parameters and high training cost in Chinese semantic image search are solved, and efficient and accurate Chinese image retrieval is achieved.

CN120104826BActive Publication Date: 2025-10-24ZTE CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510586248.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-10-24
Estimated Expiration
2045-05-08

AI Technical Summary

Technical Problem

In the existing technology, the Chinese semantic image search method needs to go through a semantic conversion model, a text encoding model, and an image encoding model, resulting in a large number of model parameters, high training costs, and low image search accuracy and efficiency.

Method used

Using the lightweight MobileCLIP model as the basis, the feature vectors of the first language encoding model and the image encoding model are aligned in the same semantic space, and image retrieval is performed directly, which reduces the number of models and parameters and reduces training costs.

Benefits of technology

It improves the accuracy and efficiency of Chinese semantic image search, reduces model training costs, is suitable for embedded devices, and enhances user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104826B_ABST
    Figure CN120104826B_ABST
Patent Text Reader

Abstract

The disclosure provides an image retrieval method, first language information to be retrieved is input into a pre-trained first language encoding model, a first feature vector of the first language information is extracted, the first feature vector output by the first language encoding model for the first language is semantically aligned with a second feature vector output by a second language encoding model for the second language in the same semantic space; the first feature vector is input into an image encoding model, and a first target image set corresponding to the first feature vector is queried from a target database; the image encoding model is trained according to images in the target database and texts in the second language corresponding to the images; the accuracy and efficiency of image retrieval of the first language are improved, the number of models and model parameters involved are less, and the model training cost is reduced. The disclosure also provides an electronic device, a computer readable medium and a computer program product.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of computer, in particular to an image retrieval method, an electronic device, a computer readable medium and a computer program product. BACKGROUND

[0002] The fusion product of FTTR (Fiber To The Room) and NAS (Network Attached Storage) supports not only local high-speed data storage, but also synchronization with cloud disk to realize end-cloud backup double protection. Users can quickly backup photos and videos by tapping the phone.

[0003] In related technologies, semantic image search based on the CLIP (Contrastive Language-Image Pre-Training) model is an advanced cross-modal retrieval technology. It encodes the features of images and texts and maps them into the same vector space, so as to realize the function of quickly retrieving relevant images through text description. However, most current researches and practical applications still focus on English semantic image search, and the support for Chinese is relatively limited. SUMMARY

[0004] The present disclosure provides an image retrieval method, an electronic device, a computer readable medium and a computer program product.

[0005] In a first aspect, the present disclosure provides an image retrieval method, which comprises:

[0006] inputting first language information to be retrieved into a pre-trained first language encoding model to extract a first feature vector of the first language information, wherein the first feature vector output by the first language encoding model for the first language is semantically aligned with a second feature vector output by a second language encoding model for a second language in the same semantic space;

[0007] inputting the first feature vector into an image encoding model to query a first target image set corresponding to the first feature vector from a target image database, wherein the image encoding model is trained according to images in the target database and texts in the second language corresponding to the images.

[0008] In a second aspect, the present disclosure also provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program executable by the processor, and the computer program is executed by the processor to implement the image retrieval method.

[0009] In a third aspect, the embodiments of the present disclosure further provide a computer readable medium, which stores a computer program, and the program is executed to implement the image retrieval method as described above.

[0010] In a fourth aspect, the embodiments of the present disclosure further provide a computer program product, which comprises a computer program, and the computer program is executed by a processor to implement the image retrieval method as described above.

[0011] The image retrieval method in the embodiments of the present disclosure comprises: inputting first language information to be retrieved into a pre-trained first language encoding model to extract a first feature vector of the first language information, wherein the first feature vector output by the first language encoding model for a first language is semantically aligned with a second feature vector output by a second language encoding model for a second language in the same semantic space; inputting the first feature vector into an image encoding model to query a first target image set corresponding to the first feature vector from a target image database; wherein the image encoding model is trained according to images in the target database and texts in the second language corresponding to the images; based on the first language encoding model, the embodiments of the present disclosure obtain the first feature vector of the first language information to be retrieved, and the first feature vector and the second feature vector of the second language are matched in the same semantic space, so that image retrieval can be directly performed in the target image database based on the first language information by using the image encoding model, thereby avoiding matching deviation caused by cross-language information loss, improving the accuracy and efficiency of first language semantic image search, and reducing the number of models and model parameters involved, thereby reducing the model training cost. BRIEF DESCRIPTION OF DRAWINGS

[0012] In the drawings of the embodiments of the present disclosure:

[0013] Figure 1 The image retrieval process schematic diagram provided for the embodiments of the present disclosure;

[0014] Figure 2 The schematic diagram of the MobileCLIP model provided for the embodiments of the present disclosure;

[0015] Figure 3 The first language encoding model training process schematic diagram provided for the embodiments of the present disclosure;

[0016] Figure 4 The first language encoding model and the second language encoding model alignment training schematic diagram provided for the embodiments of the present disclosure;

[0017] Figure 5 The image retrieval process schematic diagram provided for a specific example of the present disclosure;

[0018] Figure 6 The module composition schematic diagram of the electronic device provided for the embodiments of the present disclosure. DETAILED DESCRIPTION

[0019] To make the skilled in the art better understand the technical solutions of the present disclosure, the embodiments of the present disclosure are described in detail below.

[0020] The present disclosure will be described more fully hereinafter with reference to the accompanying drawings, in which embodiments of the present disclosure are shown. The present disclosure may, however, be embodied in many different forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the present disclosure to those skilled in the art.

[0021] The accompanying drawings, which are included to provide a further understanding of the present disclosure and are incorporated in and constitute a part of this specification, illustrate embodiments of the present disclosure and together with the detailed description serve to explain the present disclosure. The above and other features and advantages of the present disclosure will become more apparent from the following detailed description, taken in conjunction with the accompanying drawings.

[0022] The present disclosure can be described with reference to plan views and / or cross-sectional views by idealized schematic illustrations of the described present disclosure. Therefore, the example illustrations can vary depending on the manufacturing technology and / or tolerances.

[0023] The embodiments of the present disclosure and the features thereof can be combined with each other, if not in conflict.

[0024] The terms used in the present disclosure are merely used to describe particular embodiments, and are not intended to limit the present disclosure. As used in the present disclosure, the term "and / or" includes any and all combinations of one or more of the associated listed items. As used in the present disclosure, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. As used in the present disclosure, the terms "comprises," "comprising," "includes," "including," "has," "having," and the like are intended to be inclusive and allow for any other features, integers, steps, operations, elements, components, and / or groups thereof not expressly mentioned or inherent to the described present disclosure. Stated differently, the terms do not have an exclusionary, but rather an inclusive, meaning.

[0025] Unless otherwise defined, all terms (including technical and scientific terms) used in the present disclosure have the same meaning as commonly understood by one of ordinary skill in the art. It will be further understood that terms, such as those defined in commonly used dictionaries, should be interpreted as having a meaning that is consistent with their meaning in the context of the relevant art and the present disclosure, and will not be interpreted in an overly literal or overly formal sense unless expressly so defined herein.

[0026] The present disclosure is not limited to the embodiments shown in the drawings, but includes modifications of configurations formed based on a manufacturing process. Therefore, the regions exemplified in the drawings have a schematic property, and the shape of the regions shown in the drawings exemplifies a specific shape of a region of an element, but is not intended to be restrictive.

[0027] The FTTR system of the fused NAS not only supports local high-speed data storage, but also can be synchronized with a cloud disk to realize end-cloud backup double protection. At present, the fused NAS and the FTTR system support intelligent search of images, which can not only improve the speed and stability of the home network, but also improve the data management and retrieval efficiency through intelligent storage and management functions.

[0028] In the related art, the Chinese semantic image search is implemented by introducing a semantic conversion network to convert Chinese or Chinese-English mixed text into English text, and then performing image search on the English text. This Chinese semantic image search method needs to pass through three neural network models of a semantic conversion model, a text encoding model and an image encoding model, which additionally increases the semantic conversion model and the semantic conversion step, and the model parameter quantity and the model training data scale are huge, the model training process is time-consuming and costly, and at the same time, it also affects the image search accuracy and efficiency.

[0029] To solve the above problems, the embodiments of the present disclosure provide an image retrieval method, which is applied to an image retrieval model, the image retrieval model comprising a first language encoding model, a second language encoding model and an image encoding model, and the first language encoding model and the image encoding model are used to jointly complete the image retrieval task based on the first language.

[0030] Figure 1 The image retrieval process schematic diagram provided by the embodiments of the present disclosure is shown in Figure 1 The image retrieval method comprises the following steps:

[0031] Step S11, input the first language information to be retrieved into the pre-trained first language encoding model to extract the first feature vector of the first language information, wherein the first feature vector output by the first language encoding model for the first language is semantically aligned with the second feature vector output by the second language encoding model for the second language in the same semantic space.

[0032] In some embodiments, the first language is Chinese, and the second language is English, and correspondingly, the first language encoding model is a Chinese encoding model, the second language encoding model is an English encoding model, the first feature vector is a Chinese feature vector, and the second feature vector is an English feature vector.

[0033] Step S12, input the first feature vector into the image encoding model to query a first target image set corresponding to the first feature vector from a target image database; wherein the image encoding model is trained according to the images in the target database and the text in the second language corresponding to the images.

[0034] The training data for the image coding model is a dataset of image-text pairs consisting of images in the target library and text in the second language. The image coding model enables second-language semantic image search. Specifically, the second feature vector in the second language output by the second-language coding model is aligned with the image feature vectors of each image in the target library output by the image coding model. Furthermore, the first feature vector extracted by the first-language coding model and the second feature vector extracted by the second-language model are semantically aligned within the same semantic space. Therefore, the first feature vector is automatically aligned with the image feature vectors of each image in the target library, enabling direct image search of the target library based on the first language using the aforementioned image coding model.

[0035] In the disclosed embodiment, the second language encoding model and the image encoding model use the lightweight English semantic image search model MobileCLIP as the model base due to its significant advantages in parameter quantity and performance. Figure 2 A schematic diagram of the MobileCLIP model provided in the embodiment of the present disclosure is shown in FIG. Figure 2 As shown, taking the MobileCLIP-S0 model as an example, its image encoder model has only 11.4M parameters, and its English text encoder model, used as a second language encoding model, has only 42.4M parameters. The overall parameter count of the MobileCLIP model is far smaller than the hundreds of megabytes or even billions of parameters of larger models like CLIP and ChineseCLIP. Choosing a lightweight model like MobileCLIP as the model foundation not only significantly reduces the demand for computing resources, but also makes it more suitable for embedded devices or low-computing scenarios. Furthermore, the MobileCLIP model maintains excellent performance in English semantic retrieval and possesses excellent text-image alignment capabilities. It is the optimal choice for achieving both lightweight and effective retrieval, laying a solid foundation for efficient Chinese semantic image search.

[0036] The image retrieval method in the embodiments of the present disclosure comprises: inputting first language information to be retrieved into a pre-trained first language encoding model to extract a first feature vector of the first language information, wherein the first feature vector output by the first language encoding model for the first language is semantically aligned with a second feature vector output by a second language encoding model for the second language in the same semantic space; inputting the first feature vector into an image encoding model to query a first target image set corresponding to the first feature vector from a target image database; wherein the image encoding model is trained according to images in the target database and texts in the second language corresponding to the images; the embodiments of the present disclosure obtain the first feature vector of the first language information to be retrieved based on the first language encoding model, the first feature vector and the second feature vector of the second language are matched in the same semantic space, and image retrieval can be directly performed in the target image database based on the first language information based on the image encoding model, thereby avoiding matching deviation caused by cross-language information loss, improving the accuracy and efficiency of first language semantic image search, and reducing the number of models and model parameters involved, thereby reducing the model training cost.

[0037] In some embodiments, the step of inputting the first feature vector into the image encoding model to query a first target image set corresponding to the first feature vector from the target image database (i.e., step S12) comprises the following steps:

[0038] Step S121, calculating a first similarity between the first feature vector and image feature vectors of each image in the target image database by the image encoding model.

[0039] It should be noted that each image in the target image database can be input into the image encoding model of the MobileCLIP model which has been frozen in advance to obtain image feature vectors corresponding to each image , and normalize each image feature vector to obtain each third unit vector.

[0040] In some embodiments, each image feature vector can be normalized according to the following formula (1) to obtain each third unit vector :

[0041] (1)

[0042] In some embodiments, each first feature vector can be normalized according to the following formula (2) to obtain each first unit vector :

[0043] (2)

[0044] In some embodiments, the first similarity is a cosine similarity, and the higher the cosine similarity value is, the closer the semantic of the first language information to be retrieved is to the image, and the higher the matching degree is. In some embodiments, the first similarity between the first feature vector and the image feature vector of each image in the target gallery can be calculated according to the following formula (3) :

[0045] (3)

[0046] wherein, is the first unit vector, is the third unit vector.

[0047] In step S122, a first target image set corresponding to the first feature vector is determined according to the first similarities, and each image in the first target set is sorted in descending order of the first similarity.

[0048] In this step, the images in the target gallery that satisfy the preset condition of the first similarity are determined, and the first target image set is generated according to these images. For example, the target images that satisfy the preset condition of the first similarity can be images whose first similarity is greater than a preset first similarity threshold, or the top target number of images in the target gallery in terms of the first similarity. It should be noted that each image in the first target image set is sorted in descending order of the first similarity.

[0049] Since the first feature vector obtained by encoding the first language information and the image feature vector obtained by encoding the image are difficult to be 100% similar, the similarity between the image feature vectors obtained by encoding the images of the same category is higher. For example, the similarity between the first feature vector obtained by encoding the first language information "dog" and the image feature vector obtained by encoding the image of a dog is higher than the similarity between the image feature vectors obtained by encoding two images of a dog. Therefore, in the embodiments of the present disclosure, after obtaining the first target image set, in order to further optimize the sorting of each image in the first target image set, each image in the first target image set is divided into two parts according to the first similarity, and the images with lower first similarity are re-sorted in an iterative manner to ensure that the sorting of the images in the final retrieval result is more in line with the user's expectation. Therefore, in some embodiments, after the step of determining the first target image set corresponding to the first feature vector according to the first similarities (i.e., step S122), the image retrieval method can further include the following steps:

[0050] Step S123, dividing the first target set into a first image set and a second image set; wherein the first similarity corresponding to each second image in the second image set is less than the first similarity corresponding to each first image in the first image set.

[0051] In this step, according to a preset second similarity threshold, each image in the first target set greater than or equal to the preset second similarity threshold is determined as a first image, and all first images form the first image set; each image in the first target set less than the preset second similarity threshold is determined as a second image, and all second images form the second image set.

[0052] Step S124, for each second image in the second image set, calculating a second similarity corresponding to the second image according to the similarity between the second image and each first image.

[0053] In this step, for each second image in the second image set, the third similarity between the second image and each first image in the first image set is calculated respectively, and the second similarity corresponding to the second image is calculated according to each third similarity.

[0054] In some embodiments, the second similarity corresponding to the second image can be the average of each third similarity, or can be the median of each third similarity. In the embodiments of the present disclosure, the second similarity corresponding to the second image is obtained by calculating the average of each third similarity, and the specific process is as follows. The second similarity corresponding to the i th second image can be calculated according to the following formula (4):

[0055] (4)

[0056] Wherein, is the average of each third similarity corresponding to the i th second image, is the image feature vector of the i th second image, is the image feature vector of the j th first image, is the third similarity between the i th second image and the j th first image, and S is the total number of first images in the first image set.

[0057] In this step, the second similarity corresponding to each second image in the second image set is calculated according to the above formula (4).

[0058] Step S125, according to the second similarity corresponding to each second image, merging the second image set and the first image set to obtain a second target image set.

[0059] In this step, the order of each first image in the first target image set is kept unchanged, the order of each second image in the second image set is adjusted, and the second image set is merged with the first image set to obtain a second target set.

[0060] By further dividing the first target set into two image sets and reordering each second image in the second image set with a smaller first similarity, the search result can be optimized, the retrieval precision and user experience of the first language semantic search can be improved, and finally the accurate matching between the first language and the image is realized.

[0061] In some embodiments, the merging of the second image set with the first image set to obtain a second target image set according to the second similarity corresponding to each second image (i.e., step S125) comprises the following steps: determining a target second image corresponding to the highest second similarity according to the second similarity corresponding to each second image, adding the target second image to the end of all first images in the first image set to obtain an adjusted first image set and an adjusted second image set, and turning to the step of calculating the second similarity corresponding to each second image in the second image set according to the similarity between the second image and each first image until all second images are added to the first image set; and determining the first image set with all second images added as the second target image set.

[0062] From the second similarity corresponding to each second image in the second image set, the maximum value of the second similarity is determined, and the second image corresponding to the maximum value is determined as a target second image. The target second image is the most similar image to the first image set found after reordering the second image set. Therefore, the target second image is added to the end of all first images in the first image set to form a new first image set, in which the order of each first image is unchanged. Similarly, step S124 is iteratively performed on each remaining second image in the second image set, the second similarity corresponding to each remaining second image is calculated based on the current first image set, and the second image with the maximum second similarity is added to the current first image set until all second images in the second image set are added to the first image set, thereby obtaining a second target image set.

[0063] In order to reduce the number of sorting times, a preset number of second images with larger second similarity can also be selected as target second images each time. The smaller the preset number of sorting, the higher the sorting precision of the final second target image set.

[0064] After the above iteration is completed, the images in the first target image set that are more similar to the first language information to be retrieved are sorted, and the arrangement order of each second image in the second image set is also optimized, so that the arrangement order of each image in the second target image set obtained is more in line with the retrieval requirements of the user, facilitating the user to quickly lock the required image and improving the user experience.

[0065] In some embodiments, the step of dividing the first target set into the first image set and the second image set (i.e., step S123) can include the following steps: selecting a to-be-deleted image set with a first similarity less than a preset first threshold from the first target image set, and deleting the to-be-deleted image set from the first target image set to obtain a first target set after deletion; and dividing the first target set after deletion into the first image set and the second image set.

[0066] That is, the first target image set can be divided into the following three image sets:

[0067] (1) The first image set: that is, the high-similarity image set, the first similarity of each image in the first image set is greater than or equal to a preset second similarity threshold, and the first image in the set is directly retained. For example, the first image set is the image set with a first similarity greater than or equal to 0.35.

[0068] (2) The second image set: that is, the medium-similarity image set, the first similarity of each image in the first target set is less than the preset second similarity threshold and greater than the preset first threshold, and the order of the second image in the set needs to be adjusted. For example, the second image set is the image set with a first similarity less than 0.35 and greater than 0.17.

[0069] (3) The third image set: that is, the low-similarity image set, the first similarity of each image in the third image set is less than the preset first threshold, and the third image set is directly deleted and no longer displayed to the user. For example, the third image set is the image set with a first similarity less than 0.17.

[0070] Figure 3 The first language encoding model training process diagram provided by the embodiments of the present disclosure is shown in FIG. 1. Figure 3 Before the step of inputting the first language information to be retrieved into the pre-trained first language encoding model to extract the first feature vector of the first language information (i.e., step S11), the image retrieval method can further include the following steps:

[0071] Step S21 : Inputting the first language information of each preset first language and second language text pair into a preset initial first language encoding model to extract a first feature vector corresponding to each first language information; and inputting the second language information of each preset first language and second language text pair into a second language encoding model to extract a second feature vector corresponding to each second language information.

[0072] The preset first language and second language text pairs are used as training data for the first language encoding model. The preset first language and second language text pairs include first language information and its corresponding second language information. In this step, the first feature vector of the first language information in the first language and second language text pair is extracted using the preset initial first language encoding model, and the second feature vector of the second language information in the same first language and second language text pair is extracted using the second language encoding model. In other words, for a first language and second language text pair, , comparing the first language with the second language text First language information in Input preset initial first language encoding model , get the first feature vector corresponding to the first language information , ; Compare the first language and second language texts Second language information in Input second language encoding model , get the second feature vector corresponding to the second language information , .

[0073] In some embodiments, the second language encoding model It can be the English encoder (Text Encoder) of the MobileCLIP model. In the embodiment of the present disclosure, the second language encoding model The parameters remain unchanged.

[0074] Figure 4 This is a schematic diagram of the alignment training of the first language encoding model and the second language encoding model provided in the embodiment of the present disclosure, as shown in FIG. Figure 4 As shown, multiple first language information including "Australian Shepherd Puppy" is input into the preset initial first language encoding model (i.e. Chinese encoder), get the corresponding first feature vectors I1, I2, ..., I N , wherein a first language information corresponds to a first feature vector, for example, "Australian Shepherd Puppy" corresponds to the first feature vector I1; multiple second language information including "pepperthe aussie pup" are input into the second language encoding model (i.e., the Text Encoder), to obtain corresponding second feature vectors T1, T2, …, T N wherein one second language information corresponds to one second feature vector, for example, “pepper the aussie pup” corresponds to the second feature vector T1.

[0075] Step S22, constructing a loss function according to each first feature vector and each second feature vector.

[0076] In some embodiments, the loss function is selected to be a contrastive learning loss function, and its formula is as follows:

[0077] (5)

[0078] wherein N is the number of samples in a batch, represents the cosine similarity, is an adjustable temperature parameter for controlling the smoothness of the distribution.

[0079] Step S23, freezing the second language encoding model, iteratively training the preset initial first language encoding model according to the loss function, to obtain the pre-trained first language encoding model.

[0080] The preset initial first language encoding model is an initialized but untrained first language encoding model, and the parameters of the second language encoding model are fixed, i.e., the second language encoding model is frozen, the preset initial first language encoding model is trained to obtain the final model parameters, and the model parameters are configured into the preset initial first language encoding model to obtain the first language encoding model. That is, during the training of the first language encoding model, only the parameters of the preset initial first language encoding model are adjusted so that the loss function converges, while the parameters of the second language encoding model remain in a frozen state. After training for a target training number of epochs (rounds), the first language encoding model can be consistent with the second language encoding model in the same semantic space. The target training number can be set according to experience values.

[0081] As shown in Figure 4 , the output first feature vectors (I1, I2, …, I N ) of the preset initial first language encoding model are aligned with the output second feature vectors of the second language encoding model . Through this alignment process, the first feature vectors can be accurately mapped to the same vector space as the second language encoding model.

[0082] Since the second eigenvector output by the English encoding model in the MobileCLIP model has been aligned with the image eigenvector output by the image encoding model in the MobileCLIP model (e.g. Figure 2 As shown in Figure 2, the first feature vector output by the trained first language encoding model is automatically aligned with the image feature vector output by the image encoding model. This allows for direct image retrieval based on the first language information, effectively implementing the first language semantic image search function.

[0083] The disclosed embodiments employ a method of aligning a second language encoding model to train a first language encoding model, aligning a first feature vector output by the first language encoding model with a second feature vector output by the second language encoding model. This ensures that the first language information and the image encoding model match within the same semantic space. This can reduce the size of training data for the first language encoding model, thereby reducing the storage space for training data, shortening model training time, and lowering model training costs.

[0084] In order to eliminate the scale effect of the eigenvector, the first eigenvector and the second eigenvector can also be normalized to unit vectors. Therefore, in some embodiments, the loss function is constructed based on each first eigenvector and each second eigenvector (i.e., step S22), including the following steps:

[0085] Step S221 , normalizing each first eigenvector to obtain each first unit vector; and normalizing each second eigenvector to obtain each second unit vector.

[0086] The first unit vector corresponding to each first eigenvector can be calculated according to the above formula (2): .

[0087] In some embodiments, each second eigenvector can be calculated according to the following formula (6): The corresponding second unit vector :

[0088] (6)

[0089] Step S222: construct a loss function based on each first unit vector and each second unit vector.

[0090] The loss function constructed based on each first unit vector and each second unit vector is shown in the following formula (7):

[0091] (7)

[0092] In the first language encoding model, the word vector table often accounts for a large part of the model parameters, about half of the model parameters. In the related art, the parameter amount of the ChineseCLIP model is often in the order of tens of billions, making it difficult to deploy on embedded end NAS devices. Moreover, the ChineseCLIP model has high complexity, and NAS devices are usually limited by computing power and memory resources, while the inference requirements of ChineseCLIP exceed the load of most embedded systems.

[0093] In some related technologies, two-step optimization of Chinese encoding model training and Chinese text and image alignment training is adopted, and a Chinese text and image retrieval model is trained based on the CLIP model. However, this method still has the problem of excessive training data size and model parameter amount.

[0094] To solve this problem, the embodiments of the present disclosure use a word table pruning technology to construct a lightweight word table, which can significantly reduce the size of the word vector table, thereby reducing the overall parameter amount of the first language encoding model and realizing the construction of a lightweight first language encoding model.

[0095] In some embodiments, the first language encoding model can select the MobileBert model or the MobileCLIP-S0 model as the model base. In the embodiments of the present disclosure, the MobileBert model is taken as an example for illustration. In the process of constructing the Chinese encoding Bert model, the core of constructing the lightweight word table is to determine a high-quality Chinese data set covering commonly used Chinese characters. The embodiments of the present disclosure select the Chinese corpus in large-scale data sets such as “ShuSheng-WanJuan data set” and “MNBVC super large-scale Chinese data set” to construct the lightweight word table. Therefore, before the first language information in each preset first language and second language text pair is input into the preset initial first language encoding model and the first feature vector corresponding to each first language information is extracted (i.e., step S21), the image retrieval method can further include the following steps:

[0096] Step S31, statistics of the frequency of each first language character in the preset first language data set in the preset first language data set.

[0097] In this step, Chinese character extraction is performed based on frequency statistics, and the frequency of each Chinese character appearing in the entire corpus is counted. In some embodiments, a Chinese character list sorted by frequency can be further constructed according to the frequency of appearance.

[0098] In some embodiments, the step of counting the frequency of each first language character in the preset first language data set can include the following steps: data cleaning is performed on the preset first language data set to obtain a cleaned first language data set; and the frequency of each first language character in the cleaned first language data set in the cleaned first language data set is counted. That is, before counting the frequency of Chinese characters in the preset Chinese data set, data cleaning is performed on the corpus in the preset Chinese data set. For example, full-quantity word segmentation and cleaning can be performed on the Chinese corpus, and punctuation marks, non-Chinese characters and rare noise data are deleted.

[0099] In step S32, according to the frequency, a target number of target first language characters with large frequency are selected.

[0100] In some embodiments, the target number of target first language characters with large frequency can be selected by setting a quantity threshold, or the target number of target first language characters with large frequency can be selected by setting a proportion threshold, to realize the pruning of the vocabulary table. Taking the setting of the proportion threshold as an example, in order to ensure that the first language coding model can cover most of the commonly used Chinese characters in actual application, the proportion threshold can be set to 95%, that is, the Chinese characters with a frequency of more than 95% in the corpus are selected as target Chinese characters (i.e., target first language characters), and a word vector table is generated according to the target Chinese characters. The Chinese characters that are filtered out are Chinese characters that are rarely used by users.

[0101] Through statistical analysis, it is found that about 3500-5000 Chinese characters can cover most of the Chinese text. Further experiments show that the coverage rate of the target Chinese characters selected by the disclosed embodiments is similar to the coverage rate of 3500 Chinese characters in the Modern Chinese Commonly Used Character Table published by the National Language and Character Working Committee, and the distribution characteristics of high-frequency Chinese characters are consistent. That is, the size of the word vector table obtained through Chinese character screening is similar to the size of the Modern Chinese Commonly Used Character Table, and through the pruning of the vocabulary table, the size of the vocabulary table can be significantly reduced while the expression ability of the first language coding model for core language characteristics is retained, and the first language coding model can be lightened and the calculation complexity of the first language coding model can be reduced. In addition, this vocabulary pruning scheme can flexibly adapt to the corpus requirements of different fields, and provides an efficient path for the construction of field-specific Chinese Bert models.

[0102] In step S33, the target layer of the preset initial first language coding model is initialized according to the target first language characters.

[0103] In this step, the target layer of the preset initial first language coding model is initialized by using the word vector table corresponding to the target Chinese characters, and the establishment of the preset initial first language coding model is completed.

[0104] In some embodiments, before the step of inputting each first language information in the preset first language and second language text pair into the preset initial first language encoding model and extracting a first feature vector corresponding to each first language information (i.e., step S21), the image retrieval method can further include the following steps:

[0105] Step S41, extracting low-frequency words in the second language text with an occurrence frequency less than a preset second threshold.

[0106] In some embodiments, the training data used to construct the preset first language and second language text pair is selected from the second language text in the DataComp data set. The DataComp data set is large in size and covers a wide variety of English text-image pairs. The DataComp data set includes common English words, but the distribution of the words is uneven, so further processing of the English low-frequency words in the DataComp data set is needed.

[0107] In this step, a small part of the DataComp data set is selected to construct the preset first language and second language text pair. When constructing the first language and second language text pair, English texts are not directly selected at random, because this can result in insufficient proportion of low-frequency words, which in turn affects the learning effect of the first language encoding model on these words. Therefore, in this step, statistical analysis is performed on the English texts in the DataComp data set, the occurrence frequency of each English word is calculated, and low-frequency words with an occurrence frequency less than a preset second threshold are extracted to obtain a low-frequency word set.

[0108] It should be noted that the low-frequency words here do not refer to words that are used less by users. The words in the DataComp data set are basically words that are used more by users. The low-frequency words here refer to words that appear less in the DataComp data set but are not less used by users, so the number of these samples needs to be increased.

[0109] Step S42, constructing a second language sentence corresponding to at least one low-frequency word in the low-frequency words.

[0110] In this step, one or more English sentences can be reconstructed according to grammatical rules, preset context scenarios, and the low-frequency words, so that they can be integrated into various actual use scenarios.

[0111] Step S43, translating the second language sentences corresponding to each low-frequency word and the second language sentences in the second language text into first language sentences corresponding to each second language sentence.

[0112] In this step, the English sentence constructed in step S42 is translated into the corresponding Chinese sentence, and the English sentence in the DataComp dataset is translated into the corresponding Chinese sentence. A targeted prompt can be designed to translate the English sentence into the corresponding Chinese sentence through a large language model to ensure translation quality and context consistency. For example, the prompt can be: "Translate the following description in English into fluent Chinese while maintaining semantic accuracy and high matching with image content."

[0113] Step S44, generating each preset first language and second language text pair according to each second language sentence and the first language sentence corresponding to each second language sentence.

[0114] To realize the Chinese semantic search function, the embodiments of the present disclosure process the English text-image pairs used in the training of the MobileCLIP model using a large language model, translate the English text into the corresponding Chinese text, and thus generate a Chinese text-image pair dataset suitable for Chinese semantic search training, i.e., generate a first language and second language text pair of the order of ten million. The above-mentioned scheme for constructing a preset first language and second language text pair not only significantly improves the coverage of low-frequency words, but also ensures that the constructed sentences are semantically clear and close to actual scenarios, thereby providing more comprehensive support for the alignment of English and Chinese semantics during the training process of the first language encoding model.

[0115] To clearly illustrate the scheme of the embodiments of the present disclosure, the following will combine Figure 5 to illustrate the image retrieval process in detail through a specific example. Figure 5 The image retrieval process provided by a specific example of the present disclosure is shown in Figure 5 The first language information to be retrieved is "baby", and there are 10 images in the target image library. The image feature vectors of these 10 images are extracted and stored in advance using the image encoding model. After extracting the first feature vector of "baby" using the first language encoding model, the first similarity between the first feature vector corresponding to "baby" and the image feature vectors of the 10 images is calculated, and the 10 images are sorted in order of the first similarity from high to low to obtain a first target image set. The first similarity of each first image in the first target image set is in turn: 0.35, 0.28, 0.27, 0.24, 0.22, 0.21, 0.19, 0.18, 0.17, 0.14.

[0116] Step 1, dividing the first target image set into 3 image sets according to a preset threshold.

[0117] The preset threshold values are 0.25 and 0.17, and according to the two preset threshold values, the first target image set is divided into the following three image sets: the image with a first similarity greater than 0.25 is a first image, and belongs to a first image set; the image with a first similarity between 0.25 and 0.17 is a second image, and belongs to a second image set; and the image with a first similarity less than or equal to 0.17 is a third image, and belongs to a third image set. The order of each first image in the first image set is unchanged, each second image in the second image set is reordered, and the third image set is deleted.

[0118] Through steps 2-n, after the second image set is reordered, the second image set is combined with the first image set to obtain a second target image set.

[0119] Step 2: For the first second image (i.e., the image with a first similarity of 0.24) in the second image set, the third similarity of the second image with the three first images in the first image set is calculated as 0.58, 0.55, and 0.59, respectively, and the second similarity corresponding to the second image is calculated as 0.573 by averaging the above three third similarities. In this way, the second similarity corresponding to all six second images in the second image set can be calculated as 0.573, 0.302, 0.325, 0.521, 0.488, and 0.511, respectively. The maximum value of the above six second similarities is 0.573, and the second image corresponding to the maximum value of the second similarity (i.e., the second image with a first similarity of 0.24) is determined as the current target second image. The current target second image is added to the end of all first images in the first image set to obtain an updated first image set and an updated second image set. At this time, the updated first image set includes four first images, and the updated second image set includes five second images.

[0120] Step 3, for the first second image (i.e., the image with a first similarity of 0.22) in the current second image set, the third similarities of the second image and the four first images in the first image set are calculated as 0.58, 0.55, 0.59, and 0.28, respectively. According to the above four third similarities, the second similarity corresponding to the second image is calculated as 0.289 by averaging. In this way, the second similarities corresponding to all five second images in the current second image set can be calculated as 0.289, 0.302, 0.527, 0.488, and 0.510, respectively. The maximum value of the above five second similarities is 0.527, and the second image corresponding to the maximum value of the second similarity (i.e., the second image with a first similarity of 0.19) is determined as the current target second image. The current target second image is added to the end of all first images in the current first image set to obtain an updated first image set and an updated second image set. At this time, the updated first image set includes five first images, and the updated second image set includes four second images.

[0121] In this way, after nine steps, i.e., n = 9, the updated first image set includes nine first images, and the second image set is empty. The updated first image set is the second target image set.

[0122] The image search scheme of the embodiments of the present disclosure can be applied to a Home NAS product, especially in the scenario of efficiently managing and retrieving massive multimedia data in daily life for home users, greatly improving the convenience of users in managing and retrieving massive multimedia data. When a user needs to quickly find specific content in a large photo or video library, manual searching or relying on cumbersome file name and tag management is not needed. Instead, the user only needs to input a Chinese search description in natural language to achieve accurate positioning. This not only improves the image retrieval efficiency, but also greatly improves the user experience, especially for non-technical users and users who lack patience in file organization. In addition, the Chinese semantic image search function is particularly important for home users with a large number of Chinese media resources. By using natural language processing and computer vision technology, the system can understand various Chinese expressions and realize accurate matching between text and images by combining a multi-modal model. Whether it is to find travel records, family activity moments, or organize important memories, Chinese semantic image search can significantly simplify the operation and become a key step for home NAS to transform from a simple storage device to an intelligent assistant.

[0123] The embodiments of the present disclosure also propose a Chinese semantic image search method based on MobileCLIP, which realizes Chinese text-image matching by constructing Chinese-English text pairs, laying a foundation for Chinese semantic image search.

[0124] The image retrieval result optimization scheme optimizes the ranking of images with low similarity by using a rearrangement algorithm, improves the retrieval accuracy and user experience of Chinese semantic image search, and finally realizes efficient matching of Chinese text and images.

[0125] The image retrieval result optimization scheme optimizes the ranking of images with low similarity by using a rearrangement algorithm, improves the retrieval accuracy and user experience of Chinese semantic image search, and finally realizes efficient matching of Chinese text and images.

[0126] The image retrieval result optimization scheme optimizes the ranking of images with low similarity by using a rearrangement algorithm, improves the retrieval accuracy and user experience of Chinese semantic image search, and finally realizes efficient matching of Chinese text and images.

[0127] The image retrieval result optimization scheme optimizes the ranking of images with low similarity by using a rearrangement algorithm, improves the retrieval accuracy and user experience of Chinese semantic image search, and finally realizes efficient matching of Chinese text and images.

[0128] The image retrieval result optimization scheme optimizes the ranking of images with low similarity by using a rearrangement algorithm, improves the retrieval accuracy and user experience of Chinese semantic image search, and finally realizes efficient matching of Chinese text and images. Figure 6 The image retrieval result optimization scheme optimizes the ranking of images with low similarity by using a rearrangement algorithm, improves the retrieval accuracy and user experience of Chinese semantic image search, and finally realizes efficient matching of Chinese text and images.

[0129] The image retrieval result optimization scheme optimizes the ranking of images with low similarity by using a rearrangement algorithm, improves the retrieval accuracy and user experience of Chinese semantic image search, and finally realizes efficient matching of Chinese text and images.

[0130] The image retrieval result optimization scheme optimizes the ranking of images with low similarity by using a rearrangement algorithm, improves the retrieval accuracy and user experience of Chinese semantic image search, and finally realizes efficient matching of Chinese text and images.

[0131] The image retrieval result optimization scheme optimizes the ranking of images with low similarity by using a rearrangement algorithm, improves the retrieval accuracy and user experience of Chinese semantic image search, and finally realizes efficient matching of Chinese text and images.

[0132] Those of ordinary skill in the art will understand that the functional modules / units in all or some of the steps, systems, apparatuses disclosed above can be implemented as software, firmware, hardware, and appropriate combinations thereof.

[0133] In hardware implementation, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, one physical component can have multiple functions, or one function or step can be performed by several physical components in cooperation.

[0134] Some or all of the physical components can be implemented as software executed by a processor, such as a central processing unit (CPU), a digital signal processor, or a microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on computer-readable media, which can include computer storage media (or non-transitory media) and communication media (or transitory media). As is well known to those of ordinary skill in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules or other data. Computer storage media include, but are not limited to, random access memory (RAM, more specifically SDRAM, DDR, etc.), read only memory (ROM), electrically erasable programmable read only memory (EEPROM), FLASH memory or other memory technologies; compact disc read-only memory (CD-ROM), digital versatile discs (DVD) or other optical disk storage; magnetic cassettes, magnetic tapes, magnetic disk storage or other magnetic storage devices; any other medium that can be used to store the desired information and that can be accessed by a computer. Further, as is well known to those of ordinary skill in the art, communication media typically embodies computer readable instructions, data structures, program modules or other data in a modulated data signal such as a carrier wave or other transport mechanism and includes any information delivery media. As a non-limiting example, the foregoing example of a non-transitory medium is merely meant to illustrate that such can be embodied in a computer program, i.e., software, that is downloaded to or accessed from one or more computer storage media.

[0135] The present disclosure has disclosed example embodiments, and although the specific terms are employed, they are used in a generic sense only and should not be construed to be limited to the specific embodiments described herein. In some instances, it will be readily apparent to those skilled in the art that a feature, characteristic or element described in connection with a particular embodiment can be used in conjunction with other embodiments unless expressly stated otherwise. Accordingly, it will be understood that various changes in form and details can be made without departing from the scope of the disclosure as set forth in the appended claims.

Claims

1. An image retrieval method, the method comprising: inputting first language information to be retrieved into a pre-trained first language encoding model to extract a first feature vector of the first language information, wherein the first feature vector output by the first language encoding model is semantically aligned with a second feature vector output by a second language encoding model in a same semantic space; calculating, by an image encoding model, a first similarity between the first feature vector and an image feature vector of each image in a target image library; determining, according to each of the first similarities, a first target image set corresponding to the first feature vector, each image in the first target set being sorted in descending order of the first similarity; dividing the first target set into a first image set and a second image set, wherein each second image in the second image set corresponds to a first similarity smaller than a first similarity corresponding to each first image in the first image set; for each second image in the second image set, calculating a third similarity between the second image and each first image in the first image set, and calculating a second similarity corresponding to the second image according to each of the third similarities; determining, according to each second similarity corresponding to each second image, a target second image corresponding to a highest second similarity, adding the target second image to the end of all first images in the first image set to obtain an adjusted first image set and an adjusted second image set, and turning to the step of calculating, for each second image in the second image set, a second similarity corresponding to the second image according to a similarity between the second image and each first image until all second images are added to the first image set; determining the first image set to which all second images are added as the second target image set; wherein the image encoding model is trained according to images in the target database and texts in a second language corresponding to the images.

2. The method of claim 1, wherein, The dividing of the first target set into a first image set and a second image set comprises: selecting, from the first target image set, a to-be-deleted image set with a first similarity smaller than a preset first threshold, and deleting the to-be-deleted image set from the first target image set to obtain a deleted first target set; dividing the deleted first target set into a first image set and a second image set.

3. The method of claim 1 or 2, wherein, Before the inputting of the first language information to be retrieved into the pre-trained first language encoding model to extract the first feature vector of the first language information, the method further comprises: inputting first language information in each preset first language and second language text pair into a preset initial first language encoding model to extract a first feature vector corresponding to each first language information, and inputting second language information in each preset first language and second language text pair into the second language encoding model to extract a second feature vector corresponding to each second language information; constructing a loss function according to each first feature vector and each second feature vector; Freeze the second language coding model, iteratively train the preset initial first language coding model according to the loss function, and obtain the pre-trained first language coding model.

4. The method of claim 3, wherein, Before the first language information in each preset first language and second language text pair is input into the preset initial first language coding model and the first feature vector corresponding to each first language information is extracted, the method further includes: Counting the frequency of each first language character in the preset first language data set in the preset first language data set; According to the frequency, selecting a target number of target first language characters with a larger frequency; According to the target first language character, initializing the target layer of the preset initial first language coding model.

5. The method of claim 4, wherein, The frequency of each first language character in the preset first language data set in the preset first language data set is counted, including: Data cleaning is performed on the preset first language data set to obtain a cleaned first language data set; Counting the frequency of each first language character in the cleaned first language data set in the cleaned first language data set.

6. The method of claim 3, wherein, Before the first language information in each preset first language and second language text pair is input into the preset initial first language coding model and the first feature vector corresponding to each first language information is extracted, the method further includes: Extracting low-frequency words in the text of the second language with a frequency less than a preset second threshold value; For at least one low-frequency word in the low-frequency word, constructing a second language sentence corresponding to the at least one low-frequency word; Translate each second language sentence and the second language sentence in the text of the second language into a first language sentence corresponding to each second language sentence; According to each second language sentence and each first language sentence corresponding to each second language sentence, generate each preset first language and second language text pair.

7. The method of claim 3, wherein, The loss function is constructed according to each first feature vector and each second feature vector, including: Normalizing each first feature vector to obtain each first unit vector, and normalizing each second feature vector to obtain each second unit vector; According to each first unit vector and each second unit vector, the loss function is constructed.

8. The method of claim 1, wherein, The first language is Chinese, and the second language is English. 9.An electronic device comprising a memory and a processor; The memory stores a computer program executable by the processor, and the computer program is executed by the processor to implement the image retrieval method of any one of claims 1-8.

10. A computer readable medium having stored thereon a computer program, wherein, The program is executed to implement the image retrieval method of any one of claims 1-8. 11.A computer program product comprising a computer program, the computer program being executed by a processor to implement the image retrieval method of any one of claims 1-8. 11.A computer program product comprising a computer program, the computer program being executed by a processor to implement the image retrieval method of any one of claims 1-8.