Image Retrieval Method, Device, Equipment and Storage Medium Based on Chinese Data

By translating English text data into Chinese and optimizing knowledge distillation, the problem of low image recognition efficiency in Chinese data is solved, and efficient Chinese CLIP model training and image retrieval are achieved.

CN115238115BActive Publication Date: 2025-05-27PING AN TECH (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210897638.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-28
Publication Date
2025-05-27
Estimated Expiration
2042-07-28

AI Technical Summary

Technical Problem

Image recognition based on Chinese data is low efficiency. The existing CLIP model is mainly pre-trained based on English data, which leads to the need to start from scratch, which is very expensive.

Method used

By obtaining English text data, using machine translation algorithm to translate it into Chinese text data, and inputting the English text data into the text encoder model for encoding. Through knowledge distillation and loss function optimization, a Chinese text encoder is obtained, and finally image retrieval is performed based on the Chinese text image pre-trained model.

Benefits of technology

The training of the Chinese CLIP model is realized, which improves the efficiency of image recognition based on Chinese data, and avoids the huge cost of training from scratch.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115238115B_ABST
    Figure CN115238115B_ABST
Patent Text Reader

Abstract

The present invention relates to artificial intelligence technology, and discloses an image retrieval method based on Chinese data, including: translating English text data into Chinese text data; training an English text data vector set to obtain an English text training data set; distilling the English text training data set and the Chinese text data to obtain a first probability value and a second probability value; calculating the loss value of the first probability value and the second probability value; optimizing a text encoder model to obtain a Chinese text encoder; performing model inference on the Chinese text encoder; inputting the Chinese data to be analyzed into a Chinese text image pre-training model to obtain an image corresponding to the Chinese data to be analyzed. In addition, the present invention also relates to blockchain technology, and the English text data can be stored in the nodes of the blockchain. The present invention also provides an image retrieval device, an electronic device and a storage medium based on Chinese data. The present invention can improve the efficiency of image recognition based on Chinese data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence, and in particular, to an image retrieval method, device, electronic device and computer-readable storage medium based on Chinese data. Background Art

[0002] With the development of image recognition technology, more and more image recognition technologies have emerged. The CLIP model with image recognition has been continuously extended to detection, image-text retrieval, and conditional generation of images, achieving amazing effects in many fields. However, in order to improve the efficiency of image recognition, it is necessary to collect English and Chinese data for image recognition.

[0003] The existing CLIP model is pre-trained based on English data and is a pre-trained neural network model for matching images and texts. By encoding images and texts, the similarity between the image encoding and the text encoding is calculated for image-text matching. In practical applications, if you want to use the CLIP model with Chinese, you need to train from scratch, which is extremely costly, resulting in low efficiency when performing image recognition based on Chinese data. Summary of the Invention

[0004] The present invention provides an image retrieval method, device and computer-readable storage medium based on Chinese data, and its main purpose is to solve the problem of low efficiency when performing image recognition based on Chinese data.

[0005] To achieve the above object, an image retrieval method based on Chinese data provided by the present invention includes:

[0006] Obtain the English text data of the training data, and use a preset machine translation algorithm to translate the English text data into Chinese text data;

[0007] Input the English text data into a preset text encoder model for encoding to obtain an English text data vector set, and use a preset text-image pre-training model to train the English text data vector set to obtain an English text training data set;

[0008] Input the English text training data set into a preset teacher model for knowledge distillation to obtain a first probability value, and input the Chinese text data into a preset student model for knowledge distillation to obtain a second probability value;

[0009] Use a preset loss function to calculate the mean absolute error loss value of the first probability value and the second probability value, and optimize the text encoder model according to the mean absolute error loss value to obtain a Chinese text encoder;

[0010] Perform model inference according to the Chinese text encoder and a preset image encoder to obtain a pre-trained model for Chinese text images;

[0011] Obtain the Chinese data to be analyzed, and input the Chinese data to be analyzed into the pre-trained model for Chinese text images to obtain an image corresponding to the Chinese data to be analyzed.

[0012] Optionally, the step of translating the English text data into Chinese text data by using a preset machine translation algorithm includes:

[0013] Perform syntactic structure segmentation on the English text data to obtain segmented sentences;

[0014] Extract the semantic features of each of the segmented sentences;

[0015] Use the machine translation algorithm to perform Chinese translation on the segmented sentences according to the semantic features to obtain Chinese data of the segmented sentences;

[0016] Synthesize the Chinese data of the segmented sentences into Chinese text data in the order of the segmented sentences in the English text data.

[0017] Optionally, the step of inputting the English text data into a preset text encoder model for encoding to obtain an English text data vector set includes:

[0018] Convert each segmented sentence of the English text data into a unified fixed length to obtain standard sentences;

[0019] Use a preset tokenization method to perform word segmentation on the standard sentences to obtain segmented text data, and aggregate the segmented text data into a segmented sentence sequence;

[0020] Input the segmented sentence sequence into a preset text encoder for encoding to obtain word encoding, sentence encoding, and sentence position encoding;

[0021] Add the word encoding, the sentence encoding, and the sentence position encoding to obtain an English text data vector;

[0022] Aggregate the English text data vectors into an English text data vector set.

[0023] Optionally, the step of inputting the English text training data set into a preset teacher model for knowledge distillation to obtain a first probability value includes:

[0024] Use a preset sequence adversarial network model to convert the English text training data set into unlabeled English data;

[0025] Input the untagged English data into the teacher model for training to obtain untagged English training data;

[0026] Set the distillation temperature of the knowledge distillation, and perform knowledge distillation on the untagged English training data according to the distillation temperature to obtain soft-label English data;

[0027] Calculate the probability of the soft-label English data according to a preset classification function and the distillation temperature to obtain a first probability value.

[0028] Optionally, the calculating the probability of the soft-label English data according to a preset classification function and the distillation temperature to obtain a first probability value includes:

[0029] Use the following algorithm to calculate the probability of the soft-label English data according to a preset classification function and the distillation temperature to obtain a first probability value:

[0030]

[0031] where p i is the probability of the i-th class label in the soft-label data, exp is the exponential function, t is the distillation temperature parameter, z i is the i-th vector element in the soft-label English data, z j is the j-th vector element in the soft-label English data, and n is the number of vectors in the soft-label English data.

[0032] Optionally, the calculating the mean absolute error loss value of the first probability value and the second probability value using a preset loss function includes:

[0033] Use the following algorithm to calculate the mean absolute error loss value of the first probability value and the second probability value using a preset loss function:

[0034] loss = ||X - Y|| 2

[0035] where loss is the mean absolute error loss value, X is the first probability value, and Y is the second probability value.

[0036] Optionally, the performing model inference according to the Chinese text encoder and a preset image encoder to obtain a pre-trained Chinese text image model includes:

[0037] Obtain image data, and input the image data into the image encoder for encoding to obtain an image feature vector;

[0038] Input the Chinese text data into the Chinese text encoder for encoding to obtain a Chinese text feature vector;

[0039] Calculate the similarity between the image feature vector and the Chinese text feature vector;

[0040] Determine the feature vector with the highest similarity as the image corresponding to the Chinese text data;

[0041] Determine a Chinese text image pre-training model based on the image and the Chinese text data.

[0042] To solve the above problems, the present invention also provides an image retrieval device based on Chinese data, and the device includes:

[0043] A Chinese text data translation module, configured to obtain English text data of training data, and translate the English text data into Chinese text data by using a preset machine translation algorithm;

[0044] An English text data encoding module, configured to input the English text data into a preset text encoder model for encoding to obtain an English text data vector set, and train the English text data vector set by using a preset text image pre-training model to obtain an English text training data set;

[0045] A probability value acquisition module, configured to input the English text training data set into a preset teacher model for knowledge distillation to obtain a first probability value, and input the Chinese text data into a preset student model for knowledge distillation to obtain a second probability value;

[0046] A text encoder model optimization module, configured to calculate the mean absolute error loss value of the first probability value and the second probability value by using a preset loss function, and optimize the text encoder model according to the mean absolute error loss value to obtain a Chinese text encoder;

[0047] A pre-training model inference module, configured to perform model inference according to the Chinese text encoder and a preset image encoder to obtain a Chinese text image pre-training model;

[0048] An image retrieval module, configured to obtain Chinese data to be analyzed, and input the Chinese data to be analyzed into the Chinese text image pre-training model to obtain the image corresponding to the Chinese data to be analyzed.

[0049] To solve the above problems, the present invention also provides an electronic device, and the electronic device includes:

[0050] At least one processor; and,

[0051] A memory communicatively connected to the at least one processor; wherein,

[0052] The memory stores a computer program that can be executed by the at least one processor. When the computer program is executed by the at least one processor, the at least one processor is enabled to execute the above-mentioned image retrieval method based on Chinese data.

[0053] To solve the above problems, the present invention also provides a computer-readable storage medium. The computer-readable storage medium stores at least one computer program, and when the at least one computer program is executed by a processor in an electronic device, the above-mentioned image retrieval method based on Chinese data is implemented.

[0054] In an embodiment of the present invention, an English text data can be translated into Chinese text data through a machine translation model. The English text data is trained through a CLI model, and the trained English text data is used as the training data of a teacher model, while the Chinese text data is used as the training data of a student model. Based on knowledge distillation on the Chinese-English translation data, the model is optimized using the mean absolute error loss, thereby realizing the training of Chinese CLIP. Therefore, the image retrieval method based on Chinese data proposed by the present invention can solve the problem of low efficiency in image recognition according to Chinese data. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 It is a schematic flowchart of an image retrieval method based on Chinese data provided by an embodiment of the present invention;

[0056] Figure 2 It is a schematic flowchart of encoding English text data provided by an embodiment of the present invention;

[0057] Figure 3 It is a schematic flowchart of calculating a first probability value provided by an embodiment of the present invention;

[0058] Figure 4 It is a functional module diagram of an image retrieval device based on Chinese data provided by an embodiment of the present invention;

[0059] Figure 5 It is a schematic structural diagram of an electronic device for implementing the image retrieval method based on Chinese data provided by an embodiment of the present invention.

[0060] The realization, functional features and advantages of the objectives of the present invention will be further described in conjunction with embodiments with reference to the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0061] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0062] An embodiment of the present application provides an image retrieval method based on Chinese data. The execution subject of the image retrieval method based on Chinese data includes, but is not limited to, at least one of electronic devices such as a server, a terminal, etc. that can be configured to execute the method provided in the embodiment of the present application. In other words, the image retrieval method based on Chinese data can be executed by software or hardware installed on a terminal device or a server device, and the software can be a blockchain platform. The server includes, but is not limited to: a single server, a server cluster, a cloud server, or a cloud server cluster, etc. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.

[0063] Referring to Figure 1 As shown, it is a schematic flowchart of an image retrieval method based on Chinese data provided by an embodiment of the present invention. In this embodiment, the image retrieval method based on Chinese data includes:

[0064] S1. Obtain the English text data of the training data, and use a preset machine translation algorithm to translate the English text data into Chinese text data;

[0065] In the embodiment of the present invention, the English text data is the English text data in four hundred million pairs of image-text data collected by OpenAI (an artificial intelligence company) with great effort.

[0066] Specifically, computer statements with data scraping functions (such as Java statements, Python statements, etc.) can be used to scrape the stored English text data from a pre-determined storage area, and the storage area includes, but is not limited to, a database, a blockchain node, and a network cache.

[0067] In the embodiment of the present invention, the step of using a preset machine translation algorithm to translate the English text data into Chinese text data includes: splitting the sentence structure of the English text data to obtain split sentences; extracting the semantic features of each split sentence; using the machine translation algorithm to perform Chinese translation on the split sentences according to the semantic features to obtain Chinese data of the split sentences; and synthesizing the Chinese data of the split sentences into Chinese text data in the order of the split sentences in the English text data.

[0068] Specifically, the machine translation algorithm, also known as automatic translation, is a process of using a computer to convert one natural language (source language) into another natural language (target language). The machine translation process is divided into three stages: original text translation, original text-translation conversion, and translation generation. Machine systems are divided into two major categories: rule-based and corpus-based. For example, Chinese to English, English to Chinese, etc.

[0069] Specifically, in English sentences, most of the sentence strings end with commas. Therefore, commas are used as the sentence segmentation feature. For example, using CSP and CSC to represent full stops and commas respectively for sentence strings, the sentence is represented as: CSP = CSC, CSC, …, CSC. The part of the sentence after the CSC comma is cut and the cut mark “++” is added, that is, CSP = CSC, ++CSC, …, CSC.

[0070] Furthermore, a bag-of-words model with the function of extracting sentence semantic features can be used to extract the semantic features of the segmented sentences. Among them, the bag-of-words model represents the words in an English text in a way of putting them in bags to express the sentence semantic features. For example, there are two English sentences: ① John likes to watch movies, Mary likes too. ② John also likes to watch football games. The above two sentences can construct a dictionary A, A - {"John" 1, "likes" 2, "to" 3, "watch" 4, "movies" 5, "also" 6, "football" 7, "games" 8, "Mary" 9, "too" 10}. Then the vector representation of the first example sentence is: [1, 2, 1, 1, 1, 0, 0, 0, 1, 1]. The numbers in it represent the number of times the word at the current index position appears in the sentence. The 2 in the vector means that the word "likes" appears 2 times in this sentence. Therefore, the features of the sentences in the text can directly sum up the features of each word, and the semantic features can be verbs, conjunctions, grammar, tenses, etc. in English sentences.

[0071] Exemplarily, the clause translation synthesis template is used to convert the form of the subordinate clause during the translation process, and the sentence order is adjusted with the core clause as the main body according to the original meaning of the sentence. When there is a temporal parallel relationship in the sentence, "v+ing" is used to represent the progressive tense, "v+ed" is used to represent the past tense or the passive voice, and "to+v" is used to represent the purpose relationship. For example, in the sentence "the power converter supplies power to the inverter++to generate the operating voltage for the display part", two clauses are used in English, and it is translated as "The power converter supplies power to the inverter to generate the operating voltage for the display part".

[0072] S2. Input the English text data into a preset text encoder model for encoding to obtain an English text data vector set, and use a preset text image pre-training model to train the English text data vector set to obtain an English text training data set;

[0073] In the embodiment of the present invention, the text encoder model is based on the bert (pre-trained language representation model) structure, which realizes the conversion of text to dynamic word vectors and enhances the semantic information of the text vectors. Among them, the bert model is a truly bidirectional language model, and each word can utilize the context information of the word at the same time.

[0074] In the embodiment of the present invention, refer Figure 2 As shown, the inputting the English text data into a preset text encoder model for encoding to obtain an English text data vector set includes:

[0075] S21. Convert each segmented sentence of the English text data into a unified fixed length to obtain a standard sentence;

[0076] S22. Use a preset tokenization method to segment the standard sentence into word segmentation text data, and collect the word segmentation text data into a segmented sentence sequence;

[0077] S23. Input the segmented sentence sequence into a preset text encoder for encoding to obtain vocabulary encoding, sentence encoding, and sentence position encoding;

[0078] S24. Add the vocabulary encoding, the sentence encoding, and the sentence position encoding to obtain an English text data vector;

[0079] S25. Collect the English text data vectors into an English text data vector set.

[0080] Specifically, each segmented sentence of the English text data is converted into a unified fixed length and can be filled with 0 to obtain standard English text data, such as {I, love, eat, apples}, {I, love, Qtrade, 0}, and the tokenization method is Tokenization in NPL (Natural Language Technology). The process of splitting the original text into sub-units is called Tokenization.

[0081] Specifically, in BERT (Pre-trained Language Representation Model), there are a Token Embedding layer, a Segment Embeddings layer, and a Position Embeddings layer. The Token Embedding layer can achieve word encoding, the Segment Embeddings layer can achieve sentence encoding, and the Position Embeddings layer can achieve sentence position encoding.

[0082] Exemplarily, when the word encoding is {1, 0, 1, 1}, the sentence encoding is {1, 2, 0, 1}, and the position encoding is {0, 1, 1, 0}, then the vector encoding of the English text data is {2, 3, 2, 2}, and the vector encoding of the English text data is pooled to obtain an English text data vector set.

[0083] In the embodiment of the present invention, the text image pre-training model refers to the CLIP model. The CLIP model has been continuously extended to detection, image-text retrieval, and conditional generation of images, and has achieved amazing zero-shot effects in many fields. CLIP pre-training is a cross-modal contrastive learning model of the CLIP model on a large-scale image-text pair, which is trained based on English data. CLIP has a text encoder for text encoding and an image encoder for image encoding, and both encoders are based on transformers.

[0084] In the embodiment of the present invention, training the English text data vector set by using a preset text image pre-training model to obtain an English text training data set is to perform matrix multiplication on the English text data vector set and the image vector corresponding to a preset image to obtain the similarity between the two vectors, so as to complete the training of the English text data vector set and obtain an English text training data set.

[0085] S3. Input the English text training data set into a preset teacher model for knowledge distillation to obtain a first probability value, and input the Chinese text data into a preset student model for knowledge distillation to obtain a second probability value;

[0086] In the embodiments of the present invention, the Knowledge Distillation is a model compression method and a training method based on the "teacher-student network concept". The knowledge contained in the already trained model is distilled and extracted into another model. The teacher model is only used during the training process, and the same teacher model can be used to distill multiple student models.

[0087] In the embodiments of the present invention, refer Figure 3 As shown, inputting the English text training data set into a preset teacher model for knowledge distillation to obtain a first probability value includes:

[0088] S31. Using a preset sequence adversarial network model to convert the English text training data set into unlabeled English data;

[0089] S32. Inputting the unlabeled English data into the teacher model for training to obtain unlabeled English training data;

[0090] S33. Setting the distillation temperature of the knowledge distillation and performing knowledge distillation on the unlabeled English training data according to the distillation temperature to obtain soft-labeled English data;

[0091] S34. Calculating the probability of the soft-labeled English data according to a preset classification function and the distillation temperature to obtain a first probability value.

[0092] Specifically, the sequence adversarial network model (seqGAN model) is composed of a generation network and a discriminant network. Among them, the generation network imitates real data to generate a similar sample distribution to deceive the discriminant network, and the discriminant network is continuously updated during iteration to distinguish the generated samples from the real data. The generation network and the discriminant network play against each other until the data label balance is achieved.

[0093] Specifically, the distillation temperature is related to the probability distribution of the labels. When T (temperature) is 1, it is a special case of the original classification function (softmax). When the temperature is less than 1, the probability distribution is steeper than the original. When the temperature is greater than 1, the probability distribution is flatter than the original. Therefore, the higher the temperature, the more average the probability distribution of each value on the classification function.

[0094] Specifically, calculating the probability of the soft-labeled English data according to a preset classification function and the distillation temperature to obtain a first probability value includes:

[0095] Using the following algorithm to calculate the probability of the soft-labeled English data according to a preset classification function and the distillation temperature to obtain a first probability value:

[0096]

[0097] where p i is the probability of the i-th type of label in the soft label data, exp is the exponential function, t is the distillation temperature parameter, z i is the i-th vector element in the soft label English data, z j is the j-th vector element in the soft label English data, and n is the number of vectors in the soft label English data.

[0098] Specifically, the classification function is the softmax function. The soft label probability distribution of the teacher model is calculated using the softmax function. Suppose the soft label vector of the teacher model is [2.0, 1.0, 0.1]. Through the softmax function, a vector of [2.0, 1.0, 0.1] is transformed into [0.7, 0.2, 0.1], and the sum of each item is 1.

[0099] In one practical application scenario of the present invention, in a handwritten character recognition task, for a blurred picture of '3', due to the similarity of shapes, it has a certain probability of belonging to the '2' or '5' category. Therefore, during the distillation process, the trained teacher model provides the label probability distribution information of the softMax (output layer) to the student model as guidance during prediction. Among them, the probability distribution of the soft label English data contains information between categories. The features contained in this soft label are not available in the original unlabeled English training data. Therefore, by passing the soft label information to the student model, the learning ability of the student model can be improved.

[0100] In the embodiment of the present invention, the step of inputting the Chinese text data into a preset student model for knowledge distillation to obtain a second probability value is the same as the step of inputting the English text training data set into a preset teacher model for knowledge distillation to obtain a first probability value, and will not be elaborated here.

[0101] S4. Calculate the mean absolute error loss value of the first probability value and the second probability value using a preset loss function, and optimize the text encoder model according to the mean absolute error loss value to obtain a Chinese text encoder;

[0102] In the embodiment of the present invention, the first probability value may be the soft label probability output by the English text training data set in the teacher model, and the second probability value may be the hard label probability output by the Chinese text data in the student model. The hard label probability is the label output after distillation of the Chinese text data through label classification in the student model.

[0103] In the embodiment of the present invention, calculating the mean absolute error loss value of the first probability value and the second probability value using a preset loss function includes:

[0104] Use the following algorithm to calculate the mean absolute error loss value of the first probability value and the second probability value using a preset loss function:

[0105] loss = ||X - Y|| 2

[0106] where loss is the mean absolute error loss value, X is the first probability value, and Y is the second probability value.

[0107] In an embodiment of the present invention, optimizing the text encoder model according to the mean absolute error loss value to obtain a Chinese text encoder includes: calculating the optimal loss value of the mean absolute error loss value using a preset gradient descent method; when the optimal loss value is less than a preset loss threshold, determining the text encoder model as a Chinese text encoder.

[0108] Specifically, the gradient descent method is to find the gradient with the fastest descent of a curve function. Each optimization will be carried out along the direction of the fastest descent gradient. Use the gradient descent method to find the parameters that minimize the loss value of the loss function, and use the parameters to optimize the text encoder to obtain a Chinese text encoder, where the parameters include the parameters of the embedding layer, the parameters of the Transformer layer (multi-head attention mechanism), and the parameters of the softmax (classification layer).

[0109] S5. Perform model inference according to the Chinese text encoder and a preset image encoder to obtain a Chinese text image pre-training model.

[0110] In an embodiment of the present invention, the model inference is to simplify and use the capabilities of the model so that it can quickly and efficiently operate on unknown data to obtain expected results.

[0111] In an embodiment of the present invention, performing model inference according to the Chinese text encoder and a preset image encoder to obtain a Chinese text image pre-training model includes: obtaining image data, and inputting the image data into the image encoder for encoding to obtain an image feature vector; inputting the Chinese text data into the Chinese text encoder for encoding to obtain a Chinese text feature vector; calculating the similarity between the image feature vector and the Chinese text feature vector; determining the feature vector with the highest similarity as the image corresponding to the Chinese text data; determining the Chinese text image pre-training model according to the image and the Chinese text data.

[0112] Specifically, calculating the similarity between the image feature vector and the Chinese text feature vector includes:

[0113] Calculate the similarity between the image feature vector and the Chinese text feature vector using the following algorithm:

[0114]

[0115] where T is the similarity, and x i is the i-th vector element of the image feature, and y i is the i-th vector element of the Chinese text feature vector.

[0116] Specifically, the steps of inputting the Chinese text data into the Chinese text encoder for encoding to obtain the Chinese text feature vector and inputting the English text data into the preset text encoder model for encoding to obtain the English text data vector set are the same, and will not be elaborated here.

[0117] S6. Obtain the Chinese data to be analyzed, and input the Chinese data to be analyzed into the Chinese text image pre-training model to obtain the image corresponding to the Chinese data to be analyzed.

[0118] In the embodiment of the present invention, the trained Chinese CLIP (Chinese text image pre-training model) can realize the retrieval within Chinese or pictures, and can also realize the cross-modal retrieval task.

[0119] The embodiment of the present invention can translate the English text data into Chinese text data through a machine translation model, train the English text data through the CLI model, use the trained English text data as the training data of the teacher model, use the Chinese text data as the training data of the student model, and optimize the model based on knowledge distillation on the Chinese-English translation data using the mean absolute error loss, thereby realizing the training of the Chinese CLIP. Therefore, the image retrieval method based on Chinese data proposed by the present invention can solve the problem of low efficiency in image recognition according to Chinese data.

[0120] As Figure 4 shown, it is a functional module diagram of an image retrieval device based on Chinese data provided by an embodiment of the present invention.

[0121] The image retrieval device 100 based on Chinese data of the present invention can be installed in an electronic device. According to the realized functions, the image retrieval device 100 based on Chinese data can include a Chinese text data translation module 101, an English text data encoding module 102, a probability value acquisition module 103, a text encoder model optimization module 104, a pre-training model inference module 105, and an image retrieval module 106. The modules of the present invention can also be referred to as units, which refer to a series of computer program segments that can be executed by a processor of an electronic device and can complete fixed functions, and are stored in the memory of the electronic device.

[0122] In this embodiment, the functions of each module / unit are as follows:

[0123] The Chinese text data translation module 101 is used to obtain the English text data of the training data and translate the English text data into Chinese text data by using a preset machine translation algorithm;

[0124] The English text data encoding module 102 is used to input the English text data into a preset text encoder model for encoding to obtain an English text data vector set, and use a preset text image pre-training model to train the English text data vector set to obtain an English text training data set;

[0125] The probability value obtaining module 103 is used to input the English text training data set into a preset teacher model for knowledge distillation to obtain a first probability value, and input the Chinese text data into a preset student model for knowledge distillation to obtain a second probability value;

[0126] The text encoder model optimization module 104 is used to calculate the mean absolute error loss value of the first probability value and the second probability value by using a preset loss function, and optimize the text encoder model according to the mean absolute error loss value to obtain a Chinese text encoder;

[0127] The pre-trained model inference module 105 is used to perform model inference according to the Chinese text encoder and a preset image encoder to obtain a Chinese text image pre-training model;

[0128] The image retrieval module 106 is used to obtain the Chinese data to be analyzed, and input the Chinese data to be analyzed into the Chinese text image pre-training model to obtain the image corresponding to the Chinese data to be analyzed.

[0129] Specifically, each module in the image retrieval device 100 based on Chinese data in the embodiment of the present invention adopts the same technical means as those in the above Figures 1 to 3 The image retrieval method based on Chinese data described in the text, and can produce the same technical effects, which will not be elaborated here.

[0130] As Figure 5 shown, it is a schematic structural diagram of an electronic device for implementing the image retrieval method based on Chinese data provided by an embodiment of the present invention.

[0131] The electronic device 1 may include a processor 10, a memory 11, a communication bus 12, and a communication interface 13, and may further include a computer program stored in the memory 11 and executable on the processor 10, such as an image retrieval program based on Chinese data.

[0132] Among them, in some embodiments, the processor 10 may be composed of an integrated circuit. For example, it may be composed of a single packaged integrated circuit, or may be composed of multiple packaged integrated circuits with the same or different functions, including a combination of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 10 is the control core (Control Unit) of the electronic device, connecting various components of the entire electronic device through various interfaces and circuits, and by running or executing programs or modules stored in the memory 11 (such as executing an image retrieval program based on Chinese data, etc.), and calling the data stored in the memory 11, to perform various functions of the electronic device and process data.

[0133] The memory 11 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, mobile hard disks, multimedia cards, card-type memories (such as SD or DX memories, etc.), magnetic memories, magnetic disks, optical discs, etc. In some embodiments, the memory 11 may be an internal storage unit of the electronic device, such as the mobile hard disk of the electronic device. In some other embodiments, the memory 11 may also be an external storage device of the electronic device, such as a plug-in mobile hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device. Further, the memory 11 may also include both an internal storage unit and an external storage device of the electronic device. The memory 11 can be used not only to store application software installed on the electronic device and various types of data, such as the code of an image retrieval program based on Chinese data, etc., but also to temporarily store data that has been output or will be output.

[0134] The communication bus 12 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. The bus is set to realize the connection and communication between the memory 11 and at least one processor 10, etc.

[0135] The communication interface 13 is used for communication between the above-mentioned electronic device and other devices, and includes a network interface and a user interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), and is generally used to establish a communication connection between this electronic device and other electronic devices. The user interface may be a display, an input unit (such as a keyboard), and optionally, the user interface may also be a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. Among them, the display may also be appropriately referred to as a display screen or a display unit, and is used to display the information processed in the electronic device and to display a visual user interface.

[0136] Only the electronic device with components is shown in the figure. Those skilled in the art can understand that the structure shown in the figure does not constitute a limitation on the electronic device, and it may include fewer or more components than shown in the figure, or combine certain components, or have different component arrangements.

[0137] For example, although not shown, the electronic device may further include a power source (such as a battery) for supplying power to each component. Preferably, the power source may be logically connected to the at least one processor 10 through a power management device, so as to implement functions such as charge management, discharge management, and power consumption management through the power management device. The power source may also include any components such as one or more DC or AC power sources, a recharge device, a power failure detection circuit, a power converter or an inverter, and a power status indicator. The electronic device may also include a variety of sensors, a Bluetooth module, a Wi-Fi module, etc., which will not be elaborated here.

[0138] It should be understood that the above embodiments are only for illustration purposes and are not limited by this structure in the scope of the patent application.

[0139] The image retrieval program based on Chinese data stored in the memory 11 in the electronic device 1 is a combination of multiple instructions. When running in the processor 10, it can implement:

[0140] Obtain the English text data of the training data, and use a preset machine translation algorithm to translate the English text data into Chinese text data;

[0141] Input the English text data into a preset text encoder model for encoding to obtain an English text data vector set, and use a preset text-image pre-training model to train the English text data vector set to obtain an English text training data set;

[0142] Input the English text training dataset into a preset teacher model for knowledge distillation to obtain a first probability value, and input the Chinese text data into a preset student model for knowledge distillation to obtain a second probability value;

[0143] Use a preset loss function to calculate the mean absolute error loss value of the first probability value and the second probability value, and optimize the text encoder model according to the mean absolute error loss value to obtain a Chinese text encoder;

[0144] Perform model inference based on the Chinese text encoder and a preset image encoder to obtain a Chinese text image pre-training model;

[0145] Obtain the Chinese data to be analyzed, and input the Chinese data to be analyzed into the Chinese text image pre-training model to obtain the image corresponding to the Chinese data to be analyzed.

[0146] Specifically, the specific implementation method of the above instructions by the processor 10 can refer to the description of the relevant steps in the corresponding embodiment of the attached drawings, which will not be elaborated here.

[0147] Furthermore, if the modules / units integrated in the electronic device 1 are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. The computer-readable storage medium can be volatile or non-volatile. For example, the computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory).

[0148] The present invention also provides a computer-readable storage medium, and the readable storage medium stores a computer program, and when the computer program is executed by a processor of an electronic device, it can implement:

[0149] Obtain the English text data of the training data, and use a preset machine translation algorithm to translate the English text data into Chinese text data;

[0150] Input the English text data into a preset text encoder model for encoding to obtain an English text data vector set, and use a preset text image pre-training model to train the English text data vector set to obtain an English text training dataset;

[0151] Input the English text training dataset into a preset teacher model for knowledge distillation to obtain a first probability value, and input the Chinese text data into a preset student model for knowledge distillation to obtain a second probability value;

[0152] Calculate the mean absolute error loss value of the first probability value and the second probability value by using a preset loss function, and optimize the text encoder model according to the mean absolute error loss value to obtain a Chinese text encoder;

[0153] Perform model inference according to the Chinese text encoder and a preset image encoder to obtain a Chinese text-image pre-training model;

[0154] Obtain the Chinese data to be analyzed, and input the Chinese data to be analyzed into the Chinese text-image pre-training model to obtain an image corresponding to the Chinese data to be analyzed.

[0155] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation.

[0156] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0157] In addition, in each embodiment of the present invention, the functional modules can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware, or in the form of a combination of hardware and software functional modules.

[0158] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention.

[0159] Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, it is intended to cover all changes falling within the meaning and scope of the equivalent elements of the claims in the present invention. Any reference signs attached to the claims should not be regarded as limiting the claimed rights.

[0160] The blockchain referred to in the present invention is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithm. Blockchain, in essence, is a decentralized database, a series of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity of the information (anti-counterfeiting) and generate the next block. The blockchain can include the blockchain underlying platform, the platform product service layer, and the application service layer, etc.

[0161] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.

[0162] In addition, it is obvious that the term "including" does not exclude other units or steps, and the singular does not exclude the plural. The multiple units or devices stated in the system claims can also be implemented by one unit or device through software or hardware. The terms such as first and second are used to represent names and do not indicate any specific order.

[0163] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. An image retrieval method based on Chinese data, characterized in that, the method includes: obtaining the English text data of the training data, and translating the English text data into Chinese text data by using a preset machine translation algorithm; inputting the English text data into a preset text encoder model for encoding to obtain an English text data vector set, and training the English text data vector set by using a preset text-image pre-training model to obtain an English text training data set; inputting the English text training data set into a preset teacher model for knowledge distillation to obtain a first probability value, and inputting the Chinese text data into a preset student model for knowledge distillation to obtain a second probability value; calculating the mean absolute error loss value of the first probability value and the second probability value by using a preset loss function, and optimizing the text encoder model according to the mean absolute error loss value to obtain a Chinese text encoder; performing model inference according to the Chinese text encoder and a preset image encoder to obtain a Chinese text-image pre-training model; obtaining the Chinese data to be analyzed, and inputting the Chinese data to be analyzed into the Chinese text-image pre-training model to obtain the image corresponding to the Chinese data to be analyzed; wherein, the step of inputting the English text training data set into a preset teacher model for knowledge distillation to obtain a first probability value includes: converting the English text training data set into unlabeled English data by using a preset sequence adversarial network model; inputting the unlabeled English data into the teacher model for training to obtain unlabeled English training data; setting the distillation temperature of the knowledge distillation, and performing knowledge distillation on the unlabeled English training data according to the distillation temperature to obtain soft-labeled English data; calculating the probability of the soft-labeled English data according to a preset classification function and the distillation temperature to obtain a first probability value; the step of performing model inference according to the Chinese text encoder and a preset image encoder to obtain a Chinese text-image pre-training model includes: obtaining image data, and inputting the image data into the image encoder for encoding to obtain an image feature vector; inputting the Chinese text data into the Chinese text encoder for encoding to obtain a Chinese text feature vector; calculating the similarity between the image feature vector and the Chinese text feature vector; determining the feature vector with the highest similarity as the image corresponding to the Chinese text data; and determining the Chinese text-image pre-training model according to the image and the Chinese text data.

2. The image retrieval method based on Chinese data according to claim 1, characterized in that, the step of translating the English text data into Chinese text data by using a preset machine translation algorithm includes: performing syntactic structure segmentation on the English text data to obtain segmented sentences; extracting the semantic features of each segmented sentence; using the machine translation algorithm to perform Chinese translation on the segmented sentences according to the semantic features to obtain segmented sentence Chinese data; Combine the segmented statement Chinese data into Chinese text data in the order of each segmented statement in the English text data.

3. The image retrieval method based on Chinese data according to claim 1, wherein, the inputting the English text data into a preset text encoder model for encoding to obtain an English text data vector set includes: converting each segmented statement of the English text data into a unified fixed length to obtain a standard statement; performing word segmentation on the standard statement by using a preset tokenization method to obtain segmented text data, and aggregating the segmented text data into a segmented statement sequence; inputting the segmented statement sequence into a preset text encoder for encoding to obtain a vocabulary encoding, a statement encoding, and a statement position encoding; adding the vocabulary encoding, the statement encoding, and the statement position encoding to obtain an English text data vector; aggregating the English text data vectors into an English text data vector set.

4. The image retrieval method based on Chinese data according to claim 1, wherein, the calculating the probability of the soft label English data according to a preset classification function and the distillation temperature to obtain a first probability value includes: using the following algorithm to calculate the probability of the soft label English data according to a preset classification function and the distillation temperature to obtain a first probability value: Among them, is the probability of the th type of label in the soft label English data, is an exponential function, is the distillation temperature parameter, is the th vector element in the soft label English data, is the th vector element in the soft label English data, is the number of vectors in the soft label English data.

5. The image retrieval method based on Chinese data according to any one of claims 1 to 4, wherein, the calculating the mean absolute error loss value of the first probability value and the second probability value by using a preset loss function includes: using the following algorithm to calculate the mean absolute error loss value of the first probability value and the second probability value by using a preset loss function: Among them, is the mean absolute error loss value, is the first probability value, is the second probability value.

6. An image retrieval device based on Chinese data for implementing the image retrieval method based on Chinese data according to any one of claims 1 to 5, wherein, the device includes: a Chinese text data translation module for obtaining the English text data of the training data, and translating the English text data into Chinese text data by using a preset machine translation algorithm; an English text data encoding module for inputting the English text data into a preset text encoder model for encoding to obtain an English text data vector set, and training the English text data vector set by using a preset text image pre-training model to obtain an English text training data set; a probability value obtaining module for inputting the English text training data set into a preset teacher model for knowledge distillation to obtain a first probability value, and inputting the Chinese text data into a preset student model for knowledge distillation to obtain a second probability value; a text encoder model optimization module for calculating the mean absolute error loss value of the first probability value and the second probability value by using a preset loss function, and optimizing the text encoder model according to the mean absolute error loss value to obtain a Chinese text encoder; a pre-trained model inference module for performing model inference according to the Chinese text encoder and a preset image encoder to obtain a Chinese text image pre-trained model; An image retrieval module, configured to obtain Chinese data to be analyzed, and input the Chinese data to be analyzed into the pre-trained Chinese text image model to obtain an image corresponding to the Chinese data to be analyzed.

7. An electronic device, characterized in that, the electronic device includes: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the Chinese data-based image retrieval method according to any one of claims 1 to 5.

8. A computer-readable storage medium storing a computer program, characterized in that, when the computer program is executed by a processor, it implements the Chinese data-based image retrieval method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Trademark image card distribution method and device, electronic equipment and storage medium

    CN114610931A