Chinese image and text retrieval model training method, device, equipment and medium based on CLIP
By training and adjusting the pre-built Chinese encoder, replacing the English encoder in the CLIP graphic search model, the problem of insufficient graphic search performance in Chinese graphic search is solved, and more efficient Chinese graphic search ability is achieved.
Patent Information
- Application Number
- CN202210730910.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-24
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2042-06-24
AI Technical Summary
The performance of the CLIP-based graphic search model in Chinese graphic search has not yet reached the English level, and cannot be directly transferred to the learning of Chinese graphic search pairs.
By obtaining the Chinese text training set, randomly selecting positive and negative samples, training the pre-constructed Chinese encoder, calculating the similarity between the positive and negative samples vectors, adjusting the encoder parameters until the preset training conditions are reached, and the Chinese encoder that completed the training is obtained. Then, it is replaced by the English encoder in the CLIP graphic and text search model, and is subject to graphic matching training until the second preset training condition is met, and the target Chinese graphic and text CLIP retrieval model is obtained.
The performance of the CLIP graphic search model in Chinese graphic search is improved, so that it can more effectively understand and learn Chinese graphic search pairs and improve retrieval accuracy.
Smart Images

Figure CN115221276B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a CLIP-based Chinese image and text retrieval model training method, device, electronic equipment and computer-readable storage medium. Background Art
[0002] With the development of search engine technology, pure text-based search can no longer meet the needs of people's daily life or work. As graphic and text information is more intuitive and richer, graphic and text search functions that combine images and text are becoming increasingly important.
[0003] Currently, in the field of image and text search, the performance of the image and text retrieval model based on CLIP is very powerful and has been widely used. However, the image and text retrieval model based on CLIP is obtained by training 400 million pairs of English image and text. The model has a good learning and understanding ability for English, but it cannot be directly transferred to the learning of Chinese image and text pairs. Its performance in Chinese image and text search needs to be improved. Summary of the invention
[0004] The present invention provides a CLIP-based Chinese image-text retrieval model training method, device, electronic device and computer-readable storage medium, the main purpose of which is to improve the Chinese image-text retrieval performance of the CLIP image-text retrieval model.
[0005] To achieve the above object, the present invention provides a CLIP-based Chinese image-text retrieval model training method, comprising:
[0006] Step A, obtaining a Chinese text training set, randomly selecting a Chinese text from the Chinese text training set as a positive sample, and other Chinese texts as negative samples;
[0007] Step B, performing Chinese recognition training on the pre-built Chinese encoder to generate a preset number of positive sample vectors of the positive samples and a set of negative sample vectors of the negative samples;
[0008] Step C, calculating a first similarity between the positive sample vectors, and calculating a second similarity between the positive sample vector and the negative sample vector set;
[0009] Step D, adjusting the parameters of the pre-constructed Chinese encoder according to the first similarity and the second similarity, and returning to the above step B, until the pre-constructed Chinese encoder meets the first preset training condition, exiting the Chinese recognition training, and obtaining a trained Chinese encoder;
[0010] Step E: using the trained Chinese encoder to replace the English encoder in the pre-built CLIP image-text retrieval model to obtain a Chinese image-text CLIP retrieval model;
[0011] Step F, obtaining a Chinese image-text pair training set, using the Chinese image-text pair training set to perform image-text matching training on the Chinese image-text CLIP retrieval model, until the image-text matching training meets a second preset training condition, then exiting the image-text matching training to obtain a target Chinese image-text CLIP retrieval model.
[0012] Optionally, the Chinese recognition training of generating a preset number of positive sample vectors of the positive samples and generating a negative sample vector set of the negative samples using the pre-constructed Chinese encoder comprises:
[0013] Randomly generating a preset number of parameter values of preset parameters in the pre-built Chinese encoder;
[0014] Selecting one parameter value from a preset number of parameter values in turn to assign a value to the preset parameter;
[0015] Performing vector conversion on the positive sample using the pre-constructed Chinese encoder after the assignment to obtain a positive sample vector corresponding to the positive sample;
[0016] All positive sample vectors corresponding to the positive samples are collected to obtain the preset number of positive sample vectors.
[0017] Optionally, the calculating a second similarity between the positive sample vector and the negative sample vector set includes:
[0018] Performing a clustering operation on the negative sample vector set until the clustering operation meets a preset clustering condition, exiting the clustering operation, and obtaining each cluster center after clustering;
[0019] Using a preset distance function, sequentially calculating the unilateral distance between each of the cluster centers and each of the positive sample vectors;
[0020] Calculate the comprehensive distance between the negative sample vector set and the positive sample vector according to all the one-sided distances;
[0021] The comprehensive distance is inverted to obtain a second similarity between the positive sample vector and the negative sample vector set.
[0022] Optionally, before adjusting the parameters of the pre-constructed Chinese encoder according to the first similarity and the second similarity, the method further includes:
[0023] performing normalization processing on the first similarity and the second similarity;
[0024] Calculating a harmonic value between the normalized first similarity and the normalized second similarity using a preset harmonic formula;
[0025] When the reconciliation value does not meet a preset reconciliation value threshold, the parameters of the pre-built Chinese encoder are adjusted.
[0026] Optionally, the calculating the harmonic value between the normalized first similarity and the normalized second similarity by using a preset harmonic formula includes:
[0027] The harmonic value is calculated using the following preset harmonic formula:
[0028] F=α*S1+β*S2
[0029] Among them, F represents the harmonic value, S1 represents the normalized first similarity, S2 represents the normalized second similarity, α is the first harmonic coefficient, β is the second harmonic coefficient, and through the first harmonic coefficient and the second harmonic coefficient, when S1 increases and S2 decreases, the F value increases, and when S1 decreases and S2 increases, the F value decreases.
[0030] Optionally, performing image-text matching training on the Chinese image-text CLIP retrieval model using the Chinese image-text pair training set until the image-text matching training meets a second preset training condition, then exiting the image-text matching training to obtain a target Chinese image-text CLIP retrieval model, including:
[0031] Using the text editor in the Chinese image-text CLIP retrieval model, the text information in the Chinese image-text pair training set is converted into vectors to obtain a text vector set, and using the image editor in the Chinese image-text CLIP retrieval model, the image information in the Chinese image-text pair training set is converted into vectors to obtain an image vector set;
[0032] By utilizing the cross-modal contrastive learning mechanism in the Chinese image-text CLIP retrieval model, the image and text are matched according to the text vector set and the image vector set to obtain a predicted image-text matching result.
[0033] Calculating the loss value between the predicted image-text matching result and the actual result corresponding to the Chinese image-text pair training set;
[0034] Determining whether the loss value satisfies the second preset training condition;
[0035] If the loss value does not meet the second preset training condition, returning to the above-mentioned steps of using the text editor in the Chinese image-text CLIP retrieval model to perform vector conversion on the text information in the Chinese image-text pair training set to obtain a text vector set, and using the image editor in the Chinese image-text CLIP retrieval model to perform vector conversion on the image information in the Chinese image-text pair training set to obtain an image vector set;
[0036] If the loss value satisfies the second preset training condition, the image-text matching training is exited to obtain the target Chinese image-text CLIP retrieval model.
[0037] Optionally, obtaining a Chinese image-text pair training set includes:
[0038] Obtain an English picture-text pair set from a preset English picture-text pair library;
[0039] Translate the English picture-text pair into Chinese to obtain a Chinese picture-text pair training set
[0040] In order to solve the above problems, the present invention also provides a Chinese picture and text retrieval model training device based on CLIP, the device comprising:
[0041] A Chinese text training set acquisition module is used to acquire a Chinese text training set, randomly select one Chinese text from the Chinese text training set as a positive sample, and other Chinese texts as negative samples;
[0042] A Chinese encoder training module is used to perform Chinese recognition training on a pre-constructed Chinese encoder to generate a preset number of positive sample vectors of the positive samples and a negative sample vector set of the negative samples; calculate a first similarity between the positive sample vectors, and calculate a second similarity between the positive sample vectors and the negative sample vector set; adjust the parameters of the pre-constructed Chinese encoder according to the first similarity and the second similarity until the pre-constructed Chinese encoder meets the first preset training condition, exit the Chinese recognition training, and obtain a Chinese encoder that has completed the training;
[0043] A Chinese picture-text model construction module is used to replace the English encoder in the pre-constructed CLIP picture-text retrieval model with the trained Chinese encoder to obtain a Chinese picture-text CLIP retrieval model;
[0044] The Chinese picture-text model training module is used to obtain a Chinese picture-text pair training set, and use the Chinese picture-text pair training set to perform picture-text matching training on the Chinese picture-text CLIP retrieval model until the picture-text matching training meets a second preset training condition, then exit the picture-text matching training to obtain a target Chinese picture-text CLIP retrieval model.
[0045] In order to solve the above problem, the present invention further provides an electronic device, the electronic device comprising:
[0046] a memory storing at least one computer program; and
[0047] The processor executes the program stored in the memory to implement the above-mentioned CLIP-based Chinese image and text retrieval model training method.
[0048] In order to solve the above problems, the present invention also provides a computer-readable storage medium, in which at least one computer program is stored. The at least one computer program is executed by a processor in an electronic device to implement the above-mentioned CLIP-based Chinese image and text retrieval model training method.
[0049] The embodiment of the present invention uses a pre-constructed Chinese encoder to perform vector conversion on Chinese positive samples and Chinese negative samples, obtains a first similarity between positive sample vectors and a second similarity between the positive sample vectors and the negative sample vectors by calculation, uses the first similarity and the second similarity to adjust the parameters of the pre-constructed Chinese encoder and performs Chinese recognition training on the positive and negative sample vectors, so that the trained Chinese encoder has the ability to understand and learn Chinese, and then uses the trained Chinese encoder to replace the English encoder in the pre-constructed CLIP image-text retrieval model to obtain a Chinese image-text CLIP retrieval model, and then uses a Chinese image-text pair training set to train the Chinese image-text CLIP retrieval model to obtain a target Chinese image-text CLIP retrieval model with Chinese image-text pair retrieval capability. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 A schematic diagram of a flow chart of a CLIP-based Chinese image-text retrieval model training method provided in one embodiment of the present invention;
[0051] Figure 2 A schematic diagram of a detailed implementation flow of one of the steps in the CLIP-based Chinese image-text retrieval model training method provided in one embodiment of the present invention;
[0052] Figure 3 A schematic diagram of a detailed implementation flow of one of the steps in the CLIP-based Chinese image-text retrieval model training method provided in one embodiment of the present invention;
[0053] Figure 4 A schematic diagram of a detailed implementation flow of one of the steps in the CLIP-based Chinese image-text retrieval model training method provided in one embodiment of the present invention;
[0054] Figure 5A schematic diagram of a detailed implementation flow of one of the steps in the CLIP-based Chinese image-text retrieval model training method provided in one embodiment of the present invention;
[0055] Figure 6 A functional module diagram of a Chinese image-text retrieval model training device based on CLIP provided in one embodiment of the present invention;
[0056] Figure 7 A schematic diagram of the structure of an electronic device for implementing the CLIP-based Chinese image-text retrieval model training method provided by an embodiment of the present invention.
[0057] The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0058] It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.
[0059] The embodiment of the present application provides a Chinese image-text retrieval model training method based on CLIP. The execution subject of the Chinese image-text retrieval model training method based on CLIP includes but is not limited to at least one of the electronic devices such as a server and a terminal that can be configured to execute the method provided by the embodiment of the present application. In other words, the Chinese image-text retrieval model training method based on CLIP can be executed by software or hardware installed on a terminal device or a server device, and the software can be a blockchain platform. The server can be an independent server, or it can be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (Content Delivery Network, CDN), and big data and artificial intelligence platforms.
[0060] Reference Figure 1 FIG. 1 is a flow chart of a CLIP-based Chinese image-text retrieval model training method according to an embodiment of the present invention. In this embodiment, the CLIP-based Chinese image-text retrieval model training method includes:
[0061] Step A, obtaining a Chinese text training set, randomly selecting a Chinese text from the Chinese text training set as a positive sample, and other Chinese texts as negative samples;
[0062] In an embodiment of the present invention, a Chinese text data set can be obtained from a designated open source natural language learning model corpus, or a Python script with data crawling capabilities can be used to crawl text information from a designated website, and then the crawled text information can be used to construct a Chinese text data set. Furthermore, the acquired and constructed Chinese text data sets are preprocessed by removing stop words, removing useless symbols, and other preprocessing operations to obtain the Chinese text training set.
[0063] In the embodiment of the present invention, a preset random method may be used to randomly select a Chinese text from the Chinese text training set as a positive sample, and the remaining Chinese texts may be used as negative samples.
[0064] Step B, performing Chinese recognition training on the pre-built Chinese encoder to generate a preset number of positive sample vectors of the positive samples and a set of negative sample vectors of the negative samples;
[0065] It can be understood that a CLIP image-text retrieval model usually includes two parts: a text encoder and an image encoder, wherein the image encoder implements feature extraction and vector conversion of image information, and the text encoder mainly implements feature extraction and vector conversion of text information, and the current text editor is mainly for English text.
[0066] In an embodiment of the present invention, the relevant parameters of the text encoder can be initialized based on the structure of the text encoder in the current CLIP model to obtain the pre-constructed Chinese encoder, and the pre-constructed Chinese encoder can be used to perform corresponding Chinese recognition training. On the one hand, the Chinese learning ability of the text encoder can be improved. On the other hand, the trained Chinese editor can be used to cooperate with the image encoder in the CLIP image and text retrieval model to achieve the retrieval ability of Chinese images and texts.
[0067] In an embodiment of the present invention, the Chinese recognition training includes generating a preset number of positive sample vectors of the positive samples using a pre-built Chinese encoder,
[0068] For details, see Figure 2 The method of using a pre-built Chinese encoder to generate a preset number of positive sample vectors of the positive samples includes:
[0069] S21, randomly generating a preset number of parameter values of preset parameters in the pre-built Chinese encoder;
[0070] S22, selecting a parameter value from a preset number of parameter values in turn to assign a value to the preset parameter;
[0071] S23, using the pre-built Chinese encoder after the assignment to perform vector conversion on the positive sample to obtain a positive sample vector corresponding to the positive sample;
[0072] S24. Gather all positive sample vectors corresponding to the positive samples to obtain the preset number of positive sample vectors.
[0073] In an embodiment of the present invention, the preset parameter may be any parameter in the pre-built Chinese encoder that can change the output result, such as the convolution kernel size, the number of neurons or the number of neuron connections.
[0074] In the embodiment of the present invention, by assigning different values to the preset parameters each time, different vector outputs are obtained for the same positive sample under different preset parameter values.
[0075] In an embodiment of the present invention, the negative samples can be input into the pre-constructed Chinese encoder in batches or all at once according to the actual number of the negative samples, and the pre-constructed Chinese encoder can be used to extract the vector features of each negative sample to obtain the corresponding negative sample vector, and all the negative sample vectors can be collected to obtain the negative sample vector set.
[0076] Step C, calculating a first similarity between the positive sample vectors, and calculating a second similarity between the positive sample vector and the negative sample vector set;
[0077] In an embodiment of the present invention, the number of the positive sample vectors is smaller than the number of the negative sample vector set, and the distance or similarity between the positive sample vectors can be directly calculated using Euclidean distance, Manhattan distance, or cosine similarity to obtain the first similarity.
[0078] In an embodiment of the present scheme, the data volume of the negative sample vectors in the negative sample vector set may be much larger than the data volume of the positive sample vectors. In order to improve the calculation efficiency, it is preferable to first perform a clustering operation on the negative sample vector set, and then use the cluster centers obtained by the clustering results to calculate the distance or similarity between each cluster center and each of the positive sample vectors to obtain the second similarity.
[0079] For details, see Figure 3 As shown, the calculating the second similarity between the positive sample vector and the negative sample vector set includes:
[0080] S31, performing a clustering operation on the negative sample vector set until the clustering operation satisfies a preset clustering condition, then exiting the clustering operation, and obtaining each cluster center after clustering;
[0081] S32, using a preset distance function, sequentially calculating the unilateral distance between each of the cluster centers and each of the positive sample vectors;
[0082] S33, calculating the comprehensive distance between the negative sample vector set and the positive sample vector according to all the one-sided distances;
[0083] S34. Negate the comprehensive distance to obtain a second similarity between the positive sample vector and the negative sample vector set.
[0084] In the embodiment of the present invention, the preset clustering condition may be that the cluster centers obtained in each clustering operation tend to be stable, or that the number of clustering operations reaches a preset maximum number of clustering operations.
[0085] In the embodiment of the present invention, the preset distance function may be a distance function such as Euclidean distance, Manhattan distance or cosine similarity.
[0086] In the embodiment of the present invention, the comprehensive distance may be obtained by calculating the variance values of all the unilateral distances.
[0087] In another optional embodiment of the present invention, a different weight can be assigned to each of the unilateral distances, and the weighted distance assigned with the weight corresponding to each of the unilateral distances is calculated, and then the average of all the weighted distances is calculated, and finally the calculated average is used as the comprehensive distance.
[0088] It can be understood that the greater the distance between the two, the smaller the similarity between the two. Therefore, the second similarity is obtained by negating the comprehensive distance, and the second similarity can normally reflect the degree of similarity between the two.
[0089] Step D, adjusting the parameters of the pre-constructed Chinese encoder according to the first similarity and the second similarity, and returning to the above step B, until the pre-constructed Chinese encoder meets the first preset training condition, exiting the Chinese recognition training, and obtaining a trained Chinese encoder;
[0090] It can be understood that the preset number of positive sample vectors are obtained by converting the same Chinese sample through the pre-constructed Chinese encoder, and the negative sample vector set corresponds to different Chinese samples and is different from the positive sample. In theory, the stronger the Chinese text learning ability of the pre-constructed Chinese encoder, the greater the first similarity between the corresponding positive sample vectors, and the smaller the second similarity between the preset number of positive sample vectors and the negative sample vector set, that is, the learning results of the same text are very close, while the learning results between different texts are very different. Therefore, according to the principle that the first similarity value is getting larger and the corresponding second similarity value is getting smaller, constantly adjusting the parameters of the pre-constructed Chinese encoder can promote the improvement of the Chinese learning ability of the pre-constructed Chinese encoding.
[0091] For details, see Figure 4 As shown, before adjusting the parameters of the pre-built Chinese encoder according to the first similarity and the second similarity, the method may further include:
[0092] S41, normalizing the first similarity and the second similarity;
[0093] S42, calculating a harmonic value between the normalized first similarity and the normalized second similarity using a preset harmonic formula;
[0094] S43: When the harmonization value does not meet a preset harmonization value threshold, adjusting the parameters of the pre-constructed Chinese encoder.
[0095] It is understandable that the calculation methods of the first similarity and the second similarity may be the same or different, and the obtained result standards may be inconsistent. Therefore, it is necessary to perform corresponding normalization processing on the first similarity and the second similarity.
[0096] In the embodiment of the present invention, the preset reconciliation formula may adopt the following calculation formula:
[0097] F=α*S1+β*S2
[0098] Among them, F represents the harmonic value, S1 represents the normalized first similarity, S2 represents the normalized second similarity, α is the first harmonic coefficient, β is the second harmonic coefficient, and through the first harmonic coefficient and the second harmonic coefficient, when S1 increases and S2 decreases, the F value increases, and when S1 decreases and S2 increases, the F value decreases.
[0099] In the embodiment of the present invention, the preset reconciliation threshold can be set according to the actual training results of the pre-built Chinese encoder.
[0100] In the embodiment of the present invention, the first preset training condition may be to exit the corresponding training when the harmony value reaches or exceeds the preset harmony value threshold.
[0101] Step E: using the trained Chinese encoder to replace the English encoder in the pre-built CLIP image-text retrieval model to obtain a Chinese image-text CLIP retrieval model;
[0102] It can be understood that the pre-constructed CLIP image-text retrieval model obtains image-text matching learning capabilities based on training of a large number of English image-text pairs, and the difference between Chinese image-text retrieval and English image-text retrieval lies in the difference in text information. Therefore, the embodiment of the present invention uses a Chinese text training set to train the pre-constructed Chinese encoder, with the aim of improving the Chinese learning and comprehension capabilities of the corresponding text encoder in the CLIP image-text retrieval model.
[0103] In an embodiment of the present invention, the trained Chinese encoder is used to replace the English encoder in the pre-built CLIP image-text retrieval model. On the one hand, the trained Chinese editor can be applied to Chinese text learning scenarios, and on the other hand, the image-text comparison learning capability in the existing CLIP image-text retrieval model can be directly inherited.
[0104] Step F, obtaining a Chinese image-text pair training set, using the Chinese image-text pair training set to perform image-text matching training on the Chinese image-text CLIP retrieval model, until the image-text matching training meets a second preset training condition, then exiting the image-text matching training to obtain a target Chinese image-text CLIP retrieval model.
[0105] In the embodiment of the present invention, the Chinese image-text pair refers to information composed of an image and a text in pairs, wherein the text is a Chinese text.
[0106] It is understandable that the current CLIP image-text retrieval model is obtained based on the training of 400 million English image-text pairs, and the 400 million English image-text pairs can be used to generate a Chinese image-text pair training set.
[0107] In detail, the method of obtaining the Chinese picture-text pair training set includes: obtaining an English picture-text pair set from a preset English picture-text pair library; and translating the English picture-text pair set into Chinese to obtain the Chinese picture-text pair training set.
[0108] In an embodiment of the present invention, the English image-text pair can be obtained from the existing CLIP image-text retrieval model training library, and the English image-text can be translated into Chinese using a translation tool (e.g., MT) to obtain the corresponding Chinese image-text pair. This operation can reduce the difficulty and cost of obtaining the Chinese image-text pair.
[0109] For details, see Figure 5 As shown, the Chinese image-text pair training set is used to perform image-text matching training on the Chinese image-text CLIP retrieval model until the image-text matching training meets a second preset training condition, then the image-text matching training is exited to obtain a target Chinese image-text CLIP retrieval model, including:
[0110] S61, using the text editor in the Chinese image-text CLIP retrieval model to perform vector conversion on the text information in the Chinese image-text pair training set to obtain a text vector set, and using the image editor in the Chinese image-text CLIP retrieval model to perform vector conversion on the image information in the Chinese image-text pair training set to obtain an image vector set;
[0111] S62. Utilizing the cross-modal contrast learning mechanism in the Chinese image-text CLIP retrieval model, matching the image with the text is performed according to the text vector set and the image vector set to obtain a predicted image-text matching result.
[0112] S63, calculating the loss value between the predicted image-text matching result and the real result corresponding to the Chinese image-text pair training set;
[0113] S64, determining whether the loss value satisfies the second preset training condition;
[0114] S65. If the loss value does not meet the second preset training condition, return to S61;
[0115] S66. If the loss value satisfies the second preset training condition, the image-text matching training is exited to obtain a target Chinese image-text CLIP retrieval model.
[0116] In the embodiment of the present invention, the Chinese picture-text CLIP retrieval model is trained by using the Chinese picture-text pair training set, so that the Chinese picture-text CLIP retrieval model has the learning ability of comparing Chinese and pictures and texts, and can be further applied to retrieval based on Chinese picture-text information.
[0117] The embodiment of the present invention uses a pre-constructed Chinese encoder to perform vector conversion on Chinese positive samples and Chinese negative samples, obtains a first similarity between positive sample vectors and a second similarity between the positive sample vectors and the negative sample vectors by calculation, uses the first similarity and the second similarity to adjust the parameters of the pre-constructed Chinese encoder and performs Chinese recognition training on the positive and negative sample vectors, so that the trained Chinese encoder has the ability to understand and learn Chinese, and then uses the trained Chinese encoder to replace the English encoder in the pre-constructed CLIP image-text retrieval model to obtain a Chinese image-text CLIP retrieval model, and then uses a Chinese image-text pair training set to train the Chinese image-text CLIP retrieval model to obtain a target Chinese image-text CLIP retrieval model with Chinese image-text pair retrieval capability.
[0118] like Figure 6 , which is a functional module diagram of a CLIP-based Chinese image and text retrieval model training device provided in one embodiment of the present invention.
[0119] The Chinese image-text retrieval model training device 100 based on CLIP of the present invention can be installed in an electronic device. According to the functions to be implemented, the Chinese image-text retrieval model training device 100 based on CLIP can include a Chinese text training set acquisition module 101, a Chinese encoder training module 102, a Chinese image-text model construction module 103 and a Chinese image-text model training module 104. The module of the present invention can also be referred to as a unit, which refers to a series of computer program segments that can be executed by an electronic device processor and can complete fixed functions, which are stored in the memory of the electronic device.
[0120] In this embodiment, the functions of each module / unit are as follows:
[0121] The Chinese text training set acquisition module 101 is used to acquire a Chinese text training set, randomly select one Chinese text from the Chinese text training set as a positive sample, and other Chinese texts as negative samples;
[0122] The Chinese encoder training module 102 is used to perform Chinese recognition training on the pre-constructed Chinese encoder to generate a preset number of positive sample vectors for the positive samples and generate a negative sample vector set for the negative samples; calculate a first similarity between the positive sample vectors, and calculate a second similarity between the positive sample vectors and the negative sample vector set; adjust the parameters of the pre-constructed Chinese encoder according to the first similarity and the second similarity until the pre-constructed Chinese encoder meets the first preset training condition, exit the Chinese recognition training, and obtain a Chinese encoder that has completed the training;
[0123] The Chinese picture-text model construction module 103 is used to replace the English encoder in the pre-constructed CLIP picture-text retrieval model with the trained Chinese encoder to obtain the Chinese picture-text CLIP retrieval model;
[0124] The Chinese image-text model training module 104 is used to obtain a Chinese image-text pair training set, and use the Chinese image-text pair training set to perform image-text matching training on the Chinese image-text CLIP retrieval model until the image-text matching training meets a second preset training condition, then exit the image-text matching training to obtain a target Chinese image-text CLIP retrieval model.
[0125] In detail, each module in the CLIP-based Chinese image-text retrieval model training device 100 in the embodiment of the present invention is used in the same manner as described above. Figures 1 to 5 The same technical means are used as the CLIP-based Chinese image and text retrieval model training method described in , and can produce the same technical effects, so I will not go into details here.
[0126] like Figure 7, which is a schematic diagram of the structure of an electronic device for implementing a Chinese image-text retrieval model training method based on CLIP provided by an embodiment of the present invention.
[0127] The electronic device 1 may include a processor 10, a memory 11 and a bus, and may also include a computer program stored in the memory 11 and executable on the processor 10, such as a text-based entity relationship extraction program.
[0128] The memory 11 includes at least one type of readable storage medium, and the readable storage medium includes a flash memory, a mobile hard disk, a multimedia card, a card-type memory (e.g., SD or DX memory, etc.), a magnetic memory, a disk, an optical disk, etc. The memory 11 may be an internal storage unit of the electronic device 1 in some embodiments, such as a mobile hard disk of the electronic device 1. The memory 11 may also be an external storage device of the electronic device 1 in other embodiments, such as a plug-in mobile hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (SecureDigital, SD) card, a flash card (Flash Card), etc. equipped on the electronic device 1. Further, the memory 11 may also include both an internal storage unit of the electronic device 1 and an external storage device. The memory 11 may not only be used to store application software and various types of data installed in the electronic device 1, such as the code of a text-based entity relationship extraction program, but also be used to temporarily store data that has been output or is to be output.
[0129] The processor 10 may be composed of an integrated circuit in some embodiments, for example, a single packaged integrated circuit, or a plurality of packaged integrated circuits with the same or different functions, including one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and combinations of various control chips, etc. The processor 10 is the control core (Control Unit) of the electronic device, and uses various interfaces and lines to connect various components of the entire electronic device, and executes or executes programs or modules (such as text-based entity relationship extraction programs, etc.) stored in the memory 11, and calls data stored in the memory 11 to execute various functions of the electronic device 1 and process data.
[0130] The bus may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. The bus is configured to realize connection and communication between the memory 11 and at least one processor 10, etc.
[0131] Figure 7 Only an electronic device with components is shown, and those skilled in the art will understand that Figure 7 The structure shown does not constitute a limitation on the electronic device 1, and may include fewer or more components than shown in the figure, or combine certain components, or arrange the components differently.
[0132] For example, although not shown, the electronic device 1 may also include a power source (such as a battery) for supplying power to each component. Preferably, the power source may be logically connected to the at least one processor 10 through a power management device, so that the power management device can realize functions such as charging management, discharging management, and power consumption management. The power source may also include any components such as one or more DC or AC power sources, recharging devices, power failure detection circuits, power converters or inverters, power status indicators, etc. The electronic device 1 may also include a variety of sensors, Bluetooth modules, Wi-Fi modules, etc., which will not be repeated here.
[0133] Furthermore, the electronic device 1 may also include a network interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), which is generally used to establish a communication connection between the electronic device 1 and other electronic devices.
[0134] Optionally, the electronic device 1 may further include a user interface, which may be a display, an input unit (such as a keyboard), or a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, and an OLED (Organic Light-Emitting Diode) touch device. The display may also be appropriately referred to as a display screen or a display unit, which is used to display information processed in the electronic device 1 and to display a visual user interface.
[0135] It should be understood that the embodiment is for illustration only and the scope of the patent application is not limited to this structure.
[0136] The text-based entity relationship extraction program stored in the memory 11 of the electronic device 1 is a combination of multiple instructions. When running in the processor 10, it can achieve:
[0137] Step A, obtaining a Chinese text training set, randomly selecting a Chinese text from the Chinese text training set as a positive sample, and other Chinese texts as negative samples;
[0138] Step B, performing Chinese recognition training on the pre-built Chinese encoder to generate a preset number of positive sample vectors of the positive samples and a set of negative sample vectors of the negative samples;
[0139] Step C, calculating a first similarity between the positive sample vectors, and calculating a second similarity between the positive sample vector and the negative sample vector set;
[0140] Step D, adjusting the parameters of the pre-constructed Chinese encoder according to the first similarity and the second similarity, and returning to the above step B, until the pre-constructed Chinese encoder meets the first preset training condition, exiting the Chinese recognition training, and obtaining a trained Chinese encoder;
[0141] Step E: using the trained Chinese encoder to replace the English encoder in the pre-built CLIP image-text retrieval model to obtain a Chinese image-text CLIP retrieval model;
[0142] Step F, obtaining a Chinese image-text pair training set, using the Chinese image-text pair training set to perform image-text matching training on the Chinese image-text CLIP retrieval model, until the image-text matching training meets a second preset training condition, then exiting the image-text matching training to obtain a target Chinese image-text CLIP retrieval model.
[0143] Furthermore, if the module / unit integrated in the electronic device 1 is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. The computer-readable storage medium can be volatile or non-volatile. For example, the computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disk, a computer memory, and a read-only memory (ROM).
[0144] The present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor of an electronic device, the computer program can implement:
[0145] Step A, obtaining a Chinese text training set, randomly selecting a Chinese text from the Chinese text training set as a positive sample, and other Chinese texts as negative samples;
[0146] Step B, performing Chinese recognition training on the pre-built Chinese encoder to generate a preset number of positive sample vectors of the positive samples and a set of negative sample vectors of the negative samples;
[0147] Step C, calculating a first similarity between the positive sample vectors, and calculating a second similarity between the positive sample vector and the negative sample vector set;
[0148] Step D, adjusting the parameters of the pre-constructed Chinese encoder according to the first similarity and the second similarity, and returning to the above step B, until the pre-constructed Chinese encoder meets the first preset training condition, exiting the Chinese recognition training, and obtaining a trained Chinese encoder;
[0149] Step E: using the trained Chinese encoder to replace the English encoder in the pre-built CLIP image-text retrieval model to obtain a Chinese image-text CLIP retrieval model;
[0150] Step F, obtaining a Chinese image-text pair training set, using the Chinese image-text pair training set to perform image-text matching training on the Chinese image-text CLIP retrieval model, until the image-text matching training meets a second preset training condition, then exiting the image-text matching training to obtain a target Chinese image-text CLIP retrieval model.
[0151] In addition, each functional module in each embodiment of the present invention may be integrated into one processing unit, each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of hardware plus software functional modules.
[0152] It is obvious to those skilled in the art that the present invention is not limited to the details of the above exemplary embodiments, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0153] Therefore, no matter from which point of view, the embodiments should be regarded as illustrative and non-restrictive, and the scope of the present invention is limited by the appended claims rather than the above description, so it is intended that all changes falling within the meaning and scope of the equivalent elements of the claims are included in the present invention. Any attached figure mark in the claims should not be regarded as limiting the claims involved.
[0154] The blockchain referred to in this invention is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, encryption algorithm, etc. Blockchain is essentially a decentralized database, a string of data blocks generated by cryptographic methods. Each data block contains a batch of network transaction information, which is used to verify the validity of its information (anti-counterfeiting) and generate the next block. Blockchain can include the underlying blockchain platform, platform product service layer, and application service layer.
[0155] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.
[0156] In addition, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices stated in the system claim can also be implemented by one unit or device through software or hardware. The second and other words are used to indicate names, but not to indicate any particular order.
[0157] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solution of the present invention can be modified or replaced by equivalents without departing from the spirit and scope of the technical solution of the present invention.
Claims
1. A Chinese image-text retrieval model training method based on CLIP, characterized in that: The method comprises: Step A, obtaining a Chinese text training set, randomly selecting a Chinese text from the Chinese text training set as a positive sample, and other Chinese texts as negative samples; Step B, performing Chinese recognition training on the pre-built Chinese encoder to generate a preset number of positive sample vectors of the positive samples and a set of negative sample vectors of the negative samples; Step C, calculating a first similarity between the positive sample vectors, and calculating a second similarity between the positive sample vector and the negative sample vector set; Step D, adjusting the parameters of the pre-constructed Chinese encoder according to the first similarity and the second similarity, and returning to the above step B, until the pre-constructed Chinese encoder meets the first preset training condition, exiting the Chinese recognition training, and obtaining a trained Chinese encoder; Step E: using the trained Chinese encoder to replace the English encoder in the pre-built CLIP image-text retrieval model to obtain a Chinese image-text CLIP retrieval model; Step F, obtaining a Chinese image-text pair training set, using the Chinese image-text pair training set to perform image-text matching training on the Chinese image-text CLIP retrieval model, until the image-text matching training meets a second preset training condition, then exiting the image-text matching training, and obtaining a target Chinese image-text CLIP retrieval model; The calculating of the second similarity between the positive sample vector and the negative sample vector set comprises: performing a clustering operation on the negative sample vector set until the clustering operation satisfies a preset clustering condition, exiting the clustering operation, and obtaining each cluster center after clustering; using a preset distance function, sequentially calculating the unilateral distance between each cluster center and each positive sample vector; Calculate the comprehensive distance between the negative sample vector set and the positive sample vector according to all the one-sided distances; invert the comprehensive distance to obtain a second similarity between the positive sample vector and the negative sample vector set; Before adjusting the parameters of the pre-constructed Chinese encoder according to the first similarity and the second similarity, the method also includes: normalizing the first similarity and the second similarity; calculating the harmonic value between the normalized first similarity and the normalized second similarity using a preset harmonic formula; when the harmonic value does not meet the preset harmonic value threshold, adjusting the parameters of the pre-constructed Chinese encoder.
2. The CLIP-based Chinese image-text retrieval model training method according to claim 1, characterized in that: The Chinese recognition training includes using a pre-built Chinese encoder to generate a preset number of positive sample vectors of the positive samples, and using the pre-built Chinese encoder to generate a preset number of positive sample vectors of the positive samples includes: Randomly generating a preset number of parameter values of preset parameters in the pre-built Chinese encoder; Selecting one parameter value from a preset number of parameter values in turn to assign a value to the preset parameter; Performing vector conversion on the positive sample using the pre-constructed Chinese encoder after the assignment to obtain a positive sample vector corresponding to the positive sample; All positive sample vectors corresponding to the positive samples are collected to obtain the preset number of positive sample vectors.
3. The CLIP-based Chinese image-text retrieval model training method according to claim 1, characterized in that: The step of calculating the harmonic value between the normalized first similarity and the normalized second similarity by using a preset harmonic formula includes: The harmonic value is calculated using the following preset harmonic formula: in, represents the harmonic value, represents the first similarity after normalization, represents the normalized second similarity, is the first harmonic coefficient, is the second harmonic coefficient, through the first harmonic coefficient and the second harmonic coefficient, so that when Increase and When decreasing, The value increases when Reduce and When increasing, The value decreases.
4. The CLIP-based Chinese image-text retrieval model training method according to claim 1, characterized in that: The step of performing image-text matching training on the Chinese image-text CLIP retrieval model using the Chinese image-text pair training set until the image-text matching training meets a second preset training condition, then exiting the image-text matching training to obtain a target Chinese image-text CLIP retrieval model includes: Using the text editor in the Chinese image-text CLIP retrieval model, the text information in the Chinese image-text pair training set is converted into vectors to obtain a text vector set, and using the image editor in the Chinese image-text CLIP retrieval model, the image information in the Chinese image-text pair training set is converted into vectors to obtain an image vector set; Using the cross-modal contrast learning mechanism in the Chinese image-text CLIP retrieval model, matching the image and the text is performed according to the text vector set and the image vector set to obtain a predicted image-text matching result; Calculating the loss value between the predicted image-text matching result and the actual result corresponding to the Chinese image-text pair training set; Determining whether the loss value satisfies the second preset training condition; If the loss value does not meet the second preset training condition, returning to the above-mentioned steps of using the text editor in the Chinese image-text CLIP retrieval model to perform vector conversion on the text information in the Chinese image-text pair training set to obtain a text vector set, and using the image editor in the Chinese image-text CLIP retrieval model to perform vector conversion on the image information in the Chinese image-text pair training set to obtain an image vector set; If the loss value satisfies the second preset training condition, the image-text matching training is exited to obtain the target Chinese image-text CLIP retrieval model.
5. The CLIP-based Chinese image-text retrieval model training method according to claim 1, characterized in that: The step of obtaining a Chinese image-text pair training set includes: Obtain an English picture-text pair set from a preset English picture-text pair library; The English picture-text pair set is translated into Chinese to obtain a Chinese picture-text pair training set.
6. A CLIP-based Chinese image-text retrieval model training device, used to implement the CLIP-based Chinese image-text retrieval model training method as claimed in any one of claims 1 to 5, characterized in that: The device comprises: A Chinese text training set acquisition module is used to acquire a Chinese text training set, randomly select one Chinese text from the Chinese text training set as a positive sample, and other Chinese texts as negative samples; A Chinese encoder training module, used for performing Chinese recognition training on a pre-built Chinese encoder to generate a preset number of positive sample vectors of the positive samples and a set of negative sample vectors of the negative samples; Calculating a first similarity between the positive sample vectors, and calculating a second similarity between the positive sample vector and the negative sample vector set; adjusting the parameters of the pre-constructed Chinese encoder according to the first similarity and the second similarity until the pre-constructed Chinese encoder meets a first preset training condition, exiting the Chinese recognition training, and obtaining a trained Chinese encoder; A Chinese picture-text model construction module is used to replace the English encoder in the pre-constructed CLIP picture-text retrieval model with the trained Chinese encoder to obtain a Chinese picture-text CLIP retrieval model; The Chinese picture-text model training module is used to obtain a Chinese picture-text pair training set, and use the Chinese picture-text pair training set to perform picture-text matching training on the Chinese picture-text CLIP retrieval model until the picture-text matching training meets a second preset training condition, then exit the picture-text matching training to obtain a target Chinese picture-text CLIP retrieval model.
7. An electronic device, characterized in that: The electronic device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and instructions are executed by the at least one processor so that the at least one processor can execute the CLIP-based Chinese image and text retrieval model training method as described in any one of claims 1 to 5.
8. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the CLIP-based Chinese image-text retrieval model training method as described in any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Training method and training device of multi-modal pre-training model and electronic equipment
CN113283551A
Training method of image search model and image search method
CN113590865A