Image-text matching method and electronic equipment
By using knowledge distillation technology of student graphic and text models and teacher graphic and text models in miniaturized electronic devices, the problem of low graphic and text matching accuracy caused by limited equipment processing resources is solved, and higher graphic and text matching accuracy and model performance are achieved.
Patent Information
- Application Number
- CN202311563842.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-21
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2043-11-21
AI Technical Summary
The processing resources of existing miniaturized electronic devices are limited, resulting in very low accuracy of graphics and text matching processing through neural network models.
A graphic matching method is adopted to update the parameters of the student graphic and text model through knowledge distillation between the student graphic and text model and the teacher graphic and text model, thereby improving the accuracy of graphic and text matching. The specific steps include: input the graphic and text pairs into the graphic and text models of students and teachers, fuse the images and text vectors, perform knowledge distillation, and update the model parameters based on the distillation loss.
Through knowledge distillation technology, students learn more complex knowledge of teachers' graphic and text models, thereby improving the accuracy of graphic and text matching and enhancing the performance of the model under limited resources.
Smart Images

Figure CN120067238A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of neural networks, and in particular, to an image-text matching method and an electronic device. Background Art
[0002] Image-text matching (ITM) processing is a basic vision-language (VL) processing. Image-text matching processing can be used to search for the corresponding image by text, or search for the corresponding text by image. Currently, the processing resources of miniaturized electronic devices are limited, and the scale of the running neural network model is small, resulting in very low accuracy of image-text matching processing through the neural network model. Summary of the Invention
[0003] Embodiments of this application provide an image-text matching method and an electronic device, which are used to improve the accuracy of image-text matching processing by a neural network model.
[0004] To achieve the above object, the embodiments of this application adopt the following technical solutions:
[0005] In a first aspect, an image-text matching method is provided. The method includes: displaying an interface, where the interface includes a control for obtaining a search target; in response to an operation on the control, matching the search target according to a student image-text model to obtain a search result; the search target is text and the search result is visual media, or the search target is visual media and the search result is text; displaying the search result; where the training process of the student image-text model includes: inputting an image in a plurality of image-text pairs into a student image encoder in the student image-text model to obtain a student image vector; inputting the image into a teacher image encoder in a teacher image-text model to obtain a teacher image vector; inputting text in a plurality of image-text pairs into a student text encoder in the student image-text model to obtain a student text vector; inputting the text into a teacher text encoder in the teacher image-text model to obtain a teacher text vector; fusing the student image vector and the student text vector to obtain student image-text information; fusing the teacher image vector and the teacher text vector to obtain teacher image-text information; performing knowledge distillation on the student image-text information using the teacher image-text information, and updating the parameters of the student image-text model according to the distillation loss after knowledge distillation.
[0006] The graphic-text matching method provided by the embodiments of this application inputs a graphic-text pair into a student graphic-text model to obtain a student image vector and a student text vector, and inputs the same graphic-text pair into a teacher graphic-text model to obtain a teacher image vector and a teacher text vector. The student image vector and the student text vector are fused to obtain student graphic-text information; the teacher image vector and the teacher text vector are fused to obtain teacher graphic-text information. On the one hand, the teacher graphic-text information comes from a more complex teacher graphic-text model, and compared with the student graphic-text information from a simple student graphic-text model, it has higher-level graphic-text semantic information. On the other hand, the fused graphic-text information has higher-level graphic-text semantic information compared with pure image information or text information, and has higher knowledge expression ability when used for knowledge distillation. Then, the teacher graphic-text information is used to perform knowledge distillation on the student graphic-text information, and the parameters of the student graphic-text model are updated according to the distillation loss after knowledge distillation, and the updated student graphic-text model is used for graphic-text matching processing. In this way, the updated student graphic-text model learns the knowledge of the complex teacher graphic-text model, and can improve the accuracy of the neural network model for graphic-text matching processing.
[0007] In a possible implementation manner, the student image vector and the student text vector are fused to obtain student graphic-text information; the teacher image vector and the teacher text vector are fused to obtain teacher graphic-text information, including: multiplying the student image vector and the student text vector matrix to obtain a student graphic-text information matrix; multiplying the teacher image vector and the teacher text vector matrix to obtain a teacher graphic-text information matrix. The fused graphic-text information has higher-level graphic-text semantic information compared with pure image information or text information, and has higher knowledge expression ability when used for knowledge distillation.
[0008] In a possible implementation manner, the teacher graphic-text information is used to perform knowledge distillation on the student graphic-text information, and the parameters of the student graphic-text model are updated according to the distillation loss after knowledge distillation, including: obtaining an image distillation loss and a text distillation loss according to the student graphic-text information matrix and the teacher graphic-text information matrix; the image distillation loss is used to represent the loss of the student image vector relative to the teacher image vector, and the text distillation loss is used to represent the loss of the student text vector relative to the teacher text vector; obtaining a distillation loss according to the image distillation loss and the text distillation loss; obtaining a total loss according to the distillation loss and the training loss; the training loss is used to represent the loss between the model output result and the true value after training the student graphic-text model with the graphic-text pair; updating the parameters of the student graphic-text model according to the total loss. The parameters of the student graphic-text model can be updated with the knowledge from the teacher graphic-text model.
[0009] In a possible implementation manner, obtaining an image distillation loss and a text distillation loss according to the student graphic-text information matrix and the teacher graphic-text information matrix includes: obtaining the image distillation loss Distill according to the following two formulas respectivelyimage and the text distillation loss Distill text :
[0010]
[0011]
[0012] Among them, the matrix S' V represents the transpose of the student image-text information matrix S V The matrix T' V represents the transpose of the teacher image-text information matrix T V The softmax() represents the exponential normalization function, and log_softmax() represents taking the log of the result of softmax(). S V (i) represents the i-th row feature in the student image-text information matrix S V The matrix T V (i) represents the i-th row feature in the teacher image-text information matrix T V The matrix S' V (i) represents the i-th row feature in the matrix S' V The matrix T' V (i) represents the i-th row feature in the matrix T' V The number of image-text pairs is represented by N.
[0013] The softmax() function has the characteristic of normalization. The larger the value, the closer the value obtained after normalization is to 1, and the smaller the value, the closer the value obtained after normalization is to 0. Using the softmax() function to fuse the teacher image vector and the teacher text vector in the teacher image-text model can significantly strengthen the knowledge on the diagonal of the teacher image-text information matrix and enhance the distillation value of the knowledge. The reason is that the values on the diagonal of the teacher image-text information matrix are much larger than those on the non-diagonal, that is, the knowledge on the diagonal of the teacher image-text information matrix is more and has higher value.
[0014] The log_softmax() function can map the relatively large values obtained by the softmax() function to the same level, effectively avoiding the overflow problem caused by some features in the student image-text information matrix having too large values, and ensuring the stability of the student image-text model in receiving knowledge.
[0015] In a possible implementation manner, the distillation loss is obtained according to the image distillation loss and the text distillation loss, including: performing weighted summation on the image distillation loss and the text distillation loss to obtain the distillation loss. The loss of knowledge distillation includes the loss of the image after knowledge distillation and the loss of the text.
[0016] In a possible implementation manner, obtaining a total loss according to a distillation loss and a training loss includes: performing a weighted sum on the distillation loss and the training loss to obtain the total loss. The total loss includes the loss of knowledge distillation and the loss of model training.
[0017] In a possible implementation manner, updating the parameters of the student image-text model according to the total loss includes: using the gradient descent method to iteratively update the parameters of the student image-text model until the total loss converges to the lowest and a preset number of iterations is reached.
[0018] In a possible implementation manner, the teacher image-text model is superior to the student image-text model in at least one of the following: the number of model layers, the number of hidden neurons, the size of the multi-layer perceptron, and the number of multi-heads. The teacher image-text model is more complex and has a higher level of image-text semantic information compared to that from the simple student image-text model. By learning the more complex teacher image-text model, the performance of the student image-text model in performing image-text matching can be enhanced.
[0019] In a second aspect, an electronic device is provided, including a processor and a memory. Instructions are stored in the memory, and when the processor executes the instructions, the method as described in the first aspect and any of its implementation manners is executed.
[0020] In a third aspect, a computer-readable storage medium is provided, including instructions, and when the instructions run on an electronic device, the electronic device is caused to execute the method as described in the first aspect and any of its implementation manners.
[0021] In a fourth aspect, a computer program product including instructions is provided, and when the instructions run on the above-mentioned electronic device, the electronic device is caused to execute the method as described in the first aspect and any of its implementation manners.
[0022] In a fifth aspect, a chip system is provided. The chip system includes a processor for supporting the electronic device to implement the functions involved in the first aspect above. In a possible design, the device further includes an interface circuit, and the interface circuit can be used to receive signals from other devices (such as a memory), or send signals to other devices (such as a communication interface). The chip system may include a chip and may also include other discrete devices.
[0023] The technical effects of the second aspect to the fifth aspect refer to the technical effects of the first aspect and any of its implementation manners, and will not be repeated here. Description of the Drawings
[0024] Figure 1 A schematic diagram of image-text matching provided by an embodiment of the present application;
[0025] Figure 2 A schematic diagram of the interface for performing text-to-image search in the gallery application of a mobile phone provided by an embodiment of the present application;
[0026] Figure 3 It is a schematic diagram of an interface for image search in text in the intelligent recognition application of a mobile phone provided by an embodiment of the present application;
[0027] Figure 4 It is a schematic diagram of a graphic and text matching system provided by an embodiment of the present application;
[0028] Figure 5 It is a schematic diagram of the structure of a server provided by an embodiment of the present application;
[0029] Figure 6 It is a schematic diagram of the structure of an electronic device provided by an embodiment of the present application;
[0030] Figure 7 It is a schematic diagram of the software architecture of an electronic device provided by an embodiment of the present application;
[0031] Figure 8 It is a schematic diagram of the process of a graphic and text matching method provided by an embodiment of the present application;
[0032] Figure 9 It is a schematic diagram of a student graphic and text information matrix provided by an embodiment of the present application;
[0033] Figure 10 It is a schematic diagram of a teacher graphic and text information matrix provided by an embodiment of the present application;
[0034] Figure 11 It is a schematic diagram of the process of another graphic and text matching method provided by an embodiment of the present application;
[0035] Figure 12 It is a schematic diagram of a softmax() function provided by an embodiment of the present application;
[0036] Figure 13 It is a schematic diagram of a log_softmax() function provided by an embodiment of the present application;
[0037] Figure 14 It is a schematic diagram of the principle of the gradient descent method provided by an embodiment of the present application;
[0038] Figure 15 It is a schematic diagram of the process of yet another graphic and text matching method provided by an embodiment of the present application;
[0039] Figure 16 It is a schematic diagram of the structure of a chip system provided by an embodiment of the present application. Detailed implementation manners
[0040] First, some concepts related to the present application are described.
[0041] The terms "first", "second", etc. involved in the embodiments of the present application are only used for the purpose of distinguishing the same type of features, and cannot be understood as indicating relative importance, quantity, order, etc.
[0042] The terms "exemplary" or "for example" etc. involved in the embodiments of the present application are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the present application should not be construed as being more preferred or more advantageous than other embodiments or design solutions. Rather, the use of the terms "exemplary" or "for example" etc. is intended to present the relevant concepts in a specific manner.
[0043] The terms "coupled" and "connected" involved in the embodiments of the present application should be understood in a broad sense. For example, it can refer to a direct physical connection, or an indirect connection implemented through electronic devices, such as a connection implemented through resistors, inductors, capacitors or other electronic devices.
[0044] Text vector: The text can be input into a text encoder to obtain a text vector, which is used to characterize the semantic features of the entire text. The text encoder can adopt models such as transformers commonly used in natural language processing (NLP), and the present application does not limit this here.
[0045] Image vector: The image frame of a picture or video can be input into an image encoder to obtain a visual vector, which is used to characterize the semantic features of the picture or image frame. The image encoder can adopt commonly used convolutional neural network (CNN) models or vision transformer (VIT) models, etc., and this solution is not limited here. For the convenience of description in the present application, the image frames of pictures and videos are collectively referred to as images, and pictures and videos are collectively referred to as visual media. In particular, the image frame of a video can be a thumbnail of the video.
[0046] Training of the neural network model (specifically referring to the text-image model in the present application, abbreviated as the model): Each set of training data in the training data set of the neural network model includes the input of the neural network model and the corresponding label, and the label represents the true output value corresponding to the input data. The training process of the neural network model is to continuously update the parameters of the neural network model, so that after the neural network model inputs the input data of each set of training data, its output result gradually fits the label in this set of training data. The quality of the fitting is usually evaluated by a loss function. Among them, the value of the loss function is used to measure the gap between the output of the neural network model and the true output, that is, the smaller the value of the loss function, the better the fitting of the neural network model is considered.
[0047] Knowledge distillation is a teacher-student training framework that adopts the method of transfer learning. By using the output of a pre-trained complex teacher model (teacher model in this application is the teacher image-text model) as a supervision signal, another relatively simple student model (student model in this application is the student image-text model) is trained, effectively learning the student model from the teacher model, and transferring the knowledge of the complex teacher model to the relatively simple student model at the cost of a slight performance loss. The key of knowledge distillation lies in how to extract rich knowledge from the teacher model and transfer the knowledge from the teacher model to guide the training of the student model. Knowledge distillation can be used for model compression and model enhancement. Model compression means guiding the training of the student model on the same labeled dataset according to the teacher model to obtain a simple and efficient student model. Model enhancement uses other resources (such as unlabeled, cross-domain or cross-modal data) or optimization strategies (such as mutual learning and self-learning) to improve the accuracy of the student model. In this application, model enhancement is used to improve the performance of the contrastive language-image pre-training (CLIP) model in the field of image-text matching (ITM) applications (especially Chinese image-text matching).
[0048] Vision-Language (VL) is a field formed by the intersection of two research areas, computer vision (CV) and natural language processing (NLP), aiming to endow artificial intelligence (AI) systems with the ability to learn effective information from multi-modal (e.g., visual and language modalities) data. Inspired by NLP pre-trained language models (such as bidirectional encoder representation from transformer (BERT), generative pre-trained transformer (GPT), etc.), vision-language pre-training (VLP) has gradually attracted attention and become the core training method for VL processing. This method first conducts pre-training through self-supervised learning, using auxiliary tasks to automatically mine supervision signals from large-scale unlabeled data to train the model, thus obtaining a general training model. Then, in specific applications, a small amount of manually labeled data is used to fine-tune the training model, and good results can be achieved.
[0049] As a basic VL processing, image-text matching (ITM) processing has gradually received extensive attention. Image-text matching processing mainly includes two sub-tasks - image-to-text retrieval (TR) and text-to-image retrieval (IR). Image-to-text retrieval means searching for the corresponding text through an image (image frame of a picture or video). For the image frame of a video, it is equivalent to searching for the corresponding text through the video including the image frame. Therefore, image-to-text retrieval can also refer to searching for the corresponding text through visual media (i.e., including pictures and videos). For example, searching for a corresponding text "puppy" through a thumbnail of a picture or video including a puppy. Text-to-image retrieval means searching for the corresponding image through text. For the image frame of a video, it is equivalent to searching for the corresponding video including the image frame through text. Therefore, text-to-image retrieval can also refer to searching for the corresponding visual media (i.e., including pictures and videos) through text. For example, searching for a picture or video including a puppy through a text "puppy". Image-text matching processing can bridge the semantic gap between the visual and language modalities and achieve accurate semantic alignment between multiple modalities, that is, establish a cross-modal semantic consistency relationship between the image modality and the text modality.
[0050] With the rise of VLP, thanks to the training with large-scale unlabeled data, the obtained training model far outperforms the model without pre-training in image-text matching processing. A representative one is the CLIP model, which mainly includes an image encoder and a text encoder. The image encoder of the CLIP model is a visual transformer (ViT) or a residual network (ResNet), and the text encoder is the aforementioned BERT. In the pre-training stage of the CLIP model, image-text contrastive learning is adopted, and the optimization objectives include maximizing the cosine similarity score of image-text pairs and minimizing the score of unmatched image-text pairs. Through pre-training with a large amount of data, excellent zero-shot effects in multiple image-text matching processes are achieved. Zero-shot learning means training the model with a training set so that the model can classify a test set, where there is no intersection between the training set and the test set. In this application, an image-text pair refers to a data pair formed by an associated image and a piece of text (in the form of <image, text>). For example, an image including a puppy and a piece of text "puppy" can form an image-text pair. For a video, image frames (such as thumbnails) need to be extracted from the video as the images in the image-text pair. Essentially, only the image frames of the video participate in the training, rather than the entire video. Therefore, in this application, an image represents the image frames of a picture or a video.
[0051] The application of the CLIP model in image-text matching is as Figure 1 shown. When the CLIP model 11 is used for image-to-text search, the user inputs an image, and the image encoder 111 in the CLIP model extracts an image vector from the image and searches for the corresponding text vector in the image-text vector database, so as to feedback the corresponding text to the user. When the CLIP model is used for text-to-image search, the user inputs text to the CLIP model, and the text encoder 112 in the CLIP model extracts a text vector and searches for the corresponding image vector in the image-text vector database, so as to feedback the corresponding image to the user.
[0052] The following is combined with Figure 2 to illustrate the interface involved in text-to-image search in the gallery application of a mobile phone:
[0053] As Figure 2 shown in (a) of it, the mobile phone can display the main interface 201, which can also be called the desktop. The main interface 201 may include an icon 202 of the gallery application. The mobile phone can receive the operation of the user clicking the icon 202, and in response to this operation, the mobile phone can start the gallery application and display as Figure 2The interface 203 shown in (b) therein. Among them, the interface 203 can be an album interface. It should be noted that in response to the user's operation of clicking on the icon 202, the mobile phone can launch the gallery application and display the picture interface of the gallery. The picture interface includes thumbnails of the pictures in the gallery or the large picture of a certain picture. In the picture interface, in response to the user's operation on the "Album" control, the above-mentioned album interface 203 is displayed.
[0054] As Figure 2 shown in (b) therein, the interface 203 includes multiple albums. Among them, the "All Pictures" album includes 2,023 pictures, the "Camera" album includes 1,502 pictures and videos, the "Screenshots and Screen Recordings" album includes 102 pictures and videos, the "My Favorites" album includes 48 pictures and videos, the "One Recording, Multiple Gains" album has 34 pictures and videos, the "Video Editing" album has 65 videos, the "Self-created Album" has 57 pictures and videos, and the "Shared Album" has 100 pictures and videos. The interface 203 may include a search box 204.
[0055] As Figure 2 shown in (c) therein, in response to the user's operation of entering the text "sky" in the search box 204, the mobile phone can display the interface 205, and this interface 205 can be called a search interface. The interface 205 can display 100 visual media related to "sky", and this visual media can be pictures or videos (the image frames in the videos include sky elements), 32 pictures or videos containing the text "sky". Among them, the 100 pictures related to "sky" can be searched because these 100 pictures or images (collectively called visual media) are matched with the text "sky" through a text-image model (such as the CLIP model mentioned above); the 32 pictures related to the pictures containing the word "sky" can be searched because the optical character recognition (OCR) technology is used to recognize that these 32 pictures contain the word "sky". In practical applications, the mobile phone can also associate with the text input by the user to obtain associated words and perform searches based on the associated words.
[0056] Next, in combination with Figure 3 the interfaces involved in image search for text in the intelligent recognition application of the mobile phone will be described:
[0057] As Figure 3 shown in (a) therein, the mobile phone can display the search interface 301 on the negative first screen. The search interface 301 can include the intelligent recognition icon 302. The mobile phone can receive the user's operation of clicking on the intelligent recognition icon 302. In response to this operation, the mobile phone can launch the intelligent recognition application, turn on the camera, and display as Figure 3 shown in (b) therein, the shooting interface 303.
[0058] As shown Figure 3 in (b) of Figure 3 , the interface 303 includes an image 304 (such as an image of a car) captured in real time by a camera, a take photo button 305, and an album button 306. The mobile phone can receive a click operation of the user on the take photo button 305, and in response to this operation, the intelligent recognition application searches for the text corresponding to the image 304 through a text-image model (such as the CLIP model described above). Alternatively, the mobile phone can receive an operation of the user clicking the album button 306, and in response to this operation, the intelligent recognition application searches for the text corresponding to the image in the album through the text-image model.
[0059] As shown Figure 3 in (c) of Figure 3 , the text 308 ("car") corresponding to the image 304 is displayed in the interface 307, and a "Search for more" button 309 can also be displayed. The mobile phone can receive an operation of the user clicking the "Search for more" button 309, and in response to this operation, the mobile phone can call a search engine to search for more relevant content about "car".
[0060] Currently, the processing resources of miniaturized electronic devices (such as mobile phones, tablet computers, and laptop computers) are limited, and the scale of the running text-image model is small, resulting in very low accuracy of text-image matching processing (such as text-to-image search and image-to-text search) through the text-image model.
[0061] For this reason, as shown Figure 4 in Figure 4 , an embodiment of the present application provides a text-image matching system, which is used to execute the text-image matching method of the present application. The system mainly includes: a server 21 and an electronic device 22. The server 21 executes S101 - S104 to implement high-quality data screening, knowledge distillation and training of the student text-image model, and deploys the trained student text-image model on the electronic device 22. The electronic device 22 executes S105 to implement text-image matching processing according to the student text-image model.
[0062] Among them, in S101, the server 21 performs high-quality data screening on a large amount of open-source graphic and text data sets, so as to screen out multiple high-quality graphic and text pairs. It should be noted that for the image frames where the image is a video, the image frames in the video need to be extracted first for training. In S102, the server 21 trains the above-mentioned multiple graphic and text pairs with a student graphic and text model (which can be the CLIP model described above). In S103, the server 21 uses a complex teacher graphic and text model (which can be the CLIP model described above) to perform knowledge distillation on the simple student graphic and text model, and uses the distillation loss to update the parameters of the student graphic and text model, so that the student graphic and text model can learn the knowledge of the teacher graphic and text model, thereby improving the accuracy of the student graphic and text model in performing graphic and text matching processing. In S104, the server 21 deploys the trained student graphic and text model in the graphic and text matching application of the electronic device. In S105, the graphic and text matching application in the electronic device 22 performs graphic and text matching processing based on the student graphic and text model.
[0063] As Figure 5 shown, the server 21 may include at least one processor 211, a communication line 212, a memory 213, and at least one communication interface 214. The communication line 212 may include a path for transmitting information between the above components. The communication interface 214 uses any transceiver-like device for communicating with other devices. Instructions are stored in the memory 213, and when the processor 211 executes the instructions, the graphic and text matching method involved in the embodiments of the present application is executed.
[0064] The processor involved in the embodiments of the present application may be a chip. For example, it may be a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on chip (SoC), a central processing unit (CPU), a network processor (NP), a digital signal processing circuit (DSP), a micro controller unit (MCU), a programmable logic device (PLD), or other integrated chips.
[0065] The memory involved in the embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM). It should be noted that the memories of the systems and methods described herein are intended to include but not be limited to these and any other suitable types of memories.
[0066] The electronic device is an electronic device with a display screen. The electronic device can be mobile or fixed. The electronic device can be deployed on land (such as indoors or outdoors, handheld or vehicle-mounted, etc.), on water (such as a ship, etc.), or in the air (such as an airplane, balloon, and satellite, etc.). This electronic device can be referred to as a user equipment (UE), access terminal, terminal unit, subscriber unit, terminal station, mobile station (MS), mobile platform, terminal agent, or terminal device, etc. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, smart bracelet, smart screen, smart watch, virtual reality (VR) device, augmented reality (AR) device, terminal in industrial control, terminal in self-driving, terminal in remote medical, terminal in smart grid, terminal in transportation safety, terminal in smart city, terminal in smart home, etc. The embodiments of this application do not limit the specific type and structure of the electronic device. A possible structure of the electronic device will be described below.
[0067] Taking the electronic device as a mobile phone as an example, Figure 6 A possible structure of the electronic device 22 is shown. The electronic device 22 may include a processor 210, an external memory interface 220, an internal memory 221, a universal serial bus (USB) interface 230, a power management module 240, a battery 241, a wireless charging coil 242, antenna 1, antenna 2, a mobile communication module 250, a wireless communication module 260, an audio module 270, a speaker 270A, a receiver 270B, a microphone 270C, a headphone interface 270D, a sensor module 280, a key 290, a motor 291, an indicator 292, a camera 293, a display screen 294, and a subscriber identification module (SIM) card interface 295, etc. Optionally, in some embodiments, an audio digital signal processor (ADSP) 243 is further included.
[0068] It can be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the electronic device 22. In other embodiments of the present application, the electronic device 22 may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0069] The processor 210 may include one or more processing units. For example, the processor 210 may be a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on chip (SoC), a central processing unit (CPU), an application processor (AP), a network processor (NP), a digital signal processor (DSP), a microcontroller unit (MCU), a programmable logic device (PLD), a modulation and demodulation processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a baseband processor, and a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors. For example, the processor 210 may be an application processor AP. Or, the above-mentioned processor 210 may be integrated in a system on chip (SoC). Or, the above-mentioned processor 210 may be integrated in an integrated circuit (IC) chip. The processor 210 may include an analog front end (AFE) and a microcontroller unit (MCU) in the IC chip.
[0070] The processor 210 executes the display control method provided in the embodiments of the present application by executing programs and computer instructions stored in the internal memory 221.
[0071] A memory may also be provided in the processor 210 for storing computer instructions and data. In some embodiments, the memory in the processor 210 is a cache memory. This memory can store the computer instructions or data that the processor 210 has just used or recycled. If the processor 210 needs to use the computer instructions or data again, it can directly call them from the memory. This avoids repeated accesses and reduces the waiting time of the processor 210, thus improving the efficiency of the system.
[0072] In some embodiments, the processor 210 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a USB interface, etc.
[0073] The ADSP 243 can be coupled to the audio module 270 and the sensor module 280. The ADSP 243 can be used to process audio signals and also process sensor data. When the processor is in the sleep state, the ADSP 243 can still remain operational, thereby reducing the power consumption of the electronic device.
[0074] It can be understood that the interface connection relationships between the modules illustrated in the embodiments of the present application are only illustrative and do not constitute a structural limitation on the electronic device 22. In other embodiments of the present application, the electronic device 22 may also adopt different interface connection methods or a combination of multiple interface connection methods as described in the above embodiments.
[0075] The wireless communication function of the electronic device 22 can be implemented through antenna 1, antenna 2, the mobile communication module 250, the wireless communication module 260, the modulation and demodulation processor, and the baseband processor, etc.
[0076] Antenna 1 and Antenna 2 are used for transmitting and receiving electromagnetic wave signals. Each antenna in the electronic device 22 can be used to cover a single or multiple communication frequency bands. Different antennas can also be multiplexed to improve the utilization rate of the antennas. For example, Antenna 1 can be multiplexed as a diversity antenna for a wireless local area network. In some other embodiments, the antenna can be used in combination with a tuning switch.
[0077] The mobile communication module 250 can provide solutions for wireless communications including 2G / 3G / 4G / 5G, etc. applied to the electronic device 22. The wireless communication module 260 can provide solutions for wireless communications including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared technology (IR), etc. applied to the electronic device 22. In some embodiments, Antenna 1 of the electronic device 22 is coupled to the mobile communication module 250, and Antenna 2 is coupled to the wireless communication module 260, so that the electronic device 22 can communicate with the network and other devices through wireless communication technologies.
[0078] The external memory interface 220 can be used to connect an external memory card, such as a micro SanDisk (Micro SD) card, to implement the storage capacity expansion of the electronic device 22. The external memory card communicates with the processor 210 through the external memory interface 220 to implement the data storage function. For example, files such as music and videos are saved in the external memory card.
[0079] The internal memory 221 can be used to store computer-executable program codes, and the executable program codes include computer instructions. The processor 210 executes various functional applications and data processing of the electronic device 22 by running the computer instructions stored in the internal memory 221. In addition, the internal memory 221 can include high-speed random access memory, and can also include non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc.
[0080] The memory involved in the embodiments of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM). It should be noted that the memories of the systems and methods described herein are intended to include, but are not limited to, these and any other suitable types of memories.
[0081] The electronic device 22 can implement audio functions through the audio module 270, the speaker 270A, the receiver 270B, the microphone 270C, the headphone interface 270D, and the application processor, etc. Such as music playback, recording, etc.
[0082] The audio module 270 is used to convert digital audio information into an analog audio signal for output, and is also used to convert an analog audio input into a digital audio signal. In some embodiments, the audio module 270 may be disposed in the processor 210, or some functional modules of the audio module 270 may be disposed in the processor 210. The speaker 270A, also referred to as a "loudspeaker", is used to convert an audio electrical signal into a sound signal. The receiver 270B, also referred to as an "earpiece", is used to convert an audio electrical signal into a sound signal. The microphone 270C, also referred to as a "microphone" or "transmitter", is used to convert a sound signal into an electrical signal. The electronic device 22 may be provided with at least one microphone 270C. The headphone jack 270D is used to connect a wired headphone. The headphone jack 270D may be a USB interface 230, or may be a 3.5 mm open mobile terminal platform (OMTP) standard interface or a cellular telecommunications industry association of the USA (CTIA) standard interface.
[0083] The keys 290 include a power-on key, volume keys, etc. The keys 290 may be mechanical keys or may be touch keys. The electronic device 22 can receive key inputs and generate key signal inputs related to the user settings and function controls of the electronic device 22. The motor 291 can generate a vibration prompt. The motor 291 can be used for incoming call vibration prompts and can also be used for touch vibration feedback. The indicator 292 may be an indicator light and can be used to indicate the charging state, battery level change, and can also be used to indicate messages, missed calls, notifications, etc. The SIM card interface 295 is used to connect a SIM card. The SIM card can be inserted into or removed from the SIM card interface 295 to achieve contact and separation from the electronic device 22. The electronic device 22 may support 1 or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 295 can support a Nano SIM card, a Micro SIM card, a SIM card, etc. In some embodiments, the electronic device 22 uses an embedded SIM (eSIM) card, and the eSIM card can be embedded in the electronic device 22 and cannot be separated from the electronic device 22.
[0084] The electronic device 22 can implement a shooting function through an ISP, a camera 293, a video codec, a GPU, a display screen 294, and an application processor, etc. The ISP is used to process the data fed back by the camera 293. In some embodiments, the ISP may be disposed in the camera 293. The camera 293 is used to capture still images or videos. In some embodiments, the electronic device 22 may include 1 or N cameras 293, where N is a positive integer greater than 1.
[0085] The electronic device 22 can implement the display function through a GPU, a display screen 294, an application processor, etc. The GPU is a microprocessor for image processing, connected to the display screen 294 and the application processor. The GPU is used to execute mathematical and geometric calculations for graphics rendering. The processor 210 may include one or more GPUs, which execute computer instructions to generate or change display information.
[0086] The sensor module 280 may include a pressure sensor, a gyroscope sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a distance sensor, a proximity light sensor, a fingerprint sensor, a temperature sensor, a touch sensor, an ambient light sensor, an angle sensor, etc. When the display screen 294 is a foldable screen, the angle sensor can detect the folding angle of the display screen 294, and the range of the folding angle is 0 - 180 degrees.
[0087] The battery 241 may include one or more batteries to supply power to the load.
[0088] The power management module 240 is used to receive a charging input from a charger. Among them, the charger can be a wireless charger, such as a wireless charging dock, other electronic devices 22 with reverse wireless charging function, etc. The power management module 240 can receive the wireless charging input through the wireless charging coil 242 of the electronic device. The charger can also be a wired charger. For example, the power management module 240 can receive the charging input of the wired charger through the USB interface 230. The power management module 240 is also called a charging chip.
[0089] Among them, while the power management module 240 charges the battery 241, it can also supply power to the electronic device. The power management module 240 receives the input of the battery 241 and supplies power to the processor 210, the internal memory 221, the external memory interface 220, the display screen 294, the camera 293, the wireless communication module 260, etc. The power management module 240 can also be used to monitor parameters such as the capacity, voltage, number of battery cycles, and battery health status (leakage, impedance) of the battery 241. In some other embodiments, the power management module 240 can also be disposed in the processor 210.
[0090] The display screen 294 is used to display images, videos, etc. The display screen 294 includes a display panel. In some embodiments, the electronic device 22 may include one or more display screens 294.
[0091] Figure 7It is a software structure block diagram of the electronic device 22 according to an embodiment of the present application. The software system of the electronic device 22 may adopt a layered architecture. The layered architecture divides the software into several layers, and each layer has a clear role and division of labor. The layers communicate with each other through software interfaces. Taking the Android system as an example, in some embodiments, the Android system can be divided into four layers, from top to bottom, namely the application layer, the application framework layer, the Android runtime and the system libraries, and the kernel layer.
[0092] As Figure 7 shown, the application layer may include application programs such as a picture search text application and a text search picture application, and may perform picture-text matching processing by using the picture-text model described above.
[0093] The application framework layer provides application programming interfaces (APIs) and programming frameworks for the application programs in the application layer. The application framework layer includes some predefined functions.
[0094] The system libraries may include multiple functional modules. For example: surface manager, media libraries, 3D graphics processing libraries (such as: OpenGL for embedded systems (OpenGL ES), 2D graphics engine (such as: Skia graphics library (SGL)), etc. Among them, the media libraries support the playback and recording of various common audio and video formats, as well as static image files, etc. The media libraries can support various video image coding formats, such as: Moving Pictures Experts Group 4 (MPEG4), Joint Photographic Experts Group (JPG), Portable Network Graphics (PNG), etc.
[0095] The kernel layer is the layer between hardware and software.
[0096] First, Figure 4 the "high-quality data screening" in S101 and the "student picture-text model training" in S102 will be described in detail:
[0097] Taking the student image-text model as an example of the CLIP model, during the process of using the CLIP model for Chinese image-text matching processing, there are significant differences in the accuracy of the CLIP models trained using datasets of different qualities. The reason is that there are significant differences in the quality of the open-source Chinese datasets, and it is necessary to select Chinese datasets with higher quality for training.
[0098] In order to screen high-quality datasets through experiments, taking two representative Chinese datasets in the current open-source community, namely the Taisu dataset and the Wukong dataset, as examples, 11M image-text pairs can be randomly selected from these two datasets as the training set to train the CLIP model. Specifically, the image encoder of the CLIP model uses Vit, and the text encoder uses BERT. Except for the different datasets used for training, other experimental conditions are exactly the same. During the specific experimental process, the industry-standard Flickr30K-cn test set is used as the evaluation dataset, and the evaluation metrics are the text-to-image recall rate (Text-to-Image R1) and the image-to-text recall rate (Image-to-Text R1). The recall rate is also called the recall ratio, which refers to the ratio of the number of relevant targets found to the total number of all actually relevant targets. The larger the recall rate, the more comprehensive the search. As can be seen from Table 1, the CLIP model trained using the Taisu dataset is superior to the CLIP model trained using the Wukong dataset in both the text-to-image recall rate and the image-to-text recall rate. That is, the quality of the Taisu dataset is higher than that of the Wukong dataset. Therefore, the Taisu dataset can be used to train the CLIP model.
[0099] Table 1
[0100] Training set Recall rate of text-to-image search Recall rate of image-to-text search Wukong dataset - 11M 65.24 77.7 Taisu dataset - 11M 68.68 83.2
[0101] The following Figure 4 in S103 "knowledge distillation" will be described in detail. As Figure 8 shown, S103 specifically includes:
[0102] S201. Input the images in multiple image-text pairs into the student image encoder in the student image-text model to obtain student image vectors; input the images in multiple image-text pairs into the teacher image encoder in the teacher image-text model to obtain teacher image vectors; input the texts in multiple image-text pairs into the student text encoder in the student image-text model to obtain student text vectors; input the texts in multiple image-text pairs into the teacher text encoder in the teacher image-text model to obtain teacher text vectors.
[0103] The teacher image-text model can be an open-source Chinese pre-trained CLIP model based on ViT-Huge, and the student image-text model can be an open-source Chinese pre-trained CLIP model based on ViT-Base. The teacher image-text model has stronger image-text learning ability. Therefore, in this application, the learning ability of the teacher image-text model is transferred to the student image-text model through image-text information fusion and knowledge distillation to enhance the learning ability of the student image-text model. Exemplarily, as shown in Table 2, the teacher image-text model is superior to the student image-text model in at least one of the following: the number of model layers, the number of hidden neurons, the size of the multilayer perceptron (MLP), and the number of multi-heads. Among them, the number of model layers refers to the number of layers of the neural network model. The number of hidden neurons refers to the number of hidden neurons in the MLP. The MLP is the core of the model. The input of each neuron in the MLP is x, and the output is y = h w,b (x). By adjusting the parameter weights (weight, w) and biases (bias, b) to fit the input-output relationship reflected by the existing training data (x, y), the MLP size represents the model scale. The larger the MLP size, the larger the model scale. Multi-head refers to multi-head attention, which is used to help the model learn the influence of different contexts on the search results. The more the number of heads, the more conducive to capturing a larger range of features.
[0104] Table 2
[0105]
[0106] Assume that IV represents the image vector and TV represents the text vector. The images in N image-text pairs are input into the student image encoder to obtain the student image vector IV student as follows:
[0107] IV student ={IV student_1 ,IV student_2 ,...,IV student_N |IV student_i ∈R 1*J}
[0108] The texts in N image-text pairs are input into the student text encoder to obtain the student text vector TV student as follows:
[0109] TV student ={TV student_1 ,TV student_2 ,...,TV student_N |TV student_i ∈R 1*J}
[0110] Input the images in N image-text pairs into the teacher image encoder to obtain the teacher image vector IV teacher That is:
[0111] IV teacher ={IV teacher_1 ,IV teacher_2 ,...,IV teacher_N |IV teacher_i ∈R 1*K}
[0112] Input the texts in N image-text pairs into the teacher text encoder to obtain the teacher text vector TV teacher That is:
[0113] TV teacher ={TV teacher_1 ,TV teacher_2 ,...,TV teacher_N |TV teacher_i ∈R 1*K}
[0114] Where N represents the number of image-text pairs, and J < K. Exemplarily, J = 512 and K = 1024. The subscripts 1, 2,..., N in each vector represent the nth image-text pair. For example, IV student_1 represents the student image vector corresponding to the image in the 1st image-text pair, TV student_1 represents the student text vector corresponding to the text in the 1st image-text pair, IV teacher_1 represents the teacher image vector corresponding to the image in the 1st image-text pair, TV teacher_1 represents the teacher text vector corresponding to the text in the 1st image-text pair, and so on. R 1*J represents a 1*J-dimensional real vector, and R 1*K represents a 1*K-dimensional real vector.
[0115] S202. Fuse the student image vector and the student text vector to obtain the student image-text information; fuse the teacher image vector and the teacher text vector to obtain the teacher image-text information.
[0116] The cosine similarity between the student image vector and the student text vector can be used to fuse and obtain the student image-text information. Similarly, the cosine similarity between the teacher image vector and the teacher text vector can be used to fuse and obtain the teacher image-text information. The teacher image-text information comes from a more complex teacher image-text model and has higher-level image-text semantic information compared to the student image-text information from a simple student image-text model. In addition, the fused image-text information also has higher-level image-text semantic information compared to pure image information or text information, and thus has higher knowledge expression ability when used for knowledge distillation.
[0117] Specifically, multiply the student image vector by the student text vector to obtain the student image-text information matrix S V :
[0118] S V = L × IV student @ TV student
[0119] where L is a coefficient used to normalize the values of the student image-text information matrix S V within a reasonable range; "×" represents numerical multiplication, and "@" represents matrix multiplication. The student image-text information matrix S V is an N*N information fusion matrix, and its form is as Figure 9 shown.
[0120] Multiply the teacher image vector by the teacher text vector to obtain the teacher image-text information matrix T V :
[0121] T V = L × IV teacher @ TV teacher
[0122] where L is a coefficient used to normalize the values of the teacher image-text information matrix T V within a reasonable range; "×" represents numerical multiplication, and "@" represents matrix multiplication. The teacher image-text information matrix T V is an N*N information fusion matrix, and its form is as Figure 10 shown.
[0123] S203. Use the teacher image-text information for knowledge distillation of the student image-text information, and update the parameters of the student image-text model according to the distillation loss after knowledge distillation.
[0124] As Figure 11 shown, step S203 includes S2031 - S2034:
[0125] S2031. Obtain the image distillation loss Distill V and the text distillation loss Distill V based on the student image-text information matrix S image and the teacher image-text information matrix T text .
[0126] The image distillation loss Distill image is used to represent the loss of the student image vector relative to the teacher image vector, and the text distillation loss Distill text is used to represent the loss of the student text vector relative to the teacher text vector. Specifically, the image distillation loss Distill imageand the text distillation loss Distill text 。
[0127]
[0128]
[0129] Among them, the matrix S' V represents the transpose of the student image-text information matrix S V , the matrix T' V represents the transpose of the teacher image-text information matrix T V , DistillLossFunction() represents the distillation loss function, softmax() represents the exponential normalization function, and log_softmax() represents taking the log of the result of softmax(). S V (i) represents the i-th row feature in the student image-text information matrix S V , T V (i) represents the i-th row feature in the teacher image-text information matrix T V , S' V (i) represents the i-th row feature in the matrix S' V , T' V (i) represents the i-th row feature in the matrix T' V .
[0130] As Figure 12 shown, the softmax() function has the normalization feature. The larger the value, the closer the value obtained after normalization is to 1, and the smaller the value, the closer the value obtained after normalization is to 0. Using the softmax() function to fuse the teacher image vector and the teacher text vector in the teacher image-text model can significantly strengthen the knowledge on the diagonal of the teacher image-text information matrix and enhance the distillation value of the knowledge. The reason is that the values on the diagonal of the teacher image-text information matrix are much larger than those on the non-diagonal, that is, the knowledge on the diagonal of the teacher image-text information matrix is more and has higher value.
[0131] As Figure 13 shown, the log_softmax() function can map the relatively large values obtained by the softmax() function to the same level, effectively avoiding the overflow problem caused by some features in the student image-text information matrix having too large values, and ensuring the stability of the student image-text model in receiving knowledge.
[0132] S2032. According to the image distillation loss Distill image and the text distillation loss Distill text , the distillation loss Loss distill is obtained.
[0133] Distillation Loss distill It is used to represent the loss of the student image-text model relative to the teacher image-text model after knowledge distillation. Specifically, according to the following formula, the image distillation loss Distill image and the text distillation loss Distill text are weighted and summed to obtain the distillation loss Loss distill .
[0134] Loss distill = θ × Distill image + μ × Distill text
[0135] where θ and μ are coefficients, θ + μ = 1. Exemplarily, θ = 0.5 and μ = 0.5.
[0136] S2033. Obtain the total loss Loss distill based on the distillation loss Loss train and the training loss Loss total .
[0137] The training loss Loss train is used to represent the loss between the model output result and the true value after training the student image-text model with image-text pairs. Specifically, according to the following formula, the distillation loss Loss distill and the training loss Loss train are weighted and summed to obtain the total loss Loss total .
[0138] Loss total = α × Loss train + β × Loss distill
[0139] where α and β are coefficients, α + β = 1. Exemplarily, α = 0.5 and β = 0.5.
[0140] S2034. Update the parameters of the student image-text model according to the total loss Loss total .
[0141] In the prior art, the parameters of the student image-text model are updated according to the training loss Loss train . In this application, the distillation loss Loss distill compensates the training loss Loss train to obtain the total loss Loss total . According to the total loss Loss totalUpdating the parameters of the student graph model can use the knowledge from the teacher graph model to update the parameters of the student graph model. For example, the gradient descent method can be used to iteratively update the parameters of the student graph model until the total loss Loss total The final training is completed when the convergence (no longer changing or changing slowly) reaches the minimum and the preset number of iterations is reached. The parameters of the student graph model are the weights and biases of the MLP mentioned above.
[0142] The principle of gradient descent is as follows Figure 14 As shown in the figure, the loss of the model is a bowl-shaped curve with the parameters of the model as variables. Starting from point A on the curve, the gradient (partial derivative vector) of point A is calculated. The direction in which the gradient is negative is the direction in which the loss is reduced, thereby determining the next point B on the curve. Starting from point B, points C and D on the curve are determined in the above manner. The gradient of point D is approximately 0, indicating that it converges to point D. The parameters at this time are the parameters after the model converges.
[0143] Below Figure 4 S105 "Image-text matching processing" is described in detail.
[0144] Image-text matching processing includes but is not limited to the above-mentioned text-to-image and image-to-text, and can also include high-level image-to-image and image-to-text matching processing that is biased towards artificial intelligence. Text-to-image refers to inputting a complex text into the student image-text model, and the student image-text model generates an image. Image-to-text refers to inputting an image into the student image-text model, and the student image-text model generates a complex text.
[0145] like Figure 15 As shown, step S105 includes:
[0146] S301, electronic device display interface.
[0147] The interface includes controls for obtaining search targets. Search targets can be Figure 2 The text shown in or Figure 3 Visual media (including images and videos) shown in . The controls used to obtain the search target include but are not limited to Figure 2 Search box 204, Figure 3 The photo taking button 305 etc.
[0148] S302. In response to the operation on the control, the electronic device matches the search target according to the student graphic model to obtain search results.
[0149] When the search target is text, the search results are visual media (including images and videos). When the search target is visual media, the search results are text. Figure 2The operation of the user entering text in the search box 204 as shown, or, it can be Figure 3 The click operation of the user on the photographing button 305 as shown. The search results can be Figure 2 The visual media shown in (c) in Figure 3 or, it can be the text shown in (c) in
[0150] S303. The electronic device displays the search results.
[0151] Exemplarily, the electronic device displaying the search results can be Figure 2 displaying the visual media in the search interface 205 as shown in (c) in Figure 3 or, it can be displaying the text in the interface 307 as shown in (c) in
[0152] In the graphic-text matching method provided by the embodiments of the present application, the graphic-text pair is input into the student graphic-text model to obtain the student image vector and the student text vector, and the same graphic-text pair is input into the teacher graphic-text model to obtain the teacher image vector and the teacher text vector. The student image vector and the student text vector are fused to obtain the student graphic-text information; the teacher image vector and the teacher text vector are fused to obtain the teacher graphic-text information. On the one hand, the teacher graphic-text information comes from a more complex teacher graphic-text model, and has higher-level graphic-text semantic information compared with the student graphic-text information from the simple student graphic-text model. On the other hand, the fused graphic-text information has higher-level graphic-text semantic information compared with the pure image information or text information, and has higher knowledge expression ability when used for knowledge distillation. Then, the teacher graphic-text information is used to perform knowledge distillation on the student graphic-text information, the parameters of the student graphic-text model are updated according to the distillation loss after knowledge distillation, and the updated student graphic-text model is used for graphic-text matching processing. In this way, the updated student graphic-text model learns the knowledge of the complex teacher graphic-text model, and can improve the accuracy of the neural network model for graphic-text matching processing.
[0153] As Figure 16 shown, the embodiments of the present application also provide a chip system. The chip system 160 includes at least one processor 1601 and at least one interface circuit 1602. The at least one processor 1601 and the at least one interface circuit 1602 can be interconnected through a line. The processor 1601 is used to support the electronic device to implement each step in the above method embodiments, such as Figure 4 、 Figure 8 、 Figure 11 、 Figure 15 shown in the method, and the at least one interface circuit 1602 can be used to receive signals from other devices (such as a memory), or, send signals to other devices (such as a communication interface). The chip system can include a chip, and can also include other discrete devices.
[0154] The embodiments of the present application further provide a computer-readable storage medium, which includes instructions. When the instructions run on the above-mentioned electronic device or server, the electronic device or server is caused to execute each step in the method embodiments above. For example, execute Figure 4 , Figure 8 , Figure 11 , Figure 15 the methods shown.
[0155] The embodiments of the present application further provide a computer program product including instructions. When the instructions run on the above-mentioned electronic device or server, the electronic device or server is caused to execute each step in the method embodiments above. For example, execute Figure 4 , Figure 8 , Figure 11 , Figure 15 the methods shown.
[0156] Regarding the technical effects of the chip system, computer-readable storage medium, and computer program product, refer to the technical effects of the foregoing method embodiments.
[0157] It should be understood that in various embodiments of the present application, the magnitudes of the sequence numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0158] Those of ordinary skill in the art can realize that the modules and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0159] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and modules described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0160] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or modules can be in electrical, mechanical, or other forms.
[0161] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical modules, that is, they can be located in one device or distributed to multiple devices. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0162] In addition, in each embodiment of the present application, the functional modules can be integrated in one device, or each module can exist physically alone, or two or more modules can be integrated in one device.
[0163] In the above embodiments, they can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using a software program, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or data center that contains one or more integrated media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.
[0164] As described above, it is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the said claims.
Claims
1. A method for matching images and texts, It is characterized in that The method comprises: Displaying an interface, the interface including a control for obtaining a search target; In response to the operation of the control, the search target is matched according to the student graphic model to obtain search results; the search target is text and the search result is visual media, or the search target is visual media and the search result is text; displaying the search results; The training process of the student graph-text model includes: Input the images in the plurality of image-text pairs into a student image encoder in a student image-text model to obtain a student image vector; input the images into a teacher image encoder in a teacher image-text model to obtain a teacher image vector; input the texts in the plurality of image-text pairs into a student text encoder in the student image-text model to obtain a student text vector; input the texts into a teacher text encoder in the teacher image-text model to obtain a teacher text vector; The student image vector and the student text vector are merged to obtain student image and text information; the teacher image vector and the teacher text vector are merged to obtain teacher image and text information; The teacher's graphic information is used to perform knowledge distillation on the student's graphic information, and the parameters of the student's graphic model are updated according to the distillation loss after the knowledge distillation.
2. The method according to claim 1, It is characterized in that The step of fusing the student image vector with the student text vector to obtain student image-text information; and fusing the teacher image vector with the teacher text vector to obtain teacher image-text information includes: Perform matrix multiplication of the student image vector and the student text vector to obtain the student image-text information matrix; Perform matrix multiplication of the teacher image vector and the teacher text vector to obtain the teacher image and text information matrix.
3. The method according to claim 2, It is characterized in that The step of performing knowledge distillation on the student image-text information using the teacher image-text information, and updating the parameters of the student image-text model according to the distillation loss after knowledge distillation, includes: An image distillation loss and a text distillation loss are obtained according to the student image-text information matrix and the teacher image-text information matrix; the image distillation loss is used to represent the loss of the student image vector relative to the teacher image vector, and the text distillation loss is used to represent the loss of the student text vector relative to the teacher text vector; Obtaining the distillation loss according to the image distillation loss and the text distillation loss; The total loss is obtained according to the distillation loss and the training loss; the training loss is used to represent the loss between the model output result and the true value after the student image-text model is trained using the image-text pair; Parameters of the student graph-text model are updated according to the total loss.
4. The method according to claim 3, It is characterized in that The obtaining of the image distillation loss and the text distillation loss according to the student image-text information matrix and the teacher image-text information matrix includes: The image distillation loss Distill is obtained according to the following two formulas respectively image and the text distillation loss Distill text : Among them, the matrix S' V represents the transpose of the student graphic and text information matrix S V The matrix T' V represents the transpose of the teacher graphic and text information matrix T V softmax() represents the exponential normalization function, log_softmax() represents taking the log of the result of softmax(), S V (i) represents the i-th row feature in the student graphic and text information matrix S V The T V (i) represents the i-th row feature in the teacher graphic and text information matrix T V The S' V (i) represents the i-th row feature in the matrix S' V The T V '(i) represents the i-th row feature in the matrix T V ', and N represents the number of the graphic and text pairs.
5. The method according to claim 3, It is characterized in that The obtaining the distillation loss according to the image distillation loss and the text distillation loss includes: The distilled loss is obtained by performing a weighted sum of the image distilled loss and the text distilled loss.
6. The method according to claim 3, wherein, obtaining the total loss according to the distilled loss and the training loss includes: performing a weighted sum of the distilled loss and the training loss to obtain the total loss.
7. The method according to any one of claims 1-6, wherein, updating the parameters of the student image-text model according to the total loss includes: iteratively updating the parameters of the student image-text model using the gradient descent method until the total loss converges to the lowest and a preset number of iterations is reached.
8. The method according to any one of claims 1-7, wherein, the teacher image-text model is superior to the student image-text model in at least one of the following: the number of model layers, the number of hidden neurons, the size of the multi-layer perceptron, and the number of heads.
9. An electronic device, wherein, it includes a processor and a memory, and instructions are stored in the memory. When the processor executes the instructions, the method according to any one of claims 1-8 is executed.
10. A computer-readable storage medium, wherein, it includes instructions. When the instructions are executed on an electronic device, the electronic device is caused to execute the method according to any one of claims 1-8.
Citation Information
Patent Citations
Multi-modal joint representation learning method and system based on variational distillation
CN114841335A
Image-text retrieval method and system based on cross-modal cross guidance
CN116186317A
Image text model processing method and image text retrieval system
CN116521833A
Single stream multi-level alignment for vision-language pretraining
US20230281963A1
Picture-text model generation method and apparatus based on multiple experts, and device and medium
WO2023168811A1