Training Method, System and Storage Medium of Multimodal Retrieval Model on Scholar Network

By introducing data adaptation module, pre-training module and nonlinear migration module into the multimodal data retrieval model of Xuezhe.com, the problem of low data retrieval accuracy of Xuezhe.com is solved, and more efficient multimodal data retrieval effect is achieved.

CN114912576BActive Publication Date: 2025-06-10SOUTH CHINA NORMAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210360056.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-07
Publication Date
2025-06-10
Estimated Expiration
2042-04-07

AI Technical Summary

Technical Problem

The existing technology is difficult to improve the accuracy of the multimodal data retrieval results of Scholar Network, mainly because the graphic and text semantic correspondence of the Scholar Network data set is weak, the data is noisy, and the text is mainly text, which is not suitable for the existing pre-training modules and fine-tuning methods.

Method used

A training method for multimodal retrieval model of scholar network is proposed, including data adaptation module, pre-training module and nonlinear migration module. By crawling the user data of the Scholar Network, input it into the data adaptation module, obtaining continuous feature vectors, and then inputting them into the pre-training module to obtain higher-order semantic information feature vectors. Then, the predicted similarity of the graphic and text data and text data is calculated by the nonlinear migration module, and the model parameters are adjusted based on the predicted similarity and true similarity.

Benefits of technology

It effectively improves the accuracy of the data retrieval results of Scholar Network, adapts to the characteristics of multimodal data of Scholar Network, and improves the performance of the model on Scholar Network data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114912576B_ABST
    Figure CN114912576B_ABST
Patent Text Reader

Abstract

The present invention discloses a training method, system and storage medium for a multi-modal retrieval model of a scholar network, which can be widely applied to the field of retrieval technology. The method of the present invention inputs multiple user data including graphic and text data and first text data crawled into a data adaptation module to obtain continuous feature vectors that can be received by a pre-training module. Then, after inputting the continuous feature vectors into the pre-training module, high-order semantic information feature vectors are obtained. Furthermore, through a non-linear transfer module, according to the high-order semantic information feature vectors, the predicted similarity between the graphic and text data and the first text data is calculated. Then, the predicted similarity and the true similarity are used to adjust the parameters of the multi-modal retrieval model of the scholar network, so that the parameters of the multi-modal retrieval model of the scholar network can achieve a better effect for data retrieval of the scholar network, thereby effectively improving the accuracy of the scholar network data retrieval results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of retrieval technology, and in particular to a training method, system and storage medium for a multimodal retrieval model of a scholar network. Background Art

[0002] In the related art, in common multimodal data, images often correspond to text descriptions of the images, that is, images and text have a fairly strong underlying semantic association. However, the data sets collected on academic websites generally have a weaker semantic correspondence between images and text, and are more inclined to human activities rather than the correspondence between images and text in detail. At the same time, the data noise is relatively large. Compared with existing image and text data sets, the text in Scholar.com will be more common, more text-based, and not used to describe the image in detail. Therefore, the distribution of multimodal data based on Scholar.com is different from the distribution of existing data sets, which leads to the use of existing pre-training modules as initialization parameters, and then using a small amount of downstream data for fine-tuning, and then using a small amount of downstream data for fine-tuning. It is difficult to improve the accuracy of retrieval results. Summary of the invention

[0003] The present invention aims to solve at least one of the technical problems existing in the prior art. To this end, the present invention proposes a training method, system and storage medium for a scholar network multimodal retrieval model, which can effectively improve the accuracy of scholar network data retrieval results.

[0004] On the one hand, an embodiment of the present invention provides a method for training a scholar network multimodal retrieval model, wherein the scholar network multimodal retrieval model includes a data adaptation module, a pre-training module and a non-linear migration module, and the training method includes the following steps:

[0005] Crawling multiple user data of the scholar network, wherein the user data includes graphic data and first text data;

[0006] Inputting the user data into the data adaptation module to obtain a continuous feature vector;

[0007] Inputting the continuous feature vector into the pre-training module to obtain a high-order semantic information feature vector;

[0008] Inputting the high-order semantic information feature vector into the nonlinear migration module, and calculating the predicted similarity between the graphic data and the first text data;

[0009] Obtaining true similarity between the graphic data and the first text data;

[0010] The parameters of the scholar.net multimodal retrieval model are adjusted according to the predicted similarity and the true similarity.

[0011] In some embodiments, the data adaptation module includes a linear embedding layer and a text extraction layer; the step of inputting the user data into the data adaptation module to obtain a continuous feature vector includes:

[0012] Input the first text data into the linear embedding layer to obtain a first text continuous feature vector;

[0013] Input the graphic-text data into the text extraction layer to obtain second text data;

[0014] Input the second text data into the linear embedding layer to obtain a second text continuous feature vector.

[0015] In some embodiments, the pre-training module includes a text feature extractor and a visual feature extractor; the step of inputting the continuous feature vector into the pre-training module to obtain a high-order semantic information feature vector includes:

[0016] Input the first text continuous feature vector into the text feature extractor to obtain a first text high-order semantic information feature vector; and input the second text continuous feature vector into the text feature extractor to obtain a second text high-order semantic information feature vector;

[0017] Input the graphic-text data into the visual feature extractor to obtain an image high-order semantic information feature vector.

[0018] In some embodiments, the step of inputting the high-order semantic information feature vector into the non-linear transfer module to calculate the predicted similarity between the graphic-text data and the first text data includes:

[0019] Input the first text high-order semantic information feature vector and the image high-order semantic information feature vector into the non-linear transfer module to obtain a first similarity;

[0020] Input the first text high-order semantic information feature vector and the second text high-order semantic information feature vector into the non-linear transfer module to obtain a second similarity;

[0021] Calculate the predicted similarity according to the first similarity and the second similarity.

[0022] In some embodiments, the non-linear transfer module includes a fully connected layer, a BN layer, and a ReLU layer.

[0023] In some embodiments, before inputting the first text data into the linear embedding layer, the method further includes the following steps:

[0024] Convert the first text data into word text through a Chinese word segmentation tool.

[0025] In some embodiments, the text extraction layer includes an OCR module.

[0026] On the other hand, an embodiment of the present invention provides a training system for a multi-modal retrieval model of a scholar network. The multi-modal retrieval model of the scholar network includes a data adaptation module, a pre-training module, and a non-linear transfer module. The system includes:

[0027] A crawling module for crawling multiple user data of the scholar network, where the user data includes graphic and text data and first text data;

[0028] A first data processing module for inputting the user data into the data adaptation module to obtain a continuous feature vector;

[0029] A second data processing module for inputting the continuous feature vector into the pre-training module to obtain a high-order semantic information feature vector;

[0030] A calculation module for inputting the high-order semantic information feature vector into the non-linear transfer module to calculate the predicted similarity between the graphic and text data and the first text data;

[0031] An acquisition module for acquiring the true similarity between the graphic and text data and the first text data;

[0032] An adjustment module for adjusting the parameters of the multi-modal retrieval model of the scholar network according to the predicted similarity and the true similarity.

[0033] On the other hand, an embodiment of the present invention provides a training system for a multi-modal retrieval model of a scholar network, including:

[0034] At least one memory for storing a program;

[0035] At least one processor for loading the program to execute the training method of the multi-modal retrieval model of the scholar network.

[0036] On the other hand, an embodiment of the present invention provides a storage medium in which a computer-executable program is stored. When the computer-executable program is executed by a processor, it is used to implement the training method of the multi-modal retrieval model of the scholar network.

[0037] The training method of the multi-modal retrieval model of the scholar network provided by this embodiment has the following beneficial effects:

[0038] In this embodiment, a plurality of user data including graphic and text data and first text data crawled are input into a data adaptation module to obtain continuous feature vectors that can be received by a pre-training module. Then, after the continuous feature vectors are input into the pre-training module, high-order semantic information feature vectors are obtained. Next, according to the high-order semantic information feature vectors, a non-linear transfer module calculates the predicted similarity between the graphic and text data and the first text data. Then, the predicted similarity and the true similarity are used to adjust the parameters of the multi-modal retrieval model of the Scholar Network, so that the parameters of the multi-modal retrieval model of the Scholar Network can achieve a better effect for data retrieval of the Scholar Network, effectively improving the accuracy of the data retrieval results of the Scholar Network.

[0039] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. Brief Description of the Drawings

[0040] The following further describes the present invention in conjunction with the drawings and embodiments, where:

[0041] Figure 1 is a flowchart of a training method for a multi-modal retrieval model of the Scholar Network according to an embodiment of the present invention;

[0042] Figure 2 is a schematic diagram of a multi-modal retrieval model of the Scholar Network according to an embodiment of the present invention. Detailed Embodiments

[0043] The embodiments of the present invention are described in detail below. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions from beginning to end. The embodiments described below with reference to the drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation of the present invention.

[0044] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by terms such as up, down, front, back, left, right, etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention.

[0045] In the description of the present invention, the meaning of several is more than one, the meaning of multiple is more than two, greater than, less than, exceeding, etc. are understood as not including the present number, and above, below, within, etc. are understood as including the present number. If there is a description of first and second, it is only for the purpose of distinguishing technical features and should not be construed as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or implicitly indicating the sequence of the indicated technical features.

[0046] In the description of the present invention, unless otherwise clearly defined, terms such as setting, installing, connecting, etc. should be understood in a broad sense, and technicians in the relevant technical field can reasonably determine the specific meanings of the above terms in the present invention based on the specific content of the technical solution.

[0047] In the description of the present invention, the description with reference to the terms "one embodiment", "some embodiments", "illustrative embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0048] With the rapid development and widespread application of the Internet, academic social networks have received widespread attention. Scholar.com is an academic social network with a large number of academic users, which generates a large amount of academic data for practical applications, providing a data basis for academic services such as academic relationships and academic activities in real scenarios. Among these data, there are also a large number of multimodal data including pictures and texts published by scholars on Scholar.com.

[0049] In order to better learn and understand the multimodal data of real academic social networks, one approach is to collect a large amount of real data for training, but it is difficult to collect enough data. On the other hand, with the continuous expansion of computing resources, the number of network layers is getting deeper and deeper, and the number of parameters is increasing, which is data-hungry. Both of these aspects make it difficult to directly use the target data for training from scratch.

[0050] The academic community multimodal dataset for practical applications is different from the general multimodal dataset. In common multimodal datasets, images often correspond to text descriptions of the image, which has quite strong low-level semantic association information. However, the datasets collected on academic websites generally have weaker semantic correspondence between images and texts, and are more inclined to human activities rather than correspondence between images and texts in details. At the same time, the data noise is relatively large. Compared with existing image and text datasets, the text in Scholar.com will be longer and more text-based, and is not used to describe the image in detail. Images serve as supplementary information here to better illustrate the specific situation. The distribution of multimodal data based on Scholar.com is different from that of existing datasets, which makes it difficult to achieve good results by using the traditional method of using pre-trained modules as initialization parameters and then fine-tuning with a small amount of downstream data.

[0051] Based on this, refer to Figure 1, embodiments of the present invention provide a training method for a multi-modal retrieval model of a scholar network. Specifically, as Figure 2 shown, the multi-modal retrieval model of the scholar network includes a data adaptation module, a pre-training module, and a non-linear transfer module. For Figure 2 the training method of the model shown, as Figure 1 shown, the training method includes steps 110 - 150:

[0052] Step 110, crawl multiple user data of the scholar network, where the user data includes graphic and text data and first text data.

[0053] In the embodiments of the present application, mainly crawl the multi-modal data of scholars in the academic social network as user data. Among them, the multi-modal data of scholars includes text data T i and image data I i , and use the crawled text data T i as the first text data. Define to represent a set of training samples with two modalities of image and text. Since there is also a part of text data in the image data , then this part of the text data can be used as the second text data, and define In the process of multi-modal data retrieval, given a modality of data, another modality of data can be retrieved. For example, if the user gives image data, the corresponding text data of the graphic and text data can be retrieved; conversely, if the user gives text data, the image data can be retrieved.

[0054] Step 120, input the user data into the data adaptation module to obtain continuous feature vectors.

[0055] In the embodiments of the present application, as Figure 2 shown, the data adaptation module includes a linear embedding layer and a text extraction layer. Based on this structure, in this step, the first text data can be input into the linear embedding layer to convert the first text data into a first text continuous feature vector through the linear embedding layer; input the graphic and text data into the text extraction layer, extract the second text data from the image data through the text extraction layer, and then input the second text data into the linear embedding layer to convert the second text data into a second text continuous feature vector through the linear embedding layer. In this embodiment, the text extraction layer can adopt an OCR module. Among them, the OCR (Optical Character Recognition) module refers to an electronic device that monitors the characters printed on paper, determines their shapes by detecting the dark and bright patterns, and then translates the shapes into computer text using character recognition methods.

[0056] In this embodiment, since this embodiment focuses on the retrieval of images and texts, for the text modality, because the distribution of texts on the dataset is different from the distribution of other multi-modal datasets, this embodiment adds a simple linear embedding layer to obtain Chinese features on the Scholar Network. For the image visual modality, in order to make full use of the text information contained in the images, this embodiment first uses OCR technology to extract the text in the images and uses the extracted text as additional information.

[0057] Specifically, for Figure 2 's data adaptive module, in this embodiment, a Chinese word segmentation tool is first used to segment the Chinese in the dataset, and then the segmented words are used in the pre-training module Chinese RoBERTa Large to obtain the Chinese feature representation, and then connected to a simple linear layer to process the text features into the text input vector of the pre-training module. The text T i can be expressed as w 1 , w 2 ,..., w m , where w j is the feature representation of the j-th Chinese sample, and m is the total number of words in this text. Since there is a large amount of text information on the Scholar Network images, this embodiment uses an existing OCR model to extract text from the images where OCR is any optical character recognition algorithm, is the text extracted from the image. The same operations are also performed on the text extracted from the image, so as to obtain the feature representation of each word of where l is the total number of words.

[0058] Step 130: Input the continuous feature vector into the pre-training module to obtain a high-order semantic information feature vector.

[0059] In the embodiment of the present application, as Figure 2 shown, the pre-training module includes a text feature extractor and a visual feature extractor. Based on this structure, in this step, the first text continuous feature vector can be input into the text feature extractor to obtain a first text high-order semantic information feature vector; and the second text continuous feature vector can be input into the text feature extractor to obtain a second text high-order semantic information feature vector; the image-text data is input into the visual feature extractor to obtain an image high-order semantic information feature vector. Specifically, the first text continuous feature vector and the second text continuous feature vector can be encoded by the same text feature extractor to obtain the text high-order semantic information feature vector. The visual feature extractor processes the image data from pixel points into a continuous feature representation.

[0060] In the embodiments of the present application, for the pre-training module, there are mainly three input data. One is the original text T i which can be represented as w 1 , w 2 ,..., w m , the image data I i and the text extracted from the image For the visual features, the trained visual feature extractor is used to process the image from pixel points into a continuous feature representation v i = Visual Encoder(I i ). Similarly, for the two texts, the same text extractor is used to obtain the feature representations of the two texts. The feature representations of the text data are f i = Text Encoder(T i ), For the pre-training module, since the data volume of the multi-modal data of the downstream task ScholarNet is too small, the model parameters are fixed in this embodiment to improve the model utilization rate.

[0061] Step 140: Input the high-order semantic information feature vector into the non-linear transfer module, and calculate the predicted similarity between the graphic and text data and the first text data.

[0062] In this embodiment, the first text high-order semantic information feature vector and the image high-order semantic information feature vector can be input into the non-linear transfer module to obtain the first similarity; at the same time, the first text high-order semantic information feature vector and the second text high-order semantic information feature vector are input into the non-linear transfer module to obtain the second similarity; then, according to the first similarity and the second similarity, the predicted similarity is calculated.

[0063] Specifically, due to the difference in the data used by the pre-trained model and the data distribution of the downstream ScholarNet data, the existing multi-modal data pays more attention to specific low-level semantic associations, while the academic multi-modal data requires higher-level academic activities and other topics. In order to reduce low-level semantics and pay more attention to high-level semantics during the knowledge transfer process, in the training process, this embodiment uses a non-linear transfer module to better transfer the knowledge learned upstream to the ScholarNet data. At the same time, a lightweight transfer module adapter2Scholar is proposed to convert the existing pre-training module to the ScholarNet multi-modal data.

[0064] After the image and text pass through the pre-trained model, the obtained feature representations f i , v i and Then, a lightweight non - linear transfer module is used to adapt to and learn the multi - modal data of the downstream Scholar Network. This module consists of multiple fully - connected layers, BN layers, and ReLU (activation function) layers. The new feature representations of the image data and the first text data are respectively where W v and W t are two parameters to be learned. W v is for learning visual features, while W t is for learning text features. After passing through a residual neural network, multiple layers are stacked on the original features for non - linear transformation, and then the original features are added to obtain two new features of sample i. The similarity between the image data and the first text data is

[0065] According to the above network, the new feature representations of the image - supplemented text (second text data) and the existing text data (first text data) can be obtained as respectively Since the extracted text and the existing text belong to the same modality, in order to better learn the similarity between the two, the same learning parameter is used. The similarity between the two features is These new features ignore the low - level semantic - related features and focus on higher - level information such as themes and human activities, so as to better transfer the knowledge learned by the pre - trained multi - modal model to the multi - modal data retrieval of the Scholar Network. The similarities in two dimensions, namely the similarity between image and text and the similarity between text and text, are fused to obtain the final multi - similarity fusion result sim = sim1+sim2. The multi - similarity fusion result is used as the predicted similarity between the graphic - text data and the first text data.

[0066] Step 150: Obtain the true similarity between the graphic - text data and the first text data.

[0067] Step 160: Adjust the parameters of the Scholar Network multi - modal retrieval model according to the predicted similarity and the true similarity.

[0068] In the embodiment of the present application, by inversely adjusting the parameters of the model, the model can better adapt to the multi - modal retrieval process of the Scholar Network.

[0069] In the training process of this embodiment, the logistic regression loss function minL is defined sim = ∑ (i,j)y(i, j) log(sim(i, j) + (1 - y(i, j)) log(1 - sim(i, j))), where y(i, j) represents the true similarity between the image and the text, 1 indicates similarity, and 0 indicates dissimilarity. In this embodiment, the loss function is used to transmit the gradient back to the lightweight migration module through backpropagation, realizing the parameter learning of the migration module from downstream to upstream, so as to obtain the image-text retrieval result using the model.

[0070] Specifically, after completing the training process of the above model, the above model can be used for the actual retrieval process. Among them, the retrieval process includes the following steps:

[0071] Step 1: Sample the image-text modal data of the Scholar Network data and divide it into text data T i and image data I i ;

[0072] Step 2: Process the text data with a word segmentation tool, and add a simple linear embedding layer Linear Embedding before the pre-training module to process the text data into a continuous feature vector T i ={w 1 , w 2 ,..., w m};

[0073] Step 3: Use the optical character recognition algorithm OCR to extract the text data in the image to obtain the continuous feature vector of the text supplementary information of the image

[0074] Step 4: Respectively use the visual feature extractor Visual Encoder (I i ) and Text Encoder (T i ) for the image data and the text data to obtain the feature vector representations v i , f i ,

[0075] Step 5: Add a lightweight non-linear migration module at the backend of the pre-training model to adapt to the few-shot multi-modal data downstream, and then input the feature vector with high-order semantics into the migration module composed of multiple fully connected layers, BN layers and ReLU layers;

[0076] Step 6: Input the high-order semantic features of the image and text data into the migration module to obtain the new features of the image and text Obtain the image-text similarity by calculating the cosine similarity of the two new features

[0077] Step 7: Input the high-order semantic features of the text extracted from the image and the existing text into the transfer module to obtain new features Calculate the similarity between the two

[0078] Step 8: Integrate the feature similarities in two dimensions: image and text, and text and text, to obtain the multi-dimensional similarity between the text and the image: sim = sim1 + sim2;

[0079] Step 9: Use the similarity score sim obtained after image-text matching to represent the similarity between the input image and the text, so as to perform image-text information retrieval.

[0080] In summary, a training method for a multi-modal retrieval model of a scholar network provided in this embodiment has the following beneficial effects:

[0081] First, extract the text information on the image as supplementary information for the image to further improve the accuracy of image-text retrieval;

[0082] Second, add a simple linear embedding layer and optical character recognition technology OCR at the data input end to process the data into continuous feature vectors, which is convenient for better utilization of the characteristics of the scholar network data;

[0083] Third, select a pre-trained model trained by a large amount of big data to perform rapid data analysis on a small amount of academic multi-modal data, and strengthen the ability of data learning feature expression;

[0084] Fourth, fix the parameters of the pre-trained model to solve the problem of insufficient downstream scholar network multi-modal data volume, and at the same time make better use of the existing model;

[0085] Fifth, during the training process, automatically learn a non-linear model transfer module, which is used to better transfer the knowledge learned upstream to the data of the scholar network;

[0086] Sixth, add multiple non-linear layers at the backend of the pre-trained multi-modal model to further distill relevant knowledge from the upstream task to the downstream task, and better realize the knowledge transfer from upstream to downstream.

[0087] Seventh, use a lightweight transfer module to transfer the existing pre-trained model to the scholar network multi-modal data, and solve the problem of the difference between the model data and the test data.

[0088] The embodiment of the present invention provides a training system for a multi-modal retrieval model of a scholar network. The multi-modal retrieval model of the scholar network includes a data adaptation module, a pre-training module, and a non-linear transfer module. The system includes:

[0089] A crawling module, configured to crawl multiple user data of the Scholar Network, where the user data includes graphic and text data and first text data;

[0090] A first data processing module, configured to input the user data into the data adaptation module to obtain continuous feature vectors;

[0091] A second data processing module, configured to input the continuous feature vectors into the pre-training module to obtain high-order semantic information feature vectors;

[0092] A calculation module, configured to input the high-order semantic information feature vectors into the non-linear transfer module to calculate the predicted similarity between the graphic and text data and the first text data;

[0093] An acquisition module, configured to acquire the true similarity between the graphic and text data and the first text data;

[0094] An adjustment module, configured to adjust the parameters of the multi-modal retrieval model of the Scholar Network according to the predicted similarity and the true similarity.

[0095] The content of the method embodiment of the present invention is applicable to the system embodiment of the present invention. The functions specifically implemented by the system embodiment of the present invention are the same as those of the above method embodiment, and the beneficial effects achieved are also the same as those of the above method.

[0096] An embodiment of the present invention provides a training system for a multi-modal retrieval model of the Scholar Network, including:

[0097] At least one memory, configured to store a program;

[0098] At least one processor, configured to load the program to execute Figure 1 The training method of the multi-modal retrieval model of the Scholar Network as shown.

[0099] The content of the method embodiment of the present invention is applicable to the system embodiment of the present invention. The functions specifically implemented by the system embodiment of the present invention are the same as those of the above method embodiment, and the beneficial effects achieved are also the same as those of the above method.

[0100] An embodiment of the present invention provides a storage medium, in which a computer-executable program is stored, and when the computer-executable program is executed by a processor, it is used to implement Figure 1 The training method of the multi-modal retrieval model of the Scholar Network as shown.

[0101] An embodiment of the present invention also provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes Figure 1 the training method of the multi-modal retrieval model of the Scholar Network shown in

[0102] The embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those of ordinary skill in the art to which the present invention pertains, various changes can be made without departing from the gist of the present invention. In addition, without conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other.

Claims

1. A training method for a multi-modal retrieval model of a scholar network, characterized in that, the multi-modal retrieval model of the scholar network includes a data adaptation module, a pre-training module, and a non-linear transfer module, and the training method includes the following steps: Crawl multiple user data of the scholar network, where the user data includes graphic and text data and first text data; Input the user data into the data adaptation module to obtain continuous feature vectors; Input the continuous feature vectors into the pre-training module to obtain high-order semantic information feature vectors; Input the high-order semantic information feature vectors into the non-linear transfer module to calculate the predicted similarity between the graphic and text data and the first text data; Obtain the true similarity between the graphic and text data and the first text data; Adjust the parameters of the multi-modal retrieval model of the scholar network according to the predicted similarity and the true similarity; wherein, the data adaptation module includes a linear embedding layer and a text extraction layer; the step of inputting the user data into the data adaptation module to obtain continuous feature vectors includes: Input the first text data into the linear embedding layer to obtain first text continuous feature vectors; Input the graphic and text data into the text extraction layer to obtain second text data; Input the second text data into the linear embedding layer to obtain second text continuous feature vectors; wherein, the step of calculating the predicted similarity between the graphic and text data and the first text data includes: Perform dimensional fusion on the image and text similarity between the graphic and text data and the first text data and the text and text similarity between the second text data and the first text data to obtain a multi-similarity fusion result; Use the multi-similarity fusion result as the predicted similarity between the graphic and text data and the first text data.

2. The training method for a multi-modal retrieval model of a scholar network according to claim 1, characterized in that, the pre-training module includes a text feature extractor and a visual feature extractor; the step of inputting the continuous feature vectors into the pre-training module to obtain high-order semantic information feature vectors includes: Input the first text continuous feature vectors into the text feature extractor to obtain first text high-order semantic information feature vectors; and input the second text continuous feature vectors into the text feature extractor to obtain second text high-order semantic information feature vectors; Input the graphic and text data into the visual feature extractor to obtain image high-order semantic information feature vectors.

3. The training method for a multi-modal retrieval model of a scholar network according to claim 2, characterized in that, the step of inputting the high-order semantic information feature vectors into the non-linear transfer module to calculate the predicted similarity between the graphic and text data and the first text data includes: Input the first text high-order semantic information feature vectors and the image high-order semantic information feature vectors into the non-linear transfer module to obtain a first similarity; Input the first text high-order semantic information feature vectors and the second text high-order semantic information feature vectors into the non-linear transfer module to obtain a second similarity; Calculate a predicted similarity based on the first similarity and the second similarity.

4. A method for training a multi-modal retrieval model of a scholar network according to claim 3, wherein, the non-linear migration module includes a fully connected layer, a BN layer, and a ReLU layer.

5. A method for training a multi-modal retrieval model of a scholar network according to claim 1, wherein, before inputting the first text data into the linear embedding layer, the method further includes the following steps: Convert the first text data into word text through a Chinese word segmentation tool.

6. A method for training a multi-modal retrieval model of a scholar network according to claim 1, wherein, the text extraction layer includes an OCR module.

7. A training system for a multi-modal retrieval model of a scholar network, wherein, the multi-modal retrieval model of the scholar network includes a data adaptation module, a pre-training module, and a non-linear migration module, and the system includes: A crawling module for crawling multiple user data of the scholar network, where the user data includes graphic and text data and first text data; A first data processing module for inputting the user data into the data adaptation module to obtain a continuous feature vector; A second data processing module for inputting the continuous feature vector into the pre-training module to obtain a high-order semantic information feature vector; A calculation module for inputting the high-order semantic information feature vector into the non-linear migration module to calculate the predicted similarity between the graphic and text data and the first text data; An acquisition module for acquiring the true similarity between the graphic and text data and the first text data; An adjustment module for adjusting the parameters of the multi-modal retrieval model of the scholar network according to the predicted similarity and the true similarity; wherein, the data adaptation module includes a linear embedding layer and a text extraction layer; inputting the user data into the data adaptation module to obtain a continuous feature vector includes: Inputting the first text data into the linear embedding layer to obtain a first text continuous feature vector; Inputting the graphic and text data into the text extraction layer to obtain second text data; Inputting the second text data into the linear embedding layer to obtain a second text continuous feature vector; wherein, calculating the predicted similarity between the graphic and text data and the first text data includes: Performing dimension fusion on the image and text similarity between the graphic and text data and the first text data and the text and text similarity between the second text data and the first text data to obtain a multi-similarity fusion result; Using the multi-similarity fusion result as the predicted similarity between the graphic and text data and the first text data.

8. A training system for a multi-modal retrieval model of a scholar network, wherein, includes: At least one memory for storing programs; At least one processor for loading the program to execute the method for training a multi-modal retrieval model of a scholar network according to any one of claims 1-6.

9. A storage medium, wherein, It stores a computer-executable program, which is used to implement the training method of the ScholarNet multi-modal retrieval model according to any one of claims 1-6 when executed by a processor.

Citation Information

Patent Citations

  • Multi-modal pre-training model training method, application method and device thereof

    CN112990297A

  • Multimedia data searching method and device, equipment and storage medium

    CN113590850A