Joint training method of image-text vector model and image-text large model and electronic equipment

By jointly training the graphic vector model and the graphic big model, the problem of intimate model coordination is solved, multimodal retrieval is improved to enhance the accuracy of generation and reduce computing resource consumption, achieving a more efficient training process.

CN120256592AActive Publication Date: 2025-07-04INST OF AUTOMATION CHINESE ACAD OF SCI
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510743100.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-07-04
Estimated Expiration
2045-06-05

AI Technical Summary

Technical Problem

During the training process of existing graphic search and generation models, the graphic vector model and graphic large model are usually independently trained, resulting in a lack of close coordination between the two and affecting the overall performance.

Method used

The joint training method of graphic vector model and graphic big model is adopted. By obtaining the general question and answer sample set, industry question and answer sample set and industry knowledge base, the graphic big model is used to process the general question samples and graphic vector model are used to process the industry problem samples, and the model parameters are updated in combination with the loss value to achieve close cooperation of the model.

Benefits of technology

It improves multimodal retrieval and enhances the overall accuracy of generation, reduces the computational amount and hardware resource consumption during training, and improves the overall performance of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256592A_ABST
    Figure CN120256592A_ABST
Patent Text Reader

Abstract

The invention provides a joint training method for an image-text vector model and an image-text large model and electronic equipment. The method comprises the following steps: acquiring a general question and answer sample set, an industry question and answer sample set and an industry knowledge base; processing a general question sample in the general question and answer sample set by using a to-be-trained image-text large model to obtain a general answer; processing industry question samples in the industry question and answer sample set by using a to-be-trained image-text vector model, an industry knowledge base and a to-be-trained image-text large model to obtain industry answers; determining a loss value according to the general answers, the general question and answer sample set, the industry answers and the industry question and answer sample set; and updating parameters of the to-be-trained image-text large model according to the loss value, and updating parameters of the to-be-trained image-text vector model according to the industry answer to obtain the image-text large model and the image-text vector model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to the fields of artificial intelligence and deep learning, and more specifically, to a method for jointly training a graphic-text vector model and a graphic-text large model, and an electronic device. Background Art

[0002] With the development of artificial intelligence technology, graphic-text retrieval and generation models have been widely used in various fields, and retrieval-augmented generation is one of the most commonly used application scenarios. In the retrieval-augmented generation scenario, a graphic-text vector model is usually used to extract features of a question input by a user, so as to retrieve an image or text segment related to the input question as context information, and then a graphic-text large model is used to process the input question and the context information to generate an answer to the question, which can make the generated answer more accurate and rich, and enhance the performance of the generation task.

[0003] However, in existing retrieval-augmented generation methods, the graphic-text vector model and the graphic-text large model are usually trained independently, resulting in insufficient tight cooperation between the two and affecting the overall performance. Summary of the Invention

[0004] The present disclosure provides a method for jointly training a graphic-text vector model and a graphic-text large model, and an electronic device, which are used to solve at least one of the above problems.

[0005] According to a first aspect of an embodiment of the present disclosure, there is provided a method for jointly training a graphic-text vector model and a graphic-text large model, the method including: obtaining a general Q&A sample set, an industry Q&A sample set, and an industry knowledge base; using a to-be-trained graphic-text large model to process general question samples in the general Q&A sample set to obtain general answers; using the to-be-trained graphic-text vector model, the industry knowledge base, and the to-be-trained graphic-text large model to process industry question samples in the industry Q&A sample set to obtain industry answers; determining a loss value according to the general answers, the general Q&A sample set, the industry answers, and the industry Q&A sample set; updating parameters of the to-be-trained graphic-text large model according to the loss value, and updating parameters of the to-be-trained graphic-text vector model according to the industry answers to obtain the graphic-text large model and the graphic-text vector model.

[0006] Optionally, the using the to-be-trained graphic-text vector model, the industry knowledge base, and the to-be-trained graphic-text large model to process industry question samples in the industry Q&A sample set to obtain industry answers includes: using the to-be-trained graphic-text vector model to extract question feature vectors of the industry question samples; retrieving a preset number of industry knowledge segments matching the question feature vectors from the industry knowledge base; using the to-be-trained graphic-text large model to process the industry question samples and the preset number of industry knowledge segments to obtain the industry answers.

[0007] Optionally, using the to-be-trained image-text large model to process the industry problem sample and the preset number of industry knowledge segments to obtain the industry answer includes: using the to-be-trained image-text large model to process the industry problem sample and the preset number of industry knowledge segments to obtain the preset number of candidate industry answers, where each candidate industry answer corresponds to each industry knowledge segment one by one; determining one candidate industry answer from the preset number of candidate industry answers as the industry answer.

[0008] Optionally, updating the parameters of the to-be-trained image-text vector model according to the industry answer includes: determining the industry knowledge segment corresponding to the industry answer as the positive sample, and determining the other industry knowledge segments in the preset number of industry knowledge segments except the positive sample as the negative samples; updating the parameters of the to-be-trained image-text vector model using the contrastive learning method based on the positive sample and the negative samples.

[0009] Optionally, the image-text large model and the image-text vector model are obtained after multiple parameter updates, and the method further includes: when a preset update condition is satisfied, determining the latest image-text vector model as the knowledge base image-text vector model, and using the knowledge base image-text vector model to extract the segment feature vectors of each industry knowledge segment in the industry knowledge base, where the segment feature vectors are used to calculate the matching degree between the corresponding industry knowledge segment and the problem feature vector.

[0010] Optionally, the preset update condition includes at least one of the following: the update times of the image-text vector model are equal to 0 or greater than or equal to a preset number of times, and a preset time duration has elapsed since the last update of the knowledge base image-text vector model.

[0011] Optionally, the image-text large model and the image-text vector model are obtained after multiple parameter updates, and updating the parameters of the to-be-trained image-text large model according to the loss value and updating the parameters of the to-be-trained image-text vector model according to the industry answer to obtain the image-text large model and the image-text vector model includes: when a synchronous update condition is satisfied, updating the parameters of the to-be-trained image-text large model according to the loss value and updating the parameters of the to-be-trained image-text vector model according to the industry answer to obtain the image-text large model and the image-text vector model; when the synchronous update condition is not satisfied, updating the parameters of the to-be-trained image-text large model according to the loss value and keeping the parameters of the to-be-trained image-text vector model unchanged to obtain the image-text large model and the image-text vector model.

[0012] According to a second aspect of the embodiments of the present disclosure, there is provided a joint training device for a graphic-text vector model and a graphic-text large model. The device includes: an acquisition unit configured to acquire a general question-answering sample set, an industry question-answering sample set, and an industry knowledge base; a general question-answering unit configured to process general question samples in the general question-answering sample set using the graphic-text large model to be trained to obtain general answers; an industry question-answering unit configured to process industry question samples in the industry question-answering sample set using the graphic-text vector model to be trained, the industry knowledge base, and the graphic-text large model to be trained to obtain industry answers; a determination unit configured to determine a loss value according to the general answers, the general question-answering sample set, the industry answers, and the industry question-answering sample set; and an update unit configured to update parameters of the graphic-text large model to be trained according to the loss value and update parameters of the graphic-text vector model to be trained according to the industry answers to obtain the graphic-text large model and the graphic-text vector model.

[0013] Optionally, the industry question-answering unit is further configured to: use the graphic-text vector model to be trained to extract a question feature vector of the industry question sample; retrieve a preset number of industry knowledge fragments matching the question feature vector from the industry knowledge base; and use the graphic-text large model to be trained to process the industry question sample and the preset number of industry knowledge fragments to obtain the industry answer.

[0014] Optionally, the industry question-answering unit is further configured to: use the graphic-text large model to be trained to process the industry question sample and the preset number of industry knowledge fragments to obtain the preset number of candidate industry answers, where each candidate industry answer corresponds to each industry knowledge fragment one by one; and determine one candidate industry answer from the preset number of candidate industry answers as the industry answer.

[0015] Optionally, the update unit is further configured to: determine the industry knowledge fragment corresponding to the industry answer as a positive sample, and determine other industry knowledge fragments in the preset number of industry knowledge fragments except the positive sample as negative samples; and update parameters of the graphic-text vector model to be trained using a contrastive learning method based on the positive sample and the negative sample.

[0016] Optionally, the graphic-text large model and the graphic-text vector model are obtained after multiple parameter updates. The device further includes an extraction unit configured to, when a preset update condition is satisfied, determine the latest graphic-text vector model as a knowledge base graphic-text vector model, and use the knowledge base graphic-text vector model to extract fragment feature vectors of each industry knowledge fragment in the industry knowledge base, where the fragment feature vectors are used to calculate the matching degree between the corresponding industry knowledge fragment and the question feature vector.

[0017] Optionally, the preset update condition includes at least one of the following: the update times of the graphic-text vector model is equal to 0 or greater than or equal to a preset number of times, or a preset duration has elapsed since the last update of the graphic-text vector model of the knowledge base.

[0018] Optionally, the graphic-text large model and the graphic-text vector model are obtained after multiple parameter updates. Wherein, the update unit is further configured to: when the synchronous update condition is satisfied, update the parameters of the graphic-text large model to be trained according to the loss value, and update the parameters of the graphic-text vector model to be trained according to the industry answer to obtain the graphic-text large model and the graphic-text vector model; when the synchronous update condition is not satisfied, update the parameters of the graphic-text large model to be trained according to the loss value, and keep the parameters of the graphic-text vector model to be trained unchanged to obtain the graphic-text large model and the graphic-text vector model.

[0019] According to a third aspect of the embodiments of the present disclosure, there is provided an electronic device, including: at least one processor; at least one memory storing computer-executable instructions, wherein when the computer-executable instructions are run by the at least one processor, the at least one processor is caused to execute the joint training method of the graphic-text vector model and the graphic-text large model according to the exemplary embodiments of the present disclosure.

[0020] According to a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, and when the instructions in the computer-readable storage medium are run by at least one processor, the at least one processor is caused to execute the joint training method of the graphic-text vector model and the graphic-text large model according to the exemplary embodiments of the present disclosure.

[0021] According to a fifth aspect of the embodiments of the present disclosure, there is provided a computer program product, including computer instructions, and when the computer instructions are run by at least one processor, the at least one processor is caused to execute the joint training method of the graphic-text vector model and the graphic-text large model according to the exemplary embodiments of the present disclosure.

[0022] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects: According to the joint training method and electronic device of the graphic-text vector model and the graphic-text large model of the present disclosure, by jointly training the graphic-text vector model and the graphic-text large model that are associated in application, not only can the graphic-text vector model and the graphic-text large model cooperate closely, which helps to improve the overall accuracy of multi-modal retrieval enhanced generation and improve the overall performance of the model, but also the overall calculation amount in training can be reduced, the computing resource consumption of the electronic device executing the training can be reduced, and the hardware loss can be reduced.

[0023] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. Brief Description of the Drawings

[0024] The drawings herein are incorporated into and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure, and do not constitute an undue limitation of the present disclosure.

[0025] Figure 1 is a flowchart of a method for jointly training a graphic - text vector model and a graphic - text large model according to an exemplary embodiment of the present disclosure.

[0026] Figure 2 is a schematic diagram of a process for generating industry answers according to an exemplary embodiment of the present disclosure.

[0027] Figure 3 is a schematic diagram of a process for generating a fragment feature vector database according to an exemplary embodiment of the present disclosure.

[0028] Figure 4 is a schematic diagram of a process for retrieving industry knowledge fragments according to an exemplary embodiment of the present disclosure.

[0029] Figure 5 is a schematic diagram of an asynchronous update mechanism for a graphic - text vector model and a knowledge - base graphic - text vector model according to an exemplary embodiment of the present disclosure.

[0030] Figure 6 is a schematic diagram of a process for generating answers in different situations according to an exemplary embodiment of the present disclosure.

[0031] Figure 7 is a block diagram of a device for jointly training a graphic - text vector model and a graphic - text large model according to an exemplary embodiment of the present disclosure.

[0032] Figure 8 is a block diagram of an electronic device according to an exemplary embodiment of the present disclosure. Detailed Description of the Embodiments

[0033] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the drawings.

[0034] It should be noted that the terms "first", "second", etc. in the description, claims and above-mentioned drawings of the present disclosure are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following examples do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0035] It should be noted here that "at least one of several items" in the present disclosure all represents three parallel situations, including "any one of the several items", "a combination of any multiple of the several items", and "all of the several items". For example, "including at least one of A and B" includes the following three parallel situations: (1) including A; (2) including B; (3) including A and B. Another example is "performing at least one of step one and step two", which means the following three parallel situations: (1) performing step one; (2) performing step two; (3) performing step one and step two.

[0036] Next, a method for jointly training a text-image vector model and a text-image large model and an electronic device according to an exemplary embodiment of the present disclosure will be described in detail with reference to the drawings.

[0037] Figure 1 is a flowchart of a method for jointly training a text-image vector model and a text-image large model according to an exemplary embodiment of the present disclosure. This method can be executed on an electronic device with sufficient computing power.

[0038] Referring to Figure 1 , in step S101, a general Q&A sample set, an industry Q&A sample set, and an industry knowledge base are obtained.

[0039] Specifically, the general Q&A sample set includes multiple pairs of general question samples and general answer samples, and the industry Q&A sample set includes multiple pairs of industry question samples and industry answer samples. Each general question sample, each general answer sample, each industry question sample, and each industry answer sample can include only text data, that is, a question or answer in text form, or can include both text data and image data, that is, a question or answer in a text-and-image combination form. As for the differences between these two sample sets, as the name implies, industry questions are questions whose answers require the use of industry knowledge, while general questions are questions that do not require the use of industry knowledge. The industry knowledge base is used to provide industry knowledge, including multiple industry knowledge segments. Each industry knowledge segment can include only text data, only image data, or both text data and image data. The industry knowledge base can be constructed manually, automatically by an electronic device, or by combining manual and electronic device construction. The present disclosure does not limit this.

[0040] It should be understood that the training of the model generally includes multiple parameter updates, and the processes described in steps S101 to S105 are the processes executed for one update. For the case of multiple updates, the complete sample set can be divided into multiple sample subsets, and each update is executed based on one sample subset. Therefore, the general Q&A sample set and the industry Q&A sample set here are the sample subsets used for one update, and the industry knowledge base can be shared among different parameter updates.

[0041] In step S102, the general question samples in the general Q&A sample set are processed using the to-be-trained text-and-image large model to obtain general answers.

[0042] This step directly uses the to-be-trained text-and-image large model to answer the general questions.

[0043] In step S103, the industry question samples in the industry Q&A sample set are processed using the to-be-trained text-and-image vector model, the industry knowledge base, and the to-be-trained text-and-image large model to obtain industry answers.

[0044] This step uses the to-be-trained text-and-image vector model to retrieve relevant industry knowledge from the industry knowledge base as a reference for the to-be-trained text-and-image large model to answer industry questions.

[0045] In step S104, the loss value is determined based on the general answers, the general Q&A sample set, the industry answers, and the industry Q&A sample set.

[0046] Optionally, this step compares the general answer obtained in step S102 with the general answer samples corresponding to the same general question sample in the general Q&A sample set to obtain a first loss value, and compares the industry answer obtained in step S103 with the industry answer samples corresponding to the same industry question sample in the industry Q&A sample set to obtain a second loss value. Finally, the weighted value of the first loss value and the second loss value is determined as the loss value to describe the error of the answer, which is used as a reference for updating the graphic and text large model. As an example, the first loss value or the second loss value can be determined based on the loss of predicting the next Token by the graphic and text large model.

[0047] In step S105, the parameters of the graphic and text large model to be trained are updated according to the loss value, and the parameters of the graphic and text vector model to be trained are updated according to the industry answer, so as to obtain the graphic and text large model and the graphic and text vector model.

[0048] For the graphic and text large model, the previous steps belong to the forward propagation steps in model training, and this step belongs to the backward propagation steps in model training. The gradient can be calculated using the backward propagation algorithm, and the parameters of the model can be updated using an optimization algorithm (such as the gradient descent algorithm). For the graphic and text vector model, the accuracy of the industry knowledge retrieved in step S103 can be evaluated according to the industry answer, and the evaluation result can be used as the basis for updating the parameters of the graphic and text vector model.

[0049] According to the joint training method of the graphic and text vector model and the graphic and text large model of the exemplary embodiment of the present disclosure, by jointly training the graphic and text vector model and the graphic and text large model that are associated in application, not only can the graphic and text vector model and the graphic and text large model cooperate closely, which helps to improve the overall accuracy of multi-modal retrieval enhanced generation and improve the overall performance of the model, but also the overall computational amount in training can be reduced, the computational resource consumption of the electronic device executing the training can be reduced, and the hardware loss can be reduced. In addition, by distinguishing and processing general questions and industry questions, and not using the industry knowledge base to retrieve general questions, not only can the data processing amount be reduced, the computational resource consumption of the hardware be reduced, but also the interference of industry knowledge on the answers to general questions can be reduced, and the accuracy of the answers can be improved.

[0050] As an example, a pre-trained graphic and text vector model and a graphic and text large model can be used as the basic models, and steps S101 to S105 are used to jointly fine-tune the two basic models.

[0051] It should be understood that the above step numbers are for convenience of distinguishing different steps and do not limit the execution order. Under reasonable logic, Figure 1 the execution order of the steps in can be adjusted. For example, the execution order of step S102 and step S103 can be swapped, or they can be executed simultaneously when the electronic device has parallel computing capabilities. The present disclosure does not limit this.

[0052] Next, a joint training method for a graphic-text vector model and a graphic-text large model according to an exemplary embodiment of the present disclosure will be further introduced.

[0053] Regarding the answer to the industry question, optionally, step S103 includes: using the graphic-text vector model to be trained to extract the question feature vector of the industry question sample; retrieving a preset number of industry knowledge fragments matching the question feature vector from the industry knowledge base; using the graphic-text large model to be trained to process the industry question sample and the preset number of industry knowledge fragments to obtain an industry answer. By using the graphic-text vector model to be trained to extract the question feature vector, vectorization processing of the industry question sample can be achieved, and accordingly, retrieval and matching of industry knowledge fragments can be realized, providing reference information for the generation of industry answers.

[0054] As an example, as Figure 2 shown, a knowledge base graphic-text vector model similar to the graphic-text vector model can be used to uniformly perform vectorization processing on the industry knowledge fragments in the industry knowledge base. The specific processing process is as Figure 3 shown. First, a multimodal parser can be used to parse industry knowledge fragments of different modalities. The picture data is directly input into the knowledge base graphic-text vector model, and the text data (including the original text and the introduction text of the picture) is preliminarily processed by a sentence splitter and then input into the knowledge base graphic-text vector model to obtain the fragment feature vector of each industry knowledge fragment. The fragment feature vectors are summarized to obtain a fragment feature vector database, which can be regarded as the vectorized database of the industry knowledge base. As Figure 4 shown, when performing retrieval and matching (since the processing method in the model inference stage is the same, the model to be trained is not emphasized anymore), the graphic-text vector model is used to extract the question feature vector of the industry question (the industry question may include picture data and / or text data), and match it with each fragment feature vector in the fragment feature vector database. For example, the Euclidean distance between each fragment feature vector and the question feature vector can be calculated as the matching degree, and the sorting model sorts these matching degrees to select the K fragment feature vectors with the highest matching degree (for example, the shortest Euclidean distance), and returns the corresponding K industry knowledge fragments in the industry knowledge base as the Top-K industry knowledge fragments, so as to retrieve the preset number of industry knowledge fragments with the highest matching degree. Then as Figure 2 shown, the returned Top-K industry knowledge fragments and the industry question are input into the graphic-text large model to generate an industry answer.

[0055] Further optionally, the operation of using the to-be-trained graphic and text large model in step S103 to process the industry problem sample and a preset number of industry knowledge fragments to obtain an industry answer includes: using the to-be-trained graphic and text large model to process the industry problem sample and a preset number of industry knowledge fragments to obtain a preset number of candidate industry answers, where each candidate industry answer corresponds to each industry knowledge fragment one by one; determining one candidate industry answer from the preset number of candidate industry answers as the industry answer. By generating candidate industry answers for each industry knowledge fragment respectively, the mutual influence between different retrieved industry knowledge fragments can be reduced, the accuracy of each candidate industry answer itself can be improved, and then one can be selected from them to realize the answer to the industry problem.

[0056] It should be understood that the method of generating candidate industry answers by combining industry problem samples and industry knowledge fragments is an existing technology in this field and will not be elaborated here. Regarding how to determine the industry answer from the preset number of candidate industry answers, multiple factors can be comprehensively considered. As an example, content relevance can be considered, such as the semantic relevance between the candidate industry answer and the industry problem (such as including but not limited to semantic similarity), and the degree of fit between the candidate industry answer and the corresponding industry knowledge fragment; answer quality can also be considered, such as the language fluency, information integrity, logical consistency, etc. of the candidate industry answer; the credibility and authority of the information source can also be considered. Of course, other reasonable factors can also be considered and will not be listed one by one here.

[0057] Based on the embodiment of obtaining a preset number of industry knowledge fragments and the corresponding candidate industry answers, optionally, the operation of updating the parameters of the to-be-trained graphic and text vector model according to the industry answer in step S105 includes: determining the industry knowledge fragment corresponding to the industry answer as a positive sample, and determining the other industry knowledge fragments in the preset number of industry knowledge fragments except the positive sample as negative samples; based on the positive samples and negative samples, using the contrastive learning method to update the parameters of the to-be-trained graphic and text vector model. By determining the positive and negative correlations between the corresponding industry knowledge fragment and the industry answer according to whether a certain candidate industry answer is selected, it is convenient to evaluate the retrieval results obtained based on the graphic and text vector model and provide a reference for the parameter update of the graphic and text vector model. On this basis, the chain update of the graphic and text vector model is realized by using the contrastive learning method, ensuring the training quality.

[0058] Based on the embodiment of retrieving a preset number of industry knowledge segments that match the problem feature vector from the industry knowledge base, as described above, a knowledge base image-text vector model similar to the image-text vector model can be used to uniformly vectorize the industry knowledge segments in the industry knowledge base, and the obtained segment feature vectors are used to calculate the matching degree between the corresponding industry knowledge segments and the problem feature vector. In this regard, optionally, the image-text large model and the image-text vector model are obtained after multiple parameter updates. The joint training method of the image-text vector model and the image-text large model according to the exemplary embodiment of the present disclosure further includes: when the preset update condition is satisfied, determining the latest image-text vector model as the knowledge base image-text vector model, and using the knowledge base image-text vector model to extract the segment feature vectors of each industry knowledge segment in the industry knowledge base.

[0059] This embodiment proposes an asynchronous update mechanism for the image-text vector models on the problem side and the knowledge base side during the training process. Specifically, using the same image-text vector model on the problem side and the knowledge base side (i.e., the structures and parameters of the image-text vector model and the knowledge base image-text vector model are the same) can ensure the consistent extraction of the feature vectors of industry problems and industry knowledge segments. However, each time the knowledge base image-text vector model is updated, it is necessary to synchronously update the segment feature vector database corresponding to the industry knowledge base. Due to the large amount of data in the industry knowledge base, it will cause a large amount of computational effort, seriously affect the update efficiency, and consume a large amount of hardware computing resources. By adopting an asynchronous update mechanism, as Figure 5 shown, the parameters of the image-text vector model on the problem side are updated in a timely manner. For the knowledge base image-text vector model, it is copied from the problem side when the preset update condition is satisfied, and the segment feature vector database is updated, which can effectively reduce the update frequency of the segment feature vector database, reduce the computational effort and the resulting hardware computing resource consumption, and improve the model training efficiency. It should be understood that by reasonably configuring the preset update condition, the knowledge base image-text vector model can be reasonably updated, so as to balance the update frequency and the effectiveness of the segment feature vector database, and ensure the training quality.

[0060] Optionally, the preset update conditions include at least one of the following: the continuous update count of the text-image vector model is equal to 0 or greater than or equal to a preset count, and a preset duration has elapsed since the last update of the text-image vector model in the knowledge base. The preset count and the preset duration can both be set as required. It should be understood that both are positive numbers, and the preset count threshold is greater than 1, so that the update frequency of the text-image vector model is higher than that of the text-image vector model in the knowledge base, that is, the text-image vector model is updated first, and after a certain number of times (specifically, the preset count minus 1 time) of delay, the text-image vector model in the knowledge base is updated. In particular, when the continuous update count of the text-image vector model is equal to 0, it means that the parameters of the text-image vector model have not been updated yet. At this time, the unupdated text-image vector model can also be copied as the text-image vector model in the knowledge base to initialize the model and ensure the smooth progress of the initial update of the model parameters.

[0061] Optionally, the text-image large model and the text-image vector model are obtained after multiple parameter updates. This means that the text-image large model obtained in one update is used as the text-image large model to be trained in the next update, and the text-image vector model obtained in one update is used as the text-image vector model to be trained in the next update. Step S105 specifically includes: when the synchronous update condition is met, updating the parameters of the text-image large model to be trained according to the loss value, and updating the parameters of the text-image vector model to be trained according to the industry answer to obtain the text-image large model and the text-image vector model; when the synchronous update condition is not met, updating the parameters of the text-image large model to be trained according to the loss value, and keeping the parameters of the text-image vector model to be trained unchanged to obtain the text-image large model and the text-image vector model. By configuring the synchronous update condition, it is possible to update the text-image large model and the text-image vector model simultaneously only when this condition is met. Otherwise, only the text-image large model needs to be updated, which can reasonably control the computational amount of model updates and help optimize the computational resource consumption of the electronic device.

[0062] As an example, the synchronous update conditions include at least one of the following: the continuous update count of the text-image large model reaches a count threshold, and a duration threshold has elapsed since the last update of the text-image vector model. The count threshold and the duration threshold can both be set as required. It should be understood that both are positive numbers, and the count threshold can be greater than 1, so that the update frequency of the text-image large model is higher than that of the text-image vector model, that is, the text-image large model is updated first, and after a certain number of times (specifically, the count threshold minus 1 time) of delay, the text-image vector model is updated. In particular, if the count threshold is set to 1, it means that the parameters of the text-image large model and the text-image vector model are always updated synchronously.

[0063] In a specific embodiment, the multimodal retrieval enhancement generation system is divided into three main parts: knowledge base vectorization, knowledge base retrieval, and answer generation. The system uses the pre-trained graph-text vector model and the graph-text large model as the basic model to build the multimodal retrieval enhancement generation system framework. First, the industry knowledge base data is vectorized using the knowledge base graph-text vector model and stored in the segment feature vector database. The graph-text vector model is used to target industry problems, such as Figure 4 As shown in the figure, industry questions are searched in the industry knowledge base and the corresponding fragment feature vector database, and the relevant Top-K industry knowledge fragments are obtained as the search results. The large picture and text model is used to generate answers. For general questions, no search is performed and general answers are directly generated. For industry questions, the following is performed first: Figure 4 The search results are combined to generate industry answers. Figure 6 As shown, when it is determined that there are search results, it means that the question (query) is an industry question. At this time, the question and the search results are input into the large picture and text model as prompts to generate industry answers; when there are no search results, it means that the question is a general question. If there is no refusal to answer script set in advance, the question is input into the large picture and text model as a prompt to generate a general answer. If there is a refusal to answer script set in advance, the refusal to answer script is directly used as the answer.

[0064] Figure 7 is a block diagram of a joint training device for a text-image vector model and a text-image large model according to an exemplary embodiment of the present disclosure. Figure 7 The joint training device 700 of the image-text vector model and the image-text large model includes an acquisition unit 701, a general question and answer unit 702, an industry question and answer unit 703, a determination unit 704, and an update unit 705.

[0065] The acquisition unit 701 may acquire a general question and answer sample set, an industry question and answer sample set, and an industry knowledge base.

[0066] The general question and answer unit 702 can use the large picture and text model to be trained to process the general question samples in the general question and answer sample set to obtain a general answer.

[0067] The industry question and answer unit 703 can use the image and text vector model to be trained, the industry knowledge base and the image and text large model to be trained to process the industry question samples in the industry question and answer sample set to obtain industry answers.

[0068] Optionally, the industry question and answer unit 703 can also: use the image-text vector model to be trained to extract the question feature vector of the industry question sample; retrieve a preset number of industry knowledge fragments that match the question feature vector from the industry knowledge base; use the image-text large model to be trained to process the industry question samples and a preset number of industry knowledge fragments to obtain industry answers.

[0069] Optionally, the industry Q&A unit 703 may also: use the text-image large model to be trained to process industry question samples and a preset number of industry knowledge segments to obtain a preset number of candidate industry answers, where each candidate industry answer corresponds one-to-one to each industry knowledge segment; determine one candidate industry answer from the preset number of candidate industry answers as the industry answer.

[0070] The determination unit 704 may determine a loss value according to the general answer, the general Q&A sample set, the industry answer, and the industry Q&A sample set.

[0071] The update unit 705 may update the parameters of the text-image large model to be trained according to the loss value, and update the parameters of the text-image vector model to be trained according to the industry answer to obtain the text-image large model and the text-image vector model.

[0072] Optionally, the update unit 705 may also: determine the industry knowledge segment corresponding to the industry answer as a positive sample, and determine the other industry knowledge segments except the positive sample among the preset number of industry knowledge segments as negative samples; based on the positive samples and negative samples, use the contrastive learning method to update the parameters of the text-image vector model to be trained.

[0073] Optionally, the text-image large model and the text-image vector model are obtained after multiple parameter updates. Among them, the update unit 705 may also: under the condition of meeting the synchronous update condition, update the parameters of the text-image large model to be trained according to the loss value, and update the parameters of the text-image vector model to be trained according to the industry answer to obtain the text-image large model and the text-image vector model; under the condition of not meeting the synchronous update condition, update the parameters of the text-image large model to be trained according to the loss value, and keep the parameters of the text-image vector model to be trained unchanged to obtain the text-image large model and the text-image vector model.

[0074] Optionally, the text-image large model and the text-image vector model are obtained after multiple parameter updates. Among them, the joint training device 700 for the text-image vector model and the text-image large model according to the exemplary embodiments of the present disclosure further includes an extraction unit (not shown in the figure). The extraction unit may, under the condition of meeting the preset update condition, determine the latest text-image vector model as the knowledge base text-image vector model, and use the knowledge base text-image vector model to extract the segment feature vectors of each industry knowledge segment in the industry knowledge base, where the segment feature vectors are used to calculate the matching degree between the corresponding industry knowledge segment and the question feature vector.

[0075] Optionally, the preset update condition includes at least one of the following: the update times of the text-image vector model is equal to 0 or greater than or equal to the preset times, and a preset duration has elapsed since the last update of the knowledge base text-image vector model.

[0076] Regarding the device in the above embodiments, the specific manner in which each unit performs operations has been described in detail in the embodiments related to the method, and will not be elaborated herein.

[0077] Figure 8 The structural block diagram of an electronic device 800 according to an exemplary embodiment of the present disclosure is shown.

[0078] Referring to Figure 8 , the electronic device 800 includes: at least one memory 801 and at least one processor 802. Computer-executable instructions are stored in the at least one memory 801. When the computer-executable instructions are run by the at least one processor 802, the at least one processor is caused to execute the joint training method of the graphic vector model and the graphic large model as described in the above exemplary embodiments.

[0079] As an example, the electronic device 800 may be a PC computer, a tablet device, a personal digital assistant, a smart phone, or other devices capable of executing the above instruction set. Here, the electronic device 800 does not have to be a single electronic device 800, and may also be any assembly of devices or circuits that can execute the above instructions (or instruction sets) individually or jointly. The electronic device 800 may also be part of an integrated control system or system manager, or may be configured as a portable electronic device 800 that can be interconnected with a local or remote (e.g., via wireless transmission) interface.

[0080] In the electronic device 800, the processor 802 may include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller, or a microprocessor. By way of example and not limitation, the processor 802 may also include an analog processor, a digital processor, a microprocessor, a multi-core processor, a processor array, a network processor, etc.

[0081] The processor 802 may run instructions or code stored in the memory 801, where the memory 801 may also store data. The instructions and data may also be sent and received via a network interface device over a network, where the network interface device may employ any known transmission protocol.

[0082] The memory 801 may be integrated with the processor 802. For example, RAM or flash memory may be arranged within an integrated circuit microprocessor, etc. In addition, the memory 801 may include a separate device, such as an external disk drive, a storage array, or other storage devices that can be used by any database system. The memory 801 and the processor 802 may be operatively coupled or may communicate with each other, for example, through an I / O port, a network connection, etc., such that the processor 802 can read files stored in the memory.

[0083] In addition, the electronic device 800 may further include a video display (such as, a liquid crystal display) and a user interaction interface (such as, a keyboard, a mouse, a touch input device, etc.). All components of the electronic device 800 may be connected to each other via a bus and / or a network.

[0084] According to an exemplary embodiment of the present disclosure, a computer-readable storage medium storing instructions may also be provided, wherein the instructions, when run by at least one processor, cause the at least one processor to execute the joint training method of the graphic vector model and the graphic large model as described in the above exemplary embodiment. Examples of such computer-readable storage media include: read-only memory (ROM), random access programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disc memory, hard disk drive (HDD), solid state drive (SSD), cartridge memory (such as, multimedia card, secure digital (SD) card or extreme digital (XD) card), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid state disk, and any other device configured to store a computer program and any associated data, data files, and data structures in a non-transitory manner and provide the computer program and any associated data, data files, and data structures to a processor or computer such that the processor or computer can execute the computer program. The computer program in the above computer-readable storage medium may run in an environment deployed in computer devices such as clients, hosts, proxy devices, servers, etc. In addition, in one example, the computer program and any associated data, data files, and data structures are distributed on a networked computer system such that the computer program and any associated data, data files, and data structures are stored, accessed, and executed in a distributed manner by one or more processors or computers.

[0085] According to an exemplary embodiment of the present disclosure, a computer program product may also be provided, including computer instructions that, when run by at least one processor, execute the joint training method of the graphic vector model and the graphic large model as described in the above exemplary embodiment.

[0086] Other embodiments of the present disclosure will be readily apparent to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common general knowledge or conventional technical means in the technical field not disclosed by the present disclosure. The specification and examples are only to be considered as exemplary, and the true scope and spirit of the present disclosure are indicated by the appended claims.

[0087] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.

Claims

1. A joint training method for a graphic-text vector model and a graphic-text large model, characterized in that, The method includes: Obtaining a general Q&A sample set, an industry Q&A sample set, and an industry knowledge base; Processing the general question samples in the general Q&A sample set using the text and image large model to be trained to obtain general answers; Processing the industry question samples in the industry Q&A sample set using the text and image vector model to be trained, the industry knowledge base, and the text and image large model to be trained to obtain industry answers; Determining a loss value according to the general answers, the general Q&A sample set, the industry answers, and the industry Q&A sample set; Updating the parameters of the text and image large model to be trained according to the loss value, and updating the parameters of the text and image vector model to be trained according to the industry answers to obtain the text and image large model and the text and image vector model.

2. The method according to claim 1, characterized in that, The processing the industry question samples in the industry Q&A sample set using the text and image vector model to be trained, the industry knowledge base, and the text and image large model to be trained to obtain industry answers includes: Using the text and image vector model to be trained to extract the question feature vector of the industry question sample; Retrieving a preset number of industry knowledge fragments from the industry knowledge base that match the question feature vector; Using the text and image large model to be trained to process the industry question sample and the preset number of industry knowledge fragments to obtain the industry answer.

3. The method according to claim 2, characterized in that The processing the industry question sample and the preset number of industry knowledge fragments using the text and image large model to be trained to obtain the industry answer includes: Using the text and image large model to be trained to process the industry question sample and the preset number of industry knowledge fragments to obtain the preset number of candidate industry answers, where each candidate industry answer corresponds to each industry knowledge fragment; Determining one candidate industry answer from the preset number of candidate industry answers as the industry answer.

4. The method according to claim 3, characterized in that, The updating the parameters of the text and image vector model to be trained according to the industry answer includes: Determining the industry knowledge fragment corresponding to the industry answer as the positive sample, and determining the other industry knowledge fragments in the preset number of industry knowledge fragments except the positive sample as the negative samples; Updating the parameters of the text and image vector model to be trained using the contrastive learning method based on the positive sample and the negative samples.

5. The method according to claim 2, characterized in that, The text and image large model and the text and image vector model are obtained after multiple parameter updates, and the method further includes: When a preset update condition is satisfied, determining the latest text and image vector model as the knowledge base text and image vector model, and using the knowledge base text and image vector model to extract the fragment feature vectors of each industry knowledge fragment in the industry knowledge base, where the fragment feature vectors are used to calculate the matching degree between the corresponding industry knowledge fragment and the question feature vector.

6. The method according to claim 5, wherein the preset update condition includes at least one of the following: the update times of the text and image vector model are equal to 0 or greater than or equal to a preset number of times, and a preset duration has elapsed since the last update of the knowledge base text and image vector model.

7. The method according to claim 1, wherein The described text-image large model and the text-image vector model are obtained after multiple parameter updates. Among them, updating the parameters of the text-image large model to be trained according to the loss value and updating the parameters of the text-image vector model to be trained according to the industry answer to obtain the text-image large model and the text-image vector model includes: When the synchronous update condition is satisfied, updating the parameters of the text-image large model to be trained according to the loss value and updating the parameters of the text-image vector model to be trained according to the industry answer to obtain the text-image large model and the text-image vector model; When the synchronous update condition is not satisfied, updating the parameters of the text-image large model to be trained according to the loss value and keeping the parameters of the text-image vector model to be trained unchanged to obtain the text-image large model and the text-image vector model.

8. An electronic device, characterized in that, It includes: At least one processor; At least one memory storing computer-executable instructions, wherein, when the computer-executable instructions are run by the at least one processor, the at least one processor is caused to execute the joint training method of the text-image vector model and the text-image large model according to any one of claims 1 to 7.

9. A computer-readable storage medium, characterized in that, The instructions in the computer-readable storage medium, when run by at least one processor, cause the at least one processor to execute the joint training method of the text-image vector model and the text-image large model according to any one of claims 1 to 7.

10. A computer program product comprising computer instructions, characterized in that, When the computer instructions are run by at least one processor, the at least one processor is caused to execute the joint training method of the text-image vector model and the text-image large model according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Extraction type intelligent question-answering method and system introducing agricultural domain knowledge

    CN112527999A

  • Knowledge question-answering method and system based on large language model

    CN117708282A

  • Question and answer matching method and device and electronic equipment

    CN118113822A

  • Model pre-training method and device for multi-language task

    CN119293514A

  • Method, Apparatus for Determining Answer to Question, Device, Storage Medium and Program Product

    US20230214688A1