Question answering method and device based on large model, electronic equipment and storage medium
By determining the knowledge information set corresponding to the keywords of the question in the knowledge base and using a large model to process the information, the problems of low knowledge retrieval efficiency and inaccurate answer services in the existing technology are solved, and fast and accurate question answers and in-depth knowledge query are achieved.
Patent Information
- Application Number
- CN202510344524.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-06-20
AI Technical Summary
When providing question-solving services in the prior art, there are problems such as low information retrieval efficiency and accuracy, users need to compare multiple similar questions one by one, and it is difficult to inquire about knowledge points in depth.
By determining the knowledge information set corresponding to the keywords of the target question information in the knowledge base, and processing the target question information and the knowledge information set using a large model to generate answer information.
It improves the efficiency and accuracy of knowledge retrieval, can provide questions quickly and accurately, and allows users to conduct deeper knowledge query and learning on target question information.
Smart Images

Figure CN120179807A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence, and particularly to the fields of large models and natural language processing. Background Art
[0002] Currently, with the help of big data and related algorithms, many education and training platforms provide application products with question-solving functions. The question-solving function will identify the uploaded questions, and then display multiple questions similar to the uploaded questions in the question bank in a list sorting form by calling the search interface. Users can click on each of the returned questions one by one to compare whether the uploaded question is the same as the returned question, so as to find the solution information corresponding to the uploaded question. Summary of the Invention
[0003] This disclosure provides a method, apparatus, electronic device, and storage medium for question solving based on a large model.
[0004] According to one aspect of this disclosure, a method for question solving based on a large model is provided, including:
[0005] Determine the keywords of the target question information;
[0006] Determine the knowledge information set corresponding to the keywords of the target question information in the knowledge base; wherein, the knowledge base includes multiple keywords and multiple knowledge information sets corresponding to the multiple keywords;
[0007] Use the first large model to process the target question information and the knowledge information set to obtain the solution information.
[0008] According to another aspect of this disclosure, a device for question solving based on a large model is provided, including:
[0009] A keyword determination module for determining the keywords of the target question information;
[0010] A knowledge determination module for determining the knowledge information set corresponding to the keywords of the target question information in the knowledge base; wherein, the knowledge base includes multiple keywords and multiple knowledge information sets corresponding to the multiple keywords;
[0011] A solution module for using the first large model to process the target question information and the knowledge information set to obtain the solution information.
[0012] According to another aspect of this disclosure, an electronic device is provided, including:
[0013] At least one processor; and
[0014] A memory communicatively connected to the at least one processor; wherein,
[0015] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute any method in the embodiments of the present disclosure.
[0016] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute any method in the embodiments of the present disclosure.
[0017] According to another aspect of the present disclosure, there is provided a computer program product including a computer program, and the computer program implements any method in the embodiments of the present disclosure when executed by a processor.
[0018] In the technical solution of the embodiments of the present disclosure, in the above method provided by the embodiments of the present disclosure, by using a plurality of keywords included in the knowledge base and a plurality of knowledge information sets corresponding to the plurality of keywords, the knowledge information set corresponding to the keyword of the target topic information can be determined in the knowledge base, improving the efficiency and accuracy of knowledge retrieval, and by processing the target topic information and the knowledge information set through the first large model, answer information is obtained, so as to display the answer and analysis corresponding to the target topic information to the user. Therefore, the technical solution of the embodiments of the present disclosure can quickly and accurately answer and analyze the target topic information, and the user can query and learn deeper knowledge points of the target topic information.
[0019] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0021] Figure 1 is a schematic diagram of the page for taking pictures to solve problems;
[0022] Figure 2 is a schematic diagram of the page for returning the solution;
[0023] Figure 3 is a schematic flowchart of the method for answering questions based on a large model provided by an embodiment of the present disclosure;
[0024] Figure 4 is a schematic diagram of the image in the target topic information provided by an embodiment of the present disclosure;
[0025] Figure 5 is a schematic flowchart of the knowledge retrieval service provided by an embodiment of the present disclosure;
[0026] Figure 6 It is a schematic flowchart of a retrieval answer task provided by an embodiment of the present disclosure;
[0027] Figure 7 It is a schematic block diagram of a question answering device based on a large model provided by an embodiment of the present disclosure;
[0028] Figure 8 It is a schematic block diagram of a question answering device based on a large model provided by another embodiment of the present disclosure;
[0029] Figure 9 It is a block diagram of an electronic device for implementing the large model-based question answering method of the embodiments of the present disclosure. Detailed implementation manners
[0030] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0031] To facilitate understanding of the large model-based question answering method provided by the embodiments of the present disclosure, the related technologies of the embodiments of the present disclosure are described below. The following related technologies can be arbitrarily combined with the technical solutions of the embodiments of the present disclosure as optional solutions, and they all fall within the protection scope of the embodiments of the present disclosure.
[0032] In the related art, the user can click to take a photo to solve the problem and jump to a secondary page. After the user takes a photo and uploads the photo, multiple questions with high similarity to the target question are returned and listed.
[0033] In one example, when the user clicks the photo-taking problem-solving button as shown in Figure 1 and jumps to the photo-taking interface to take a photo and upload the target question, and waits for the answer to the target corresponding question to be returned, but is guided to another seemingly information-rich page as shown in Figure 2The secondary page shown above. However, such a problem-solving return hides significant uncertainties and inefficiencies. What users expect to see is the definite answer and analysis of the question, but what is returned are several similar questions. It cannot guarantee to correspond to the question uploaded by the user. At the same time, the user needs to click one by one and shuttle between each question to determine which question is the one they uploaded. If the user needs to conduct in-depth research on the knowledge points involved in the question, they also need to go to other search tools to search. This process is not only cumbersome but also exposes limitations. Although a large number of question returns are provided, it invisibly exacerbates the inefficiency of information browsing, wasting the user's precious time and being difficult to meet their desire for highly consistent, accurate question answers and in-depth knowledge point excavation.
[0034] It can be seen that when users keep clicking and scrolling down, they are very likely to encounter repeated exposure of question content. Similar question stems and similar answer analyses make the already tight time quietly pass by while browsing and comparing repeated questions, greatly reducing the efficiency and value of knowledge acquisition. The quality and depth of the question answers and analysis content cannot be guaranteed in this similar and overlapping browsing method. Returning multiple similar questions aims to attract clicks rather than accurately presenting the target question answers. At the same time, due to the lack of functions, users cannot query and learn deeper knowledge points of the questions, and finally can only obtain superficial question information fragments.
[0035] In addition, when users take pictures of questions in the question bank to solve problems, the above scheme can be implemented, but it cannot accurately answer questions outside the question bank, especially newly emerging questions, and manual collation of the question bank is required to output answers.
[0036] In the technical solution of the embodiment of the present disclosure, for the above method provided by the embodiment of the present disclosure, by using multiple keywords included in the knowledge base and multiple knowledge information sets corresponding to the multiple keywords one by one, the knowledge information set corresponding to the keyword of the target question information can be determined in the knowledge base, improving the efficiency and accuracy of knowledge retrieval. And through the first large model, the target question information and the knowledge information set are processed to obtain the answer information, so as to display the answer and analysis corresponding to the target question information to the user. Therefore, the technical solution of the embodiment of the present disclosure can quickly and accurately answer and analyze the target question information, and the user can query and learn deeper knowledge points of the target question information.
[0037] Figure 3shows a method for answering questions based on a large model provided by an embodiment of the present disclosure. This method can be applied to a question answering device based on a large model, and this device can be deployed in an electronic device. The electronic device is, for example, a single or multi-machine terminal, a server, or other processing devices. Among them, the terminal can be a user equipment (UE) such as a mobile device, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc.; the server can be a single-machine server or a server cluster. In some possible implementation manners, this method can also be implemented by a processor calling computer-readable instructions stored in a memory. As Figure 3 shown, this method can include the following steps S310 to S330.
[0038] Step S310, determine the keywords of the target question information.
[0039] Exemplarily, the target question information can include the text input by the user and / or the uploaded picture. Optionally, the user can input the text of the question in the electronic device or upload a picture including the question.
[0040] In one implementation manner, a multi-modal model can be used to determine the keywords in the target question information. For example, in some application scenarios, the target question information can include one or more types of information such as text, picture, voice, etc., and a multi-modal model can be used to process the target question information. In the embodiments of the present disclosure, the multi-modal model can understand text and / or pictures, so as to extract the keywords of the target question information. Among them, the keywords are used to reflect the theme and core content of the target question information, and the target question information can correspond to one or more keywords. For example, the keywords can include some content in the target question information or the information obtained by summarizing the target question information.
[0041] Optionally, the target question information can be any type of question. For example, the target question information can be a multiple-choice question, a true-false question, a short-answer question, etc.
[0042] Optionally, the multi-modal model can also process the text and / or picture, and convert the question in the text and / or picture into a preset code language to better extract keywords.
[0043] Optionally, the multi-modal model can adopt a GLM (Generalized Linear Model) structure, which can better understand the semantics, grammar, and context information in pictures and natural languages.
[0044] Step S320: Determine the knowledge information set corresponding to the keyword of the target question information in the knowledge base; wherein, the knowledge base includes multiple keywords and multiple knowledge information sets corresponding to the multiple keywords.
[0045] In the embodiments of the present disclosure, the knowledge information set may include information for explaining the keyword. For example, the knowledge information set may include the definition, application, etc. of the keyword, so that the subsequent knowledge information set can be used to answer the question. The knowledge base may include the mapping relationship between multiple keywords and multiple knowledge information sets. The keyword of the target question information can be retrieved from the multiple keywords in the knowledge base, and the knowledge information set corresponding to the retrieved keyword is determined as the knowledge information set corresponding to the keyword of the target question information.
[0046] Optionally, the multiple keywords included in the knowledge base and the multiple knowledge information sets corresponding to the multiple keywords may come from textbooks, teaching videos, courseware, databases, etc.
[0047] For example, if the keyword of the target question information includes Zn(OH) 2 (zinc hydroxide) reacting with ammonia water, the corresponding knowledge information set may include the following information: "Zinc hydroxide is an inorganic compound with the chemical formula Zn(OH) 2 , composed of divalent zinc and two hydroxide ions, is an amphoteric hydroxide, insoluble in water, soluble in acid, alkali solutions and ammonia water. Ammonia water is an aqueous solution of gaseous ammonia, etc.".
[0048] Step S330: Process the target question information and the knowledge information set using the first large model to obtain the answer information.
[0049] In the embodiments of the present disclosure, the first large model is used to understand the target question information and the knowledge information set, generate the answer information. The answer information may include the answer and analysis corresponding to the target question information, and the answer information can be presented to the user to complete the answering of the target question information.
[0050] Optionally, the large models in the embodiments of the present disclosure, including the first large model, the second large model, etc., may be large language models (LLMs).
[0051] The above method provided by the embodiments of the present disclosure can use multiple keywords included in the knowledge base and multiple knowledge information sets corresponding to the multiple keywords one by one to determine the knowledge information set corresponding to the keyword of the target topic information in the knowledge base, improving the efficiency and accuracy of knowledge retrieval. And by processing the target topic information and the knowledge information set through the first large model, answer information is obtained, so as to display the answer and analysis corresponding to the target topic information to the user. Therefore, the technical solution of the embodiments of the present disclosure can quickly and accurately answer and analyze the target topic information, and the user can query and learn deeper knowledge points of the target topic information.
[0052] In some embodiments, the method for answering questions based on a large model may further include:
[0053] Using a second large model to extract the keywords of the reference topic information in the question bank;
[0054] Retrieving in at least one database to obtain a knowledge information set corresponding to the keyword of the reference topic information;
[0055] Constructing a knowledge base based on the knowledge information set corresponding to the keyword of the reference topic information.
[0056] In the embodiments of the present disclosure, the reference topic information can be understood as the text of any question in the question bank. The second large model can extract the keywords of the reference topic information. The keywords are used to reflect the theme and core content of the reference topic information. One reference topic information can correspond to one or more keywords. For example, the keywords can include part of the content in the reference topic information, or can include the information obtained by summarizing the reference topic information.
[0057] Optionally, the reference topic information can be any type of question. For example, the reference topic information can be multiple-choice questions, true or false questions, short-answer questions, etc.
[0058] In the embodiments of the present disclosure, the database can include multiple knowledge information. The keywords of the reference topic information can be retrieved in the database, and the knowledge information set corresponding to the keyword can be recalled from the multiple knowledge information in the database. Thus, the keywords and the knowledge information set corresponding to the keywords can be added to the knowledge base. The knowledge base can include multiple keywords and multiple knowledge information sets corresponding to the multiple keywords one by one.
[0059] Optionally, the number of databases can be multiple, and the knowledge information in each database can be different. Thus, the knowledge information recalled from each database can also be different. By summarizing the knowledge information recalled from each database, a knowledge information set is obtained.
[0060] Optionally, the number of reference topic information can be multiple. The second large model can extract keywords of each reference topic information respectively, so as to retrieve a set of knowledge information corresponding to the keywords of each reference topic information, and thus a knowledge base can be constructed based on the set of knowledge information corresponding one by one to the keywords of each reference topic information.
[0061] Optionally, based on the ERNIE (Enhanced Representation through kNowledgeIntEgration, a knowledge-enhanced continuous learning semantic understanding framework) pre-training framework, the second large model can be obtained. Through a series of model compression and optimization techniques, while maintaining high accuracy, the size and computational complexity of the second large model are significantly reduced.
[0062] According to the above embodiments, the second large model can be used to extract keywords of reference topic information, and a set of knowledge information corresponding to the keywords of reference topic information can be retrieved in the database, so as to construct a knowledge base that maps keywords to sets of knowledge information. This provides a retrieval basis for determining the set of knowledge information corresponding to the keywords of the target topic information in the knowledge base later, and further improves the efficiency and accuracy of knowledge retrieval.
[0063] In some embodiments, the set of knowledge information includes general knowledge information. At least one database includes a general information database. Correspondingly, retrieving the set of knowledge information corresponding to the keywords of reference topic information in at least one database includes:
[0064] Retrieving multiple candidate knowledge information in the general information database based on the keywords of reference topic information;
[0065] Determining the general knowledge information corresponding to the keywords of reference topic information among the multiple candidate knowledge information based on at least one of the citation times, search volume, and publication date of the candidate knowledge information.
[0066] In the embodiments of the present disclosure, the general information database can include multiple general knowledge information. The general knowledge information has characteristics such as universal applicability, wide dissemination, and easy understanding. The keywords of reference topic information can be retrieved in the general information database, so as to recall multiple candidate knowledge information corresponding to the keywords among the multiple general knowledge information in the general information database. Then, the multiple candidate knowledge information can be screened based on at least one of the citation times, search volume, and publication date of the candidate knowledge information, and the candidate knowledge information that meets the preset conditions is determined as the general knowledge information corresponding to the keywords of reference topic information. The set of knowledge information corresponding to the keywords can be obtained based on the general knowledge information corresponding to the keywords.
[0067] Exemplarily, candidate knowledge information with a higher citation count, a higher search volume, and a more recent publication date can be determined as the general knowledge information corresponding to the keywords of the reference topic information. It is also possible to set weights for the citation count, search volume, and publication date of the candidate knowledge information respectively, calculate the scores of each candidate knowledge information, and determine the candidate knowledge information with a higher score as the general knowledge information corresponding to the keywords of the reference topic information.
[0068] Optionally, the number of general information databases can be one or more. For example, the general information database can include an encyclopedia information database based on a knowledge graph, or a library database including a vast amount of literature data.
[0069] According to the above embodiments, multiple candidate knowledge information can be retrieved from the general information database based on the keywords of the reference topic information, and the general knowledge information corresponding to the keywords can be determined among the multiple candidate knowledge information, thus ensuring the accuracy and comprehensiveness of the knowledge information set corresponding to the keywords.
[0070] In some embodiments, the knowledge information set includes subject knowledge information. At least one database includes a subject knowledge database. Correspondingly, retrieving the knowledge information set corresponding to the keywords of the reference topic information from at least one database includes:
[0071] Encoding the reference topic information using a text encoding model to obtain a topic vector corresponding to the reference topic information;
[0072] Based on the vector similarity between the topic vector and the knowledge vectors corresponding to each subject knowledge information in the subject knowledge database, determine the subject knowledge information corresponding to the keywords of the reference topic information in the subject knowledge database.
[0073] Optionally, the subject knowledge database can include subject knowledge information of multiple subjects. Each subject knowledge information can include in-depth and professional knowledge information in each subject, and the subject knowledge information in each subject can be obtained from corresponding textbooks, teaching videos, and courseware. The subject knowledge information with a higher vector similarity between the topic vector and the knowledge vector can be determined as the subject knowledge information corresponding to the keywords of the reference topic information, so that the knowledge information set corresponding to the keywords can be obtained based on the subject knowledge information corresponding to the keywords.
[0074] Optionally, the text encoding model can also encode each subject knowledge information to obtain a knowledge vector corresponding to each subject knowledge information, so as to use the topic vector and the knowledge vector to determine the subject knowledge information corresponding to the keywords of the reference topic information in the subject knowledge database.
[0075] Optionally, the subject knowledge database may include indication vectors corresponding to subject knowledge information of multiple subjects. After encoding each subject knowledge information using a text encoding model, the encoded knowledge vectors can be stored in each subject knowledge database.
[0076] Optionally, the text encoding model may adopt a text representation model in related technologies, which can convert text into a vector form represented by numerical values.
[0077] Optionally, multiple subject knowledge information with high vector similarity can be sorted according to custom rules (such as timestamp information and / or similarity calculation values), so as to determine the subject knowledge information corresponding to the keywords of the reference question information among the multiple subject knowledge information with high vector similarity.
[0078] According to the above embodiments, the reference question information can be encoded using a text encoding model, and based on the vector similarity between the question vector and the knowledge vectors corresponding to each subject knowledge information in the subject knowledge database, the subject knowledge information corresponding to the keywords of the reference question information can be determined in the subject knowledge database, thus ensuring the professionalism and comprehensiveness of the knowledge information set corresponding to the keywords.
[0079] In some embodiments, step S310, determining the keywords of the target question information using a multi-modal model, includes:
[0080] Interpret the image in the target question information using a multi-modal model to obtain the description information of the image;
[0081] Extract the text in the image using a multi-modal model;
[0082] Perform keyword extraction based on the description information of the image and the text in the image to obtain the keywords of the target question information.
[0083] In the embodiments of the present disclosure, the image in the target question information can be interpreted using a multi-modal model to obtain the description information of the image, and the text in the image can be extracted using a multi-modal model, taking into account both text and image information for keyword extraction.
[0084] Optionally, the subject corresponding to the target question information can be determined in advance, and the multi-modal model can interpret the image in the target question information based on the subject corresponding to the target question information to obtain the description information of the image.
[0085] In one example, the image in the target question information can be as Figure 4 shown, and obtaining the description information of the image using a multi-modal model may include:
[0086] "The picture contains a single-choice chemistry question asking about Zn(OH) 2The standard equilibrium constant expression for the reaction dissolved in ammonia water, and the options involve combinations of K (potassium) and K[Zn(OH) 2 , K[Zn(NH3) 4 2+ and other chemical formulas and constants (specific options are not shown), and it is required to select the correct answer from the given options.
[0087] In addition, the text in the image extracted by the multimodal model may include: "24. Zn(OH) 2 The reaction of the precipitate dissolved in ammonia water is..."
[0088] According to the above embodiments, the multimodal model can be used to interpret the image in the target question information and extract the text in the image, while taking into account both text and image types of information to extract keywords, improving the accuracy of the keywords.
[0089] In some embodiments, the method for answering questions based on a large model may further include:
[0090] Based on the sample question information, perform supervised fine-tuning and / or reinforcement learning alignment on the pre-trained model to obtain a multimodal model.
[0091] In the embodiments of the present disclosure, the sample question information may include sample images, description information of the sample images determined in the sample images, and the text extracted and labeled in the sample images, so that based on the above information, supervised fine-tuning and / or reinforcement learning alignment can be performed on the pre-trained model to obtain a multimodal model.
[0092] Optionally, supervised fine-tuning can adopt SFT (Supervised Fine-Tuning, supervised fine-tuning), and reinforcement learning alignment can adopt DPO (Direct Preference Optimization, direct preference optimization), thereby further optimizing the effect of the pre-trained model in the text extraction task to obtain a multimodal model.
[0093] In an example, for the question as Figure 4 shown, if a general pre-trained model is directly used for recognition, only text can be extracted, and specific information such as superscripts, subscripts, numerators, and denominators of formulas cannot be recognized. By using sample question information with formulas and annotating the sample question information with a specific code language, performing supervised fine-tuning and / or reinforcement learning alignment on the pre-trained model can enable the multimodal model to recognize complex formulas and represent specific information such as superscripts, subscripts, numerators, and denominators in the formulas through a specific code language.
[0094] According to the above embodiments, by performing supervised fine-tuning and / or reinforcement learning alignment on the pre-trained model through sample question information, the consistency and accuracy of the multimodal model in interpreting images and extracting text can be improved.
[0095] In some embodiments, in step S330, the first large model is used to process the target question information and the knowledge information set to obtain answer information, including:
[0096] The image in the target question information and the text in the image are input into the first large model to obtain the answer and analysis corresponding to the target question information.
[0097] In the embodiments of the present disclosure, the first large model can understand the image in the target question information and the text in the image, so as to more accurately understand the target question information, and then more precisely combine the knowledge information set to obtain the answer and analysis corresponding to the target question information.
[0098] Optionally, the subject corresponding to the target question information can be determined in advance, and the subject corresponding to the target question information, the image in the target question information, and the text in the image are input into the first large model, so as to obtain more accurate answers and analyses.
[0099] According to the above embodiments, when the target question information includes an image, the first large model can understand the image in the target question information and the text in the image, further improving the accuracy of the answer and analysis.
[0100] To more clearly understand the technical solution of the embodiments of the present disclosure, a specific application example is provided below. In this application example, the method for answering questions based on a large model includes two parts: constructing a knowledge retrieval service and a retrieval answering task.
[0101] Figure 5 The flowchart of constructing the knowledge retrieval service is shown, as Figure 5 shown, constructing the knowledge retrieval service may include:
[0102] Step 1: Use the second large model to extract keywords for the reference question information;
[0103] In one example, the reference question information may be the question information as Figure 4 shown.
[0104] Correspondingly, the keyword extraction results may include:
[0105] 1. Zn(OH) 2 reacting with ammonia water, 2. standard equilibrium constant expression, 3. K and K[Zn(OH) 2 , K[Zn(NH3) 4 2+.
[0106] Step 2: Call at least one database for each keyword of the reference topic information to recall knowledge information, and obtain a set of knowledge information corresponding to the keywords of the reference topic information;
[0107] In one example, the set of knowledge information recalled by the general information database may include:
[0108] "Zinc hydroxide is an inorganic compound with the chemical formula Zn(OH) 2 , composed of divalent zinc and two hydroxide ions, is an amphoteric hydroxide, insoluble in water, soluble in acid, alkali solutions and ammonia water. Ammonia water is an aqueous solution of gaseous ammonia,... The standard equilibrium constant is the equilibrium constant calculated from standard thermodynamic functions, also known as the thermodynamic equilibrium constant."
[0109] The set of knowledge information recalled by the subject knowledge database may include:
[0110] "The negative electrode of the voltaic cell is zinc (Zn), the positive electrode is copper (Cu), and the electrolyte is sulfuric acid (H2SO4)..."
[0111] Step 3: Use a text encoding model to encode (Embedding) the reference topic information to obtain a topic vector corresponding to the reference topic information;
[0112] In one example, the reference topic information can be encoded in vector form. For example, the topic vector is [0.0233, 0.13343,..., 0.0912].
[0113] Step 4: Use a text encoding model to encode each subject knowledge information, obtain the knowledge vectors corresponding to each subject knowledge information, and store them in the subject knowledge database;
[0114] In one example, the subject knowledge information is encoded as {"time": "2024-11-09", "text": "Volatile, ammonia water easily volatilizes ammonia gas, and the evaporation rate increases with increasing temperature and prolonging the storage time, and the evaporation amount also increases with increasing concentration..."}, indicating that the timestamp is 2024-11-09, and the text includes "Volatile, ammonia water easily volatilizes ammonia gas, and the evaporation rate increases with increasing temperature and prolonging the storage time, and the evaporation amount also increases with increasing concentration...". The knowledge vector is encoded as [0.56437, 0.84534,..., 0.08568].
[0115] Step 5: By calculating the vector similarity between the topic vector and the knowledge, retrieve the results with higher similarity, and sort them according to custom rules (such as timestamp information or similarity calculation values) to obtain several subject knowledge information.
[0116] After building the knowledge retrieval service, users can upload images, and the electronic device can perform retrieval and answering tasks. Figure 6 The flowchart of the retrieval and answering task is shown, as Figure 6 shown, the retrieval and answering task may include:
[0117] Step 1: Obtain the image and the corresponding subject in the target question information;
[0118] In one example, the image in the target question information is as Figure 4 shown, and the subject is chemistry.
[0119] Step 2: Interpret the image through a multimodal model to obtain the description information of the image, and extract the text in the image, taking into account both text and image information;
[0120] In one example, as Figure 4 shown, the description information of the image may include: The picture contains a multiple-choice chemistry question asking for the standard equilibrium constant expression of the reaction of Zn(OH) 2 dissolved in ammonia water, and the options involve combinations of chemical formulas and constants such as K and K[Zn(OH) 2 , K[Zn(NH3) 4 2+ etc. (specific options are not shown), and it is required to select the correct answer from the given options.
[0121] The text extracted from the image may include: "24. Zn(OH) 2 The reaction of the precipitate dissolved in ammonia water is ……".
[0122] Step 3: Invoke the knowledge retrieval service to obtain the knowledge information set corresponding to the keywords of the target question information;
[0123] Step 4: Inject the image in the target question information, the text in the image, and the knowledge information set into the first large model to generate answers and explanations.
[0124] It can be seen that for the problem-solving method based on a large model provided by the present disclosure, first, a multimodal model can be used to accurately identify the images in the target problem information. Secondly, by using multiple keywords included in the knowledge base and multiple knowledge information sets corresponding to the multiple keywords one by one, the knowledge information set corresponding to the keywords of the target problem information can be determined in the knowledge base, improving the efficiency and accuracy of knowledge retrieval. And by processing the target problem information and the knowledge information set through the first large model, solution information is obtained, so as to display the answer and analysis corresponding to the target problem information to the user. Therefore, the technical solution of the embodiments of the present disclosure can quickly and accurately answer and analyze the target problem information, and the user can query and learn deeper knowledge points of the target problem information.
[0125] According to an embodiment of the present disclosure, the present disclosure also provides an apparatus for solving problems based on a large model. Figure 7 The schematic block diagram of an apparatus for solving problems based on a large model provided by an embodiment of the present disclosure is shown, as Figure 7 shown, the apparatus includes:
[0126] A keyword determination module 710, configured to determine the keywords of the target problem information;
[0127] A knowledge determination module 720, configured to determine the knowledge information set corresponding to the keywords of the target problem information in the knowledge base; wherein, the knowledge base includes multiple keywords and multiple knowledge information sets corresponding to the multiple keywords;
[0128] A solution module 730, configured to process the target problem information and the knowledge information set by using the first large model to obtain solution information.
[0129] In some embodiments, as Figure 8 shown, the apparatus for solving problems based on a large model further includes a knowledge base construction module 810, and the knowledge base construction module 810 is configured to:
[0130] Use the second large model to extract the keywords of the reference problem information in the question bank;
[0131] Retrieve in at least one database to obtain the knowledge information set corresponding to the keywords of the reference problem information;
[0132] Construct a knowledge base based on the knowledge information set corresponding to the keywords of the reference problem information.
[0133] In some embodiments, the knowledge information set includes general knowledge information, and the knowledge base construction module 810 is specifically configured to:
[0134] Based on the keywords of the reference problem information, retrieve multiple candidate knowledge information in the general information database;
[0135] Determine general knowledge information corresponding to the keywords of the reference topic information from multiple pieces of candidate knowledge information based on at least one of the citation count, search volume, and publication date of the candidate knowledge information.
[0136] In some embodiments, the knowledge information set includes subject knowledge information, and the knowledge base construction module 810 is specifically configured to:
[0137] Encode the reference topic information using a text encoding model to obtain a topic vector corresponding to the reference topic information;
[0138] Based on the vector similarity between the topic vector and the knowledge vectors corresponding to each piece of subject knowledge information in the subject knowledge database, determine the subject knowledge information corresponding to the keywords of the reference topic information in the subject knowledge database.
[0139] In some embodiments, the keyword determination module 710 is specifically configured to:
[0140] Interpret the image in the target topic information using a multimodal model to obtain a description information of the image;
[0141] Use the multimodal model to extract the text in the image;
[0142] Extract keywords based on the description information of the image and the text in the image to obtain the keywords of the target topic information.
[0143] In some embodiments, as Figure 8 shown, the topic answering device based on the large model further includes:
[0144] A model training module 820, configured to perform supervised fine-tuning and / or reinforcement learning alignment on a pre-trained model based on sample topic information to obtain a multimodal model.
[0145] In some embodiments, the answering module 730 is specifically configured to:
[0146] Input the image in the target topic information and the text in the image into the first large model to obtain the answer and analysis corresponding to the target topic information.
[0147] For the specific functions and examples of each module and sub-module of the device in the embodiments of the present disclosure, reference may be made to the relevant descriptions of the corresponding steps in the above method embodiments, which will not be elaborated here.
[0148] In the technical solution of the present disclosure, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0149] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0150] Figure 9 FIG. shows a schematic block diagram of an exemplary electronic device 900 that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, personal digital assistants, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0151] As Figure 9 shown, the device 900 includes a computing unit 901 that can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 907 into a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the device 900 can also be stored. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0152] A plurality of components in the device 900 are connected to the I / O interface 905, including: an input unit 906, such as a keyboard, a mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a magnetic disk, an optical disk, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the device 900 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0153] The computing unit 901 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), various dedicated Artificial Intelligence (AI) computing chips, various computing units running machine learning model algorithms, a Digital Signal Processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 executes the various methods and processes described above, such as the method for solving problems based on a large model. For example, in some embodiments, the method for solving problems based on a large model can be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as the storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the computing unit 901, one or more steps of the method for solving problems based on a large model described above can be executed. Alternatively, in other embodiments, the computing unit 901 can be configured to execute the method for solving problems based on a large model in any other suitable manner (e.g., by means of firmware).
[0154] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.
[0155] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code may execute entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine or entirely on the remote machine or server.
[0156] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0157] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a cathode ray tube (CRT) or a liquid crystal display (LCD) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0158] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0159] A computer system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The relationship between the client and the server is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, or a server of a distributed system, or a server incorporating a blockchain.
[0160] It should be understood that various forms of the processes shown above may be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure may be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this is not limited herein.
[0161] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A method for solving a problem based on a large model, comprising: Determine the keywords of the target topic information; Determine a knowledge information set corresponding to the keyword of the target topic information in a knowledge base; wherein the knowledge base includes a plurality of keywords and a plurality of knowledge information sets corresponding to the plurality of keywords; The target question information and the knowledge information set are processed using the first large model to obtain answer information.
2. The method according to claim 1, further comprising: Using the second largest model, extract keywords from the reference question information in the question bank; Retrieving in at least one database a knowledge information set corresponding to the keywords of the reference topic information; The knowledge base is constructed based on a set of knowledge information corresponding to the keywords of the reference topic information.
3. The method according to claim 2, wherein: The knowledge information set includes general knowledge information; The step of retrieving a knowledge information set corresponding to the keywords of the reference title information from at least one database includes: Based on the keywords of the reference topic information, a plurality of candidate knowledge information is retrieved from a general information database; Based on at least one of the number of citations, the search volume, and the release date of the candidate knowledge information, general knowledge information corresponding to the keyword of the reference title information is determined from the plurality of candidate knowledge information.
4. The method according to claim 2 or 3, wherein the knowledge information set includes subject knowledge information; The step of retrieving a knowledge information set corresponding to the keywords of the reference title information from at least one database includes: Encoding the reference title information using a text encoding model to obtain a title vector corresponding to the reference title information; Based on the vector similarity between the title vector and the knowledge vectors corresponding to each subject knowledge information in the subject knowledge database, the subject knowledge information corresponding to the keyword of the reference title information is determined in the subject knowledge database.
5. The method according to any one of claims 1 to 4, wherein: The keywords for determining the target topic information include: Using a multimodal model to interpret the image in the target title information to obtain description information of the image; Extracting text from the image using the multimodal model; Keywords are extracted based on the description information of the image and the text in the image to obtain keywords of the target title information.
6. The method according to claim 5, further comprising: Based on the sample topic information, supervised fine-tuning and / or reinforcement learning alignment are performed on the pre-trained model to obtain the multimodal model.
7. The method according to any one of claims 1 to 6, wherein: The method of using the first large model to process the target question information and the knowledge information set to obtain answer information includes: The image in the target question information and the text in the image are input into the first large model to obtain the answer and analysis corresponding to the target question information.
8. A large model-based question-solving device, comprising: A keyword determination module, used to determine keywords of target topic information; A knowledge determination module, used to determine a knowledge information set corresponding to a keyword of the target topic information in a knowledge base; wherein the knowledge base includes a plurality of keywords and a plurality of knowledge information sets corresponding to the plurality of keywords; The answer module is used to process the target question information and the knowledge information set using the first large model to obtain answer information.
9. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-7.
11. A computer program product, comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 7.