Question answering method, optimization method, device and question answering system for photovoltaic cell experiment
By identifying user question texts and combining graph walking algorithms and multi-agent technology to generate knowledge graphs and datasets of photovoltaic cell materials, the large-model question-answering process is optimized, solving the problem of insufficient accuracy in question-answering and experimental optimization of large models in photovoltaic cell research and development, and achieving efficient and accurate knowledge updating and experimental optimization.
Patent Information
- Application Number
- CN202510769155.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2045-06-09
AI Technical Summary
The existing large models lack accuracy in question-answering and experimental optimization in photovoltaic cell research and development, and the cost of knowledge updating is high, which is incompatible with the rapid iteration of photovoltaic cell research and development.
By recognizing the question text entered by the user, combining graph walking algorithms and multi-agent technology to generate a knowledge graph and domain dataset of photovoltaic cell materials, the question-answering process of the large model is optimized, including keyword screening, semantic verification and information extraction, to build a large model in the photovoltaic cell field and improve the efficiency and accuracy of knowledge updating.
It improves the question-answering accuracy and experimental optimization capabilities of large models in the field of photovoltaic cell research and development, reduces the cost of knowledge updating, and enhances the matching degree between large models and the photovoltaic cell field.
Smart Images

Figure CN120653740A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of large-scale model design applications, and in particular to a question-and-answer method, optimization method, device, and question-and-answer system for photovoltaic cell experiments. Background Art
[0002] As research on large models continues to grow, their application in the materials field, particularly in photovoltaic cell development, has also seen significant growth. Due to the high cost of pre-training large models, most large models currently used in photovoltaic cell development focus on fine-tuning during pre-training. This involves fine-tuning the general large model using a domain-specific corpus to acquire foundational domain knowledge.
[0003] The general application of large models for photovoltaic cell research and development is to use large models, combined with domain knowledge graphs, to give photovoltaic cell experimental suggestions, so that corresponding photovoltaic cell research experiments can be carried out according to the suggestions. However, this application method has the following problems: (1) The current general large models have limited training data for the field of photovoltaic cell research and development, resulting in a large amount of resources being consumed for fine-tuning all parameters each time the knowledge of the large models is updated. This is not only costly but also incompatible with the current situation of rapid iteration of photovoltaic cell research and development, resulting in untimely knowledge updates and affecting the accuracy of large model question answering; (2) Although the current large models can obtain the ability to answer questions in the field of photovoltaic cell research and development through fine-tuning, after the experimental results are obtained by conducting experiments on the experimental suggestions for photovoltaic cell research and development based on the large models, the analysis and reasoning of the experimental result data are insufficient, and there is also a lack of feedback optimization of the experimental suggestions based on the experimental results, resulting in the lack of experimental analysis and optimization capabilities of the large models. The accuracy of the large models in experimental analysis and optimization is not high. Therefore, how to improve the accuracy of large models in question answering and experimental optimization in the field of photovoltaic cell research and development is still a technical problem that needs to be solved urgently in the existing technology. Summary of the Invention
[0004] The present application provides a question-answering method, an optimization method, an apparatus, and a question-answering system for photovoltaic cell experiments to solve the technical problem of insufficient accuracy of question-answering and experimental optimization of existing large models in the field of photovoltaic cell research and development.
[0005] According to a first aspect of the embodiments of the present application, a question-answering method for a photovoltaic cell experiment is provided, comprising:
[0006] Identify the first question text input by the user and obtain domain keywords, technical keywords and keyword weight ratios;
[0007] Based on the field keywords and the technical keywords, a preset photovoltaic cell material knowledge graph is searched and screened in combination with a graph walking algorithm to obtain multiple field search paths; wherein the photovoltaic cell material knowledge graph is generated based on the extraction of photovoltaic cell field data;
[0008] According to the multiple search paths, a plurality of domain knowledge texts are obtained from the photovoltaic cell material knowledge graph; wherein the ratio of texts satisfying domain keywords to texts satisfying technical keywords in the multiple domain knowledge texts conforms to the keyword weight ratio;
[0009] Based on the first question text and the multiple domain knowledge texts, questions are asked to the preset photovoltaic cell domain large model to obtain a first answer text; wherein, the photovoltaic cell domain large model is constructed based on a preset domain data set; the preset domain data set is extracted and constructed by multiple intelligent agents from photovoltaic cell domain data.
[0010] This application first identifies the first question text entered by the user, obtains different types of keywords and keyword weight ratios, and then searches and filters the photovoltaic cell material knowledge graph based on the keywords to obtain multiple field search paths, thereby obtaining multiple field knowledge texts from the photovoltaic cell material knowledge graph according to the multiple field search paths, obtains search keywords by identifying the user's question, and then obtains the prior field knowledge for the large model question based on the keywords. Based on the prior field knowledge, the answer of the large model can be limited to a specific range, reducing the generation of large model hallucinations, thereby improving the accuracy of the large model's question and answer in the field of photovoltaic cell research and development when the first question text and multiple field knowledge texts are combined to ask the photovoltaic cell field large model to obtain the first answer text; at the same time, by extracting photovoltaic cell field data and multi-agent extraction, a photovoltaic cell material knowledge graph and a preset field data set are constructed, thereby obtaining a large model in the field of photovoltaic cells. Compared with the existing technology, by increasing the knowledge graph that can obtain prior field knowledge and the field data set for training the large model, the knowledge update cost of the large model can be reduced, the knowledge update efficiency of the large model can be improved, and the accuracy of the large model's question and answer in the field of photovoltaic cell research and development can be improved.
[0011] In certain embodiments of the present application, the field keywords and the technical keywords are combined with a graph walking algorithm to search and screen the preset photovoltaic cell material knowledge graph to obtain multiple field search paths, specifically including:
[0012] Search the photovoltaic cell material knowledge graph according to the domain keywords to obtain a first domain subgraph;
[0013] Filtering the first domain subgraph according to the technical keywords to obtain multiple first search paths;
[0014] The semantic coherence of the plurality of first search paths is verified based on a graph walk algorithm, and filtering is performed according to the verification result to obtain a plurality of domain search paths.
[0015] This application first searches the photovoltaic cell material knowledge graph through domain keywords to obtain the first domain subgraph, which can first locate the domain where the keyword is located and narrow the search scope, and then filter the first domain subgraph in combination with technical keywords to obtain multiple first search paths, which can further narrow the scope and focus on the corresponding technology of the keyword, so as to improve the matching degree between the multiple domain search paths and the current user's question when verifying its semantic coherence based on the graph walking algorithm and filtering to obtain multiple domain search paths, and improve the accuracy of the prior domain knowledge when obtaining the prior domain knowledge based on multiple domain search paths.
[0016] In certain embodiments of the present application, the photovoltaic cell material knowledge graph is generated based on the extraction of photovoltaic cell field data, specifically including:
[0017] Based on the preset graph mode and combined with the preset first model, the photovoltaic cell field data is screened by preset keywords to obtain multiple first field text fragments;
[0018] According to the different keyword weights of the preset keywords in each first-domain text segment and the number of citations of each first-domain text segment, the plurality of first-domain text segments are screened to obtain a plurality of second-domain text segments;
[0019] According to the graph pattern and based on the first large model, knowledge extraction is performed on the plurality of second domain text fragments to obtain a plurality of domain entity subgraphs;
[0020] Combined with the semantic similarity between the multiple domain entity subgraphs, the multiple domain entity subgraphs are disambiguated and merged to obtain a photovoltaic cell material knowledge graph.
[0021] This application first filters the photovoltaic cell field data based on the preset graph pattern and the preset first model through preset keywords. It can efficiently obtain multiple first-field text fragments that conform to the graph pattern based on the first model, and then filter according to the weight of the preset keywords and the number of citations of the first-field text fragments, and can exclude text fragments with low relevance to the photovoltaic cell field. Then, according to the graph pattern, knowledge extraction is performed based on the first model to obtain multiple domain entity subgraphs, and disambiguation and merging are performed according to the mutual semantic similarity, which can improve the accuracy of the knowledge graph and the quality of the domain knowledge therein.
[0022] In certain embodiments of the present application, the preset domain dataset is extracted and constructed by multi-agents from photovoltaic cell domain data, specifically including:
[0023] By using the information extraction agent in the multi-agent, the photovoltaic cell field data is structuredly parsed to extract a plurality of first-field information, and contextual reasoning is performed on the plurality of first-field information to construct a plurality of information association relationships;
[0024] By using a quality verification agent in the multi-agents, the plurality of first-domain information is cross-validated and screened by combining the photovoltaic cell material knowledge graph with a comparison of the plurality of information association relationships to obtain a plurality of second-domain information;
[0025] Performing hierarchical summarization on the plurality of second domain information by a document summarizing agent in the multi-agent, and generating a plurality of multi-level summary information corresponding to the plurality of second domain information;
[0026] The plurality of second domain information are integrated with the corresponding plurality of multi-level summary information, and matched and integrated with a plurality of preset research categories to obtain a preset domain data set.
[0027] This application first performs structured analysis and extraction of photovoltaic cell field data based on the information extraction agent in the multi-agent to obtain multiple first-field information, and performs contextual reasoning to construct multiple information association relationships. Then, through the quality verification agent in the multi-agent, combined with the photovoltaic cell material knowledge graph and the comparison of multiple information association relationships, cross-validation and screening are performed to obtain multiple second-field information. It can screen out information with low correlation with photovoltaic cell materials in the first-field information and improve the accuracy of the information. Then, the document summary agent in the multi-agent performs hierarchical summary to generate multiple multi-level summary information, and then matches and integrates with multiple second-field information and multiple research categories to obtain a preset field data set. It can add a summary dimension to the second field information, improve information density, and thus improve the matching degree between the preset field data set and the photovoltaic cell field. Therefore, when constructing a large model of the photovoltaic cell field based on the preset field data set, the large model is more matched with the photovoltaic cell field, thereby improving the accuracy of the large model when conducting photovoltaic cell experimental question and answer.
[0028] According to a second aspect of the embodiments of the present application, a method for optimizing a photovoltaic cell experiment is provided, comprising:
[0029] Obtaining a first experimental result of a user; wherein the first experimental result is obtained by the user performing an experiment on a first experimental suggestion; and the first experimental suggestion is obtained by a photovoltaic cell experiment question-and-answer method described in this application;
[0030] According to the first experimental result and the first experimental suggestion, a question is asked to the preset photovoltaic cell experimental model to obtain a second answer text, and the first experimental suggestion is optimized according to the second answer text; wherein, the photovoltaic cell experimental model is constructed based on the preset experimental question and answer data set.
[0031] This application first obtains the first experimental result of the experiment based on the first experimental suggestion obtained by the user based on the question and answer of the photovoltaic cell experiment, and then asks questions to the photovoltaic cell experiment big model according to the first experimental result and the first experimental suggestion, and optimizes the first experimental suggestion according to the question result. Compared with the existing technology, by training a big model that can analyze the experimental results and adding a process for optimizing the experimental suggestions according to the big model, it is possible to realize the analysis and reasoning of the experimental results, as well as the feedback optimization of the experimental suggestions, thereby improving the accuracy of the experimental optimization of the big model in the field of photovoltaic cell research and development.
[0032] In certain embodiments of the present application, the photovoltaic cell experimental large model is constructed based on a preset experimental question-and-answer dataset, specifically including:
[0033] Acquire an initial experimental document dataset, and convert the title of each figure and the text content of each paragraph of each initial document in the initial experimental document dataset into vector representations, thereby obtaining multiple figure title vectors and multiple paragraph text vectors for each initial document;
[0034] Traversing each initial document, calculating a first similarity between each figure title vector in the current initial document and each paragraph text vector in the current initial document, and screening the initial documents in the initial experimental document dataset based on the first similarity to obtain a first experimental document dataset;
[0035] Based on the pre-trained graph-text matching model, the text content of each paragraph and each graph of each first document in the first experimental document dataset are converted into cross-modal vector representations, thereby obtaining multiple image vectors and multiple paragraph vectors for each first document.
[0036] Traversing each first document, calculating a second similarity between each image vector in the current first document and each paragraph vector in the current first document during the traversal, and combining image vectors and paragraph vectors whose second similarity exceeds a preset first threshold into a first image-text vector pair;
[0037] Combining a plurality of first image-text vector pairs of all first documents to obtain a first image-text vector set;
[0038] Ask questions to a preset second largest model based on the first image-text vector set, construct corresponding question-answer pairs for each vector pair in the first image-text vector set, and construct a preset experimental question-answer dataset based on the question-answer pairs;
[0039] According to the preset experimental question-answering dataset, the initial photovoltaic cell experimental large model is adjusted and trained to obtain the photovoltaic cell experimental large model.
[0040] This application first obtains an initial experimental document data set and converts it to obtain multiple chart title vectors and multiple paragraph text vectors for each initial document, then screens the initial documents based on the first similarity to obtain the first experimental document data set, which can ensure the correlation between the charts and texts of each initial document in the first experimental document data set, and then converts multiple image vectors and multiple paragraph vectors for each first document in the first experimental document data set based on the pre-trained chart-text matching model, and then obtains a first image-text vector set based on the second similarity screening combination, which can ensure the correlation between the image and paragraph text of each first image-text vector pair in the first image-text vector set, so that when asking questions to the second large model according to the first image-text vector set to construct the corresponding question-answer pair for each vector pair, it is ensured that the generated question matches the current vector pair, so that when constructing the photovoltaic cell experimental large model according to the preset experimental question-answer data set, the accuracy of the photovoltaic cell experimental large model is improved.
[0041] In certain embodiments of the present application, the step of asking questions to a preset second large model based on the first image-text vector set, constructing a corresponding question-answer pair for each vector pair of the first image-text vector set, and constructing a preset experimental question-answer dataset based on the question-answer pairs specifically includes:
[0042] For each vector pair of the first image-text vector set, ask the second model a question based on a preset question-answer prompt word to obtain a corresponding question for each vector pair;
[0043] Combining each vector pair of the first image-text vector set with the corresponding question to obtain a question-answer pair corresponding to each vector pair;
[0044] Combine all question-answer pairs to obtain the preset experimental question-answering dataset.
[0045] This application first constructs corresponding questions for each vector pair in the first image-text vector set based on the question-answer prompt words, then performs corresponding combinations to obtain question-answer pairs, and then constructs an experimental question-answering dataset, which can ensure the correspondence between questions and vector pairs, thereby obtaining diverse questions based on the diversity of vector pairs, thereby improving the data richness of the preset experimental question-answering dataset.
[0046] According to a third aspect of the embodiments of the present application, there is provided a question-answering device for photovoltaic cell experiments, comprising a question recognition module, a graph retrieval module, a knowledge acquisition module, and a model questioning module;
[0047] The question recognition module is used to recognize the first question text input by the user and obtain the domain keywords, technical keywords and keyword weight ratios;
[0048] The graph retrieval module is used to search and filter a preset photovoltaic cell material knowledge graph based on the field keywords and the technical keywords in combination with a graph walking algorithm to obtain multiple field search paths; wherein the photovoltaic cell material knowledge graph is generated based on the extraction of photovoltaic cell field data;
[0049] The knowledge acquisition module is configured to acquire a plurality of domain knowledge texts from the photovoltaic cell material knowledge graph according to the plurality of search paths; wherein the ratio of texts satisfying domain keywords to texts satisfying technical keywords in the plurality of domain knowledge texts conforms to the keyword weight ratio;
[0050] The model questioning module is used to ask questions to a preset photovoltaic cell field large model based on the first question text and the multiple domain knowledge texts to obtain a first answer text; wherein, the photovoltaic cell field large model is constructed based on a preset domain data set; the preset domain data set is extracted and constructed by multiple intelligent agents from photovoltaic cell field data.
[0051] In certain embodiments of the present application, the graph search module includes a keyword search unit, a keyword screening unit, and a path verification unit;
[0052] The keyword search unit is configured to search the photovoltaic cell material knowledge graph based on the field keywords to obtain a first field subgraph;
[0053] The keyword screening unit is configured to screen the first domain subgraph according to the technical keywords to obtain a plurality of first search paths;
[0054] The path verification unit is used to verify the semantic coherence of the multiple first search paths based on a graph walk algorithm, and filter according to the verification results to obtain multiple domain search paths.
[0055] In certain embodiments of the present application, the photovoltaic cell material knowledge graph is generated based on the extraction of photovoltaic cell field data, specifically including:
[0056] Based on the preset graph mode and combined with the preset first model, the photovoltaic cell field data is screened by preset keywords to obtain multiple first field text fragments;
[0057] According to the different keyword weights of the preset keywords in each first-domain text segment and the number of citations of each first-domain text segment, the plurality of first-domain text segments are screened to obtain a plurality of second-domain text segments;
[0058] According to the graph pattern and based on the first large model, knowledge extraction is performed on the plurality of second domain text fragments to obtain a plurality of domain entity subgraphs;
[0059] Combined with the semantic similarity between the multiple domain entity subgraphs, the multiple domain entity subgraphs are disambiguated and merged to obtain a photovoltaic cell material knowledge graph.
[0060] In certain embodiments of the present application, the preset domain dataset is extracted and constructed by multi-agents from photovoltaic cell domain data, specifically including:
[0061] By using the information extraction agent in the multi-agent, the photovoltaic cell field data is structuredly parsed to extract a plurality of first-field information, and contextual reasoning is performed on the plurality of first-field information to construct a plurality of information association relationships;
[0062] By using a quality verification agent in the multi-agents, the plurality of first-domain information is cross-validated and screened by combining the photovoltaic cell material knowledge graph with a comparison of the plurality of information association relationships to obtain a plurality of second-domain information;
[0063] Performing hierarchical summarization on the plurality of second domain information by a document summarizing agent in the multi-agent, and generating a plurality of multi-level summary information corresponding to the plurality of second domain information;
[0064] The plurality of second domain information are integrated with the corresponding plurality of multi-level summary information, and matched and integrated with a plurality of preset research categories to obtain a preset domain data set.
[0065] This application first identifies the first question text entered by the user, obtains different types of keywords and keyword weight ratios, and then searches and filters the photovoltaic cell material knowledge graph based on the keywords to obtain multiple field search paths, thereby obtaining multiple field knowledge texts from the photovoltaic cell material knowledge graph according to the multiple field search paths, obtains search keywords by identifying the user's question, and then obtains the prior field knowledge for the large model question based on the keywords. Based on the prior field knowledge, the answer of the large model can be limited to a specific range, reducing the generation of large model hallucinations, thereby improving the accuracy of the large model's question and answer in the field of photovoltaic cell research and development when the first question text and multiple field knowledge texts are combined to ask the photovoltaic cell field large model to obtain the first answer text; at the same time, by extracting photovoltaic cell field data and multi-agent extraction, a photovoltaic cell material knowledge graph and a preset field data set are constructed, thereby obtaining a large model in the field of photovoltaic cells. Compared with the existing technology, by increasing the knowledge graph that can obtain prior field knowledge and the field data set for training the large model, the knowledge update cost of the large model can be reduced, the knowledge update efficiency of the large model can be improved, and the accuracy of the large model's question and answer in the field of photovoltaic cell research and development can be improved.
[0066] According to a fourth aspect of the embodiments of the present application, there is provided a photovoltaic cell experiment optimization device, comprising a result acquisition module and a question optimization module;
[0067] The result acquisition module is configured to acquire a first experimental result of the user; wherein the first experimental result is obtained by the user performing an experiment on a first experimental suggestion; and the first experimental suggestion is obtained by using the photovoltaic cell experiment question-and-answer device according to claim 8;
[0068] The question optimization module is used to ask questions to a preset photovoltaic cell experimental large model based on the first experimental results and the first experimental suggestions, obtain a second answer text, and optimize the first experimental suggestions based on the second answer text; wherein, the photovoltaic cell experimental large model is constructed based on a preset experimental question and answer dataset.
[0069] In certain embodiments of the present application, the photovoltaic cell experimental large model is constructed based on a preset experimental question-and-answer dataset, specifically including:
[0070] Acquire an initial experimental document dataset, and convert the title of each figure and the text content of each paragraph of each initial document in the initial experimental document dataset into vector representations, thereby obtaining multiple figure title vectors and multiple paragraph text vectors for each initial document;
[0071] Traversing each initial document, calculating a first similarity between each figure title vector in the current initial document and each paragraph text vector in the current initial document, and screening the initial documents in the initial experimental document dataset based on the first similarity to obtain a first experimental document dataset;
[0072] Based on the pre-trained graph-text matching model, the text content of each paragraph and each graph of each first document in the first experimental document dataset are converted into cross-modal vector representations, thereby obtaining multiple image vectors and multiple paragraph vectors for each first document.
[0073] Traversing each first document, calculating a second similarity between each image vector in the current first document and each paragraph vector in the current first document during the traversal, and combining image vectors and paragraph vectors whose second similarity exceeds a preset first threshold into a first image-text vector pair;
[0074] Combining a plurality of first image-text vector pairs of all first documents to obtain a first image-text vector set;
[0075] Ask questions to a preset second largest model based on the first image-text vector set, construct corresponding question-answer pairs for each vector pair in the first image-text vector set, and construct a preset experimental question-answer dataset based on the question-answer pairs;
[0076] According to the preset experimental question-answering dataset, the initial photovoltaic cell experimental large model is adjusted and trained to obtain the photovoltaic cell experimental large model.
[0077] In certain embodiments of the present application, the step of asking questions to a preset second large model based on the first image-text vector set, constructing a corresponding question-answer pair for each vector pair of the first image-text vector set, and constructing a preset experimental question-answer dataset based on the question-answer pairs specifically includes:
[0078] For each vector pair of the first image-text vector set, ask the second model a question based on a preset question-answer prompt word to obtain a corresponding question for each vector pair;
[0079] Combining each vector pair of the first image-text vector set with the corresponding question to obtain a question-answer pair corresponding to each vector pair;
[0080] Combine all question-answer pairs to obtain the preset experimental question-answering dataset.
[0081] This application first obtains the first experimental result of the experiment based on the first experimental suggestion obtained by the user based on the question and answer of the photovoltaic cell experiment, and then asks questions to the photovoltaic cell experiment big model according to the first experimental result and the first experimental suggestion, and optimizes the first experimental suggestion according to the question result. Compared with the existing technology, by training a big model that can analyze the experimental results and adding a process for optimizing the experimental suggestions according to the big model, it is possible to realize the analysis and reasoning of the experimental results, as well as the feedback optimization of the experimental suggestions, thereby improving the accuracy of the experimental optimization of the big model in the field of photovoltaic cell research and development.
[0082] According to a fifth aspect of the embodiments of the present application, there is provided a question-answering system for photovoltaic cell experiments, comprising an experiment suggestion module and an experiment optimization module;
[0083] The experiment suggestion module is used to obtain user experiment questions and execute a photovoltaic cell experiment question-answering method as described in the present application to generate corresponding user experiment suggestions according to the user experiment questions;
[0084] The experiment optimization module is used to obtain the user experiment results after the user conducts the experiment according to the user experiment suggestion, and execute a photovoltaic cell experiment optimization method as described in the present application to obtain corresponding experiment optimization suggestions based on the user experiment suggestion and the user experiment results. BRIEF DESCRIPTION OF THE DRAWINGS
[0085] Figure 1 : A flowchart of a question-and-answer method for a photovoltaic cell experiment shown in certain embodiments of the present application;
[0086] Figure 2 : A schematic flow chart of a photovoltaic cell experiment optimization method according to certain embodiments of the present application;
[0087] Figure 3 : A module structure diagram of a question-and-answer device for photovoltaic cell experiments shown in certain embodiments of the present application;
[0088] Figure 4 : A module structure diagram of an optimization device for photovoltaic cell experiments shown in certain embodiments of the present application;
[0089] Figure 5 : A system architecture diagram of a question-and-answer system for photovoltaic cell experiments shown in certain embodiments of the present application. DETAILED DESCRIPTION
[0090] The embodiments of the present application are described in detail below. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions from beginning to end. The embodiments described below in conjunction with the accompanying drawings are exemplary and are only used to explain some embodiments of the present application and should not be understood as limiting the embodiments of the present application. Based on the embodiments shown in this application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0091] In the description of this application, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly indicate the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of such features. In the description of this application, unless otherwise clearly specified, "multiple" and "several" mean two or more.
[0092] The existing large models used for photovoltaic cell research and development have the following problems: (1) Current general large models, such as Deepseek and GPT-4, have limited training data for the field of photovoltaic cell research and development. As a result, each time the knowledge of the large model is updated, a large amount of resources need to be consumed for fine-tuning all parameters. This is not only costly, but also cannot match the current situation of rapid iteration of photovoltaic cell research and development, resulting in untimely knowledge updates and affecting the accuracy of large model question answering; (2) Although the current large model can obtain the ability to answer questions in the field of photovoltaic cell research and development through fine-tuning, after the experimental results are obtained by conducting experiments on experimental suggestions for photovoltaic cell research and development based on the large model, the analysis and reasoning of the experimental result data are insufficient, and there is also a lack of feedback optimization of the experimental suggestions based on the experimental results, resulting in the lack of experimental analysis and optimization capabilities of the large model. The accuracy of the large model during experimental analysis and optimization is not high. Therefore, how to improve the accuracy of large models in question answering and experimental optimization in the field of photovoltaic cell research and development is still a technical problem that needs to be solved urgently in the existing technology.
[0093] Based on the above technical background, please refer to Figure 1 The embodiment of the present application provides a question-answering method for a photovoltaic cell experiment, including steps S101 to S104, each of which is specifically as follows:
[0094] Step S101: Identify the first question text input by the user and obtain domain keywords, technical keywords and keyword weight ratios.
[0095] In certain embodiments of the present application, the recognition of the first question text input by the user is based on a pre-trained classification model, and a preferred embodiment of the pre-trained classification model is a Bert classifier.
[0096] Considering the recognition of the first question text input by the user, we obtain the domain keywords and technical keywords and their keyword weight ratios. This is because the domain keywords summarize the fields corresponding to the question text, such as "Perovskite Solar Cell Stability Research" and "Perovskite Quantum Dot Luminescence Mechanism", while the technical keywords focus on the specific technologies corresponding to the question text, such as "Thermal Stability", "Filling Factor" and "Titanium Dioxide". Searching through the combination of domain keywords and technical keywords can optimize the search efficiency and improve the search accuracy.
[0097] Step S102: Based on the field keywords and the technical keywords, a preset photovoltaic cell material knowledge graph is searched and screened in combination with a graph walking algorithm to obtain multiple field search paths; wherein, the photovoltaic cell material knowledge graph is generated based on the extraction of photovoltaic cell field data.
[0098] In certain embodiments of the present application, the field keywords and the technical keywords are combined with a graph walking algorithm to search and screen the preset photovoltaic cell material knowledge graph to obtain multiple field search paths, specifically including:
[0099] Search the photovoltaic cell material knowledge graph according to the domain keywords to obtain a first domain subgraph;
[0100] Filtering the first domain subgraph according to the technical keywords to obtain multiple first search paths;
[0101] The semantic coherence of the plurality of first search paths is verified based on a graph walk algorithm, and filtering is performed according to the verification result to obtain a plurality of domain search paths.
[0102] In certain embodiments of the present application, the photovoltaic cell material knowledge graph is searched according to the field keywords to obtain a first field subgraph, specifically: according to the field keywords, the photovoltaic cell material knowledge graph is searched based on a subgraph screening algorithm to obtain a first field subgraph.
[0103] This application first searches the photovoltaic cell material knowledge graph through domain keywords to obtain the first domain subgraph, which can first locate the domain where the keyword is located and narrow the search scope, and then filter the first domain subgraph in combination with technical keywords to obtain multiple first search paths, which can further narrow the scope and focus on the corresponding technology of the keyword, so as to improve the matching degree between the multiple domain search paths and the current user's question when verifying its semantic coherence based on the graph walking algorithm and filtering to obtain multiple domain search paths, and improve the accuracy of the prior domain knowledge when obtaining the prior domain knowledge based on multiple domain search paths.
[0104] In certain embodiments of the present application, the photovoltaic cell material knowledge graph is generated based on the extraction of photovoltaic cell field data, specifically including:
[0105] Based on the preset graph mode and combined with the preset first model, the photovoltaic cell field data is screened by preset keywords to obtain multiple first field text fragments;
[0106] According to the different keyword weights of the preset keywords in each first-domain text segment and the number of citations of each first-domain text segment, the plurality of first-domain text segments are screened to obtain a plurality of second-domain text segments;
[0107] According to the graph pattern and based on the first large model, knowledge extraction is performed on the plurality of second domain text fragments to obtain a plurality of domain entity subgraphs;
[0108] Combined with the semantic similarity between the multiple domain entity subgraphs, the multiple domain entity subgraphs are disambiguated and merged to obtain a photovoltaic cell material knowledge graph.
[0109] In certain embodiments of the present application, implementations of the first large model include but are not limited to Wenxin Yiyan, Tongyi Qianwen, iFlytek Spark, Deepseek, Gemini, GPT-4 or Kimi, with Deepseek being a preferred implementation.
[0110] In certain embodiments of the present application, the plurality of first-domain text segments are screened based on the different keyword weights of the preset keywords in each first-domain text segment, combined with the number of citations of each first-domain text segment, to obtain a plurality of second-domain text segments, specifically:
[0111] Calculating a relevance score for each first-domain text segment based on different keyword weights of the preset keyword in each first-domain text segment and the number of citations of each first-domain text segment, and screening the plurality of first-domain text segments based on the relevance scores to obtain a plurality of second-domain text segments;
[0112] The correlation score is specifically:
[0113] Score(d)=α·TF-IDF(d)+β·CitationCount(d);
[0114] Among them, α and β are weight parameters, TF-IDF(d) is the keyword weight of the preset keyword in the current first field text segment d, and CitationCount(d) is the number of citations of the current first field text segment d.
[0115] In the embodiment of the present application, the preferred implementation of the semantic similarity between the multiple domain entity subgraphs is cosine similarity; the semantic similarity between the multiple domain entity subgraphs is combined to disambiguate and merge the multiple domain entity subgraphs to obtain a photovoltaic cell material knowledge graph, specifically:
[0116] A plurality of domain entity subgraphs whose semantic similarity among the multiple domain entity subgraphs is higher than a preset semantic similarity threshold are disambiguated and merged to obtain a photovoltaic cell material knowledge graph.
[0117] This application first filters the photovoltaic cell field data based on the preset graph pattern and the preset first model through preset keywords. It can efficiently obtain multiple first-field text fragments that conform to the graph pattern based on the first model, and then filter according to the weight of the preset keywords and the number of citations of the first-field text fragments, and can exclude text fragments with low relevance to the photovoltaic cell field. Then, according to the graph pattern, knowledge extraction is performed based on the first model to obtain multiple domain entity subgraphs, and disambiguation and merging are performed according to the mutual semantic similarity, which can improve the accuracy of the knowledge graph and the quality of the domain knowledge therein.
[0118] Step S103: Acquire multiple domain knowledge texts from the photovoltaic cell material knowledge graph according to the multiple search paths; wherein the ratio of texts satisfying domain keywords to texts satisfying technical keywords in the multiple domain knowledge texts conforms to the keyword weight ratio.
[0119] In certain embodiments of the present application, the acquiring of multiple domain knowledge texts from the photovoltaic cell material knowledge graph according to the multiple search paths further includes:
[0120] Based on a preset retrieval language model, conflicting texts in the plurality of domain knowledge texts are marked to obtain a plurality of groups of conflicting knowledge texts;
[0121] Calculate the confidence score of each knowledge text in each group of conflicting knowledge texts based on its age, impact factor, and number of citations. Keep the knowledge text with the highest confidence score in each group of conflicting knowledge texts and remove the knowledge texts with lower confidence scores.
[0122] The confidence score is specifically:
[0123]
[0124] Among them, C is the confidence score, R is the total number of domain knowledge texts, r is the current knowledge text, AuthScore(r) is the authority score, t a is the literature age of the current knowledge text, IF ais the impact factor of the current knowledge text, IF max is the maximum impact factor of all knowledge texts in the current group of conflicting knowledge texts, C a is the number of citations of the current knowledge text, ∑C r The total number of citations of all knowledge texts in the current group of conflicting knowledge texts.
[0125] Step S104: Based on the first question text and the multiple domain knowledge texts, questions are asked to the preset photovoltaic cell domain model to obtain a first answer text; wherein, the photovoltaic cell domain model is constructed based on a preset domain data set; the preset domain data set is extracted and constructed by multiple agents from photovoltaic cell domain data.
[0126] In certain embodiments of the present application, implementations of the photovoltaic cell field large model include but are not limited to Wenxin Yiyan, Tongyi Qianwen, iFlytek Spark, Deepseek, Gemini, GPT-4 or Kimi, and the preferred implementation is Tongyi Qianwen.
[0127] In certain embodiments of the present application, the preset domain dataset is extracted and constructed by multi-agents from photovoltaic cell domain data, specifically including:
[0128] By using the information extraction agent in the multi-agent, the photovoltaic cell field data is structuredly parsed to extract a plurality of first-field information, and contextual reasoning is performed on the plurality of first-field information to construct a plurality of information association relationships;
[0129] By using a quality verification agent in the multi-agents, the plurality of first-domain information is cross-validated and screened by combining the photovoltaic cell material knowledge graph with a comparison of the plurality of information association relationships to obtain a plurality of second-domain information;
[0130] Performing hierarchical summarization on the plurality of second domain information by a document summarizing agent in the multi-agent, and generating a plurality of multi-level summary information corresponding to the plurality of second domain information;
[0131] The plurality of second domain information are integrated with the corresponding plurality of multi-level summary information, and matched and integrated with a plurality of preset research categories to obtain a preset domain data set.
[0132] In certain embodiments of the present application, the quality verification agent in the multi-agent system, in combination with the photovoltaic cell material knowledge graph and the comparison of the plurality of information associations, performs cross-validation screening on the plurality of first-domain information to obtain a plurality of second-domain information, specifically:
[0133] By means of a quality verification agent in the multi-agent, the confidence score of each piece of first-domain information is calculated by combining the photovoltaic cell material knowledge graph with a comparison of the plurality of information association relationships, and the plurality of first-domain information are cross-validated and screened according to the confidence scores to obtain a plurality of second-domain information;
[0134] The confidence score is specifically:
[0135] C(i)=w1·C internal(i) +w2·C external(i) +w3·C KG(i)
[0136]
[0137] Among them, w1, w2, w3 are weight coefficients, C internal(i) To score internal consistency, C external(i) Scoring external validation, C KG(i) is the matching degree of the information association relationship between the photovoltaic cell material knowledge graph and the current first-domain information; S is the set of relevant statements in the literature where the current first-domain information is located, sim(i, s) is the semantic similarity between the current first-domain information and the relevant statement s, and rel(s) is the relevance weight of the relevant statement s.
[0138] In certain embodiments of the present application, C external(i) The preferred implementation is external manual annotation scoring, C KG(i) The preferred implementation method is to calculate the matching degree of the information association relationship between the photovoltaic cell material knowledge graph and the current first field information based on similarity.
[0139] This application first performs structured analysis and extraction of photovoltaic cell field data based on the information extraction agent in the multi-agent to obtain multiple first-field information, and performs contextual reasoning to construct multiple information association relationships. Then, through the quality verification agent in the multi-agent, combined with the photovoltaic cell material knowledge graph and the comparison of multiple information association relationships, cross-validation and screening are performed to obtain multiple second-field information. It can screen out information with low correlation with photovoltaic cell materials in the first-field information and improve the accuracy of the information. Then, the document summary agent in the multi-agent performs hierarchical summary to generate multiple multi-level summary information, and then matches and integrates with multiple second-field information and multiple research categories to obtain a preset field data set. It can add a summary dimension to the second field information, improve information density, and thus improve the matching degree between the preset field data set and the photovoltaic cell field. Therefore, when constructing a large model of the photovoltaic cell field based on the preset field data set, the large model is more matched with the photovoltaic cell field, thereby improving the accuracy of the large model when conducting photovoltaic cell experimental question and answer.
[0140] Compared with the prior art, the present application first identifies the first question text input by the user, obtains different types of keywords and keyword weight ratios, and then searches and filters the photovoltaic cell material knowledge graph based on the keywords to obtain multiple field search paths, thereby obtaining multiple field knowledge texts from the photovoltaic cell material knowledge graph according to the multiple field search paths. By identifying the user's question, the search keywords are obtained, and then the prior field knowledge for the large model question is obtained based on the keywords. The answer of the large model can be limited to a specific range based on the prior field knowledge, reducing the generation of the large model hallucination. When the first question text and multiple field knowledge texts are combined to ask the photovoltaic cell field large model to obtain the first answer text, the accuracy of the large model's question and answer in the field of photovoltaic cell research and development is improved. At the same time, by extracting photovoltaic cell field data and multi-agent extraction, a photovoltaic cell material knowledge graph and a preset field data set are constructed, and then a large model in the field of photovoltaic cell is obtained. Compared with the prior art, by increasing the knowledge graph that can obtain prior field knowledge and the field data set for training the large model, the knowledge update cost of the large model can be reduced, the knowledge update efficiency of the large model can be improved, and the accuracy of the large model's question and answer in the field of photovoltaic cell research and development can be improved.
[0141] See Figure 2 The embodiment of the present application further provides a method for optimizing a photovoltaic cell experiment, including steps S201 to S202, each of which is specifically as follows:
[0142] Step S201: Obtain a first experimental result of a user; wherein the first experimental result is obtained by the user conducting an experiment on a first experimental suggestion; the first experimental suggestion is obtained through a question-and-answer method for a photovoltaic cell experiment described in this application.
[0143] Step S202: Based on the first experimental result and the first experimental suggestion, questions are asked to the preset photovoltaic cell experimental large model to obtain a second answer text, and the first experimental suggestion is optimized based on the second answer text; wherein, the photovoltaic cell experimental large model is constructed based on a preset experimental question and answer dataset.
[0144] In certain embodiments of the present application, the photovoltaic cell experimental large model is implemented by methods including but not limited to Wenxin Yiyan, Tongyi Qianwen, iFlytek Spark, Deepseek, Gemini, GPT-4 or Kimi, with Tongyi Qianwen being the preferred implementation method.
[0145] In certain embodiments of the present application, the photovoltaic cell experimental large model is constructed based on a preset experimental question-and-answer dataset, specifically including:
[0146] Get the initial experimental literature dataset D chart , and the initial experimental literature dataset Dchart The title t of each figure and the text content p of each paragraph of each original document d are converted into vector representations respectively, and multiple figure title vectors v(t) and multiple paragraph text vectors v(p) of each original document are obtained;
[0147] Traverse each initial document d, calculate the first similarity sim(t, p) between each chart title vector v(t) in the current initial document and each paragraph text vector v(p) in the current initial document, and according to the first similarity sim(t, p), perform the initial experimental document dataset D chart The initial literature d in the dataset is screened to obtain the first experimental literature dataset D refined ;
[0148] Based on the pre-trained graph-text matching model, the first experimental document dataset D refined The text content of each paragraph and each chart of each first document in the is converted into a cross-modal vector representation, and multiple image vectors and multiple paragraph vectors are obtained for each first document;
[0149] Traversing each first document, calculating a second similarity between each image vector in the current first document and each paragraph vector in the current first document during the traversal, and combining image vectors and paragraph vectors whose second similarity exceeds a preset first threshold into a first image-text vector pair;
[0150] Combining a plurality of first image-text vector pairs of all first documents to obtain a first image-text vector set;
[0151] Ask questions to a preset second largest model based on the first image-text vector set, construct corresponding question-answer pairs for each vector pair in the first image-text vector set, and construct a preset experimental question-answer dataset based on the question-answer pairs;
[0152] According to the preset experimental question-answering dataset, the initial photovoltaic cell experimental large model is adjusted and trained to obtain the photovoltaic cell experimental large model.
[0153] In certain embodiments of the present application, the initial documents in the initial experimental document dataset are screened according to the first similarity to obtain a first experimental document dataset, specifically:
[0154] Filtering out a paragraph text set corresponding to each figure title vector in the current initial document, wherein each element in the paragraph text set satisfies a first similarity with the corresponding figure title vector that is higher than a preset threshold;
[0155] If each paragraph text set of the current initial document is an empty set, the current initial document will be removed from the initial experimental document data set.
[0156] In certain embodiments of the present application, a preferred embodiment of the graph-text matching model is the CLIP model.
[0157] This application first obtains an initial experimental document data set and converts it to obtain multiple chart title vectors and multiple paragraph text vectors for each initial document, then screens the initial documents based on the first similarity to obtain the first experimental document data set, which can ensure the correlation between the charts and texts of each initial document in the first experimental document data set, and then converts multiple image vectors and multiple paragraph vectors for each first document in the first experimental document data set based on the pre-trained chart-text matching model, and then obtains a first image-text vector set based on the second similarity screening combination, which can ensure the correlation between the image and paragraph text of each first image-text vector pair in the first image-text vector set, so that when asking questions to the second large model according to the first image-text vector set to construct the corresponding question-answer pair for each vector pair, it is ensured that the generated question matches the current vector pair, so that when constructing the photovoltaic cell experimental large model according to the preset experimental question-answer data set, the accuracy of the photovoltaic cell experimental large model is improved.
[0158] In certain embodiments of the present application, the step of asking questions to a preset second large model based on the first image-text vector set, constructing a corresponding question-answer pair for each vector pair of the first image-text vector set, and constructing a preset experimental question-answer dataset based on the question-answer pairs specifically includes:
[0159] For each vector pair of the first image-text vector set, ask the second model a question based on a preset question-answer prompt word to obtain a corresponding question for each vector pair;
[0160] Combining each vector pair of the first image-text vector set with the corresponding question to obtain a question-answer pair corresponding to each vector pair;
[0161] Combine all question-answer pairs to obtain the preset experimental question-answering dataset.
[0162] This application first constructs corresponding questions for each vector pair in the first image-text vector set based on the question-answer prompt words, then performs corresponding combinations to obtain question-answer pairs, and then constructs an experimental question-answering dataset, which can ensure the correspondence between questions and vector pairs, thereby obtaining diverse questions based on the diversity of vector pairs, thereby improving the data richness of the preset experimental question-answering dataset.
[0163] Compared with the existing technology, this application first obtains the first experimental result of the experiment based on the first experimental suggestion obtained by the user based on the question and answer of the photovoltaic cell experiment, and then asks questions to the photovoltaic cell experiment big model according to the first experimental result and the first experimental suggestion, and optimizes the first experimental suggestion according to the question result. Compared with the existing technology, by training a big model that can analyze the experimental results and adding a process for optimizing the experimental suggestions based on the big model, it is possible to realize the analysis and reasoning of the experimental results, as well as the feedback optimization of the experimental suggestions, thereby improving the accuracy of the experimental optimization of the big model in the field of photovoltaic cell research and development.
[0164] Corresponding to the above Q&A method, please see Figure 3 , the embodiment of the present application provides a question-answering device for photovoltaic cell experiments, including a question recognition module 310, a graph retrieval module 320, a knowledge acquisition module 330 and a model questioning module 340;
[0165] The question recognition module 310 is used to recognize the first question text input by the user and obtain the domain keywords, technical keywords and keyword weight ratios;
[0166] The graph retrieval module 320 is configured to search and filter a preset photovoltaic cell material knowledge graph based on the domain keywords and the technical keywords in combination with a graph walking algorithm to obtain multiple domain search paths; wherein the photovoltaic cell material knowledge graph is generated based on the extraction of photovoltaic cell domain data;
[0167] The knowledge acquisition module 330 is configured to acquire a plurality of domain knowledge texts from the photovoltaic cell material knowledge graph according to the plurality of search paths; wherein the ratio of texts satisfying domain keywords to texts satisfying technical keywords in the plurality of domain knowledge texts conforms to the keyword weight ratio;
[0168] The model questioning module 340 is used to ask questions to the preset photovoltaic cell field large model based on the first question text and the multiple domain knowledge texts to obtain a first answer text; wherein, the photovoltaic cell field large model is constructed based on a preset domain data set; the preset domain data set is extracted and constructed by multiple intelligent agents from photovoltaic cell field data.
[0169] In certain embodiments of the present application, the graph search module 320 includes a keyword search unit, a keyword screening unit, and a path verification unit;
[0170] The keyword search unit is configured to search the photovoltaic cell material knowledge graph based on the field keywords to obtain a first field subgraph;
[0171] The keyword screening unit is configured to screen the first domain subgraph according to the technical keywords to obtain a plurality of first search paths;
[0172] The path verification unit is used to verify the semantic coherence of the multiple first search paths based on a graph walk algorithm, and filter according to the verification results to obtain multiple domain search paths.
[0173] In certain embodiments of the present application, the photovoltaic cell material knowledge graph is generated based on the extraction of photovoltaic cell field data, specifically including:
[0174] Based on the preset graph mode and combined with the preset first model, the photovoltaic cell field data is screened by preset keywords to obtain multiple first field text fragments;
[0175] According to the different keyword weights of the preset keywords in each first-domain text segment and the number of citations of each first-domain text segment, the plurality of first-domain text segments are screened to obtain a plurality of second-domain text segments;
[0176] According to the graph pattern and based on the first large model, knowledge extraction is performed on the plurality of second domain text fragments to obtain a plurality of domain entity subgraphs;
[0177] Combined with the semantic similarity between the multiple domain entity subgraphs, the multiple domain entity subgraphs are disambiguated and merged to obtain a photovoltaic cell material knowledge graph.
[0178] In certain embodiments of the present application, the preset domain dataset is extracted and constructed by multi-agents from photovoltaic cell domain data, specifically including:
[0179] By using the information extraction agent in the multi-agent, the photovoltaic cell field data is structuredly parsed to extract a plurality of first-field information, and contextual reasoning is performed on the plurality of first-field information to construct a plurality of information association relationships;
[0180] By using a quality verification agent in the multi-agents, the plurality of first-domain information is cross-validated and screened by combining the photovoltaic cell material knowledge graph with a comparison of the plurality of information association relationships to obtain a plurality of second-domain information;
[0181] Performing hierarchical summarization on the plurality of second domain information by a document summarizing agent in the multi-agent, and generating a plurality of multi-level summary information corresponding to the plurality of second domain information;
[0182] The plurality of second domain information are integrated with the corresponding plurality of multi-level summary information, and matched and integrated with a plurality of preset research categories to obtain a preset domain data set.
[0183] This application first identifies the first question text entered by the user, obtains different types of keywords and keyword weight ratios, and then searches and filters the photovoltaic cell material knowledge graph based on the keywords to obtain multiple field search paths, thereby obtaining multiple field knowledge texts from the photovoltaic cell material knowledge graph according to the multiple field search paths, obtains search keywords by identifying the user's question, and then obtains the prior field knowledge for the large model question based on the keywords. Based on the prior field knowledge, the answer of the large model can be limited to a specific range, reducing the generation of large model hallucinations, thereby improving the accuracy of the large model's question and answer in the field of photovoltaic cell research and development when the first question text and multiple field knowledge texts are combined to ask the photovoltaic cell field large model to obtain the first answer text; at the same time, by extracting photovoltaic cell field data and multi-agent extraction, a photovoltaic cell material knowledge graph and a preset field data set are constructed, thereby obtaining a large model in the field of photovoltaic cells. Compared with the existing technology, by increasing the knowledge graph that can obtain prior field knowledge and the field data set for training the large model, the knowledge update cost of the large model can be reduced, the knowledge update efficiency of the large model can be improved, and the accuracy of the large model's question and answer in the field of photovoltaic cell research and development can be improved.
[0184] Corresponding to the above optimization method, see Figure 4 , the embodiment of the present application also provides an optimization device for photovoltaic cell experiments, including a result acquisition module 410 and a question optimization module 420;
[0185] The result acquisition module 410 is configured to acquire a first experimental result of a user; wherein the first experimental result is obtained by the user performing an experiment on a first experimental suggestion; the first experimental suggestion is obtained by using a photovoltaic cell experiment question-and-answer device as described in this application;
[0186] The question optimization module 420 is used to ask questions to a preset photovoltaic cell experimental large model based on the first experimental results and the first experimental suggestions, obtain a second answer text, and optimize the first experimental suggestions based on the second answer text; wherein, the photovoltaic cell experimental large model is constructed based on a preset experimental question and answer data set.
[0187] In certain embodiments of the present application, the photovoltaic cell experimental large model is constructed based on a preset experimental question-and-answer dataset, specifically including:
[0188] Acquire an initial experimental document dataset, and convert the title of each figure and the text content of each paragraph of each initial document in the initial experimental document dataset into vector representations, thereby obtaining multiple figure title vectors and multiple paragraph text vectors for each initial document;
[0189] Traversing each initial document, calculating a first similarity between each figure title vector in the current initial document and each paragraph text vector in the current initial document, and screening the initial documents in the initial experimental document dataset based on the first similarity to obtain a first experimental document dataset;
[0190] Based on the pre-trained graph-text matching model, the text content of each paragraph and each graph of each first document in the first experimental document dataset are converted into cross-modal vector representations, thereby obtaining multiple image vectors and multiple paragraph vectors for each first document.
[0191] Traversing each first document, calculating a second similarity between each image vector in the current first document and each paragraph vector in the current first document during the traversal, and combining image vectors and paragraph vectors whose second similarity exceeds a preset first threshold into a first image-text vector pair;
[0192] Combining a plurality of first image-text vector pairs of all first documents to obtain a first image-text vector set;
[0193] Ask questions to a preset second largest model based on the first image-text vector set, construct corresponding question-answer pairs for each vector pair in the first image-text vector set, and construct a preset experimental question-answer dataset based on the question-answer pairs;
[0194] According to the preset experimental question-answering dataset, the initial photovoltaic cell experimental large model is adjusted and trained to obtain the photovoltaic cell experimental large model.
[0195] In certain embodiments of the present application, the step of asking questions to a preset second large model based on the first image-text vector set, constructing a corresponding question-answer pair for each vector pair of the first image-text vector set, and constructing a preset experimental question-answer dataset based on the question-answer pairs specifically includes:
[0196] For each vector pair of the first image-text vector set, ask the second model a question based on a preset question-answer prompt word to obtain a corresponding question for each vector pair;
[0197] Combining each vector pair of the first image-text vector set with the corresponding question to obtain a question-answer pair corresponding to each vector pair;
[0198] Combine all question-answer pairs to obtain the preset experimental question-answering dataset.
[0199] This application first obtains the first experimental result of the experiment based on the first experimental suggestion obtained by the user based on the question and answer of the photovoltaic cell experiment, and then asks questions to the photovoltaic cell experiment big model according to the first experimental result and the first experimental suggestion, and optimizes the first experimental suggestion according to the question result. Compared with the existing technology, by training a big model that can analyze the experimental results and adding a process for optimizing the experimental suggestions according to the big model, it is possible to realize the analysis and reasoning of the experimental results, as well as the feedback optimization of the experimental suggestions, thereby improving the accuracy of the experimental optimization of the big model in the field of photovoltaic cell research and development.
[0200] Adaptively, see Figure 5 , the embodiment of the present application also provides a question-answering system for photovoltaic cell experiments, including an experiment suggestion module 510 and an experiment optimization module 520;
[0201] The experiment suggestion module 510 is configured to obtain a user experiment question and execute a photovoltaic cell experiment question-answering method as described in the present application to generate corresponding user experiment suggestions based on the user experiment question;
[0202] The experiment optimization module 520 is used to obtain the user experiment results after the user conducts the experiment according to the user experiment suggestion, and execute a photovoltaic cell experiment optimization method as described in the present application to obtain corresponding experiment optimization suggestions based on the user experiment suggestion and the user experiment results.
[0203] It should be understood that the device provided in the embodiment of the present application corresponds to the aforementioned method. The question-and-answer device for a photovoltaic cell experiment provided in the embodiment of the present application can implement the question-and-answer method for a photovoltaic cell experiment provided in any embodiment of the present application; the question-and-answer device for a photovoltaic cell experiment provided in the embodiment of the present application can implement the question-and-answer method for a photovoltaic cell experiment provided in any embodiment of the present application.
[0204] Adaptively, the embodiments of the present application further provide a computer device and a computer-readable storage medium.
[0205] The computer device comprises: a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor;
[0206] Wherein, when the processor executes the computer program, a question-and-answer method for a photovoltaic cell experiment or an optimization method for a photovoltaic cell experiment of the present application is implemented.
[0207] The computer-readable storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor to execute a question-and-answer method for a photovoltaic cell experiment or a photovoltaic cell experiment optimization method of the present application.
[0208] The above description is a partial embodiment of the present application, which further describes in detail the purpose, technical solutions, and beneficial effects of the present application. It should be understood that the above description of the partial embodiment of the present application is not to be construed as limiting the present application. In particular, it is pointed out that for those skilled in the art, any changes, modifications, equivalent substitutions, and variations made within the spirit and principles of the present application should be included within the scope of protection of the present application.
Claims
1. A question-answering method for photovoltaic cell experiments, characterized in that: include: Identify the first question text input by the user and obtain domain keywords, technical keywords and keyword weight ratios; Based on the field keywords and the technical keywords, a preset photovoltaic cell material knowledge graph is searched and screened in combination with a graph walking algorithm to obtain multiple field search paths; wherein the photovoltaic cell material knowledge graph is generated based on the extraction of photovoltaic cell field data; According to the multiple search paths, a plurality of domain knowledge texts are obtained from the photovoltaic cell material knowledge graph; wherein the ratio of texts satisfying domain keywords to texts satisfying technical keywords in the multiple domain knowledge texts conforms to the keyword weight ratio; Based on the first question text and the multiple domain knowledge texts, questions are asked to the preset photovoltaic cell domain large model to obtain a first answer text; wherein, the photovoltaic cell domain large model is constructed based on a preset domain data set; the preset domain data set is extracted and constructed by multiple intelligent agents from photovoltaic cell domain data.
2. A question-and-answer method for photovoltaic cell experiments according to claim 1, characterized in that: The preset photovoltaic cell material knowledge graph is searched and screened based on the field keywords and the technical keywords in combination with the graph walking algorithm to obtain multiple field search paths, specifically including: Search the photovoltaic cell material knowledge graph according to the domain keywords to obtain a first domain subgraph; Filtering the first domain subgraph according to the technical keywords to obtain multiple first search paths; The semantic coherence of the plurality of first search paths is verified based on a graph walk algorithm, and filtering is performed according to the verification result to obtain a plurality of domain search paths.
3. The question-answering method for photovoltaic cell experiments according to claim 1, characterized in that: The photovoltaic cell material knowledge graph is generated based on the extraction of photovoltaic cell field data, specifically including: Based on the preset graph mode and combined with the preset first model, the photovoltaic cell field data is screened by preset keywords to obtain multiple first field text fragments; According to the different keyword weights of the preset keywords in each first-domain text segment and the number of citations of each first-domain text segment, the plurality of first-domain text segments are screened to obtain a plurality of second-domain text segments; According to the graph pattern and based on the first large model, knowledge extraction is performed on the plurality of second domain text fragments to obtain a plurality of domain entity subgraphs; Combined with the semantic similarity between the multiple domain entity subgraphs, the multiple domain entity subgraphs are disambiguated and merged to obtain a photovoltaic cell material knowledge graph.
4. The question-answering method for photovoltaic cell experiments according to claim 1, characterized in that: The preset domain dataset is obtained by extracting and constructing photovoltaic cell domain data through multi-agents, and specifically includes: By using the information extraction agent in the multi-agent, the photovoltaic cell field data is structuredly parsed to extract a plurality of first-field information, and contextual reasoning is performed on the plurality of first-field information to construct a plurality of information association relationships; By using a quality verification agent in the multi-agents, the plurality of first-domain information is cross-validated and screened by combining the photovoltaic cell material knowledge graph with a comparison of the plurality of information association relationships to obtain a plurality of second-domain information; Performing hierarchical summarization on the plurality of second domain information by a document summarizing agent in the multi-agent, and generating a plurality of multi-level summary information corresponding to the plurality of second domain information; The plurality of second domain information are integrated with the corresponding plurality of multi-level summary information, and matched and integrated with a plurality of preset research categories to obtain a preset domain data set.
5. A photovoltaic cell experiment optimization method, characterized in that: include: Obtaining a first experimental result of a user; wherein the first experimental result is obtained by the user performing an experiment on a first experimental suggestion; and the first experimental suggestion is obtained by the photovoltaic cell experiment question-and-answer method according to any one of claims 1 to 4; Based on the first experimental results and the first experimental suggestions, questions are asked to a preset photovoltaic cell experimental large model to obtain a second answer text, and the first experimental suggestion is optimized based on the second answer text; wherein, the photovoltaic cell experimental large model is constructed based on a preset experimental question and answer data set.
6. The photovoltaic cell experiment optimization method according to claim 5, characterized in that: The photovoltaic cell experimental model is constructed based on a preset experimental question-answering dataset, specifically including: Acquire an initial experimental document dataset, and convert the title of each figure and the text content of each paragraph of each initial document in the initial experimental document dataset into vector representations, thereby obtaining multiple figure title vectors and multiple paragraph text vectors for each initial document; Traversing each initial document, calculating a first similarity between each figure title vector in the current initial document and each paragraph text vector in the current initial document, and screening the initial documents in the initial experimental document dataset based on the first similarity to obtain a first experimental document dataset; Based on the pre-trained graph-text matching model, the text content of each paragraph and each graph of each first document in the first experimental document dataset are converted into cross-modal vector representations, thereby obtaining multiple image vectors and multiple paragraph vectors for each first document. Traversing each first document, calculating a second similarity between each image vector in the current first document and each paragraph vector in the current first document during the traversal, and combining image vectors and paragraph vectors whose second similarity exceeds a preset first threshold into a first image-text vector pair; Combining a plurality of first image-text vector pairs of all first documents to obtain a first image-text vector set; Ask questions to a preset second largest model based on the first image-text vector set, construct corresponding question-answer pairs for each vector pair in the first image-text vector set, and construct a preset experimental question-answer dataset based on the question-answer pairs; According to the preset experimental question-answering dataset, the initial photovoltaic cell experimental large model is adjusted and trained to obtain the photovoltaic cell experimental large model.
7. The photovoltaic cell experiment optimization method according to claim 6, characterized in that: The step of asking questions to a preset second model based on the first image-text vector set, constructing a corresponding question-answer pair for each vector pair of the first image-text vector set, and constructing a preset experimental question-answer dataset based on the question-answer pairs specifically includes: For each vector pair of the first image-text vector set, ask the second model a question based on a preset question-answer prompt word to obtain a corresponding question for each vector pair; Combining each vector pair of the first image-text vector set with the corresponding question to obtain a question-answer pair corresponding to each vector pair; Combine all question-answer pairs to obtain the preset experimental question-answering dataset.
8. A question-answering device for photovoltaic cell experiments, characterized in that: It includes question recognition module, graph retrieval module, knowledge acquisition module and model question module; The question recognition module is used to recognize the first question text input by the user and obtain the domain keywords, technical keywords and keyword weight ratios; The graph retrieval module is used to search and filter a preset photovoltaic cell material knowledge graph based on the field keywords and the technical keywords in combination with a graph walking algorithm to obtain multiple field search paths; wherein the photovoltaic cell material knowledge graph is generated based on the extraction of photovoltaic cell field data; The knowledge acquisition module is configured to acquire a plurality of domain knowledge texts from the photovoltaic cell material knowledge graph according to the plurality of search paths; wherein the ratio of texts satisfying domain keywords to texts satisfying technical keywords in the plurality of domain knowledge texts conforms to the keyword weight ratio; The model questioning module is used to ask questions to a preset photovoltaic cell field large model based on the first question text and the multiple domain knowledge texts to obtain a first answer text; wherein, the photovoltaic cell field large model is constructed based on a preset domain data set; the preset domain data set is extracted and constructed by multiple intelligent agents from photovoltaic cell field data.
9. An optimization device for photovoltaic cell experiments, characterized in that: Including result acquisition module and question optimization module; The result acquisition module is configured to acquire a first experimental result of the user; wherein the first experimental result is obtained by the user performing an experiment on a first experimental suggestion; and the first experimental suggestion is obtained by using the photovoltaic cell experiment question-and-answer device according to claim 8; The question optimization module is used to ask questions to a preset photovoltaic cell experimental large model based on the first experimental results and the first experimental suggestions, obtain a second answer text, and optimize the first experimental suggestions based on the second answer text; wherein, the photovoltaic cell experimental large model is constructed based on a preset experimental question and answer dataset.
10. A question-answering system for photovoltaic cell experiments, characterized in that: Includes experiment suggestion module and experiment optimization module; The experiment suggestion module is configured to obtain user experiment questions and execute the photovoltaic cell experiment question-answering method according to any one of claims 1 to 4 to generate corresponding user experiment suggestions according to the user experiment questions; The experiment optimization module is used to obtain the user experiment results after the user conducts the experiment according to the user experiment suggestion, and execute the optimization method of the photovoltaic cell experiment as described in any one of claims 5 to 7 to obtain corresponding experiment optimization suggestions based on the user experiment suggestion and the user experiment results.
Citation Information
Patent Citations
Domain intelligent question-answering system and method based on knowledge graph library and text vector library
CN119128095A
Analysis method and system fusing large model and knowledge graph technology
CN119150972A
Multi-hop agricultural question-answering system based on knowledge graph reasoning and large model and text fragment retrieval and answer generation method thereof
CN119782467A
Knowledge graph and vector retrieval enhancement-based photovoltaic field large model efficiency improvement method
CN119988639A
Method and device for generating online question paths from existing question banks using a knowledge graph
US20180246952A1