Question-and-answer methods, optimization methods, apparatus, and question-and-answer systems for photovoltaic cell experiments

By recognizing user question text and combining graph walking algorithms and multi-agent technology to generate photovoltaic cell material knowledge graphs and datasets, the question-answering process of large models is optimized, solving the problem of insufficient question-answering accuracy of large models in photovoltaic cell research and development, and realizing efficient knowledge updates and experimental optimization.

CN120653740BActive Publication Date: 2026-04-03HONG KONG UNIV OF SCI & TECH (GUANGZHOU)
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing large-scale models lack accuracy in question-and-answer and experimental optimization in photovoltaic cell R&D, and the cost of knowledge updates is high, which is inconsistent with the rapid iteration of photovoltaic cell R&D.

Method used

By recognizing user-inputted question text, combining graph walking algorithms and multi-agent technology, a photovoltaic cell material knowledge graph and domain dataset are generated. The question-answering process of the large model is optimized, including keyword selection, semantic verification, and information merging, to construct a large model in the photovoltaic cell field.

Benefits of technology

It improves the accuracy of question answering and the efficiency of knowledge updating in the field of photovoltaic cell research and development, reduces the cost of knowledge updating, and enhances experimental optimization capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120653740B_ABST
    Figure CN120653740B_ABST
Patent Text Reader

Abstract

This application discloses a question-and-answer method, optimization method, apparatus, and question-and-answer system for photovoltaic cell experiments, relating to the field of large-scale model design applications. The question-and-answer method includes: identifying a first question text input by a user, obtaining domain keywords, technical keywords, and keyword weight ratios; based on the domain keywords and technical keywords, using a graph walk algorithm to search and filter a preset photovoltaic cell material knowledge graph, obtaining multiple domain search paths; based on the multiple search paths, obtaining multiple domain knowledge texts from the photovoltaic cell material knowledge graph; and based on the first question text and the multiple domain knowledge texts, asking a question to a preset photovoltaic cell domain large-scale model, obtaining a first answer text. The implementation of this application can solve the technical problem of insufficient accuracy in question-and-answer and experimental optimization of existing large-scale models in the field of photovoltaic cell R&D.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of large model design applications, and in particular to a question-and-answer method, optimization method, apparatus and question-and-answer system for photovoltaic cell experiments. Background Technology

[0002] With the increasing research and development of large-scale models, their application in the materials field, especially in photovoltaic cell research and development, has also seen significant growth. Due to the high cost of pre-training large-scale models, most large-scale models currently used in photovoltaic cell research and development focus on fine-tuning during the pre-training process. This involves fine-tuning general-purpose large-scale models using domain-specific corpora to enable them to acquire fundamental knowledge within the domain.

[0003] The general application of existing large-scale models for photovoltaic cell research and development is to provide experimental suggestions for photovoltaic cells by combining the large-scale model with a domain knowledge graph, so that corresponding photovoltaic cell research experiments can be carried out according to the suggestions. However, this application method has the following problems: (1) Current general-purpose large-scale models, such as Deepseek and GPT-4, have limited training data for the photovoltaic cell research and development field. This results in a large amount of resources being spent on full parameter fine-tuning for each knowledge update of the large-scale model. This is not only costly, but also does not match the rapid iteration of photovoltaic cell research and development, resulting in untimely knowledge updates and affecting the accuracy of the large-scale model's question answering; (2) Although current large-scale models can obtain domain question answering capabilities for photovoltaic cell research and development through fine-tuning, after conducting experiments based on the large-scale model's experimental suggestions for photovoltaic cell research and development, the analysis and reasoning of the experimental results data are insufficient, and there is a lack of feedback optimization of the experimental suggestions based on the experimental results. This results in the large-scale model lacking experimental analysis and optimization capabilities, and the accuracy of the large-scale model in conducting experimental analysis and optimization is not high. Therefore, how to improve the accuracy of large-scale models in question answering and experimental optimization in the field of photovoltaic cell research and development is still a technical problem that needs to be solved by the existing technology. Summary of the Invention

[0004] This application provides a question-and-answer method, optimization method, apparatus, and question-and-answer system for photovoltaic cell experiments, in order to solve the technical problem of insufficient accuracy in question-and-answer and experimental optimization of existing large models in the field of photovoltaic cell research and development.

[0005] According to a first aspect of the embodiments of this application, a question-and-answer method for photovoltaic cell experiments is provided, comprising:

[0006] Identify the first question text entered by the user and obtain domain keywords, technical keywords, and keyword weight ratios;

[0007] Based on the domain keywords and the technical keywords, a graph walking algorithm is used to search and filter the preset photovoltaic cell material knowledge graph to obtain multiple domain search paths; wherein, the photovoltaic cell material knowledge graph is generated based on the extraction of photovoltaic cell domain data;

[0008] Based on the multiple domain retrieval paths, multiple domain knowledge texts are obtained from the photovoltaic cell material knowledge graph; wherein, the ratio of texts satisfying domain keywords to texts satisfying technical keywords in the multiple domain knowledge texts conforms to the keyword weight ratio;

[0009] Based on the first question text and the multiple domain knowledge texts, a question is posed to the preset photovoltaic cell domain model to obtain the first answer text; wherein, the photovoltaic cell domain model is constructed based on the preset domain dataset; the preset domain dataset is constructed by extracting photovoltaic cell domain data through multiple agents.

[0010] This application first identifies the initial question text input by the user, obtains different types of keywords and their weight ratios, and then searches and filters the photovoltaic cell material knowledge graph based on the keywords to obtain multiple domain retrieval paths. Based on these multiple retrieval paths, it retrieves multiple domain knowledge texts from the photovoltaic cell material knowledge graph. By identifying the user's question, it obtains retrieval keywords and then acquires the prerequisite domain knowledge used for large-scale model questioning. This prerequisite domain knowledge restricts the large-scale model's answers to a specific range, reducing the occurrence of large-scale model illusions. Therefore, when combining the initial question text and multiple domain knowledge texts to ask a question to a large-scale photovoltaic cell model and obtain the first answer text, it improves the accuracy of the large-scale model's question-and-answer function in the photovoltaic cell R&D field. Simultaneously, by extracting photovoltaic cell domain data and using multi-agent extraction, it constructs a photovoltaic cell material knowledge graph and a pre-defined domain dataset, thus obtaining a large-scale photovoltaic cell model. Compared to existing technologies, by adding a knowledge graph with available prerequisite domain knowledge and a domain dataset for training the large-scale model, it reduces the knowledge update cost of the large-scale model, improves its knowledge update efficiency, and thus enhances the accuracy of the large-scale model's question-and-answer function in the photovoltaic cell R&D field.

[0011] In some embodiments of this application, the step of searching and filtering a preset photovoltaic cell material knowledge graph based on the domain keywords and the technical keywords, combined with a graph walk algorithm, to obtain multiple domain search paths specifically includes:

[0012] Based on the domain keywords, the photovoltaic cell material knowledge graph is retrieved to obtain a first domain subgraph;

[0013] Based on the technical keywords, the first domain subgraph is filtered to obtain multiple first search paths;

[0014] The semantic coherence of the multiple first retrieval paths is verified based on the graph walking algorithm, and the results are filtered to obtain multiple domain retrieval paths.

[0015] This application first retrieves a first domain subgraph from the photovoltaic cell material knowledge graph using domain keywords. This allows for the initial identification of the domain containing the keywords, narrowing the search scope. Then, by combining technical keywords with the first domain subgraph, multiple first search paths are obtained, further narrowing the scope and focusing on the corresponding technologies of the keywords. As a result, when verifying the semantic coherence of the multiple domain search paths based on the graph walk algorithm and filtering to obtain multiple domain search paths, the matching degree between the obtained multiple domain search paths and the current user's query is improved, thereby increasing the accuracy of prior domain knowledge when obtaining prior domain knowledge based on multiple domain search paths.

[0016] In some embodiments of this application, the photovoltaic cell material knowledge graph is generated based on the extraction of data in the field of photovoltaic cells, specifically including:

[0017] Based on the preset graph pattern and combined with the preset first major model, the data in the photovoltaic cell field is filtered through preset keywords to obtain multiple text fragments in the first field;

[0018] Based on the different keyword weights of the preset keywords in each first domain text segment, and combined with the number of references of each first domain text segment, the multiple first domain text segments are filtered to obtain multiple second domain text segments;

[0019] Based on the graph pattern and the first large model, knowledge extraction is performed on the multiple second domain text fragments to obtain multiple domain entity subgraphs;

[0020] By combining the semantic similarity among the multiple domain entity subgraphs, disambiguation merging is performed on the multiple domain entity subgraphs to obtain a photovoltaic cell material knowledge graph.

[0021] This application first uses a preset graph pattern and a preset first major model to filter data in the photovoltaic cell field using preset keywords. This allows for the efficient generation of multiple first-domain text fragments that conform to the graph pattern based on the first major model. Then, based on the weight of preset keywords and the citation frequency of first-domain text fragments, text fragments with low relevance to the photovoltaic cell field are filtered out. Subsequently, based on the graph pattern and the first major model, knowledge extraction is performed to obtain multiple domain entity subgraphs. These subgraphs are then disambiguated and merged based on their semantic similarity, thereby improving the accuracy of the knowledge graph and the quality of the domain knowledge within it.

[0022] In some embodiments of this application, the preset domain dataset is obtained by extracting and constructing photovoltaic cell domain data through multiple agents, specifically including:

[0023] By using an information extraction agent in a multi-agent system, structured parsing is performed on photovoltaic cell data to extract multiple pieces of first-domain information. Contextual reasoning is then performed on these multiple pieces of first-domain information to construct multiple information association relationships.

[0024] By using the quality verification agent in the multi-agent system, combined with the photovoltaic cell material knowledge graph and the comparison of the multiple information relationships, the multiple pieces of first domain information are cross-validated and filtered to obtain multiple pieces of second domain information;

[0025] By using a document summarizing agent in a multi-agent system, the multiple pieces of second-domain information are hierarchically summarized to generate multiple multi-level summary information corresponding to the multiple pieces of second-domain information;

[0026] The multiple pieces of second-domain information are integrated with the corresponding multiple multi-level summary information, and then matched and integrated with multiple preset research categories to obtain a preset domain dataset.

[0027] This application first uses an information extraction agent in a multi-agent system to perform structured parsing of photovoltaic cell data to extract multiple pieces of first-domain information and constructs multiple information associations through contextual reasoning. Then, a quality verification agent in the multi-agent system performs cross-validation and filtering by combining a photovoltaic cell material knowledge graph with the comparison of multiple information associations to obtain multiple pieces of second-domain information. This can filter out information in the first-domain information that has low relevance to photovoltaic cell materials, improving the accuracy of the information. Next, a document summarization agent in the multi-agent system performs hierarchical summarization to generate multiple multi-level summary information. This is then matched and integrated with multiple pieces of second-domain information and multiple research categories to obtain a preset domain dataset. This adds a summary dimension to the second-domain information, increases information density, and improves the matching degree between the preset domain dataset and the photovoltaic cell field. Therefore, when constructing a large-scale photovoltaic cell model based on the preset domain dataset, the large model is more closely matched to the photovoltaic cell field, thereby improving the accuracy of the large model in photovoltaic cell experimental question answering.

[0028] According to a second aspect of the embodiments of this application, an optimization method for photovoltaic cell experiments is provided, comprising:

[0029] Obtain the user's first experimental result; wherein the first experimental result is obtained by the user conducting experiments on the first experimental suggestion; the first experimental suggestion is obtained through a question-and-answer method for photovoltaic cell experiments as described in this application;

[0030] Based on the first experimental results and the first experimental suggestions, a question is posed to the pre-set large-scale photovoltaic cell experimental model to obtain a second response text. The first experimental suggestions are then optimized based on the second response text. The large-scale photovoltaic cell experimental model is constructed based on a pre-set experimental question-and-answer dataset.

[0031] This application first obtains the first experimental results of the experiment based on the user's question-and-answer session regarding photovoltaic cell experiments, and then asks questions about the large-scale photovoltaic cell experiment model based on the first experimental results and the first experimental suggestions. Based on the question-and-answer results, the first experimental suggestions are optimized. Compared with existing technologies, by training a large-scale model that can analyze experimental results and adding a process of optimizing experimental suggestions based on the large-scale model, it is possible to realize the analysis and reasoning of experimental results, as well as the feedback optimization of experimental suggestions, thereby improving the accuracy of experimental optimization of the large-scale model in the field of photovoltaic cell research and development.

[0032] In some embodiments of this application, the large-scale experimental model of the photovoltaic cell is constructed based on a pre-set experimental question-and-answer dataset, specifically including:

[0033] Obtain the initial experimental literature dataset, and convert the title of each figure and the text content of each paragraph of each initial literature in the initial experimental literature dataset into vector representations, to obtain multiple figure title vectors and multiple paragraph text vectors for each initial literature;

[0034] Each initial document is traversed. During traversal, the first similarity between the vector of each figure title in the current initial document and the vector of each paragraph text in the current initial document is calculated. Based on the first similarity, the initial documents in the initial experimental document dataset are filtered to obtain the first experimental document dataset.

[0035] Based on the pre-trained chart-text matching model, the text content of each paragraph and each chart of each first document in the first experimental literature dataset are converted into cross-modal vector representations, resulting in multiple image vectors and multiple paragraph vectors for each first document.

[0036] Traverse each first document, and calculate the second similarity between each image vector in the current first document and each paragraph vector in the current first document during the traversal. Combine the image vectors and paragraph vectors whose second similarity exceeds a preset first threshold into a first image-text vector pair.

[0037] Combine several first image-text vector pairs from all first documents to obtain the first image-text vector set;

[0038] Based on the first set of image and text vectors, questions are asked to the second preset model, and corresponding question-and-answer pairs are constructed for each vector pair in the first set of image and text vectors. A preset experimental question-and-answer dataset is then constructed based on the question-and-answer pairs.

[0039] Based on the pre-set experimental question and answer dataset, the initial large-scale photovoltaic cell experimental model is adjusted and trained to obtain the large-scale photovoltaic cell experimental model.

[0040] This application first obtains an initial experimental literature dataset and converts it to obtain multiple figure and table title vectors and multiple paragraph text vectors for each initial literature. Then, it filters the initial literature based on a first similarity to obtain a first experimental literature dataset, which ensures the relevance of figures and text in each initial literature in the first experimental literature dataset. Then, based on a pre-trained figure-text matching model, it converts it to obtain multiple image vectors and multiple paragraph vectors for each first literature in the first experimental literature dataset. Then, it filters and combines them based on a second similarity to obtain a first image-text vector set, which ensures the relevance of images and paragraph text in each first image-text vector pair in the first image-text vector set. Thus, when asking questions to the second large model based on the first image-text vector set to construct the corresponding question-and-answer pair for each vector pair, it ensures the matching degree between the generated question and the current vector pair, thereby improving the accuracy of the photovoltaic cell experimental large model when constructing the photovoltaic cell experimental large model based on the preset experimental question-and-answer dataset.

[0041] In some embodiments of this application, the step of asking questions about a preset second large model based on the first image-text vector set, constructing a corresponding question-and-answer pair for each vector pair in the first image-text vector set, and constructing a preset experimental question-and-answer dataset based on the question-and-answer pairs specifically includes:

[0042] For each vector pair in the first image and text vector set, the second large model is asked a question based on preset question-and-answer prompts to obtain the corresponding question for each vector pair;

[0043] Combine each vector pair in the first set of image and text vectors with the corresponding question to obtain a question-and-answer pair that corresponds one-to-one with each vector pair;

[0044] Combine all question-answer pairs to obtain the preset experimental question-answer dataset.

[0045] This application first constructs a corresponding question for each vector pair in the first image-text vector set based on question-and-answer prompts, and then combines them to obtain question-and-answer pairs, thereby constructing an experimental question-and-answer dataset. This ensures the correspondence between the question and the vector pair, and thus obtains diverse questions based on the diversity of the vector pairs, improving the data richness of the preset experimental question-and-answer dataset.

[0046] According to a third aspect of the embodiments of this application, a question-and-answer device for photovoltaic cell experiments is provided, including a question recognition module, a spectrum retrieval module, a knowledge acquisition module, and a model questioning module;

[0047] The question recognition module is used to recognize the first question text input by the user and obtain domain keywords, technical keywords and keyword weight ratios;

[0048] The graph retrieval module is used to retrieve and filter a preset photovoltaic cell material knowledge graph based on the domain keywords and the technical keywords, combined with a graph walk algorithm, to obtain multiple domain retrieval paths; wherein, the photovoltaic cell material knowledge graph is generated based on the extraction of photovoltaic cell domain data;

[0049] The knowledge acquisition module is used to acquire multiple domain knowledge texts from the photovoltaic cell material knowledge graph according to the multiple domain retrieval paths; wherein, the ratio of texts that satisfy domain keywords to texts that satisfy technical keywords in the multiple domain knowledge texts conforms to the keyword weight ratio;

[0050] The model questioning module is used to ask questions to a preset photovoltaic cell domain model based on the first question text and the multiple domain knowledge texts, and obtain a first answer text; wherein, the photovoltaic cell domain model is constructed based on a preset domain dataset; the preset domain dataset is constructed by extracting photovoltaic cell domain data through multiple agents.

[0051] In some embodiments of this application, the map retrieval module includes a keyword retrieval unit, a keyword filtering unit, and a path verification unit;

[0052] The keyword retrieval unit is used to retrieve the photovoltaic cell material knowledge graph based on the domain keywords to obtain a first domain subgraph;

[0053] The keyword filtering unit is used to filter the first domain subgraph according to the technical keywords to obtain multiple first search paths;

[0054] The path verification unit is used to verify the semantic coherence of the multiple first retrieval paths based on the graph walking algorithm, and to filter the results based on the verification results to obtain multiple domain retrieval paths.

[0055] In some embodiments of this application, the photovoltaic cell material knowledge graph is generated based on the extraction of data in the field of photovoltaic cells, specifically including:

[0056] Based on the preset graph pattern and combined with the preset first major model, the data in the photovoltaic cell field is filtered through preset keywords to obtain multiple text fragments in the first field;

[0057] Based on the different keyword weights of the preset keywords in each first domain text segment, and combined with the number of references of each first domain text segment, the multiple first domain text segments are filtered to obtain multiple second domain text segments;

[0058] Based on the graph pattern and the first large model, knowledge extraction is performed on the multiple second domain text fragments to obtain multiple domain entity subgraphs;

[0059] By combining the semantic similarity among the multiple domain entity subgraphs, disambiguation merging is performed on the multiple domain entity subgraphs to obtain a photovoltaic cell material knowledge graph.

[0060] In some embodiments of this application, the preset domain dataset is obtained by extracting and constructing photovoltaic cell domain data through multiple agents, specifically including:

[0061] By using an information extraction agent in a multi-agent system, structured parsing is performed on photovoltaic cell data to extract multiple pieces of first-domain information. Contextual reasoning is then performed on these multiple pieces of first-domain information to construct multiple information association relationships.

[0062] By using the quality verification agent in the multi-agent system, combined with the photovoltaic cell material knowledge graph and the comparison of the multiple information relationships, the multiple pieces of first domain information are cross-validated and filtered to obtain multiple pieces of second domain information;

[0063] By using a document summarizing agent in a multi-agent system, the multiple pieces of second-domain information are hierarchically summarized to generate multiple multi-level summary information corresponding to the multiple pieces of second-domain information;

[0064] The multiple pieces of second-domain information are integrated with the corresponding multiple multi-level summary information, and then matched and integrated with multiple preset research categories to obtain a preset domain dataset.

[0065] This application first identifies the initial question text input by the user, obtains different types of keywords and their weight ratios, and then searches and filters the photovoltaic cell material knowledge graph based on the keywords to obtain multiple domain retrieval paths. Based on these multiple retrieval paths, it retrieves multiple domain knowledge texts from the photovoltaic cell material knowledge graph. By identifying the user's question, it obtains retrieval keywords and then acquires the prerequisite domain knowledge used for large-scale model questioning. This prerequisite domain knowledge restricts the large-scale model's answers to a specific range, reducing the occurrence of large-scale model illusions. Therefore, when combining the initial question text and multiple domain knowledge texts to ask a question to a large-scale photovoltaic cell model and obtain the first answer text, it improves the accuracy of the large-scale model's question-and-answer function in the photovoltaic cell R&D field. Simultaneously, by extracting photovoltaic cell domain data and using multi-agent extraction, it constructs a photovoltaic cell material knowledge graph and a pre-defined domain dataset, thus obtaining a large-scale photovoltaic cell model. Compared to existing technologies, by adding a knowledge graph with available prerequisite domain knowledge and a domain dataset for training the large-scale model, it reduces the knowledge update cost of the large-scale model, improves its knowledge update efficiency, and thus enhances the accuracy of the large-scale model's question-and-answer function in the photovoltaic cell R&D field.

[0066] According to a fourth aspect of the embodiments of this application, an optimization apparatus for photovoltaic cell experiments is provided, including a result acquisition module and a question optimization module;

[0067] The result acquisition module is used to acquire the user's first experimental result; wherein, the first experimental result is obtained by the user conducting an experiment based on a first experimental suggestion; the first experimental suggestion is obtained through a question-and-answer device for photovoltaic cell experiments;

[0068] The question optimization module is used to ask questions to a preset photovoltaic cell experimental model based on the first experimental results and the first experimental suggestions, obtain a second answer text, and optimize the first experimental suggestions based on the second answer text; wherein, the photovoltaic cell experimental model is constructed based on a preset experimental question and answer dataset.

[0069] In some embodiments of this application, the large-scale experimental model of the photovoltaic cell is constructed based on a pre-set experimental question-and-answer dataset, specifically including:

[0070] Obtain the initial experimental literature dataset, and convert the title of each figure and the text content of each paragraph of each initial literature in the initial experimental literature dataset into vector representations, to obtain multiple figure title vectors and multiple paragraph text vectors for each initial literature;

[0071] Each initial document is traversed. During traversal, the first similarity between the vector of each figure title in the current initial document and the vector of each paragraph text in the current initial document is calculated. Based on the first similarity, the initial documents in the initial experimental document dataset are filtered to obtain the first experimental document dataset.

[0072] Based on the pre-trained chart-text matching model, the text content of each paragraph and each chart of each first document in the first experimental literature dataset are converted into cross-modal vector representations, resulting in multiple image vectors and multiple paragraph vectors for each first document.

[0073] Traverse each first document, and calculate the second similarity between each image vector in the current first document and each paragraph vector in the current first document during the traversal. Combine the image vectors and paragraph vectors whose second similarity exceeds a preset first threshold into a first image-text vector pair.

[0074] Combine several first image-text vector pairs from all first documents to obtain the first image-text vector set;

[0075] Based on the first set of image and text vectors, questions are asked to the second preset model, and corresponding question-and-answer pairs are constructed for each vector pair in the first set of image and text vectors. A preset experimental question-and-answer dataset is then constructed based on the question-and-answer pairs.

[0076] Based on the pre-set experimental question and answer dataset, the initial large-scale photovoltaic cell experimental model is adjusted and trained to obtain the large-scale photovoltaic cell experimental model.

[0077] In some embodiments of this application, the step of asking questions about a preset second large model based on the first image-text vector set, constructing a corresponding question-and-answer pair for each vector pair in the first image-text vector set, and constructing a preset experimental question-and-answer dataset based on the question-and-answer pairs specifically includes:

[0078] For each vector pair in the first image and text vector set, the second large model is asked a question based on preset question-and-answer prompts to obtain the corresponding question for each vector pair;

[0079] Combine each vector pair in the first set of image and text vectors with the corresponding question to obtain a question-and-answer pair that corresponds one-to-one with each vector pair;

[0080] Combine all question-answer pairs to obtain the preset experimental question-answer dataset.

[0081] This application first obtains the first experimental results of the experiment based on the user's question-and-answer session regarding photovoltaic cell experiments, and then asks questions about the large-scale photovoltaic cell experiment model based on the first experimental results and the first experimental suggestions. Based on the question-and-answer results, the first experimental suggestions are optimized. Compared with existing technologies, by training a large-scale model that can analyze experimental results and adding a process of optimizing experimental suggestions based on the large-scale model, it is possible to realize the analysis and reasoning of experimental results, as well as the feedback optimization of experimental suggestions, thereby improving the accuracy of experimental optimization of the large-scale model in the field of photovoltaic cell research and development.

[0082] According to a fifth aspect of the embodiments of this application, a question-and-answer system for photovoltaic cell experiments is provided, including an experiment suggestion module and an experiment optimization module;

[0083] The experiment suggestion module is used to obtain user experiment questions and execute a question-and-answer method for photovoltaic cell experiments as described in this application to generate corresponding user experiment suggestions based on the user experiment questions.

[0084] The experiment optimization module is used to obtain the user experiment results after the user conducts the experiment according to the user experiment suggestion, and execute an optimization method for photovoltaic cell experiment as described in this application, so as to obtain corresponding experiment optimization suggestions based on the user experiment suggestion and the user experiment results. Attached Figure Description

[0085] Figure 1 This is a flowchart illustrating a question-and-answer method for a photovoltaic cell experiment, as shown in some embodiments of this application.

[0086] Figure 2 This is a flowchart illustrating an optimization method for photovoltaic cell experiments according to certain embodiments of this application.

[0087] Figure 3 This is a modular structure diagram of a question-and-answer device for a photovoltaic cell experiment, as shown in some embodiments of this application.

[0088] Figure 4 This is a modular structure diagram of an optimized apparatus for photovoltaic cell experiments shown in certain embodiments of this application;

[0089] Figure 5 This is a system architecture diagram of a question-and-answer system for a photovoltaic cell experiment, as shown in some embodiments of this application. Detailed Implementation

[0090] The embodiments of this application are described in detail below. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below in conjunction with the accompanying drawings are exemplary and are only used to explain some embodiments of this application, and should not be construed as limiting the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments shown in this application without inventive effort are within the protection scope of this application.

[0091] In the description of this application, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, unless otherwise explicitly specified, "a plurality of" or "several" means two or more.

[0092] The existing large-scale models used for photovoltaic cell research and development have the following problems: (1) Current general-purpose large-scale models, such as Deepseek and GPT-4, have limited training data for the photovoltaic cell research and development field. This results in a large amount of resources being spent on full parameter fine-tuning for each knowledge update of the large-scale model. This is not only costly, but also incompatible with the rapid iteration of photovoltaic cell research and development, leading to untimely knowledge updates and affecting the accuracy of the large-scale model's question answering; (2) Although the current large-scale models can obtain the domain question answering ability of photovoltaic cell research and development through fine-tuning, after conducting experiments based on the experimental suggestions of the large-scale model for photovoltaic cell research and development, there is insufficient analysis and reasoning of the experimental results data, and there is also a lack of feedback optimization of the experimental suggestions based on the experimental results. This results in the large-scale model lacking experimental analysis and optimization capabilities, and the accuracy of the large-scale model in conducting experimental analysis and optimization is not high. Therefore, how to improve the accuracy of large-scale models in question answering and experimental optimization in the field of photovoltaic cell research and development is still a technical problem that needs to be solved urgently in the current technology.

[0093] Based on the above technical background, please refer to Figure 1 This application provides a question-and-answer method for photovoltaic cell experiments, including steps S101 to S104, each step as follows:

[0094] Step S101: Identify the first question text entered by the user and obtain the domain keywords, technical keywords, and keyword weight ratios.

[0095] In some embodiments of this application, the identification of the first question text input by the user is based on a pre-trained classification model, and a preferred embodiment of the pre-trained classification model is the BERT classifier.

[0096] The reason for considering identifying the first question text entered by the user and obtaining domain keywords and technical keywords and their weight ratios is that domain keywords summarize the domain corresponding to the question text, such as "stability research of perovskite solar cells" and "luminescence mechanism of perovskite quantum dots", while technical keywords focus on the specific technology corresponding to the question text, such as "thermal stability", "fill factor" and "titanium dioxide". By combining domain keywords and technical keywords for retrieval, retrieval efficiency can be optimized and retrieval accuracy can be improved.

[0097] Step S102: Based on the domain keywords and the technical keywords, a graph walk algorithm is used to search and filter the preset photovoltaic cell material knowledge graph to obtain multiple domain search paths; wherein, the photovoltaic cell material knowledge graph is generated based on the extraction of photovoltaic cell domain data.

[0098] In some embodiments of this application, the step of searching and filtering a preset photovoltaic cell material knowledge graph based on the domain keywords and the technical keywords, combined with a graph walk algorithm, to obtain multiple domain search paths specifically includes:

[0099] Based on the domain keywords, the photovoltaic cell material knowledge graph is retrieved to obtain a first domain subgraph;

[0100] Based on the technical keywords, the first domain subgraph is filtered to obtain multiple first search paths;

[0101] The semantic coherence of the multiple first retrieval paths is verified based on the graph walking algorithm, and the results are filtered to obtain multiple domain retrieval paths.

[0102] In some embodiments of this application, the photovoltaic cell material knowledge graph is retrieved based on the domain keywords to obtain a first domain subgraph. Specifically, the photovoltaic cell material knowledge graph is retrieved based on the domain keywords using a subgraph filtering algorithm to obtain a first domain subgraph.

[0103] This application first retrieves a first domain subgraph from the photovoltaic cell material knowledge graph using domain keywords. This allows for the initial identification of the domain containing the keywords, narrowing the search scope. Then, by combining technical keywords with the first domain subgraph, multiple first search paths are obtained, further narrowing the scope and focusing on the corresponding technologies of the keywords. As a result, when verifying the semantic coherence of the multiple domain search paths based on the graph walk algorithm and filtering to obtain multiple domain search paths, the matching degree between the obtained multiple domain search paths and the current user's query is improved, thereby increasing the accuracy of prior domain knowledge when obtaining prior domain knowledge based on multiple domain search paths.

[0104] In some embodiments of this application, the photovoltaic cell material knowledge graph is generated based on the extraction of data in the field of photovoltaic cells, specifically including:

[0105] Based on the preset graph pattern and combined with the preset first major model, the data in the photovoltaic cell field is filtered through preset keywords to obtain multiple text fragments in the first field;

[0106] Based on the different keyword weights of the preset keywords in each first domain text segment, and combined with the number of references of each first domain text segment, the multiple first domain text segments are filtered to obtain multiple second domain text segments;

[0107] Based on the graph pattern and the first large model, knowledge extraction is performed on the multiple second domain text fragments to obtain multiple domain entity subgraphs;

[0108] By combining the semantic similarity among the multiple domain entity subgraphs, disambiguation merging is performed on the multiple domain entity subgraphs to obtain a photovoltaic cell material knowledge graph.

[0109] In some embodiments of this application, the implementation of the first large model includes, but is not limited to, Wenxin Yiyan, Tongyi Qianwen, Xunfei Xinghuo, Deepseek, Gemini, GPT-4 or Kimi, with Deepseek being the preferred implementation.

[0110] In some embodiments of this application, the step of filtering the multiple first-domain text fragments based on the different keyword weights of the preset keywords in each first-domain text fragment, combined with the citation count of each first-domain text fragment, to obtain multiple second-domain text fragments specifically involves:

[0111] Based on the different keyword weights of the preset keywords in each first domain text segment and the number of times each first domain text segment is cited, the relevance score of each first domain text segment is calculated, and the multiple first domain text segments are filtered based on the relevance scores to obtain multiple second domain text segments.

[0112] Specifically, the relevance score is as follows:

[0113] ;

[0114] in, For weight parameters, To preset keywords in the current first domain text fragment Keyword weight in For the current first domain text fragment The number of times it is cited.

[0115] In the embodiments of this application, the preferred implementation of the semantic similarity among the multiple domain entity subgraphs is cosine similarity; the step of combining the semantic similarity among the multiple domain entity subgraphs to disambiguate and merge the multiple domain entity subgraphs to obtain a photovoltaic cell material knowledge graph specifically involves:

[0116] By disambiguating and merging several domain entity subgraphs whose semantic similarity is higher than a preset semantic similarity threshold, a photovoltaic cell material knowledge graph is obtained.

[0117] This application first uses a preset graph pattern and a preset first major model to filter data in the photovoltaic cell field using preset keywords. This allows for the efficient generation of multiple first-domain text fragments that conform to the graph pattern based on the first major model. Then, based on the weight of preset keywords and the citation frequency of first-domain text fragments, text fragments with low relevance to the photovoltaic cell field are filtered out. Subsequently, based on the graph pattern and the first major model, knowledge extraction is performed to obtain multiple domain entity subgraphs. These subgraphs are then disambiguated and merged based on their semantic similarity, thereby improving the accuracy of the knowledge graph and the quality of the domain knowledge within it.

[0118] Step S103: Based on the multiple domain retrieval paths, obtain multiple domain knowledge texts from the photovoltaic cell material knowledge graph; wherein, the ratio of texts satisfying domain keywords to texts satisfying technical keywords in the multiple domain knowledge texts conforms to the keyword weight ratio.

[0119] In some embodiments of this application, the step of obtaining multiple domain knowledge texts from the photovoltaic cell material knowledge graph based on the multiple domain retrieval paths further includes:

[0120] Based on a pre-defined large language retrieval model, conflicting texts in the multiple domain knowledge texts are marked to obtain multiple sets of conflicting knowledge texts;

[0121] Based on the document age, impact factor, and citation count of each knowledge text in each group of conflicting knowledge texts, the confidence score of each knowledge text is calculated, and the knowledge text with the highest confidence score in each group of conflicting knowledge texts is retained, while knowledge texts with non-highest confidence scores are removed.

[0122] Specifically, the confidence score is as follows:

[0123] ;

[0124] ;

[0125] in, Score the confidence level. The total number of domain knowledge texts. For the current knowledge text, For authoritative rating, The document age of the current knowledge text. As the impact factor of current knowledge texts, The largest influence factor among all knowledge texts in the current group of conflicting knowledge texts. This represents the number of times the current knowledge text has been cited. This represents the total number of references to all knowledge texts within the current group of conflicting knowledge texts.

[0126] Step S104: Based on the first question text and the multiple domain knowledge texts, ask a question to the preset photovoltaic cell domain big model to obtain the first answer text; wherein, the photovoltaic cell domain big model is constructed based on the preset domain dataset; the preset domain dataset is constructed by extracting photovoltaic cell domain data through multiple agents.

[0127] In some embodiments of this application, the implementation of the large model in the field of photovoltaic cells includes, but is not limited to, Wenxin Yiyan, Tongyi Qianwen, Xunfei Xinghuo, Deepseek, Gemini, GPT-4 or Kimi, with Tongyi Qianwen being the preferred implementation.

[0128] In some embodiments of this application, the preset domain dataset is obtained by extracting and constructing photovoltaic cell domain data through multiple agents, specifically including:

[0129] By using an information extraction agent in a multi-agent system, structured parsing is performed on photovoltaic cell data to extract multiple pieces of first-domain information. Contextual reasoning is then performed on these multiple pieces of first-domain information to construct multiple information association relationships.

[0130] By using the quality verification agent in the multi-agent system, combined with the photovoltaic cell material knowledge graph and the comparison of the multiple information relationships, the multiple pieces of first domain information are cross-validated and filtered to obtain multiple pieces of second domain information;

[0131] By using a document summarizing agent in a multi-agent system, the multiple pieces of second-domain information are hierarchically summarized to generate multiple multi-level summary information corresponding to the multiple pieces of second-domain information;

[0132] The multiple pieces of second-domain information are integrated with the corresponding multiple multi-level summary information, and then matched and integrated with multiple preset research categories to obtain a preset domain dataset.

[0133] In some embodiments of this application, the quality verification agent in a multi-agent system, combined with the photovoltaic cell material knowledge graph and the comparison of the multiple information associations, performs cross-validation and filtering on the multiple pieces of first-domain information to obtain multiple pieces of second-domain information, specifically as follows:

[0134] By using a quality verification agent in a multi-agent system, combined with the photovoltaic cell material knowledge graph and the comparison of the multiple information relationships, the confidence score of each piece of first domain information is calculated, and the multiple pieces of first domain information are cross-validated and filtered based on the confidence score to obtain multiple pieces of second domain information;

[0135] Specifically, the confidence score is as follows:

[0136] ;

[0137] ;

[0138] in, These are the weighting coefficients. For internal consistency scoring, For external validation scoring, The degree of matching between the photovoltaic cell material knowledge graph and the information related to the current primary field; This is the set of relevant statements in the literature containing the information in the current primary domain. For the current first domain information and related statements semantic similarity, For related statements The relevance weights.

[0139] In some embodiments of this application, The preferred implementation method is external manual annotation and scoring. The preferred implementation is to calculate the matching degree of the information association relationship between the photovoltaic cell material knowledge graph and the current first domain information based on similarity.

[0140] This application first uses an information extraction agent in a multi-agent system to perform structured parsing of photovoltaic cell data to extract multiple pieces of first-domain information and constructs multiple information associations through contextual reasoning. Then, a quality verification agent in the multi-agent system performs cross-validation and filtering by combining a photovoltaic cell material knowledge graph with the comparison of multiple information associations to obtain multiple pieces of second-domain information. This can filter out information in the first-domain information that has low relevance to photovoltaic cell materials, improving the accuracy of the information. Next, a document summarization agent in the multi-agent system performs hierarchical summarization to generate multiple multi-level summary information. This is then matched and integrated with multiple pieces of second-domain information and multiple research categories to obtain a preset domain dataset. This adds a summary dimension to the second-domain information, increases information density, and improves the matching degree between the preset domain dataset and the photovoltaic cell field. Therefore, when constructing a large-scale photovoltaic cell model based on the preset domain dataset, the large model is more closely matched to the photovoltaic cell field, thereby improving the accuracy of the large model in photovoltaic cell experimental question answering.

[0141] Compared to existing technologies, this application first identifies the user's initial question text, obtains different types of keywords and their weight ratios, and then searches and filters the photovoltaic cell material knowledge graph based on the keywords to obtain multiple domain retrieval paths. Based on these multiple retrieval paths, it retrieves multiple domain knowledge texts from the photovoltaic cell material knowledge graph. By identifying the user's question, it obtains retrieval keywords and then acquires prior domain knowledge for large-scale model questioning. This prior domain knowledge restricts the large-scale model's answers to a specific range, reducing the generation of large-scale model illusions. Therefore, when combining the initial question text and multiple domain knowledge texts to ask a question to a large-scale photovoltaic cell model and obtain the first answer text, it improves the accuracy of the large-scale model's question-and-answer function in the photovoltaic cell R&D field. Simultaneously, by extracting photovoltaic cell domain data and using multi-agent extraction, it constructs a photovoltaic cell material knowledge graph and a pre-defined domain dataset, thus obtaining a large-scale photovoltaic cell model. Compared to existing technologies, by adding a knowledge graph with available prior domain knowledge and a domain dataset for training the large-scale model, it reduces the knowledge update cost and improves the knowledge update efficiency of the large-scale model, thereby improving the accuracy of the large-scale model's question-and-answer function in the photovoltaic cell R&D field.

[0142] Please see Figure 2 The present application also provides an optimization method for photovoltaic cell experiments, including steps S201 to S202, the specific steps of which are as follows:

[0143] Step S201: Obtain the user's first experimental result; wherein, the first experimental result is obtained by the user conducting experiments on the first experimental suggestion; the first experimental suggestion is obtained through a question-and-answer method for photovoltaic cell experiments as described in this application.

[0144] Step S202: Based on the first experimental results and the first experimental suggestion, ask questions to the preset photovoltaic cell experimental model to obtain the second answer text, and optimize the first experimental suggestion based on the second answer text; wherein, the photovoltaic cell experimental model is constructed based on the preset experimental question and answer dataset.

[0145] In some embodiments of this application, the implementation of the large-scale photovoltaic cell experimental model includes, but is not limited to, Wenxin Yiyan, Tongyi Qianwen, Xunfei Xinghuo, Deepseek, Gemini, GPT-4, or Kimi, with Tongyi Qianwen being the preferred implementation.

[0146] In some embodiments of this application, the large-scale experimental model of the photovoltaic cell is constructed based on a pre-set experimental question-and-answer dataset, specifically including:

[0147] Obtain the initial experimental literature dataset and the initial experimental literature dataset Each initial document Title of each chart and the text content of each paragraph Convert each of these values ​​into vector representations to obtain multiple figure and table title vectors for each initial document. and multiple paragraph text vectors ;

[0148] Traverse each initial document During traversal, calculate the vector of each figure title in the current initial document. Compared with the text vector of each paragraph in the current initial document First similarity And based on the first similarity For the initial experimental literature dataset The initial document in After filtering, the first experimental literature dataset was obtained. ;

[0149] Based on a pre-trained chart-text matching model, the first experimental literature dataset was used. The text content of each paragraph and each chart in each first-level document are converted into cross-modal vector representations, resulting in multiple image vectors and multiple paragraph vectors for each first-level document;

[0150] Traverse each first document, and calculate the second similarity between each image vector in the current first document and each paragraph vector in the current first document during the traversal. Combine the image vectors and paragraph vectors whose second similarity exceeds a preset first threshold into a first image-text vector pair.

[0151] Combine several first image-text vector pairs from all first documents to obtain the first image-text vector set;

[0152] Based on the first set of image and text vectors, questions are asked to the second preset model, and corresponding question-and-answer pairs are constructed for each vector pair in the first set of image and text vectors. A preset experimental question-and-answer dataset is then constructed based on the question-and-answer pairs.

[0153] Based on the pre-set experimental question and answer dataset, the initial large-scale photovoltaic cell experimental model is adjusted and trained to obtain the large-scale photovoltaic cell experimental model.

[0154] In some embodiments of this application, the step of filtering the initial documents in the initial experimental literature dataset based on the first similarity to obtain the first experimental literature dataset specifically involves:

[0155] Filter out the paragraph text set corresponding to each figure title vector in the current initial document, wherein each element in the paragraph text set satisfies the first similarity with the corresponding figure title vector being higher than a preset threshold;

[0156] If the text set of each paragraph of the current initial document is empty, the current initial document will be removed from the initial experimental document dataset.

[0157] In some embodiments of this application, the preferred embodiment of the chart text matching model is the CLIP model.

[0158] This application first obtains an initial experimental literature dataset and converts it to obtain multiple figure and table title vectors and multiple paragraph text vectors for each initial literature. Then, it filters the initial literature based on a first similarity to obtain a first experimental literature dataset, which ensures the relevance of figures and text in each initial literature in the first experimental literature dataset. Then, based on a pre-trained figure-text matching model, it converts it to obtain multiple image vectors and multiple paragraph vectors for each first literature in the first experimental literature dataset. Then, it filters and combines them based on a second similarity to obtain a first image-text vector set, which ensures the relevance of images and paragraph text in each first image-text vector pair in the first image-text vector set. Thus, when asking questions to the second large model based on the first image-text vector set to construct the corresponding question-and-answer pair for each vector pair, it ensures the matching degree between the generated question and the current vector pair, thereby improving the accuracy of the photovoltaic cell experimental large model when constructing the photovoltaic cell experimental large model based on the preset experimental question-and-answer dataset.

[0159] In some embodiments of this application, the step of asking questions about a preset second large model based on the first image-text vector set, constructing a corresponding question-and-answer pair for each vector pair in the first image-text vector set, and constructing a preset experimental question-and-answer dataset based on the question-and-answer pairs specifically includes:

[0160] For each vector pair in the first image and text vector set, the second large model is asked a question based on preset question-and-answer prompts to obtain the corresponding question for each vector pair;

[0161] Combine each vector pair in the first set of image and text vectors with the corresponding question to obtain a question-and-answer pair that corresponds one-to-one with each vector pair;

[0162] Combine all question-answer pairs to obtain the preset experimental question-answer dataset.

[0163] This application first constructs a corresponding question for each vector pair in the first image-text vector set based on question-and-answer prompts, and then combines them to obtain question-and-answer pairs, thereby constructing an experimental question-and-answer dataset. This ensures the correspondence between the question and the vector pair, and thus obtains diverse questions based on the diversity of the vector pairs, improving the data richness of the preset experimental question-and-answer dataset.

[0164] Compared to existing technologies, this application first obtains the first experimental results of the experiment based on the user's question-and-answer session regarding photovoltaic cell experiments, and then asks questions about the large-scale photovoltaic cell experiment model based on the first experimental results and the first experimental suggestions. Based on the question-and-answer results, the first experimental suggestions are optimized. Compared to existing technologies, by training a large-scale model capable of analyzing experimental results and adding a process for optimizing experimental suggestions based on the large-scale model, it is possible to achieve analysis and reasoning of experimental results, as well as feedback optimization of experimental suggestions, thereby improving the accuracy of experimental optimization of the large-scale model in the field of photovoltaic cell research and development.

[0165] For a question-and-answer method corresponding to the one described above, please refer to [link / reference]. Figure 3 This application provides a question-and-answer device for photovoltaic cell experiments, including a question recognition module 310, a spectrum retrieval module 320, a knowledge acquisition module 330, and a model questioning module 340.

[0166] The question recognition module 310 is used to recognize the first question text input by the user and obtain domain keywords, technical keywords and keyword weight ratios;

[0167] The graph retrieval module 320 is used to retrieve and filter a preset photovoltaic cell material knowledge graph based on the domain keywords and the technical keywords, combined with a graph walk algorithm, to obtain multiple domain retrieval paths; wherein, the photovoltaic cell material knowledge graph is generated based on the extraction of photovoltaic cell domain data;

[0168] The knowledge acquisition module 330 is used to acquire multiple domain knowledge texts from the photovoltaic cell material knowledge graph according to the multiple domain retrieval paths; wherein, the ratio of texts that satisfy domain keywords and texts that satisfy technical keywords in the multiple domain knowledge texts conforms to the keyword weight ratio;

[0169] The model questioning module 340 is used to ask questions to a preset photovoltaic cell domain model based on the first question text and the multiple domain knowledge texts, and obtain a first answer text; wherein, the photovoltaic cell domain model is constructed based on a preset domain dataset; the preset domain dataset is constructed by extracting and constructing photovoltaic cell domain data through multiple agents.

[0170] In some embodiments of this application, the map retrieval module 320 includes a keyword retrieval unit, a keyword filtering unit, and a path verification unit;

[0171] The keyword retrieval unit is used to retrieve the photovoltaic cell material knowledge graph based on the domain keywords to obtain a first domain subgraph;

[0172] The keyword filtering unit is used to filter the first domain subgraph according to the technical keywords to obtain multiple first search paths;

[0173] The path verification unit is used to verify the semantic coherence of the multiple first retrieval paths based on the graph walking algorithm, and to filter the results based on the verification results to obtain multiple domain retrieval paths.

[0174] In some embodiments of this application, the photovoltaic cell material knowledge graph is generated based on the extraction of data in the field of photovoltaic cells, specifically including:

[0175] Based on the preset graph pattern and combined with the preset first major model, the data in the photovoltaic cell field is filtered through preset keywords to obtain multiple text fragments in the first field;

[0176] Based on the different keyword weights of the preset keywords in each first domain text segment, and combined with the number of references of each first domain text segment, the multiple first domain text segments are filtered to obtain multiple second domain text segments;

[0177] Based on the graph pattern and the first large model, knowledge extraction is performed on the multiple second domain text fragments to obtain multiple domain entity subgraphs;

[0178] By combining the semantic similarity among the multiple domain entity subgraphs, disambiguation merging is performed on the multiple domain entity subgraphs to obtain a photovoltaic cell material knowledge graph.

[0179] In some embodiments of this application, the preset domain dataset is obtained by extracting and constructing photovoltaic cell domain data through multiple agents, specifically including:

[0180] By using an information extraction agent in a multi-agent system, structured parsing is performed on photovoltaic cell data to extract multiple pieces of first-domain information. Contextual reasoning is then performed on these multiple pieces of first-domain information to construct multiple information association relationships.

[0181] By using the quality verification agent in the multi-agent system, combined with the photovoltaic cell material knowledge graph and the comparison of the multiple information relationships, the multiple pieces of first domain information are cross-validated and filtered to obtain multiple pieces of second domain information;

[0182] By using a document summarizing agent in a multi-agent system, the multiple pieces of second-domain information are hierarchically summarized to generate multiple multi-level summary information corresponding to the multiple pieces of second-domain information;

[0183] The multiple pieces of second-domain information are integrated with the corresponding multiple multi-level summary information, and then matched and integrated with multiple preset research categories to obtain a preset domain dataset.

[0184] This application first identifies the initial question text input by the user, obtains different types of keywords and their weight ratios, and then searches and filters the photovoltaic cell material knowledge graph based on the keywords to obtain multiple domain retrieval paths. Based on these multiple retrieval paths, it retrieves multiple domain knowledge texts from the photovoltaic cell material knowledge graph. By identifying the user's question, it obtains retrieval keywords and then acquires the prerequisite domain knowledge used for large-scale model questioning. This prerequisite domain knowledge restricts the large-scale model's answers to a specific range, reducing the occurrence of large-scale model illusions. Therefore, when combining the initial question text and multiple domain knowledge texts to ask a question to a large-scale photovoltaic cell model and obtain the first answer text, it improves the accuracy of the large-scale model's question-and-answer function in the photovoltaic cell R&D field. Simultaneously, by extracting photovoltaic cell domain data and using multi-agent extraction, it constructs a photovoltaic cell material knowledge graph and a pre-defined domain dataset, thus obtaining a large-scale photovoltaic cell model. Compared to existing technologies, by adding a knowledge graph with available prerequisite domain knowledge and a domain dataset for training the large-scale model, it reduces the knowledge update cost of the large-scale model, improves its knowledge update efficiency, and thus enhances the accuracy of the large-scale model's question-and-answer function in the photovoltaic cell R&D field.

[0185] For a corresponding optimization method, please refer to [link / reference]. Figure 4 The present application also provides an optimization device for photovoltaic cell experiments, including a result acquisition module 410 and a question optimization module 420.

[0186] The result acquisition module 410 is used to acquire the user's first experimental result; wherein, the first experimental result is obtained by the user conducting an experiment based on the first experimental suggestion; the first experimental suggestion is obtained through a question-and-answer device for photovoltaic cell experiments as described in this application;

[0187] The question optimization module 420 is used to ask questions to a preset photovoltaic cell experimental model based on the first experimental results and the first experimental suggestions, obtain a second answer text, and optimize the first experimental suggestions based on the second answer text; wherein, the photovoltaic cell experimental model is constructed based on a preset experimental question and answer dataset.

[0188] In some embodiments of this application, the large-scale experimental model of the photovoltaic cell is constructed based on a pre-set experimental question-and-answer dataset, specifically including:

[0189] Obtain the initial experimental literature dataset, and convert the title of each figure and the text content of each paragraph of each initial literature in the initial experimental literature dataset into vector representations, to obtain multiple figure title vectors and multiple paragraph text vectors for each initial literature;

[0190] Each initial document is traversed. During traversal, the first similarity between the vector of each figure title in the current initial document and the vector of each paragraph text in the current initial document is calculated. Based on the first similarity, the initial documents in the initial experimental document dataset are filtered to obtain the first experimental document dataset.

[0191] Based on the pre-trained chart-text matching model, the text content of each paragraph and each chart of each first document in the first experimental literature dataset are converted into cross-modal vector representations, resulting in multiple image vectors and multiple paragraph vectors for each first document.

[0192] Traverse each first document, and calculate the second similarity between each image vector in the current first document and each paragraph vector in the current first document during the traversal. Combine the image vectors and paragraph vectors whose second similarity exceeds a preset first threshold into a first image-text vector pair.

[0193] Combine several first image-text vector pairs from all first documents to obtain the first image-text vector set;

[0194] Based on the first set of image and text vectors, questions are asked to the second preset model, and corresponding question-and-answer pairs are constructed for each vector pair in the first set of image and text vectors. A preset experimental question-and-answer dataset is then constructed based on the question-and-answer pairs.

[0195] Based on the pre-set experimental question and answer dataset, the initial large-scale photovoltaic cell experimental model is adjusted and trained to obtain the large-scale photovoltaic cell experimental model.

[0196] In some embodiments of this application, the step of asking questions about a preset second large model based on the first image-text vector set, constructing a corresponding question-and-answer pair for each vector pair in the first image-text vector set, and constructing a preset experimental question-and-answer dataset based on the question-and-answer pairs specifically includes:

[0197] For each vector pair in the first image and text vector set, the second large model is asked a question based on preset question-and-answer prompts to obtain the corresponding question for each vector pair;

[0198] Combine each vector pair in the first set of image and text vectors with the corresponding question to obtain a question-and-answer pair that corresponds one-to-one with each vector pair;

[0199] Combine all question-answer pairs to obtain the preset experimental question-answer dataset.

[0200] This application first obtains the first experimental results of the experiment based on the user's question-and-answer session regarding photovoltaic cell experiments, and then asks questions about the large-scale photovoltaic cell experiment model based on the first experimental results and the first experimental suggestions. Based on the question-and-answer results, the first experimental suggestions are optimized. Compared with existing technologies, by training a large-scale model that can analyze experimental results and adding a process of optimizing experimental suggestions based on the large-scale model, it is possible to realize the analysis and reasoning of experimental results, as well as the feedback optimization of experimental suggestions, thereby improving the accuracy of experimental optimization of the large-scale model in the field of photovoltaic cell research and development.

[0201] Adaptively, please see Figure 5 The present application also provides a question-and-answer system for photovoltaic cell experiments, including an experiment suggestion module 510 and an experiment optimization module 520.

[0202] The experiment suggestion module 510 is used to obtain user experiment questions and execute a question-and-answer method for photovoltaic cell experiments as described in this application to generate corresponding user experiment suggestions based on the user experiment questions.

[0203] The experiment optimization module 520 is used to obtain the user experiment results after the user conducts the experiment according to the user experiment suggestion, and execute an optimization method for photovoltaic cell experiment as described in this application, so as to obtain corresponding experiment optimization suggestions based on the user experiment suggestion and the user experiment results.

[0204] It should be understood that the apparatus provided in the embodiments of this application corresponds to the aforementioned method. The question-and-answer apparatus for photovoltaic cell experiments provided in the embodiments of this application can implement the question-and-answer method for photovoltaic cell experiments provided in any embodiment of this application.

[0205] Adaptively, embodiments of this application also provide a computer device and a computer-readable storage medium.

[0206] The computer device includes: a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor;

[0207] When the processor executes the computer program, it implements a question-and-answer method for a photovoltaic cell experiment or an optimization method for a photovoltaic cell experiment as described in this application.

[0208] The computer-readable storage medium stores multiple instructions adapted for loading by a processor to execute a question-and-answer method for a photovoltaic cell experiment or an optimization method for a photovoltaic cell experiment according to this application.

[0209] The above description represents some embodiments of this application, providing a further detailed explanation of the purpose, technical solution, and beneficial effects of this application. It should be understood that the above-described embodiments of this application should not be construed as limiting this application. In particular, any changes, modifications, equivalent substitutions, and variations made by those skilled in the art within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A question-and-answer method for photovoltaic cell experiments, characterized in that, include: Identify the first question text entered by the user and obtain domain keywords, technical keywords, and keyword weight ratios; Based on the domain keywords and technical keywords, a graph walking algorithm is used to search and filter a preset photovoltaic cell material knowledge graph to obtain multiple domain search paths. Specifically, the photovoltaic cell material knowledge graph is searched based on the domain keywords to obtain a first domain subgraph; the first domain subgraph is filtered based on the technical keywords to obtain multiple first search paths; the semantic coherence of the multiple first search paths is verified based on the graph walking algorithm, and the paths are filtered based on the verification results to obtain multiple domain search paths; wherein, the photovoltaic cell material knowledge graph is generated based on the extraction of photovoltaic cell domain data. Based on the multiple domain retrieval paths, multiple domain knowledge texts are obtained from the photovoltaic cell material knowledge graph; wherein, the ratio of texts satisfying domain keywords to texts satisfying technical keywords in the multiple domain knowledge texts conforms to the keyword weight ratio; Based on the first question text and the multiple domain knowledge texts, a question is posed to the preset photovoltaic cell domain model to obtain the first answer text; wherein, the photovoltaic cell domain model is constructed based on the preset domain dataset; the preset domain dataset is constructed by extracting photovoltaic cell domain data through multiple agents.

2. The question-and-answer method for photovoltaic cell experiments according to claim 1, characterized in that, The photovoltaic cell material knowledge graph is generated based on the extraction of data in the photovoltaic cell field, and specifically includes: Based on the preset graph pattern and combined with the preset first major model, the data in the photovoltaic cell field is filtered through preset keywords to obtain multiple text fragments in the first field; Based on the different keyword weights of the preset keywords in each first domain text segment, and combined with the number of references of each first domain text segment, the multiple first domain text segments are filtered to obtain multiple second domain text segments; Based on the graph pattern and the first large model, knowledge extraction is performed on the multiple second domain text fragments to obtain multiple domain entity subgraphs; By combining the semantic similarity among the multiple domain entity subgraphs, disambiguation merging is performed on the multiple domain entity subgraphs to obtain a photovoltaic cell material knowledge graph.

3. The question-and-answer method for photovoltaic cell experiments according to claim 1, characterized in that, The preset domain dataset is obtained by extracting and constructing photovoltaic cell domain data through multiple agents, specifically including: By using an information extraction agent in a multi-agent system, structured parsing is performed on photovoltaic cell data to extract multiple pieces of first-domain information. Contextual reasoning is then performed on these multiple pieces of first-domain information to construct multiple information association relationships. By using the quality verification agent in the multi-agent system, combined with the photovoltaic cell material knowledge graph and the comparison of the multiple information relationships, the multiple pieces of first domain information are cross-validated and filtered to obtain multiple pieces of second domain information; By using a document summarizing agent in a multi-agent system, the multiple pieces of second-domain information are hierarchically summarized to generate multiple multi-level summary information corresponding to the multiple pieces of second-domain information; The multiple pieces of second-domain information are integrated with the corresponding multiple multi-level summary information, and then matched and integrated with multiple preset research categories to obtain a preset domain dataset.

4. An optimized method for photovoltaic cell experiments, characterized in that, include: Obtain the user's first experimental result; wherein the first experimental result is obtained by the user conducting an experiment based on the first experimental suggestion; the first experimental suggestion is obtained through a question-and-answer method for photovoltaic cell experiments as described in any one of claims 1 to 3; Based on the first experimental results and the first experimental suggestion, a question is posed to the preset photovoltaic cell experimental model to obtain a second answer text, and the first experimental suggestion is optimized based on the second answer text; wherein, the photovoltaic cell experimental model is constructed based on the preset experimental question and answer dataset.

5. The method for optimizing a photovoltaic cell experiment according to claim 4, characterized in that, The large-scale experimental model for photovoltaic cells is constructed based on a pre-set experimental question-and-answer dataset, specifically including: Obtain the initial experimental literature dataset, and convert the title of each figure and the text content of each paragraph of each initial literature in the initial experimental literature dataset into vector representations, to obtain multiple figure title vectors and multiple paragraph text vectors for each initial literature; Each initial document is traversed. During traversal, the first similarity between the vector of each figure title in the current initial document and the vector of each paragraph text in the current initial document is calculated. Based on the first similarity, the initial documents in the initial experimental document dataset are filtered to obtain the first experimental document dataset. Based on the pre-trained chart-text matching model, the text content of each paragraph and each chart of each first document in the first experimental literature dataset are converted into cross-modal vector representations, resulting in multiple image vectors and multiple paragraph vectors for each first document. Traverse each first document, and calculate the second similarity between each image vector in the current first document and each paragraph vector in the current first document during the traversal. Combine the image vectors and paragraph vectors whose second similarity exceeds a preset first threshold into a first image-text vector pair. Combine several first image-text vector pairs from all first documents to obtain the first image-text vector set; Based on the first set of image and text vectors, questions are asked to the second preset model, and corresponding question-and-answer pairs are constructed for each vector pair in the first set of image and text vectors. A preset experimental question-and-answer dataset is then constructed based on the question-and-answer pairs. Based on the pre-set experimental question and answer dataset, the initial large-scale photovoltaic cell experimental model is adjusted and trained to obtain the large-scale photovoltaic cell experimental model.

6. The optimization method for photovoltaic cell experiments according to claim 5, characterized in that, The step of asking questions to a preset second model based on the first set of image and text vectors, constructing a corresponding question-and-answer pair for each vector pair in the first set of image and text vectors, and constructing a preset experimental question-and-answer dataset based on the question-and-answer pairs specifically includes: For each vector pair in the first image and text vector set, the second large model is asked a question based on preset question-and-answer prompts to obtain the corresponding question for each vector pair; Combine each vector pair in the first set of image and text vectors with the corresponding question to obtain a question-and-answer pair that corresponds one-to-one with each vector pair; Combine all question-answer pairs to obtain the preset experimental question-answer dataset.

7. A question-and-answer device for photovoltaic cell experiments, characterized in that, It includes a question recognition module, a graph retrieval module, a knowledge acquisition module, and a model questioning module; The question recognition module is used to recognize the first question text input by the user and obtain domain keywords, technical keywords and keyword weight ratios; The graph retrieval module is used to retrieve and filter a preset photovoltaic cell material knowledge graph based on the domain keywords and the technical keywords, combined with a graph walk algorithm, to obtain multiple domain retrieval paths; wherein, the photovoltaic cell material knowledge graph is generated based on the extraction of photovoltaic cell domain data; The knowledge acquisition module is used to acquire multiple domain knowledge texts from the photovoltaic cell material knowledge graph according to the multiple domain retrieval paths; wherein, the ratio of texts that satisfy domain keywords to texts that satisfy technical keywords in the multiple domain knowledge texts conforms to the keyword weight ratio; The model questioning module is used to ask questions to a preset photovoltaic cell domain model based on the first question text and the multiple domain knowledge texts, and obtain a first answer text; wherein, the photovoltaic cell domain model is constructed based on a preset domain dataset; the preset domain dataset is constructed by extracting and constructing photovoltaic cell domain data through multiple agents; The graph retrieval module includes a keyword retrieval unit, a keyword filtering unit, and a path verification unit. The keyword retrieval unit is used to retrieve the photovoltaic cell material knowledge graph based on the domain keywords to obtain a first domain subgraph. The keyword filtering unit is used to filter the first domain subgraph based on the technical keywords to obtain multiple first retrieval paths. The path verification unit is used to verify the semantic coherence of the multiple first retrieval paths based on a graph walk algorithm, and to filter them based on the verification results to obtain multiple domain retrieval paths.

8. An optimized apparatus for photovoltaic cell experiments, characterized in that, This includes a results acquisition module and a question optimization module; The result acquisition module is used to acquire the user's first experimental result; wherein, the first experimental result is obtained by the user conducting an experiment based on the first experimental suggestion; the first experimental suggestion is obtained through a question-and-answer device for a photovoltaic cell experiment as described in claim 7; The question optimization module is used to ask questions to a preset photovoltaic cell experimental model based on the first experimental results and the first experimental suggestions, obtain a second answer text, and optimize the first experimental suggestions based on the second answer text; wherein, the photovoltaic cell experimental model is constructed based on a preset experimental question and answer dataset.

9. A question-and-answer system for photovoltaic cell experiments, characterized in that, Includes an experiment suggestion module and an experiment optimization module; The experiment suggestion module is used to obtain user experiment questions and execute a question-and-answer method for photovoltaic cell experiments as described in any one of claims 1 to 3, so as to generate corresponding user experiment suggestions based on the user experiment questions. The experiment optimization module is used to obtain the user experiment results after the user conducts the experiment according to the user experiment suggestion, and execute the photovoltaic cell experiment optimization method as described in any one of claims 4 to 6, so as to obtain corresponding experiment optimization suggestions based on the user experiment suggestion and the user experiment results.

Citation Information

Patent Citations

  • Multi-hop agricultural question-answering system based on knowledge graph reasoning and large model and text fragment retrieval and answer generation method thereof

    CN119782467A

  • Knowledge graph and vector retrieval enhancement-based photovoltaic field large model efficiency improvement method

    CN119988639A