Research related data summary processing system
Patent Information
- Application Number
- KR1020250013474
- Authority / Receiving Office
- KR · KR
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-02-04
- Publication Date
- 2026-08-11
Smart Images

Figure PAT00004_ABST
Abstract
Description
Technology Field
[0001] The present invention relates to the field of natural language processing technology. It also relates to the field of database-based document management technology. Background Technology
[0002] With the advancement of technology, vast amounts of research-related data are accumulating. Consequently, difficulties are arising in systematically recording and classifying this data, as well as in quickly searching for the summary results researchers need.
[0003] In the past, experimental processes, results, and ideas were primarily recorded by hand to document research data. Although research data has been digitized today, a significant amount of it is still written by hand. The process of recording data manually consumes a considerable amount of time and resources. Furthermore, there is the issue of low readability due to variations in the writer's handwriting. Additionally, there is the problem of difficulty in preserving data for the long term due to the risk of document damage.
[0004] In the classification process of research-related data, there is a problem in that it is difficult to effectively link related data due to the vast volume. Additionally, since classification criteria may be applied differently by each researcher, there is a demand for a system that reduces the time consumed in the classification process.
[0005] In the process of summarizing research-related data, there is a problem in that the desired results may vary depending on the user's purpose of the summary. Additionally, since it is difficult for researchers to familiarize themselves with all the data, it is challenging to accurately extract the intended summary results. Furthermore, because research-related data have different structures, the structure of the summary results may also vary, making it difficult to read and understand them quickly. The problem to be solved
[0006] The present invention aims to reduce the time and resources consumed in the process of manually recording data.
[0007] In addition, it aims to resolve readability issues that may arise depending on the author.
[0008] In addition, it aims to solve the problem of difficulty in preserving research data for the long term due to document damage.
[0009] In addition, it aims to provide a system that effectively examines similarities between data during the data classification process.
[0010] Furthermore, it aims to reduce the time and resources required for the data classification process. Along with this, it aims to resolve the issue of diversity in classification criteria based on users.
[0011] In addition, it aims to solve the problem where the user's summary purpose can be diverse.
[0012] In addition, it aims to effectively reflect the user's intentions during the summarization process.
[0013] In addition, it aims to resolve the readability issues of summary results caused by unstructured data.
[0014] However, these tasks are exemplary and do not limit the scope of the invention. means of solving the problem
[0015] The present invention relates to a method for generating a research summary report based on one or more research-related data—the one or more research-related data being stored in a database or input by a user—comprising: receiving a user query; extracting one or more keywords from the one or more research-related data; and generating a research summary report based on the one or more research-related data and the one or more keywords. The step of generating the research summary report comprises: searching for at least one of the one or more research-related data based on the user query; searching for at least one of the one or more keywords based on the user query; and generating a research summary report through an artificial intelligence model for generating summary reports based on the at least one of the searched research-related data and the at least one of the searched keywords.
[0016] In addition, the format of the above one or more research-related data is multimodal.
[0017] In addition, one or more of the extracted keywords include metadata.
[0018] Additionally, the step of generating the above research summary report further includes the step of generating an extended query based on the input user query.
[0019] In addition, one or more of the above research-related data include embedding vector data.
[0020] Additionally, the step of generating the research summary report further includes the step of generating an extended query based on the input user query; and the step of generating the extended query further includes the step of inputting the user query into an extended query generating AI model; and the step of the extended query generating AI model generating an extended query based on the embedding vector data; and in the step of searching for at least one research-related data; and the step of searching for at least one keyword, wherein the searched at least one research-related data and the searched at least one keyword are searched based on the generated extended query.
[0021] Additionally, the step of generating an extended query further includes the step in which the extended query generating AI model is a natural language processing-based AI model.
[0022] In addition, the above natural language processing-based artificial intelligence model is at least one of Transformer-based, Seq2Seq, and Hybrid LLMs, and the LLM includes a Refine algorithm or a Map-Reduce algorithm.
[0023] In addition, the above Transformer-based LLM is at least one of GPT, Gemini, LLaMA, and Claude LLM.
[0024] In addition, the AI model for generating the summary report mentioned above is a natural language processing-based AI model.
[0025] In addition, the above natural language processing-based artificial intelligence model is at least one of Transformer-based, Seq2Seq, and Hybrid LLMs.
[0026] In addition, the above Transformer-based LLM is at least one of GPT, Gemini, LLaMA, and Claude LLM.
[0027] The present invention relates to a program for generating a research summary report based on one or more research-related data—said that the one or more research-related data are stored in a database or are entered by a user—and includes an instruction for performing at least one of the methods of claims 1 to 6.
[0028] The present invention relates to an apparatus for generating a research summary report based on one or more research-related data—said that the one or more research-related data are stored in a database or are input by a user—comprising: a data recording module in which the one or more research-related data are recorded; and a data summary module that performs at least one of the methods of claims 1 to 6. Effects of the invention
[0029] The present invention can provide the effect of reducing the time and resources consumed in the process of manually recording data.
[0030] The present invention can solve readability issues that may occur depending on the author and the problem of difficulty in preserving research data for a long period due to document damage.
[0031] The present invention provides a system that effectively examines similarity between data during the data classification process, thereby providing the effect of reducing the time and resources required during the data classification process.
[0032] The present invention can solve the problem of diversity in classification criteria according to the user and the problem of diversity in summary purposes according to the user.
[0033] The present invention can effectively reflect the user's intent during the summarization process and solve the readability problem of the summary result caused by unstructured data. Brief explanation of the drawing
[0034] FIG. 1 is a block diagram illustrating the configuration of a research-related data summary processing system according to one embodiment of the invention. FIG. 2 is a flowchart of the data recording process of a data recording module according to one embodiment of the invention. FIG. 3 is a diagram illustrating the configuration of a data recording module according to an embodiment of the invention. FIG. 4 is a flowchart of the data summarization process of a data summarization module according to one embodiment of the invention. FIG. 5 is a diagram illustrating the configuration of a data summary module according to an embodiment of the invention. FIG. 6 is a diagram illustrating the result obtained by undergoing the step of expanding a query through an expanded query generation artificial intelligence model according to one embodiment of the invention. FIGS. 7a and 7b are drawings for explaining a summary result generated based on a query input into a data summary module according to an embodiment of the invention. Specific details for implementing the invention
[0035] The present invention is capable of various modifications and may have various embodiments; specific embodiments are illustrated in the drawings and described in detail in the detailed description. The effects and features of the present invention, and the methods for achieving them, will become clear by referring to the embodiments described below in detail together with the drawings. However, the present invention is not limited to the embodiments disclosed below but can be implemented in various forms.
[0036] Hereinafter, embodiments of the present invention will be described in detail with reference to the attached drawings. When describing with reference to the drawings, identical or corresponding components are given the same reference numerals, and redundant descriptions thereof will be omitted.
[0037] Additionally, terms such as "module" as described in the specification refer to a unit that processes at least one function or operation, which may be implemented in hardware or software, or as a combination of hardware and software.
[0038] In addition, terms such as "research-related data" may refer to information generated or used during the research process. For example, "research-related data" may include data collected from research experiments, data analyzed based on collected data, data documenting the research process and results, data referenced or utilized in research, and code and algorithm data used for data analysis and processing.
[0039] In addition, "research-related data" may include data such as research notes. Furthermore, the meaning of terms such as "keyword" is not limited to keywords but may include context, such as the background, semantic relationships, or context of the text.
[0040] FIG. 1 is a block diagram illustrating the configuration of a research-related data summary processing system (10) according to one embodiment of the invention.
[0041] Referring to FIG. 1, the research-related data summary processing system (10) may include a data recording module (100) and a data summary module (200).
[0042] According to an embodiment of the present invention, the original data and processed data of a research notebook may be stored in a data recording module (100). For example, the text conversion results of the research notebook through OCR or Parser may be stored in the data recording module (100). Accordingly, based on the data stored in the data recording module (100), the data summary module (200) can generate a research summary report for a query entered by a user. Here, the research summary report may refer to a report prepared by adapting research-related data from a database related to the user query into a specific format according to the user query.
[0043] FIG. 2 is a flowchart of the data recording process of a data recording module (100) according to one embodiment of the invention.
[0044] A step (S101) of inputting research-related data according to one embodiment of the present invention may include an operation in which an input subject inputs research-related data into a data recording module (100) for the return of the final result of the data recording module (100).
[0045] Furthermore, the input entities may include researchers or research-related data input systems. Additionally, the format of the research-related data may include a multimodal format. For example, the format of the research-related data may include image formats, photograph formats, and text formats related to the research. Furthermore, the research-related data may include photographs of handwritten research records. Additionally, the research-related data may include existing publicly available databases, public data, etc.
[0046] A step (S102) of preprocessing input data according to an embodiment of the present invention may include a process of organizing and converting research-related data input by an input subject into a predetermined data format so that it can be utilized in subsequent work. Specifically, the step (S102) of preprocessing input data may include preliminary work to convert the format of research-related data input by a user or remove unnecessary information to store it in a database in a standardized form. For example, an image file input by a user may be converted into text data through OCR (Optical Character Recognition) technology. Additionally, text files input by a user may have duplicate content removed or be sorted into a specific format. Furthermore, the step (S102) of preprocessing input data may be omitted.
[0047] A step (S103) of extracting keywords through a keyword extraction AI model according to an embodiment of the present invention may include a process of extracting important information from data input in a step (S101) of inputting research-related data or from data preprocessed in a step (S102) of preprocessing input data. The "important information" extracted in the step (S103) of extracting keywords through a keyword extraction AI model may include major topics, core concepts, high-frequency terms, or words that are meaningful in the context of the data. Additionally, the method may further include a step of classifying keywords of data input in a step (S101) of inputting research-related data or data preprocessed in a step (S102) of preprocessing input data. For example, keywords for classifying research-related data may include information such as title, author, date of creation, date of modification, research field, research purpose, experimental stage, data format, data size, file path, collaborator, institution, and research funding information. Additionally, in the step (S103) of extracting keywords through a keyword extraction AI model, the keyword extraction AI model may include a natural language processing-based AI model. The natural language processing-based AI model may include one or more LLMs among Transformer-based, Seq2Seq, and Hybrid.
[0048] A step (S104) of storing research-related data and extracted keywords in a database according to an embodiment of the present invention may include a process of storing research-related data and extracted keywords. Specifically, the step (S104) of storing research-related data and extracted keywords in a database may be a step of grouping and storing research-related data entered by a user in the step (S101) of inputting research-related data and keywords extracted through a keyword extraction artificial intelligence model in the step (S103). For example, the data stored in the database may be data that groups the original research-related data and the metadata of said data. "Grouping" is for the purpose of facilitating the storage, classification, or summarization of research-related data. It may be a task of grouping research-related data and extracted keywords according to specific criteria.
[0049] The step (S105) of storing an embedding vector in a database according to an embodiment of the present invention may include a process of vector embedding data stored in the database and grouping and storing it with the data in the database. The module for vector embedding may include a natural language processing-based artificial intelligence model. In addition, the natural language processing-based artificial intelligence model may be at least one of Transformer-based, Seq2Seq, and Hybrid LLMs. The Transformer-based LLM may be at least one of GPT, Gemini, LLaMA, and Claude LLMs. In addition, the LLM can be used according to user needs through adaptability. For example, the LLM can be used as a personalized model in facilities where information security is critical, facilities isolated from external networks, or facilities where data security is extremely critical. In addition, the LLM model may include a fine-tuned model according to user settings, considering performance, cost, system environment, etc., as it allows for the parallel use of multiple LLMs. Additionally, the step of storing the embedding vector in the database (S105) may be omitted.
[0050] FIG. 3 is a drawing for explaining the configuration of a data recording module (100) according to one embodiment of the invention.
[0051] Referring to FIG. 3, the data recording module (100) may include an upload processing module (110), a keyword classification module (120), a database (130), research note metadata (140), an embedding processing module (150), and a vector database (160).
[0052] A step (S101) of inputting research-related data according to an embodiment of the present invention may be a step in which a user inputs research-related data into an upload processing module (110). Specifically, the user may input data in a format required by the upload processing module (110). For example, the format required by the upload processing module (110) may include an image format, a photograph format, and a text format. Additionally, the research-related data may include photographs of handwritten research records. Furthermore, the research-related data may include existing publicly available databases, public data, etc.
[0053] According to one embodiment of the present invention, the step of preprocessing input data (S102) may be a step of removing unnecessary information or converting the data format through a process in which the upload processing module (110) organizes and standardizes the data input by the user into the upload processing module (110). For example, in the step of preprocessing input data (S102), image data may be converted into text through OCR technology, and text data may be processed to remove duplicate content so that it can be used in subsequent steps.
[0054] According to one embodiment of the present invention, the step (S103) of extracting keywords through a keyword extraction artificial intelligence model may be a step in which the upload processing module (110) extracts important information from data received from a user or preprocessed data through a keyword classification module (120). Additionally, the important information may be classified as metadata (140). Specifically, the metadata (140) may include information such as the title of the research notebook, author, date of creation, date of modification, research field, research purpose, experimental stage, data format, data size, file path, collaborator, institution, and research funding information. Data received from a user may be classified according to the author, which is one of the metadata (140). However, the process of classifying research-related data into metadata (140) may not be performed in the data recording module (100) but may be performed in the data summary module (200).
[0055] According to one embodiment of the present invention, the step of storing research-related data and extracted keywords in a database (S104) may be a step in which the upload processing module (110) groups the data received from the user and the keywords extracted by the keyword classification module (120) and stores them in the database (130). Additionally, the step of storing research-related data and extracted keywords in a database (S104) may be a step of storing research-related data separately from the keywords.
[0056] The step of storing embedding vectors in a database (S105) may be a step in which an embedding processing module (150) embedding vectorizes data stored in a database (130) and stores the data and embedding vectors of the database (130) that are the target of embedding vectorization in a vector database (160). Specifically, the embedding processing module (150) may convert the data of the database (130) into embedding vectors through a natural language processing-based artificial intelligence model and store them in the vector database (160). Additionally, the vector data stored in the vector database (160) may be transmitted to a data summary module (200) upon a user's request and used for data similarity search and data correlation analysis. The step of storing embedding vectors in a database (S105) may be omitted.
[0057] FIG. 4 is a flowchart of the data summarization process of a data summarization module (200) according to one embodiment of the invention.
[0058] A step (S201) of inputting a query regarding research-related data according to an embodiment of the present invention may include an action in which a user inputs a query regarding research-related data into the data summary module (200) to return the final result of the data summary module (200). Specifically, the step (S201) of inputting a query regarding research-related data may be a step in which a user inputs at least one of the components for generating a research summary report into the data summary module (200). For example, the content of the query may include questions related not only to research data but also to analysis results, specific research topics, and specific researchers.
[0059] The step (S202) of expanding a query through an expanded query generation AI model according to an embodiment of the present invention may be a step of supplementing the content of a query entered by a user in the step (S201) of entering a query regarding research-related data and converting it into clear and specific content. Specifically, the step (S202) of expanding a query through an expanded query generation AI model may use a Retrieval-Augmented Generation (RAG) method. For example, in the step (S202) of expanding a query through an expanded query generation AI model, the supplementation of the query may include a process of obtaining additional context through an external database or information source. Additionally, the step (S202) of expanding a query through an expanded query generation AI model may include a process of obtaining a specific query according to the user's intent through query augmentation. Furthermore, the RAG method may use self-training data generated by the user and may include a process of converting a query tailored to the user that reflects the latest information or specific context by using external data. Additionally, the step of expanding the query through an expanded query generation AI model (S202) can be omitted.
[0060] The step (S203) of searching for research-related data and keywords related to a query from a database according to an embodiment of the present invention may be a step of searching for research-related data and keywords from a database based on a user query. "User query" may refer to a query entered in the step of entering a query for research-related data (S201) or a query expanded in the step of expanding the query through an expanded query generation AI model (S202). Additionally, research-related data and keywords related to the user query may be searched from the database in the step of storing research-related data and extracted keywords in a database (S104) or from the database in the step of storing embedding vectors in a database (S105). For example, the similarity between vectors in the database can be calculated by vectorizing the user query according to the RAG method, and research-related data and keywords of vectors with high similarity can be searched.
[0061] The step (S204) in which an AI model for generating a summary report according to an embodiment of the present invention generates a research summary report may be a step of converting the research-related data and keywords retrieved in the step (S203) of searching for research-related data and keywords related to a query from a database into structured text and delivering it to a user. "Structuring" may refer to a text structure pre-set by the user or a structure commonly used in the relevant research field. Specifically, the AI model for generating a summary report may separate the content of the research-related data and keywords by component and organize the separated components into a logical hierarchical structure. For example, in the case of a paper, the AI model for generating a summary report may separate and organize the content using the title, abstract, introduction, body text, conclusion, and references as components. Additionally, the summary report may be sorted by chronological order, researcher name, or keywords, and may be provided in a structured format commonly used in the research field according to the user's requirements.
[0062] In addition, the summary report can be provided in a format based on user settings. For example, users can select whether to extract key keywords or enable the highlighting feature in the summary report. Users can choose from formats such as a general report, a paper abstract, or a proposal as the summary report format.
[0063] FIG. 5 is a diagram illustrating the configuration of a data summary module (200) according to one embodiment of the invention.
[0064] Referring to FIG. 5, the data summary module (200) may include a related data search module (210), a query augmenter (220), and a research note summary device (230).
[0065] The step (S201) of inputting a query regarding research-related data according to one embodiment of the present invention may include the operation of inputting a query to the related data search module (210) to receive a research summary report. Additionally, the query input by the user may include data in the form required by the related data search module (210). For example, the data form required by the related data search module (210) may include text form. Additionally, the query input by the user may include questions related to research data, analysis results, specific research topics, or specific researchers. For example, the user may input specific queries such as "2024 research data summary" or "major trends in AI-based research." Additionally, the user may search for specific data or request a summary by inputting keywords or sentences in the form of natural language.
[0066] The step (S202) of expanding a query through an expanded query generation AI model according to an embodiment of the present invention may be a step of expanding a query entered by a user through a related data search module (210). Specifically, the related data search module (210) can search for the user query in a vector database (160) during the retrieval stage of the Retrieval-Augmented Generation (RAG) method, and can transform the query by obtaining additional context based on synonyms, related keywords, etc. In addition, the query entered by the user can be supplemented by the query augmentation of the query augmenter (220) and transformed into a clear and specific form. For example, if a user enters a vague query such as "AI research," the query augmenter (220) can expand it into a specific form that supplements the user's unexpressed intent, such as "key keywords and analysis results of AI research in 2024." Furthermore, the text generation model can generate a more sophisticated response by reflecting the latest information or specific context, rather than relying solely on the self-training data generated by the user.
[0067] The step (S203) of searching for research-related data and keywords related to a query from a database according to an embodiment of the present invention may be a step of searching for research-related data and keywords having similarity to a query in a database (130) according to a vector database (160) based on a query expanded in a related data search module (210). Specifically, the query expanded in the related data search module (210) is vectorized in the Generation stage of the Retrieval-Augmented Generation (RAG) method, so that the similarity between vectors in the vector database (160) can be calculated, and research-related data and keywords of vectors with high similarity can be searched from the database (130).
[0068] The step (S204) in which an AI model for generating a summary report according to an embodiment of the present invention generates a research summary report may be a step in which a research summary module (230) converts query-related data searched from a related data search module (210) through the AI model for generating a summary report into structured text and delivers it to a user. Additionally, a summary report may be generated according to the user's settings. Specifically, the user may select whether to extract key keywords for the summary report, whether to use a highlighting function, and the summary report format. The AI model for generating a summary report may request data from an LLM and a database by generating a prompt based on the user's selection. If the token size or data received by the LLM cannot be processed at once, multiple data may be processed through the algorithm of the AI model for generating a summary report. For example, the algorithm of the AI model for generating a summary report may include Refine or Map-Reduce. Additionally, the algorithm of the AI model for generating a summary report may include Chunking, Sliding Window, or Recursive Summarization.
[0069] FIG. 6 is a diagram illustrating the result obtained by undergoing the step (S202) of expanding a query through an expanded query generation artificial intelligence model according to one embodiment of the invention.
[0070] Through query expansion according to the present embodiment, ambiguous or non-specific queries by the user can improve the quality of search results by supplementing the user's intent that has not been clearly expressed. For example, the query "deep learning paper recommendation" may be supplemented with synonyms and specific research fields of researchers. Additionally, the query "autonomous driving related papers" may be filtered by citation count and include influence indicators.
[0071] FIGS. 7a and 7b are drawings for explaining a summary result generated based on a query input into a data summary module according to an embodiment of the invention.
[0072] As illustrated in FIG. 7a, when a user inputs a query for research-related data in the step (S201) of inputting a query, such as "Summarize the papers related to data processing and analysis researched by Kim Gu-no and Park Gu-no this year," the user can receive a response from the vector database (160) through the related data search module (210) that satisfies conditions such as "papers published after 2024," "select only topics such as machine learning, natural language processing, and autonomous driving," and "summarize only papers by researchers Kim Gu-no and Park Gu-no." The data is transmitted to the research note summarizing device (230), and the research note summarizing device (230) can refine the data with conditions such as "organize summary results into researcher, research date, research topic, main content, and main keywords" and "display summary sorted in a timeline format" through the step (S204) in which an artificial intelligence model for generating summary reports generates a research summary report.
[0073] As illustrated in FIG. 7b, the data summary module (200) can return the final research summary report to the user in response to the query of FIG. 7a. Referring to FIG. 7A and FIG. 7B, even if the query entered by the user is ambiguous or not specific, the present invention can generate a clear summary report by supplementing the user's query. Furthermore, the present invention can provide the user with necessary research materials among various research fields by analyzing the user's intent. Additionally, the present invention can automatically search for and summarize a large amount of research-related data. Therefore, the process of the user directly searching for or summarizing research-related data can be omitted.
[0074] The present invention provides a method for generating a research summary report based on one or more research-related data—said to be stored in a database or entered by a user. Specifically, the method for generating a research summary report based on one or more research-related data—said to be stored in a database or entered by a user—includes the steps of receiving a user query, extracting one or more keywords from one or more research-related data, and generating a research summary report based on one or more research-related data and one or more keywords. Additionally, the step of generating a research summary report includes the steps of searching for at least one of the one or more research-related data based on the user query, searching for at least one of the one or more keywords based on the user query, and generating a research summary report through an artificial intelligence model for generating summary reports based on the at least one searched research-related data and the at least one searched keyword.
[0075] The present invention provides a method for generating a research summary report based on one or more research-related data—said that the one or more research-related data are either stored in a database or entered by a user. Specifically, the format of the one or more research-related data may be multimodal. Additionally, one or more extracted keywords may include metadata. Furthermore, the step of generating a research summary report may further include the step of generating an extended query based on an input user query. Additionally, the one or more research-related data may include embedding vector data. Furthermore, the step of generating a research summary report may further include the step of generating an extended query based on an input user query. In this case, the step of generating an extended query further includes the step of inputting the user query into an extended query generation AI model and the step of the extended query generation AI model generating an extended query based on the embedding vector data.
[0076] The present invention provides a method for generating a research summary report based on one or more research-related data—said to be stored in a database or entered by a user. Specifically, in the step of searching for at least one research-related data and the step of searching for at least one keyword, the searched at least one research-related data and the searched at least one keyword may be searched based on the generated extended query. Additionally, the step of generating the extended query may further include the step of the extended query generating AI model being a natural language processing-based AI model. Furthermore, the natural language processing-based AI model is at least one or more of Transformer-based, Seq2Seq, and Hybrid LLMs, and the LLM may include a Refine algorithm or a Map-Reduce algorithm. Additionally, the Transformer-based LLM is at least one or more of GPT, Gemini, LLaMA, and Claude LLMs.
[0077] The present invention may provide a program that generates a research summary report based on one or more research-related data—said to be stored in a database or entered by a user. Specifically, the program according to the present invention may include instructions for performing a method of generating a research summary report based on one or more research-related data—said to be stored in a database or entered by a user.
[0078] The present invention may provide an apparatus for generating a research summary report based on one or more research-related data—said to be stored in a database or entered by a user. Specifically, the apparatus according to the present invention may include a data recording module in which the one or more research-related data is recorded, and a data summary module that performs a method for generating a research summary report based on the one or more research-related data—said to be stored in a database or entered by a user.
[0079] The scope of the present invention is not limited to the embodiments described above but may be implemented in various forms of embodiments within the scope of the appended claims. It is deemed that the scope of the claims of the present invention includes various modifications that are possible by anyone with ordinary knowledge in the technical field to which the invention pertains, without departing from the essence of the invention claimed in the claims. Explanation of the symbols
[0080] 10: Research-related data summary processing system 100: Data logging module 110: Upload processing module 120: Keyword Classification Module 130: Database 140: Metadata 150: Embedding processing module 160: Vector database 200: Data Summary Module S101: Research-related data input step S102: Input data preprocessing step S103: Keyword extraction step using a keyword extraction AI model S104: Step of storing research-related data and extracted keywords in the database S105: Step to save embedding vectors to the database
Claims
Claim 1 A method for generating a research summary report based on one or more research-related data—said that the one or more research-related data are stored in a database or are input by a user—comprising: receiving a user query; extracting one or more keywords from the one or more research-related data; and generating a research summary report based on the one or more research-related data and the one or more keywords; wherein the step of generating the research summary report comprises: searching for at least one research-related data among the one or more research-related data based on the user query; searching for at least one keyword among the one or more keywords based on the user query; and generating a research summary report through a summary report generating artificial intelligence model based on the at least one research-related data and the at least one keyword searched. Claim 2 A method for generating a research summary report, wherein the format of one or more of the research-related data is multimodal in claim 1. Claim 3 A method for generating a research summary report, wherein one or more extracted keywords include metadata, in accordance with claim 1. Claim 4 A method for generating a research summary report according to claim 1, wherein the step of generating the research summary report further comprises the step of generating an extended query based on the input user query. Claim 5 A method for generating a research summary report, wherein one or more of the research-related data include embedding vector data, in accordance with claim 1. Claim 6 A method for generating a research summary report according to claim 5, wherein the step of generating the research summary report further comprises the step of generating an extended query based on the input user query; wherein the step of generating the extended query further comprises the step of inputting the user query into an extended query generating AI model; and the step of the extended query generating AI model generating an extended query based on the embedding vector data; wherein, in the step of searching for at least one research-related data and the step of searching for at least one keyword, the searched at least one research-related data and the searched at least one keyword are searched based on the generated extended query. Claim 7 A method for generating a research summary report, wherein, in claim 6, the step of generating an extended query; further comprises the step of the extended query generating AI model being a natural language processing-based AI model. Claim 8 In claim 7, the natural language processing-based artificial intelligence model is at least one of Transformer-based, Seq2Seq, and Hybrid LLMs, and the LLM includes a Refine algorithm or a Map-Reduce algorithm, a method for generating a research summary report. Claim 9 In claim 8, the Transformer-based LLM is at least one of GPT, Gemini, LLaMA, and Claude, a method for generating a research summary report. Claim 10 In claim 1, the summary report generating AI model is a natural language processing-based AI model, a method for generating a research summary report. Claim 11 In claim 10, a method for generating a research summary report, wherein the natural language processing-based artificial intelligence model is at least one of Transformer-based, Seq2Seq, and Hybrid LLMs. Claim 12 In claim 11, the Transformer-based LLM is at least one of GPT, Gemini, LLaMA, and Claude, a method for generating a research summary report. Claim 13 A program for generating a research summary report based on one or more research-related data—said that the one or more research-related data are stored in a database or entered by a user—a program comprising an instruction for performing at least one of the methods of claims 1 to 6. Claim 14 An apparatus for generating a research summary report based on one or more research-related data—said that the one or more research-related data are stored in a database or entered by a user—comprising: a data recording module in which the one or more research-related data are recorded; and a data summary module that performs at least one of the methods of claims 1 to 6.