Method and system for topic-based classification of scientific papers to research proposal

The system provides a two-stage approach to classify scientific papers into research proposals by using a topic-based retriever and classifier model, addressing the challenge of monolithic reviews and enhancing the mapping of papers to proposal topics, resulting in improved readability and relevance.

JP2025175973APending Publication Date: 2025-12-03TATA CONSULTANCY SERVICES LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025082718
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-17
Filing Date
2025-05-16
Publication Date
2025-12-03

AI Technical Summary

Technical Problem

Existing automated literature review generation methods fail to provide structured, topic-based classification of scientific articles relevant to research proposals, lacking adequate mapping to intent and relying on citation text that is often unavailable during proposal writing, leading to monolithic and unreadable reviews.

Method used

A processor-implemented method involving a two-stage approach using a topic-based retriever (TR) and classifier (TC) model to classify scientific papers into research proposals, utilizing false-positive and false-negative reference text spans, and a parsing technique to extract in-text citations, enhancing the mapping of papers to proposal topics.

Benefits of technology

Enables the generation of comprehensive, well-structured literature reviews by accurately mapping scientific papers to proposal topics, improving readability and relevance, even without prior knowledge of the intent behind citing papers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025175973000001_ABST
    Figure 2025175973000001_ABST
Patent Text Reader

Abstract

To solve the following problem that: for categorizing literature for a specific target proposal, authors face challenges while organizing their papers in diverse ways, and to disclose embodiments of method and system for classification of scientific papers to research proposal based on topic.SOLUTION: A constructed dataset with a positive and negative reference text spans is augmented to obtain an extended dataset. Top-k chunks from a reference paper relevant to the citation text are considered as the reference text spans. A topic-based retriever model is trained on subset of the extended dataset using the positive reference text span, and the negative reference text span. A topic classifier model is trained using research proposal title, the proposal topic, and the reference text span from the reference paper to classify whether the reference paper is aligned with the topic. A reference paper relevant to the proposal topic is classified with corresponding topics in the research proposal.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application claims priority to Indian Application No. 202421038907 filed in India on May 17, 2024.

[0002] (Technical field) The present disclosure relates generally to data classification techniques, and more particularly to methods and systems for topic-based classification of scientific articles for research proposals. [Background technology]

[0003] The conception of a research problem typically begins through the writing of a detailed research proposal that highlights the relevance of the proposed idea and its corresponding novelty with respect to previous research. Researchers often write draft research proposals to present their current novel ideas, define the research problem, and seek funding. An essential aspect of the proposal writing process is reviewing the relevant literature and relating it to different aspects of the proposal: motivating the specificity of the research problem, establishing a baseline, integrating methodology, and, last but not least, automatically generating a literature review. A literature review discusses published information in a specific subject area, and sometimes within a specific period of time. The quality of a research proposal often depends on whether it adequately links to existing articles in the literature that help support the proposed research problem. Several approaches designed for the automated search of scientific literature can be applied to identify articles relevant to a proposal with a detailed description (e.g., abstract) of the proposal serving as a query. However, to gain a nuanced view of why scientific articles are relevant to the research proposal, it is necessary to map the retrieved scientific articles to high-level thematic categories (from this point on referred to as topics) that are relevant to the proposal. This mapping can be used for the downstream generation of a comprehensive topical literature review, as opposed to a monolithic review, thereby improving readability.

[0004] Automated literature review generation is essential for understanding and writing scientific documents and extracting information to synthesize comprehensive summaries. Existing approaches for automated literature review generation either independently summarize scientific papers or independently generate citation text for each scientific paper related to the target manuscript without considering its relationship to other related papers. Existing approaches often generate monolithic, abstract reviews that lack structured subsections linked to specific topics, adversely affecting readability. Citation text generation aims to construct sentences citing referenced papers based on their corresponding summaries and combine them into a literature review. However, existing methods assume accurate knowledge of the intent behind citing papers (e.g., background, methodology), which is not available in the early stages of writing a research proposal. Furthermore, these approaches rely solely on the summaries of referenced papers for text generation, potentially lacking adequate information for proper mapping to intent or topic.

[0005] Previous approaches for extracting text spans from research papers use citation text or queries. Citation intent detection presupposes the presence of citation text to classify papers into predefined genres, such as motivation and background. However, citation text is unavailable during the proposal writing stage. Furthermore, classification categories for potential reference papers vary for each target proposal. Existing approaches for retrieving scientific papers from corpora rely on the target paper's abstract and title, detailed text queries, or specific aspects, such as the topic and methodology. These approaches use strategies for generating appropriate embeddings for the query and research paper, often focusing on coarse-level "aspects" of the target paper. Bibliography generation aims to automatically generate chapter headings for a target proposal based on a set of reference papers. When literature for a specific target paper is classified, authors face challenges, such as introducing biases while those papers are organized in various ways. Summary of the Invention [Means for solving the problem]

[0006]

[0006] Embodiments of the present disclosure represent improvements over existing technologies as solutions to one or more of the above-mentioned technical problems recognized by the inventors in conventional systems. For example, in one embodiment, a processor-implemented method for classifying one or more scientific papers into research proposals based on topics is provided. The processor-implemented method includes receiving, via one or more hardware processors, as input data a dataset constructed using one or more existing research papers; augmenting, via the one or more hardware processors, the constructed dataset with one or more reference text spans to obtain an augmented dataset; training, via the one or more hardware processors, a topic-based retriever (TR) model based on a subset of the augmented dataset with one or more false-positive reference text spans and one or more false-negative reference text spans; and classifying, via the one or more hardware processors, the research proposal titles, one or more proposal topics, and the reference text spans labeled as positive. training a topic classifier (TC) model on a subset of the augmented dataset to classify whether the one or more reference papers are aligned with the proposal topic using one or more reference text spans from one or more reference papers obtained using the one or more false-positive reference text spans obtained using the text-based retriever model for the one or more reference papers labeled as negative and one or more false-negative reference text spans for the one or more reference papers labeled as negative; and classifying, via one or more hardware processors, one or more reference papers associated with the one or more proposal topics with the topic classifier (TC) model in the context of the one or more reference text spans obtained using the topic-based retriever (TR) model. The constructed dataset is associated with (i) a research proposal title, or (ii) an abstract, or (iii) a set of proposal topics, or (iv) one or more reference papers associated with the research proposal, or (v) one or more quoted texts under one or more proposal topics from the set of proposal topics corresponding to the one or more related reference papers.The one or more reference text spans are associated with one or more false positive reference text spans and one or more false negative reference text spans. One or more top-k chunks are identified from one or more reference papers associated with one or more citation texts based on the retriever model and are recognized as one or more reference text spans.

[0007] In one embodiment, one or more section headings are identified from one or more target papers to extract one or more in-text citations. In one embodiment, the one or more in-text citations are extracted using a parsing technique involving one or more extended regular expression (Regex) statements. In one embodiment, the extracted in-text citations and one or more cited reference papers associated with the set of proposal topics are utilized to extract one or more cited texts. In one embodiment, the one or more cited texts are related to one or more texts surrounding the one or more in-text citations. In one embodiment, a similarity score is calculated between the one or more cited texts and chunks of the reference papers. In one embodiment, the one or more chunks of text are ranked based on the similarity scores to identify one or more top-k chunks. In one embodiment, during the inference phase, the step of classifying one or more scientific papers into one or more research proposals includes: (i) the research proposal title, abstract R, and proposal topic from the test set are input into a TR model along with each chunk of reference papers in the test set tagged as related to R to calculate a similarity score; (ii) one or more top k chunks are obtained from the reference papers related to the proposal topic using the TR model based on the similarity score; (iii) the one or more top k chunks from the reference papers are input into a TC model along with the topic and proposal; and (iv) the reference papers are iteratively classified as positive or negative to highlight whether the reference paper is related to one or more proposal topics in the context of the one or more top k chunks.

[0008] In another aspect, a system for classifying one or more scientific papers into research proposals based on topics is provided, the system including a memory for storing instructions, one or more communication interfaces, and one or more hardware processors coupled to the memory via the one or more communication interfaces, the one or more hardware processors receiving a constructed dataset using one or more existing research papers as input data, augmenting the constructed dataset with one or more reference text spans to obtain an augmented dataset, training a topic-based retriever (TR) model on a subset of the augmented dataset with one or more false positive reference text spans and one or more false negative reference text spans, using the research proposal titles, one or more proposal topics, and the research proposals labeled as positive. The method further comprises instructions for training a topic classifier (TC) model on a subset of the expanded dataset to classify whether one or more reference papers are aligned with a proposal topic using one or more reference text spans from the one or more reference papers retrieved using the citation text-based retriever model and one or more pseudo-negative reference texts for one or more reference papers labeled as negative, and using the topic-based retriever (TR) model to classify one or more reference papers associated with the one or more proposal topics in the context of the retrieved one or more reference text spans. The constructed dataset is associated with (i) a research proposal title, or (ii) an abstract, or (iii) a set of proposal topics, or (iv) one or more reference papers associated with the research proposal, and (v) one or more citation texts under one or more proposal topics from the set of proposal topics corresponding to the one or more associated reference papers. The one or more reference text spans are associated with one or more false positive reference text spans and one or more false negative reference text spans. The one or more top k chunks identified from the one or more reference papers that are associated with the one or more cited texts based on the retriever model are considered to be one or more reference text spans.

[0009] In one embodiment, one or more section headings are identified from one or more target papers to extract one or more in-text citations. In one embodiment, the one or more in-text citations are extracted by using a parsing technique with one or more extended regular expression (Regex) statements. In one embodiment, the extracted in-text citations associated with the set of proposal topics and one or more citing reference papers are used to extract one or more cited texts. In one embodiment, the one or more cited texts are related to one or more texts surrounding the one or more in-text citations. In one embodiment, a similarity score is calculated between the one or more cited texts and chunks of the reference paper. In one embodiment, the one or more chunks of text are ranked based on the similarity scores to identify one or more top-k chunks. In one embodiment, during the inference phase, the step of classifying one or more scientific papers into one or more research proposals includes: (i) inputting the research proposal title, abstract R, and proposal topic from the test set into a TR model along with each chunk of reference papers in the test set tagged as related to R to calculate a similarity score; (ii) obtaining one or more top k chunks from the reference papers related to the proposal topic using the TR model based on the similarity score; (iii) inputting one or more top k chunks from the reference papers along with the topic and proposal into a TC model; and (iv) iteratively classifying the reference papers as positive or negative to highlight whether the reference paper is related to one or more proposal topics in the context of the one or more top k chunks.

[0010] In yet another aspect, a non-transitory computer-readable medium, when executed by one or more hardware processors, includes receiving as input data a dataset constructed using one or more existing research papers; augmenting the constructed dataset with one or more reference text spans to obtain an augmented dataset; training a topic-based retriever (TR) model on a subset of the augmented dataset with one or more false positive reference text spans and one or more false negative reference text spans; and augmenting the one or more false positive reference text spans retrieved using the citation text-based retriever model using research proposal titles, one or more proposal topics, and labeled as positive. The method includes one or more instructions for causing at least one of: training a topic classifier (TC) model on a subset of the expanded dataset to classify whether one or more reference papers are aligned with a proposal topic using one or more reference text spans from one or more reference papers retrieved using false-positive reference text spans and one or more false-negative reference texts for one or more reference papers labeled as negative; and using a topic-based retriever (TR) model to classify one or more reference papers associated with the one or more proposal topics in the context of the retrieved one or more reference text spans with the topic classifier (TC) model. The constructed dataset relates to (i) a research proposal title, or (ii) an abstract, or (iii) a set of proposal topics, or (iv) one or more reference papers related to the research proposal, or (v) one or more quoted texts under one or more proposal topics from a set of proposal topics corresponding to the one or more related reference papers. The one or more reference text spans are associated with one or more false-positive reference text spans and one or more false-negative reference text spans. One or more top-k chunks are identified from one or more reference papers that are related to one or more citation texts based on the retriever model and are identified as one or more relevant text spans.

[0011] In one embodiment, one or more section headings are identified from one or more target papers to extract one or more in-text citations. In one embodiment, the one or more in-text citations are extracted using a parsing technique involving one or more extended regular expression (Regex) statements. In one embodiment, the extracted in-text citations associated with a set of proposal topics and one or more citing reference papers are utilized to extract one or more cited texts. In one embodiment, the one or more cited texts are related to one or more texts surrounding the one or more in-text citations. In one embodiment, a similarity score is calculated between the one or more cited texts and chunks of the reference paper. In one embodiment, the one or more chunks of text are ranked based on the similarity score to identify one or more top-k chunks. In one embodiment, during the inference phase, the step of classifying one or more scientific papers into one or more research proposals includes: (i) the research proposal title, abstract R, and proposal topic from the test set are input into a TR model along with each chunk of reference papers in the test set tagged as related to R to calculate a similarity score; (ii) one or more top k chunks are obtained from the reference papers related to the proposal topic using the TR model based on the similarity score; (iii) the one or more top k chunks from the reference papers are input into a TC model along with the topic and proposal; and (iv) the reference papers are iteratively classified as positive or negative to highlight whether the reference paper is related to one or more proposal topics in the context of the one or more top k chunks.

[0012] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the invention, as claimed.

[0013] The accompanying drawings, which are incorporated in and constitute a part of this disclosure, illustrate exemplary embodiments and, together with the description, serve to explain the disclosed principles. [Brief explanation of the drawings]

[0014] [Figure 1] FIG. 1 illustrates a system for classifying topic bases of one or more scientific articles into one or more research proposals, according to an embodiment of the present disclosure. [Figure 2A] FIG. 2A is an exemplary functional block diagram of the system of FIG. 1 for classifying topic bases of one or more scientific articles into one or more research proposals, according to an embodiment of the present disclosure. [Figure 2B] FIG. 2B is an exemplary functional block diagram of a dataset construction unit of the system of FIG. 2A, according to an embodiment of the present disclosure. [Figure 3] FIG. 3 is an exemplary graphical representation illustrating the process of negative sampling of reference text spans for one or more topics, according to an embodiment of the present disclosure. [Figure 4A] FIG. 4A is an exemplary functional block diagram illustrating a process for augmenting data according to an embodiment of the present disclosure. [Figure 4B] FIG. 4B is an exemplary functional block diagram illustrating the training process of a topic-based reference text span retriever (TR) model and a proposal topic-based reference paper classifier (TC) model, according to an embodiment of the present disclosure. [Figure 4C] FIG. 4C is an exemplary functional block diagram illustrating an inference phase for classifying one or more scientific articles into research proposals based on topics, according to an embodiment of the present disclosure. [Figure 5A] FIG. 5A is an exemplary flow diagram illustrating a method for classifying one or more scientific articles into research proposals based on topics, according to an embodiment of the present disclosure. [Figure 5B] FIG. 5B is an exemplary flow diagram illustrating a method for classifying one or more scientific articles into research proposals based on topics, according to an embodiment of the present disclosure. [Figure 5C] FIG. 5C is an exemplary flow diagram illustrating a method for classifying one or more scientific articles into research proposals based on topics, according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0015] Exemplary embodiments are described with reference to the accompanying drawings. In the figures, the left-most digit(s) of a reference number identifies the figure in which the reference number first appears. Wherever convenient, the same reference numbers will be used throughout the drawings to refer to the same or like parts. Examples and features of the disclosed principles are described herein, and modifications, adaptations, and other implementations are possible without departing from the scope of the disclosed embodiments.

[0016] There is a need for an approach to generate comprehensive literature reviews featuring well-structured subsections or categories, each grouping papers related to a particular topic. Embodiments of the present disclosure provide a method and system for topic-based classification of scientific papers related to a research proposal. Human and large-scale language model (LLM) baselines are established for a task. Embodiments of the present disclosure provide a two-stage approach for topic-based classification of papers. During the first stage, one or more reference text spans are retrieved from scientific papers related to each topic by a retriever. A trained large-scale language model that uses pseudo-labels of the synthesized topic text spans is utilized. The second stage aligns scientific papers related to one or more topics and the retrieved text spans from the scientific papers in the context of the proposal, formulating the problem as a classification task. Research proposals with corresponding titles and abstracts R and a set of topics are classified into T. R = t1, t2..., t k A corpus of reference research papers is retrieved using the proposal title and high-level abstract as a query. R =p1, p2..., p n The task is to i ∈PR for each topic t j ∈T R The mapping that maps to the binary label y ijThe goal is to classify using ∈ 0,1. The dataset for this task is constructed using available target research papers as target proposals, section headings of related research in the proposals as topics, and papers cited within the sections as reference papers related to the topic. ij is a topic t in R j Related Papers p i Training on a dataset of citation texts that cite <R、p i , t j , c ij > contains a sample, but in the inference phase, ij The availability of the research proposal is not assumed. j In the context of research paper p i Refer to the reference paper p i The text is considered to be related to the proposal and the topic. ij It is called as S ij The availability of is not assumed for training as well as for inference.

[0017] Referring now to the drawings, and particularly to FIGS. 1-5C, in which like reference characters indicate corresponding features consistently throughout the figures, there are preferred embodiments shown, which will be described in connection with the following exemplary systems and / or methods.

[0018] FIG. 1 illustrates a system 100 for topic-based classification of scientific articles into research proposals, according to an embodiment of the present disclosure. In one embodiment, the system 100 includes one or more processor(s) 102, communication interface device(s) or input / output (I / O) interface(s) 106, and one or more data storage or memory 104 operably coupled to the one or more processor(s) 102. The memory 104 includes a database. The one or more processor(s) 102, memory 104, and I / O interface(s) 106 may be coupled by a system bus, such as a bus 108, or a similar mechanism. The one or more processor(s) 102, being hardware processors, may be implemented as one or more microprocessors, microcomputers, microcontrollers, digital signal processors, central processing units, state machines, logic circuits, and / or any device that processes signals based on operational instructions. Among other functions, the one or more processor(s) 102 are configured to retrieve and execute computer-readable instructions stored in the memory 104. In one embodiment, the system 100 can be implemented in a variety of computing systems, such as a laptop computer, a notebook, a handheld device, a workstation, a mainframe computer, a server, a network cloud, and the like.

[0019] The I / O interface device(s) 106 may include various software and hardware interfaces, such as a web interface, a graphical user interface, etc. The I / O interface device(s) 106 may include various software and hardware interfaces, such as interfaces for peripherals, such as a keyboard, a mouse, external memory, a camera device, and a printer. Additionally, the I / O interface device(s) 106 may enable the system 100 to communicate with other devices, such as a web server and an external database. The I / O interface device(s) 106 may enable multiple communications over a variety of networks and protocol types, including wired networks, such as a local area network (LAN) or cable, and wireless networks, such as a wireless LAN (WLAN), cellular, or satellite. In one embodiment, the I / O interface device(s) 106 may include one or more ports for connecting multiple devices to each other or to another server.

[0020] The memory 104 may include any computer-readable medium known in the art, such as volatile memory, such as static random access memory (SRAM) and dynamic random access memory (DRAM), and / or non-volatile memory, such as read-only memory (ROM), erasable programmable ROM, flash memory, hard disks, optical disks, and magnetic tape. In one embodiment, the memory 104 includes a repository 112 for storing modules 110 and data processed, received, or generated by the modules 110. The modules 110 may include routines, programs, objects, components, data structures, etc., that perform particular tasks or implement particular abstract data types.

[0021] Additionally, the database stores information relating to inputs entered into and / or outputs generated by the system 100 (e.g., data / outputs generated at each stage of data processing) specific to the methodology described herein. More specifically, the database stores information being processed at each stage of the proposed methodology.

[0022] Additionally, the modules 110 may include programs or coded instructions that complement the applications and functionality of the system 100. The repository 112 includes, among other things, a system database 114 and other data 116. The other data 116 may include data generated as a result of the execution of one or more modules in the modules 110. As used herein, computer program code configured with a memory, e.g., memory 104, and a hardware processor, e.g., processor 102, causes the system 100 to perform various functions described herein below.

[0023] FIG. 2A is an exemplary functional block diagram of the system 100 of FIG. 1 for topic-based classification of one or more scientific papers into one or more research proposals, according to an embodiment of the present disclosure. The system 200 can be an example of the system 100 (FIG. 1). In an example embodiment, the system 200 can be implemented in a system, such as the system 100 (FIG. 1), or can be in direct communication with that system. The system 100 includes a dataset construction unit 202, a data augmentation unit 204, a topic-based reference text span retriever (TR) unit 206, and a proposal topic-based reference paper classification (TC) unit 208. FIG. 2B is an exemplary functional block diagram of the dataset construction unit 202 of the system 200 of FIG. 2A, according to an embodiment of the present disclosure. The dataset construction unit 202 constructs a dataset using at least one existing research paper. The dataset construction unit 202 uses one or more UnArxiv IDs to retrieve articles from ArXiv, retrieves one or more LaTeX sources, and further extracts one or more in-text citations by using an improved parsing technique with one or more extended regular expression (Regex) statements. The LaTeX parser 202A identifies one or more section headings from one or more target articles in LaTeX format, i.e., from the headings of subsections in the "Related Work" section or the "Literature Review" section, j or t k It is configured to extract one or more in-text citations, such as those belonging to one or more topics and referencing papers, i.e. p i or p l One or more extracted in-text citations that cite c ij or c lk In one embodiment, the extracted in-text citations and one or more citing reference articles associated with the set of proposal topics are utilized to extract one or more cited texts. In one embodiment, the one or more cited texts relate to one or more texts surrounding the one or more in-text citations.j or t k One or more reference papers related to p i Or, p l The text of one or more in-text citations is identified and matched with entries in the references section to derive the titles of one or more referenced articles. In one embodiment, one or more topics, i.e., t j or t k and one or more corresponding reference papers, i.e., p i , or p l A relationship is established between the target proposals and the Portable Document Format (PDF) of one or more reference papers related to the target proposals and the topic(s) is obtained from various sources, namely, ArXiv, ACL, and Semantic Scholar. In one embodiment, the content of the one or more reference papers is extracted by a Portable Document Format (PDF) extractor.

[0024] The constructed dataset relates to at least one of (i) the title of a research proposal, or (ii) an abstract, or (iii) a set of proposal topics, or (iv) at least one reference paper related to the research proposal, and (v) at least one citation text under at least one proposal topic from the set of proposal topics corresponding to at least one related reference paper. The constructed dataset includes a set of research papers from one or more domains selected from UnArxiv as target proposals R, one or more target proposals T, and ... T. R Topic, Target Proposal P RThe ground truth labels for the task are derived from in-text citations with respect to consistency, i.e., classification of one or more reference papers into one or more topics, and are clearly specified by one or more authors of the target paper. In-text citations that are part of one or more topics, i.e., subsections, and cite the reference papers act as "links" between one or more topics and each reference paper. Certain assumptions for the training set are selected, i.e., each t j and p i Quote text for all pairs of c ij is available. Therefore, one example of the training data is the tuple <R、p i , t j , c ij Similarly, the samples for the test set are <R、p i , t j For example, a total of 2,417 target proposals were considered in the dataset, citing 49,506 reference papers, covering 7,123 topics, and containing 62,520 proposal topic-reference paper tuples. The target proposals were divided into training, validation, and test sets to prevent information leakage across splits. The resulting corpus included target proposals related to artificial intelligence (AI) (2.69%), machine learning (ML) (15.56%), computational linguistics (CL) (7.28%), computer vision (CV) (73.23%), and a combination of CL and CV (1.24%). The constructed dataset was tailored to map research papers to proposal topics, and it can be seamlessly extended to tasks such as catalog generation, citation text generation, and literature review generation. The dataset statistics are summarized in Table 1 below.

[0025] [Table 1]

[0026] In one embodiment, topic-based classification of one or more scientific papers into one or more research proposals is demonstrated in three stages: (i) using a retriever model to augment a dataset with one or more positive and negative reference text spans whose citation text is obtained as a query from one or more reference papers; (ii) using the augmented data to train a topic-based reference text span retriever (TR) and a proposal topic-based reference paper classifier (TC); and (iii) an inference pipeline to retrieve one or more reference text spans related to the proposal topic from the reference papers using the trained TR and classify the reference papers related to the proposal topic using the topic classifier (TC) in the context of the reference text spans obtained using the topic-based retriever model (TR).

[0027] One or more classifications of scientific papers, i.e., proposal topics t j Reference paper p i The classification of proposal topics is j Related papers p i Reference text spans from s ij The retriever model (TR) is a model for the topic t j and the reference text span s ij In one embodiment, the citation text c ij is the proposal topic for research proposal R j As part of the reference paper p i Quote text c ij is c ij As a result of j Related p i Reference text spans from s ij is used as a "link" to obtain the false positive pair <t j , s ij > to form a reference text span s ij is the quoted text c ijGiven,(,),, the performance of one or more existing retriever models is obtained,on equivalent tasks in the scientific domain: (i) retrieval of,paragraphs from scientific documents given a task, and (ii) retrieval of,text spans given citation text.,The retriever model (SR) with the best zero-shot performance for both,tasks is identified.

[0028] 3 is an exemplary graphical representation 300 illustrating the process of negatively sampling reference text spans for one or more topics, according to an embodiment of the present disclosure. The data augmentation unit 204 is configured to augment the constructed dataset with one or more reference text spans to obtain an augmented dataset. The one or more reference text spans relate to one or more false positive reference text spans and one or more false negative reference text spans. The data augmentation unit 204 uses a sliding window approach to augment the reference articles p i For example, seven sentences are selected as one chunk with a stride of three. ij The top k chunks from the most relevant reference papers are retrieved using the best-performing retriever (SR), and these are used to retrieve the topic t. j Paper for p i Reference text span of

number

number

number

[0029] The training dataset is augmented with false positive and false negative text spans for each topic leading to a resulting dataset, where the samples are

number

[0030] [Table 2]

[0031] During the training phase, the topic-based reference text span retriever (TR) unit 206 is configured to train a topic-based retriever model TR, which is a language model (LM) using positive and negative topic and reference text span pairs. j , s ij ) and Pr(false|t j , s ij ) is maximized using cross-entropy error. Because there are significantly more negative samples than positive samples in the dataset, all positive samples are included, but negative samples of each type are randomly subsampled to maintain a balanced training dataset. Sampling is performed alternately every epoch to ensure the model sees all negative samples.

[0032] The proposal topic-based reference paper classifier (TC) unit 208 is configured to train a classifier model TC on a subset of the augmented training dataset.

number

number

number

[0033] During the inference phase, the proposal title, abstract R, and topics t from the test set are used. j is the number of papers in the test set tagged as relevant to R that refer to paper p i Each chunk of c i The model TR is input together with each c i Similarity score sim ij =TR(t j , c i ) is obtained by applying a softmax function on the logits of the "true" and "false" tokens as depicted, yielding the probability P(true|t j , c i ) In one embodiment, c iis the similarity score sim ij Based on the given R and t j are ranked for the R topic t j For thesis p i Reference text span taken from

number

number

number

[0034] FIG. 4A is an exemplary functional block diagram 400 illustrating a data augmentation process according to an embodiment of the present disclosure. FIG. 4B is an exemplary functional block diagram 400 illustrating the process of training a topic-based reference text span retriever (TR) model and a proposal topic-based reference paper classifier (TC) model according to an embodiment of the present disclosure. FIG. 4C is an exemplary functional block diagram 400 illustrating the inference phase of classifying one or more scientific papers into research proposals based on topics according to an embodiment of the present disclosure. In one embodiment, one or more stages of a proposal topic-based reference paper matching task are included, where R is the proposal title and abstract, PR is a paper related to R, TR is a topic related to R, and p i ∈PR, t j , t k ∈TR, and p i and p l is t j and t k Cited in c ij , c lkis a topic in R j and t k Related Papers p i and p l is a quoted text that quotes s ij is a topic in R j For the related paper p i Reference text span from c ij the top k chunks similar to the query), and ij} is the topic t j is a negatively sampled text span for the query c ij p not similar to i The bottom k chunks from query c lk Similar to p l The top k chunks from query c ij Similar to p l (the sum of the top k chunks from p in PR) l is t k where k≠j and c lk is a topic in R k Related Papers p l is a quoted text that quotes y ij is the label, and paper p i is the topic j is 1 if it is consistent with

[0035] 5A-5C are exemplary flow diagrams illustrating a method 500 for classifying one or more scientific articles into research proposals based on topics, according to an embodiment of the present disclosure. In one embodiment, the system 100 includes one or more data storage or memory 104 operatively coupled to one or more hardware processors 102 and configured to store instructions for execution of the method steps by the one or more processors 102. The depicted flow diagrams will be better understood by way of the following explanation / description. Next, the steps of the disclosed method will be described with reference to the components of the system as depicted in FIGS. 1 and 2A.

[0036] In step 502, a constructed dataset (as depicted in the corresponding description in FIG. 2B ) is received as input data using one or more existing research papers. The constructed dataset is associated with (i) a research proposal title, or (ii) an abstract, or (iii) a set of proposal topics, or (iv) one or more reference papers related to the research proposal and (v) one or more citation texts under one or more proposal topics from the set of proposal topics corresponding to the one or more related reference papers. In step 504, the constructed dataset is augmented with one or more reference text spans to obtain an expanded dataset (as depicted in the corresponding description in FIG. 4A ). In one embodiment, one or more section headings are identified from one or more target papers to extract one or more in-text citations. In one embodiment, the one or more in-text citations are extracted by using a parsing technique with one or more extended regular expression (Regex) statements. In one embodiment, the extracted in-text citations and one or more cited reference papers associated with the set of proposal topics are utilized to extract one or more citation texts. In one embodiment, the one or more cited texts are related to one or more texts surrounding the one or more in-text citations. One or more top-k chunks are identified from one or more reference papers related to the one or more cited texts based on a retriever model and are considered as one or more reference text spans. In one embodiment, a similarity score is calculated between the one or more cited texts and the chunks of the reference papers. In one embodiment, the one or more chunks of text are ranked based on the similarity score to identify the one or more top-k chunks.

[0037] In step 506, a topic-based retriever (TR) model is trained (as depicted by the corresponding description in FIG. 4B) on a subset of the dataset augmented by one or more false positive reference text spans and one or more false negative reference text spans (as depicted by the corresponding description in FIG. 3). The one or more reference text spans are associated with one or more false positive reference text spans and one or more false negative reference text spans. In step 508, a topic classifier (TC) model is trained (as depicted by the corresponding description in FIG. 4B) to classify whether one or more reference papers align with the proposal topic in the augmented subset of the dataset using the research proposal title, one or more proposal topics, and one or more reference text spans from one or more reference papers obtained using the one or more false positive reference text spans obtained using the citation text-based retriever model labeled as positive and one or more false negative reference text spans for the one or more reference papers labeled as negative. In step 510, one or more reference papers are classified with a topic classifier (TC) as being relevant to one or more proposal topics in the context of one or more reference text spans retrieved using a topic-based retriever (TR) model (as depicted with corresponding descriptions in Figure 4C).

[0038] During the inference phase, the steps of classifying one or more scientific papers into one or more research proposals include: (i) in step 512, the research proposal title, abstract R, and proposal topic from the test set are input into the TR model along with each chunk of reference papers in the test set tagged as related to R to calculate a similarity score; (ii) in step 514, one or more top k chunks are obtained from the reference papers related to the topic of the proposal using the TR model based on the similarity score; (iii) in step 516, the one or more top k chunks from the reference papers are input into the TC model along with the topic and proposal; and (iv) in step 518, the reference papers are iteratively classified as positive or negative to highlight whether the reference paper is related to one or more proposal topics in the context of the one or more top k chunks.

[0039] (Experimental results) For example, we evaluate three models, namely, a retriever model (SR), a topic-based reference text span retriever (TR) model, and a proposal topic-based reference paper classifier (TC), by estimating their corresponding performances. The performance of the SR model in the evidence and reference text span retriever task is evaluated by calculating the evidence F1 score using the ground truth evidence and reference text spans available in the dataset. The performance of the TR model is evaluated by calculating the evidence F1 score. Using the TR model trained with a topic as a query, the top k most similar chunks retrieved from the reference paper are used to generate the predicted reference text spans for that topic.

number

[0040] To establish a baseline, the "topic classification of reference papers" task using human- and LLM-generated annotations has been evaluated. The evaluation subset was constructed by uniformly sampling four target proposals randomly from each of five domains from the test set to ensure balanced domain representation. The results highlight the selection of 20 target proposals with 52 topics that cite 362 reference papers, forming 378 positive topic-reference-paper pairs. Human annotators and LLM-generated annotations were used to classify the target proposal title, abstract, and topic. ij , Reference paper title p i and (i) the citation text c using the SR model to evaluate the performance of the TC model alone. ij and (ii) topic t using a trained TR model to evaluate the performance of the full inference pipeline. ij Reference papers obtained using ij The participants were provided with information about the reference text span of a paper. The task is to evaluate the relevance of the reference paper to a given topic, and if the reference paper is found to be relevant in the topic to cite the given reference text span, it is labeled as 1, otherwise it is labeled as 0.

[0041] Embodiments of the present disclosure provide a task of mapping relevant scientific papers to research proposal topics as a precursor to the task of generating literature reviews for new research proposals. For example, the introduction of a large-scale dataset for the task, the establishment of a dominant baseline using experts and large-scale language models (LLMs), and the feasibility of the task are demonstrated. The task assumes a real-world scenario in which citation texts or detailed topic descriptions are unavailable during the proposal creation phase (i.e., during inference) to obtain text spans from reference papers related to the topic; these text spans are required to establish the consistency of the reference papers to the topic. Citation texts (i.e., available for training data) are used as links between topics and text spans to generate pseudo-labels for training a retriever. The disclosed pipeline produces performance comparable to most LLM baselines, demonstrating corresponding effectiveness, using a majority of smaller language models (LMs) trained with pseudo-labels.

[0042] Embodiments of the present disclosure herein address the unsolvable challenges of existing assumptions made by current scientific literature retriever methods in generating automated literature reviews. Embodiments of the present disclosure herein provide a framework for topic-based classification of scientific or reference papers related to a research proposal. Instead of generating a monolithic review, the present disclosure serves as a precursor to generating a comprehensive literature review for research papers that is well-organized into a set of topics. Thus, embodiments of the present disclosure focus on the task of automatically mapping relevant scientific papers to one or more topics (e.g., proposal-specific topics defined by a user) without assuming the existence of text used to cite the papers (i.e., citation text) in the research proposal. This task precedes the task of generating a literature review with well-structured subsections in which related papers are grouped by specific topics of interest. The method for mapping research papers to one or more topics extracts "text spans" from the entire content of the reference papers, ensuring the availability of comprehensive information for mapping. Embodiments of the present disclosure extract text spans without relying on the availability of citation text or detailed queries, but assume the existence of higher-level topic names. Embodiments of the present disclosure assume the availability of research papers relevant to the proposal and define the task of mapping these papers to a user-defined set of fine-grained topics relevant to the proposal. Embodiments of the present disclosure operate under the assumptions that (a) there exists a corpus of scientific papers relevant to the proposal obtained using the title and abstract of the provided proposal, and (b) there exists a user-provided catalog of thematic categories for the target proposal consisting of a list of high-level topics required for matching scientific papers in the corpus.

[0043] In embodiments of the present disclosure, a more realistic setting is considered, assuming the availability of not only retrieved reference papers related to the proposal, but also high-level topics provided by the researcher. Based on these topics, the researcher wishes to focus on the task of categorizing related literature and reference papers into corresponding topics. A dataset is constructed to benchmark solutions to the proposed task. The dataset for the task is constructed using available target research papers as the target proposal, the section headings of the proposal's related research sections as "topics," and the papers cited under these sections as reference papers related to these "topics." The result is a more personalized method for the classification and cataloging process, closely aligned with the researcher's unique perspective. The disclosed framework / method serves as an upstream task for the task of generating a comprehensive literature review. Categorizing the retrieved papers into one or more topics further enables the generation of a cohesive summary of research related to each topic, taking into account diverse perspectives in the relevance of each reference paper to the proposal. The disclosed framework / method maps each retrieved scientific paper to one or more topics in the catalog, providing a comprehensive understanding of the different contributions corresponding to the target paper. During the proposal generation stage, we assume a realistic scenario in which citation text or detailed topic descriptions are unavailable from reference papers related to the topic to retrieve the reference text spans required to establish the consistency of the reference paper to the topic (i.e., during the inference stage). The citation text (i.e., available for training data) is used as a link between the topic and the text spans to generate pseudo-labels for training the retriever. The disclosed embodiment surpasses the baseline LLM with a 13.74% increase in F1 score for the classification task, demonstrating corresponding effectiveness.

[0044] This specification describes the subject matter herein to enable any person skilled in the art to make and use embodiments. The scope of the subject matter embodiments is defined by the claims and may include other modifications that occur to those skilled in the art. Such other modifications are intended to be within the scope of the claims if they have similar elements that do not differ from the literal language of the claims, or if they contain elements that are equivalent to the literal language of the claims but make insubstantial differences from the literal language of the claims.

[0045] It should be understood that the scope of protection herein extends not only to the above-mentioned program and computer-readable storage means containing messages, but also to the above-mentioned computer-readable storage means containing program code means for implementing one or more steps of the method when the program is executed on a server, mobile device, or any suitable programmable device. The hardware device can be any type of programmable device, including any type of computer, such as a server or personal computer, or a combination thereof. The device can also include hardware means, such as an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or a combination of hardware and software means, such as an ASIC and an FPGA, or means such as at least one microprocessor and at least one memory with software processing components disposed therein. Thus, the means can include both hardware and software means. The method embodiments described herein can be implemented in hardware and software. The device can also include software means. Alternatively, the embodiments can be implemented in different hardware devices, for example, using multiple CPUs.

[0046] Embodiments herein may include hardware and software elements. Embodiments implemented in software include, but are not limited to, firmware, resident software, microcode, etc. The functions performed by various components described herein may be implemented in other components or combinations of other components. For purposes of this description, a computer-usable medium or computer-readable medium may be any device that can configure, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.

[0047] The steps shown are designed to explain the depicted exemplary embodiments, and it is to be expected that ongoing technological developments will change the way in which particular functions are performed. These examples are provided herein for purposes of illustration, not limitation. Moreover, boundaries of functional components have been arbitrarily defined herein for convenience of description. Alternative boundaries may be defined so long as the specified functions and relationships are appropriately performed. Alternatives (including equivalents, extensions, variations, and deviations thereof described herein) will be apparent to those skilled in the relevant art(s) based on the teachings contained herein. Such alternatives are included within the scope of the disclosed embodiments. Additionally, the terms "comprising," "having," "containing," and "including," as well as other similar forms, are intended to be equivalent in meaning and open-ended, in that the item or items following any one of these words are not intended to be an inclusive list of such item or items or to be limited to only the listed item or items. It should also be noted that as used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural references unless the context clearly dictates otherwise.

[0048] Additionally, one or more computer-readable storage media can be utilized in implementing embodiments consistent with the present disclosure. A computer-readable storage medium refers to any type of physical memory in which information or data readable by a processor can be stored. Thus, a computer-readable storage medium can store instructions for execution by one or more processors, including instructions for causing the processor to perform steps or stages consistent with the embodiments described herein. The term "computer-readable storage medium" should be understood to include tangible items and exclude carrier waves and transient signals, i.e., non-transitory. Examples include random access memory (RAM), read-only memory (ROM), volatile memory, non-volatile memory, hard drives, CD-ROMs, DVDs, flash drives, disks, and any other known physical storage medium.

[0049] It is intended that the present disclosure and examples be considered as exemplary only, with the true scope of the disclosed embodiments being indicated by the following claims.

Claims

1. receiving, via one or more hardware processors, a constructed dataset as input data using at least one existing research paper, the constructed dataset relating to at least one of: (i) a research proposal title, or (ii) an abstract, or (iii) a set of proposal topics, or (iv) at least one reference paper related to the research proposal, and (v) at least one citation text under at least one proposal topic from the set of proposal topics corresponding to the at least one related reference paper; Augmenting the constructed dataset with at least one reference text span via one or more hardware processors to obtain an augmented dataset, wherein the at least one reference text span is related to at least one false positive reference text span and at least one false negative reference text span, and at least one top-k chunk is identified from at least one reference paper related to the at least one cited text based on the retriever model and is considered as the at least one reference text span; training, via one or more hardware processors, a topic-based retriever (TR) model on a subset of the augmented dataset with at least one false positive reference text span and at least one false negative reference text span; training, via one or more hardware processors, a topic classifier (TC) model on the subset of the augmented dataset to classify whether at least one reference paper is aligned with the proposal topic using the title of the research proposal, at least one proposal topic, and at least one reference text span from at least one reference paper obtained using a citation text-based retriever model labeled as positive, and at least one pseudo-positive reference text span for the at least one reference paper labeled as negative; classifying, via one or more hardware processors, at least one reference paper related to at least one proposal topic in the context of at least one reference text span retrieved using the topic-based retriever (TR) model with the topic classifier (TC) model; A processor-implemented method comprising:

2. 1. A processor implementation method, comprising:

2. The method of claim 1, wherein at least one section heading is identified from at least one subject article for extraction with at least one in-text citation, and the at least one in-text citation is extracted by using a parsing technique involving one or more extended regular expression (Regex) statements.

3. 3. The method of claim 2, wherein the extracted at least one in-text citation related to at least one proposal topic and at least one cited reference paper is utilized to extract at least one cited text, the at least one cited text being related to one or more texts surrounding the at least one in-text citation.

4. 10. The method of claim 1, wherein a similarity score is calculated between at least one cited text and a chunk of the reference paper, and one or more chunks of text are ranked based on the similarity score to identify at least one top-k chunk.

5. 1. A processor implementation method, comprising: During the inference phase, the step of classifying at least one scientific article into at least one research proposal comprises: inputting the research proposal title, the abstract R, and the proposal topic from a test set, along with each chunk of reference papers in the test set tagged as related to R, into the TR model via one or more hardware processors to calculate a similarity score; obtaining, via one or more hardware processors, at least one top-k chunk from the reference papers related to the topic of the proposal using the TR model based on the similarity scores; and inputting at least one top-k chunk from the reference paper along with the topic and the proposal into the TC model via one or more hardware processors; iteratively classifying, via one or more hardware processors, the reference papers as positive or negative, highlighting whether the reference papers are relevant to at least one proposal topic in the context of at least one top-k chunk; 10. The method of claim 1, comprising:

6. a memory for storing a plurality of instructions; one or more communication interfaces; one or more hardware processors coupled to the memory via the one or more communication interfaces; the one or more hardware processors receive as input data a constructed dataset using at least one existing research paper, the constructed dataset relating to at least one of: (i) a research proposal title, or (ii) an abstract, or (iii) a set of proposal topics, or (iv) at least one reference paper related to the research proposal, and (v) at least one citation text under at least one proposal topic from the set of proposal topics corresponding to at least one related reference paper; Augmenting the constructed dataset with at least one reference text span to obtain an extended dataset, wherein the at least one reference text span is related to at least one false positive reference text span and at least one false negative reference text span, and at least one top-K chunk is identified from at least one reference paper related to at least one citation text based on a retriever model and is considered as the at least one reference text span; training a topic-based retriever (TR) model on the subset of the dataset augmented with at least one false positive reference text span and at least one false negative reference text span; training a topic classifier (TC) model on the subset of the augmented dataset to classify whether at least one reference paper is aligned with the proposal topic using the title of the research proposal, at least one proposal topic, and at least one reference text span from at least one reference paper obtained using the citation text-based retriever model labeled as positive, and at least one pseudo-positive reference text span for the at least one reference paper labeled as negative; The system is configured by the instructions for classifying, with the topic classifier (TC) model, at least one reference paper related to at least one proposal topic in the context of at least one reference text span obtained using the topic-based retriever (TR) model.

7. 10. The system of claim 6, wherein at least one section heading is identified from the at least one subject article for extraction with at least one in-text citation, and the at least one in-text citation is extracted by using a parsing technique involving one or more extended regular expression (Regex) statements.

8. The extracted at least one in-text citation related to the at least one proposal topic and the at least one cited reference paper is utilized to extract at least one cited text, the at least one cited text being related to one or more texts surrounding the at least one in-text citation. The system of claim 7.

9. a similarity score is calculated between at least one citation text and a chunk of the reference paper, and one or more chunks of text are ranked based on the similarity score to identify at least one top-k chunk; The system of claim 6.

10. the one or more hardware processors are configured with instructions for classifying at least one scientific article into at least one research proposal during an inference phase; inputting the research proposal title, the abstract R, and proposal topics from a test set into the TR model, along with each chunk of reference papers in the test set tagged as related to R, to calculate the similarity score; obtaining at least one top-k chunk from the reference papers related to the topic of the proposal using the TR model based on the similarity score; inputting at least one top-k chunk from the reference paper along with the topic and the proposal into the TC model; Iteratively classifying the reference papers as positive or negative to highlight whether the reference papers are relevant to at least one proposal topic in the context of at least one top-k chunk; The system of claim 6 , comprising:

11. One or more non-transitory machine-readable information storage media containing one or more instructions for execution by one or more hardware processors, the one or more instructions causing: receiving a constructed dataset using at least one existing research paper as input data, the constructed dataset being related to at least one of: (i) a research proposal title, or (ii) an abstract, or (iii) a set of proposal topics, or (iv) at least one reference paper related to the research proposal, and (v) at least one citation text under at least one proposal topic from the set of proposal topics corresponding to at least one related reference paper; Augmenting the constructed dataset with at least one reference text span to obtain an expanded dataset, wherein the at least one reference text span is related to at least one false positive reference text span and at least one false negative reference text span, and at least one top-K chunk is identified from at least one reference paper related to the at least one cited text based on a retriever model and is considered as the at least one reference text span; training a topic-based retriever (TR) model on a subset of the augmented dataset with at least one false positive reference text span and at least one false negative reference text span; training a topic classifier (TC) model on the subset of the augmented dataset to classify whether at least one reference paper is aligned with the proposal topic using a title of the research proposal, at least one proposal topic, and at least one reference text span from at least one reference paper obtained using the citation text-based retriever model labeled as positive, and at least one pseudo-positive reference text span for the at least one reference paper labeled as negative; classifying, with the topic classifier (TC) model, at least one reference paper related to at least one proposal topic in the context of at least one reference text span retrieved using the topic-based retriever (TR) model; one or more non-transitory machine-readable information storage media,

12. At least one section heading is identified from at least one target article for extraction with at least one in-text citation, and the at least one in-text citation is extracted by using a parsing technique involving one or more extended regular expression (Regex) statements; 12. One or more non-transitory machine-readable information storage media according to claim 11.

13. The extracted at least one in-text citation related to the at least one proposal topic and the at least one cited reference paper is utilized to extract at least one cited text, the at least one cited text being related to one or more texts surrounding the at least one in-text citation.

13. One or more non-transitory machine-readable information storage media according to claim 12.

14. a similarity score is calculated between at least one citation text and a chunk of the reference paper, and one or more chunks of text are ranked based on the similarity score to identify at least one top-k chunk; 12. One or more non-transitory machine-readable information storage media according to claim 11.

15. During the inference phase, the step of classifying at least one scientific article into at least one research proposal comprises: inputting the research proposal title, the abstract R, and the topic of the proposal from a test set, along with each chunk of reference papers in the test set tagged as related to R, into the TR model to calculate a similarity score; obtaining at least one top-k chunk from the reference papers related to the topic of the proposal using the TR model based on the similarity score; and inputting at least one top-k chunk from the reference paper along with the topic and the proposal into the TC model; Iteratively classifying the reference papers as positive or negative to highlight whether the reference papers are relevant to at least one proposal topic in the context of at least one top-k chunk; 12. One or more non-transitory machine-readable information storage media according to claim 11, comprising: