Intelligent literature screening method and system based on zero sample voting device

Through the intelligent literature screening method based on the zero-sample voting machine, combined with multiple zero-sample voting machines and dynamic threshold optimization, the problems of low efficiency and insufficient accuracy in traditional literature search methods are solved, and efficient and accurate literature screening and expansion are achieved, suitable for academic research and interdisciplinary literature mining.

CN120277210APending Publication Date: 2025-07-08NANJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510465419.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

Traditional literature search methods have too many search results and are low in correlation when processing large-scale literature, which makes it difficult to meet user needs. Especially in emerging technology fields or interdisciplinary research, the lack of labeled data has led to limited model performance and lack of dynamic adaptability, making it difficult to adapt to literature in different fields or different semantic structures.

Method used

Using the intelligent literature screening method based on zero-sample voting machine, through the introduction of multiple zero-sample voting machines, combined with the majority voting mechanism and an improved binary search algorithm, the semantic similarity threshold is dynamically optimized to achieve literature classification and expansion.

Benefits of technology

It significantly improves the efficiency and accuracy of literature screening, can quickly locate high-value literature in massive academic data, adapt to different search needs, and is suitable for comprehensive literature analysis and interdisciplinary literature mining.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277210A_ABST
    Figure CN120277210A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent literature screening method and system based on zero sample voters. The method comprises the following steps: expanding a literature set based on seed literatures, constructing a classification system comprising a plurality of zero sample voters, screening literatures related to a target topic through a majority voting mechanism, and iteratively expanding the literatures to obtain more literatures related to the target topic. The zero sample voting device integrates subject term matching, language model pre-training, multi-large language model voting and a semantic embedding vector technology, adaptively determines a threshold value in combination with an improved binary search algorithm, and realizes high-precision classification without training data. The system supports parallel processing and intelligent API management, and the processing efficiency of a large-scale literature set is remarkably improved. The method is suitable for the fields of academic research, literature retrieval and the like, can quickly construct a large-scale high-correlation literature information base, and particularly meets the retrieval requirements of emerging technologies or interdisciplinary themes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of literature retrieval, and particularly relates to a method and system for intelligent literature screening based on a zero-shot voter. Background Art

[0002] With the rapid growth of the amount of literature information, users' demand for accurate retrieval is getting higher and higher. Traditional literature retrieval methods have too many retrieval results and low relevance when dealing with large-scale literature, making it difficult to meet users' needs. In fields such as academic research and technical intelligence mining, it is crucial to retrieve and screen relevant literature efficiently and accurately. When a large number of relevant literature to the target topic needs to be obtained, traditional literature retrieval methods mainly rely on keyword matching or rule-based classification systems. However, these methods have obvious limitations: Relying on manual annotation and training data: Traditional machine learning methods need a large amount of labeled training data to work effectively. However, for emerging technology fields or interdisciplinary research topics, there is often a lack of sufficient labeled data, resulting in limited model performance. In addition, the cost of manual annotation is high and it is difficult to adapt to rapidly changing retrieval needs.

[0003] Low retrieval efficiency and high noise: Methods based on citation expansion or simple keyword matching can expand the literature set, but they are prone to introducing a large number of literature irrelevant to the research topic, reducing the accuracy of retrieval results. And keyword matching is easily affected by term ambiguity or expression differences.

[0004] Lack of dynamic adaptability: Existing methods based on semantic similarity matching usually adopt fixed similarity thresholds or classification rules, making it difficult to adapt to literature in different fields or with different semantic structures. For example, some technical topics may require a high semantic matching accuracy, while others may require a more relaxed matching strategy, and traditional methods cannot be flexibly adjusted.

[0005] In recent years, the development of zero-shot learning and pre-trained language models has provided new possibilities for unsupervised literature classification. However, the zero-shot classification results of a single model may be biased, and the performance of different models varies greatly on different types of literature. In addition, how to efficiently integrate multiple zero-shot classification methods and combine dynamic threshold optimization strategies is still a challenge that has not been fully solved.

[0006] Therefore, there is an urgent need for a system that can automatically expand the literature library, intelligently screen highly relevant literature, and adapt to different retrieval needs. Summary of the Invention

[0007] To solve the problems of low retrieval efficiency and insufficient accuracy in the prior art, the purpose of the present invention is to provide a literature intelligent screening method and system based on a zero-shot voter. By introducing multiple zero-shot voters, the system is given stronger literature classification capabilities and the retrieval efficiency is improved.

[0008] The purpose of the present invention is achieved through the following technical solutions: A literature intelligent screening method based on a zero-shot voter, comprising the following steps: Step S1: Obtain the metadata of the target literature. Based on the given literature title, retrieve matching literature through the academic database API and obtain its unique identifier. Then use this identifier to further query the metadata of the literature, including the publication year, the number of references, and the number of citations, for subsequent analysis.

[0009] Step S2: Expand based on the seed literature, obtain the references and cited literature of the seed literature, construct a preliminary literature information library and remove duplicate literature; use the literature obtained in step S1 as the seed literature, and batch obtain the references and cited literature sets of each seed literature through the API, extract the complete metadata including the unique identifier, title, abstract, publication source, year, and citation metrics of these literatures, and construct a preliminary literature information library. Then merge the literature information library constructed this time with the existing literature information library in the system and remove duplicate literature.

[0010] Step S3: Construct multiple zero-shot voters: Use the literature information library obtained in step S2 to establish multiple zero-shot voters. Given a topic, the zero-shot voter can classify unseen literature without any training samples and determine whether it is relevant to the given topic.

[0011] Step S4. Screen out the literatures that meet the requirements: Integrate the classification results of multiple zero-shot voters in step S3 and screen out the literatures that meet the requirements. Among them, the output results of multiple zero-shot voters in step S3 are used as inputs, and a majority voting mechanism is used for comprehensive determination. First, count the number of votes obtained by each literature in each classifier. When the cumulative number of votes of a certain literature exceeds half of the total number of votes, classify this literature as a literature that meets the topic, retain it in the literature information library, otherwise remove it from the literature information library. The final output result of this step is multiple literatures related to the given topic.

[0012] Step S5: Expand the literature: Based on the multiple literatures related to the given topic obtained in step S4, use these literatures as seed literatures again, and repeat steps S2 to S4 for these seed literatures to expand the literature information library to obtain more literatures related to the given topic. The number of repetitions is not limited and depends on user needs. Finally, an expanded literature information library is obtained.

[0013] In step S3, there are a total of six zero-shot voters. The first zero-shot voter performs matching and screening based on the subject words. Specifically, first, a list of target subject words is set, and then each document in the literature information library obtained in step S2 is traversed to check whether any of the subject words are included in its title and abstract. For the documents that match successfully, they are recorded as 1 in the subject word marking dictionary, meaning that the document is related to the subject word; those that do not match are marked as 0, meaning that the document is not related to the subject word. This step finally outputs a dictionary containing the marking status of all documents as the voting result of this zero-shot voter.

[0014] The second zero-shot voter is based on the pre-trained language model BART. The pre-trained language model BART-large-MNLI is used to achieve zero-shot classification. First, a candidate label set is constructed, and then for each document in the literature information library obtained in step S2, the combined text of its title and abstract is processed in parallel. There are two labels in the candidate label set. The label 1 indicates that it is related to the subject, and the label 0 indicates that it is not related to the subject. Through the multi-thread acceleration technology, the model automatically determines the relevance of the document to the target technical subject, marks the matching documents as 1, and the non-matching documents as 0. Finally, a dictionary containing the marking status of all documents is output as the voting result of this zero-shot voter.

[0015] The third zero-shot voter is based on the pre-trained language model RoBERTa. The RoBERTa-large-MNLI pre-trained model is used to achieve zero-shot classification. First, a candidate label set including two labels is constructed. The label 1 indicates that it is related to the subject, and the label 0 indicates that it is not related to the subject. Then, for each document in the literature information library obtained in step S2, the combined text of its title and abstract is processed in parallel. The model automatically generates the relevance score of the document to the target subject, marks the matching documents as 1, and the non-matching documents as 0. Finally, a dictionary containing the marking status of all documents is output as the voting result of this zero-shot voter.

[0016] The fourth zero-shot voter is based on the voting of multiple large language models. An voting committee is formed by integrating three large language models, GPT-3.5-turbo, GPT-4, and O1-mini. Each document in the literature information library obtained in step S2 is input into each large model, and each model is required to perform binary classification judgment on the document to determine whether each document is related to the target subject. If 2 or 3 models think it is related, then each document is marked as 1, otherwise it is marked as 0. Finally, a dictionary containing the marking status of all documents is output as the voting result of this zero-shot voter.

[0017] The fifth zero-shot voter is based on Jina embedding vectors. In this step, based on the literature dataset constructed in step S2, the deep semantic representation analysis of the literature is realized through the Jina Embeddings V3 model. This model is specifically optimized for retrieval tasks, generates high-dimensional vector representations, and can capture complex semantic features in the literature title and abstract. This zero-shot voter first normalizes the text content of each piece of literature, generates embedding vectors in batches through parallel requests, and then generates embedding vectors for the target subject terms. Then, it calculates the similarity between the embedding vectors of the target subject terms and the text embedding vectors of the literature. When the similarity exceeds the threshold, the literature is considered relevant to the target subject, and at this time, it is marked as 1 for this literature, otherwise it is marked as 0. Finally, a dictionary containing the marking status of all literatures is output as the voting result of this zero-shot voter.

[0018] The threshold used by the fifth zero-shot voter is determined by an improved binary search algorithm, which can intelligently determine the optimal similarity threshold for literature classification. This method first sorts the text embedding vectors of all literatures in descending order according to their semantic similarity with the embedding vectors of the target subject terms, forming an ordered literature sequence. The system first checks the key literatures at both ends of the sequence: if the literature with the highest similarity is determined to be irrelevant, the standard is relaxed and the threshold is lowered; if the literature with the lowest similarity is determined to be relevant, the standard is tightened and the threshold is raised. Under normal circumstances, the system adopts a binary search strategy and quickly locates the classification boundary by continuously dividing the search range in half. In each iteration, the system selects the literature at the middle position and uses the fourth zero-shot voter in step S3 for determination, and adjusts the search range according to the determination result. When the preliminary boundary point is found, the system will check several adjacent literatures and determine the final threshold by calculating the average similarity of these boundary literatures to ensure the stability and reliability of the result. The advantage of this method is that for a dataset containing N pieces of literature, only about log2N key determinations are required to determine the optimal threshold, which improves the processing efficiency from the linear level to the logarithmic level. For example, when processing 100,000 pieces of literature, only about 17 determinations are required to complete the optimization.

[0019] The sixth zero-shot voter is based on OpenAI embedding vectors. In this step, OpenAI's text-embedding-3-small model is used to represent and classify literature data. First, the system batch processes the title and abstract information of each piece of literature through the API interface to generate high-quality text embedding vectors. After the embedding generation is completed, the semantic similarity between each piece of literature and the target topic words is calculated based on the generated embedding vectors. The threshold determination algorithm in the fifth zero-shot voter is reused to achieve literature classification through the dynamically determined optimal threshold. When the similarity exceeds the threshold, the literature is considered relevant to the target topic, and at this time, the literature is marked as 1, otherwise marked as 0. The final output is a dictionary containing the marking status of all literatures as the voting result of this zero-shot voter.

[0020] A system for a literature intelligent screening method based on a zero-shot voter, characterized by comprising: Literature metadata collection module: used to obtain the metadata related to the literature, such as title, abstract, year, etc., when a literature title or the unique identifier of the literature is given; Literature information library module: For the seed literature and the literature expanded according to the seed literature, they are saved in the literature information library module. When the literature in the literature information library module is determined by the zero-shot voting module to be not in line with the theme, the literature is excluded, and when it is determined to be relevant to the theme, it is retained.

[0021] Zero-shot voting module: includes six zero-shot voters. Each zero-shot voter respectively judges whether the literature is relevant to the given theme. The zero-shot voting module uses the majority voting mechanism for comprehensive determination based on the output results of these zero-shot voters to screen out the literatures that meet the requirements.

[0022] The present invention has the following beneficial effects: The literature intelligent screening method and system based on the zero-shot voter provided by the present invention. First, through the multi-round iterative literature expansion mechanism combined with the collaborative decision-making of six zero-shot voters, the effect of screening literatures related to the target theme is significantly improved, and it can quickly locate high-value literatures in a large amount of academic data. Secondly, the improved binary search algorithm is used to dynamically optimize the semantic similarity threshold, enabling the system to have the ability to adaptively adjust the threshold. This system is particularly suitable for application scenarios such as literature comprehensive analysis and interdisciplinary literature mining that require quickly constructing a high-quality literature library.

[0023] The present invention is an efficient, accurate and training data-free literature retrieval solution. Brief Description of the Drawings

[0024] Figure 1 It is the method flow chart in the embodiments of this specification Detailed Embodiments

[0025] The following will further elaborate on the content of the present invention in conjunction with embodiments, but it is not a limitation to the present invention. Embodiment

[0026] Taking the search for relevant English literature with the title "THE INVERSION PROBLEM: Why Algorithms Should Infer Mental State and Not Just Predict Behavior" as an example, the technical solution of the present invention will be further described in detail.

[0027] An intelligent literature screening method and system based on a zero-shot voter, comprising the following steps: Step S1. Obtain the academic data of the target literature: Based on the given literature title "THE INVERSION PROBLEM: Why Algorithms Should Infer Mental State and Not Just Predict Behavior", retrieve the matching literature through the academic database API - SemanticScholar API, and obtain its unique identifier "783ba3f3bfc5656a4d2d7fdea48f4cab6b61a87d". Then use this identifier to further query the metadata of the literature, including the publication year, the number of references, and the number of citations, for subsequent analysis.

[0028] Step S2. Expand and screen the relevant literature set based on the seed literature: Use the literature obtained in Step S1 as the seed literature, and batch obtain the reference literature and citation literature sets of each seed literature through SemanticScholar API. Extract the complete metadata including the unique identifier, title, abstract, publication source, year, and citation metrics of these literatures, and construct a preliminary literature information database. Then merge the literature information database constructed this time with the existing literature information database of the system and remove duplicate literatures.

[0029] Step S3. Construct multiple zero-shot voters: Use the literature information database obtained in Step S2 to establish multiple zero-shot voters. Given the topics "Behavior Prediction, Cognitive Science, Ethical AI", the zero-shot voter can classify the unseen literature without any training samples and determine whether it is relevant to the topic.

[0030] Step S4. Filter out the eligible documents: Integrate the classification results of multiple zero-shot voters in Step S3, and filter out the eligible documents. Among them, the output results of multiple zero-shot voters in Step S3 are used as the input, and a majority voting mechanism is adopted for comprehensive determination. First, count the number of votes each document receives in each classifier. When the cumulative number of votes of a certain document exceeds half of the total number of votes, classify this document as a document that conforms to the theme, retain it in the document information database, otherwise remove it from the document information database. The final output result of this step is multiple documents related to the theme "Behavior Prediction, Cognitive Science, Ethical AI".

[0031] Step S5: Expand the documents: Based on the multiple documents related to the theme obtained in Step S4, use these documents as seed documents again, and repeat Steps S2 - S4 for these seed documents to expand the document information database and obtain more documents related to the theme. The number of repetitions is not limited and depends on user requirements. Finally, an expanded document information database is obtained.

[0032] In the said Step S3, there are a total of six zero-shot voters. The first zero-shot voter performs matching and screening based on the given topic words "Behavior Prediction, Cognitive Science, Ethical AI". Specifically, first set three topic words "Behavior Prediction, Cognitive Science, Ethical AI", and then traverse each document in the document information database obtained in Step S2 to check whether any of the topic words are included in its title and abstract. For the documents that match successfully, record them as 1 in the topic word marking dictionary, meaning that this document is related to the topic word; those that do not match are marked as 0, meaning that this document is not related to the topic word. The final output of this step is a dictionary containing the marking status of all documents, which is used as the voting result of this zero-shot voter.

[0033] The second zero-shot voter is based on the pre-trained language model BART. Use the pre-trained language model BART-large-MNLI to achieve zero-shot classification. First, construct a candidate label set, and then for each document in the document information database obtained in Step S2, process the combined text of its title and abstract in parallel. There are two labels in the candidate label set. The label 1 indicates that it is related to the theme, and the label 0 indicates that it is not related to the theme. Through the multi-threaded acceleration technology, the model automatically determines the relevance of the document to the target technical theme, marks the matching documents as 1, and the non-matching documents as 0. The final output is a dictionary containing the marking status of all documents, which is used as the voting result of this zero-shot voter.

[0034] The third zero-shot voter is based on the pre-trained language model RoBERTa. The RoBERTa-large-MNLI pre-trained model is used to achieve zero-shot classification. First, a candidate label set is constructed, including two labels. A label of 1 indicates relevance to the topic, and a label of 0 indicates irrelevance to the topic. Then, for each document in the literature information library obtained in step S2, the combined text of its title and abstract is processed in parallel. The model automatically generates a relevance score of the document to the target topic, marks the matching documents as 1, and the non-matching documents as 0. Finally, a dictionary containing the marking status of all documents is output as the voting result of this zero-shot voter.

[0035] The fourth zero-shot voter is based on the voting of multiple large language models. The voting committee is composed of three large language models, GPT-3.5-turbo, GPT-4, and O1-mini. Each document in the literature information library obtained in step 2 is input into each large model, and the models are required to perform binary classification on the documents through prompt words to determine whether each document is relevant to the target topic. For example, the input prompt word can be, but is not limited to, "Is this document related to 'Behavior Prediction, Cognitive Science, Ethical AI'?". If two or three models consider it relevant, each document is marked as 1, otherwise it is marked as 0. Finally, a dictionary containing the marking status of all documents is output as the voting result of this zero-shot voter.

[0036] The fifth zero-shot voter is based on Jina embedding vectors. In this step, based on the literature dataset constructed in step S2, the Jina Embeddings V3 model is used to achieve in-depth semantic representation analysis of the literature. This model is optimized specifically for retrieval tasks, generates high-dimensional vector representations, and can capture complex semantic features in the literature title and abstract. This zero-shot voter first normalizes the text content of each document, generates embedding vectors in batches through parallel requests, and then generates embedding vectors for the target topic words "Behavior Prediction, Cognitive Science, Ethical AI". Then, the similarity between the embedding vector of the target topic word and the text embedding vector of the document is calculated. When the similarity exceeds the threshold, the document is considered relevant to the target topic, and at this time, the document is marked as 1, otherwise it is marked as 0. Finally, a dictionary containing the marking status of all documents is output as the voting result of this zero-shot voter.

[0037] The threshold used by the fifth zero-shot voter is determined by an improved binary search algorithm, which can intelligently determine the optimal similarity threshold for document classification. This method first sorts the text embedding vectors of all documents in descending order according to their semantic similarity with the embedding vector of the target topic word, forming an ordered document sequence. The system first checks the key documents at both ends of the sequence: if the document with the highest similarity is determined to be irrelevant, the standard is relaxed and the threshold is lowered; if the document with the lowest similarity is determined to be relevant, the standard is tightened and the threshold is raised. Under normal circumstances, the system adopts a binary search strategy to quickly locate the classification boundary by continuously dividing the search range in half. In each iteration, the system selects the document at the middle position and uses the fourth zero-shot voter in step S3 for determination, and adjusts the search range according to the determination result. When a preliminary boundary point is found, the system checks multiple adjacent documents and determines the final threshold by calculating the average similarity of these boundary documents to ensure the stability and reliability of the result.

[0038] The sixth zero-shot voter is based on OpenAI embedding vectors. In this step, OpenAI's text-embedding-3-small model is used to represent and classify document data. First, the system batch processes the title and abstract information of each document through the API interface to generate high-quality text embedding vectors. After the embedding generation is completed, the semantic similarity between each document and the target topic words "Behavior Prediction, Cognitive Science, EthicalAI" is calculated based on the generated embedding vectors. The threshold determination algorithm in the fifth zero-shot voter is reused to achieve document classification through the dynamically determined optimal threshold. When the similarity exceeds the threshold, the document is considered relevant to the topic, and at this time, the document is marked as 1, otherwise it is marked as 0. The final output is a dictionary containing the marked status of all documents, which is used as the voting result of this zero-shot voter.

[0039] A system for an intelligent document screening method based on zero-shot voters, including: Document metadata collection module: used to obtain the metadata related to the document, such as title, abstract, year, etc., when a document title or the unique identifier of the document is given; Document information library module: For seed documents and documents expanded according to seed documents, they are saved in the document information library module. When a document in the document information library module is determined by the zero-shot voting module to be not in line with the theme, the document is excluded, and when it is determined to be relevant to the theme, it is retained.

[0040] Zero-shot voting module: It includes six zero-shot voters. Each zero-shot voter respectively determines whether a document is relevant to a given topic. The zero-shot voting module makes a comprehensive determination using the majority voting mechanism based on the output results of these zero-shot voters, and filters out the documents that meet the requirements.

[0041] Through the present invention, a large number of documents related to the retrieval topic are obtained. By introducing multiple zero-shot voters, the limitations of traditional keyword retrieval are avoided, making the retrieval more efficient and the retrieval results more accurate.

Claims

1. A literature intelligent screening method based on a zero-shot voter, characterized in that It includes the following steps: Step S1: Obtain the metadata of the target literature. Based on the given literature title, retrieve the matching literature through the academic database API and obtain its unique identifier. Then, use this identifier to further query the metadata of the literature, including the publication year, the number of references, and the number of citations, for subsequent analysis; Step S2: Expand based on the seed literature. Obtain the references and cited literature of the seed literature, construct a preliminary literature information library, and remove duplicate literature; Step S3: Construct multiple zero-shot voters to classify the literature and determine whether it is relevant to the given topic respectively; Step S4: Integrate the classification results of multiple zero-shot voters, adopt the majority voting mechanism to screen the literature that meets the requirements, and remove the literature that is not relevant to the given topic from the literature information library; Step S5: Use the literature screened in Step S4 as the new seed literature to iteratively expand the literature information library.

2. The literature intelligent screening method based on a zero-shot voter according to claim 1, wherein In Step S2, use the literature obtained in Step S1 as the seed literature. Batch obtain the reference and cited literature sets of each seed literature through the API, extract the complete metadata including the unique identifier, title, abstract, publication source, year, and citation index of these literatures, construct a preliminary literature information library, and then merge the literature information library constructed this time with the existing literature information library in the system and remove duplicate literature.

3. The literature intelligent screening method based on a zero-shot voter according to claim 1, characterized in that The multiple zero-shot voters constructed in Step S3 include the following six types: Voter based on subject term matching: For each literature in the literature information library in Step S2, check whether the literature title / abstract contains the target subject term to determine whether it is relevant to the target topic; Voter based on the pre-trained language model BART model: Adopt BART-large-MNLI for zero-shot classification to determine whether each literature in the literature information library in Step S2 is relevant to the target topic; Voter based on the pre-trained language model RoBERTa model: Adopt RoBERTa-large-MNLI for zero-shot classification to determine whether each literature in the literature information library in Step S2 is relevant to the target topic; Voter based on voting of multiple large language models: Integrate GPT-3.5, GPT-4, and O1-mini for majority voting to determine whether each literature in the literature information library in Step S2 is relevant to the target topic; Voter based on Jina embedding vectors: Use the Jina Embeddings V3 model to perform semantic representation on each literature in the literature information library in Step S2, and then perform semantic representation on the target subject term; Calculate the semantic similarity between each literature and the subject term based on the obtained embedding vectors, and compare the similarity with the threshold to determine whether each literature is relevant to the target topic; Voter based on OpenAI embedding vectors: The text-embedding-3-small model of OpenAI is used to perform semantic representation on each document in the document information library in step S2, and then semantic representation is performed on the target subject term; The semantic similarity between each document and the subject term is calculated based on the obtained embedding vectors, and the similarity is compared with a threshold to determine whether each document is relevant to the target subject.

4. The method for intelligent literature screening based on a zero-shot voter according to claim 3, wherein, The thresholds adopted by the voter based on Jina embedding vectors and the voter based on OpenAI embedding vectors are determined by an improved binary search algorithm, which can intelligently determine the optimal similarity threshold for document classification, as follows: First, the text embedding vectors of all documents are sorted in descending order according to their semantic similarity with the target subject term embedding vector, forming an ordered document sequence. Check the key documents at both ends of the sequence: If the document with the highest similarity is determined to be irrelevant, relax the standard and lower the threshold; If the document with the lowest similarity is determined to be relevant, tighten the standard and raise the threshold; Adopt a binary search strategy to quickly locate the classification boundary by continuously dividing the search range in half. In each iteration, the system selects the document at the middle position and uses the voter based on voting of multiple large language models in step S3 for determination, and adjusts the search range according to the determination result; After finding the preliminary boundary point, check multiple adjacent documents, and determine the final threshold by calculating the average value of the similarities of these boundary documents to ensure the stability and reliability of the result.

5. The method for intelligent literature screening based on a zero-shot voter according to claim 3, characterized in that The classification process of the zero-shot voter adopts multi-threaded parallel processing, including batch generation of Jina embedding vectors and OpenAI embedding vectors; Multi-threaded acceleration of zero-shot classification of the BART / RoBERTa model; Asynchronous invocation of large language models to achieve multi-threaded acceleration.

6. The intelligent literature screening method based on a zero-shot voter according to claim 1, wherein, In step S4, the classification results of multiple zero-shot voters in step S3 are integrated, and the documents that meet the requirements are screened out. Among them, the output results of multiple zero-shot voters in step S3 are used as inputs, and a majority voting mechanism is adopted for comprehensive determination. First, count the number of votes obtained by each document in each classifier. When the cumulative number of votes of a certain document exceeds half of the total number of votes, the document is classified as a document that meets the theme and is retained in the document information library, otherwise it is removed from the document information library.

7. The method for intelligent literature screening based on a zero-shot voter according to claim 1, characterized in that, The iterative expansion of documents in step S5 includes: Using the screened documents as new seed documents, repeating steps S2, S3, and S4; After each iteration, merge the document information library and remove duplicate documents; The number of iterations is determined by user requirements until the requirements for the scale or coverage of the document library are met.

8. A system for an intelligent literature screening method based on a zero-shot voter, characterized in that, It includes the following modules: Document metadata collection module: Used to obtain the metadata related to a document when a document title or the unique identifier of the document is given; Document information library module: For seed documents and documents expanded based on seed documents, save them to the document information library module. When a document in the document information library module is determined by the zero-shot voting module to be not in line with the theme, the document is removed, and when it is determined to be relevant to the theme, it is retained; Zero-shot voting module: It includes six zero-shot voters. Each zero-shot voter respectively determines whether a document is relevant to a given topic. The zero-shot voting module makes a comprehensive judgment by adopting a majority voting mechanism based on the output results of these zero-shot voters, and screens out the documents that meet the requirements.

Citation Information

Cited By

  • Literature review generation method, electronic equipment, storage medium and program product

    CN120781846A

  • Document review generation method, electronic device, storage medium, and program product

    CN120781846B

  • Model training method and device, electronic equipment and storage medium

    CN121434770A

  • Active learning model training method and device for constructing dataset

    CN121434770B