A large model-based collaborative and innovative batch analysis system and method
By using a collaborative, innovative batch analysis system based on a large model, the problems of inaccurate classification and limitations of large models in literature novelty searches are solved, enabling efficient batch literature analysis and novelty judgment, and is suitable for science and technology and patent novelty searches.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHEJIANG UNIV
- Filing Date
- 2026-03-06
- Publication Date
- 2026-07-03
Smart Images

Figure CN122332580A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of artificial intelligence and literature novelty search, and in particular to a collaborative, innovative batch analysis system and method based on a large model. Background Technology
[0002] In traditional literature and technology novelty searches, searchers rely on manual methods to conduct large-scale literature searches and content analysis, which involves repetitive and arduous work.
[0003] Chinese patent application number CN201910226293.8 discloses a method for classifying and storing historical documents based on big data analysis. By setting up storage, update, and backup modules, it facilitates the storage, replacement, and backup of document information, preventing duplicate storage, disordered backups, and wasting staff time on reorganization, thus reducing work efficiency. However, this method still has shortcomings, such as incomplete document classification and retrieval functions, lack of retrieval comparison analysis, and insufficient classification accuracy, failing to meet the needs of users.
[0004] Chinese patent application number CN202110554334.3 discloses a knowledge graph-based method for classifying scientific and technological documents, including the following steps: document acquisition step: acquiring scientific and technological documents to be classified; text preprocessing step: performing lexical analysis on the scientific and technological documents to obtain part-of-speech tags, and filtering based on the part-of-speech tags. However, this method still has shortcomings. It only classifies scientific and technological documents by extracting relevant keywords from the documents. Therefore, users can only search for documents through the knowledge base using relatively specific keywords. Thus, when users do not have specific keywords they want to search for, or when the knowledge base cannot detect relevant content after entering corresponding keywords, the knowledge base that only classifies by keywords will not meet the needs of the users.
[0005] In recent years, large language models, with their strong text understanding and generation capabilities, have been able to assist in document retrieval and content analysis to a certain extent. However, they still have significant limitations in practical applications: First, general-purpose large models may have biases in their analysis of innovative points and lack domain adaptability; second, models generally suffer from "illusion" phenomena, which may generate false information, making it difficult for them to independently undertake all novelty search tasks; third, conventional conversational interaction modes are limited by the length of the context, making it difficult to support efficient batch processing of multiple documents, and the cost of human-computer communication is relatively high. Summary of the Invention
[0006] The purpose of this invention is to solve the problems in the prior art by proposing a collaborative batch analysis system and method for innovativeness based on a large model. This system can perform hierarchical refinement and keyword extraction of innovative points proposed by users, obtain search results on a designated information platform, perform batch content analysis and relevance comparison between user innovative points and search results, and finally form a conclusion on whether the innovativeness is achieved.
[0007] To achieve the above objectives, this invention proposes a collaborative, innovative batch analysis system based on a large model, comprising: Innovation Point Layering Module: Based on the actual content of the innovation points, a large model is used to divide the innovation points into core innovation points, general innovation points, and auxiliary innovation points, and intelligently assign weight values from high to low. Keyword extraction module: Based on hierarchical novelty points, keywords are extracted using a large model, synonyms and near-synonyms are expanded, search elements are constructed, and Chinese and English search formulas are written for different search categories and platforms; Content retrieval module: Performs multi-platform searches based on search queries, and summarizes and deduplicates the results; Content comparison module: Batch analysis of search results content, using a large model and hierarchical innovation points to compare content relevance, and combining weighted intelligent calculation of final relevance for three-level classification: highly relevant, generally relevant, and irrelevant; Content summary module: The large model is used to summarize and refine the results, integrate them to form the final conclusion, and judge whether the user's innovation points have actual innovation. User interaction module: This module provides a user interface and feeds back the processing results of each module to the user, enabling full-process calibration through human-machine collaboration.
[0008] As a preferred embodiment, the innovation point layering module, based on prompt word engineering, RAG and model fine-tuning technology, optimizes the vertical scenario capabilities of the general large model through novelty search case data and industry datasets of novelty search specifications, divides user innovation points into multiple levels such as core innovation points, general innovation points and auxiliary innovation points, and intelligently allocates weights.
[0009] As a preferred embodiment, the keyword extraction module: based on prompt word engineering, RAG and model fine-tuning technology, optimizes the vertical scenario capabilities of the general large model through a dataset of retrieval specification and retrieval example datasets, extracts keywords based on hierarchical innovation points, expands synonyms and near-synonyms, constructs retrieval elements, and writes Chinese and English retrieval expressions.
[0010] Preferably, the content retrieval module: connects to an open and supported retrieval platform based on the MCP protocol, matches the corresponding retrieval formula, and obtains retrieval results; at the same time, it receives results uploaded by users from their own searches on the retrieval platform and performs deduplication and aggregation on all results.
[0011] As a preferred embodiment, the content comparison module: splits the full search results using regular expression rules; based on prompt word engineering, RAG and model fine-tuning technology, it optimizes the capabilities of the general large model for vertical scenarios using new search case data and industry datasets of new search specifications, builds a multi-model parallel cyclic processing workflow for batch content extraction and comparison, and combines weighted intelligent calculation of the final relevance, dividing it into a three-layer structure of highly relevant, generally relevant and irrelevant.
[0012] As a preferred embodiment, the content summary module: splits the comparison results specified by the user using regular expression rules; based on prompt word engineering, RAG, and model fine-tuning technology, it optimizes the vertical scenario capabilities of the general large model using novelty search case data and industry datasets of novelty search specifications, builds a multi-model parallel cyclic processing workflow, analyzes and summarizes the comparison results, and forms sub-conclusions; and integrates multiple sub-conclusions again through the large model to form the final conclusion.
[0013] Preferably, the user interaction module provides an intuitive graphical interface and integrates with the innovation point layering module, keyword extraction module, content retrieval module, content comparison module, and content summary module via API. This module receives user innovation points and manual calibration results from each stage and feeds back the model processing results to the user.
[0014] This invention proposes an analytical method for a collaborative and innovative batch analysis system based on a large model, comprising the following steps: S1: Deploy the server and client. The server is deployed on a local or cloud computing server cluster, and the client is deployed on various user terminal systems using a B / S or C / S architecture. S2: Call the user interaction module to receive the user's original innovative ideas; S3: Call the innovation point layering module to divide the user's original innovation points into core innovation points, general innovation points and auxiliary innovation points according to their criticality and intelligently allocate weights. Feed back the layered innovation points to the user interaction module, and the user can choose to calibrate or directly proceed to the next step. S4: Call the keyword extraction module, extract keywords based on the hierarchical innovation points, expand synonyms and near-synonyms, construct search elements, and feed the search query back to the user interaction module according to the written Chinese and English search query. The user can choose to calibrate or directly proceed to the next step. S5: Call the content retrieval module, initiate a search on the open and supported retrieval platform based on Chinese and English search terms through the MCP plugin, and feed the search results back to the user interaction module. The user can choose to supplement and upload external search results and initiate deduplication and integration. The final results are fed back to the user interaction module again, and the user can choose to continue to supplement, modify, filter or directly proceed to the next step. S6: Call the content comparison module to perform batch content relevance analysis between the multi-layer innovation points and the final search result set. Combine the weight intelligent calculation of the final relevance and divide it into three categories: highly relevant, generally relevant and irrelevant. Feed the comparison result set back to the user interaction module. The user can choose to re-initiate the comparison, modify, filter or directly proceed to the next step. S7: Call the content summary module to perform batch sub-conclusion analysis and integration of the final comparison result set, form the final conclusion and feed it back to the user interaction module, where the user can choose to re-initiate the summary or make manual corrections.
[0015] Preferably, in step S6, the content comparison module splits the final retrieval result set and the multi-layer innovation point comparison into multiple innovation point subsets, performs content relevance analysis on each innovation point subset, summarizes the content relevance analysis results of all innovation point subsets, and combines the weights to filter the result set to obtain the comparison result set.
[0016] Preferably, in step S7, the content summary module splits the final comparison result set into multiple comparison result subsets using regular expression rules, writes sub-conclusions for each comparison result subset, and summarizes the sub-conclusions of all comparison result subsets to write a general conclusion, thus obtaining the final conclusion.
[0017] The beneficial effects of this invention are as follows: By integrating multiple models and related technologies, this invention optimizes innovative analysis tasks in a targeted manner and introduces a human-machine collaborative verification mechanism at key analysis nodes. Based on this, the entire process can be systematically encapsulated, forming a tool-based product that combines batch automation and result reliability. Based on the capabilities of a large model, this invention performs hierarchical refinement and keyword extraction on user-proposed innovative points, obtains search results on a designated information platform, and performs batch content analysis and relevance comparison between the user's innovative points and the search results to ultimately determine whether the innovation is valid. This invention enhances the large model's ability to analyze and compare novelty points through hierarchical novelty point analysis and avoids problems such as frequent "illusions," limited context windows, and easy loss of long-range memory inherent in large models through human-machine collaboration and batch processing workflows. This significantly alleviates repetitive labor in manual retrieval and can be directly applied to practical business scenarios such as technology novelty searches and patent novelty searches.
[0018] The features and advantages of the present invention will be described in detail through embodiments and in conjunction with the accompanying drawings. Attached Figure Description
[0019] Figure 1 This is an architecture diagram of a collaborative and innovative batch analysis system based on a large model, according to the present invention. Figure 2 This is a flowchart of the analysis method of a collaborative and innovative batch analysis system based on a large model, according to the present invention. Detailed Implementation
[0020] See Figure 1 and Figure 2 This invention discloses a collaborative and innovative batch analysis system based on a large model, comprising: Innovation Point Layering Module: Based on the actual content of the innovation points, a large model is used to divide them into multiple layers such as core innovation points, general innovation points, and auxiliary innovation points, and intelligently assign weight values from high to low. Keyword extraction module: Based on hierarchical novelty points, keywords are extracted using a large model, synonyms and near-synonyms are expanded, search elements are constructed, and Chinese and English search formulas are written for different search categories and platforms; Content retrieval module: Enables multi-platform retrieval based on Chinese and English search queries, and summarizes and deduplicates the results; Content comparison module: It uses a large model and hierarchical innovation points to compare the relevance of content, and combines weighted intelligent calculation to achieve a three-level division of highly relevant, generally relevant and irrelevant. Content summary module: The results are summarized and refined using large model comparisons, and integrated to form a final conclusion, judging whether the user's innovative points have actual innovativeness; User interaction module: Used to provide a user-friendly interface, feed back the processing results of each module to the user, and realize the whole process calibration of human-machine collaboration; The innovation point layering module, based on prompt word engineering, RAG, and model fine-tuning techniques, optimizes the capabilities of a general-purpose model for vertical scenarios using industry datasets such as novelty search case data and novelty search specifications. It categorizes user innovation points into multiple levels, including core innovation points, general innovation points, and auxiliary innovation points, and intelligently assigns weights accordingly. The keyword extraction module, also based on prompt word engineering, RAG, and model fine-tuning techniques, optimizes the capabilities of a general-purpose model for vertical scenarios using datasets such as retrieval specification and retrieval examples. Based on the layered innovation points, it extracts keywords, expands synonyms and near-synonyms, constructs retrieval elements, and writes Chinese and English retrieval expressions. The hierarchical module and keyword extraction module are integrated into an intelligent agent for innovation point analysis. The analysis results are fed back to the user interaction module via API, where users can choose to manually correct them or proceed directly to the next step. The content retrieval module connects to open and supported retrieval platforms such as ArXiv based on the MCP protocol, matches the corresponding retrieval formula specifications, obtains retrieval results, and feeds them back to the user. It also receives results uploaded by users from platforms such as CNKI and WOS, and performs deduplication and aggregation on all results. The content comparison module splits all retrieval results using regular expressions and other rules. Based on prompt word engineering, RAG, and model fine-tuning technology, it analyzes novelty search case data, Industry datasets such as novelty search standards are used to optimize the capabilities of general-purpose large models for vertical scenarios. A multi-model parallel cyclic processing workflow is built to achieve batch content extraction and comparison. The final relevance is calculated intelligently using weights and divided into three layers: highly relevant, generally relevant, and irrelevant, and then fed back to the user. The content retrieval module and content comparison module are integrated into a content analysis batch processing workflow. Retrieval and comparison results are fed back to the user interaction module via API. Users can choose to re-initiate the comparison, manually supplement the retrieval / modify / filter, or directly proceed to the next step. The content summarization module splits the user-specified comparison results using regular expressions and other rules; based on prompt words... The engineering, RAG, and model fine-tuning technologies utilize industry datasets such as novelty search case data and novelty search specifications to optimize the capabilities of a general-purpose large model for vertical scenarios. A multi-model parallel cyclic processing workflow is established, and the results are compared, analyzed, and summarized to form sub-conclusions. These sub-conclusions are then further integrated through the large model to form the final conclusion, which is fed back to the user interaction module via API. Users can choose to re-initiate the summary or make manual corrections. The user interaction module provides an intuitive graphical interface and integrates the above five functional modules via API. It receives user innovation points and manual calibration results from each stage and feeds back the model processing results to the user.
[0021] The present invention discloses an analytical method for a collaborative, innovative batch analysis system based on a large model, comprising the following steps: S1: Deploy the server and client. The server is deployed on a local or cloud computing server cluster, and the client is deployed on various user terminal systems using a B / S or C / S architecture. S2: The user interaction module is invoked to first receive the user's original innovative ideas; S3: Call the innovation point layering module to divide the user's original innovation points into multiple layers such as core innovation points, general innovation points, and auxiliary innovation points according to their criticality and intelligently allocate weights. Feed back the hierarchical innovation points to the user interaction module, where the user can choose to calibrate or directly proceed to the next step. S4: Call the keyword extraction module, extract keywords based on the hierarchical innovation points, expand synonyms and near-synonyms, construct search elements, and feed the search query back to the user interaction module according to the written Chinese and English search query. The user can choose to calibrate or directly proceed to the next step. S5: Call the content retrieval module, initiate a search on open and supported retrieval platforms such as ArXiv based on Chinese and English search terms through the MCP plugin, and feed the search results back to the user interaction module. Users can supplement and upload external search results and initiate deduplication and integration. The final results are fed back to the user interaction module again, and users can continue to supplement / modify / filter or directly proceed to the next step. S6: Call the content comparison module to perform batch content relevance analysis between the multi-layered innovation points and the final search result set. Combine the weight intelligent calculation of the final relevance and divide it into three categories: highly relevant, generally relevant, and irrelevant. Feed the comparison result set back to the user interaction module. Users can choose to re-initiate the comparison, modify / filter, or directly proceed to the next step. S7: Call the content summary module to perform batch sub-conclusion analysis and integration of the final comparison result set, form the final conclusion and feed it back to the user interaction module. Users can choose to re-initiate the summary or make manual corrections. In step S6, the content comparison module splits the final retrieval result set and the multi-layer innovation point comparison into multiple innovation point subsets, performs content relevance analysis on each innovation point subset, summarizes the content relevance analysis results of all innovation point subsets, and filters the result set by combining weights to obtain the comparison result set. In step S7, the content summary module splits the final comparison result set into multiple comparison result subsets according to the rules of regular expressions, writes sub-conclusions for each comparison result subset, summarizes the sub-conclusions of all comparison result subsets to write the final conclusion.
[0022] This invention, based on the capabilities of a large model, performs hierarchical refinement and keyword extraction on user-proposed innovative points, and obtains search results on a designated information platform. It then performs batch content analysis and relevance comparison between the user's innovative points and the search results to ultimately determine whether the innovativeness is achieved. This invention enhances the large model's ability to analyze and compare novelty points by hierarchically refining them, and avoids problems such as frequent "illusions," limited context windows, and easy loss of long-range memory that exist in large models through human-computer collaboration and batch processing workflows. It significantly alleviates repetitive work in manual retrieval and can be directly applied to practical business scenarios such as technology novelty searches and patent novelty searches.
[0023] The above embodiments are illustrative of the present invention and are not intended to limit the present invention. Any simple modifications to the present invention are within the scope of protection of the present invention.
Claims
1. A large model based collaboratively innovative batch analysis system, characterized by: include: Innovation Point Layering Module: Based on the actual content of the innovation points, a large model is used to divide the innovation points into core innovation points, general innovation points, and auxiliary innovation points, and intelligently assign weight values from high to low. Keyword extraction module: Based on hierarchical novelty points, keywords are extracted using a large model, synonyms and near-synonyms are expanded, search elements are constructed, and Chinese and English search formulas are written for different search categories and platforms; Content retrieval module: Performs multi-platform searches based on search queries, and summarizes and deduplicates the results; Content comparison module: Batch analysis of search results content, using a large model and hierarchical innovation points to compare content relevance, and combining weighted intelligent calculation of final relevance for three-level classification: highly relevant, generally relevant, and irrelevant; Content summary module: The large model is used to summarize and refine the results, integrate them to form the final conclusion, and judge whether the user's innovation points have actual innovation. User interaction module: This module provides a user interface and feeds back the processing results of each module to the user, enabling full-process calibration through human-machine collaboration.
2. The large model based collaborable innovative batch analysis system of claim 1, wherein: The innovation point layering module, based on prompt word engineering, RAG and model fine-tuning technology, optimizes the vertical scenario capabilities of the general large model through novelty search case data and industry datasets of novelty search standards, divides user innovation points into multiple levels such as core innovation points, general innovation points and auxiliary innovation points, and intelligently allocates weights.
3. The large model based collaborable innovative batch analysis system of claim 1, wherein: The keyword extraction module, based on prompt word engineering, RAG, and model fine-tuning technology, optimizes the capabilities of a general large model for vertical scenarios through a dataset of search query specifications and examples. Based on hierarchical innovation points, it extracts keywords, expands synonyms and near-synonyms, constructs search elements, and writes Chinese and English search queries.
4. The large model based collaborable innovative batch analysis system of claim 1, wherein: The content retrieval module: Based on the MCP protocol, it connects to open and supported retrieval platforms, matches the corresponding retrieval specifications, and obtains retrieval results; at the same time, it receives the results that users upload and retrieve on the retrieval platform, and performs deduplication and aggregation on all results.
5. The large model based collaborable innovative batch analysis system of claim 1, wherein: The content comparison module splits the full search results using regular expression rules; based on prompt word engineering, RAG, and model fine-tuning technology, it optimizes the capabilities of the general large model for vertical scenarios using novelty search case data and industry datasets of novelty search specifications, and builds a multi-model parallel cyclic processing workflow for batch content extraction and comparison. It combines weighted intelligent calculation of the final relevance and divides it into a three-layer structure of highly relevant, generally relevant, and irrelevant.
6. The large model based collaborable innovative batch analysis system of claim 1, wherein: The content summary module: splits the comparison results specified by the user using regular expression rules; based on prompt word engineering, RAG and model fine-tuning technology, it optimizes the vertical scenario capabilities of the general large model through new search case data and industry datasets of new search specifications, builds a multi-model parallel loop processing workflow, analyzes and summarizes the comparison results, and forms sub-conclusions; and integrates multiple sub-conclusions again through the large model to form the final conclusion.
7. The large model based collaborable innovative batch analysis system of claim 1, wherein: The user interaction module provides an intuitive graphical interface and integrates with the innovation point layering module, keyword extraction module, content retrieval module, content comparison module, and content summary module via API. It is used to receive user innovation points and manual calibration results at each stage, and to feed back the model processing results to the user.
8. The analysis method of a collaborative and innovative batch analysis system based on a large model according to claims 1-7, characterized in that: Includes the following steps: S1: Deploy the server and client. The server is deployed on a local or cloud computing server cluster, and the client is deployed on various user terminal systems using a B / S or C / S architecture. S2: Call the user interaction module to receive the user's original innovative ideas; S3: Call the innovation point layering module to divide the user's original innovation points into core innovation points, general innovation points and auxiliary innovation points according to their criticality and intelligently allocate weights. Feed back the layered innovation points to the user interaction module, and the user can choose to calibrate or directly proceed to the next step. S4: Call the keyword extraction module, extract keywords based on the hierarchical innovation points, expand synonyms and near-synonyms, construct search elements, and feed the search query back to the user interaction module according to the written Chinese and English search query. The user can choose to calibrate or directly proceed to the next step. S5: Call the content retrieval module, initiate a search on the open and supported retrieval platform based on Chinese and English search terms through the MCP plugin, and feed the search results back to the user interaction module. The user can choose to supplement and upload external search results and initiate deduplication and integration. The final results are fed back to the user interaction module again, and the user can choose to continue to supplement, modify, filter or directly proceed to the next step. S6: Call the content comparison module to perform batch content relevance analysis between the multi-layer innovation points and the final search result set. Combine the weight intelligent calculation of the final relevance and divide it into three categories: highly relevant, generally relevant and irrelevant. Feed the comparison result set back to the user interaction module. The user can choose to re-initiate the comparison, modify, filter or directly proceed to the next step. S7: Call the content summary module to perform batch sub-conclusion analysis and integration of the final comparison result set, form the final conclusion and feed it back to the user interaction module, where the user can choose to re-initiate the summary or make manual corrections.
9. The analysis method of a collaborative innovative batch analysis system based on a large model as described in claim 8, characterized in that: In step S6, the content comparison module splits the final retrieval result set and the multi-layer innovation point comparison into multiple innovation point subsets, performs content relevance analysis on each innovation point subset, summarizes the content relevance analysis results of all innovation point subsets, and combines the weights to filter the result set to obtain the comparison result set.
10. The analysis method of a collaborative innovative batch analysis system based on a large model as described in claim 8, characterized in that: In step S7, the content summary module splits the final comparison result set into multiple comparison result subsets using regular expression rules, writes sub-conclusions for each comparison result subset, and summarizes the sub-conclusions of all comparison result subsets to write a general conclusion, thus obtaining the final conclusion.
Citation Information
Patent Citations
Historical document classified storage method based on big data analysis
CN109977076A
Scientific and technical literature classification method based on knowledge graph
CN113239201A