Medical literature intelligent duplicate checking and hierarchical recommendation interaction system based on subject service
By constructing an intelligent plagiarism detection and hierarchical recommendation interactive system for medical literature based on subject services, the system solves the problems of accuracy and collaboration in the medical field of existing systems, realizes efficient plagiarism detection and accurate recommendation of medical literature, adapts to the needs of medical research development, and improves research efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TIANJIN MEDICAL UNIV
- Filing Date
- 2026-02-02
- Publication Date
- 2026-05-12
AI Technical Summary
Existing literature processing systems lack specific adaptation designs for the medical field, making it difficult to accurately capture semantic relationships in medical terms and differences in clinical trial designs. This leads to biases in similarity judgments during plagiarism checks, making it impossible to effectively distinguish between repetitive and innovative points in core research elements of medical literature. Furthermore, the coordination between the various functional modules of the system is insufficient, failing to meet the full-process needs of researchers.
A subject-based intelligent plagiarism detection and hierarchical recommendation interactive system for medical literature is constructed, including a medical-specific plagiarism detection module, a subject-adaptive hierarchical recommendation module, a research scenario interaction module, and a dynamic subject resource module. Data interaction and parameter synchronization are achieved through internal standardized data interfaces. Combined with the medical ontology knowledge base and the dynamic subject resource module, multi-dimensional similarity calculation and hierarchical recommendation are performed.
Significantly improves the accuracy of plagiarism detection and recommendation, adapts to the characteristics of medical disciplines, dynamically adapts to the needs of medical research development, realizes a closed-loop service throughout the entire process, and improves the efficiency of medical research work.
Smart Images

Figure CN122019722A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of document service technology, and in particular to an intelligent plagiarism detection and hierarchical recommendation interactive system for medical documents based on subject services. Background Technology
[0002] The number of medical literatures continues to surge with the development of scientific research, and researchers' need for literature plagiarism detection and accurate recommendation is becoming increasingly urgent.
[0003] Existing literature processing systems are mostly general-purpose architectures, lacking specific designs tailored to the medical field. This makes it difficult to accurately capture core disciplinary features such as semantic relationships in medical terminology and differences in clinical trial designs, leading to biases in similarity assessments and an inability to effectively distinguish between repetitive and innovative points in core research elements of medical literature. Furthermore, the algorithm models and parameter settings of existing systems largely rely on human pre-setting, with insufficient disclosure of technical implementation details, creating a "black box" problem. This fails to meet the requirements of the latest review guidelines, and the resource databases are outdated, unable to adapt to the dynamic changes in medical terminology, standards, and research hotspots. In addition, the functional modules of existing systems lack coordination, data interaction rules are inconsistent, and plagiarism detection results are disconnected from recommendation services, failing to meet researchers' needs throughout the entire process from plagiarism detection to accurate acquisition of suitable literature, thus hindering the improvement of medical research efficiency. Summary of the Invention
[0004] To address the technical problems existing in the prior art, embodiments of the present invention provide an intelligent plagiarism detection and hierarchical recommendation interactive system for medical literature based on subject services. The technical solution is as follows:
[0005] On the one hand, it provides an intelligent plagiarism detection and hierarchical recommendation interactive system for medical literature based on subject services, including a medical-specific plagiarism detection module, a subject-adaptive hierarchical recommendation module, a research scenario interaction module, a medical data preprocessing module, and a dynamic subject resource module. Each module achieves data interaction and parameter synchronization through internal standardized data interfaces. The data interfaces adopt the system's built-in medical data interaction communication rules, and the data transmission format is the system's preset structured format. The data processed by each module is stored in a distributed database.
[0006] The medical data preprocessing module receives medical literature input by users and, relying on the system's built-in medical terminology standard library and clinical trial data specifications, performs format standardization, professional terminology normalization, scientific research element extraction and noise filtering operations to generate structured literature data containing research topics, experimental designs, data indicators and conclusion statements. After the preprocessing operation is completed, the data is transmitted to the medical-specific plagiarism detection module through the internal data interface.
[0007] The medical-specific plagiarism detection module receives structured literature data, calls the medical ontology knowledge base and subject-specific feature base of the dynamic subject resource module, integrates medical semantic association analysis and multi-dimensional similarity calculation, completes similarity detection at the level of full text of literature and scientific research elements, locates duplicate content and associates the subject classification and evidence level of similar literature, and pushes the detection results to the subject adaptation hierarchical recommendation module and scientific research scenario interaction module through internal data interface.
[0008] The Dynamic Discipline Resource Module constructs a four-tiered medical discipline system, using first-level disciplines as a framework to divide second-level branches. These second-level branches are further subdivided by disease subtypes, with each subtype corresponding to a specific research direction. The discipline system covers basic medicine, clinical medicine, preventive medicine, pharmacy, public health, and their respective sub-fields. The system establishes exclusive classification and coding rules for the four-tiered medical discipline system. Simultaneously, the Dynamic Discipline Resource Module constructs exclusive research feature databases and research hotspot maps for each field. The research hotspot maps are updated in real-time with the latest data from the disciplines. The Dynamic Discipline Resource Module provides semantic parsing support data to the medical-specific plagiarism detection module and graded filtering support data to the discipline-adaptive graded recommendation module through fixed data interfaces. When the support data is updated, the parameters of the associated modules are automatically synchronized and adapted.
[0009] The subject-adapted hierarchical recommendation module receives the similarity detection results from the medical-specific plagiarism detection module and the user's research direction tags transmitted by the scientific research scenario interaction module. Combining the four-level medical subject system and research hotspot map of the dynamic subject resource module, it completes multi-dimensional hierarchical screening of literature through the system's built-in medical scientific research scenario adaptation algorithm, generates a hierarchical recommended literature list, and pushes the hierarchical recommended literature list to the scientific research scenario interaction module through the internal data interface.
[0010] The research scenario interaction module receives user operation commands, transmits parameter setting commands to the medical-specific plagiarism detection module and the subject-adaptive hierarchical recommendation module, transforms the result display requirements into a visual data presentation format, and transmits user research direction tags and interaction behavior data to the subject-adaptive hierarchical recommendation module and the dynamic subject resource module, respectively. After being parsed by the research scenario interaction module, the user operation commands are transmitted to the corresponding functional modules, and the corresponding functional modules feed back the processing results to the research scenario interaction module for visual presentation.
[0011] One or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages:
[0012] Adapting to the characteristics of medical disciplines, significantly improving the accuracy of plagiarism detection and recommendation: By constructing a four-level medical discipline system of primary disciplines, secondary branches, disease subtypes, and research directions, coupled with a dedicated feature library of dynamic discipline resource modules, and combined with an attention mechanism semantic model that integrates medical ontology knowledge base, it accurately captures the semantic associations of medical terms, clinical trial design, and other core research elements, effectively solving the problem of semantic parsing bias in general systems, ensuring that similarity judgments are aligned with the characteristics of different medical branches, and hierarchical recommendations can accurately match users' research directions and research stages, providing researchers with highly adaptable literature resources.
[0013] This system clearly defines the implementation logic of each module, data processing rules, and model training mechanism throughout the entire process. There are no manually preset parameters or judgment rules. All weight matrices and similarity thresholds are generated based on training and practical data statistics from the medical annotation corpus. The system fully discloses technical details, completely avoids the black box problem, and ensures the reproducibility of the technical solution.
[0014] With outstanding dynamic adaptability, it meets the needs of medical research development: relying on the real-time capture and topic clustering update mechanism of the dynamic subject resource module, it can keep up with the iteration of medical terminology, standard revision and research hotspot changes. Combined with the system's subject-specific iteration mechanism, it automatically optimizes semantic model parameters, exclusive feature library content and weight configuration. It can achieve dynamic adaptation with medical research trends without manual intervention, and maintain the consistency between the system's service capabilities and the development of the subject in the long term.
[0015] Highly efficient module collaboration enables a closed-loop service across the entire process: Through built-in standardized data interfaces and a unified medical data interaction protocol, it achieves efficient data exchange and parameter synchronization among modules such as medical data preprocessing, deduplication, hierarchical recommendation, research scenario interaction, and dynamic resource updates, constructing a closed-loop service from literature input to result feedback and algorithm optimization. It also supports external platform integration and multi-research scenario export adaptation, balancing service universality and specialization, effectively improving the overall efficiency of medical research. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 A flowchart of the interactive system for intelligent plagiarism detection and hierarchical recommendation of medical literature based on subject services provided in this application embodiment. Detailed Implementation
[0018] The technical solution provided in this application will now be described in conjunction with the accompanying drawings.
[0019] To facilitate understanding of the embodiments of this application, the following points will be explained first: First, in this application, "at least one" means one or more, and "more than one" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character "" generally indicates an "or" relationship between the preceding and following related objects, but it does not exclude the possibility of indicating an "and" relationship; the specific meaning can be understood in context. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can mean: a, b, c; a and b; a and c; b and c; or a and b and c. Here, a, b, and c can be single or multiple.
[0020] Second, the use of prefixes such as "first" and "second" in this application is merely for the purpose of distinguishing and describing different things belonging to the same category, and does not constrain the order, size, or quantity of things. For example, "first message" and "second message" are simply different messages, and there is no chronological, size, or priority relationship between them.
[0021] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.
[0022] like Figure 1 The diagram shown is a flowchart of the interactive system for intelligent plagiarism detection and hierarchical recommendation of medical literature based on subject services provided in this application embodiment. It includes a medical-specific plagiarism detection module, a subject-adaptive hierarchical recommendation module, a research scenario interaction module, a medical data preprocessing module, and a dynamic subject resource module. Each module achieves data interaction and parameter synchronization through an internal standardized data interface. The data interface adopts the system's built-in medical data interaction communication rules, and the data transmission format is a pre-set structured format. The data processed by each module is stored in a distributed database.
[0023] The medical data preprocessing module receives medical literature input by users and, relying on the system's built-in medical terminology standard library and clinical trial data specifications, performs format standardization, professional terminology normalization, scientific research element extraction and noise filtering operations to generate structured literature data containing research topics, experimental designs, data indicators and conclusion statements. After the preprocessing operation is completed, the data is transmitted to the medical-specific plagiarism detection module through the internal data interface.
[0024] The medical-specific plagiarism detection module receives structured literature data, calls the medical ontology knowledge base and subject-specific feature base of the dynamic subject resource module, integrates medical semantic association analysis and multi-dimensional similarity calculation, completes similarity detection at the level of full text of literature and scientific research elements, locates duplicate content and associates the subject classification and evidence level of similar literature, and pushes the detection results to the subject adaptation hierarchical recommendation module and scientific research scenario interaction module through internal data interface.
[0025] The Dynamic Discipline Resource Module constructs a four-tiered medical discipline system, using first-level disciplines as a framework to divide second-level branches. These second-level branches are further subdivided by disease subtypes, with each subtype corresponding to a specific research direction. The discipline system covers basic medicine, clinical medicine, preventive medicine, pharmacy, public health, and their respective sub-fields. The system establishes exclusive classification and coding rules for the four-tiered medical discipline system. Simultaneously, the Dynamic Discipline Resource Module constructs exclusive research feature databases and research hotspot maps for each field. The research hotspot maps are updated in real-time with the latest data from the disciplines. The Dynamic Discipline Resource Module provides semantic parsing support data to the medical-specific plagiarism detection module and graded filtering support data to the discipline-adaptive graded recommendation module through fixed data interfaces. When the support data is updated, the parameters of the associated modules are automatically synchronized and adapted.
[0026] The subject-adapted hierarchical recommendation module receives the similarity detection results from the medical-specific plagiarism detection module and the user's research direction tags transmitted by the scientific research scenario interaction module. Combining the four-level medical subject system and research hotspot map of the dynamic subject resource module, it completes multi-dimensional hierarchical screening of literature through the system's built-in medical scientific research scenario adaptation algorithm, generates a hierarchical recommended literature list, and pushes the hierarchical recommended literature list to the scientific research scenario interaction module through the internal data interface.
[0027] The research scenario interaction module receives user operation commands, transmits parameter setting commands to the medical-specific plagiarism detection module and the subject-adaptive hierarchical recommendation module, transforms the result display requirements into a visual data presentation format, and transmits user research direction tags and interaction behavior data to the subject-adaptive hierarchical recommendation module and the dynamic subject resource module, respectively. After being parsed by the research scenario interaction module, the user operation commands are transmitted to the corresponding functional modules, and the corresponding functional modules feed back the processing results to the research scenario interaction module for visual presentation.
[0028] As an optional implementation, when the medical-specific plagiarism detection module performs multi-dimensional similarity calculations, it sequentially completes the following operations: The system performs medical-specific word segmentation on structured literature data, extracting core research elements such as medical terminology, clinical trial design elements, data indicator systems, ethical statements, and conclusion derivation logic. The word segmentation rules are constructed based on the terminology association relationships in the medical ontology knowledge base of the dynamic discipline resource module and are built into the medical-specific plagiarism detection module.
[0029] A semantic model based on an attention mechanism and integrated with a medical ontology knowledge base is adopted. This model is trained and generated by the system’s built-in medical annotation corpus, which covers Chinese and English literature in the fields of basic medicine, clinical medicine, preventive medicine, pharmacy, and public health. The model takes the core research elements of structured literature data as input, captures the domain-specific semantic associations, clinical trial design similarity features, and data result consistency features of the core research elements through attention weight allocation, and outputs the semantic association feature values of each core research element.
[0030] A composite similarity calculation model is constructed. The composite similarity calculation model has built-in weighted calculation logic for semantic similarity of medical terms, matching degree of clinical trial design, consistency of data indicators, and similarity of conclusion expression. The weight matrix for weighted calculation is a subject-specific weight matrix. This matrix is generated by training the annotated corpus of the corresponding medical field through multi-dimensional similarity contribution statistics and stored in the subject-specific feature library of the dynamic subject resource module. The composite similarity calculation model automatically retrieves the corresponding weight matrix according to the subject classification label of the input document and obtains the comprehensive similarity value through weighted calculation.
[0031] The composite similarity calculation model compares the comprehensive similarity threshold with the subject-specific threshold. The subject-specific threshold is generated by quantile statistics from the plagiarism detection practice data of each medical branch and stored in the subject-specific feature library of the dynamic subject resource module. It locates duplicate content according to three dimensions: full text, chapter, and research elements, and associates the evidence level of the publication journal, subject branch classification, and data source reliability information of similar literature. It then performs differentiated annotation, and the format of the differentiated annotation is adapted to the visualization presentation of the research scenario interaction module.
[0032] As an optional implementation, when the subject-adaptive hierarchical recommendation module performs hierarchical filtering, it completes the following operations in sequence:
[0033] A medical-specific grading dimension is set up, which includes the degree of fit between discipline branches, the relevance of research hotspots, the level of evidence-based medicine, the suitability of clinical trial design, and the matching degree of the user's research stage. Each grading dimension is associated with a fixed data mapping between the four-level medical discipline system and the research hotspot map of the dynamic discipline resource module.
[0034] The four-level medical discipline system based on dynamic discipline resource modules calculates the degree of fit between the research direction of input literature and medical sub-disciplines, disease subtypes, and drug classifications through discipline classification coding mapping. The coding mapping rules are constructed based on the classification coding rules formulated by the system for the four-level medical discipline system.
[0035] By combining the research hotspot map of the dynamic subject resource module, the correlation between the input literature topic and the research hotspot of the subject field is generated through semantic matching of topic keywords, correlation analysis of research methods, and similarity comparison of core conclusions. The judgment rules of each analysis method are built into the subject adaptation hierarchical recommendation module.
[0036] The system has built-in evidence-based medicine evidence grading rules, classifies the literature to be recommended according to the level of evidence, integrates the citation frequency of the literature, the journal impact factor, and the peer review opinions to construct a quality assessment index system. The weight allocation of the quality assessment index system is generated by training the evidence-based medicine related annotation corpus and stored in the subject-specific feature library of the dynamic subject resource module.
[0037] The subject-adaptive hierarchical recommendation module has a built-in research stage hierarchical weight matrix. This matrix is generated by training a corpus of subject-adaptive annotations for different research stages and stored in the subject-specific feature library of the dynamic subject resource module. The subject-adaptive hierarchical recommendation module retrieves the corresponding matrix based on the user's research stage tags, adjusts the weight ratio of each hierarchical dimension, sorts the literature according to the weight ratio, and generates a hierarchical recommended literature list that includes recommendation reasons and subject-adaptive explanations.
[0038] As an optional embodiment, when the medical data preprocessing module performs scientific research element extraction, it sequentially completes the following operations:
[0039] The system performs format standardization operations on the input literature, unifying text encoding, figure and table labeling rules, reference formats, and clinical trial data presentation standards. These standards are built into the system and are adapted to the research content presentation needs of medical literature.
[0040] Based on the system's built-in general medical thesaurus, redundant information, format marks, and meaningless characters in the input documents are removed. Administrative statements and author biographies that are irrelevant to the research content are eliminated through regular expression matching. The regular expression matching rules are constructed based on the text structure and expression features of administrative statements and author biographies and are built into the medical data preprocessing module. The administrative statement matching rule identifies text segments containing funding information and acknowledgments, while the author biographies matching rule identifies text segments containing institutional information and contact information.
[0041] The system calls the built-in medical terminology standard library to perform normalization processing on disease names, generic drug names, and diagnostic and treatment technical terms in the literature. The term matching adopts a combination of exact matching and semantic association matching, giving priority to exact matching. When there is no exact matching result, semantic association matching is performed to ensure semantic consistency of terms.
[0042] A named entity recognition and relation extraction model is adopted. This model is trained and optimized by the system's built-in medical annotation corpus. The model takes standardized literature data as input and extracts research elements such as research type, experimental design type, sample size, intervention measures, outcome indicators, statistical methods, and key conclusions. The extraction logic is constructed based on the description position of each research element in medical literature and the guiding features of related keywords. After extraction, a structured data dictionary is formed, and the data dictionary format is consistent with the storage format of the structured literature data.
[0043] As an optional implementation, when the dynamic subject resource module performs updates and adaptations, it completes the following operations in sequence:
[0044] Each field's dedicated research feature database is constructed according to the discipline branches of the four-level medical discipline system. The feature data includes commonly used research methods, clinical trial design specifications, core indicator systems, ethical review points, and high-frequency related terms in each field. The feature data is extracted through systematic analysis of highly cited literature, authoritative guidelines, and standard operating procedures in each discipline. The feature data format is also adapted to the semantic parsing requirements of the medical-specific plagiarism detection module and the hierarchical screening requirements of the discipline-adapted hierarchical recommendation module.
[0045] The dynamic subject resource module connects to medical literature publishing platforms, research project publishing platforms, and academic conference publication platforms. Through the system's built-in data crawling technology, it captures the latest published results, project proposal directions, and guideline updates in real time. When new literature is added to a single platform and subject, forming a topic cluster, it triggers an update of the research hotspot map. The rules for determining topic clustering are built into the dynamic subject resource module. The updated research hotspot map is synchronized to the subject adaptation and hierarchical recommendation module through a triggering mechanism. The updated data of the research hotspot map and the hierarchical screening process of the subject adaptation and hierarchical recommendation module achieve real-time data adaptation.
[0046] As an optional embodiment, the subject-adaptive hierarchical recommendation module is equipped with a two-way optimization mechanism, which performs the following operations in sequence when the mechanism is executed:
[0047] The system receives user literature viewing time, download behavior, marking operations, plagiarism check parameter settings, and recommendation feedback evaluation data transmitted from the scientific research scenario interaction module. Based on this data, a user scientific research behavior profile is constructed, which includes the subject classification, research methods, and evidence-level feature tags of the user's frequently interacted literature.
[0048] Based on user research behavior profiles, the system clusters literature topics and keywords that match the characteristics of in-depth literature reading by user viewing time, statistically analyzes the research methods of downloaded literature, sets plagiarism detection parameters, and matches corresponding research scenario data to uncover users' core research directions, preferred research methods, and research stages. The execution rules of each discovery method are built into the subject-adaptive hierarchical recommendation module, which updates users' exclusive research tags and interest keyword library. The format of the tag and keyword library is consistent with the format of the user's research direction tags.
[0049] The system receives data on the topics of duplicate content and the subject distribution of similar literature from the medical-specific plagiarism detection module. This data is then fed back into the algorithm model of the hierarchical recommendation strategy. The model automatically adjusts the weight matrix of the hierarchical dimension, and the adjusted weight matrix is stored in the subject-specific feature library of the dynamic subject resource module.
[0050] The system receives user interaction data from the scientific research scenario interaction module, including clicks, downloads, favorites, and feedback evaluations of recommended documents. It then uses this data to iteratively optimize the subject-specific scientific research scenario adaptation algorithm. The algorithm optimization is achieved by analyzing the correlation between the interaction behavior data and the feature data of the recommended documents. The optimized algorithm parameters are synchronized to the hierarchical screening process and stored in the algorithm parameter library of the subject-specific hierarchical recommendation module.
[0051] As an optional embodiment, when the scientific research scenario interaction module implements its functions, it performs the following operations in sequence:
[0052] The interface for setting up custom plagiarism detection parameters includes options for adjusting the weight of subject branches, selecting the granularity of duplicate annotation, and limiting the range of evidence levels for the compared literature. Adjustable parameters include the weight coefficients of each similarity dimension, annotation granularity options, and the range of evidence levels for the compared literature. The parameter adjustment logic establishes a fixed data association with the similarity calculation process of the medical-specific plagiarism detection module.
[0053] The system employs a hierarchical layout based on information priority to visualize the plagiarism detection results. The displayed content includes a comprehensive similarity score, three levels of duplicate content annotation, comparison of evidence levels of similar literature, and a correlation graph of duplicate research elements. Core duplicate content and high-evidence-level similar literature are presented first in the display layout. The color coding adopts the system's built-in medical literature visualization color scheme, and the color saturation corresponding to the duplicate features is dynamically adjusted in a positive correlation with the degree of duplication.
[0054] The interface for filtering recommended literature is set up in multiple dimensions. The interface allows users to filter recommended literature by combination of conditions such as evidence level, subject branch, experimental design type, publication time, and sample size. The filtering logic is implemented by precise matching and range matching of the corresponding fields of the conditions. When filtering with multiple conditions, the logic AND operation is performed. The filtering conditions correspond one-to-one with the hierarchical dimensions of the subject-adaptive hierarchical recommendation module.
[0055] It adopts a distributed database architecture to store users' entire research process data. The stored content includes plagiarism check history, modification trajectory of duplicate content, interaction records of recommended literature, and parameter setting schemes. Data indexes are established according to research project numbers to realize the classification, backtracking, and export of users' entire research process data. The data storage format is compatible with both structured literature data and hierarchical recommended literature lists.
[0056] As an optional implementation, when the medical-specific plagiarism detection module performs differential annotation of duplicate content, it completes the following operations in sequence:
[0057] Full-text annotation distinguishes similarity levels through color gradients. The correspondence between color gradients and similarity levels is set according to the medical literature visual recognition rules constructed by the system. It also displays the subject classification, evidence level, and data source reliability score of similar literature, and adapts the display content format to the visualization layout of the scientific research scenario interaction module.
[0058] Chapter-level annotations add subject-adaptation identifiers next to the titles of repeated chapters. The annotation information includes the core repeated elements of the chapter, and the identifier format is consistent with the display specifications of the interactive interface of the scientific research scenario interaction module.
[0059] The research element-level annotation targets duplicate clinical trial designs, core data indicators, and conclusion derivation logic. It displays detailed comparisons and differences with similar literature through pop-up windows. The comparison logic is constructed based on the correlation characteristics of research elements and is built into a dedicated medical plagiarism detection module. The content displayed in the pop-up windows corresponds to the filtering function condition fields of the research scenario interaction module.
[0060] All annotations have associated viewing entry points, allowing users to view full-text links to similar documents, reference lists, and related recommended documents; an annotation objection submission entry point is also provided. After a user submits an objection, a re-detection is triggered. The re-detection performs the entire process of multi-dimensional similarity calculation using the medical-specific plagiarism detection module, and the results of the re-detection are synchronously updated to the visualization display interface of the scientific research scenario interaction module.
[0061] As an optional implementation, the system sets up a subject-specific iteration mechanism, which performs the following operations sequentially when the mechanism is executed:
[0062] Collect user feedback on plagiarism check results and data on the suitability of recommended literature transmitted from the scientific research scenario interaction module. Combine this with industry data on updates to medical terminology, revisions to clinical trial guidelines, and upgrades to guideline versions to optimize the semantic parsing model and similarity calculation parameters of the medical-specific plagiarism check module. The optimization direction is determined based on the correlation analysis results between feedback issues and updated industry data. The optimized model and parameters are synchronously stored in the database of the corresponding module.
[0063] The dynamic subject resource module updates the field-specific scientific research feature databases, supplements the feature data of newly added subject branches, the judgment criteria of new research methods, and the exclusive indicators of rare disease research. The supplementary data is extracted by analyzing the dynamic literature of subject development. The supplementary feature data is synchronously adapted to the similarity calculation process of the medical-specific plagiarism detection module and the hierarchical screening process of the subject-adapted hierarchical recommendation module.
[0064] Based on data analysis results of changes in medical research trends, the rules for extracting research elements in the medical data preprocessing module, the weights of the hierarchical dimensions in the subject-matching hierarchical recommendation module, and the visualization and interaction settings in the research scenario interaction module were adjusted. The adjusted rules, weights, and settings were synchronized to the corresponding modules, and the system was put into operation after the adjustment parameters of each module were synchronized.
[0065] As an optional implementation, when the scientific research scenario interaction module adapts to the medical scientific research scenario interface, it performs the following operations in sequence:
[0066] It provides standardized interfaces for connecting to medical literature platforms, clinical trial registration platforms, and scientific research management platforms, supporting direct retrieval and duplication comparison of literature data and experimental design schemes. The interface data format is consistent with the structured literature data format generated by the medical data preprocessing module, and the interface communication rules are consistent with the system's built-in medical data interaction communication rules.
[0067] It supports integration with medical paper writing tools and evidence-based medicine analysis tools, enabling embedded applications for plagiarism detection and annotation, and recommended literature citation. It integrates data transmission format to match the detection results of the medical-specific plagiarism detection module, and subject-specific hierarchical recommendation literature list format to match the hierarchical recommendation module. It also integrates interactive logic that is consistent with the user operation process of the scientific research scenario interaction module.
[0068] It offers multiple export formats tailored to different research scenarios. The exported content includes plagiarism reports, recommended literature lists, and comparison maps of duplicate elements. All exported content undergoes standardized processing, and the export formats are the commonly used medical research document formats built into the system. The fields of the exported content correspond to the visualization display content and filtering function condition fields of the research scenario interaction module. Each export format is adapted to the data needs of research scenarios such as paper publication, project application, and review writing.
[0069] 1. This system includes a medical-specific plagiarism detection module, a subject-matched hierarchical recommendation module, a research scenario interaction module, a medical data preprocessing module, and a dynamic subject resource module. Each module adopts a distributed deployment model, achieving cross-module data interaction and parameter synchronization through internal standardized data interfaces. The interface communication rules are based on the system's built-in medical data-specific interaction protocol, which supports data transmission verification, abnormal retry, and breakpoint resumption. Data transmission uniformly adopts a preset structured format, with fields including module identifier, data type, timestamp, content body, and checksum. During data transmission, field verification and consistency checks are performed; if verification fails, data retransmission is triggered. All processed data is stored in a distributed database, partitioned according to module function, using a master-slave replication architecture. The master database stores real-time data, while the slave database is used for data backup and query distribution.
[0070] The overall collaborative operation process of the system is as follows:
[0071] The scientific research scenario interaction module receives medical literature and operation instructions input by the user, transmits the original literature data to the medical data preprocessing module, and at the same time parses the user operation instructions, extracts parameter setting information and research direction tags, and temporarily stores them in the module's local cache. The cached data is retained until the end of the current operation process.
[0072] The medical data preprocessing module performs format standardization, terminology normalization, research element extraction, and noise filtering on the original documents to generate structured document data. After passing the field integrity verification, the data is transmitted to the medical-specific plagiarism detection module through the internal data interface, and the structured data is simultaneously stored in the corresponding partition of the distributed database.
[0073] The medical-specific plagiarism detection module calls upon the supporting data from the dynamic subject resource module to perform multi-dimensional similarity calculations and duplicate content location, generate detection results, and synchronously push them to the subject-adaptive hierarchical recommendation module and the scientific research scenario interaction module through internal data interfaces.
[0074] The dynamic subject resource module updates the subject system, feature library, and research hotspot map in real time. After the data is updated, it supports the synchronous adaptation of parameters between the medical-specific plagiarism detection module and the subject adaptation and hierarchical recommendation module through the interface callback mechanism. After the adaptation is completed, the adaptation results are fed back to the dynamic subject resource module.
[0075] The subject-adaptive hierarchical recommendation module combines similarity detection results, user research direction tags, and dynamic resource data to execute a hierarchical screening algorithm, generate a hierarchical recommended literature list, and push it to the scientific research scenario interaction module through an internal data interface.
[0076] The scientific research scenario interaction module transforms the detection results and recommendation list into a visual presentation, while recording user interaction data. This data is then fed back to the subject-adaptive hierarchical recommendation module and the dynamic subject resource module through an internal data interface for algorithm optimization and resource updates.
[0077] 2. Implementation of Scientific Research Element Extraction in Medical Data Preprocessing Module
[0078] The core of the medical data preprocessing module is to convert unstructured medical literature into standardized structured data. The specific implementation steps are as follows:
[0079] 2.1 Standardized Formatting: The system adopts its built-in medical literature format specifications, uniformly encoding text to UTF-8. Figure / table labels uniformly use the format of figure / table number, subject classification, and core content. References uniformly use the author's name, journal name, publication year, volume (issue): page number format. Clinical trial data is uniformly presented in the field order of trial subjects, intervention measures, observation indicators, and statistical results, achieving uniformity in formatting across different sources and adapting to subsequent semantic analysis and similarity calculation.
[0080] 2.2 Redundant Information and Noise Filtering: Based on the system's built-in medical stop word database, which includes non-core medical research terms, common function words, and meaningless interjections, the database is expanded synchronously with updates to medical terminology. Redundant information and formatting tags are removed through character matching. Regular expression matching is used to remove administrative statements and author biographies. The matching rules for administrative statements are constructed based on the funding, support, and acknowledgments statements guided by "This Research / Project," accurately identifying text segments containing funding information and acknowledgments. The matching rules for author biographies are constructed based on the affiliation, postal code, email address, and telephone number information guided by "Author / Corresponding Author," identifying text segments containing affiliation information and contact details.
[0081] 2.3 Terminology Normalization: The system's built-in medical terminology standard library is invoked. This library covers core medical terms such as diseases, drugs, and diagnostic and treatment techniques. Each term is associated with a unique code, semantic explanation, and a set of synonyms. Term matching employs a mechanism of precise matching followed by semantic association matching and completion. First, precise character matching is performed between the document terminology and the standard library terminology. If a match is successful, the terminology is replaced with the standard term. If no precise match result is found, synonyms are associated through semantic similarity calculation. The semantic similarity calculation uses a cosine similarity algorithm. If a match is successful, normalization is completed, ensuring semantic consistency of the terms.
[0082] 2.4 Research Element Extraction and Structured Transformation: A named entity recognition and relation extraction model trained and optimized using the system's built-in medical annotated corpus was employed. The model was trained using cross-validation, iterating until the loss function converged, and the training results were stored in the module model library. The model uses standardized literature data as input and extracts core elements based on the description position and associated keywords of each research element. Research type is identified using keywords such as retrospective study, prospective study, and randomized controlled trial; trial design type is identified using keywords such as single-blind, double-blind, and cohort study; sample size, intervention measures, outcome indicators, statistical methods, and key conclusions are extracted using specific keywords and contextual semantic analysis. After extraction, a structured data dictionary is formed. The data dictionary fields are completely consistent with the structured literature data storage format, ensuring that the data can be directly used for subsequent similarity calculations.
[0083] 3. Construction, updating, and adaptation of dynamic subject resource modules
[0084] The dynamic subject resource module provides core resource support for the system, including a four-level medical subject system, a dedicated research feature database, and a research hotspot map. The specific implementation steps are as follows:
[0085] 3.1 Construction of a Four-Tier Medical Discipline System: Based on core disciplines in the medical field, a four-tier system is constructed, comprising first-level disciplines, second-level branches, disease subtypes, and research directions. This system covers basic medicine, clinical medicine, preventive medicine, pharmacy, and public health, extending to sub-fields within each discipline. The hierarchical relationships are constructed according to the inherent logic of each discipline. The system establishes a unique classification coding rule for this system. The code consists of 8 characters: the first two characters represent the first-level discipline, the middle two characters represent the second-level branch, and the last four characters correspond to the disease subtype and research direction, respectively. The code corresponds one-to-one with the discipline name, ensuring the uniqueness and identifiability of the discipline classification. After the discipline system is constructed, it is stored in the resource library of the dynamic discipline resource module, supporting the expansion of codes for new discipline branches and the adjustment of existing branches. After expansion and adjustment, the system is synchronously updated to the associated modules.
[0086] 3.2 Construction of a Dedicated Scientific Research Feature Database: A feature database is constructed according to the four-tiered medical discipline system, with each branch corresponding to an independent feature dataset. This dataset includes commonly used research methods, clinical trial design guidelines, core indicator systems, ethical review points, and frequently used related terms. Feature data is extracted through systematic analysis of highly cited literature, authoritative guidelines, and standard operating procedures across various disciplines. The extraction process employs keyword clustering and semantic association analysis. Keyword clustering determines the clustering dimension based on the characteristics of each discipline branch, while semantic association analysis establishes the inherent connections between feature data, ensuring the relevance and comprehensiveness of the feature data. The feature data format is adapted to both the semantic parsing requirements of the medical-specific plagiarism detection module and the hierarchical filtering requirements of the discipline-adaptive hierarchical recommendation module. Each feature data entry is associated with a corresponding discipline code, supporting rapid retrieval by code.
[0087] 3.3 Research Hotspot Map Construction and Real-time Updates: The system integrates with medical literature publishing platforms, research project publishing platforms, and academic conference publication platforms via built-in data crawling technology. The crawling frequency is set according to the platform's data update cycle. Crawled content includes the latest published results, project proposals, and updated guidelines. After deduplication, the crawled data is analyzed using a density-based topic clustering algorithm. The topic clustering rules are built into the dynamic subject resource module. When new literature is added to a single platform and subject, forming a topic cluster, the research hotspot map is updated. The update process includes adding new hotspot topic annotations, ranking hotspot popularity, and iterating on old hotspot topics. Hotspot popularity ranking is calculated based on literature publication time and the number of related literatures. The updated research hotspot map is synchronized to the subject-adaptive hierarchical recommendation module via an interface callback mechanism, achieving real-time data adaptation with the hierarchical screening process.
[0088] 4. Implementation of multi-dimensional similarity calculation in a medical-specific plagiarism detection module
[0089] The medical-specific plagiarism detection module performs document plagiarism checks through multi-dimensional similarity calculations. The specific implementation steps are as follows:
[0090] 4.1 Core Research Element Extraction and Word Segmentation: The system receives structured literature data from the medical data preprocessing module and performs medical-specific word segmentation on medical terms, clinical trial design elements, data indicator systems, ethical statements, and conclusion derivation logic within the data. The word segmentation rules are constructed based on terminological relationships within the medical ontology knowledge base of the dynamic subject resource module. Segmentation is prioritized for complete terms, followed by semantic pause segmentation for non-terminological content, avoiding the splitting of core terms. The segmentation results generate a terminology list and a semantic fragment set for subsequent similarity calculations.
[0091] 4.2 Semantic Relationship Feature Extraction: An attention-based semantic model integrating a medical ontology knowledge base is adopted. The model is trained and generated from the system's built-in medical annotated corpus, which covers Chinese and English literature in five core disciplines, and the corpus size continues to expand with the development of disciplines. The model input is a list of segmented terms and a set of semantic fragments. Through attention weight allocation, the semantic capture of core research elements such as clinical trial design and data indicators is strengthened. The attention weight of core elements is higher than that of other elements. The model outputs semantic relationship feature values for each core research element. The feature value ranges from 0 to 1, with higher values indicating stronger semantic relationship.
[0092] 4.3 Comprehensive Similarity Calculation: A comprehensive similarity calculation model is constructed, incorporating weighted calculation logic across four dimensions: semantic similarity of medical terminology, matching degree of clinical trial design, consistency of data indicators, and similarity of conclusion statements. The weight matrix for weighted calculation is a subject-specific matrix, generated by statistical training on the contribution of multi-dimensional similarity from annotated corpora in the corresponding medical field. The weight ratio of each dimension is determined by analyzing the influence of different dimensions on the plagiarism detection results, and the weight matrix is automatically iterated as the corpus is updated. The weight matrix is stored in the subject-specific feature library of the dynamic subject resource module. The model automatically retrieves the corresponding matrix based on the subject classification tags of the input document, substitutes the semantic association feature values of each dimension, and calculates the comprehensive similarity value, which ranges from 0 to 1.
[0093] 4.4 Duplicate Content Location and Differentiated Annotation: The composite similarity calculation model compares a comprehensive similarity threshold with a subject-specific threshold. These thresholds are generated from the upper quartile statistical analysis of plagiarism detection data from various medical branches. Different medical branches generate their own thresholds based on their specific disciplinary characteristics, and these thresholds are stored in the subject-specific feature library of the dynamic subject resource module. Content with a comprehensive similarity value higher than the corresponding threshold is considered duplicate. Duplicate content is located at three levels: full text, chapter, and research element. A full text similarity value higher than the threshold indicates full text duplication; a chapter similarity value higher than the threshold indicates chapter duplication; and a single research element similarity value higher than the threshold indicates element duplication. Simultaneously, the model associates the publication journal evidence level, subject branch classification, and data source reliability information of similar literature, performing differentiated annotation. The annotation format is adapted to the visualization requirements of the research scenario interaction module, ensuring that the annotation information is clear and identifiable.
[0094] 5. Implementation of hierarchical screening and two-way optimization in the subject-adaptive hierarchical recommendation module
[0095] 5.1 Implementation steps for tiered screening:
[0096] The system sets up medical-specific hierarchical dimensions, including the relevance of discipline branches, the relevance of research hotspots, the level of evidence-based medicine, the suitability of clinical trial design, and the matching degree of the user's research stage. Each dimension is associated with a fixed data mapping of the four-level medical discipline system and research hotspot map in the dynamic discipline resource module. The mapping relationship is stored in the module's local configuration file and supports dynamic adjustment.
[0097] The four-level medical discipline system based on dynamic discipline resource modules calculates the fit value through discipline classification coding mapping. The discipline codes of the input documents are compared with the discipline codes of the recommended documents according to the level. If the first-level discipline is consistent, a basic score is obtained. If the second-level branches are consistent, the scores are accumulated. If the disease subtypes are consistent, the scores are accumulated. If the research directions are consistent, the scores are accumulated. The total score ranges from 0 to 1, which is the discipline branch fit value.
[0098] The correlation score of research hotspots is calculated by combining the research hotspot map. The calculation is carried out through three sub-tasks: semantic matching of topic keywords, correlation analysis of research methods, and similarity comparison of core conclusions. Each sub-task has equal weight, and the total score ranges from 0 to 1. The higher the score, the closer the correlation with the hotspot. The judgment rules of each analysis method are built into the module to ensure the consistency of the calculation.
[0099] Literature Evidence Level Classification and Quality Assessment: The system incorporates evidence-based medicine evidence grading rules, classifying recommended literature into Levels I and V, with Level I being the highest. A quality assessment index system is constructed by integrating citation frequency, journal impact factor, and peer review comments. The index weights are generated through training on an evidence-based medicine-related annotated corpus, and the corresponding numerical values are used to calculate the quality assessment score, ranging from 0 to 1.
[0100] Weight Adjustment and Hierarchical Sorting: The module incorporates a hierarchical weight matrix for research stages. The weight proportions for each dimension differ between basic research and clinical trial stages, generated based on the research needs characteristics of each stage. The matrix is stored in the subject-specific feature library of the dynamic subject resource module. The module retrieves the corresponding matrix based on the user's research stage tags, multiplies the values of each dimension by their corresponding weights, and sums the results to obtain a comprehensive literature recommendation score. The scores are then sorted from highest to lowest to generate a hierarchical recommended literature list. The list includes basic literature information, reasons for recommendation, and subject suitability explanations, formatted to meet the display requirements of the research scenario interaction module.
[0101] 5.2 Implementation steps of the two-way optimization mechanism:
[0102] Data Collection and Behavioral Profile Construction: This involves receiving user data from the research scenario interaction module, including literature viewing time, download behavior, tagging operations, plagiarism check parameter settings, and recommendation feedback evaluation data. Based on this data, a user research behavior profile is constructed. By statistically analyzing the subject classification, research methods, and evidence level of frequently interacted literature, feature tags are generated. These tags are sorted by weight, with the weight determined by interaction frequency and depth. Tags whose viewing time meets the deep interaction criteria have their weight increased.
[0103] User Needs Mining and Tag Updates: Based on user research behavior profiles, core needs are mined through topic keyword clustering, research method statistics, and research scenario matching. Clustering algorithms are used to cluster topic keywords in deeply interacting literature to determine the user's core research direction. The distribution of research methods in downloaded literature is statistically analyzed to determine the user's preferred research methods. By matching plagiarism detection parameters with preset research scenarios, the user's research stage is determined. Based on the mining results, user-specific research tags and interest keyword libraries are updated. The tag format is consistent with the user's research direction tags, ensuring the data can be directly used for tiered filtering.
[0104] Reverse weight adjustment: Receive data on the topic of duplicate content and the subject distribution of similar literature from the medical-specific plagiarism detection module, and input this data back into the algorithm model of the hierarchical recommendation strategy. The model automatically reduces the relevance weight of the subject branch corresponding to the duplicate content. The adjustment range is determined based on the proportion of duplicate literature. The adjusted weight matrix is stored in the subject-specific feature library of the dynamic subject resource module.
[0105] Algorithm Iteration and Optimization: The algorithm receives user data on clicks, downloads, favorites, and feedback evaluations of recommended documents. It uses Pearson correlation coefficient to analyze the correlation between interactive behavior and the features of recommended documents. Based on the analysis results, it adjusts the weights of each feature and iteratively optimizes the algorithm to adapt to the subject research scenario. The optimized algorithm parameters are synchronized to the hierarchical screening process and stored in the module's algorithm parameter library to ensure continuous improvement in recommendation accuracy.
[0106] 6. Implementation of Functionality and Interface Adaptation for the Scientific Research Scenario Interaction Module
[0107] 6.1 Core Function Implementation Steps:
[0108] The plagiarism detection parameter customization interface is implemented as follows: The interactive settings interface includes options for adjusting the weight of subject branches, selecting the granularity of duplicate annotation, and limiting the range of evidence levels for compared literature. Adjustable parameters include the weight coefficients for each similarity dimension, annotation granularity options, and the range of evidence levels for compared literature. The parameter adjustment logic establishes a fixed data association with the similarity calculation process of the medical-specific plagiarism detection module. After the user adjusts the parameters, the changes are transmitted in real-time to the medical-specific plagiarism detection module through an internal interface, triggering parameter synchronization updates.
[0109] The plagiarism detection results are visualized using a hierarchical layout prioritizing information. The top of the interface displays the overall similarity score and a summary of the core duplicated content; the middle section shows the three levels of duplicate content annotations and a comparison of the evidence levels of similar literature; and the bottom displays a graph showing the relationship between duplicated research elements. Core duplicated content and high-evidence-level similar literature are presented first. Color coding uses the system's built-in medical literature visualization color scheme, with the degree of duplication dynamically adjusted in a positive correlation with color saturation to ensure users can quickly identify core information.
[0110] Multi-dimensional filtering of recommended literature: An interactive interface is provided, allowing users to filter by evidence level, subject branch, experimental design type, publication date, and sample size. The filtering logic is implemented through exact and range matching of the corresponding fields. Publication date and sample size support range matching, while other conditions support exact matching. When filtering based on multiple conditions, a logical AND operation is performed. The filtering conditions correspond one-to-one with the hierarchical dimensions of the subject-adaptive hierarchical recommendation module, and the filtering results are updated and displayed in real time.
[0111] User Data Storage and Retrospection: A distributed database architecture is used to store user research process data. Data indexes are created according to research project numbers, with each project corresponding to an independent data folder. This folder stores plagiarism check history, modification records of duplicate content, recommended literature interaction records, and parameter settings. Data can be retrieved by project number and time range, enabling categorized retrospection and export. The data storage format is compatible with both structured literature data and hierarchical recommended literature lists, ensuring compatibility of exported data.
[0112] 6.2 Implementation of Interface Adaptation for Medical Research Scenarios:
[0113] External Platform Integration and Adaptation: Standardized interfaces are provided for integration with medical literature platforms, clinical trial registration platforms, and research management platforms. The interface data format is consistent with the structured literature data format generated by the medical data preprocessing module, and the interface communication rules are consistent with the system's built-in medical data interaction communication rules. The interface supports bidirectional data transmission, allowing direct access to literature data and experimental design schemes from external platforms for plagiarism checks and comparisons. It can also send plagiarism check results and recommendation lists back to external platforms, achieving data interoperability.
[0114] External tool integration and adaptation: Supports integration with medical paper writing tools and evidence-based medicine analysis tools, enabling embedded applications for plagiarism detection, annotation, and recommended literature citation. The integrated data transmission format adapts to the detection results of the medical-specific plagiarism detection module and the recommendation list format of the subject-specific hierarchical recommendation module. The integrated interaction logic is consistent with the user operation flow of the research scenario interaction module, maintaining the user's original operation flow after integration without requiring additional adaptation or learning.
[0115] Research-specific export adaptation: Offers three dedicated export formats for paper publication, project application, and review article writing. Exported content includes plagiarism reports, recommended literature lists, and comparison graphs of duplicate elements. All exported content undergoes standardized processing: the paper publication format is optimized for journal submission requirements; the project application format highlights core data and conclusions; and the review article format supplements literature correlation analysis. Exported content fields correspond to visualization display content and filtering function condition fields, ensuring complete exported data and meeting the needs of different research scenarios.
[0116] 7. Implementation of a dedicated iterative mechanism for system disciplines
[0117] The system employs a subject-specific iteration mechanism to ensure it adapts to technological advancements and evolving user needs in the medical field. The specific implementation steps are as follows:
[0118] Feedback Data Collection and Industry Information Integration: This involves collecting user feedback on plagiarism check results and evaluation data on the suitability of recommended literature transmitted through the research scenario interaction module. Data is categorized and statistically analyzed by issue type, including similarity judgment bias, insufficient recommendation accuracy, and annotation errors. Simultaneously, it integrates industry data on medical terminology updates, clinical trial specification revisions, and guideline version upgrades to establish an industry update database, synchronizing with the latest industry standards in real time.
[0119] Model and parameter optimization: Based on the correlation analysis results between feedback issues and updated industry data, the semantic parsing model and similarity calculation parameters of the medical-specific plagiarism detection module were optimized. To address similarity judgment bias issues, the attention weights for semantic association feature extraction were adjusted. In response to industry standard updates, the terminology matching rules and similarity thresholds were optimized. The optimized model and parameters were synchronously stored in the corresponding module's database, overwriting older version data.
[0120] Feature Database Update and Adaptation: The dedicated research feature database for the dynamic subject resource module is updated, supplementing it with feature data for newly added subject branches, criteria for judging new research methods, and dedicated indicators for rare disease research. Supplementary data is extracted from literature analyzing the dynamic development of the subject, verified by experts, and then entered into the feature database. The supplementary feature data is synchronously adapted to the similarity calculation process of the medical-specific plagiarism detection module and the hierarchical screening process of the subject-specific adaptation and recommendation module, ensuring collaborative adaptation across modules.
[0121] System-wide parameter adjustment and synchronization: Based on data analysis results of changes in medical research trends, adjust the research element extraction rules of the medical data preprocessing module, the hierarchical dimension weights of the discipline adaptation and hierarchical recommendation module, and the visualization presentation and interaction settings of the research scenario interaction module.
[0122] After the adjustment is completed, the adjusted rules, weights and settings are pushed to the corresponding modules through the system's unified parameter synchronization mechanism. After the parameters of all modules are synchronized, the system is put into operation to ensure that the system as a whole is adapted to the trend of medical research.
[0123] In this specification, multiple instances may be described in real time as components, operations, or structures of a single instance. Although individual operations of one or more methods are shown and described as separate operations, one or more of the separate operations may be executed simultaneously and do not need to be executed in the order shown. Structures and functions presented as separate components in the example configuration may be implemented as composite structures or components. Similarly, structures and functions presented as single components may be implemented as separate components. These and other variations, modifications, additions, and improvements fall within the scope of this document.
[0124] While an overview of the subject matter has been described with reference to specific example embodiments, various modifications and changes can be made to these embodiments without departing from the broader scope of embodiments of this disclosure. Such embodiments of the subject matter are referred to herein, individually or collectively, by the term "invention," and are used for convenience only and are not intended to limit the scope of this application to any single disclosure or concept, should more than one disclosure or concept be disclosed in fact.
[0125] The embodiments described herein have been described in sufficient detail to enable those skilled in the art to practice the disclosed teachings. Other embodiments may be used and derived therefrom, such that structural and logical substitutions and changes may be made without departing from the scope of this disclosure. Therefore, the detailed description should not be construed as limiting, and the scope of the various embodiments is defined only by the appended claims and the full scope of their equivalents.
Claims
1. A medical literature intelligent plagiarism detection and hierarchical recommendation interactive system based on subject services, characterized in that: It includes a medical-specific plagiarism detection module, a subject-adaptive hierarchical recommendation module, a research scenario interaction module, a medical data preprocessing module, and a dynamic subject resource module; Each module achieves data interaction and parameter synchronization through internal standardized data interfaces. The data interfaces adopt the system's built-in medical data interaction and communication rules, and the data transmission format is the system's preset structured format. The data processed by each module is stored in a distributed database. The medical data preprocessing module receives medical literature input by users and, relying on the system's built-in medical terminology standard library and clinical trial data specifications, performs format standardization, professional terminology normalization, scientific research element extraction and noise filtering operations to generate structured literature data containing research topics, experimental designs, data indicators and conclusion statements. After the preprocessing operation is completed, the data is transmitted to the medical-specific plagiarism detection module through the internal data interface. The medical-specific plagiarism detection module receives structured literature data, calls the medical ontology knowledge base and subject-specific feature base of the dynamic subject resource module, integrates medical semantic association analysis and multi-dimensional similarity calculation, completes similarity detection at the level of full text of literature and scientific research elements, locates duplicate content and associates the subject classification and evidence level of similar literature, and pushes the detection results to the subject adaptation hierarchical recommendation module and scientific research scenario interaction module through internal data interface. The dynamic discipline resource module constructs a four-level medical discipline system, which is divided into second-level branches based on first-level disciplines. The second-level branches are further subdivided according to disease subtypes. Each disease subtype corresponds to a specific research direction. The discipline system covers basic medicine, clinical medicine, preventive medicine, pharmacy, public health, and the sub-directions under each field. The system formulates exclusive classification coding rules for the four-level medical discipline system. The dynamic subject resource module simultaneously constructs a dedicated scientific research feature library and research hotspot map for each field. The research hotspot map is updated in real time with the latest data from the subject field. The dynamic subject resource module provides semantic parsing support data to the medical-specific plagiarism detection module and graded filtering support data to the subject adaptation and graded recommendation module through a fixed data interface. When the support data is updated, the parameters of the related modules are automatically synchronized and adapted. The subject-adapted hierarchical recommendation module receives the similarity detection results from the medical-specific plagiarism detection module and the user's research direction tags transmitted by the scientific research scenario interaction module. Combining the four-level medical subject system and research hotspot map of the dynamic subject resource module, it completes multi-dimensional hierarchical screening of literature through the system's built-in medical scientific research scenario adaptation algorithm, generates a hierarchical recommended literature list, and pushes the hierarchical recommended literature list to the scientific research scenario interaction module through the internal data interface. The scientific research scenario interaction module receives user operation instructions, transmits parameter setting instructions to the medical-specific plagiarism detection module and the subject-adaptive hierarchical recommendation module, transforms the result display requirements into a visual data presentation format, and transmits user research direction tags and interaction behavior data to the subject-adaptive hierarchical recommendation module and the dynamic subject resource module, respectively.
2. The intelligent plagiarism detection and hierarchical recommendation interactive system for medical literature based on subject services as described in claim 1, characterized in that, When performing multi-dimensional similarity calculations, the medical-specific plagiarism detection module completes the following operations in sequence: The system performs medical-specific word segmentation on structured literature data, extracting core research elements such as medical terminology, clinical trial design elements, data indicator systems, ethical statements, and conclusion derivation logic. The word segmentation rules are constructed based on the terminology association relationships in the medical ontology knowledge base of the dynamic discipline resource module and are built into the medical-specific plagiarism detection module. A semantic model based on an attention mechanism and integrated with a medical ontology knowledge base is adopted. This model is trained and generated by the system’s built-in medical annotation corpus, which covers Chinese and English literature in the fields of basic medicine, clinical medicine, preventive medicine, pharmacy, and public health. The model takes the core research elements of structured literature data as input, captures the domain-specific semantic associations, clinical trial design similarity features, and data result consistency features of the core research elements through attention weight allocation, and outputs the semantic association feature values of each core research element. A composite similarity calculation model is constructed. The composite similarity calculation model has built-in weighted calculation logic for semantic similarity of medical terms, matching degree of clinical trial design, consistency of data indicators, and similarity of conclusion expression. The weight matrix for weighted calculation is a subject-specific weight matrix. This matrix is generated by training the annotated corpus of the corresponding medical field through multi-dimensional similarity contribution statistics and stored in the subject-specific feature library of the dynamic subject resource module. The composite similarity calculation model automatically retrieves the corresponding weight matrix according to the subject classification label of the input document and obtains the comprehensive similarity value through weighted calculation. The composite similarity calculation model compares the comprehensive similarity threshold with the subject-specific threshold. The subject-specific threshold is generated by quantile statistics from the plagiarism detection practice data of each medical branch and stored in the subject-specific feature library of the dynamic subject resource module. It locates duplicate content according to three dimensions: full text, chapter, and research elements, and associates the evidence level of the publication journal, subject branch classification, and data source reliability information of similar literature. It then performs differentiated annotation, and the format of the differentiated annotation is adapted to the visualization presentation of the research scenario interaction module.
3. The intelligent plagiarism detection and hierarchical recommendation interactive system for medical literature based on subject services as described in claim 1, characterized in that, When the subject-matching hierarchical recommendation module performs hierarchical filtering, it completes the following operations in sequence: A medical-specific grading dimension is set up, which includes the degree of fit between discipline branches, the relevance of research hotspots, the level of evidence-based medicine, the suitability of clinical trial design, and the matching degree of the user's research stage. Each grading dimension is associated with a fixed data mapping between the four-level medical discipline system and the research hotspot map of the dynamic discipline resource module. The four-level medical discipline system based on dynamic discipline resource modules calculates the degree of fit between the research direction of input literature and medical sub-disciplines, disease subtypes, and drug classifications through discipline classification coding mapping. The coding mapping rules are constructed based on the classification coding rules formulated by the system for the four-level medical discipline system. By combining the research hotspot map of the dynamic subject resource module, the correlation between the input literature topic and the research hotspot of the subject field is generated through semantic matching of topic keywords, correlation analysis of research methods, and similarity comparison of core conclusions. The judgment rules of each analysis method are built into the subject adaptation hierarchical recommendation module. The system has built-in evidence-based medicine evidence grading rules, classifies the literature to be recommended according to the level of evidence, integrates the citation frequency of the literature, the journal impact factor, and the peer review opinions to construct a quality assessment index system. The weight allocation of the quality assessment index system is generated by training the evidence-based medicine related annotation corpus and stored in the subject-specific feature library of the dynamic subject resource module. The subject-adaptive hierarchical recommendation module has a built-in research stage hierarchical weight matrix. This matrix is generated by training a corpus of subject-adaptive annotations for different research stages and stored in the subject-specific feature library of the dynamic subject resource module. The subject-adaptive hierarchical recommendation module retrieves the corresponding matrix based on the user's research stage tags, adjusts the weight ratio of each hierarchical dimension, sorts the literature according to the weight ratio, and generates a hierarchical recommended literature list that includes recommendation reasons and subject-adaptive explanations.
4. The intelligent plagiarism detection and hierarchical recommendation interactive system for medical literature based on subject services as described in claim 1, characterized in that, medical... When the data preprocessing module performs scientific research element extraction, it completes the following operations in sequence: The system performs format standardization operations on the input literature, unifying text encoding, figure and table labeling rules, reference formats, and clinical trial data presentation standards. These standards are built into the system and are adapted to the research content presentation needs of medical literature. Based on the system's built-in general medical thesaurus, redundant information, format marks, and meaningless characters in the input documents are removed. Administrative statements and author biographies that are irrelevant to the research content are eliminated through regular expression matching. The regular expression matching rules are constructed based on the text structure and expression features of administrative statements and author biographies and are built into the medical data preprocessing module. The administrative statement matching rule identifies text segments containing funding information and acknowledgments, while the author biographies matching rule identifies text segments containing institutional information and contact information. The system calls the built-in medical terminology standard library to perform normalization processing on disease names, generic drug names, and diagnostic and treatment technical terms in the literature. The term matching adopts a combination of exact matching and semantic association matching, giving priority to exact matching. When there is no exact matching result, semantic association matching is performed to ensure semantic consistency of terms. A named entity recognition and relation extraction model is adopted. This model is trained and optimized by the system's built-in medical annotation corpus. The model takes standardized literature data as input and extracts research elements such as research type, experimental design type, sample size, intervention measures, outcome indicators, statistical methods, and key conclusions. The extraction logic is constructed based on the description position of each research element in medical literature and the guiding features of related keywords. After extraction, a structured data dictionary is formed, and the data dictionary format is consistent with the storage format of the structured literature data.
5. The intelligent plagiarism detection and hierarchical recommendation interactive system for medical literature based on subject services as described in claim 1, characterized in that, When the dynamic subject resource module performs updates and adaptations, the following operations are completed in sequence: Each field's dedicated research feature database is constructed according to the discipline branches of the four-level medical discipline system. The feature data includes commonly used research methods, clinical trial design specifications, core indicator systems, ethical review points, and high-frequency related terms in each field. The feature data is extracted through systematic analysis of highly cited literature, authoritative guidelines, and standard operating procedures in each discipline. The feature data format is also adapted to the semantic parsing requirements of the medical-specific plagiarism detection module and the hierarchical screening requirements of the discipline-adapted hierarchical recommendation module. The dynamic subject resource module connects to medical literature publishing platforms, research project publishing platforms, and academic conference publication platforms. Through the system's built-in data crawling technology, it captures the latest published results, project proposal directions, and guideline updates in real time. When new literature is added to a single platform and subject, forming a topic cluster, it triggers an update of the research hotspot map. The rules for determining topic clustering are built into the dynamic subject resource module. The updated research hotspot map is synchronized to the subject adaptation and hierarchical recommendation module through a triggering mechanism. The updated data of the research hotspot map and the hierarchical screening process of the subject adaptation and hierarchical recommendation module achieve real-time data adaptation.
6. The intelligent plagiarism detection and hierarchical recommendation interactive system for medical literature based on subject services as described in claim 3, characterized in that, The subject-adaptive hierarchical recommendation module is equipped with a two-way optimization mechanism. When this mechanism is executed, the following operations are performed sequentially: The system receives user literature viewing time, download behavior, marking operations, plagiarism check parameter settings, and recommendation feedback evaluation data transmitted from the scientific research scenario interaction module. Based on this data, a user scientific research behavior profile is constructed, which includes the subject classification, research methods, and evidence-level feature tags of the user's frequently interacted literature. Based on user research behavior profiles, the system clusters literature topics and keywords that match the characteristics of in-depth literature reading by user viewing time, statistically analyzes the research methods of downloaded literature, sets plagiarism detection parameters, and matches corresponding research scenario data to uncover users' core research directions, preferred research methods, and research stages. The execution rules of each discovery method are built into the subject-adaptive hierarchical recommendation module, which updates users' exclusive research tags and interest keyword library. The format of the tag and keyword library is consistent with the format of the user's research direction tags. The system receives data on the topics of duplicate content and the subject distribution of similar literature from the medical-specific plagiarism detection module. This data is then fed back into the algorithm model of the hierarchical recommendation strategy. The model automatically adjusts the weight matrix of the hierarchical dimension, and the adjusted weight matrix is stored in the subject-specific feature library of the dynamic subject resource module. The system receives user interaction data from the scientific research scenario interaction module, including clicks, downloads, favorites, and feedback evaluations of recommended documents. It then uses this data to iteratively optimize the subject-specific scientific research scenario adaptation algorithm. The algorithm optimization is achieved by analyzing the correlation between the interaction behavior data and the feature data of the recommended documents. The optimized algorithm parameters are synchronized to the hierarchical screening process and stored in the algorithm parameter library of the subject-specific hierarchical recommendation module.
7. The intelligent plagiarism detection and hierarchical recommendation interactive system for medical literature based on subject services as described in claim 1, characterized in that, When implementing the functions of the scientific research scenario interaction module, the following operations are performed in sequence: The interface for setting up custom plagiarism detection parameters includes options for adjusting the weight of subject branches, selecting the granularity of duplicate annotation, and limiting the range of evidence levels for the compared literature. Adjustable parameters include the weight coefficients of each similarity dimension, annotation granularity options, and the range of evidence levels for the compared literature. The parameter adjustment logic establishes a fixed data association with the similarity calculation process of the medical-specific plagiarism detection module. The system employs a hierarchical layout based on information priority to visualize the plagiarism detection results. The displayed content includes a comprehensive similarity score, three levels of duplicate content annotation, comparison of evidence levels of similar literature, and a correlation graph of duplicate research elements. Core duplicate content and high-evidence-level similar literature are presented first in the display layout. The color coding adopts the system's built-in medical literature visualization color scheme, and the color saturation corresponding to the duplicate features is dynamically adjusted in a positive correlation with the degree of duplication. The interface for filtering recommended literature is set up in multiple dimensions. The interface allows users to filter recommended literature by combination of conditions such as evidence level, subject branch, experimental design type, publication time, and sample size. The filtering logic is implemented by precise matching and range matching of the corresponding fields of the conditions. When filtering with multiple conditions, the logic AND operation is performed. The filtering conditions correspond one-to-one with the hierarchical dimensions of the subject-adaptive hierarchical recommendation module. It adopts a distributed database architecture to store users' entire research process data. The stored content includes plagiarism check history, modification trajectory of duplicate content, interaction records of recommended literature, and parameter setting schemes. Data indexes are established according to research project numbers to realize the classification, backtracking, and export of users' entire research process data. The data storage format is compatible with both structured literature data and hierarchical recommended literature lists.
8. The intelligent plagiarism detection and hierarchical recommendation interactive system for medical literature based on subject services as described in claim 2, characterized in that, When the medical-specific plagiarism detection module performs differential annotation of duplicate content, it completes the following operations in sequence: Full-text annotation distinguishes similarity levels through color gradients. The correspondence between color gradients and similarity levels is set according to the medical literature visual recognition rules constructed by the system. It also displays the subject classification, evidence level, and data source reliability score of similar literature, and adapts the display content format to the visualization layout of the scientific research scenario interaction module. Chapter-level annotations add subject-adaptation identifiers next to the titles of repeated chapters. The annotation information includes the core repeated elements of the chapter, and the identifier format is consistent with the display specifications of the interactive interface of the scientific research scenario interaction module. The research element-level annotation targets duplicate clinical trial designs, core data indicators, and conclusion derivation logic. It displays detailed comparisons and differences with similar literature through pop-up windows. The comparison logic is constructed based on the correlation characteristics of research elements and is built into a dedicated medical plagiarism detection module. The content displayed in the pop-up windows corresponds to the filtering function condition fields of the research scenario interaction module. All annotations have associated viewing links, allowing users to view full-text links to similar documents, reference lists, and related recommended documents; An objection submission portal is set up. After a user submits an objection, a re-detection is triggered. The re-detection performs the entire process of multi-dimensional similarity calculation using the medical-specific plagiarism detection module. The results of the re-detection are simultaneously updated to the visualization display interface of the scientific research scenario interaction module.
9. The intelligent plagiarism detection and hierarchical recommendation interactive system for medical literature based on subject services as described in claim 1, characterized in that, The system is configured with a subject-specific iteration mechanism. When this mechanism is executed, the following operations are performed sequentially: Collect user feedback on plagiarism check results and data on the suitability of recommended literature transmitted from the scientific research scenario interaction module. Combine this with industry data on updates to medical terminology, revisions to clinical trial guidelines, and upgrades to guideline versions to optimize the semantic parsing model and similarity calculation parameters of the medical-specific plagiarism check module. The optimization direction is determined based on the correlation analysis results between feedback issues and updated industry data. The optimized model and parameters are synchronously stored in the database of the corresponding module. The dynamic subject resource module updates the field-specific scientific research feature databases, supplements the feature data of newly added subject branches, the judgment criteria of new research methods, and the exclusive indicators of rare disease research. The supplementary data is extracted by analyzing the dynamic literature of subject development. The supplementary feature data is synchronously adapted to the similarity calculation process of the medical-specific plagiarism detection module and the hierarchical screening process of the subject-adapted hierarchical recommendation module. Based on data analysis results of changes in medical research trends, the rules for extracting research elements in the medical data preprocessing module, the weights of the hierarchical dimensions in the subject-matching hierarchical recommendation module, and the visualization and interaction settings in the research scenario interaction module were adjusted. The adjusted rules, weights, and settings were synchronized to the corresponding modules, and the system was put into operation after the adjustment parameters of each module were synchronized.
10. The intelligent plagiarism detection and hierarchical recommendation interactive system for medical literature based on subject services as described in claim 7, characterized in that, When adapting the scientific research scenario interaction module to the interface of medical scientific research scenarios, the following operations are performed in sequence: It provides standardized interfaces for connecting to medical literature platforms, clinical trial registration platforms, and scientific research management platforms, supporting direct retrieval and duplication comparison of literature data and experimental design schemes. The interface data format is consistent with the structured literature data format generated by the medical data preprocessing module, and the interface communication rules are consistent with the system's built-in medical data interaction communication rules. It supports integration with medical paper writing tools and evidence-based medicine analysis tools, enabling embedded applications for plagiarism detection and annotation, and recommended literature citation. It integrates data transmission format to match the detection results of the medical-specific plagiarism detection module, and subject-specific hierarchical recommendation literature list format to match the hierarchical recommendation module. It also integrates interactive logic that is consistent with the user operation process of the scientific research scenario interaction module. It offers multiple export formats tailored to different research scenarios. The exported content includes plagiarism reports, recommended literature lists, and comparison maps of duplicate elements. All exported content undergoes standardized processing, and the export formats are the commonly used medical research document formats built into the system. The fields of the exported content correspond to the visualization display content and filtering function condition fields of the research scenario interaction module. Each export format is adapted to the data needs of research scenarios such as paper publication, project application, and review writing.