Medical literature processing method and device applied to evidence-based medicine
By analyzing the text describing the purpose of medical research to determine keywords, generating search queries, and constructing a search strategy configuration table, this approach solves the problems of non-standard processes and low accuracy in traditional medical literature management. It achieves efficient and accurate literature retrieval and management, and supports systematic reviews and meta-analysis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-18
- Publication Date
- 2026-04-14
AI Technical Summary
Traditional medical literature management methods suffer from subjective human factors, non-standard management processes, low accuracy in literature collection and identification, high learning costs, and low efficiency in literature retrieval.
By analyzing the text describing the purpose of medical research to identify keywords, generating search queries, and constructing a search strategy configuration table based on relevance, the system automatically populates the literature database search items. Combined with log records and multi-level database field lists, it achieves standardized and efficient management of literature retrieval.
It achieves objectivity and efficiency in medical literature retrieval, reduces learning costs, improves the accuracy and efficiency of literature retrieval, ensures accurate reflection of the PRISMA view, and supports systematic reviews and meta-analysis.
Smart Images

Figure CN117112877B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of literature management technology, and in particular to a method and apparatus for processing medical literature applied to evidence-based medicine. Background Technology
[0002] Medical literature management allows for the effective organization, tracking, and citation of various literature resources acquired during the research process. Literature management enables researchers to organize these documents systematically, facilitating retrieval and use; it helps researchers track their research progress and the latest findings in related fields; and it helps researchers automatically generate reference and citation lists, ensuring the accuracy and consistency of citation formats, reducing the time and error rate associated with manually compiling and writing citations.
[0003] Currently, traditional medical literature management methods are subject to human subjectivity, resulting in non-standardized overall management processes and requiring significant learning costs. Furthermore, due to the instability and diversity of literature, the accuracy of literature collection and identification is relatively low. Summary of the Invention
[0004] This disclosure provides a method and apparatus for processing medical literature applicable to evidence-based medicine.
[0005] According to one aspect of this application, a method for processing medical literature applied to evidence-based medicine is provided, comprising:
[0006] Obtain the text describing the purpose of medical research and parse it to identify multiple keywords that characterize the purpose of medical research;
[0007] Obtain the URL address corresponding to the literature database that matches the stated medical research objective, and generate a medical literature source configuration table based on it;
[0008] Extract the semantic features of each keyword and determine the degree of association between different keywords based on the semantic features of different keywords;
[0009] For any given keyword, obtain its relevance to other keywords;
[0010] Other keywords whose relevance exceeds a set relevance threshold are combined with any of the keywords to generate a search query for medical literature retrieval in the literature database;
[0011] All the search terms corresponding to the keywords are added to a pre-built search strategy configuration table. Based on the search strategy configuration table, the medical literature source configuration table is called to access the URL address corresponding to the literature database. The search terms are then automatically populated into the search items of the literature database to retrieve medical literature from the literature database.
[0012] Optionally, the method further includes:
[0013] Generate a search operation log, which includes the number of medical documents retrieved from each literature database using all search terms;
[0014] Obtain the medical literature maintenance list, which records the estimated number of medical literatures containing any keyword;
[0015] For any keyword, if the following conditions are met, the document feature data of the medical documents containing that keyword are obtained and mapped to the first database list to form a first-level database field list:
[0016] Based on its corresponding search formula, the number of medical documents retrieved from a literature database is equal to the estimated number of medical documents containing any of the keywords recorded in the medical document maintenance list.
[0017] Optionally, the method further includes:
[0018] The medical literature maintenance list reading component is invoked to read the estimated quantity from the medical literature maintenance list;
[0019] The retrieval operation log reading component is invoked to read the number of medical documents retrieved from each literature database using all search terms from the retrieval operation log, and compared with the estimated number to determine whether the two are equal.
[0020] Optionally, the method further includes:
[0021] Based on the established document database information maintenance component, each field in the primary database field list is checked to perform the following steps to maintain the document feature status of each medical document, and the primary database field list is marked after the maintenance is completed:
[0022] Determine whether there are any missing document feature data and / or duplicate document feature data; if there are missing document feature data, complete them; if there are duplicate document feature data, remove duplicates, until the document feature status maintenance of all medical documents is completed.
[0023] Optionally, the method further includes: extracting data from the first-level database field list that has been marked to obtain the literature feature data corresponding to the fields therein, and mapping it to a second database list in a manner corresponding to medical literature to form a second-level database field list.
[0024] Optionally, the method further includes: obtaining a screening status marker obtained by screening the fields in the secondary library field list, and maintaining the document feature status of each medical document based on the screening status marker until the document feature status maintenance of all medical documents is completed.
[0025] Optionally, the method further includes: extracting data from the secondary database field list that has completed the status maintenance to obtain the literature feature data corresponding to the fields therein, and mapping it to the third database list in a manner corresponding to medical literature to form a tertiary database field list.
[0026] Optionally, the method further includes:
[0027] Generate a first log that checks each field in the first-level database field list to maintain the document feature status of each medical document.
[0028] A second log is generated to re-screen the fields in the secondary library field list in order to maintain the document feature status of each medical document.
[0029] Optionally, the method further includes:
[0030] Obtain the first log and parse it to obtain the first operation performed for corresponding maintenance and the object pointed to by the first operation;
[0031] Obtain the second log and parse it to obtain the second operation performed for the corresponding rescreening and the object pointed to by the second operation;
[0032] Statistical results are obtained by performing statistics on the first operation and the object pointed to by the first operation, the second operation and the object pointed to by the second operation;
[0033] The statistical results are rendered and displayed on a pre-built PRISMA component to form a PRISMA view.
[0034] According to one aspect of this application, a medical literature processing device for evidence-based medicine is provided, comprising:
[0035] The data parsing unit is used to obtain the text describing the purpose of medical research and to parse it to determine multiple keywords that characterize the purpose of medical research.
[0036] The configuration table generation unit is used to obtain the URL address corresponding to the literature database that matches the medical research purpose, and generate a medical literature source configuration table based on it.
[0037] The data extraction unit is used to extract the semantic features of each keyword and determine the degree of association between different keywords based on their semantic features.
[0038] The relevance determination unit is used to obtain the relevance between any given keyword and other keywords.
[0039] The retrieval formula generation unit is used to take other keywords whose relevance exceeds a set relevance threshold and combine them with any of the keywords to generate a retrieval formula for medical literature retrieval in the literature database;
[0040] The literature retrieval unit is used to add the search terms corresponding to all keywords to a pre-built search strategy configuration table, and call the medical literature source configuration table based on the search strategy configuration table to access the URL address corresponding to the literature database, and automatically populate the search terms into the search items of the literature database to retrieve medical literature in the literature database.
[0041] In this application, the following steps are taken: First, the text describing the purpose of medical research is obtained and parsed to determine multiple keywords representing that purpose. Second, the URLs of literature databases matching the stated purpose are obtained, and a medical literature source configuration table is generated based on these URLs. Third, the semantic features of each keyword are extracted, and the semantic features between different keywords are determined to establish their relevance. Fourth, for any given keyword, its relevance to other keywords is determined. Fifth, keywords with relevance exceeding a set threshold are combined with the given keyword to generate a search query for medical literature retrieval in the literature database. Sixth, all search queries corresponding to the keywords are added to a pre-built search strategy configuration table, and the process is based on... The search strategy configuration table calls the medical literature source configuration table to access the URL address corresponding to the literature database, and automatically fills the search terms into the search items of the literature database to retrieve medical literature. This achieves the objective determination of the direction of literature retrieval based on keywords in the medical research purpose during the medical literature retrieval process. At the same time, it constructs the search terms based on the relevance between keywords, and further constructs the search strategy configuration table interface to call the medical literature source configuration table. This achieves literature retrieval quickly and with a clear purpose, thereby improving the efficiency of medical literature retrieval, standardizing the literature retrieval process, eliminating the influence of human subjectivity, and reducing the learning cost of literature retrieval. Furthermore, throughout the process, the refinement of literature feature data is divided into processing stages such as retrieval, initial screening, and secondary screening. Each stage has a corresponding log to record the process data of that stage. This ensures that all data operations and their corresponding objects in the literature retrieval process, screening process, formation of primary, secondary, and tertiary database field lists are completely clear and error-free. This guarantees that the generated PRISMA view can accurately reflect the results of all operations in the above-mentioned process, realizing the mutual embedding of systematic reviews and meta-analysis with literature management, improving the accuracy and efficiency of literature processing, and providing a reliable data foundation for subsequent clinical research.
[0042] It should be understood that the description in the Summary of the Invention is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0043] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. The drawings are provided for a better understanding of the invention and are not intended to limit the scope of this disclosure. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:
[0044] Figure 1 This is a flowchart illustrating a medical literature processing method applied to evidence-based medicine, as described in an embodiment of this application.
[0045] Figure 2 This is a schematic diagram of a medical literature processing device applied to evidence-based medicine, according to an embodiment of this application. Detailed Implementation
[0046] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.
[0047] Furthermore, the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.
[0048] Figure 1 This is a schematic flowchart illustrating a medical literature processing method applied to evidence-based medicine, as described in an embodiment of this application. Figure 1 As shown, it includes:
[0049] S101. Obtain the text describing the purpose of the medical research and parse it to determine multiple keywords that characterize the purpose of the medical research.
[0050] Optionally, in this embodiment, the description text can be obtained from other databases, such as by retrieving the description text from a predetermined data interface for direct access to the database.
[0051] Optionally, the format and size of the stated text are not limited.
[0052] Optionally, in this embodiment, in order to improve the efficiency of parsing, annotations can be added to statements or words in the text that reflect the purpose of medical research before acquisition, so that the annotations can be read directly in step S101 to quickly determine the keywords.
[0053] Optionally, in this embodiment, after obtaining the keywords, the keywords can be added to the keyword maintenance list to maintain and adjust the keywords, such as performing keyword supplementation operations, thereby improving the accuracy of the keywords.
[0054] S102. Obtain the URL address corresponding to the literature database that matches the medical research objective, and generate a medical literature source configuration table based on it.
[0055] Optionally, in this embodiment, the number of documents in the literature database, the language of expression of the medical documents included, and the format of the medical documents are not limited.
[0056] Furthermore, in order to fully utilize the medical literature source configuration table and improve its effectiveness, in addition to including the URL address corresponding to the literature database, the medical literature source configuration table can also include the name of the literature database and the language used to describe the medical literature in the literature database. In addition, in order to distinguish different literature databases and facilitate subsequent medical literature retrieval, the medical literature source configuration table can also include a unique ID assigned to each literature database, thereby providing multiple options for accessing literature databases and improving the diversity of medical literature retrieval.
[0057] Specifically, for example, a medical literature source configuration table with empty content can be pre-built. This table can then include ID, database name, URL, and language fields. After obtaining the URL address corresponding to the literature database, the name of the literature database, and the language used to describe the medical literature in the database, the data corresponding to this information can be directly assigned to the corresponding fields. Furthermore, a unique ID can be assigned to different literature databases according to ID allocation rules. These ID allocation rules, for example, assign unique IDs based on the importance of the literature database; the higher the importance, the smaller the ID value. This allows users to directly select important databases based on their IDs when using the literature database, thereby improving the focus of medical literature retrieval.
[0058] Optionally, after completing the medical literature source configuration table, it can be stored in a database table to facilitate the maintenance of the medical literature source configuration table, including adding, deleting, modifying, and querying.
[0059] S103. Extract the semantic features of each keyword and determine the degree of association between different keywords.
[0060] Optionally, in this embodiment, the semantic features of each keyword can be extracted using a trained semantic feature analysis model. The specific semantic feature model can be selected based on the application scenario; for example, it can be a neural network model.
[0061] Optionally, the semantic features between different keywords can be determined using an attention model to determine the degree of association between them. Such an attention model could be, for example, a transformer model.
[0062] In this embodiment, in order to improve the accuracy of relevance determination, sample keywords are extracted from historical retrieval samples that have been used in advance. The attention model is trained based on these sample keywords, so that the attention model learns the relevance features between different keywords. This ensures that the accuracy of relevance determination is improved when using the attention model, and avoids identifying different keywords with no relevance as having high relevance.
[0063] In this embodiment, by analyzing the relevance, the likelihood of different keywords being used in combination during retrieval can be determined, so as to generate a search query in the future, combining the most likely keywords to be used together to form a search query.
[0064] S104. For any given keyword, obtain its relevance to other keywords.
[0065] Optionally, in this embodiment, the degree of relevance between any keyword and other keywords can be determined by forming a relevance queue.
[0066] S105. Take other keywords whose relevance exceeds the set relevance threshold and combine them with any of the keywords to generate a search formula for medical literature retrieval in the literature database.
[0067] Optionally, if the relevance in the relevance queue is sorted from largest to smallest along the queue head to tail, then the minimum relevance greater than and above the relevance threshold can be found. The keywords corresponding to the minimum relevance and all relevances in the relevance queue preceding the minimum relevance can be directly taken and combined with any keyword to obtain the corresponding search expression.
[0068] It should be noted that in some embodiments, a search query can also be generated from a single keyword.
[0069] Furthermore, when generating search queries, you can specify the fields that are valid when using keywords for retrieval, such as whether it is an abstract or the full text.
[0070] S106. Add the search terms corresponding to all keywords to the pre-built search strategy configuration table, and call the medical literature source configuration table based on the search strategy configuration table to access the URL address corresponding to the literature database, and automatically populate the search terms into the search items of the literature database to retrieve medical literature in the literature database.
[0071] Optionally, in this embodiment, the search query can be copied to the corresponding field of a pre-built search strategy configuration table. In addition, a keyword combination field is configured in the search strategy configuration table to store the keyword combinations used in the corresponding search query. Furthermore, a search query ID is also configured in the search strategy configuration table to quickly retrieve the corresponding search query for retrieval through the search query ID.
[0072] In this embodiment, the retrieval strategy configuration is implemented in a concise and clear manner through the above-mentioned retrieval strategy configuration table. When conducting retrieval in the future, the retrieval can be triggered directly based on the retrieval strategy configuration table, which reduces the difficulty of algorithm implementation and saves the cost of process interaction.
[0073] Optionally, the retrieval strategy configuration table can be, for example, an Excel file.
[0074] In this embodiment, a well-written control script can be used to make the retrieval strategy configuration table call the medical literature source configuration table. In addition, the priority of the retrieval formula and the retrieval priority of the literature database can be defined in the control script.
[0075] Optionally, the method further includes:
[0076] Generate a search operation log, which includes the number of medical documents retrieved from each literature database using all search terms.
[0077] In this embodiment, the process data of medical literature retrieval is recorded by the retrieval operation log, such as the search terms used, the literature databases searched, the number of literatures retrieved, and the time span of the literatures. This allows for fine-grained recording of the retrieval process, making it easier to directly obtain relevant data from the retrieval operation log when generating the PRISMA diagram later.
[0078] The above-mentioned retrieval operation log can be generated in parallel during document retrieval to ensure the real-time nature of the retrieval operation log.
[0079] Optionally, the method further includes:
[0080] Obtain the medical literature maintenance list, which records the estimated number of medical literatures containing any keyword;
[0081] For any keyword, if the following conditions are met, the document feature data of the medical documents containing that keyword are obtained and mapped to the first database list to form a first-level database field list:
[0082] Based on its corresponding search formula, the number of medical documents retrieved from a literature database is equal to the estimated number of medical documents containing any of the keywords recorded in the medical document maintenance list.
[0083] Optionally, in this embodiment, the step of obtaining the medical literature maintenance list can be performed after step S106, or simultaneously with or before step S106.
[0084] Optionally, in this embodiment, the medical literature maintenance list can be a list maintained for results retrieved manually using a search query. The recorded data can be pre-extracted using a literature analysis component, for example, by using another literature database as a unit. This includes information such as the journal in which the literature was published, the name of the literature, the publication date, the publication address, and relevant paragraphs, etc., and is also categorized and statistically analyzed based on keywords. Therefore, this medical literature maintenance list provides a reference for searchable medical literature, which can be compared with the retrieved medical literature obtained after executing step S106, for example, by directly comparing the quantities, to determine the accuracy of the search in S106, and thus the accuracy of the search strategy in the search strategy configuration table. When the number of medical literature retrieved from a literature database is equal to the estimated number of medical literature containing any keyword recorded in the medical literature maintenance list, it indicates that the data in the medical literature maintenance list is correct. Then, the literature feature data of medical literature containing any keyword is obtained from the medical literature maintenance list and mapped to the first database list to form a first-level database field list.
[0085] Optionally, to quickly achieve the above-mentioned direct quantity-based comparison, the method further includes:
[0086] The medical literature maintenance list reading component is invoked to read the estimated quantity from the medical literature maintenance list;
[0087] The retrieval operation log reading component is invoked to read the number of medical documents retrieved from each literature database using all search terms from the retrieval operation log, and compared with the estimated number to determine whether the two are equal.
[0088] Specifically, the medical literature maintenance list reading component parses the medical literature maintenance list according to the table header during reading, thereby reading the description field of the estimated quantity, and directly obtains the value of this field to read the estimated quantity.
[0089] Similarly, the retrieval operation log reading component parses the retrieval operation log to read the number of medical documents retrieved from each literature database using all search terms. For example, a comparison function is used to compare these numbers to determine if they are equal.
[0090] Optionally, the method further includes:
[0091] Based on the established document database information maintenance component, each field in the primary database field list is checked to perform the following steps to maintain the document feature status of each medical document, and the primary database field list is marked after the maintenance is completed:
[0092] Determine whether there are any missing document feature data and / or duplicate document feature data; if there are missing document feature data, complete them; if there are duplicate document feature data, remove duplicates, until the document feature status maintenance of all medical documents is completed.
[0093] Optionally, in this embodiment, after the above-mentioned document retrieval is completed, the document database information maintenance component can be enabled to complete and deduplicate the document feature data.
[0094] It should be noted here that the characteristic data of the document may include at least one of the following: Key, Itemtype, Publication, Author, Title, ISBN, ISSN, DOI, URL, Abstract, Date, Pages, Issue, and Volume. The specific data can be determined according to the application scenario.
[0095] In this embodiment, the document feature data is improved by completing and deduplicating the document feature data through the above-mentioned document database information maintenance component.
[0096] Optionally, the method further includes: extracting data from the first-level database field list that has been marked to obtain the literature feature data corresponding to the fields therein, and mapping it to a second database list in a manner corresponding to medical literature to form a second-level database field list.
[0097] Optionally, the above-mentioned steps for forming the secondary database field list can be performed after the above-mentioned steps for forming the primary database field list, such as after completing the above-mentioned steps for supplementing and deduplicating the literature feature data.
[0098] Optionally, the method further includes: obtaining a screening status marker obtained by screening the fields in the secondary library field list, and maintaining the document feature status of each medical document based on the screening status marker until the document feature status maintenance of all medical documents is completed.
[0099] Optionally, the above-mentioned document feature status maintenance can be performed after the secondary database field list is formed, so as to further improve the secondary database field list and further improve the accuracy of document feature data.
[0100] Optionally, the method further includes: extracting data from the secondary database field list that has completed the status maintenance to obtain the literature feature data corresponding to the fields therein, and mapping it to the third database list in a manner corresponding to medical literature to form a tertiary database field list.
[0101] In this embodiment, the most complete document feature data is saved by mapping the document feature data in the secondary library field list that has completed the status maintenance to the tertiary library field list, so as to facilitate subsequent use.
[0102] Optionally, the method further includes:
[0103] Generate a first log that checks each field in the first-level database field list to maintain the document feature status of each medical document.
[0104] A second log is generated to re-screen the fields in the secondary library field list in order to maintain the document feature status of each medical document.
[0105] Specifically, the steps for generating the first log can be executed when forming the first-level database field list. Similarly, the steps for generating the second log can be executed when forming the second-level database field list, thus ensuring the real-time nature and accuracy of the data recorded in the logs.
[0106] Optionally, the method further includes:
[0107] Obtain the first log and parse it to obtain the first operation performed for corresponding maintenance and the object pointed to by the first operation;
[0108] Obtain the second log and parse it to obtain the second operation performed for the corresponding rescreening and the object pointed to by the second operation;
[0109] Statistical results are obtained by performing statistics on the first operation and the object pointed to by the first operation, the second operation and the object pointed to by the second operation;
[0110] The statistical results are rendered and displayed on a pre-built PRISMA component to form a PRISMA view.
[0111] The process of obtaining the first log, second log, etc., to form the PRISMA view can be executed after the second-level library field list is formed. This ensures that all data operations and corresponding objects in the literature retrieval process, filtering process, and the formation of the first-level, second-level, and third-level library field lists are completely clear and error-free. This guarantees that the generated PRISMA view can accurately reflect the results of all operations in the above scheme execution process.
[0112] In this embodiment, considering that the PRISMA view (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) can describe and display the methods and results of systematic reviews and meta-analyses, statistics are performed on all data operations in the literature retrieval process, screening process, formation of the first-level library field list, the second-level library field list, and the third-level library field list in the above-mentioned scheme of this application embodiment. These operations include supplementation, deduplication, rescreening, and status marking. This makes the results of all operations in the execution process of the above scheme obvious and traceable.
[0113] In a specific application scenario, the document retrieval process, the filtering process, and the formation of the first-level database field list, the second-level database field list, and the third-level database field list in the above scheme can be executed on the backend server, while the process of obtaining the first log, the second log, etc. to form the PRISMA view can be executed on the frontend device, so that users can directly view the PRISMA view and understand the results of all operations in the execution process of the above scheme.
[0114] Figure 2 This is a schematic diagram of a medical literature processing device applied to evidence-based medicine, provided as an embodiment of this application. Figure 2 As shown, it includes:
[0115] The data parsing unit is used to obtain the text describing the purpose of medical research and to parse it to determine multiple keywords that characterize the purpose of medical research.
[0116] The configuration table generation unit is used to obtain the URL address corresponding to the literature database that matches the medical research purpose, and generate a medical literature source configuration table based on it.
[0117] The data extraction unit is used to extract the semantic features of each keyword and determine the degree of association between different keywords based on their semantic features.
[0118] The relevance determination unit is used to obtain the relevance between any given keyword and other keywords.
[0119] The retrieval formula generation unit is used to take other keywords whose relevance exceeds a set relevance threshold and combine them with any of the keywords to generate a retrieval formula for medical literature retrieval in the literature database;
[0120] The literature retrieval unit is used to add the search terms corresponding to all keywords to a pre-built search strategy configuration table, and call the medical literature source configuration table based on the search strategy configuration table to access the URL address corresponding to the literature database, and automatically populate the search terms into the search items of the literature database to retrieve medical literature in the literature database.
[0121] Optionally, the device further includes:
[0122] The first log unit is used to generate a retrieval operation log, which includes the number of medical documents retrieved from each literature database using all search terms;
[0123] The maintenance list acquisition unit is used to acquire a medical literature maintenance list, which records the estimated number of medical literatures containing any keyword.
[0124] The feature data acquisition unit is used to acquire the feature data of medical documents containing any keyword and map it to a first database list to form a first-level database field list if the following conditions are met:
[0125] Based on its corresponding search formula, the number of medical documents retrieved from a literature database is equal to the estimated number of medical documents containing any of the keywords recorded in the medical document maintenance list.
[0126] Optionally, the device further includes:
[0127] The first component invocation unit is used to invoke the medical literature maintenance list reading component to read the estimated quantity from the medical literature maintenance list;
[0128] The second component calling unit is used to call the retrieval operation log reading component to read the number of medical documents retrieved from each literature database using all search terms from the retrieval operation log, and compare it with the estimated number to determine whether the two are equal.
[0129] Optionally, the apparatus further includes: a first filtering unit, configured to check each field in the primary database field list based on a set literature database information maintenance component to perform the following steps to maintain the literature feature status of each medical document, and after completing the maintenance, mark the primary database field list: determine whether there is missing literature feature data, and / or whether there is duplicate literature feature data; if there is missing literature feature data, perform completion processing; if there is duplicate literature feature data, perform deduplication processing, until the literature feature status maintenance of all medical documents is completed.
[0130] Optionally, the apparatus further includes: a second filtering unit, configured to extract data from the first-level database field list that has been marked, to obtain the literature feature data corresponding to the fields therein, and to map it to a second database list in a manner corresponding to medical literature to form a second-level database field list.
[0131] Optionally, the apparatus further includes: a status marking unit, used to obtain a screening status mark obtained by screening the fields in the secondary library field list, and to maintain the document feature status of each medical document based on the screening status mark until the document feature status maintenance of all medical documents is completed.
[0132] Optionally, the device further includes: a third filtering unit, which extracts data from the secondary database field list that has completed the status maintenance to obtain the literature feature data corresponding to the fields therein, and maps it to the third database list in a manner corresponding to medical literature to form a tertiary database field list.
[0133] Optionally, the device further includes:
[0134] The second log unit is used to generate a first log that checks each field in the primary database field list to maintain the document feature status of each medical document.
[0135] The third log unit is used to generate a second log that re-screens the fields in the secondary library field list to maintain the document feature status of each medical document.
[0136] Optionally, the device further includes:
[0137] The data acquisition unit is configured to: acquire the first log and parse it to obtain the first operation performed for corresponding maintenance and the object pointed to by the first operation; acquire the second log and parse it to obtain the second operation performed for corresponding rescreening and the object pointed to by the second operation;
[0138] The view drawing unit is used to perform statistics on the first operation and the object pointed to by the first operation, the second operation and the object pointed to by the second operation to obtain statistical results, and render the statistical results on a pre-built PRISMA component to form a PRISMA view.
[0139] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0140] In the context of this disclosure, a readable storage medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A readable storage medium can be a machine-readable signal medium or a machine-readable storage medium. A readable storage medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0141] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including voice input, speech input, or tactile input).
[0142] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0143] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0144] It should be understood that the various forms of processes described above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this disclosure can be achieved, and this is not limited herein.
[0145] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A method for processing medical literature applied to evidence-based medicine, characterized in that, include: Obtain the text describing the purpose of medical research and parse it to identify multiple keywords that characterize the purpose of medical research; Obtain the URL address corresponding to the literature database that matches the stated medical research objective, and generate a medical literature source configuration table based on it; Extract the semantic features of each keyword and determine the degree of association between different keywords based on the semantic features of different keywords; For any given keyword, obtain its relevance to other keywords; Other keywords whose relevance exceeds a set relevance threshold are combined with any of the keywords to generate a search query for medical literature retrieval in the literature database; Add the search terms corresponding to all keywords to a pre-built search strategy configuration table, and call the medical literature source configuration table based on the search strategy configuration table to access the URL address corresponding to the literature database, and automatically populate the search terms into the search items of the literature database to retrieve medical literature in the literature database; The method further includes: Generate a search operation log, which includes the number of medical documents retrieved from each literature database using all search terms; Obtain the medical literature maintenance list, which records the estimated number of medical literatures containing any keyword; For any keyword, if the following conditions are met, the document feature data of the medical documents containing that keyword are obtained and mapped to the first database list to form a first-level database field list: Based on its corresponding search formula, the number of medical documents retrieved from a literature database is equal to the estimated number of medical documents containing any of the keywords recorded in the medical document maintenance list. Determine whether there are missing document feature data and / or duplicate document feature data in the primary database field list; if there are missing document feature data, complete them; if there are duplicate document feature data, remove duplicates, until the document feature status maintenance of all medical documents is completed.
2. The method according to claim 1, characterized in that, The method further includes: The medical literature maintenance list reading component is invoked to read the estimated quantity from the medical literature maintenance list; The retrieval operation log reading component is invoked to read the number of medical documents retrieved from each literature database using all search terms from the retrieval operation log, and compared with the estimated number to determine whether the two are equal.
3. The method according to claim 1, characterized in that, The method further includes: Based on the established document database information maintenance component, the fields in the primary database field list are checked one by one to perform the following steps to maintain the document feature status of each medical document, and the primary database field list is marked after the maintenance is completed.
4. The method according to claim 3, characterized in that, The method further includes: extracting data from the first-level database field list that has been marked to obtain the literature feature data corresponding to the fields, and mapping it to a second database list in a manner corresponding to medical literature to form a second-level database field list.
5. The method according to claim 4, characterized in that, The method further includes: obtaining a screening status marker obtained by screening the fields in the secondary library field list, and maintaining the document feature status of each medical document based on the screening status marker until the document feature status maintenance of all medical documents is completed.
6. The method according to claim 5, characterized in that, The method further includes: extracting data from the secondary database field list that has completed the status maintenance to obtain the literature feature data corresponding to the fields, and mapping it to the third database list in a manner corresponding to medical literature to form a tertiary database field list.
7. The method according to claim 6, characterized in that, The method further includes: Generate a first log that checks each field in the first-level database field list to maintain the document feature status of each medical document. A second log is generated to re-screen the fields in the secondary library field list in order to maintain the document feature status of each medical document.
8. The method according to claim 7, characterized in that, The method further includes: Obtain the first log and parse it to obtain the first operation performed for corresponding maintenance and the object pointed to by the first operation; Obtain the second log and parse it to obtain the second operation performed for the corresponding rescreening and the object pointed to by the second operation; Statistical results are obtained by performing statistics on the first operation and the object pointed to by the first operation, the second operation and the object pointed to by the second operation; The statistical results are rendered and displayed on a pre-built PRISMA component to form a PRISMA view.
9. A medical literature processing device for evidence-based medicine, characterized in that, include: The data parsing unit is used to obtain the text describing the purpose of medical research and to parse it to determine multiple keywords that characterize the purpose of medical research. The configuration table generation unit is used to obtain the URL address corresponding to the literature database that matches the medical research purpose, and generate a medical literature source configuration table based on it. The data extraction unit is used to extract the semantic features of each keyword and determine the degree of association between different keywords based on their semantic features. The relevance determination unit is used to obtain the relevance between any given keyword and other keywords. The retrieval formula generation unit is used to take other keywords whose relevance exceeds a set relevance threshold and combine them with any of the keywords to generate a retrieval formula for medical literature retrieval in the literature database; The literature retrieval unit is used to add the search terms corresponding to all keywords to a pre-built search strategy configuration table, and call the medical literature source configuration table based on the search strategy configuration table to access the URL address corresponding to the literature database, and automatically fill the search terms into the search items of the literature database to retrieve medical literature in the literature database. The first log unit is used to generate a retrieval operation log, which includes the number of medical documents retrieved from each literature database using all search terms; The maintenance list acquisition unit is used to acquire a medical literature maintenance list, which records the estimated number of medical literatures containing any keyword. The feature data acquisition unit is used to acquire the feature data of medical documents containing any keyword and map it to a first database list to form a first-level database field list if the following conditions are met: Based on its corresponding search formula, the number of medical documents retrieved from a literature database is equal to the estimated number of medical documents containing any of the keywords recorded in the medical document maintenance list. The first filtering unit is used to determine whether there are missing literature feature data in the primary database field list, and / or whether there are duplicate literature feature data. If there are missing literature feature data, it is filled in; if there are duplicate literature feature data, it is deduplicated, until the literature feature status maintenance of all medical literature is completed.
Citation Information
Patent Citations
Semantic retrieval method and system for patent literatures
CN111581349A
Document normalization method, document searching method, corresponding apparatuses, device, and storage medium
WO2017096777A1