Construction method of nursing field text annotation corpus based on deep learning

By building a text labeling corpus in nursing fields based on deep learning, the efficiency and accuracy problems in the existing technology are solved, and efficient and accurate text labeling and corpus management are achieved.

CN120045724APending Publication Date: 2025-05-27CHONGQING MEDICAL UNIVERSITY
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202510179111.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The prior art faces efficiency and accuracy issues when building text labeling corpus in nursing, especially due to the lack of high-quality labeling data and dependence on expertise.

Method used

Using a deep learning-based approach, a BERT-based entity recognition model for nursing field is constructed by collecting and preprocessing nursing text data, and combining manual annotation and post-processing steps to automatically annotate and optimize the corpus.

Benefits of technology

It improves the efficiency and accuracy of text labeling, reduces manual intervention, and realizes continuous update and optimization of the corpus, adapts to the needs of different nursing fields and application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045724A_ABST
    Figure CN120045724A_ABST
Patent Text Reader

Abstract

The invention discloses a method for constructing a nursing field text annotation corpus based on deep learning, and relates to the field of artificial intelligence technology and medical information processing. Comprising a data collection and preprocessing step, a BERT-based nursing field entity recognition model construction and training step, an automatic labeling and post-processing step, a corpus construction and management step and an application and intelligent support step. According to the method, the problems of efficiency and accuracy of text labeling in the nursing field are solved, and the continuously developing and changing text data processing requirements in the nursing field can be better met. The entity recognition model architecture specially aiming at the text characteristics in the nursing field is constructed based on deep learning, and challenges can be effectively handled when complex and diversified nursing text data is processed. And meanwhile, by utilizing the automatic labeling capability of deep learning, the labeling efficiency is improved, the manpower and time cost is reduced, and a new direction, a new mode and new experience are provided for nursing informatization development.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of artificial intelligence technology and medical information processing, and particularly to a method for constructing a text annotation corpus in the nursing field based on deep learning. Background Art

[0002] With the rapid development of the healthcare industry, the role of the nursing discipline in the entire medical system has become increasingly prominent. Currently, the construction of the nursing field corpus lags behind, lacking large-scale and high-quality annotated data, and it is difficult to meet the growing application requirements. Therefore, developing a nursing corpus that can automatically annotate texts and continuously update has become a key research issue.

[0003] Traditional text annotation methods in the nursing field mainly rely on manual annotation. This method is not only time-consuming and laborious but also inefficient, and it is difficult to ensure the consistency of annotation quality. In contrast, text annotation methods based on deep learning have significant features. Deep learning technology can automatically learn from a massive database, automatically adjust rule parameters, and optimize rules and models, thereby realizing the automation of text annotation. In medical practice, the two common model architectures of deep learning are mainly convolutional neural networks and recurrent neural networks, which perform well in text classification and entity recognition tasks, greatly improving the annotation efficiency and quality and reducing the uncertainty brought by manual intervention.

[0004] Deep learning, that is, deep network learning, refers to a collection of algorithms that can extract features of input data from low-level to high-level by simulating the hierarchical structure of the human brain, thereby being able to interpret the input data. In text annotation, it can automatically learn the complex semantics and semantic understanding of texts without the need for manual design of complex feature rules, reducing the dependence on domain expertise. Deep learning models, especially models based on Transformer, such as BERT, have achieved remarkable results in multiple natural language processing tasks. The advantage of these models lies in their ability to capture context information in texts and improve the accuracy of entity recognition and classification. In the nursing field, the application status of deep learning is still in its initial stage. With the sharp increase in the research popularity related to deep learning in the medical field, its accuracy, systematicness, and effectiveness have been preliminarily verified.

[0005] At present, there are still multiple challenges in constructing a text annotation corpus in the nursing field based on deep learning. First, the complexity of knowledge and the large amount of data in the nursing field pose certain resistance to the progress of various tasks of text information processing. Second, there is a lack of high-quality annotation data, especially for the annotation data of specific nursing tasks, which affects the training effect of deep learning models and is difficult to achieve the ideal annotation accuracy. In addition, the professional nature of the nursing field requires annotators to have certain medical background knowledge, which increases the difficulty and cost of annotation. In summary, developing a text annotation corpus in the nursing field based on deep learning has important research value and practical application prospects. Summary of the Invention

[0006] The purpose of the present invention is to overcome the deficiencies of the prior art and provide a construction method for a text annotation corpus in the nursing field based on deep learning, so as to solve the problems of the efficiency and accuracy of text annotation in the nursing field.

[0007] The purpose of the present invention is achieved through the following technical solutions: A construction method for a text annotation corpus in the nursing field based on deep learning, comprising: Data collection and preprocessing step: Collect nursing text data from professional databases, online communities, and different medical information systems, clean and standardize the collected nursing text data, and then perform word segmentation and term extraction on the nursing text data; Construction and training step of an entity recognition model in the nursing field based on BERT: Construct and train an entity recognition model in the nursing field based on BERT, and use the manually annotated text data as the training set during the training process for model training; Automatic annotation and post-processing step: Use the trained entity recognition model in the nursing field to automatically annotate the nursing text data, and use a learning model based on support vector machines to perform post-processing operations on the annotation results; Construction and management step of the corpus: Store the annotated nursing text data in the database in a standardized format, and construct a multi-level corpus table; Divide the nursing text into different categories according to the nature of the nursing text content, and assign a label that matches its content to each data record; Conduct regular data audits and manual verification to check whether the annotated data is accurate; Regularly evaluate the effect of data quality management, and according to the quality evaluation results, perform data correction and optimization.

[0008] Furthermore, the data collection and preprocessing step specifically includes: Data collection: Collect nursing text data from professional databases, online communities, and different medical information systems. The nursing text data includes nursing professional literature, clinical nursing guidelines, nursing-related data in online communities and forums, medical records, nursing records, nursing plans, drug usage records, audio and video materials, nursing textbooks, and teaching syllabi. Data cleaning and standardization: Clean the collected nursing text data to remove redundant information, incorrect data, or parts with non-standard formats; convert the collected nursing text data into a standardized format, perform desensitization processing on data that may involve privacy, and unify character encoding and format. Word segmentation and term extraction: Perform word segmentation on the nursing text to split the text into independent words or term units; extract medical terms, nursing terms, and vocabulary of common diseases using rule matching, word frequency statistics, and dictionary-based term matching, and construct a professional term library.

[0009] Furthermore, the specific steps of performing word segmentation on the nursing text to split the text into independent words or term units include the following: Use existing medical term dictionaries for accurate word segmentation. Use a word segmentation tool based on a probability model to supplement word segmentation for out-of-vocabulary words. Utilize the sub-word decomposition ability in the pre-trained model to split professional terms into meaningful units.

[0010] Furthermore, the specific steps of constructing and training the BERT-based entity recognition model in the nursing field specifically include: Construction and design of the BERT-based entity recognition model in the nursing field: Use the BERT network as the basic model and incorporate an advanced context awareness mechanism; add a multi-scale feature fusion module based on a convolutional neural network to capture local and global features in the text; add a custom feature layer in the nursing field, including a medical term embedding module and a relationship extraction module; generate medical term embeddings based on a medical knowledge graph or nursing process professional knowledge, combine them with BERT word vectors to identify specific entities in the nursing text; use the ClinicalBERT pre-trained model for initialization and perform secondary pre-training on nursing field data; perform lightweight processing on the BERT network using pruning or distillation methods, and adopt a parameter sharing mechanism for some layers in the entity recognition model in the nursing field. Data annotation and standardization: Organize a relevant expert team in the field to manually annotate the nursing text data. The annotation content includes professional terms, nursing measures, nursing records, and medication information; convert the annotation results into a standard format, perform data augmentation using synonym replacement, text transformation, and noise injection methods, and combine an active learning strategy to achieve dynamic expansion of the data; collect new nursing field terms or knowledge points and dynamically update the label system. Model training and optimization: Use the hierarchical partitioning method to divide the dataset into a training set, a validation set, and a test set with a ratio of 8:1:1; adjust the model parameters, use the backpropagation algorithm to update the model weights layer by layer based on the error, and use the cross-entropy loss function to calculate the difference between the predicted label and the true label; set up a fine-grained classification system, subdivide the nursing text entity types into multiple subcategories, design unique labels for each entity type and subcategory, and formulate detailed annotation specifications. Parameter adjustment: During the training process, adopt the AdaGrad adaptive learning rate adjustment algorithm to dynamically adjust the learning rate; introduce the L2 regularization term and the Dropout strategy respectively used to prevent the model from overfitting and enhance the generalization ability of the model; adopt the early stopping strategy to monitor the training process. When the performance on the validation set no longer improves, stop the training to avoid the model overfitting on the training set.

[0011] Further, the artificial annotation of the nursing text data specifically includes: Select professional text annotation tools to perform artificial annotation on the nursing text data. The annotation tools include Doccano, Prodigy, and BRAT. Configure the label system and task type of the annotation tool so that the annotation tool supports entity annotation and relation extraction functions. Perform artificial annotation according to the formulated detailed annotation rules, and define the annotation label system and establish the relationship definition between labels. Use the annotation tool to automatically check the annotation consistency, and manually review the parts with large annotation differences.

[0012] Further, the automatic annotation and post-processing steps include: Input the preprocessed large-scale nursing text data into the trained nursing domain entity recognition model for batch processing. The entity recognition model identifies and annotates the named entities in the text. During the automatic annotation process, use the monitoring feedback method for real-time monitoring, and at the same time compare the matching degree between the model prediction results and the annotation rules. Once any annotation errors or inconsistencies are found, make timely adjustments and corrections. Use natural language processing (NLP) tools to perform post-processing operations on the annotation results, including correcting typos and grammar; use the professional term standard library in the nursing field to proofread and correct the non-standard annotation data; conduct manual review on the post-processed annotation results, and let third-party experts make necessary decisions and corrections for the parts that the model fails to accurately annotate or have doubts about; feedback the problems found in the post-processing and manual review to the model training process to continuously optimize the model performance.

[0013] Further, the steps for constructing and managing the corpus specifically include: Corpus Storage and Management: Store the labeled nursing text data in the database in a standardized format. Use a relational database to store structured data and a NoSQL database to store unstructured data. Establish indexes for commonly queried fields, including keyword indexes, annotation item indexes, time indexes, and patient ID indexes. The corpus tables can be divided into a main table and auxiliary tables. The main table includes a nursing record table, a medication information table, and an annotation information table. The auxiliary tables include a nursing staff table and a patient information table. Classification and Tagging Management: Divide the nursing texts into different categories according to the nature of the content. Assign tags that match the content to each data record to endow semantic information. Set tag levels to support multi-level management so that each nursing text data can be accurately annotated. Use a one-to-many or many-to-many association form to associate each text record with its corresponding tag information. Design the data structure to achieve one-to-many or many-to-many tag associations. Data Quality Control and Evaluation: Regularly conduct data audits and manual verification to check whether the annotated data is accurate. Set up a data quality detection framework and data verification rules for quality assessment to ensure that the data quality meets the requirements of subsequent applications. Establish a data correction process to promptly correct or update incorrect data. Regularly evaluate the effectiveness of data quality management and make optimization adjustments according to the evaluation results.

[0014] Furthermore, the nursing record table is used to store the core information of each nursing record, the medication information table is used to store the patient's drug usage records and their annotations, and the annotation information table is used to store the detailed content of all annotations. The nursing staff table is used to store the basic information of the nursing staff, including the name and employee number of the nursing staff, and the patient information table is used to store the basic data of the patient, including the patient ID, gender, and age.

[0015] Furthermore, it also includes application and intelligent support steps, specifically including: Data Application: Apply the nursing field text annotation corpus to the intelligent nursing management system, intelligent medical decision support, and nursing quality assessment and reporting fields. Update and Optimization of the Deep Learning Model: Regularly update the deep learning model, use the newly annotated data for retraining and model fine-tuning, and improve the accuracy and adaptability of the model through hyperparameter optimization, data augmentation techniques, transfer learning, and model integration methods. Regularly evaluate the model performance and continuously monitor the running effect of the model.

[0016] The beneficial effects of the present invention are: 1) Improve the text annotation efficiency By designing a customized model in the nursing field based on BERT, leveraging its powerful computing capabilities and automated learning features, common text patterns in the nursing field are recognized. In the data collection and preprocessing stage, various word segmentation and term extraction methods are adopted to lay the foundation for annotation. Using the preprocessed large-scale data for learning, and then quickly annotating new nursing texts. This significantly improves work efficiency compared to traditional manual annotation methods, providing strong support for subsequent data analysis, clinical decision-making, and nursing research.

[0017] 2) Improve the accuracy of text annotation Due to the high professionalism and complexity of nursing field texts, text annotation methods based on deep learning can deeply learn the semantic, syntactic, and professional knowledge features of texts through training on a large-scale and diverse nursing text corpus, avoiding errors caused by one-sided understanding or knowledge limitations in manual annotation, and thus implementing annotation more accurately. In addition, in the design of deep learning models, specific knowledge in the nursing field is introduced, and various optimization strategies are adopted, such as hierarchical partitioning, parameter adjustment, regularization, etc., effectively improving the annotation accuracy. A professional team formulates detailed rules, defines annotation specifications and label systems, accurately annotates entities and processes fuzzy and implicit information, while performing data augmentation and dynamic updates. By real-time monitoring and detecting annotation results, errors are corrected in a timely manner, improving the quality of annotation. At the same time, cross-checking and expert review are implemented. Through cross-validation among different annotators and review by domain experts, the accuracy and authority of annotation results are ensured.

[0018] 3) Adaptability and generalizability The text annotation corpus based on deep learning has good adaptability and can be appropriately adjusted and optimized according to the text characteristics of different nursing fields and application scenarios. By collecting nursing text data from different sources and of different types for training, the model learns the features and patterns of various texts. At the same time, by adjusting the structure and parameters of the model, it can adapt to the needs of text annotation in different scenarios. In addition, the nursing field text annotation corpus can also be applied to multiple fields such as intelligent nursing management systems, intelligent medical decision support, nursing quality assessment and reporting, etc., providing data support and decision-making basis for different fields. Whether it is a comprehensive hospital or a university education institution, this method can be used to build a corpus according to its own needs, promoting the effective utilization and sharing of nursing text data and driving the informatization development of the entire nursing industry.

[0019] 4) Data standardization in the nursing field Constructing a text annotation corpus in the nursing field helps to standardize data. Deep learning models can annotate and transform nursing text data based on a pre-established annotation specification system and the classification of entity types, etc., which can improve the performance of natural language processing tasks and provide richer and deeper data resources for information extraction and knowledge discovery in the nursing field. Brief Description of the Drawings

[0020] Figure 1 It is a flowchart of a construction method for a text annotation corpus in the nursing field based on deep learning. Detailed Implementation Manner

[0021] Next, the technical solutions of the present invention will be clearly and completely described in conjunction with the embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts fall within the protection scope of the present invention.

[0022] Refer to Figure 1 , the present invention provides a technical solution: A construction method for a text annotation corpus in the nursing field based on deep learning, including: S1. Data collection and preprocessing step: Collect nursing text data from professional databases, online communities, and different medical information systems, clean and standardize the collected nursing text data, and then perform word segmentation and term extraction on the nursing text data; The data collection and preprocessing step specifically includes: S11. Data collection: Collect nursing text data from professional databases, online communities, and different medical information systems. The nursing text data includes nursing professional literature, clinical nursing guidelines, nursing-related data in online communities and forums, medical records, nursing records, nursing plans, drug usage records, audio and video materials, nursing textbooks, and teaching syllabuses; The medical information systems include electronic medical record systems, medication management systems, nursing information systems, etc.

[0023] By screening different nursing data sources, the comprehensiveness and diversity of the data are ensured, specifically including: (1) Obtain patient nursing records, medical record data information, etc. through the hospital information system (HIS) and the patient's electronic medical record (EMR), etc.

[0024] (2) Obtain standardized nursing operation procedures and nursing methods through relevant nursing professional books, operation manuals, guidelines, etc.

[0025] (3)Analyze the structures of data from different sources, which are divided into structured (such as database records) and unstructured (such as PDF documents), and determine suitable data extraction methods.

[0026] (4)Audio and video resources in the clinical environment, teaching materials and syllabuses in nursing teaching.

[0027] For different data sources, different data collection methods are used, specifically including: (1)For unstructured data, such as PDF files and web content, use web crawler technology to capture the required content.

[0028] (2)For hospital information systems, use dedicated data interfaces to obtain internal data.

[0029] (3)Directly obtain and collect various types of texts and structured materials.

[0030] S12. Data cleaning and standardization: Clean the collected nursing text data to remove redundant information, incorrect data, or parts with non-standard formats, such as clause splitting, removing extra spaces and symbols. Convert the collected nursing text data into a standardized format, such as consistent time format and units, desensitize the data that may involve privacy, and unify the character encoding and format; Data cleaning can ensure the legality and compliance of the data, and the purpose of standardization is to convert the data into a format suitable for subsequent deep learning models.

[0031] S13. Word segmentation and term extraction: Perform word segmentation on the nursing text to split the text into independent word or term units; Extract medical terms, nursing terms, and vocabulary of common diseases in the form of rule matching, word frequency statistics, and dictionary-based term matching, and build a professional term library.

[0032] Word segmentation is to facilitate subsequent term extraction and annotation. The methods mainly include dictionary-based word segmentation, statistical word segmentation, and deep learning word segmentation methods. In this embodiment, the step of performing word segmentation on the nursing text to split the text into independent word or term units specifically includes the following steps: Use existing medical term dictionaries (such as UMLS, ICD-10, domestic medical vocabulary sets) for accurate word segmentation; Use word segmentation tools based on probability models (such as HMM, CRF) (such as Jieba segmentation) to supplement the word segmentation of out-of-vocabulary words; Utilize the sub-word decomposition ability in pre-trained models (such as BERT, RoBERTa) to split professional terms into meaningful units.

[0033] S2. Construction and Training Steps of the Entity Recognition Model in the Nursing Field Based on BERT: Construct and train an entity recognition model in the nursing field based on BERT. During the training process, use the manually annotated text data as the training set for model training. Further, the construction and training steps of the entity recognition model in the nursing field based on BERT specifically include: S21. Construction and Design of the Entity Recognition Model in the Nursing Field Based on BERT: Use the BERT network as the basic model and incorporate an advanced context awareness mechanism; add a multi-scale feature fusion module based on the convolutional neural network to capture local and global features in the text; add a custom feature layer for the nursing field, including a medical term embedding module and a relation extraction module; generate medical term embeddings based on a medical knowledge graph or nursing process expertise, combine them with the BERT word vectors to identify specific entities in nursing texts; use the ClinicalBERT pre-trained model for initialization and perform secondary pre-training on nursing field data; use pruning or distillation methods to lightweight the BERT network and adopt a parameter sharing mechanism for some layers in the entity recognition model in the nursing field.

[0034] In the text annotation task in the nursing field, to adapt to the characteristics and annotation requirements of nursing texts, introduce specific knowledge such as medical terms and nursing processes in the nursing field, design and train an entity recognition model in the nursing field based on BERT (Bidirectional Encoder Representations from Transformers). Incorporate a more advanced context awareness mechanism into the model so that the model can more accurately capture the context information in the text to improve the recognition of implicit relationships in nursing texts and improve the annotation accuracy. Generate medical term embeddings based on a medical knowledge graph (such as UMLS) or nursing process expertise, and combining these domain knowledge embeddings with the BERT word vectors can enhance the understanding of proprietary nouns in the nursing field. To improve the training and inference efficiency, incorporating lightweight design can improve the training and inference efficiency of the model, and designing a parameter sharing mechanism for some layers can reduce redundant calculations. Through the above architecture design and optimization, it is possible to ensure the improvement of the accuracy of the text annotation task in the nursing field while reducing the complexity and inference cost of the model, providing an efficient and professional natural language processing solution for the nursing field.

[0035] S22. Data annotation and standardization: Organize a team of relevant experts in the field (such as front-line clinical nurses and doctors, clinical nursing experts, nursing education and researchers, etc.) to manually annotate the nursing text data. The annotation content includes professional terms, nursing measures, nursing records, and medication information. Convert the annotation results into a standard format, such as the BIO format (B-Disease, I-Disease, etc.) or the JSON format, for subsequent training use. Perform data augmentation by means of synonym replacement, text transformation, and noise injection, and combine active learning strategies to achieve dynamic expansion of the data. Collect new nursing field terms or knowledge points and dynamically update the label system. Among them, professional terms include disease names, symptoms, anatomical terms, treatment methods, drug names, etc., nursing measures include nursing operation procedures, nursing plans, etc., nursing records include the current situation of the patient, a brief description of the condition, the diagnosis and treatment process, and medication information includes drug names, dosages, administration methods, and times, etc.

[0036] Furthermore, the manual annotation of the nursing text data specifically includes: Select a professional text annotation tool to manually annotate the nursing text data. The annotation tools include Doccano, Prodigy, and BRAT. Configure the label system and task type of the annotation tool so that the annotation tool supports entity annotation and relationship extraction functions. Perform manual annotation according to the formulated detailed annotation rules, and define the annotation label system and establish the relationship definition between labels, such as the treatment relationship of "drug - symptom improvement" and the corresponding relationship of "nursing measure - nursing diagnosis". Use the annotation tool to automatically check the annotation consistency, and manually review the parts with large annotation differences.

[0037] According to the above data annotation and standardization process, construct a high-quality, professional, and dynamically updated nursing field annotation dataset to provide reliable data support for the subsequent deep learning model training.

[0038] S23. Model training and optimization: Use the stratified division method to divide the dataset into a training set, a validation set, and a test set with a ratio of 8:1:1. Adjust the model parameters, use the backpropagation algorithm to update the model weights layer by layer based on the error, and use the cross-entropy loss function to calculate the difference between the predicted label and the true label. Set up a fine-grained classification system, subdivide the nursing text entity types into multiple subcategories, design a unique label for each entity type and subcategory, and formulate detailed annotation specifications.

[0039] During training, the Adam optimizer is used to improve the convergence speed and stability. The cross-entropy loss function calculates the difference between the predicted labels and the true labels to optimize the accuracy of the annotation task. Configure the model training parameters, set the initial learning rate (such as 1e-4), and dynamically adjust it in combination with a learning rate scheduler; select an appropriate batch size (such as 16 or 32) according to the data scale and hardware resources; dynamically observe the performance of the validation set during training and select an appropriate number of training epochs. Use accuracy, recall, and F1 score to measure the model training effect.

[0040] S24. Parameter adjustment: During the training process, the AdaGrad adaptive learning rate adjustment algorithm is used to dynamically adjust the learning rate; introduce the L2 regularization term and Dropout strategy respectively used to prevent model overfitting and enhance the generalization ability of the model; adopt the early stopping strategy to monitor the training process. When the performance on the validation set no longer improves, stop the training to avoid model overfitting on the training set.

[0041] In a specific embodiment, the process of setting up the fine-grained classification system is as follows: Subdivide the entity types in the nursing text, such as patient information (name, age, gender, medical history, etc.), diagnosis information (disease name, symptom description, etc.), medication information (drug name, dosage, usage, etc.), nursing operations (drug administration, dressing change, monitoring, etc.). The following is a specific case: Patient information (Name: Zhang San, Age: 35 years old, Gender: Male, Medical history: Previous history of diabetes); Diagnosis information (Disease name: Coronary heart disease, Symptom description: Chest tightness, Shortness of breath); Medication information (Drug name: Aspirin, Dosage: 100mg, Usage: Once a day / qd); Nursing operations (Drug administration: Intravenous injection of cefuroxime; Dressing change: Perform a wound dressing change once; Monitoring: Electrocardiogram monitoring for 24 hours).

[0042] On this basis, a unique label is designed for each entity type and subcategory, and a detailed annotation specification is formulated to ensure the consistency and distinguishability of the annotation results, which provides convenience for subsequent text analysis and applications. For example, for the BIO annotation format, it can be set as B-Disease: The starting part of the disease name; I-Medication: The middle part of the drug name; B-Procedure: The starting part of the nursing operation.

[0043] During the annotation process, define the annotation specifications, clarify the annotation standards and examples, ensure that annotators have a consistent understanding, and explain the handling methods for boundary cases such as ambiguous descriptions or implicit information; accurately annotate each entity in the text. For example: The patient's body temperature shows an increase to 39°C → The patient O body temperature B-Symptom shows O increase O to O 39°C I-Symptom; maintain clear label distinctions for multiple entities or nested entities. Finally, use an annotation tool to automatically check the annotation consistency (such as entity boundaries, label types), and manually review the parts with significant annotation differences to improve the accuracy and practicality of the annotation results.

[0044] S3. Automatic annotation and post-processing steps: Use the trained entity recognition model in the nursing field to automatically annotate the nursing text data, and use a learning model based on support vector machines to perform post-processing operations on the annotation results. Specifically, it includes: S31. Input the preprocessed large-scale nursing text data into the trained entity recognition model in the nursing field for batch processing. The entity recognition model identifies and annotates the named entities in the text; during the automatic annotation process, a monitoring and feedback method is adopted for real-time monitoring, and at the same time, the matching degree between the model prediction results and the annotation rules is compared. Once annotation errors or inconsistencies are found, they are adjusted and corrected in a timely manner; To improve the efficiency and accuracy of automatic annotation, in this embodiment, parallel computing technology is used to divide the collected large-scale nursing text data into several small batches and process multiple batches simultaneously; an automatic annotation tool (such as an NER model) is used to process multiple data segments simultaneously, effectively improving the annotation efficiency and shortening the annotation cycle. Implementing monitoring and feedback ensures the stability of the annotation quality.

[0045] S32. Use natural language processing NLP tools (such as the grammar correction model of Hugging Face or a dedicated spelling check tool) to perform post-processing operations on the annotation results, including correcting typos and grammar; use the professional term standard library in the nursing field to proofread and correct the annotation data that does not meet the standards; for example, detect and remove duplicate entities or relationships in the annotation data, and merge the annotation results for the same entity that appears in different positions. Use the professional term standard library in the nursing field (such as UMLS, domestic nursing term sets) to match and calibrate the annotation results, convert the terms that do not meet the standards, and ensure the unity and standardization of the terms in the annotation results. For example, if the annotation result is hypertension, but the standard term is essential hypertension, then use the standard term as the reference and replace the original annotation result.

[0046] Manually review the post - processed annotation results. For parts that the model fails to accurately annotate or where there are doubts, third - party experts make necessary decisions and corrections. Feed back the problems found in post - processing and manual review to the model training process to continuously optimize the model performance. Adopt sampling inspection, full - volume inspection, and double - person review. Extract samples from the post - processed annotation results, and have experts manually verify the correctness of the annotations. For key data sets or important annotation types, conduct item - by - item manual verification. Have two annotators independently review the annotation results, and for inconsistent parts, third - party experts make decisions. The above strategies improve the annotation quality, enhance the data utilization rate, and strengthen the reliability of the results, providing a reliable data foundation for subsequent applications such as model training and text analysis.

[0047] S4. Steps for corpus construction and management: Store the annotated nursing text data in the database in a standardized format, and construct a multi - level corpus table; According to the nature of the nursing text content, divide the nursing text into different categories, and assign a label that matches its content to each data record; Conduct regular data audits and manual verification to check whether the annotated data is accurate; Regularly evaluate the effectiveness of data quality management, and based on the quality assessment results, make data corrections and optimizations.

[0048] Specifically, the steps for corpus construction and management include: S41. Corpus storage and management: Store the annotated nursing text data in the database in a standardized format (such as JSON, CSV, or SQL table). Use a relational database (such as MySQL, PostgreSQL) to store structured data and a NoSQL database (such as MongoDB) to store unstructured data; Establish indexes for commonly queried fields, including keyword indexes, annotation item indexes, time indexes, and patient ID indexes for quick retrieval, update, and maintenance; The corpus table can be divided into a main table and auxiliary tables. The main table includes a nursing record table, a medication information table, and an annotation information table, and the auxiliary tables include a nursing staff table and a patient information table.

[0049] Among them, the nursing record table is used to store the core information of each nursing record, the medication information table is used to store the patient's medication records and their annotations, and the annotation information table is used to store the detailed content of all annotations; The nursing staff table is used to store the basic information of nursing staff, including the names and staff numbers of nursing staff, and the patient information table is used to store the basic data of patients, including patient ID, gender, and age.

[0050] S42. Classification and tagging management: Classify nursing texts into different categories according to the nature of the content; assign tags that match the content to each data record to endow semantic information; set tag levels to support multi-level management so that each piece of nursing text data can be accurately labeled; use one-to-many or many-to-many associations to associate each text record with its corresponding tag information; design a data structure to achieve one-to-many or many-to-many tag associations. In addition, use a deep learning model to automatically assign predefined tags to text data, and conduct manual review on the results of automatic tagging and make adjustments if necessary. Regularly update and maintain the tag system to adapt to new nursing needs and practices.

[0051] S43. Data quality control and evaluation: Regularly conduct data audits and manual verification to check whether the labeled data is accurate; set up a data quality detection framework and data verification rules for quality assessment to ensure that the data quality meets the requirements of subsequent applications; establish a data correction process to correct or update incorrect data in a timely manner; regularly evaluate the effectiveness of data quality management and make optimization adjustments according to the evaluation results.

[0052] Furthermore, it also includes application and intelligent support steps, specifically including: Data application: Apply the annotated corpus of nursing texts to the fields of intelligent nursing management systems, intelligent medical decision support, and nursing quality assessment and reporting; use the annotated corpus as an educational resource for the education of nursing students and the continuing education of in-service nurses, providing practical case analysis and learning. Analyze the data in the annotated corpus to discover new nursing knowledge and best practices, which can promote innovation in the nursing field.

[0053] The intelligent nursing management system can help nurses and medical managers efficiently process patient care information, track care plans, and analyze nursing behaviors by integrating nursing text data, providing data support, such as automatically assisting in generating nursing records and nursing operation reminders, recommending personalized care plans, and nursing human resource management. The intelligent medical decision support system CDSS that relies on nursing texts and annotated data can help medical staff make accurate and timely treatment decisions, such as assisting in accurate diagnosis, personalized medication guidance, and clinical pathway optimization. The nursing quality assessment system evaluates the nursing quality by analyzing nursing text data and generates relevant reports for feedback.

[0054] Update and optimization of the deep learning model: Regularly update the deep learning model, conduct retraining and model fine-tuning using new annotated data, and improve the accuracy and adaptability of the model through hyperparameter optimization, data augmentation techniques, transfer learning, and model integration methods; regularly evaluate the model performance and continuously monitor the running effect of the model.

[0055] The present invention efficiently creates and continuously updates a text annotation corpus through a deep learning model. The specific beneficial effects include: 1. Improve text annotation efficiency By designing a customized model for the nursing field based on BERT and leveraging its powerful computing capabilities and automated learning features, common text patterns in the nursing field are recognized. In the data collection and preprocessing stage, various word segmentation and term extraction methods are adopted to lay the foundation for annotation. The preprocessed large-scale data is used for learning, and then new nursing texts are quickly annotated. This significantly improves the work efficiency compared to traditional manual annotation methods, providing strong support for subsequent data analysis, clinical decision-making, and nursing research.

[0056] 2. Enhance the accuracy of text annotation Due to the high professionalism and complexity of nursing field texts, the deep learning-based text annotation method can deeply learn the semantic, syntactic, and professional knowledge features of texts through training on a large-scale and diverse nursing text corpus, avoiding errors caused by one-sided understanding or knowledge limitations in manual annotation, and thus implementing annotation more accurately. In addition, in the design of the deep learning model, specific knowledge in the nursing field is introduced, and various optimization strategies are adopted, such as hierarchical partitioning, parameter adjustment, regularization, etc., effectively improving the annotation accuracy. A professional team formulates detailed rules, defines annotation specifications and label systems, accurately annotates entities and processes ambiguous and implicit information, while performing data augmentation and dynamic updates. By real-time monitoring and detecting annotation results, errors are corrected in a timely manner, improving the quality of annotation. At the same time, cross-checks and expert reviews are implemented. Through cross-validation among different annotators and reviews by domain experts, the accuracy and authority of the annotation results are ensured.

[0057] 3. Adaptability and generalizability The deep learning-based text annotation corpus has good adaptability and can be appropriately adjusted and optimized according to the text characteristics of different nursing fields and application scenarios. By collecting nursing text data from different sources and of different types for training, the model learns the features and patterns of various texts. At the same time, by adjusting the structure and parameters of the model, it can adapt to the text annotation requirements in different scenarios. In addition, the nursing field text annotation corpus can also be applied to multiple fields such as intelligent nursing management systems, intelligent medical decision support, nursing quality assessment and reporting, etc., providing data support and decision-making basis for different fields. Whether it is a comprehensive hospital or a university education institution, this method can be used to build a corpus according to its own needs, promoting the effective utilization and sharing of nursing text data and driving the informatization development of the entire nursing industry.

[0058] 4. Data standardization in the nursing field Constructing a text annotation corpus in the nursing field helps to standardize data. Deep learning models can annotate and transform nursing text data according to a pre-established annotation specification system and the division of entity types, etc., which can improve the performance of natural language processing tasks and provide richer and deeper data resources for information extraction and knowledge discovery in the nursing field.

[0059] The above are only the preferred embodiments of the present invention. It should be understood that the present invention is not limited to the form disclosed herein, should not be regarded as excluding other embodiments, but can be used in various other combinations, modifications and environments, and can be changed within the scope of the concept described herein through the above teachings or the technology or knowledge in related fields. And the changes and alterations made by those skilled in the art without departing from the spirit and scope of the present invention shall fall within the protection scope of the appended claims of the present invention.

Claims

1. A method for constructing a text annotation corpus in the nursing field based on deep learning, characterized in that: include: Data collection and preprocessing steps: Collect nursing text data from professional databases, online communities, and different medical information systems, clean and standardize the collected nursing text data, and then perform word segmentation and term extraction on the nursing text data; Steps for building and training a BERT-based nursing entity recognition model: Build and train a BERT-based nursing entity recognition model. During the training process, manually annotated text data is used as the training set in the model. Automatic labeling and post-processing steps: Use the trained nursing entity recognition model to automatically label the nursing text data, and use the support vector machine-based learning model to post-process the labeling results; The steps for constructing and managing the corpus are as follows: store the annotated nursing text data in a standardized format in the database and construct a multi-level corpus table; divide the nursing text into different categories according to the nature of the nursing text content, and assign a label that matches its content to each data record; conduct regular data audits and manual verifications to check whether the annotated data is accurate; regularly evaluate the effectiveness of data quality management, and perform data correction and optimization based on the quality assessment results.

2. According to a method for constructing a nursing field text annotation corpus based on deep learning according to claim 1, it is characterized in that: The data collection and preprocessing steps specifically include: Data collection: Collect nursing text data from professional databases, online communities, and different medical information systems. The nursing text data includes nursing professional literature, clinical nursing guidelines, nursing-related data in online communities and forums, medical records, nursing records, nursing plans, drug use records, audio and video materials, nursing teaching materials, and teaching syllabuses; Data cleaning and standardization: Clean the collected nursing text data to remove redundant information, erroneous data or parts with irregular formats; convert the collected nursing text data into a standardized format, desensitize the data that may involve privacy, and unify the character encoding and format; Word segmentation and term extraction: perform word segmentation on the nursing text and split the text into independent words or term units; use rule matching, word frequency statistics, and dictionary-based term matching to extract medical terms, nursing terms, and common disease vocabulary, and build a professional term library.

3. The method for constructing a nursing field text annotation corpus based on deep learning according to claim 1, characterized in that: The word segmentation process of the nursing text to split the text into independent words or term units specifically includes the following steps: Use existing medical terminology dictionaries for accurate word segmentation; Use a word segmentation tool based on a probability model to supplement the segmentation of unregistered words; Leverage the subword decomposition capabilities of the pre-trained model to segment professional terms into meaningful units.

4. The method for constructing a nursing field text annotation corpus based on deep learning according to claim 1, characterized in that: The steps of constructing and training the BERT-based nursing entity recognition model specifically include: Construction and design of BERT-based entity recognition model in the nursing field: Using the BERT network as the basic model, it integrates advanced context-aware mechanisms; adds a multi-scale feature fusion module based on a convolutional neural network to capture local and global features in the text; adds a custom feature layer in the nursing field, including a medical term embedding module and a relationship extraction module; generates medical term embeddings based on medical knowledge graphs or nursing process expertise, and combines them with BERT word vectors to identify specific entities in nursing texts; uses the ClinicalBERT pre-trained model as initialization and performs secondary pre-training on nursing field data; uses pruning or distillation methods to lightweight the BERT network, and adopts a parameter sharing mechanism for some layers in the entity recognition model in the nursing field; Data labeling and standardization: Organize a team of experts in this field to manually label the nursing text data, including professional terms, nursing measures, nursing records, and medication information; convert the labeling results into a standard format, use synonym replacement, text transformation, and noise injection to enhance the data, and combine active learning strategies to achieve dynamic expansion of data; collect new nursing terms or knowledge points, and dynamically update the label system; Model training and optimization: Use the hierarchical partitioning method to divide the data set into a training set, a validation set, and a test set with a ratio of 8:1:1; adjust the model parameters, use the back propagation algorithm to update the model weights layer by layer based on the error, and use the cross entropy loss function to calculate the difference between the predicted label and the true label; set up a fine-grained classification system, subdivide the nursing text entity types into multiple subcategories, design unique labels for each entity type and subcategory, and formulate detailed annotation specifications; Parameter adjustment: During the training process, the AdaGrad adaptive learning rate adjustment algorithm is used to dynamically adjust the learning rate. The L2 regularization term and Dropout strategy are introduced to prevent model overfitting and enhance the generalization ability of the model respectively. The early stopping strategy is used to monitor the training process. When the performance on the validation set no longer improves, the training is stopped to avoid overfitting of the model on the training set.

5. The method for constructing a nursing field text annotation corpus based on deep learning according to claim 4 is characterized in that: The manual labeling of the nursing text data specifically includes: Select professional text annotation tools to manually annotate nursing text data, including Doccano, Prodigy, and BRAT; Configure the labeling system and task type of the annotation tool to enable the annotation tool to support entity annotation and relationship extraction functions; Perform manual labeling according to the detailed labeling rules, define the labeling system and establish the relationship between labels; Use annotation tools to automatically check annotation consistency and manually review parts with large annotation differences.

6. The method for constructing a nursing field text annotation corpus based on deep learning according to claim 1, characterized in that: The automatic marking and post-processing steps include: The pre-processed large-scale nursing text data is input into the trained nursing field entity recognition model for batch processing. The entity recognition model identifies and labels the named entities in the text. During the automatic labeling process, the monitoring feedback method is used for real-time monitoring, and the matching degree between the model prediction results and the labeling rules is compared. Once labeling errors or inconsistencies are found, adjustments and corrections are made in a timely manner. Use natural language processing (NLP) tools to post-process the annotation results, including spelling and grammar correction; use the standard library of professional terminology in the nursing field to proofread and correct the annotation data that does not meet the standards; manually review the post-processed annotation results, and third-party experts make necessary decisions and corrections for parts that the model failed to accurately label or have doubts; feedback the problems found in post-processing and manual review to the model training process to continuously optimize model performance.

7. The method for constructing a nursing field text annotation corpus based on deep learning according to claim 1, characterized in that: The steps of constructing and managing the corpus specifically include: Corpus storage and management: Store the annotated nursing text data in a standardized format in the database, use a relational database to store structured data, and use a NoSQL database to store unstructured data; create indexes for commonly used query fields, including keyword indexes, annotation item indexes, time indexes, and patient ID indexes; divide the corpus tables into primary tables and auxiliary tables. The primary tables include nursing record tables, medication information tables, and annotation information tables, and the auxiliary tables include nursing staff tables and patient information tables. Classification and labeling management: According to the nature of the nursing text content, the nursing text is divided into different categories; each data record is assigned a label that matches its content and given semantic information; the label level is set to support multi-level management so that each nursing text data can be accurately labeled; each text record is associated with its corresponding label information in the form of one-to-many association or many-to-many association; the data structure is designed to realize one-to-many or many-to-many label association; Data quality control and assessment: Conduct data audits and manual verifications regularly to check whether the labeled data is accurate; set up a data quality detection framework and data verification rules to conduct quality assessments to ensure that the data quality meets the requirements of subsequent applications; establish a data correction process to correct or update erroneous data in a timely manner; regularly evaluate the effectiveness of data quality management and make optimization adjustments based on the evaluation results.

8. The method for constructing a nursing field text annotation corpus based on deep learning according to claim 7, characterized in that: The nursing record table is used to store the core information of each nursing record, the medication information table is used to store the patient's medication use records and their annotations, and the annotation information table is used to store the detailed contents of all annotations; the nursing staff table is used to store the basic information of the nursing staff, including the nursing staff's name and work number, and the patient information table is used to store the patient's basic data, including the patient ID, gender, and age.

9. The method for constructing a nursing field text annotation corpus based on deep learning according to claim 1, characterized in that: It also includes application and intelligent support steps, including: Data application: Apply the nursing text annotation corpus to the fields of intelligent nursing management system, intelligent medical decision support, and nursing quality assessment and reporting; Update and optimization of deep learning models: Regularly update deep learning models, use newly annotated data for retraining and model fine-tuning, and improve the accuracy and adaptability of models through hyperparameter optimization, data enhancement technology, transfer learning, and model integration; regularly evaluate model performance and continuously monitor model operation results.

Citation Information

Cited By

  • Label quality control method and device with active learning ability

    CN120782332A

  • Home disabled elderly safety care education content system construction method and system

    CN121093939A

  • A home disabled elderly safety care education content system construction method and system

    CN121093939B

  • Corpus construction method and system

    CN121122547A

  • Nursing teaching task-oriented nursing field text annotation corpus construction method

    CN122020186A