Content detection model training and fine-tuning method, content detection method, device, equipment, medium and product

CN120950979BActive Publication Date: 2026-08-21BEIJING HENGAN JIAXIN SAFETY TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511404375.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-28
Publication Date
2026-08-21
Estimated Expiration
2045-09-28

AI Technical Summary

Technical Problem

以医疗场景为例,对于医疗文件的跨境传输监管,对日益复杂的文件类型、多语义环境、动态变化的信息下,为跨境监管带来具大挑战

Benefits of technology

[0034] The technical solution of this invention involves acquiring sample medical documents and sample tag data; training an initial detection model using the sample medical documents and sample tag data to obtain a basic content detection model; wherein the initial detection model is a BERT model; and optimizing the basic content detection model based on a medical knowledge base to obtain an optimized content detection model. This technical solution improves the efficiency and accuracy of medical document content recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120950979B_ABST
    Figure CN120950979B_ABST
Patent Text Reader

Abstract

The application discloses a content detection model training and fine-tuning method, a content detection method, a device, equipment, a medium and a product, and relates to the technical fields of large models, deep learning and medical treatment. The method comprises the following steps: acquiring sample medical documents and sample label data; training an initial detection model by using the sample medical documents and the sample label data, so as to obtain a basic content detection model; wherein the initial detection model is a BERT model; and optimizing the basic content detection model based on a medical knowledge base, so as to obtain an optimized content detection model. Through the technical solution, the content detection efficiency and accuracy of medical documents can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of large models, deep learning, and medical technology, and in particular to a method for training and fine-tuning a content detection model, a content detection method, apparatus, equipment, medium, and product. Background Technology

[0002] With the accelerating digitalization of the global economy, cross-border data flow has become a fundamental requirement for international business activities. Ensuring data security while promoting the orderly flow of data across borders is crucial. However, the explosive growth in data volume and the diversification of data types pose significant challenges to cross-border data supervision. Existing methods for cross-border data transmission scenarios are insufficient in terms of compliance, robustness, and cross-language adaptability. Taking the medical scenario as an example, the supervision of cross-border transmission of medical documents presents significant challenges due to increasingly complex document types, multi-semantic environments, and dynamically changing information. Summary of the Invention

[0003] This invention provides a method for training and fine-tuning a content detection model, a content detection method, an apparatus, a device, a medium, and a product to efficiently and accurately identify the content of medical documents.

[0004] According to one aspect of the present invention, a method for training a content detection model is provided, the method comprising:

[0005] Obtain sample medical documents and sample label data;

[0006] The initial detection model is trained using the sample medical documents and the sample tag data to obtain a basic content detection model; wherein, the initial detection model is a BERT model;

[0007] Based on a medical knowledge base, the basic content detection model is optimized to obtain an optimized content detection model.

[0008] According to another aspect of the present invention, a method for fine-tuning a content detection model is provided, the method comprising:

[0009] Obtain labeled medical documents;

[0010] The target content detection model is obtained by fine-tuning the annotated medical documents to obtain the optimized content detection model; wherein the optimized content detection model is trained by the content detection model training method provided by the present invention.

[0011] According to another aspect of the present invention, a content detection method is provided, the method comprising:

[0012] Retrieve the medical documents to be tested;

[0013] The medical document to be detected is input into the target content detection model to obtain the final detection result; wherein, the target content detection model is fine-tuned according to the fine-tuning method of the content detection model provided by the present invention;

[0014] Alternatively, create a medical prompt template;

[0015] The medical document to be detected and the medical prompt word template are input into the large model to obtain the final detection result; wherein, the final detection result includes the final predicted document type and the final predicted privacy entity.

[0016] According to another aspect of the present invention, a training apparatus for a content detection model is provided, the apparatus comprising:

[0017] The sample data acquisition module is used to acquire sample medical documents and sample label data;

[0018] The basic detection model determination module is used to train the initial detection model using the sample medical documents and the sample label data to obtain a basic content detection model; wherein, the initial detection model is a BERT model;

[0019] An optimized content detection module is used to optimize the basic content detection model based on a medical knowledge base, resulting in an optimized content detection model.

[0020] According to another aspect of the present invention, a fine-tuning apparatus for a content detection model is provided, the apparatus comprising:

[0021] The labeled document acquisition module is used to acquire labeled medical documents;

[0022] The target detection model fine-tuning module is used to fine-tune the optimized content detection model using the labeled medical documents to obtain the target content detection model; wherein the optimized content detection model is trained using the content detection model training method provided by the present invention.

[0023] According to another aspect of the present invention, a content detection apparatus is provided, the apparatus comprising:

[0024] The document acquisition module is used to acquire medical documents to be inspected.

[0025] The final detection result determination module is used to input the medical document to be detected into the target content detection model to obtain the final detection result; wherein, the target content detection model is fine-tuned according to the fine-tuning method of the content detection model provided by the present invention;

[0026] The prompt word template building module is used to build medical prompt word templates;

[0027] The final detection result determination module is also used to input the medical document to be detected and the medical prompt word template into the large model to obtain the final detection result; wherein, the final detection result includes the final predicted document type and the final predicted privacy entity.

[0028] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0029] At least one processor; and

[0030] A memory communicatively connected to the at least one processor; wherein,

[0031] The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to execute the training method of the content detection model, or the fine-tuning method of the content detection model, or the content detection method according to any embodiment of the present invention.

[0032] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions, the computer instructions being configured to cause a processor to execute and implement a training method for a content detection model, a fine-tuning method for a content detection model, or a content detection method according to any embodiment of the present invention.

[0033] According to another aspect of the present invention, a computer program product is provided, the computer program product comprising a computer program that, when executed by a processor, implements a training method for a content detection model, a fine-tuning method for a content detection model, or a content detection method according to any embodiment of the present invention.

[0034] The technical solution of this invention involves acquiring sample medical documents and sample tag data; training an initial detection model using the sample medical documents and sample tag data to obtain a basic content detection model; wherein the initial detection model is a BERT model; and optimizing the basic content detection model based on a medical knowledge base to obtain an optimized content detection model. This technical solution improves the efficiency and accuracy of medical document content recognition.

[0035] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0036] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0037] Figure 1 This is a flowchart of a training method for a content detection model according to an embodiment of the present invention;

[0038] Figure 2 This is a flowchart of a training method for a content detection model according to an embodiment of the present invention;

[0039] Figure 3 This is a flowchart of a fine-tuning method for a content detection model according to an embodiment of the present invention;

[0040] Figure 4 This is a flowchart of a content detection method provided according to an embodiment of the present invention;

[0041] Figure 5 This is a schematic diagram of the structure of a training device for a content detection model according to an embodiment of the present invention;

[0042] Figure 6 This is a schematic diagram of the structure of a fine-tuning device for a content detection model according to an embodiment of the present invention;

[0043] Figure 7 This is a schematic diagram of the structure of a content detection device according to an embodiment of the present invention;

[0044] Figure 8 This is a schematic diagram of the structure of an electronic device that implements the training method, fine-tuning method, or content detection method of the content detection model in the embodiments of the present invention. Detailed Implementation

[0045] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0046] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0047] Furthermore, it should be noted that the collection, storage, use, processing, transmission, provision, and disclosure of medical documents and other related data involved in the technical solution of this invention all comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0048] In the healthcare industry, the accurate extraction of patient privacy information (patient ID / personnel ID, name, ID number, mobile phone number) and the identification of the nature of medical documents (such as lab reports, inpatient medical records, etc.) are core requirements for compliant management and business flow of cross-border data transmission. Currently, there are three major pain points: diverse formats of privacy entities (e.g., patient IDs may be "inpatient number + 6 digits" or "outpatient ID + letters"), confusing file types (e.g., both lab reports and imaging reports contain the keyword "diagnosis"), and interference from non-medical documents (e.g., administrative notices mixed into documents awaiting processing). This invention achieves the dual objectives of "privacy information extraction + medical document identification" with 5-10 labeled samples through "entity-classification dual-task collaborative meta-learning + multi-dimensional thinking chain prompts," meeting the high-precision requirements of small-sample scenarios.

[0049] Figure 1 This is a flowchart illustrating a training method for a content detection model according to an embodiment of the present invention. This embodiment is applicable to situations involving cross-border transmission of medical documents under conditions of small sample size or zero sample size, addressing how to train a content detection model to achieve content detection of medical documents. This method can be executed by a content detection model training device, which can be implemented in hardware and / or software. This device can be configured in an electronic device that carries the training function of the content detection model, such as a server. Figure 1 As shown, the method includes:

[0050] S110. Obtain sample medical documents and sample label data.

[0051] In this embodiment, the sample medical documents refer to medical documents and non-medical documents in different languages obtained from the Internet, including but not limited to the text in scanned test reports, handwritten case records, printed hospitalization case records, etc. The sample label data refers to the label data of the sample medical documents, including but not limited to the document type and the real privacy entities in the medical documents. It should be noted that the number of sample medical documents is small, usually including 5-8 labeled samples in each type of document.

[0052] An optional method for obtaining sample medical documents includes: performing text recognition on the original medical document based on OCR technology to obtain the recognized text; correcting the recognized text based on the term matching knowledge base to obtain the corrected text; performing format verification on the corrected text to obtain the parsable text; performing desensitization processing on the parsable text to obtain the desensitized text; and automatically segmenting the desensitized text according to the medical document structure to obtain the sample medical documents.

[0053] Among them, the original medical document refers to an unprocessed medical document, including but not limited to pictures.

[0054] Specifically, use OCR technology to perform text recognition on the original medical document to obtain the recognized text, and correct the recognized text based on the term matching library to obtain the corrected text. For example, "cardiospasm" is corrected to "achalasia of cardia". Then perform format verification on the corrected text, such as "JY" prefix rule verification, to obtain the parsable text. Furthermore, perform desensitization processing on the privacy information in the parsable text, such as ID card numbers and mobile phone numbers, through masking processing to obtain the desensitized text, such as "110101XXXX". Finally, automatically segment the desensitized text according to the medical document structure to obtain the sample medical documents. For example, automatically segment according to the "patient information column", "diagnosis content column", and "test result column" to obtain the sample medical document associated with the keywords "name / ID / ID card number" in the "patient information column".

[0055] It can be understood that by using OCR technology, the text information in medical documents containing pictures can be accurately recognized, and at the same time, by performing desensitization processing on the privacy information in the parsable text, privacy leakage during model training and inference can be avoided.

[0056] S120. Train the initial detection model using the sample medical documents and the sample label data to obtain the basic content detection model.

[0057] In this embodiment, the initial detection model is the BERT model; optionally, it includes a pre-training layer, a classification branch, and an entity branch. Among them, the pre-training layer is used for feature extraction; the classification branch is used for identifying the document type; the entity branch is used for identifying the privacy entities in the document. The so-called basic content detection model refers to the content detection model obtained by iteratively training the initial detection model.

[0058] An alternative approach is to input sample medical documents into an initial detection model to obtain sample prediction results. The sample prediction results include the predicted document type and the predicted privacy entity. Based on the predicted document type and the document tags in the sample tag data, a classification loss is determined. Based on the predicted privacy entity and the entity tags in the sample tag data, an entity loss is determined. Based on the classification loss and the entity loss, a training loss is determined. The initial detection model is iteratively trained using the training loss to obtain a basic content detection model.

[0059] The classification loss optimizes the accuracy of the classification branch; the entity loss measures the overall error of the entity branch, ensuring that it accurately classifies entities and identifies synonyms. The training loss prioritizes privacy-preserving entity extraction (e.g., core medical compliance) while maintaining document classification accuracy.

[0060] Specifically, sample medical documents are input into the initial detection model to obtain sample prediction results. These prediction results include predicted document types and predicted privacy entities. Based on the cross-entropy loss function, a classification loss is determined according to the predicted document types and document tags in the sample label data. An entity loss is determined according to the predicted privacy entities and entity tags in the sample label data. The classification loss and entity loss are weighted and summed to determine the training loss; for example, the entity loss weight is 0.6, and the classification loss weight is 0.4. A higher entity loss weight better reflects the core user needs. The initial detection model is iteratively trained using the training loss until a training stopping condition is met, at which point training stops, resulting in a basic content detection model. It should be noted that the training stopping condition includes reaching a set number of iterations or the training loss stabilizing within a set range. The set number of iterations and the set range can be set by those skilled in the art based on actual circumstances.

[0061] Understandably, joint training with two branches can achieve collaborative optimization of entity extraction and document classification.

[0062] It should be noted that the entity loss weight of 0.6 and the classification loss weight of 0.4 can be verified through multiple sets of experiments, as shown in the table below:

[0063]

[0064] For example, determining entity loss based on the sample predicted privacy entity and the entity label in the sample label data includes: determining entity recognition loss based on the sample predicted privacy entity and the entity label in the label data; determining synonym comparison loss based on the sample predicted privacy entity and the temperature coefficient; and determining entity loss based on entity recognition loss, synonym comparison loss, and synonym entity weight.

[0065] The entity recognition loss optimizes the accuracy of entity classification in entity branches. The same entity contrast loss forces the aggregation of synonymous entity features, such as "patient ID" and "hospitalization number," and the separation of dissimilar entity features, such as "ID" and "name." The synonymous entity weight balances the classification and contrast losses; a larger weight emphasizes the differentiation of synonymous entities, preventing "hospitalization number" from being misclassified as "doctor ID." Its value range is [0, 1.2], with an optimal value of 0.8-0.9 in medical document content recognition scenarios. The temperature coefficient controls feature discrimination; a smaller coefficient amplifies feature differences more significantly, preventing the model from confusing "ICD code" with "patient ID." Its value range is [0.08, 0.2], with an optimal value of 0.1-0.12 in medical document content recognition scenarios.

[0066] Specifically, based on the cross-entropy loss function, the entity recognition loss is determined according to the sample predicted privacy entities and the entity labels in the label data. The synonymous similarity between the sample predicted privacy entities is calculated; for example, the feature vectors of the sample predicted privacy entities can be determined, and the cosine similarity between the feature vectors can be calculated. Then, based on the pre-defined similarity and temperature coefficient, the synonymous entity comparison loss is determined, for example, through the following formula:

[0067] ;in, Predict features for privacy entity i, such as the "ID" feature, for the sample. Features for synonyms of entity i, such as the feature "hospital number"; τ is the temperature coefficient; Predict the privacy entity for the j-th sample.

[0068] Then, the result of multiplying the synonym entity comparison loss and the synonym entity weight is added to the entity recognition loss to obtain the entity loss.

[0069] Understandably, by introducing entity recognition loss and synonym comparison loss to determine entity loss, entity branches can better identify entities and synonyms.

[0070] S130. Based on the medical knowledge base, the basic content detection model is optimized to obtain the optimized content detection model.

[0071] In this embodiment, the optimized content detection model refers to the detection model after optimizing the basic content detection model. The so-called medical knowledge base refers to a knowledge base containing a large amount of medical-related information; optionally, the medical knowledge base includes publicly available medical knowledge graphs and private hospital-related document data, as well as medical rules; the medical knowledge base can be automatically updated periodically.

[0072] Specifically, real feature vectors and entity prefix feature vectors are extracted from the medical knowledge base, and the basic content detection model is optimized and trained to obtain an optimized content detection model.

[0073] The technical solution of this invention involves acquiring sample medical documents and sample tag data; training an initial detection model using the sample medical documents and sample tag data to obtain a basic content detection model; wherein the initial detection model is a BERT model; and optimizing the basic content detection model based on a medical knowledge base to obtain an optimized content detection model. This technical solution improves the efficiency and accuracy of medical document content recognition.

[0074] Figure 2 This is a flowchart illustrating a training method for a content detection model according to an embodiment of the present invention. Based on the above embodiments, this embodiment further optimizes the process of "optimizing the basic content detection model based on a medical knowledge base to obtain an optimized content detection model," providing an optional implementation scheme. For example... Figure 2 As shown, the method includes:

[0075] S210. Obtain sample medical documents and sample label data.

[0076] S220. The initial detection model is trained using sample medical documents and sample label data to obtain the basic content detection model.

[0077] The initial detection model was the BERT model.

[0078] S230. Based on the medical knowledge base, the basic content detection model is optimized to obtain an optimized content detection model.

[0079] An alternative approach involves optimizing a basic content detection model based on a medical knowledge base to obtain an optimized content detection model. This includes: extracting entity feature vectors and entity prefix feature vectors from the medical knowledge base; determining the entity vector similarity between the entity feature vectors and the entity prefix feature vectors; using the absolute value of the difference between the entity vector similarity and the similarity target value as the distillation loss; determining the optimization loss based on the training loss, the distillation loss, and the distillation loss weights; and iteratively training the basic content detection model using the optimization loss to obtain the optimized content detection model.

[0080] The distillation loss is used to force entity features to conform to medical domain rules, reducing the formatting error rate; its value ranges from [0, ∞). The distillation loss weight is used to determine the constraint strength of the rule on the model, and its value ranges from [0.5, 0.6]. The larger the distillation loss weight, the stricter the constraint of the rule on the entity features. The optimization loss is used to force entity features to conform to medical domain rules, reducing the formatting error rate. The similarity target value can be set based on the actual situation to ensure that the entity conforms to the formatting compliance, and its value ranges from [0.7, 0.9], with the optimal value for the medical scenario being 0.7-0.8.

[0081] Specifically, BERT encodings of entity feature vectors (e.g., “JY250903”, dimension 512) and entity prefix feature vectors (e.g., “JY”, dimension 512) are extracted from the medical knowledge base. The entity vector similarity between the entity feature vectors and the entity prefix feature vectors is determined. The absolute value of the difference between the entity vector similarity and the target similarity value is used as the distillation loss. The result of multiplying the distillation loss by its weights is added to the training loss to obtain the optimization loss. The optimization loss is used to iteratively train the basic content detection model until the training termination condition is met, resulting in the optimized content detection model. The training termination condition is that the average entity extraction score is greater than a first set value (e.g., 88%), the document classification accuracy is greater than a second set value (e.g., 90%), and there is no improvement after a set number of consecutive rounds. The set number of rounds can be set according to actual needs, for example, 5 rounds.

[0082] Understandably, by introducing a medical knowledge base for knowledge distillation constraints, it can be ensured that the optimized content detection model can ensure that the document content meets the format requirements.

[0083] As another optional aspect of the present invention, the iterative training process further includes: for each iteration, dynamically adjusting the distillation loss weight and temperature coefficient corresponding to the iteration based on the classification confidence score output by the model in the previous iteration.

[0084] Specifically, for each iteration, if the classification confidence score of the model output in the previous iteration is greater than the third set value, such as 0.7, the temperature coefficient is automatically lowered, for example, from 0.12 to 0.1; and the distillation loss weight is increased, for example, from 0.5 to 0.6; otherwise, no adjustment is made.

[0085] Understandably, the performance of the content detection model can be improved by dynamically adjusting the distillation loss weights and temperature coefficients during iterative training rounds.

[0086] The technical solution of this invention involves acquiring sample medical documents and sample tag data; training an initial detection model using the sample medical documents and sample tag data to obtain a basic content detection model; wherein the initial detection model is a BERT model; and optimizing the basic content detection model based on a medical knowledge base to obtain an optimized content detection model. This technical solution improves the efficiency and accuracy of medical document content recognition.

[0087] Figure 3 This is a flowchart illustrating a method for fine-tuning a content detection model according to an embodiment of the present invention. This embodiment is applicable to situations involving cross-border transmission of medical documents under conditions of small or zero samples, addressing how to train a content detection model to achieve content detection in medical documents. This method can be executed by a content detection model fine-tuning device, which can be implemented in hardware and / or software. This device can be configured in an electronic device that carries the fine-tuning function of the content detection model, such as a server. Figure 3 As shown, the method includes:

[0088] S310. Obtain labeled medical documents.

[0089] In this embodiment, labeled medical documents refer to a small number of labeled medical and non-medical documents from the target hospital, numbering 5-8. For example, a medical document (inpatient medical record): "Inpatient number (patient ID): ZY20250901, Name: Zhang San, ID number: 110101XXXX, Mobile number: 138XXXX1234, Admission diagnosis: Pneumonia" → Entity labeled + Classified as "Inpatient Medical Record"; a non-medical document (administrative notice): "Notice on equipment maintenance, Contact person: Li Si, Telephone: 139XXXX5678" → No privacy entity labeled + Classified as "Non-medical document".

[0090] S320. Fine-tune the optimized content detection model using labeled medical documents to obtain the target content detection model.

[0091] The optimized content detection model is trained using the content detection model training method provided in the above embodiments of the present invention. The target content detection model refers to the finally trained model used for actual content detection.

[0092] An optional approach involves fine-tuning an optimized content detection model using labeled medical documents to obtain a target content detection model. This includes: inputting labeled medical documents into the optimized content detection model to obtain target prediction results; the target prediction results include the target predicted document type and the target predicted privacy entity; constructing a fine-tuning loss based on the target prediction results and the label data corresponding to the labeled sample documents; determining parameter regularization constraint values ​​based on the model parameters and parameter complexity coefficients of the optimized content detection model; determining an adaptation loss based on the parameter regularization constraint values ​​and the fine-tuning loss; and fine-tuning the optimized content detection model based on the adaptation loss to obtain the target content detection model.

[0093] The fine-tuning loss is used to allow the model to quickly adapt to the document format of the target hospital. It should be noted that the fine-tuning loss is constructed in the same way as the training loss, and will not be elaborated here. The adaptation loss is used to enable the optimized content detection model to adapt quickly. The parameter complexity coefficient, also known as the L2 regularization coefficient, is used to control the model parameter complexity and avoid overfitting when adapting to small samples.

[0094] Specifically, labeled medical documents are input into the optimized content detection model to obtain target prediction results. These results include the predicted document type and the predicted privacy entity. A fine-tuning loss is constructed based on the target prediction results and the corresponding label data of the labeled sample documents. Then, the L2 norm of the optimized content detection model's parameters is calculated. The L2 norm is multiplied by the parameter complexity coefficient to obtain the parameter regularization constraint value. This regularization constraint value is then added to the fine-tuning loss to obtain the adaptation loss. The adaptation loss is used to fine-tune the optimized content detection model until the fine-tuning training stopping condition is met, at which point training stops, and the target content detection model is obtained. It should be noted that the fine-tuning training stopping condition is the same as the training stopping condition in the above embodiment, or the training stopping condition can be fine-tuned (e.g., by adjusting accuracy and iteration rounds) to obtain the fine-tuning training stopping condition.

[0095] Understandably, by fine-tuning the optimized content detection model using a small number of standard medical documents from the target hospitals, the target content detection model can be quickly adapted to the target hospitals for medical document content detection.

[0096] As another optional aspect of the present invention, the fine-tuning process further includes: employing gradient descent for fine-tuning; and dynamically adjusting the learning rate during gradient descent.

[0097] Specifically, during the fine-tuning process, gradient descent is used for iterative training. In the first iteration, the learning rate is set to 0.002, and in the second iteration, the learning rate is increased, for example, to 0.004.

[0098] Understandably, increasing the learning rate can help avoid overfitting on small samples.

[0099] As another optional aspect of the present invention, the fine-tuning process further includes: during the fine-tuning process, determining that the pre-training layer parameters of the optimized content detection model remain unchanged, and adjusting the entity branch, classification branch, and dynamic parameter mapping layer of the optimized content detection model; wherein, the dynamic parameter mapping layer includes distillation loss weights and temperature coefficients.

[0100] Specifically, during the fine-tuning process, the parameters of the pre-trained layer of the optimized content detection model are kept unchanged, and the fine-tuning only adjusts the entity branch, classification branch, and dynamic parameter mapping layer of the optimized content detection model.

[0101] It is understandable that freezing the parameters of the key layers can prevent the destruction of the underlying features and prevent overfitting.

[0102] The results of the experiments to verify the overfitting prevention are shown in the table below:

[0103]

[0104] The technical solution provided in this invention obtains a target content detection model by acquiring labeled medical documents and then fine-tuning the optimized content detection model using these labeled documents. This method, which uses labeled medical documents to fine-tune the optimized content detection model, improves the accuracy of content detection in medical documents.

[0105] Figure 4 This is a flowchart of a content detection method according to an embodiment of the present invention. This embodiment is applicable to situations involving cross-border transmission of medical documents under conditions of small sample size or zero sample size, addressing how to train a content detection model to achieve content detection of medical documents. This method can be executed by a content detection device, which can be implemented in hardware and / or software. This device can be configured in an electronic device carrying content detection functionality, such as a server. Figure 4 As shown, the method includes:

[0106] S410. Obtain the medical document to be tested.

[0107] In this embodiment, the medical document to be detected refers to a medical document that needs to be detected in real time.

[0108] S420. Input the medical document to be detected into the target content detection model to obtain the final detection result.

[0109] The target content detection model is fine-tuned according to the fine-tuning method of the content detection model provided in any of the above embodiments. The so-called final detection result includes the final predicted document type and the final predicted privacy entity.

[0110] Specifically, the medical document to be detected is input into the target content detection model, processed by the model, and the final detection result is output.

[0111] S430, Construct a medical prompt template.

[0112] S440. Input the medical document to be detected and the medical prompt word template into the large model to obtain the final detection result.

[0113] It should be noted that in the case of zero samples, that is, when the target content detection model has a low recognition rate, the S430-S440 method can be used for medical document content detection.

[0114] Specifically, based on the prompt word engineering, a medical prompt word template is built, which may include the following:

[0115] 1. Medical document identification: Check whether the document contains medical keywords such as "Patient ID", "Diagnosis", "LabTest", and "Medical Order". If the classification confidence score is ≥0.85, it is a medical document; otherwise, it is a non-medical document.

[0116] 2. Medical record breakdown:

[0117] - Including "LabItems", "Reference Range", and "Test Date" → categorized as "LabReport";

[0118] - Including "Admission Diagnosis", "Hospitalization ID", and "Medical Record" → Classified as "Hospitalization Record";

[0119] -Including "CT / MRI", "ImagingDiagnosis", and "FilmNumber" → categorized as "ImagingReport";

[0120] - Including "Dosage and Usage", "Daily Frequency", and "Medication Duration" → categorized as "Medical Order";

[0121] 3. Entity extraction and location: In medical documents, locate the "Patient Information" or "Personal Information" sections or paragraphs containing keywords such as "name / ID / ID card / telephone", excluding irrelevant paragraphs such as "Diagnosis Content" and "Test Results";

[0122] 4. Privacy Entity Filtering:

[0123] -Patient ID: Extract alphanumeric combinations containing prefixes such as "ZY / JY / outpatient" (e.g., ZY20250901, JY250820).

[0124] - Name: Extract 2-4 Chinese characters after "Name:" (excluding "Doctor's Name" and "Contact Person's Name");

[0125] - ID card number: Extract 18 digits / letters (the last digit can be X, excluding 11-digit mobile phone numbers);

[0126] -Mobile number: Extract 11 consecutive digits (excluding the numeric string in the patient ID);

[0127] 5. Entity verification: Check whether the extracted ID number conforms to the rule of "first 6 digits address code + 8 digits birth date code", and whether the mobile phone number conforms to the rule of "starting with 13 / 14 / 15 / 17 / 18 / 19";

[0128] 6. Output Results: Non-medical files will output "Non-medical file, no extraction required"; medical files will output "File type: XXX, privacy entity: {ID:XXX, name:XXX, ID number:XXX, mobile number:XXX}".

[0129] Specifically, the medical documents to be detected and the medical prompt word templates are input into the large model, and the final detection results are obtained through model inference.

[0130] A concrete example, input document fragment 1 (lab report, medical document):

[0131] "Laboratory report number (patient ID): JY20250903, name: Li Si, gender: male, ID number: 3101011990XXXX, contact number: 135XXXX6789, test item: complete blood count, reference value: 4-10×10^9 / L."

[0132] An instantiated prompt template is automatically generated based on the medical prompt word template, as follows:

[0133] 1. Medical document identification: Documents containing "Laboratory Order Number (Patient ID)" and "Laboratory Items" with a classification confidence level of 0.92 are considered medical documents.

[0134] 2. Medical document breakdown: Includes "Test Items" and "Reference Values" → Classified as "Lab Report";

[0135] 3. Entity extraction and location: Locate the paragraph containing "Laboratory Report Number (Patient ID)", "Name", "ID Number", and "Contact Number";

[0136] 4. Privacy Entity Filtering:

[0137] - Patient ID: Extract “JY20250903” (containing the “JY” prefix, conforming to the lab report ID format);

[0138] - Name: Extract "Li Si" (the last two Chinese characters after "Name:", not doctor information);

[0139] - ID number: Extract “3101011990XXXX” (18 digits, the first 6 digits “310101” is the Shanghai address code);

[0140] - Mobile phone number: Extract "135XXXX6789" (11 digits, starting with 135);

[0141] 5. Entity verification: The ID number conforms to the birthdate code rule (starting with 1990), and the mobile phone number conforms to the rule of starting with 135;

[0142] 6. Result Output: File Type: Laboratory Report, Privacy Entity: {ID:JY20250903, Name: Li Si, ID Number:3101011990XXXX, Mobile Number:135XXXX6789}.

[0143] Another concrete example is input document fragment 2 (administrative notice, non-medical document):

[0144] "Notice on Conducting Hospital Fire Safety Training, Date: 2025-09-10, Contact Person: Wang Wu, Phone: 136XXXX9012, Participants: Medical and nursing staff from all departments."

[0145] An instantiated prompt template is automatically generated based on the medical prompt word template, as follows:

[0146] 1. Medical document identification: Documents containing only "hospital" and "medical staff" but lacking core medical keywords such as "patient ID" and "diagnosis" have a classification confidence level of 0.3 and are classified as non-medical documents.

[0147] 2. Medical document breakdown: No execution required;

[0148] 3. Entity extraction and location: No need to execute;

[0149] 4. Privacy Entity Filtering: No need to perform this step;

[0150] 5. Entity verification: No need to perform;

[0151] 6. Output results: Non-medical files, no extraction required.

[0152] The technical solution provided in this invention obtains the medical document to be detected, inputs the document into a target content detection model, and obtains the final detection result; or it constructs a medical prompt word template, inputs the medical document to be detected and the prompt word template into a large model, and obtains the final detection result. Through the above technical solutions, the detection efficiency and accuracy of medical document content can be improved.

[0153] The technical solutions of this invention include: 1) Improved accuracy and efficiency: Through small-sample learning and adaptive optimization, the system significantly reduces its reliance on labeled data while maintaining high accuracy. In actual testing, the F1 score reaches 0.91, which is about 30% higher than traditional regular expression-based methods and about 20% higher than supervised learning-based methods in small-sample scenarios. 2) Adaptability: Traditional systems require regular manual updates to the rule base (usually quarterly or monthly), while this system can adapt to new data patterns and threat types in real time, greatly reducing maintenance costs and response latency. 3) Cross-domain generalization ability: Through a meta-learning framework, the system can quickly adapt to the specific needs of different industries and fields, such as customer information protection in the financial industry, medical record confidentiality in the medical industry, and intellectual property protection in the manufacturing industry. In cross-domain testing, the system achieves an F1 score of over 0.85 with only 10 samples, demonstrating strong generalization ability. 4) Explanatory power and credibility: The system provides confidence scores and decision explanations to help administrators understand the basis of content recognition, enhancing the credibility and transparency of the system. This is particularly important for compliance audits and troubleshooting.

[0154] Figure 5 This is a schematic diagram of a training device for a content detection model according to an embodiment of the present invention. This embodiment is applicable to situations involving cross-border transmission of medical documents under conditions of small sample size or zero sample size, addressing how to train a content detection model to achieve content detection of medical documents. The training device for the content detection model can be implemented in hardware and / or software, and can be configured in an electronic device that carries the training function of the content detection model, such as a server. Figure 5 As shown, the device includes:

[0155] The sample data acquisition module 510 is used to acquire sample medical documents and sample label data;

[0156] The basic detection model determination module 520 is used to train the initial detection model using sample medical documents and sample label data to obtain the basic content detection model; wherein, the initial detection model is the BERT model;

[0157] The optimized content detection module 530 is used to optimize the basic content detection model based on the medical knowledge base, resulting in an optimized content detection model.

[0158] The technical solution of this invention involves acquiring sample medical documents and sample tag data; training an initial detection model using the sample medical documents and sample tag data to obtain a basic content detection model; wherein the initial detection model is a BERT model; and optimizing the basic content detection model based on a medical knowledge base to obtain an optimized content detection model. This technical solution improves the efficiency and accuracy of medical document content recognition.

[0159] Optionally, the basic detection model determination module 520 includes:

[0160] The sample prediction result determination unit is used to input the sample medical document into the initial detection model to obtain the sample prediction result; the sample prediction result includes the sample predicted document type and the sample predicted privacy entity.

[0161] The classification loss determination unit is used to determine the classification loss based on the predicted document type of the sample and the document labels in the sample label data;

[0162] The entity loss determination unit is used to determine the entity loss based on the sample predicted privacy entity and the entity label in the sample label data;

[0163] The training loss determination unit is used to determine the training loss based on the classification loss and the entity loss.

[0164] The model training unit is used to iteratively train the initial detection model using training loss to obtain the basic content detection model.

[0165] Optionally, the entity loss determination unit is used for:

[0166] Based on the sample prediction of privacy entities and entity labels in the label data, determine the entity recognition loss;

[0167] Based on the sample prediction of privacy entities and temperature coefficients, determine the synonym contrast loss;

[0168] The entity loss is determined based on the entity recognition loss, the synonym entity comparison loss, and the synonym entity weight.

[0169] Optionally, the optimized content detection module 530 is specifically used for:

[0170] Extract entity feature vectors and entity prefix feature vectors from the medical knowledge base;

[0171] Determine the entity vector similarity between the entity feature vector and the entity prefix feature vector;

[0172] The absolute value of the difference between the entity vector similarity and the target similarity value is used as the distillation loss;

[0173] The optimization loss is determined based on the training loss, distillation loss, and distillation loss weights.

[0174] An optimized content detection model is obtained by iteratively training the basic content detection model using an optimization loss.

[0175] Optionally, the device also includes a dynamic parameter adjustment module for:

[0176] During the iterative training process, for each iteration, the distillation loss weight and temperature coefficient corresponding to that iteration are dynamically adjusted based on the classification confidence of the model output in the previous iteration.

[0177] Optionally, the sample data acquisition module 510 is specifically used for:

[0178] The original medical document is used to perform text recognition based on OCR technology to obtain the recognized text.

[0179] Based on a terminology matching knowledge base, the identified text is corrected to obtain the corrected text;

[0180] Perform format validation on the corrected text to obtain parsable text;

[0181] Desensitize the parsable text to obtain desensitized text;

[0182] Based on the medical document structure, the anonymized text is automatically segmented to obtain sample medical documents.

[0183] The training apparatus for the content detection model provided in this embodiment of the invention can execute the training method for the content detection model provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0184] Figure 6 This is a schematic diagram of a content detection model fine-tuning device according to an embodiment of the present invention. This embodiment is applicable to situations involving cross-border transmission of medical documents under conditions of small or zero samples, addressing how to train a content detection model to achieve content detection of medical documents. The content detection model fine-tuning device can be implemented in hardware and / or software, and can be configured in an electronic device that carries the content detection model fine-tuning function, such as a server. Figure 6 As shown, the device includes:

[0185] The labeled document acquisition module 610 is used to acquire labeled medical documents;

[0186] The target detection model fine-tuning module 620 is used to fine-tune the optimized content detection model using labeled medical documents to obtain the target content detection model; wherein, the optimized content detection model is trained by the training method of the content detection model provided in any embodiment of the present invention.

[0187] The technical solution provided in this invention obtains a target content detection model by acquiring labeled medical documents and then fine-tuning the optimized content detection model using these labeled documents. This method, which uses labeled medical documents to fine-tune the optimized content detection model, improves the accuracy of content detection in medical documents.

[0188] Optionally, the target detection model fine-tuning module 620 is specifically used for:

[0189] The labeled medical documents are input into the optimized content detection model to obtain the target prediction results; the target prediction results include the target predicted document type and the target predicted privacy entity.

[0190] Based on the target prediction results and the label data corresponding to the labeled sample documents, a fine-tuning loss is constructed;

[0191] Based on the model parameters and parameter complexity coefficients of the optimized content detection model, determine the parameter regularization constraint values;

[0192] The adaptation loss is determined based on the parameter regularization constraint value and the fine-tuning loss;

[0193] The optimized content detection model is fine-tuned based on the adaptation loss to obtain the target content detection model.

[0194] Optionally, the device also includes a learning rate adjustment module for:

[0195] During the fine-tuning process, gradient descent is used.

[0196] The learning rate is dynamically adjusted during gradient descent.

[0197] Optionally, the device also includes a learning rate adjustment module for:

[0198] During the fine-tuning process, the pre-training layer parameters of the optimized content detection model are kept unchanged, while the entity branch, classification branch, and dynamic parameter mapping layer of the optimized content detection model are adjusted; among them, the dynamic parameter mapping layer includes distillation loss weights and temperature coefficients.

[0199] The fine-tuning device for the content detection model provided in this embodiment of the invention can execute the fine-tuning method for the content detection model provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0200] Figure 7This is a schematic diagram of a content detection device according to an embodiment of the present invention. This embodiment is applicable to situations involving cross-border transmission of medical documents under conditions of small sample or zero sample size, addressing how to train a content detection model to achieve content detection of medical documents. The content detection device can be implemented in hardware and / or software, and can be configured in an electronic device carrying content detection functionality, such as a server. Figure 7 As shown, the device includes:

[0201] The document acquisition module 710 is used to acquire medical documents to be detected.

[0202] The final detection result determination module 720 is used to input the medical document to be detected into the target content detection model to obtain the final detection result; wherein, the target content detection model is fine-tuned by the fine-tuning method of the content detection model provided in any embodiment of the present invention.

[0203] The prompt word template building module 730 is used to build medical prompt word templates;

[0204] The final detection result determination module 720 is also used to input the medical document to be detected and the medical prompt word template into the large model to obtain the final detection result; wherein, the final detection result includes the final predicted document type and the final predicted privacy entity.

[0205] The technical solution provided in this invention obtains the medical document to be detected, inputs the document into a target content detection model, and obtains the final detection result; or it constructs a medical prompt word template, inputs the medical document to be detected and the prompt word template into a large model, and obtains the final detection result. Through the above technical solutions, the detection efficiency and accuracy of medical document content can be improved.

[0206] The content detection device provided in this embodiment of the invention can execute the content detection method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the method.

[0207] According to embodiments of the present invention, the present invention also provides an electronic device, a readable storage medium, and a computer program product.

[0208] Figure 8 This is a schematic diagram of the structure of an electronic device that implements the training method, fine-tuning method, or content detection method of the content detection model in the embodiments of the present invention. Figure 8A schematic diagram of an electronic device 10, which can be used to implement embodiments of the present invention, is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0209] like Figure 8 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0210] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0211] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as training methods for content detection models or fine-tuning methods for content detection models or content detection methods.

[0212] In some embodiments, the method for training or fine-tuning a content detection model, or the content detection method, may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the method for training or fine-tuning a content detection model, or the content detection method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the method for training or fine-tuning a content detection model, or the content detection method, by any other suitable means (e.g., by means of firmware).

[0213] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0214] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0215] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0216] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0217] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0218] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0219] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0220] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A training method for a content detection model, characterized in that, include: Obtain sample medical documents and sample label data; The initial detection model is trained using the sample medical documents and the sample tag data to obtain a basic content detection model, including: inputting the sample medical documents into the initial detection model to obtain sample prediction results; the sample prediction results include sample predicted document type and sample predicted privacy entity; The classification loss is determined based on the predicted document type of the sample and the document tags in the sample tag data; Determining entity loss based on the predicted privacy entities from the samples and the entity labels in the sample label data includes: determining entity recognition loss based on the predicted privacy entities from the samples and the entity labels in the label data; Based on the sample, predict privacy entities and temperature coefficients, and determine the synonym comparison loss; The entity loss is determined based on the entity recognition loss, the synonym entity comparison loss, and the synonym entity weights. The training loss is determined based on the classification loss and the entity loss. The initial detection model is iteratively trained using the training loss to obtain a basic content detection model; wherein, the initial detection model is a BERT model; Based on a medical knowledge base, the basic content detection model is optimized to obtain an optimized content detection model.

2. The method according to claim 1, characterized in that, Based on a medical knowledge base, the basic content detection model is optimized to obtain an optimized content detection model, including: Extract entity feature vectors and entity prefix feature vectors from the medical knowledge base; Determine the entity vector similarity between the entity feature vector and the entity prefix feature vector; The absolute value of the difference between the entity vector similarity and the target similarity value is used as the distillation loss; The optimization loss is determined based on the training loss, the distillation loss, and the distillation loss weights. The optimized content detection model is obtained by iteratively training the basic content detection model using the optimized loss.

3. The method according to claim 2, characterized in that, The iterative training process also includes: For each iteration, the distillation loss weight and temperature coefficient are dynamically adjusted based on the classification confidence score output by the model in the previous iteration.

4. The method according to claim 1, characterized in that, Obtain sample medical documents, including: The original medical document is used to perform text recognition based on OCR technology to obtain the recognized text. Based on a terminology matching knowledge base, the identified text is corrected to obtain the corrected text; The corrected text is formatted and validated to obtain parsable text; The parsable text is then de-identified to obtain de-identified text; Based on the medical document structure, the desensitized text is automatically segmented to obtain sample medical documents.

5. A method for fine-tuning a content detection model, characterized in that, include: Obtain labeled medical documents; The target content detection model is obtained by fine-tuning the annotated medical documents to obtain the optimized content detection model; wherein the optimized content detection model is trained by the training method of the content detection model according to any one of claims 1-4.

6. The method according to claim 5, characterized in that, The target content detection model is obtained by fine-tuning the optimized content detection model using the annotated medical documents, including: The labeled medical document is input into the optimized content detection model to obtain the target prediction result; the target prediction result includes the target predicted document type and the target predicted privacy entity. Based on the target prediction results and the label data corresponding to the labeled sample documents, a fine-tuning loss is constructed; Based on the model parameters and parameter complexity coefficients of the optimized content detection model, determine the parameter regularization constraint values; The adaptation loss is determined based on the regularization constraint value of the parameter and the fine-tuning loss; The optimized content detection model is fine-tuned based on the adaptation loss to obtain the target content detection model.

7. The method according to claim 5 or 6, characterized in that, The fine-tuning process also includes: Fine-tuning uses gradient descent. The learning rate is dynamically adjusted during gradient descent.

8. The method according to claim 5 or 6, characterized in that, The fine-tuning process also includes: The pre-training layer parameters of the optimized content detection model are kept unchanged, and the entity branch, classification branch, and dynamic parameter mapping layer of the optimized content detection model are adjusted; wherein, the dynamic parameter mapping layer includes distillation loss weights and temperature coefficients.

9. A content detection method, characterized in that, include: Retrieve the medical documents to be tested; The medical document to be detected is input into the target content detection model to obtain the final detection result; wherein the target content detection model is fine-tuned using the fine-tuning method of the content detection model according to any one of claims 5-8.

10. A training device for a content detection model, characterized in that, include: The sample data acquisition module is used to acquire sample medical documents and sample label data; The basic detection model determination module is used to train an initial detection model using the sample medical document and the sample tag data to obtain a basic content detection model, including: inputting the sample medical document into the initial detection model to obtain sample prediction results; the sample prediction results include the sample predicted document type and the sample predicted privacy entity; The classification loss is determined based on the predicted document type of the sample and the document tags in the sample tag data; Based on the predicted privacy entities from the samples and the entity labels in the sample label data, the entity loss is determined, including: Based on the predicted privacy entities from the samples and the entity labels in the labeled data, the entity recognition loss is determined. Based on the sample, predict privacy entities and temperature coefficients, and determine the synonym comparison loss; The entity loss is determined based on the entity recognition loss, the synonym entity comparison loss, and the synonym entity weights. The training loss is determined based on the classification loss and the entity loss. The initial detection model is iteratively trained using the training loss to obtain a basic content detection model; wherein, the initial detection model is a BERT model; An optimized content detection module is used to optimize the basic content detection model based on a medical knowledge base, resulting in an optimized content detection model.

11. A fine-tuning device for a content detection model, characterized in that, include: The labeled document acquisition module is used to acquire labeled medical documents; The target detection model fine-tuning module is used to fine-tune the optimized content detection model using the labeled medical documents to obtain the target content detection model; wherein the optimized content detection model is trained using the training method for the content detection model according to any one of claims 1-4.

12. A content detection device, characterized in that, include: The document acquisition module is used to acquire medical documents to be inspected. The final detection result determination module is used to input the medical document to be detected into the target content detection model to obtain the final detection result; wherein, the target content detection model is fine-tuned by the fine-tuning method of the content detection model according to any one of claims 5-8.

13. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the training method of the content detection model according to any one of claims 1-4, or the fine-tuning method of the content detection model according to any one of claims 5-8, or the content detection method according to claim 9.

14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute and implement the training method of the content detection model according to any one of claims 1-4, the fine-tuning method of the content detection model according to any one of claims 5-8, or the content detection method according to claim 9.

15. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the training method of the content detection model according to any one of claims 1-4, the fine-tuning method of the content detection model according to any one of claims 5-8, or the content detection method according to claim 9.

Citation Information

Patent Citations

  • Medical entity labeling method and device, equipment and storage medium

    CN116705345A

  • Knowledge graph construction model training method, knowledge graph construction method, knowledge graph construction device and equipment

    CN118734948A