Short text matching method based on enhanced contrast learning and multi-task optimization

By using the ALBERT model with multi-source data, supervised contrastive learning, and character-level noise simulation, the semantic understanding and robustness issues of short text matching models in the medical field are solved, achieving high-precision semantic matching and noise tolerance, and is applicable to multiple vertical fields.

CN120911474APending Publication Date: 2025-11-07CHENGDU HARIT MEDICAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511098604.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-06
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing technologies in the medical field suffer from problems such as ambiguous semantic boundaries, obstacles to knowledge transfer, scarcity of high-quality supervised data, insufficient model robustness, and lack of multi-task collaborative optimization mechanisms. As a result, short text matching models have weak semantic understanding capabilities, are sensitive to interference from similar-looking characters, and have low training efficiency in professional fields.

Method used

We employ the ALBERT Chinese pre-trained model, combined with high-quality training sets from multiple sources, supervised contrastive learning, and masked language modeling tasks. Through generative hard negative examples and character-level noise simulation, we construct an end-to-end unified training process to enhance the model's semantic matching ability and noise robustness in the medical field.

Benefits of technology

It significantly improves the model's semantic discrimination ability and robustness in noisy scenarios in the medical field, enhances training efficiency and multi-task collaboration, and is suitable for applications in various vertical fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120911474A_ABST
    Figure CN120911474A_ABST
Patent Text Reader

Abstract

The invention discloses a short text matching method based on enhanced contrast learning and multi-task optimization, and belongs to the technical field of natural language processing. According to the method, a multi-source data set is constructed by fusing general field and medical field texts, a generative language model is adopted to generate a high-quality difficult case to form a training triple, and a character-level noise enhancement mechanism based on a medical similar character dictionary is introduced to simulate an OCR error; aLBERT is used as an encoder, supervised comparative learning and mask language modeling tasks are jointly optimized, multi-task cooperative training is achieved through dynamic gradient balance, and finally a short text matching model with high semantic discrimination ability and noise robustness is obtained. On a test set containing 50% OCR noise in the cardiovascular field, the accuracy rate of the method reaches 94.3% and is improved by 12.2% compared with that of a traditional BERT-base model.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of natural language processing, and particularly relates to a short text matching method based on enhanced contrast learning and multi-task optimization, which is especially suitable for semantic matching scenes in vertical fields with dense professional terms. BACKGROUND

[0002] Short text matching technology, as one of the core tasks in the field of natural language processing (NLP), is widely used in intelligent customer service, information retrieval, question and answer systems, and medical auxiliary diagnosis, etc. The goal of this task is to judge the semantic similarity or relevance between two short texts. In the medical field, accurate text matching can significantly improve the intelligent level of systems such as disease auxiliary decision-making, case archiving, and symptom recognition.

[0003] Traditional methods rely on artificial rules or shallow semantic models (such as TF-IDF model), which have limited generalization ability in cross-domain scenarios. While methods based on deep pre-training models (such as BERT) are effective in general fields, they are difficult to fully adapt to the problems of professional terms, noise interference, and data scarcity in vertical fields such as medicine. Therefore, in vertical fields with strong professionalism and data scarcity such as medicine, existing technologies still face many challenges: 1. Fuzzy semantic boundaries and knowledge transfer barriers: There are many professional terms in medical texts and the semantics are fine, such as the semantic differences between "angina", "troponin", and "atrial fibrillation", which are crucial for diagnosis and judgment. However, general pre-training models lack targeted knowledge and are difficult to accurately grasp the semantic boundaries of professional terms.

[0004] 2. Lack of high-quality supervised data: Building high-quality matching corpus in the medical field is costly, and existing research relies on general open-source datasets, which lack domain adaptability and are difficult to cover actual problems such as professional test index nouns in test report forms or colloquial expressions in medical conversations.

[0005] 3. Insufficient model robustness: In actual applications, medical texts are often obtained through OCR devices, and there are a lot of "muscle / limb" and "heart / essential" type of similar character misrecognition interference. In addition, there is a lot of unstructured colloquial communication between doctors and patients, with noise such as pinyin abbreviations and misspelled words in the text, resulting in a serious decline in the robustness of existing models in actual scenarios.

[0006] 4. Poor quality of difficult example construction: Contrastive learning (such as SimCSE) has become an important means to improve the quality of semantic representation, but mainstream methods mostly use unsupervised or random negative sample methods, which lack effective characterization of domain semantic boundaries, resulting in insufficient discrimination ability of the trained model.

[0007] 5. Lack of multi-task synergistic optimization mechanism: Most existing methods separate the text matching and language modeling tasks, and fail to effectively utilize the pre-training tasks such as MLM to assist in constraining the model, resulting in insufficient generalization ability and tolerance to local interference. SUMMARY

[0008] 1. Technical problem to be solved: In view of the problems in the prior art, the purpose of the present application is to provide a short text matching method based on enhanced contrast learning and multi-task optimization. The scheme is based on a pre-trained language model, and through domain-specific data construction, joint training of contrast learning and language modeling tasks, diversified difficult negative example generation strategy and character-level data enhancement mechanism, it realizes high-precision and high-robustness text matching capability, especially suitable for semantic matching tasks in professional scenarios such as cardiovascular diseases in the medical field.

[0009] 2. Technical solution: To solve the above problems, the core technical scheme of the present application includes the following aspects: (1) Bottom semantic encoding structure based on ALBERT Chinese pre-training model: The present application adopts the Chinese version of the lightweight pre-training language model ALBERT as the bottom encoder, taking into account the model performance and resource efficiency. Through fine-tuning adaptation of medical short text context, the extraction and modeling of domain semantic features are realized.

[0010] (2) Construction of multi-source high-quality training set integrating general corpus and medical corpus: The present application integrates multi-source heterogeneous corpus to construct a high-quality training dataset suitable for short text semantic matching tasks. First, a large amount of open-source general semantic matching datasets and company-internal accumulated cardiovascular medical question and answer short text data are collected and cleaned, and efficient deduplication is performed through MinHash and Locality-Sensitive Hashing (LSH) method to remove redundant samples and improve data quality.

[0011] Subsequently, based on the semantic indexing system constructed by the vector database, the vector similarity retrieval technology is used to automatically mine semantically similar text pairs from the corpus, generating the initial "anchor-positive" pairs. For these similar samples, a generative large language model is introduced, combined with a designed prompt template (Prompt Template), to automatically generate "hard negative" difficult negative examples with similar semantics but significant expression differences, expanding the sample format to a triple (anchor, positive, hardnegative). This generation process not only improves the semantic discriminability of the training samples, but also significantly enhances the robustness of the model to boundary samples.

[0012] To ensure the accuracy and semantic integrity of the samples, all automatically generated samples are manually reviewed or annotated to filter semantic deviations and unreasonable generation, ensuring that the training data has high discriminability and high credibility at the semantic level. The multi-source sample construction process significantly improves the quality of the supervised signal for model training, providing a solid data foundation for subsequent contrast learning frameworks.

[0013] (3) Joint training mechanism of supervised contrast learning (SimCSE) and mask language modeling (MLM) task: During the model training phase, the present application jointly introduces the SimCSE supervised contrast learning goal and the MLM task, and designs a double-task loss function structure. SimCSE is used to optimize the converging structure of sentence vectors, improving the semantic matching discrimination ability; the MLM task is used to strengthen the model's perception of character-level context, enhancing language modeling and error correction capabilities. To avoid gradient conflict or imbalance between the two tasks, the present application further introduces a gradient normalization mechanism to improve training stability and multi-task collaboration efficiency.

[0014] (4) Character-level robustness enhancement mechanism for real scene OCR noise: The present application is specifically aimed at the problem of shape similar character errors caused by OCR recognition in medical scenarios, and introduces a character-level data enhancement mechanism based on a shape similar character dictionary. In the training data, characters are randomly selected and replaced through a shape similar character dictionary (such as "muscle" and "limb", "Shaanxi" and "narrow"), or redundant words are randomly inserted between characters, thereby simulating interference factors in the actual recognition process and improving the robustness of the model under noisy text. This mechanism can be synchronized with contrast learning and MLM tasks, so that the model has anti-interference ability at both semantic abstraction and character repair levels.

[0015] (5) Unified training framework and adaptation process: The present application forms an end-to-end unified training process, covering data construction, enhancement strategy, pre-trained model fine-tuning, and multi-task joint training. The finally trained model can be directly deployed in medical field question and answer systems, diagnosis recommendation systems, report verification systems, etc. scenes, realizing accurate and highly robust short text semantic matching.

[0016] Through the above technical solutions, the present application effectively solves the problems of weak semantic understanding ability, sensitivity to shape similar character interference, and low training efficiency of existing short text matching models in professional fields, and has clear innovation, applicability and promotional value.

[0017] 3. Advantages: Compared with the prior art, the technical solutions provided by the present application have the following advantages: (1) Improve the semantic discrimination ability in professional fields: Unlike traditional methods that rely solely on general pre-training corpus, the present invention fully integrates open-source general data and vertical field medical data, and generates high-quality difficult negative example triples using a generative large language model, effectively expanding the training semantic boundary. This solution not only improves the model's semantic recognition ability for professional terms such as "chest tightness" and "angina pectoris", but also enhances the ability to distinguish complex medical expressions (such as the mixing of spoken language and professional terms), significantly improving the model's precision performance in fine-grained semantic matching tasks.

[0018] (2) Enhance the robustness of the model in noisy scenarios: Compared to traditional models that only focus on semantic encoding, the present invention first introduces a character-level similar character perturbation mechanism to simulate common character recognition errors in OCR scenarios, and enhances the model's fault tolerance to structural perturbations through random replacement and insertion. This mechanism, combined with MLM task collaborative training, enables the model to recognize and correct character-level abnormalities, providing stronger robustness and generalization ability in practical applications such as paper document recognition and medical question answering automatic matching.

[0019] (3) Optimize training efficiency and multi-task collaborative effect: The present invention uses SimCSE supervised contrastive learning and MLM task joint training, automatically coordinates the update amplitude of the two types of loss functions through gradient normalization mechanism, and overcomes the performance bottleneck caused by "single target training" in existing methods. The model can simultaneously learn global semantic relationships and local character modeling during the training phase, resulting in faster training convergence, more robust semantic vectors, and significantly improved overall modeling efficiency.

[0020] (4) Enhance the diversity and practicality of sample generation: The present invention introduces diversified generation prompt word templates extracted from databases, and automatically constructs multiple types and high complexity difficult negative examples using the capabilities of large models, effectively solving the problems of low efficiency and poor coverage of traditional manual sample construction. The generated samples are reliable after manual verification, and the trained model can better adapt to the matching needs of diverse expressions in actual contexts.

[0021] (5) Strong scalability, suitable for multiple vertical field applications: The technical framework proposed by the present invention is not only applicable to the cardiovascular medical field, but also can be flexibly applied in other vertical industries (such as law, finance, industrial question answering, etc.). By replacing field data and prompt templates, the model can quickly adapt to field-specific terminology systems and expression methods, and has good universality and generalizability.

[0022] In summary, this invention has significant advantages in improving semantic modeling quality, enhancing noise robustness, optimizing training mechanisms, and expanding practical applications, breaking through the performance bottleneck of existing short text matching technologies in vertical fields.

[0023] It should be noted that the structures not described in this invention are not related to the design points and improvement directions of this invention, and are the same as or can be implemented using existing technologies, so they will not be elaborated here. Attached Figure Description

[0024] Figure 1 This is a flowchart illustrating the data construction process of the present invention. Figure 2 This is a flowchart of the model training process of the present invention; Figure 3 This is a diagram of the multi-task model architecture of the present invention. Detailed Implementation

[0025] To facilitate understanding of the present invention, a more complete description of the invention will be given below with reference to the accompanying drawings, which illustrate several embodiments of the invention. However, the invention can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided so that the disclosure of the invention will be more thorough and complete.

[0026] like Figure 1 As shown, multi-source data fusion: General corpora (STS-B, LCQMC) and de-identified medical data (processed with differential privacy); Triple generation: Positive samples are retrieved based on bge-large-zh-v1.5 and Faiss index, and difficult negative examples are generated by combining qwen; like Figure 2 As shown, during model training: The gradient normalization mechanism dynamically adjusts the weights (λ=0.7) to avoid single-task dominance.

[0027] Test results: On a cardiovascular test set of 10,000 samples (including 50% OCR noise), the traditional method based on the BERT-base model achieved an accuracy of 82.1%, while the accuracy of this invention reached 94.3%.

[0028] 1. Data preparation stage 1.1. Multi-source data fusion Data source: General corpus: Chinese STS-B (120,000 pairs), LCQMC (260,000 pairs); Medical data: 80,000 cardiovascular outpatient records were anonymized using differential privacy technology (Gaussian noise added, privacy budget ε=0.1): [Chief Complaint] Squeezing pain behind the sternum for 3 hours; [Diagnosis] Acute myocardial infarction; All medical data is processed using differential privacy technology (Gaussian noise with a mean of 0 and a standard deviation of σ is added, with a privacy budget of ε=0.1), sensitive information such as patient names and ID numbers is removed, and irreversible anonymized data is generated.

[0029] Deduplication: Using MinHash + Local Sensitive Hash (LSH) and setting the Jaccard similarity threshold to 0.98, redundant data is reduced by 20%.

[0030] 1.2. Triple Sample Generation Positive example mining: Using bge-large-zh-v1.5 encoded text, the Faiss index was used to retrieve the top-3 semantically similar pairs (such as "chest pain" and "precordial discomfort").

[0031] Difficult-to-bear example generation: creating semantic differences by rewriting key medical terms; Template example: Command: Generate a statement that differs from the medical significance of the anchor point; Input anchor: "Elevated troponin levels"; Output generated: "Crotonin levels decreased"; Quality control: Samples are validated by 3 doctors, and samples with semantic equivalence or professional errors are rejected (error rate <2%).

[0032] 2. Implementation of key technologies 2.1. Noise Enhancement Module Replacement of similar-looking characters: Based on a pre-built medical dictionary (such as {"muscle":"limb","narrow":"Shaanxi"}), 5% of the characters are randomly replaced.

[0033] Example: "Coronary artery stenosis" → "Coronary artery narrowing"; Symbol insertion: Randomly insert spaces / punctuation marks (2% probability) to simulate OCR errors.

[0034] 2.2. Multi-task model architecture Model architecture design: Base model: The ALBERT pre-trained model is used, which reduces the number of parameters through factorized embedding and cross-layer parameter sharing; Dual-head design: Task 1: Supervised contrastive learning (SimCSE), optimize text representation alignment; Task 2: Masked language modeling (MLM), enhance language understanding ability.

[0035] 2.3. Joint training and gradient regulation Loss function design: Contrastive loss (SimCSE): ; Where is the text vector representation, is the temperature hyperparameter (preferably 0.05), and sim(·) is the cosine similarity; MLM loss function definition: ; a. The input sequence is ; b. The masked sequence is , some of which are replaced by ; c. The set of masked positions is ; d. The model predicts the probability of the word at the i-th position ; Gradient normalization: Dynamically adjust the gradient ratio of the two tasks to avoid the dominance of a certain task: ; Where is the balance coefficient (preferably 0.7), and are the gradient vectors of the loss function.

[0036] Key parameters:

[0037] 3. Training and deployment 3.1. Training process Batch organization: Each batch contains 32 groups of triplets (anchor / positive / difficult negative = 1:1:1); Stopping condition: F1 of the validation set does not improve for 3 consecutive rounds; 3.2. Medical triage application Actual effect:

[0038] The above-described embodiments only express some implementation manners of the present application, which are described in a more specific and detailed manner, but should not be understood as a limitation on the patent scope of the present application; it should be noted that, for ordinary skilled persons in the art, several modifications and improvements can be made without departing from the concept of the present application, which all belong to the protection scope of the present application; therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. A short text matching method based on enhanced contrast learning and multi-task optimization, characterized in that, The method comprises the following steps: S1, constructing a multi-source training dataset: fusing a general field open source semantic matching dataset and a vertical field medical short text data, and performing deduplication processing through MinHash and local sensitive hashing (LSH); S2, generating high-quality triple samples: retrieving positive samples (s, s + ) based on vector similarity, and generating semantically related but non-equivalent hard negative examples s - using a generative language model combined with a preset prompt template, forming a triple (s, s + , s - ); S3, introducing character-level noise enhancement: based on a medical homograph dictionary, randomly replacing characters or inserting noise symbols in the training text to simulate optical character recognition (OCR) errors; S4, designing a multi-task joint training framework, using an ALBERT encoder to simultaneously perform supervised contrastive learning (SimCSE) and mask language modeling (MLM) tasks, and dynamically balancing the gradient updates of the two tasks through a gradient normalization mechanism; S5, end-to-end model training: inputting a triple batch to jointly train the model, optimizing the contrastive loss function and the MLM loss function, and outputting a robust short text semantic matching model.

2. The short text matching method based on enhanced contrast learning and multi-task optimization according to claim 1, characterized in that, The generation of difficult negative examples in the S2 step comprises: Retrieving Top-K positive samples of anchor texts from the corpus based on vector similarity; Inputting the anchor text and the prompt template into the generative language model to generate negative samples with key semantic differences based on the preset medical terminology constraints; All generated samples need to be manually verified to filter out samples with semantic bias.

3. The short text matching method based on enhanced contrast learning and multi-task optimization according to claim 1, characterized in that, The character-level noise enhancement in the S3 step is specifically: Based on the common misrecognition statistics of medical OCR, a homograph mapping dictionary is constructed, including similar pairs of characters (shuan, sai, ji, fu, zhai, shan); Randomly select 5% of the characters in the training text and replace them with homographs according to the dictionary; Insert spaces or punctuation marks at random positions to simulate OCR recognition errors.

4. The short text matching method based on enhanced contrastive learning and multi-task optimization according to claim 1, characterized in that, The joint training framework in the S4 step satisfies: Using a contrastive loss function The contrastive loss function is defined as follows: ; wherein is a text vector representation, is a temperature hyperparameter, sim( ) is a cosine similarity function; Using MLM loss function Reinforcing character-level modeling, the MLM loss function is defined as: ; wherein is a set of mask positions, is a post-mask text; Using a gradient normalization mechanism to balance the gradient updates of the two tasks, the gradient normalization formula is: ; wherein is a gradient weight coefficient for the contrastive learning task, and are gradient vectors of the loss function, respectively.

5. The short text matching method based on enhanced contrastive learning and multi-task optimization according to claim 1, characterized in that, The training configuration in the S5 step includes: Using the AdamW optimizer, setting the learning rate to 2e-5 and the weight decay to 0.01; Each batch contains 32 triplets, each containing one anchor text, one positive sample, and one difficult negative example; The training termination condition is that the F1 score of the validation set does not improve for 3 consecutive rounds.

6. The short text matching method based on enhanced contrastive learning and multi-task optimization according to claim 1, characterized in that, The vertical field medical short text data includes cardiovascular clinic records, test reports, and doctor-patient conversation texts, and all data have been desensitized to remove patient privacy information. 7.An electronic device comprising a memory and a processor, the memory storing a computer program, wherein, The processor executes the program to implement the steps of the method of any one of claims 1-6.

8. A computer readable storage medium having stored thereon computer instructions, wherein, The instructions are executed by the processor to implement the steps of the method of any one of claims 1-6.