Large language model alignment method oriented to specific professional field

By constructing training data sets in professional fields and performing supervision and fine-tuning, combined with the use of literature recommendation modules, the problems of insufficient knowledge accuracy and illusion of literature recommendation in large language models in specific professional fields are solved, and more efficient task processing capabilities and more reliable literature recommendations are achieved.

CN120218254APending Publication Date: 2025-06-27TONGJI UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510400789.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Large language models face problems such as insufficient knowledge accuracy and illusion of literature recommendation in the application of specific professional fields, and it takes a lot of time and effort to obtain training data in professional fields.

Method used

By constructing training data sets in professional fields, using the knowledge reserves and text generation capabilities of high-level large language models, training samples are generated and supervised fine-tuned, and combining the literature recommendation module to use a matching algorithm based on text and text vectors for reference recommendation.

Benefits of technology

It significantly improves the task processing capabilities of large language models in specific professional fields, improves knowledge accuracy and reliability of literature recommendations, and reduces the cost of data collection and model adaptation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218254A_ABST
    Figure CN120218254A_ABST
Patent Text Reader

Abstract

The invention relates to a large language model alignment method oriented to a specific professional field, and the method is realized based on a data set construction assembly line, a supervision fine tuning technology and a literature recommendation module, and aims to enhance the understanding and expression ability of a large language model in the specific field. The method comprises the following steps: firstly, in a data set construction stage, dividing a target domain into a plurality of sub-domains by requesting a high-level large language model, and generating a training data set for all the sub-domains according to various task scenes; then, the target model is optimized through a supervision fine tuning technology in a training stage, so that the target model can adapt to a specific task in a professional field; and finally, real and accurate references are provided for the questions input by the user through a literature matching module. According to the large language model alignment method oriented to the specific professional field, the performance of the large language model in the professional field task can be remarkably improved at low cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of natural language processing, and in particular, to a method and technical framework for aligning large language models for specific professional fields. Background Art

[0002] Large language models have broad application prospects in vertical fields, especially in professional fields such as medicine, finance, and law. Through customized training, large language models can provide more accurate and efficient solutions. Although open-source models have great potential, they still face many challenges when applied to downstream tasks in specific fields. On the one hand, the pre-training corpus of open-source models usually covers a vast amount of network data, which generally contains information irrelevant or even incorrect to the professional field, diluting the truly effective professional knowledge and thus limiting the performance of the model. On the other hand, users often need the model to provide task-related reference documents. However, the generative nature of language models leads to the "hallucination" problem, that is, the model may generate reference materials that do not exist in reality, thus affecting the reliability of the model. In addition, how to obtain training data in the professional field remains an important challenge, and developers need to invest a lot of time and effort in data collection, model adaptation, and training. Summary of the Invention

[0003] The purpose of the present invention is to provide a method for aligning large language models for specific professional fields in view of the deficiencies of the prior art.

[0004] The purpose of the present invention is achieved by the following technical solutions: A method for aligning large language models for specific professional fields, characterized by including the following steps: (1) Preparation stage Determine the target field and its specific task types, as well as the target model to be aligned, obtain the model parameter file of the target model and the access right to the advanced large language model; (2) Dataset construction stage For various tasks in the target field described in step (1), utilize the knowledge reserve and text generation ability of the advanced large language model to form training samples, where each question and its corresponding answer constitute a training sample; construct a training dataset from the training samples and provide it to step (3); (3) Supervised fine-tuning stage Input the question in each training sample in the dataset in step (2) into the target model in step (1), calculate the cross-entropy loss for the answer output by the target model and the answer in the training sample, and optimize the parameters of the target model until convergence to obtain a professional field model; provide it to step (4); (4)The literature recommendation module uses a text-based matching algorithm and a text vector-based matching algorithm to recommend reference documents respectively according to the current question, and the final literature recommendation result is obtained after merging; it is provided to step (5). (5)Use the professional field model described in step (3) to generate text, and recommend reference documents through the literature recommendation module described in step (4). The results of both are merged and the final answer is output.

[0005] Preferably, step (2), Specifically includes the following sub-steps: (2.1)Divide the target field into several sub-fields by requesting an advanced large language model; (2.2)For each task type in the target field, request the advanced large language model to generate several training samples for the sub-fields in step (2.1), and each training sample includes a question and its corresponding answer; (2.3)Randomly sort and merge all the training samples obtained in step (2.2) to obtain the training data set for the target field.

[0006] Preferably, step (2), The number of sub-fields divided by the advanced large language model for the target field is a variable, and the generation process of the sub-fields can be expressed as: Among them, represents the set composed of all sub-fields, represents all sub-fields, represents the process of the advanced large language model generating sub-fields, represents the target field, represents the number of sub-fields to be generated.

[0007] Preferably, step (2), The specific task types of the target field need to be set in advance, and are expressed as: Among them, represents the set composed of various tasks in the target field, represents the specific task types included in the target field set in advance.

[0008] Preferably, in step (2), The number of training samples generated by the advanced large language model for various tasks and sub-fields in the target field is a variable, and the generation process is as follows: Among them, and respectively represent the current sub - domain and the specific task type, represents the number of training samples expected to be generated for each task type and each sub - domain, represents the process of the advanced large - language model generating training samples, represents the set of training samples generated by the advanced large - language model for the current sub - domain and task type, represents all the training samples included in the current sub - domain and task type.

[0009] Subsequently, all the training samples are randomly sorted and then merged to obtain the final training dataset, as shown below: Among them, represents the finally constructed training dataset.

[0010] Preferably, in step (3), the supervised fine - tuning calculates the cross - entropy loss between the answer output by the target model and the answer in the training sample, so as to update the model parameters. The loss function is as follows: Among them, is the loss of supervised fine - tuning, and respectively represent the sets of word tokens corresponding to the question and answer in the training sample after text embedding, represents all the updatable parameters in the model, represents the first word tokens in

[0011] Preferably, step (4), specifically includes the following sub - steps: (4.1) For the target domain, a sufficient number of reference documents are pre - screened to form a reference document library; (4.2) According to the current question, use a text - based matching algorithm to screen out several candidate reference documents with the highest matching degree in the document library in step (4.1); (4.3) According to the current question, use a text - vector - based matching algorithm to screen out several candidate reference documents with the highest matching degree in the reference document library in step (4.1); (4.4) Perform a union operation on all the candidate reference documents in steps (4.2) and (4.3) to obtain the final document recommendation result.

[0012] Preferably, in step (4), The text-based matching algorithm calculates the matching degree of a document by computing the proportion of the same text between the current question and the title of the reference document, as follows: Among them, represents the text-based matching degree, and represent the current question and the title of the document to be matched, respectively, means converting the text into a set of individual words; Subsequently, the 3 reference documents with the highest matching degrees are selected as candidate documents, as follows: Among them, represents the candidate documents obtained through the text-based matching algorithm, represents the preset reference document library.

[0013] Preferably, in step (4), the text vector-based matching algorithm calculates the cosine similarity between the current question and the word vectors converted from the title of the reference document as the matching degree, as follows: Among them, represents the text vector-based matching degree, and represent the current question and the title of the document to be matched, respectively, means the text embedding model that can convert the text into vector representation; Subsequently, the 3 reference documents with the highest matching degrees are selected as candidate documents, as follows: Among them, represents the candidate documents obtained through the text-based matching algorithm, represents the preset reference document library; Finally, and are subjected to a union operation to obtain the final document recommendation result, as follows: Among them represents the final document recommendation result.

[0014] Through the above technical solution, by constructing a training dataset in a specific professional field, supervised fine-tuning, and combining a literature recommendation module, it is possible to solve problems such as insufficient knowledge accuracy of large language models in specific professional fields and literature recommendation hallucinations, and achieve the purpose of expanding their professional field task processing capabilities at a relatively low cost. Compared with the prior art, the dataset construction method of the present invention utilizes the knowledge reserve and text generation ability of advanced large language models, and through sub-field division and training sample generation, can efficiently construct a high-quality training dataset for specific professional fields, providing sufficient data support for model alignment. In the supervised fine-tuning stage of the present invention, through cross-entropy loss calculation and parameter optimization, the comprehensive task processing ability of the model in a specific field can be significantly improved. The literature recommendation module of the present invention combines a text-based matching algorithm and a text vector-based matching algorithm, and can more accurately screen and recommend relevant reference documents to ensure the authenticity and relevance of the literature.

[0015] Through the above technical solution, the present invention can effectively improve the task processing ability of large language models in specific professional fields at a relatively low cost. Brief Description of the Drawings

[0016] Figure 1 is a schematic structural diagram of the present invention; Figure 2 is a description of the construction process of the training dataset involved in an embodiment of the present invention; Figure 3 is a schematic flow diagram of the present invention; Figure 4 is a legal knowledge understanding task in the field of copyright law according to an embodiment of the present invention. Detailed Embodiments

[0017] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. This embodiment is implemented on the premise of the technical solution of the present invention, and gives detailed implementation manners and specific operation processes, but the protection scope of the present invention is not limited to the following embodiments.

[0018] The present invention aims at the application problems of large language models in specific professional fields, and proposes a large language model training method adapted to professional field tasks and an accurate and authentic reference document recommendation method.

[0019] The advanced large language model of the present invention is an advanced natural language processing model known in the prior art, such as GPT-4 or DeepSeek-V3; the target model of the present invention is an open-source pre-trained language model known in the prior art, such as Llama series models or Qwen series models.

[0020] Embodiment: Step 1: Determine that the target domain is "Copyright Law", and the specific task types included are: legal knowledge understanding, legal knowledge application, and open-ended knowledge Q&A. The target models to be aligned include: Llama3-8B, Qwen2-7B, Mistral-7B, and Baichuan2-7B. The advanced large language model used to construct the training dataset is DeepSeek-V2.5. Obtain the model parameter files of the target models and the access rights to the advanced large language model.

[0021] Step 2: In the dataset construction phase, utilize the rich professional knowledge reserve and powerful text generation and reasoning capabilities of the advanced large language model to construct a training dataset for various tasks in the target domain. Specifically, it includes the following sub-steps: Step 2.1: Through prompt engineering, request the advanced large language model to divide the target domain into N sub-domains (N is a hyperparameter), ensuring a high degree of differentiation between each sub-domain, and obtain the set of sub-domains : Among them, represents the set composed of all sub-domains, represents all sub-domains, represents the process of the advanced large language model generating sub-domains, represents the target domain, represents the number of sub-domains to be generated.

[0022] Step 2.2: For all task types in the target domain (represented as ), request the advanced large language model to generate K training samples (K is a hyperparameter) for each sub-domain in Step 2.1. Each training sample contains a question and its corresponding answer, and obtain the training sample sets for each sub-domain : Among them, and respectively represent the current sub-domain and the specific task type, represents the expected number of training samples to be generated for each task type and each sub-domain, represents the process of the advanced large language model generating training samples, represents the training sample set generated by the advanced large language model for the current sub-domain and task type, represents all the training samples included in the current sub-domain and task type.

[0023] Step 2.3: Randomly sort and merge all the generated training samples in Step 2.2 to form the training dataset for the target domain : Step 3: In the supervised fine-tuning phase, use the training dataset constructed in Step 2 to train the target model. Input the questions in each training sample of the training dataset into the target model to obtain the answers output by the model, and calculate the cross-entropy loss for the answers output by the model and the answers in the training samples : wherein, is the loss of supervised fine-tuning, and respectively represent the sets of word tokens corresponding to the questions and answers in the training samples after text embedding, represents all the updatable parameters in the model, represents the first word tokens in

[0024] Step 4: Construct a reference library for the target domain, and use a word-based matching algorithm and a text vector-based matching algorithm to recommend references according to the current question. Specifically, it includes the following sub-steps: Step 4.1: For the "Copyright Law" target domain, pre-screen a sufficient amount of references related to "Copyright Law" to form a reference library ; Step 4.2: According to the current question, use a word-based matching algorithm to screen out several candidate references with the highest matching degree from the reference library : wherein, represents the word-based matching degree, and respectively represent the current question and the title of the document to be matched, represents converting the text into a set of individual words. represents the candidate documents obtained through the word-based matching algorithm, represents the preset reference library.

[0025] Step 4.3: According to the current question, use a text vector-based matching algorithm to screen out several candidate references with the highest matching degree from the reference library : Among them, represents the matching degree based on text vectors, and represent the current question and the title of the literature to be matched respectively, represents a text embedding model that can convert text into vector representations. represents the candidate literature obtained through a word-based matching algorithm, represents a preset reference literature library.

[0026] Step 4.4: Perform a union operation on the candidate reference literatures in Step 4.2 and Step 4.3 to obtain the final reference literature result : Step 5: Use the trained "Copyright Law" professional field model to generate text to obtain a preliminary answer, recommend reference literatures for the current question through the literature recommendation module, and connect the recommendation result to the preliminary answer output by the model to form the final answer, as follows: Figure 4 As shown, legal knowledge understanding task experiments in the field of copyright law were carried out on four types of open-source models, and specific tasks in the "Copyright Law" field were completed using the method of the present invention, thereby verifying the effectiveness of the method.

[0027] Figure 4 shows the accuracy and stability of the official fine-tuned version of the open-source model and the case after alignment using the method of the present invention in the task of legal knowledge understanding.

[0028] Figure 4 uses accuracy and information entropy as the presentation methods of model performance.

Claims

1. A large language model alignment method for a specific professional field, characterized in that: The following steps are involved: (1) Preparation stage Determine the target domain and its specific task types, as well as the target model to be aligned, obtain the model parameter file of the target model and access to the high-level large language model; (2) Dataset construction phase For various tasks in the target domain described in step (1), the knowledge reserve and text generation capability of the advanced large language model are used to form training samples, where each question and its corresponding answer constitute the training sample; Construct a training data set from the training samples and provide it to step (3); (3) Supervised fine-tuning stage Input the questions in each training sample in the data set in step (2) into the target model in step (1), calculate the cross entropy loss between the answers output by the target model and the answers in the training samples, optimize the parameters of the target model until convergence, and obtain a professional domain model; and provide it to step (4); (4) The literature recommendation module uses a text-based matching algorithm and a text vector-based matching algorithm to recommend references based on the current question, and then merges them to obtain the final literature recommendation result; Provided to step (5); (5) Use the professional domain model described in step (3) to generate text, and use the literature recommendation module described in step (4) to recommend references. The results of the two are combined to output the final answer.

2. The large language model alignment method for a specific professional field according to claim 1, characterized in that: The step (2) It includes the following sub-steps: (2.1) Divide the target domain into several subdomains by requesting a high-level large language model; (2.2) For each task type in the target domain, request the high-level large language model to generate several training samples for the sub-domains in step (2.1), each training sample contains a question and its corresponding answer; (2.3) All training samples obtained in step (2.2) are randomly sorted and merged to obtain the training data set of the target domain.

3. The large language model alignment method for a specific professional field according to claim 1, characterized in that: The step (2) The number of sub-domains divided by the requesting high-level large language model for the target domain is a variable, and the generation process of the sub-domains can be expressed as: in, Represents the set of all sub-domains, Represents all sub-domains, Represents the process of generating sub-domains of high-level large language models, represents the target area, Indicates the number of sub-domains that need to be generated.

4. The large language model alignment method for a specific professional field according to claim 1, characterized in that: The step (2) The specific task type in the target area needs to be pre-set, expressed as: in, Represents a set of various tasks in the target domain, Indicates the specific task types included in the preset target area.

5. The large language model alignment method for a specific professional field according to claim 1, characterized in that: In step (2), The number of training samples generated by the advanced large language model for various tasks and sub-fields in the target field is a variable, and the generation process is as follows: in, and Respectively represent the current sub-field and specific task type, represents the number of training samples expected to be generated for each task type and each sub-field, Represents the process of generating training samples by advanced large language models. Represents the set of training samples generated by the high-level large language model for the current sub-domain and task type. Represents all training samples contained in the current sub-domain and task type; Then all training samples are randomly sorted and merged to obtain the final training data set, as shown below: in, Indicates the training dataset that is finally constructed.

6. The large language model alignment method for a specific professional field according to claim 1, characterized in that: In step (3), The supervised fine-tuning calculates the cross entropy loss for the answer output by the target model and the answer in the training sample, thereby updating the model parameters. The loss function is as follows: in, To supervise the fine-tuning loss, and Respectively represent the word sets corresponding to the questions and answers in the training samples after text embedding, represents all updateable parameters in the model, express The front A word.

7. The large language model alignment method for a specific professional field according to claim 1, characterized in that: The step (4), It includes the following sub-steps: (4.1) Pre-screen a sufficient number of references for the target field to form a reference library; (4.2) Based on the current problem, use a text-based matching algorithm to select several candidate references with the highest matching degree in the literature library in step (4.1); (4.3) Based on the current problem, use a text vector-based matching algorithm to select several candidate references with the highest matching degree in the reference library in step (4.1); (4.4) Perform a union operation on all candidate references in step (4.2) and step (4.3) to obtain the final document recommendation result.

8. The large language model alignment method for a specific professional field according to claim 1, characterized in that: In step (4), The text-based matching algorithm calculates the proportion of the same text in the current question and the reference title to obtain the matching degree of the reference, as shown below: in, Represents the degree of text-based matching, and Represent the current question and the title of the document to be matched, respectively. Indicates the conversion of text into a collection of individual characters; Then, the three references with the highest matching degree are taken as candidate documents, as shown below: in, represents the candidate documents obtained by the text-based matching algorithm, Represents a preset reference library.

9. The large language model alignment method for a specific professional field according to claim 1, characterized in that: In step (4), The text vector-based matching algorithm calculates the cosine similarity between the current question and the word vector converted from the reference title as the matching degree, as shown below: in, Represents the matching degree based on text vector, and Represent the current question and the title of the document to be matched, respectively. It represents the text embedding model, which can convert text into vector representation; then, the three references with the highest matching degree are taken as candidate documents, as shown below: in, represents the candidate documents obtained by the text-based matching algorithm, Indicates the preset reference library; Finally and Perform a union operation to obtain the final document recommendation result, as shown below: in Represents the final literature recommendation result.

Citation Information

Cited By

  • Cross-domain ship large model training method and device, computer equipment, storage medium and computer program product

    CN121614874A