Target domain model training method and device, computer equipment and storage medium

By performing structured processing and reinforcement learning training on target domain corpus data, a model with structured logic and professional knowledge in the target domain is generated. This solves the problem of accuracy and compliance of general large models in specific domain question-answering scenarios, and enables the model to perform efficient and reliable reasoning under complex tasks.

CN121809579APending Publication Date: 2026-04-07SHENZHEN WEIBAO INFORMATION SERVICE CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-22
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

General-purpose large models cannot guarantee the accuracy and compliance of business reasoning in specific domain question-answering scenarios. Existing methods rely heavily on context engineering, making it difficult to guarantee the correctness of the models.

Method used

By acquiring corpus data in the target domain and processing it into a structured thought chain, and using business logic and compliance rules to generate reasoning thought chain corpus data, the initial large model is trained. Then, reinforcement learning training with multi-dimensional rewards is carried out in combination with reinforcement learning training data to form a target large model in the target domain.

Benefits of technology

It improves the correctness and compliance of business reasoning in the target domain of the target large model, provides a reliable reasoning framework, and significantly improves the stability and correctness of the model's answers under complex tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121809579A_ABST
    Figure CN121809579A_ABST
Patent Text Reader

Abstract

The invention relates to a target domain model training method and device, computer equipment, a computer readable storage medium and a computer program product. The method comprises the steps of obtaining first corpus data of a target field, and performing structured thinking chain processing on the first corpus data of the target field to obtain reasoning thinking chain corpus data; training an initial large model based on the reasoning thinking chain corpus data to obtain a structured thinking chain large model of the target domain; obtaining second corpus data of the target field, and processing the second corpus data of the target field to obtain reinforcement learning training data; and performing multi-dimensional reward reinforcement learning training on the structured thinking chain large model of the target domain based on the reinforcement learning training data to obtain a target large model of the target domain. And the business reasoning correctness of the target field of the target large model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field, and in particular to a method, apparatus, computer device, computer-readable storage medium, and computer program product for training a target domain model. Background Technology

[0002] With the development of general-purpose large-scale models, more and more users are starting to use them. For questions in specific domains, it is necessary to carefully design prompts and provide a few examples to guide the general-purpose large-scale model to generate applicable responses in that specific domain's question-answering scenario. However, this method relies heavily on context engineering, and the correctness of the model's business reasoning is difficult to guarantee. Summary of the Invention

[0003] Therefore, it is necessary to provide a training method, apparatus, computer device, computer-readable storage medium, and computer program product for a target domain model that can improve the accuracy of model business reasoning, in order to address the above-mentioned technical problems.

[0004] Firstly, this application provides a method for training a target domain model, the method comprising:

[0005] Obtain the first corpus data of the target domain, and perform structured thinking chain processing on the first corpus data of the target domain to obtain reasoning thinking chain corpus data;

[0006] The initial large model is trained based on the reasoning thought chain corpus data to obtain the structured thought chain large model of the target domain;

[0007] Obtain second corpus data in the target domain, and process the second corpus data in the target domain to obtain reinforcement learning training data;

[0008] Based on the reinforcement learning training data, a multi-dimensional reward reinforcement learning training is performed on the structured thinking chain model of the target domain to obtain the target model of the target domain.

[0009] In one embodiment, the process of performing structured thought chain processing on the first corpus data of the target domain to obtain inference thought chain corpus data includes:

[0010] By utilizing the business logic and compliance rules of the target domain, the first corpus data of the target domain is processed using structured thinking chain processing to obtain reasoning thinking chain corpus data.

[0011] In one embodiment, the step of using the business logic and compliance rules of the target domain to perform structured thought chain processing on the first corpus data of the target domain to obtain inference thought chain corpus data includes:

[0012] Semantic analysis is performed on user questions in the first corpus data of the target domain to obtain semantic analysis results;

[0013] Obtain the knowledge data needed to answer the user's question;

[0014] Based on the semantic analysis results and the knowledge data, reasoning is performed according to the business logic of the target domain to obtain the logical reasoning results between the question and the answer.

[0015] The logical reasoning results between the questions and answers are evaluated for reasonableness and compliance with rules to obtain reasoning chain corpus data.

[0016] In one embodiment, training the initial large model based on the reasoning thought chain corpus data to obtain the structured thought chain large model of the target domain includes:

[0017] Based on the reasoning thought chain corpus data, and combined with the thought chain supervision instructions, the initial large model is trained under supervision to obtain a structured thought chain large model for the target domain.

[0018] In one embodiment, training the initial large model based on the reasoning thought chain corpus data to obtain the structured thought chain large model of the target domain includes:

[0019] For each reasoning thought chain corpus data, the reasoning thought chain corpus data is processed through an initial large model to obtain the target output loss and thought chain loss;

[0020] The multi-task joint loss is obtained based on the target output loss and the thought chain loss.

[0021] Based on the multi-task joint loss, the parameters of the initial large model are adjusted to obtain the structured thought chain large model of the target domain.

[0022] In one embodiment, the processing of the second corpus data in the target domain to obtain reinforcement learning training data includes:

[0023] The second corpus data in the target domain is processed using structured thought chain processing to obtain reinforcement learning training data.

[0024] In one embodiment, the multi-dimensional reward includes outcome reward and process reward; the step of performing multi-dimensional reward reinforcement learning training on the structured thought chain model of the target domain based on the reinforcement learning training data to obtain the target domain's target model includes:

[0025] Based on the reinforcement learning training data, the model parameters of the structured thought chain model in the target domain are optimized using a group policy optimization algorithm that includes outcome rewards and process rewards, thereby obtaining the target model in the target domain.

[0026] In one embodiment, the target field is the insurance field, the financial field, the medical field, the education field, or the transportation field.

[0027] Secondly, this application also provides a training apparatus for a target domain model, the apparatus comprising:

[0028] The structured thinking chain module is used to acquire the first corpus data of the target domain, and to perform structured thinking chain processing on the first corpus data of the target domain to obtain reasoning thinking chain corpus data.

[0029] The first training module is used to train the initial large model based on the reasoning thought chain corpus data to obtain a structured thought chain large model for the target domain.

[0030] The training data determination module is used to acquire the second corpus data of the target domain and process the second corpus data of the target domain to obtain reinforcement learning training data.

[0031] The second training module is used to perform multi-dimensional reward reinforcement learning training on the structured thinking chain model of the target domain based on the reinforcement learning training data, so as to obtain the target model of the target domain.

[0032] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the method described in the first aspect.

[0033] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in the first aspect.

[0034] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the method described in the first aspect.

[0035] The aforementioned training method, apparatus, computer equipment, computer-readable storage medium, and computer program product for the target domain model acquire first corpus data of the target domain, process the first corpus data into a structured thinking chain to obtain inference thinking chain corpus data, and use the inference thinking chain corpus data to train the initial large model to obtain a structured thinking chain large model of the target domain. This structured thinking chain large model possesses the structured thinking logic and professional knowledge understanding ability of the target domain. Then, a second corpus data is acquired, processed to obtain reinforcement learning training data, and the structured thinking chain large model of the target domain is trained with multi-dimensional reward reinforcement learning using the reinforcement learning training data to obtain a target large model of the target domain, thereby improving the business reasoning correctness of the target large model in the target domain. Attached Figure Description

[0036] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0037] Figure 1 This is an application environment diagram of a training method for a target domain model in one embodiment;

[0038] Figure 2 This is a flowchart illustrating the training method for a target domain model in one embodiment;

[0039] Figure 3 This is a schematic diagram of the hierarchical structure of the structured thinking chain module in one embodiment;

[0040] Figure 4 This is a schematic diagram of supervised instruction training for a structured thought chain in a target domain, as shown in one embodiment.

[0041] Figure 5 This is a schematic diagram of a multi-dimensional reward mechanism module in one embodiment;

[0042] Figure 6 This is a schematic diagram of reinforcement learning training for multi-dimensional rewards in a target domain in one embodiment;

[0043] Figure 7 This is a schematic diagram of the system architecture for training an insurance domain model in one embodiment;

[0044] Figure 8 This is a schematic diagram of supervisory instruction training for a structured thought chain in the insurance field in one embodiment;

[0045] Figure 9This is a schematic diagram of reinforcement learning training for multi-dimensional rewards in the insurance field in one embodiment;

[0046] Figure 10 This is a structural block diagram of a training device for a target domain model in one embodiment;

[0047] Figure 11 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0048] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0049] It should be noted that the terms "first," "second," etc., used in this application can be used to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from the second element. The terms "comprising" and "having," and any variations thereof, used in this application, are intended to cover non-exclusive inclusion. The term "multiple" used in this application refers to two or more. The term "and / or" used in this application refers to one of the embodiments, or any combination of multiple embodiments.

[0050] The method provided in this application embodiment can be applied to, for example... Figure 1 In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104 or located on the cloud or other network servers. Terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, drones, low-altitude aircraft, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, projection devices, etc. Portable wearable devices can include smartwatches, smart bracelets, head-mounted devices, etc. Head-mounted devices can be virtual reality (VR) devices, augmented reality (AR) devices, smart glasses, etc. Server 104 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.

[0051] In one exemplary embodiment, such as Figure 2 As shown, a method for training a target domain model is provided, which can be applied to... Figure 1Taking a terminal or server as an example, the explanation includes the following steps 202 to 208. Wherein:

[0052] Step 202: Obtain the first corpus data of the target domain, perform structured thinking chain processing on the first corpus data of the target domain, and obtain reasoning thinking chain corpus data.

[0053] The target domain can be insurance, finance, healthcare, education, sports, or transportation, among others, but is not limited to these. The first corpus data can be multimedia data from the target domain. Multimedia data can include one or more of the following: text, image, audio, and video data. Text data can be conversational text from the target domain. Image data can be used to extract text data. Audio data can be used to perform speech recognition to obtain text data. Video data can be used to extract text data. The first corpus data can be generated by cleaning, filtering, and compliance-checking the original corpus data from the target domain, combined with data augmentation and the construction of positive and negative sample pairs. The amount of first corpus data required for each training batch can be determined as needed.

[0054] The reasoning thought chain corpus data can be textual information with structured logical thinking. A structured thought chain refers to using a structured thinking framework in the reasoning process to improve the accuracy of model reasoning.

[0055] Structured thinking chain processing of the first corpus data in the target domain can include analyzing user questions in the first corpus data, obtaining knowledge data on responses to user questions, performing logical reasoning based on business logic, and judging the rationality of the reasoning logic and whether there is any content with compliance risks.

[0056] For example, the processor performs structured thought chain processing on each first corpus data in the target domain to obtain the corresponding reasoning thought chain corpus data.

[0057] Step 204: Train the initial large model based on the reasoning thought chain corpus data to obtain a structured thought chain large model for the target domain.

[0058] The initial large model can be selected as needed, such as Qwen, deepseek, hunyuan, llama, and other network structures. The initial large model exemplified in this solution can be expanded or simplified based on the limitations of model inference speed, memory usage, and the requirements for task accuracy and security in practical applications.

[0059] For example, the processor can input inference thought chain corpus data into an initial large model for processing, calculate the multi-task joint loss using a preset loss function, and adjust the model parameters of the initial large model based on the multi-task joint loss to obtain a structured thought chain large model for the target domain. The preset loss function includes parameters such as the target output sequence length, thought chain sequence length, input, preceding output, preceding thought chain, and weight coefficients of the thought chain loss.

[0060] Step 206: Obtain the second corpus data of the target domain, and process the second corpus data of the target domain to obtain reinforcement learning training data.

[0061] The second corpus data can be multimedia data from the target domain. Multimedia data can include one or more types of data, such as text, images, audio, and video. Text data can be conversational text from the target domain. Image data can be used to extract text data. Audio data can be used to perform speech recognition to obtain text data. Video data can be used to extract text data. The second corpus data can be obtained by cleaning, filtering, and compliance checks on the original target domain corpus data, as well as sampling positive and negative samples. The amount of second corpus data required for each training batch can be determined as needed. The second corpus data can be the same as or different from the first corpus data.

[0062] For example, the processor processes the second corpus data to obtain reinforcement learning training data. This can be done by directly using the second corpus data as reinforcement learning training data, or by performing structured thought chain processing on the second corpus data. The reinforcement learning training data can cover user question sets, large model instructions, and recalled knowledge corpora, etc.

[0063] Step 208: Based on the reinforcement learning training data, perform multi-dimensional reward reinforcement learning training on the structured thinking chain model of the target domain to obtain the target model of the target domain.

[0064] Reinforcement learning refers to training and optimizing large models using rule-based rewards or feedback signals (such as user satisfaction or rewarding model results). A reward model is a model based on quantified scoring, typically combining rule matching, semantic similarity, and other factors to calculate the reward signal, used to evaluate the quality of the output during the reinforcement learning process.

[0065] Multi-dimensional rewards refer to a composite reward mechanism that comprehensively and evenly guides model behavior during reinforcement learning training. It integrates multiple reward signals from different levels and perspectives to form a comprehensive optimization objective. Multi-dimensional rewards can include outcome rewards and process rewards. Outcome rewards are reward signals for the model's final output (such as insurance question-and-answer answers) during reinforcement learning training, used to evaluate the accuracy, relevance, security, and user experience of the answer. They can be combined with process rewards to optimize overall performance. Process rewards are reward signals for intermediate steps in the model's reasoning process during reinforcement learning training, used to guide improvements in the correctness of reasoning logic.

[0066] The target domain big model refers to a domain-specific model formed by injecting target domain expertise, business rules, and compliance requirements into a general big language model through post-training techniques. It can be used in scenarios such as target domain knowledge question answering.

[0067] For example, for each reinforcement learning training data, the processor inputs the reinforcement learning training data into a structured mind chain model of the target domain, calculates the group consistency loss through a group policy loss function that includes multi-dimensional rewards, and adjusts the model parameters of the structured mind chain model according to the group consistency loss to obtain a target model of the target domain.

[0068] In this embodiment, a first corpus of data in the target domain is acquired, and the first corpus data is processed into a structured thinking chain to obtain reasoning thinking chain corpus data. The initial large model is trained using the reasoning thinking chain corpus data to obtain a structured thinking chain large model in the target domain. This structured thinking chain large model possesses the structured thinking logic and professional knowledge understanding ability of the target domain. Then, a second corpus of data is acquired, and the second corpus data is processed to obtain reinforcement learning training data. The structured thinking chain large model in the target domain is trained using reinforcement learning with multi-dimensional rewards using reinforcement learning data to obtain a target large model in the target domain. This improves the business reasoning accuracy of the target large model in the target domain.

[0069] In an exemplary embodiment, the first corpus data of the target domain is processed by a structured thinking chain to obtain inference thinking chain corpus data, including: using the business logic and compliance rules of the target domain to process the first corpus data of the target domain into a structured thinking chain to obtain inference thinking chain corpus data.

[0070] For example, the processor utilizes the business logic and compliance rules of the target domain to perform structured thought chain processing on the first corpus data. This involves adding the business logic and compliance rules of the target domain to the analysis process, using the business logic to reason about the first corpus data, and then using the compliance rules to check it, resulting in reasoning thought chain corpus data. The reasoning thought chain corpus data can be text information with structured thought logic, containing intermediate reasoning process data between questions and answers.

[0071] By adding business logic and compliance rules from the target domain to the analysis process, the resulting reasoning thought chain corpus data contains textual information with structured thinking logic. The structured thinking chain model trained using this reasoning thought chain corpus data has the ability to understand the structured thinking logic and professional knowledge of the target domain.

[0072] In an exemplary embodiment, the first corpus data of the target domain is processed using structured thought chain processing based on the business logic and compliance rules of the target domain to obtain inference thought chain corpus data. This includes: performing semantic analysis on user questions in the first corpus data of the target domain to obtain semantic analysis results; obtaining knowledge data required to answer user questions; performing inference based on the semantic analysis results and knowledge data according to the business logic of the target domain to obtain logical inference results between questions and answers; and conducting a reasonableness assessment and compliance rule verification on the logical inference results between questions and answers to obtain inference thought chain corpus data.

[0073] Reasonableness assessment determines whether the logical reasoning between the question and the answer is correct. Compliance rule verification determines whether there are compliance risks in the logical reasoning between the question and the answer.

[0074] like Figure 3 As shown, the structured thinking chain module comprises three layers: a problem analysis layer, an information processing layer, and an output delivery layer. The structured thinking chain processing includes four core stages: premise verification, concept clarification, logical reasoning, and reflection and review. The problem analysis layer performs premise verification, which may involve semantic analysis of user questions in the first corpus data of the target domain to obtain semantic analysis results. The information processing layer performs concept clarification and logical reasoning. Concept clarification may involve obtaining the knowledge data needed to answer user questions; logical reasoning involves reasoning according to the business logic of the target domain based on the semantic analysis results and knowledge data to obtain the logical reasoning result between the question and the answer. The output delivery layer performs reflection and review. Reflection and review may involve evaluating the reasonableness and compliance rules of the logical reasoning result between the question and the answer to obtain the reasoning thinking chain corpus data. Knowledge data includes concepts and / or knowledge points.

[0075] Taking the insurance field as an example, the processing steps of the structured thinking chain may include: <step1 Premise Verification> Analyze the user's question and the implicit assumptions of the user's question; <step2 Concept Clarification> Obtain the concepts or knowledge points required to answer the user's question. <step3 Logical Reasoning> Reason according to the insurance business logic, such as calculation, qualification judgment, etc., and form a clear and accurate answer logic based on the reasoning process. <step4 Reflection and Review> Judge the reasonableness of the reasoning logic and whether there are content of compliance risks. If there are mistakes, re-analyze.

[0076] Through the process of premise verification - concept clarification - logical reasoning - reflection and review, the pre-trained model conducts step-by-step and in-depth analysis, effectively optimizing the problems of logical confusion and factual errors of the general model in complex tasks, providing a reliable reasoning framework for the model, significantly improving the stability and correctness of the answers, especially suitable for scenarios with complex professional knowledge, and making explicit the business logic rules and compliance rules as necessary steps in the reasoning analysis, making the reasoning process of the large model traceable and the compliance inspection programmable, ensuring the accuracy and security of the content generated by the large model's reasoning from the mechanism; generating reasoning thinking chain corpus data with structured thinking chains for the corpus data of the target field through the reasoning framework, and using this reasoning thinking chain corpus data to train the large model, enabling the large model to have the concept and structured reasoning logic ability of the target field, and improving the accuracy and stability of the structured thinking chain large model in the target field.

[0077] In an exemplary embodiment, training the initial large model based on the reasoning thinking chain corpus data to obtain the structured thinking chain large model in the target field includes: Supervising and training the initial large model based on the reasoning thinking chain corpus data in combination with the thinking chain supervision instruction to obtain the structured thinking chain large model in the target field.

[0078] Reasoning chain corpus data can include questions, thought chains, and answers. A question is a complex task to be solved. A thought chain refers to a detailed, step-by-step reasoning process or intermediate steps. The answer is a conclusion drawn from the reasoning process. Supervised Fine-Tuning with Chain of Thought is a training method aimed at improving the complex reasoning ability of language models. Its core idea is to teach pre-trained models how to perform step-by-step, explicit logical reasoning like humans, rather than directly generating the final answer, by using labeled data in the format of "question-reasoning process-answer". Supervised chain of thought instructions are a prompting technique aimed at improving the reliability and accuracy of large language models in complex reasoning tasks. Its core idea is to guide the model to self-evaluate its reasoning process while generating the final answer. Supervised chain of thought instructions include thought chain generation instructions and self-evaluation instructions. Thought chain generation instructions guide the model to generate detailed steps for solving the problem, i.e., the thought chain. Self-evaluation instructions guide the model to evaluate the generated thought chain, judging the correctness, completeness, and logical consistency of the reasoning process.

[0079] like Figure 4 As shown, a structured thought chain is performed on the professional knowledge text of the target domain, namely, premise verification, concept clarification, logical reasoning, and reflection and review, to obtain reasoning thought chain corpus data. Using the reasoning thought chain corpus data, a general model is trained under the thought chain supervision instruction to adjust the model parameters of the general model, so as to obtain a large structured thought chain model of the target domain.

[0080] By using reasoning thought chain corpus data combined with thought chain supervision instructions to supervise the training of the initial large model, we can reduce the "illusion" problem that may occur in the model during the reasoning process, reduce the probability of outputting wrong answers, improve the reliability of the final answer, and significantly improve the reliability of the model. Moreover, it not only shows the model's reasoning path, but also shows the model's judgment of its own reasoning quality, which greatly enhances the interpretability of the model. The self-evaluation generated by the model can serve as high-quality feedback data for further training and optimization of the model, forming a self-improving closed loop.

[0081] In an exemplary embodiment, training an initial large model based on inference thought chain corpus data to obtain a structured thought chain large model for the target domain includes: processing the inference thought chain corpus data for each inference thought chain corpus data using the initial large model to obtain target output loss and thought chain loss; obtaining multi-task joint loss based on target output loss and thought chain loss; and adjusting the parameters of the initial large model based on multi-task joint loss to obtain a structured thought chain large model for the target domain.

[0082] During the initial large model training, the target output loss and the thought chain loss are calculated using a preset loss function. The preset loss function includes parameters such as the target output sequence length, the thought chain sequence length, the input, the preceding output, the preceding thought chain, and the weight coefficients of the thought chain loss.

[0083] By adjusting the loss across multiple tasks, the performance of a large structured thought chain model in the target domain can be improved in complex tasks.

[0084] In an exemplary embodiment, the reasoning thought chain corpus data is processed using an initial large model to obtain the target output loss and the thought chain loss. This includes: treating each reasoning thought chain corpus data as a target sample; processing the target sample using the initial large model to obtain the target output sequence length, thought chain sequence length, input, preceding output, preceding thought chain, and weight coefficients of the thought chain loss; for each target sample, traversing each position of the target output sequence and calculating the first probability of generating the current word given the input and preceding output; obtaining the total number of samples and calculating the target output loss based on the total number of samples, the target output sequence length, and the first probability; for each target sample, traversing each position of the thought chain sequence and calculating the second probability of generating the current reasoning step given the input and preceding thought chain; and calculating the thought chain loss based on the total number of samples, the thought chain sequence length, the second probability, and the weight coefficients of the thought chain loss.

[0085] The total number of samples refers to the training batch size N. Calculate the logarithm of the first probability, then divide the logarithm by the target output sequence length, and then divide by the total number of samples to obtain the target output loss. Calculate the logarithm of the second probability, then divide the logarithm by the thought chain sequence length, and then divide by the total number of samples to obtain the thought chain loss.

[0086] For example, the preset loss function is as follows:

[0087] Formula (1)

[0088] Where N is the batch size for model training. This represents the target output sequence length of the i-th sample. This represents the length of the thought chain sequence of the i-th sample. This represents the model input for the i-th sample. This represents the thought chain sequence of the i-th sample. This represents the thought chain of the i-th sample at the t-th time step of the sequence. Let represent the preceding thought chain, λ represent the weight coefficients of the thought chain loss, θ represent the weights of the model to be trained, and P(.) represent the probability distribution of the model generating the next word in this sequence. This represents the target word (token) of the i-th sample at the t-th time step of the sequence. This represents the complete target output sequence corresponding to the i-th sample, consisting of multiple... There are two components, where t represents the time step in the sequence. This indicates the preceding output.

[0089] The calculation process of the target output loss includes: for the i-th sample, traversing each position t (i.e., time step) of the target output sequence, and calculating the loss given the input. and preceding output At that time, generate the current word element. The logarithm of the first probability, then for all words in the i-th sample (divided by) The average log-likelihood of the target output per word is obtained by averaging the log-likelihood of all samples (divided by N).

[0090] The calculation process of the thought chain loss includes: for the i-th sample, traversing each position t of its thought chain sequence, and calculating "given input..." and preceding thought chains At that time, generate the logarithm of the second probability of the current inference step; then divide all inference steps in the sample by . The weighted average of the log-likelihood of each word in the thought chain is obtained by averaging the log-likelihood of all samples (divided by N) and multiplying it by the weight λ of the thought chain.

[0091] The target output loss ensures the correctness of the model's generated results, while the thought chain loss forces the model to learn "how to reason," thereby improving interpretability and generalization ability for complex tasks. The hyperparameter λ is used to adjust the balance between the two.

[0092] In one exemplary embodiment, processing second corpus data in the target domain to obtain reinforcement learning training data includes: performing structured thought chain processing on the second corpus data in the target domain to obtain reinforcement learning training data.

[0093] Reinforcement learning training data can be textual information with structured logical thinking, including data on the intermediate reasoning process between questions and answers. A structured thought chain refers to employing a structured thinking framework during the reasoning process to improve the accuracy of the model's reasoning.

[0094] Structured thinking chain processing of the second corpus data in the target domain can include analyzing user questions in the second corpus data, obtaining knowledge data on responses to user questions, performing logical reasoning based on business logic, and judging the rationality of the reasoning logic and whether there is any content with compliance risks.

[0095] The second corpus data in the target domain undergoes structured thought chain processing to obtain reinforcement learning training data. This includes: utilizing the business logic and compliance rules of the target domain, the second corpus data in the target domain undergoes structured thought chain processing to obtain reinforcement learning training data. The reinforcement learning training data may be the same as or different from the reasoning thought chain corpus data.

[0096] In an exemplary embodiment, structured thought chain processing is performed on the second corpus data of the target domain to obtain reinforcement learning training data, including: performing semantic analysis on user questions in the second corpus data of the target domain to obtain semantic analysis results; obtaining knowledge data required to answer user questions; performing reasoning according to the business logic of the target domain based on the semantic analysis results and knowledge data to obtain the logical reasoning results between questions and answers; and conducting a reasonableness assessment and compliance rule verification on the logical reasoning results between questions and answers to obtain reinforcement learning training data.

[0097] Reasonableness assessment determines whether the logical reasoning between the question and the answer is correct. Compliance rule verification determines whether there are compliance risks in the logical reasoning between the question and the answer.

[0098] Through a process of premise verification, concept clarification, logical reasoning, and reflective review, the pre-model undergoes step-by-step and in-depth analysis, further providing a reliable reasoning framework for the model and significantly improving the stability and correctness of the answers.

[0099] In an exemplary embodiment, the multi-dimensional reward includes outcome reward and process reward; the target domain's structured thinking chain model is trained using reinforcement learning with multi-dimensional rewards based on reinforcement learning training data to obtain the target domain's target model, including: optimizing the model parameters of the target domain's structured thinking chain model based on reinforcement learning training data using a group policy optimization algorithm that includes outcome reward and process reward, to obtain the target domain's target model.

[0100] Outcome rewards can include correctness rewards, compliance rewards, and format rewards. Correctness rewards verify whether the model output conforms to knowledge. Compliance rewards verify whether the model output violates grammatical rules. Format rewards are a reward mechanism designed to ensure that the model output conforms to specific structural or format requirements. Process rewards can include inference format rewards and process accuracy rewards. Inference format rewards ensure that the model performs correct inference and is placed in a correct manner. <think>The process accuracy reward validates the model's reasoning according to a pre-defined thought process chain. For example... Figure 5 As shown, a multi-dimensional reward mechanism module is provided. This module includes outcome rewards and process rewards. Outcome rewards are used to evaluate the model's output answer, while process rewards are used to evaluate the reasoning process of the model's output. This reward mechanism can simulate human evaluation of the model's output, thereby providing refined guidance in reinforcement learning training and significantly improving the model's reasoning ability, output quality, and user experience in complex tasks.

[0101] The loss function used when training a large structured thought chain model can be a group policy loss function that includes both outcome rewards and process rewards. The group policy loss function can be a group policy optimization algorithm (GRPO).

[0102] The specific calculation process of the GRPO algorithm based on multi-dimensional rewards is as follows: Formula (2):

[0103] Formula (2)

[0104] θ represents the policy parameters, which are the weight parameters of the model to be optimized. This represents the state, the observational information or situation presented to the model by the environment or task at time step t. Indicates an action, in a state The model then selects the operation to execute based on the current strategy. The current policy is represented by the policy function with parameter θ, which indicates the state. Select action The probability distribution. This indicates the old strategy, i.e., the strategy used during sampling, which is used to calculate the importance sampling ratio to evaluate the change of the new strategy relative to the old strategy. This represents the advantage function, used to evaluate the performance of a given state. Next action How good is the performance relative to the average performance in this state? The advantage function A_t is decomposed into a weighted sum of process reward and outcome reward. Weighting coefficients representing process advantages Weighting coefficients representing the advantage of the outcome; It represents process advantage, measuring the advantage of the current action relative to the average process reward; It represents the advantage of the outcome, measuring the advantage of the current state relative to the final reward. This represents the pruning hyperparameter, which typically ranges from 0.1 to 0.3. It limits the update range between the old and new strategies, thereby improving training stability. The entropy regularization coefficient is used to encourage policy exploration and prevent premature policy convergence; H represents policy entropy; η represents the group loss weight. This represents the loss of group consistency. This represents the process reward obtained at the k-th time step, which is accumulated from time step t to the end through a discount factor. This represents the process value function, used to estimate the value of a process in a given state. The expected return of future process rewards is used as a baseline when calculating process advantage. It is the loss function for generalized proximal policy optimization. It is the expectation operator, representing the expectation of a policy. All possible state-action pairs ( , Take the expected value. It is in state Next, according to the strategy Take action The probability of. It is in state Next, according to the old strategy Take action The probability of. It is a clipping function that will... Limited to the range That is, if Less than Then return ,if Greater than Then return Otherwise return . It is the policy entropy, representing the policy. In state Uncertainty or diversity under the circumstances.

[0105] like Figure 6 As shown, a structured thinking chain is performed on the professional knowledge text of the target domain, namely premise verification, concept clarification, logical reasoning, and reflection and review, to obtain reinforcement learning training data. The structured thinking chain model is trained using the reinforcement learning training data and a multi-dimensional reward mechanism, and the model parameters of the structured thinking chain model are adjusted to obtain the target model of the target domain.

[0106] The following describes the training method for the target domain model, taking the insurance domain as an example. The implementation environment for the insurance domain model training method provided in this application includes the following hardware components: an AI high-performance computing cluster, a distributed storage system, and a GPU inference server. The AI ​​high-performance computing cluster can be configured with multiple high-performance servers locally, each equipped with eight NVIDIA H20 GPU cards, constructing a distributed training architecture to support parallel training of large-scale models. Each H20 GPU possesses high computing performance and large memory bandwidth. The configuration of a single server (computing node) can be as follows:

[0107]

[0108] Distributed storage systems can employ highly scalable distributed storage architectures to centrally manage insurance domain knowledge bases, model training datasets, and model weight files, providing a stable and efficient data foundation for large-scale data processing and model development. GPU inference servers can be dedicated inference clusters built on high-performance GPU architectures, providing high-concurrency, low-latency model inference capabilities for online services, ensuring the real-time performance and stability of online services.

[0109] like Figure 7 As shown, the system architecture for training insurance domain models comprises two stages: data construction and model training. The system architecture includes a structured thinking chain construction module, a structured thinking chain supervised instruction training module, a multi-dimensional reward mechanism module, and a multi-dimensional reward reinforcement learning training module. The structured thinking chain construction module processes insurance domain professional knowledge text to obtain text information with structured thinking logic. Then, the structured thinking chain supervised instruction training module uses this structured thinking logic text information to train a general-purpose large model, enabling the general-purpose large model to possess the structured thinking logic and insurance professional knowledge understanding capabilities of the insurance domain, resulting in a structured thinking chain large model for the insurance domain. Finally, the multi-dimensional reward reinforcement learning training module uses reinforcement learning training data combined with a multi-dimensional reward mechanism to train the structured thinking chain large model, resulting in the target large model for the insurance domain.

[0110] Methods for training insurance domain models using system architecture include:

[0111] (1) Obtain the first corpus data in the insurance field, perform semantic analysis on the user questions in the first corpus data in the insurance field, obtain the semantic analysis results, and obtain the knowledge data required to answer the user questions; based on the semantic analysis results and knowledge data, perform reasoning according to the insurance business logic to obtain the logical reasoning results between the questions and answers; conduct a reasonableness assessment and compliance rule test on the logical reasoning results between the questions and answers to obtain the reasoning thought chain corpus data.

[0112] The first corpus data can be professional knowledge texts in the insurance field, or high-quality insurance data generated by cleaning, screening, and compliance assessment of professional knowledge texts in the insurance field, combined with data augmentation technology and the construction of positive and negative sample pairs.

[0113] (2) Based on the reasoning thought chain corpus data, the initial large model is trained under supervision by combining the thought chain supervision instructions to obtain a structured thought chain large model in the insurance field;

[0114] like Figure 8 As shown, a structured thought chain is applied to professional knowledge texts in the insurance field, namely, premise verification, concept clarification, logical reasoning, and reflective review, to obtain reasoning thought chain corpus data. Using the reasoning thought chain corpus data, a general model is trained under the thought chain supervision instruction. Multi-task joint loss is calculated through a preset loss function. Based on the multi-task joint loss, the model parameters of the general model are adjusted, such as merging model weights, to obtain a large structured thought chain model in the insurance field.

[0115] (3) The second corpus data in the insurance field is processed by structured thinking chain to obtain reinforcement learning training data;

[0116] (4) Based on reinforcement learning training data, the model parameters of the structured thinking chain model in the insurance field are optimized by a group policy optimization algorithm that includes outcome rewards and process rewards, so as to obtain the target model in the insurance field.

[0117] The multi-dimensional rewards include process rewards and outcome rewards. By employing reinforcement learning training methods combined with multi-dimensional rewards, a structured thinking chain-based insurance domain model with higher accuracy and lower compliance risk is obtained, serving as the target model for the insurance industry.

[0118] like Figure 9 As shown, a structured thinking chain is applied to professional knowledge texts in the insurance field, namely, premise verification, concept clarification, logical reasoning, and reflective review, to obtain reinforcement learning training data. The structured thinking chain model is trained using the reinforcement learning training data and a multi-dimensional reward mechanism. The group policy loss is calculated through the group policy loss function, and the model parameters of the structured thinking chain model are adjusted through the group policy loss to obtain the target model in the insurance field.

[0119] In this embodiment, firstly, for professional textual knowledge in the insurance field, a four-step structured thinking chain reasoning framework of "premise verification - concept clarification - logical reasoning - reflection and review" is constructed to explicitly incorporate insurance business logic and compliance rules into the reasoning logic, generating insurance conversation corpus with a structured thinking chain. Secondly, combined with professional knowledge corpus in the insurance field, a pre-trained large model is trained under supervised instruction of the structured thinking chain, enabling the large model to master insurance concepts and reasoning abilities, while also possessing structured reasoning logic capabilities. Finally, a multi-dimensional reward reinforcement learning approach is used to improve the correctness of the large model's insurance business reasoning, reduce insurance compliance risks, and lower user comprehension costs in insurance intelligent customer service scenarios (such as insurance Q&A scenarios), thereby improving user experience satisfaction. Combining process rewards (rewarding the correctness of each step of the thinking chain) and result rewards (rewarding the correctness of the final answer) guides the model to not only pursue the correct answer but also the correct reasoning process. This solves the logical bias and "shortcut" problems that may be caused by the single reward signal in traditional reinforcement learning schemes. Through two clearly defined and sequentially connected training phases, this approach systematically builds the model's domain expert capabilities, effectively avoiding the "rote memorization" and overfitting problems caused by single-supervised training. It also addresses the issue of reinforcement learning easily "fabricating" data when the knowledge base is weak. This solution overcomes the problems of supervised training being prone to overfitting to static data, lacking dynamic optimization and deep reasoning capabilities, and struggling to handle complex insurance cases. It also overcomes the strong reliance on rewards, logical biases, and lack of deep reasoning support for complex insurance clauses in reinforcement learning. Furthermore, it solves the problem that general models may skip key verification steps (such as clause premise checks), leading to incorrect or high-risk answers (such as misinterpreting insurance liability). Based on a structured thinking chain, the target-oriented large-scale model for the insurance domain was comprehensively evaluated on multiple self-built evaluation datasets. These datasets cover core insurance tasks such as insurance product inquiry, claims process analysis, clause understanding and compliance judgment, personalized user consultation, and complex scenario reasoning, comprehensively examining the model's performance in terms of professionalism, accuracy, and practicality. The results show that the target large model has achieved significant improvements in tasks such as question answering in the insurance field, outperforming existing general models (deepseek, hunyuan-turbos, etc.) by 5-10%.

[0120] Furthermore, in intelligent customer service scenarios, the trained target model in the insurance field effectively integrates recalled knowledge corpora with generative responses, ensuring the accuracy and logical consistency of the question-and-answer content. Through premise verification, the system filters and confirms the context of user questions and the recalled knowledge; through concept clarification, it further analyzes insurance terminology and semantics; through logical reasoning, it constructs a rigorous reasoning process from facts to conclusions; and finally, through reflection and review, it conducts quality assessment and compliance checks on the generated results. This design significantly improves the model's reasoning ability, especially in terms of accuracy, professionalism, and user experience in complex insurance scenarios.

[0121] In one exemplary embodiment, the method further includes: acquiring an insurance question input by the user; analyzing the insurance question using a target large model to obtain an answer containing a thought process chain, and displaying it to the user. The large model of the insurance domain obtained through training can provide correct and stable answers.

[0122] It is understandable that the above description uses the insurance field as an example. Similar processing can be used in the financial, medical, educational, or transportation fields. The first and second corpora can be implemented using multimedia data corresponding to the respective field. For example, in the financial field, the first and second corpora can be multimedia data on financial transactions. In the medical field, the first and second corpora can be multimedia data on medical knowledge or consultation. In the transportation field, the first and second corpora can be multimedia data on transportation knowledge. In the education field, the first and second corpora can be multimedia data on educational knowledge. Multimedia data can include one or more of the following: text data, image data, audio data, and video data. Multimedia data can include question-and-answer data.

[0123] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps. It is understood that the steps in different embodiments can be freely combined as needed, and all non-contradictory solutions formed by such combinations are within the scope of protection of this application.

[0124] Based on the same inventive concept, this application also provides an apparatus for implementing the method described above. The solution provided by this apparatus is similar to the solution described in the above method; therefore, specific limitations in one or more apparatus embodiments provided below can be found in the limitations of the method described above, and will not be repeated here.

[0125] In one exemplary embodiment, such as Figure 10 As shown, a training device 1000 for a target domain model is provided. The device includes: a structured thought chain module 1002, a first training module 1004, a training data determination module 1006, and a second training module 1008.

[0126] The structured thinking chain module 1002 is used to acquire the first corpus data of the target domain, and to perform structured thinking chain processing on the first corpus data of the target domain to obtain reasoning thinking chain corpus data.

[0127] The first training module 1004 is used to train the initial large model based on the reasoning thought chain corpus data to obtain a structured thought chain large model of the target domain.

[0128] The training data determination module 1006 is used to acquire the second corpus data of the target domain and process the second corpus data of the target domain to obtain reinforcement learning training data.

[0129] The second training module 1008 is used to perform multi-dimensional reward reinforcement learning training on the structured thinking chain model of the target domain based on the reinforcement learning training data, so as to obtain the target model of the target domain.

[0130] In an exemplary embodiment, the structured thinking chain module 1002 is further configured to utilize the business logic and compliance rules of the target domain to perform structured thinking chain processing on the first corpus data of the target domain to obtain inference thinking chain corpus data.

[0131] In an exemplary embodiment, the structured thinking chain module 1002 is further configured to perform semantic analysis on user questions in the first corpus data of the target domain to obtain semantic analysis results; obtain knowledge data required to answer the user questions; perform reasoning according to the business logic of the target domain based on the semantic analysis results and the knowledge data to obtain logical reasoning results between questions and answers; and perform a reasonableness assessment and compliance rule verification on the logical reasoning results between questions and answers to obtain reasoning thinking chain corpus data.

[0132] In an exemplary embodiment, the first training module 1004 is used to supervise the training of the initial large model based on the reasoning thought chain corpus data and combined with the thought chain supervision instructions to obtain a structured thought chain large model of the target domain.

[0133] In an exemplary embodiment, the first training module 1004 is further configured to process the reasoning thought chain corpus data for each reasoning thought chain corpus data using an initial large model to obtain a target output loss and a thought chain loss; obtain a multi-task joint loss based on the target output loss and the thought chain loss; and adjust the parameters of the initial large model based on the multi-task joint loss to obtain a structured thought chain large model for the target domain.

[0134] In an exemplary embodiment, the first training module 1004 is further configured to treat each inference thought chain corpus data as a target sample, process the target sample through an initial large model, and obtain the target output sequence length, thought chain sequence length, input, preceding output, preceding thought chain, and weight coefficients of thought chain loss for the target sample; for the target sample, traverse each position of the target output sequence and calculate the first probability of generating the current word given the input and preceding output; obtain the total number of samples, and calculate the target output loss based on the total number of samples, the target output sequence length, and the first probability; for the target sample, traverse each position of the thought chain sequence and calculate the second probability of generating the current inference step given the input and preceding thought chain; and calculate the thought chain loss based on the total number of samples, the thought chain sequence length, the second probability, and the weight coefficients of thought chain loss.

[0135] In an exemplary embodiment, the training data determination module 1006 is further configured to perform structured thought chain processing on the second corpus data of the target domain to obtain reinforcement learning training data. Exemplarily, the training data determination module 1006 is also configured to utilize the business logic and compliance rules of the target domain to perform structured thought chain processing on the second corpus data of the target domain to obtain reinforcement learning training data. The reinforcement learning training data may be the same as or different from the inference thought chain corpus data.

[0136] In an exemplary embodiment, the training data determination module 1006 is further configured to perform semantic analysis on user questions in the second corpus data of the target domain to obtain semantic analysis results; obtain knowledge data required to answer user questions; perform reasoning according to the business logic of the target domain based on the semantic analysis results and knowledge data to obtain logical reasoning results between questions and answers; and conduct a reasonableness assessment and compliance rule verification on the logical reasoning results between questions and answers to obtain reinforcement learning training data.

[0137] In an exemplary embodiment, the multi-dimensional reward includes outcome reward and process reward; the second training module 1008 is further configured to optimize the model parameters of the structured thought chain model of the target domain based on reinforcement learning training data using a group policy optimization algorithm that includes outcome reward and process reward, so as to obtain the target large model of the target domain.

[0138] In one exemplary embodiment, the target field is the insurance field, the financial field, the medical field, the education field, or the transportation field.

[0139] Each module in the above-mentioned device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of a computer device in hardware form or independent of it, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0140] In one exemplary embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 11 As shown, this computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores data. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. When the computer program is executed by the processor, it implements a method.

[0141] Those skilled in the art will understand that Figure 11 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0142] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the method described above.

[0143] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the above-described method.

[0144] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the methods described above.

[0145] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0146] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.

[0147] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0148] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.< / think>

Claims

1. A method for training a target domain model, characterized in that, The method includes: Obtain the first corpus data of the target domain, and perform structured thinking chain processing on the first corpus data of the target domain to obtain reasoning thinking chain corpus data; The initial large model is trained based on the reasoning thought chain corpus data to obtain the structured thought chain large model of the target domain; Obtain second corpus data in the target domain, and process the second corpus data in the target domain to obtain reinforcement learning training data; Based on the reinforcement learning training data, a multi-dimensional reward reinforcement learning training is performed on the structured thinking chain model of the target domain to obtain the target model of the target domain.

2. The method according to claim 1, characterized in that, The first corpus data in the target domain is processed using structured thought chain processing to obtain inference thought chain corpus data, including: By utilizing the business logic and compliance rules of the target domain, the first corpus data of the target domain is processed using structured thinking chain processing to obtain reasoning thinking chain corpus data.

3. The method according to claim 2, characterized in that, The process involves using the business logic and compliance rules of the target domain to perform structured thought chain processing on the first corpus data of the target domain, resulting in inference thought chain corpus data, including: Semantic analysis is performed on user questions in the first corpus data of the target domain to obtain semantic analysis results; Obtain the knowledge data needed to answer the user's question; Based on the semantic analysis results and the knowledge data, reasoning is performed according to the business logic of the target domain to obtain the logical reasoning results between the question and the answer. The logical reasoning results between the questions and answers are evaluated for reasonableness and compliance with rules to obtain reasoning chain corpus data.

4. The method according to claim 1, characterized in that, The process of training the initial large model based on the reasoning thought chain corpus data to obtain the structured thought chain large model of the target domain includes: Based on the reasoning thought chain corpus data, and combined with the thought chain supervision instructions, the initial large model is trained under supervision to obtain a structured thought chain large model for the target domain.

5. The method according to claim 1, characterized in that, The process of training the initial large model based on the reasoning thought chain corpus data to obtain the structured thought chain large model of the target domain includes: For each reasoning thought chain corpus data, the reasoning thought chain corpus data is processed through an initial large model to obtain the target output loss and thought chain loss; The multi-task joint loss is obtained based on the target output loss and the thought chain loss. Based on the multi-task joint loss, the parameters of the initial large model are adjusted to obtain the structured thought chain large model of the target domain.

6. The method according to claim 1, characterized in that, The reinforcement learning training data is obtained by processing the second corpus data in the target domain, including: The second corpus data in the target domain is processed using structured thought chain processing to obtain reinforcement learning training data.

7. The method according to claim 1, characterized in that, Multi-dimensional rewards include outcome rewards and process rewards; the multi-dimensional reward reinforcement learning training of the structured thought chain model of the target domain based on the reinforcement learning training data, to obtain the target domain's target model, includes: Based on the reinforcement learning training data, the model parameters of the structured thought chain model in the target domain are optimized using a group policy optimization algorithm that includes outcome rewards and process rewards, thereby obtaining the target model in the target domain.

8. The method according to any one of claims 1 to 7, characterized in that, The target areas are insurance, finance, healthcare, education, or transportation.

9. A training device for a target domain model, characterized in that, The device includes: The structured thinking chain module is used to acquire the first corpus data of the target domain, and to perform structured thinking chain processing on the first corpus data of the target domain to obtain reasoning thinking chain corpus data. The first training module is used to train the initial large model based on the reasoning thought chain corpus data to obtain a structured thought chain large model for the target domain. The training data determination module is used to acquire the second corpus data of the target domain and process the second corpus data of the target domain to obtain reinforcement learning training data. The second training module is used to perform multi-dimensional reward reinforcement learning training on the structured thinking chain model of the target domain based on the reinforcement learning training data, so as to obtain the target model of the target domain.

10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 8.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 8.

12. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 8.