Strategy question-answering system and method based on vertical domain large model

By fine-tuning large language models and introducing low-rank structures, the problems of weak answering ability and training disadvantages in the application of large models in vertical fields are solved, achieving higher question-and-answer accuracy and correlation, as well as better generalization ability.

CN120104730APending Publication Date: 2025-06-06ENG UNIV OF THE CHINESE PEOPLES ARMED POLICE FORCE
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202510070776.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The existing large models lack domain data sets in vertical fields, poor data quality, weak answering ability of large models, and obvious disadvantages in indicators such as training convergence accuracy, training time and memory usage.

Method used

By receiving natural language input from users, searching open source data sets, cleaning and preprocessing, fine-tuning the pre-processed basic large language model using training data, introducing a low-rank structure to modify its weight matrix, forming the final model.

Benefits of technology

It improves the model's understanding of terms and concepts in specific fields, enhances the accuracy and relevance of question-and-answer, reduces the complexity of the model, reduces the risk of overfitting, and improves generalization ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104730A_ABST
    Figure CN120104730A_ABST
Patent Text Reader

Abstract

The invention discloses a strategy question-answering system and method based on a vertical domain large model, and the method comprises the steps: S1, receiving the natural language input of a user, and retrieving an open source data set to obtain an original data set with the comprehensive similarity meeting the requirements; s2, performing necessary data cleaning and preprocessing on the original data set to obtain a training set; s3, performing fine tuning on the pre-processed basic large language model by adopting the training data set, and introducing a low-rank structure to modify a weight matrix to obtain a final model; s4, the problem enters a final model for post-processing to generate prediction output; the system takes a large language basic model subjected to fine tuning training as a core, introduces a low-rank structure to perform fine tuning on model parameters, reduces parameter quantity required by training while keeping model performance, reduces model complexity, reduces an overfitting risk, improves generalization ability, supports deep fusion of a specific field knowledge base and universal field data, and has a wide application prospect. And the accuracy and correlation of questions and answers are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a strategic question answering technology, and in particular to a strategic question answering system and method based on a vertical field big model. Background Art

[0002] Disruptive technology clusters represented by artificial intelligence are influencing a new round of intelligent technology changes with unprecedented breadth and depth. In the field of natural language processing, large language models have made significant progress with their powerful language understanding and generation capabilities. In 2017, Google launched the Transformer network architecture for processing natural language tasks. The following year, OpenAI released GPT-1 (Generative Pre-trained Transformer-1), which can generate fluent natural language text. In 2022, OpenAI released the large language model ChatGPT, which has a very high level of human-computer interaction and a very wide range of application scenarios, which has attracted widespread attention from the whole society. These large models such as BERT, GPT, T5, etc. are usually based on the Transformer architecture and are pre-trained on large-scale unlabeled text data to learn general language representations, and then fine-tuned on downstream tasks through supervised learning, achieving excellent results.

[0003] Nowadays, artificial intelligence has a wide range of application needs in various natural language processing fields. For example, physical training is a basic way to enhance physical fitness. However, in the actual training process, training injuries are frequent due to factors such as unscientific organizational methods, inadequate protective measures, irregular training movements, and poor psychological quality. Therefore, for intelligent question-and-answer needs such as training injury prevention and treatment, it is of great significance to give full play to the advantages of big models in vertical fields and realize more professional and effective question-and-answer "assistants". However, the existing big models lack domain data sets, have poor data quality, weak answering capabilities, and have obvious disadvantages in indicators such as convergence accuracy, training time, and video memory occupancy of big model training. Summary of the invention

[0004] The purpose of the present invention is to provide a strategic question-answering system and method based on a vertical field big model, which can provide auxiliary information for the question-answering process based on domain-specific knowledge bases and general domain knowledge bases, and enhance the model's understanding of specific field terms and concepts.

[0005] The present invention is achieved through the following technical solutions:

[0006] A strategic question-answering method based on a vertical domain big model, including:

[0007] S1: Receive natural language input from the user, retrieve open source datasets to obtain original datasets whose comprehensive similarity meets the requirements;

[0008] S2: Perform necessary data cleaning and preprocessing on the original data set to obtain the training set;

[0009] S3: Use the training dataset to fine-tune the pre-processed basic large language model, introduce a low-rank structure to modify its weight matrix to obtain the final model;

[0010] S4: The problem enters the final model for post-processing to generate prediction output.

[0011] Preferably, S2 also includes combing relevant documents and network resources disclosed on the Internet to construct and integrate high-quality data question-answer pairs related to the problem.

[0012] Preferably, the full parameter fine-tuning method is used in S3 to form a series of empirical hyperparameters that affect the accuracy and convergence speed of the model, and the gradient G of the weight matrix W (W∈R m×n ) Slowly changing low-rank structure, calculate two projection matrices P∈R m×r and Q∈R n×r , project the gradient matrix G into a low-rank form P T GQ,

[0013]

[0014] Where η is the learning rate, W T is the weight matrix, W 0 is the initial weight matrix, G t is the gradient matrix, P t and Q t are two projection matrices, which are used to perform full-parameter learning by establishing multiple subspaces and switching between different subspaces during training.

[0015] W t =W 0 +ΔW T1 +ΔW T2 +ΔW Tn

[0016] Among them, W t is the weight matrix, W 0 is the initial weight matrix, ΔW Tn is the weight matrix updated in each round of iteration.

[0017] Preferably, the large language base model in S3 adopts the Qwen2-0.5B-Instruct base model, and based on the Transformer architecture, the expressiveness of the model is enhanced by the SwiGLU activation function, and it is trained on a context length of 32k, and expanded to a longer context length through technologies such as Dual Chunk Attention.

[0018] Furthermore, the Qwen2-0.5B-Instruct base model also uses QKV bias and group query attention to enhance the model's understanding of complex language structures and its ability to process natural language.

[0019] A strategic question-answering system based on a large vertical domain model, including a human-computer interaction module, a data retrieval module, a data processing module and a model training module;

[0020] The human-computer interaction module is used to receive natural language input from the user and provide corresponding output;

[0021] The data retrieval module is used to retrieve and collect original data sets whose comprehensive similarity meets the requirements;

[0022] The data processing module is used to clean and preprocess the original data set to obtain a training set, which is divided into training data, verification data and test data;

[0023] The model training module is used to fine-tune the pre-trained large language base model.

[0024] Furthermore, the original data set includes public literature resources, network resources and field databases.

[0025] Compared with the prior art, the present invention has the following beneficial effects:

[0026] (1) The present invention proposes to use a fine-tuned large language basic model as the core to understand and generate natural language input, introduce a low-rank structure to fine-tune the model parameters, maintain the model performance while reducing the number of parameters required for training, reduce the model complexity and reduce the risk of overfitting, improve the generalization ability, support the deep integration of specific domain knowledge bases and general domain data, and improve the accuracy and relevance of questions and answers. The large model trained by the proposed technology has obvious advantages in training professional knowledge understanding in the field of injury prevention and treatment and overcoming the "illusion" of large model answers. The relevant results can provide a reference for the design and application of large model systems for vertical field question answering;

[0027] (2) The large language base model is trained through the SwiGLU activation function to enhance the model's expressiveness, and adopts QKV bias and group query attention to improve the model's understanding of complex language structures and its ability to process natural language. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 is an information processing flow chart of the present invention;

[0029] Figure 2 It is the overall architecture diagram of the system of the present invention;

[0030] Figure 3 Update graph for weight matrix;

[0031] Figure 4 A training data format diagram for the system of the present invention;

[0032] Figure 5 is the training loss curve of different fine-tuning models in this embodiment. DETAILED DESCRIPTION

[0033] The present invention is further described in detail below in conjunction with specific embodiments, which are intended to explain the present invention rather than to limit it.

[0034] refer to Figure 1 and Figure 2 As shown, the present invention provides a strategy question answering method based on a vertical field big model, including:

[0035] S1: Receive natural language input from the user, retrieve open source datasets to obtain original datasets whose comprehensive similarity meets the requirements;

[0036] S2: Perform necessary data cleaning and preprocessing on the original data set to obtain the training set;

[0037] S3: Use the training dataset to fine-tune the pre-processed basic large language model, introduce a low-rank structure to modify its weight matrix to obtain the final model;

[0038] S4: The problem enters the final model for post-processing to generate prediction output.

[0039] A strategic question-answering system based on a large model in a vertical field includes a human-computer interaction module, a data retrieval module, a data processing module, and a model training module. The specific architecture of the system consists of a top-down question-answering layer, a training layer, a model layer, and a data layer. The question-answering layer is the interface for direct interaction between the system and the user, responsible for receiving the user's natural language input and providing corresponding output. The training layer is the core of the performance optimization of the large model, providing strategies and tools for model optimization, and adapting to various fine-tuning methods such as GaLore strategy full parameter fine-tuning, LoRA, P-Tuning, etc., so that the large model can adapt to specific application scenario knowledge and data. Features: The training layer relies on powerful hardware computing support and fine-tuning tools to ensure the efficiency and stability of the training process. The design of the model fully considers the diversity and complementarity of the model to meet different user needs and application scenarios, and supports major large language models such as Tongyi Qianwen, ChatGLM and Llama. Through fine-tuning of the training layer, the performance of the adapted pre-trained model on specific tasks can be further improved. The data layer focuses on the diversity and coverage of the data, including strictly screened and pre-processed public literature resources, network resources and domain databases, which can provide rich and high-quality training data for large language models.

[0040] In the information processing flow of the strategic question-answering system, the fine-tuned large language model is used as the core to understand and generate natural language input, and supports the deep integration of specific domain knowledge bases and general domain data to improve the accuracy and relevance of questions and answers. Figure 1 As shown, the system receives natural language questions input by users. After cleaning and standardization preprocessing, the questions enter the encoder-decoder architecture of the large language model to complete the encoding conversion of the input questions to high-dimensional word vectors. The encoded word vectors can capture the semantic information of the questions. The large language model uses prediction mechanisms and contextual information to generate prediction outputs. In addition, based on domain-specific knowledge bases and general domain knowledge bases, it can provide auxiliary information for the question-answering process and enhance the model's understanding of domain-specific terms and concepts.

[0041] For the data layer of the strategic question-answering system architecture, training data sets related to vertical field tasks are retrieved, collected and prepared, and necessary data cleaning and preprocessing are performed to ensure data set quality and annotation accuracy. Taking the data in the field of training injury prevention and treatment as an example, the open source Chinese medical data set cMedQA2 is used, which is divided into training data, verification data and test data, with a total of 108,000 questions and 203,569 answers. In the specific implementation, after data cleaning, more than 2,000 data items are finally screened and retained from the original data entries. The content of these data is closely related to the diagnosis, treatment and prevention of training injuries. At the same time, by combing through relevant literature and network resources disclosed on the Internet, more than 50 high-quality injury data question-answer pairs related to training practice are constructed and integrated, and data enhancement of the question-answer pairs is completed, which enhances the pertinence and practicality of the data set and ensures the quality of the data set and the reliability of the analysis results.

[0042] For the model layer of the strategic question-answering system architecture, in terms of comprehensive computing power requirements and Chinese language reasoning capabilities, in terms of open source large language models, Qwen2 can be used as a basic large language model, and Qwen2-0.5B-Instruct is an instruction fine-tuning language model with 50 million parameters in the Qwen2 series. Based on the Transformer architecture, the SwigLU activation function is used to enhance the model's expressiveness, and it is trained on a 32k context length, and expanded to a longer context length through technologies such as Dual ChunkAttention. In addition, the Qwen2-0.5B-Instruct base model also uses QKV bias and group query attention to enhance the model's understanding of complex language structures and its ability to process natural language.

[0043] For the training layer of the strategic question-answering system architecture, a specific data set is used to further train the pre-trained basic large language model to adapt the model to specific tasks or fields. Specifically, a full-parameter fine-tuning method based on the gradient low-rank projection strategy is adopted to form a series of empirical hyperparameters that affect the accuracy and convergence speed of the model. The gradient G(W∈R m×n ) The slowly changing low-rank structure, that is, the rank of the gradient matrix G will become lower during the training process, and the gradient matrix can be approximated by a smaller subspace, by calculating the two projection matrices P∈R m×r and Q∈R n×r , project the gradient matrix G into a low-rank form P T GQ,

[0044]

[0045] Where η is the learning rate, W T is the weight matrix, W 0 is the initial weight matrix, G t is the gradient matrix, P t and Q t are two projection matrices, which are used to perform full-parameter learning by establishing multiple subspaces and switching between different subspaces during training.

[0046] W t =W 0 +ΔW T1 +ΔW T2 +ΔW Tn

[0047] Among them, W t is the weight matrix, W 0 is the initial weight matrix, ΔW Tn is the weight matrix updated in each round of iteration.

[0048] Specific implementation example: Since the amount of computational effort required for fine-tuning large language models is usually very large, the full-parameter fine-tuning of Qwen2-0.5B-Instruct based on the GaLore strategy still requires 17 to 20 GB of video memory for training under the condition of batch_size of 8. Therefore, in order to facilitate the promotion of the experience of full-parameter fine-tuning of large models, the experimental case environment adopts the Alibaba Cloud Artificial Intelligence Platform (PAI), configured with a single-card NVIDIA A10-24GB GPU. The required software dependency libraries and experimental recommended versions are shown in Table 1.

[0049] Table 1 Dependencies and version numbers

[0050]

[0051] The large model fine-tuning training tool uses the LLaMA-Factory platform, and the training data uses the OpenAI format, such as Figure 4 As shown in the figure, the dataset is the open source Chinese medical dataset cMedQA2, which has undergone rigorous data cleaning and manual data enhancement, totaling more than 2,000 records.

[0052] Taking Qwen2-0.5B-Instruct as the basic model, the fine-tuning of GaLore strategy parameters and LoRA fine-tuning are compared. The experience of fine-tuning hyperparameters is that a lower learning rate helps the model learn new tasks better while retaining the knowledge of old tasks. A larger batch size will make the model more likely to forget the knowledge of old tasks. Therefore, by adjusting the learning rate, batch size and number of training rounds, the amount of training parameters, training time and video memory usage during training are adjusted to optimize model performance. The hyperparameter settings in the experiment are shown in Table 2.

[0053] Table 2 Training parameters

[0054]

[0055] By setting different ranks, the training time, memory usage, and loss comparison results of the GaLore strategy full parameter fine-tuning and LoRA fine-tuning are shown in Table 3. From the perspective of training time, as the Rank value of the GaLore strategy increases, the training time increases slightly, but the overall change is not large, remaining at around 11 minutes. This shows that the training efficiency of the GaLore strategy is relatively stable under different Rank values, while the training time of LoRA is significantly longer, reaching 17.66 minutes, which may be because LoRA requires more computing resources during the fine-tuning process.

[0056] In terms of video memory usage, as the Rank value of the GaLore strategy increases, the video memory usage also increases accordingly. This is because a higher Rank value means an increase in model parameters, and more video memory is required to store these parameters. LoRA has a slightly higher video memory usage than GaLore at the same Rank value. This may be because the parameter structure or optimization algorithm of the LoRA model leads to slightly different video memory usage efficiency.

[0057] In addition, loss is an important indicator for measuring model performance. The lower the loss, the higher the prediction accuracy of the model. From the data, the loss of the GaLore strategy gradually decreases as the Rank value increases, which shows that the prediction accuracy of the model increases with the increase in the number of parameters, while LoRA has a higher loss at the same Rank value, which may mean that the LoRA model does not achieve the best performance under the current fine-tuning strategy.

[0058] In general, the GaLore strategy shows good training efficiency and memory usage efficiency under different Rank values, and as the Rank value increases, the prediction accuracy of the model improves. LoRA takes a long time to train at the same Rank value, occupies a slightly higher memory, and has a higher loss, which proves that LoRA's performance is not as good as the GaLore strategy under the current fine-tuning strategy.

[0059] Table 3 Analysis of training effect

[0060]

[0061] Figure 5 The loss curve of the entire training process is shown in detail. The X-axis represents the number of iterations in the training process, and the Y-axis represents the loss value of the model at each global step. It can be seen that the training loss of the GaLore strategy decreases faster, and the four different fine-tuning models tend to be stable after 120 steps of training, indicating that under the current training conditions, the GaLore strategy converges faster and reaches the convergence state at the same time as LoRA; at the same time, the downward trend and final value of the training loss are important indicators for evaluating model performance. The loss value of the GaLore strategy at any global step is much lower than that of LoRA, which shows that even if the low-rank GaLore strategy is used, its effect is far better than that of the high-rank LoRA.

[0062] Then the accuracy of the system is analyzed, and bilingual evaluation is used. The indicators are used to measure the accuracy. The given standard is reference, the model is generated as candidate, the length is n, and there are m words in the candidate that appear in the reference. m / n is the calculation formula of BLEU 1-gram; according to the continuous word model n-gram, BLEU can be divided into multiple evaluation indicators, such as BLEU-1, BLEU-2, BLEU-3, BLEU-4, and the formula is:

[0063]

[0064] Among them, BP is a penalty factor that penalizes short output sentences to prevent the training results from being biased towards short sentences, c is the number of words in the candidate sentence, and r is the number of words in the reference sentence:

[0065]

[0066] P n is the precision of n-gram, and its formula is:

[0067]

[0068] ROUGE is an indicator for evaluating automatic summarization and machine translation. It compares the generated text with the reference and obtains the corresponding score to measure the "similarity" between the automatic generation and the reference. ROUGE-N mainly counts the recall rate on N-gram. For N-gram, the ROUGE-N score can be calculated. The calculation formula is:

[0069]

[0070] In ROUGE-L, L refers to the longest common subsequence, which is used to calculate the longest common subsequence between C and reference S. The calculation formula is:

[0071]

[0072] R in the formula LCS Recall rate, P LCS Indicates accuracy:

[0073]

[0074] Table 4 shows the BLEU and ROUGE data indicators of the GaLore strategy full parameter fine-tuning and LoRA fine-tuning on the evaluation set. The higher the indicator value, the closer the model is to expectations in text generation and the better the effect. The results show that with the increase of Rank value, the GaLore strategy has significant improvements in all evaluation indicators, especially BLEU-4 and ROUGE-2, which indicates that the ability to handle complex text structures and phrase generation is enhanced.

[0075] Table 4 Performance comparison

[0076]

[0077]

[0078] The experiment found that when the Rank of the GaLore strategy was 1024, the evaluation index was lower than that of Rank 512. Although the increase in the Rank value increases the number of parameters, theoretically it should be able to capture more complex patterns. However, when the Rank was 512, the index data performed best. This may be due to the following factors: (1) Overfitting: Increasing the number of model parameters can improve the expressiveness of the model, but it also increases the risk of overfitting. When the model is too complex, it may over-adapt to the noise and details in the training data instead of learning generalized patterns, resulting in decreased performance in the test set or actual applications. (2) Insufficient training: For more complex models, more training data and longer training time may be required for adequate training. (3) Difficult optimization: As the number of model parameters increases, the optimization problem becomes more complex and it may be more difficult to find the global optimal solution. This may lead to local minima during training, thus affecting the final performance of the model. At the same Rank value, LoRA performs far worse than the GaLore strategy, which shows that LoRA fine-tuning is not effective enough under the current setting and requires further optimization and adjustment.

[0079] Application examples such as training injury prevention and treatment scenarios generally have extremely high real-time, accuracy, professionalism and operability. According to the information flow of the strategy question-answering system on this side, through question input, the large model of fine-tuning training is used to perform reasoning ability prediction output, and feedback is given to the user in text form to complete the intelligent question-answering of training injuries.

[0080] Table 5 below shows the empirical case 1 for role confirmation "Who are you?". The comparison models are three Qwen2-0.5B base fine-tuning models based on traditional full-parameter fine-tuning, GaLore full-parameter fine-tuning, and LoRA fine-tuning. In the empirical results, the traditional full-parameter fine-tuning model gave answers that were irrelevant to "Who are you?"; since traditional full-parameter fine-tuning involves the adjustment of all model parameters, this may cause the model to be too complex, thereby increasing the risk of overfitting. At the same time, too many training rounds and insufficient training data may also cause the model to not have enough information to learn the generalization rules, but to over-adapt to specific samples in limited data; in contrast, the answers of the GaLore full-parameter fine-tuning model and the LoRA fine-tuning model are both objectively normal and effective. This may be because the complexity of the model is reduced by introducing a low-rank structure during training, thereby reducing the risk of overfitting and improving generalization ability.

[0081] Table 5 Empirical Case 1 (Role Confirmation)

[0082]

[0083] As shown in Table 6 below, in Empirical Cases 2 and 3, the specific questions about training injuries are: "What should I do if I have a muscle strain?" "What should I do if I have a fracture during training?" The answers given by the GaLore strategy with full parameter fine-tuning are highly professional and operational. Each question is combined with the actual training of the troops, the cause is summarized, and treatment is guided according to different situations, showing good training effects. The answers given by LoRA with fine-tuning also have the answer paradigm of the basic model, showing general universality. This shows that the model after fine-tuning based on the GaLore strategy with full parameters can meet the expected results of the experiment, greatly enhancing the professionalism and operability of the answers in the field of training injuries.

[0084] Table 6 Empirical Case 2 and Empirical Case 3 (Questions and Answers on Training Injuries)

[0085]

[0086]

[0087]

[0088] The above contents are further detailed descriptions of the present invention in combination with specific embodiments, and it cannot be determined that the specific implementation of the present invention is limited to these inventions. For ordinary technicians in the technical field of the present invention, without departing from the concept of the present invention, they can also make several simple deductions or substitutions, which should be regarded as belonging to the protection scope of the present invention.

Claims

1. A strategic question-answering method based on a vertical domain big model, characterized in that: include: S1: Receive natural language input from the user, retrieve open source datasets to obtain original datasets whose comprehensive similarity meets the requirements; S2: Perform necessary data cleaning and preprocessing on the original data set to obtain the training set; S3: Use the training dataset to fine-tune the pre-processed basic large language model, introduce a low-rank structure to modify its weight matrix to obtain the final model; S4: The problem enters the final model for post-processing to generate prediction output.

2. According to claim 1, a strategic question-answering method based on a vertical domain big model is characterized in that: The S2 also includes sorting out relevant documents and network resources disclosed on the Internet to construct and integrate high-quality data question-answer pairs related to the problem.

3. According to claim 1, a strategic question-answering method based on a vertical domain big model is characterized in that: In S3, the full parameter fine-tuning method is used for training, forming a series of empirical hyperparameters that affect the accuracy and convergence speed of the model. The gradient G of the weight matrix W (W∈R m×n ) Slowly changing low-rank structure, calculate two projection matrices P∈R m×r and Q∈R n×r , project the gradient matrix G into a low-rank form P T GQ, Where η is the learning rate, W T is the weight matrix, W0 is the initial weight matrix, G t is the gradient matrix, P t and Q t are two projection matrices, which are used to perform full-parameter learning by establishing multiple subspaces and switching between different subspaces during training. IN t =W0+ΔW T1 +ΔW T2 +ΔW Tn Among them, W t is the weight matrix, W0 is the initial weight matrix, ΔW Tn is the weight matrix updated in each round of iteration.

4. According to claim 1, a strategic question-answering method based on a vertical domain big model is characterized in that: The large language base model in the S3 adopts the Qwen2-0.5B-Instruct base model, and based on the Transformer architecture, the expressiveness of the model is enhanced by the SwiGLU activation function, and it is trained on a context length of 32k, and expanded to a longer context length through technologies such as DualChunk Attention.

5. According to claim 4, a strategic question-answering method based on a vertical domain big model is characterized in that: The Qwen2-0.5B-Instruct base model also uses QKV bias and group query attention to enhance the model’s understanding of complex language structures and its ability to process natural language.

6. A strategic question-answering system based on a vertical domain big model, characterized in that: It includes human-computer interaction module, data retrieval module, data processing module and model training module; The human-computer interaction module is used to receive natural language input from the user and provide corresponding output; The data retrieval module is used to retrieve and collect original data sets whose comprehensive similarity meets the requirements; The data processing module is used to clean and preprocess the original data set to obtain a training set, which is divided into training data, verification data and test data; The model training module is used to fine-tune the pre-trained large language base model.

7. A strategic question-answering system based on a vertical domain big model according to claim 6, characterized in that: The original data set includes public literature resources, network resources and field databases.

Citation Information

Cited By

  • Chinese semantic matching enhancement method and system based on BAAI-bge model

    CN120578756A

  • Training method and device of vertical domain question and answer model, equipment and storage medium

    CN121117615A

  • Large language model structured multi-mode response method and device oriented to vertical field

    CN121168669A

  • Large language model structured multi-modal response method and device for vertical field

    CN121168669B

  • Question and answer method based on large model and data agent

    CN121413751A