Enterprise ip chat system construction method based on large model fine tuning

Through layered fine-tuning and multimodal data fusion, the personalization and cost control problems of enterprise IP chat systems under the limitation of computing resources are solved, efficient and accurate corporate knowledge answers and user interaction are achieved, and brand image and business optimization are improved.

CN120258042APending Publication Date: 2025-07-04KUATULI (GUANGZHOU) TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510331372.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In the scenario where computing resources are limited, it is difficult for the existing technology to build a personalized enterprise IP chat system without causing high computing costs, and it is impossible to accurately understand complex business scenarios and enterprise professional knowledge.

Method used

The hierarchical fine-tuning strategy is adopted to divide the large language model into the basic layer, the adaptation layer and the output layer, and LoRA fine-tuning and full-parameter fine-tuning are performed respectively. Combined with multimodal data fusion and data balance, the closed-source model is used to generate training data and fine-tune the open source model.

Benefits of technology

While reducing computing costs, it improves the conversational capabilities of the chat system, can accurately answer corporate professional knowledge, enhance the quality of interaction with users, and improve brand image and business optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120258042A_ABST
    Figure CN120258042A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of natural language processing, and discloses an enterprise ip chat system construction method based on large model fine tuning, which comprises the following steps: obtaining enterprise knowledge data and chat dialogue data, respectively inputting the enterprise knowledge data and the chat dialogue data into a first large language model, and correspondingly generating an enterprise knowledge question and answer data set and a chat dialogue question and answer data set; inputting the enterprise knowledge question and answer data set and the chat dialogue question and answer data set into a pre-trained second large language model for fine tuning, performing hierarchical configuration on the second large language model in the fine tuning process, and performing independent optimization on parameters of each layer through multiple stages; according to the method, during fine tuning, hierarchical configuration is performed on the model, and parameters of each layer are independently optimized in stages, so that the constructed dialogue system can exceed a pure LoRA fine tuning effect in the dialogue capability, and meanwhile, the construction cost is lower than that of a full-amount fine tuning scheme.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and more specifically, to a method for constructing an enterprise IP chat system based on large model fine-tuning. Background Art

[0002] Many companies have their own mascots, which use large language models to achieve conversations with people. As a vivid symbol of the company's image, enterprise mascots can create many benefits for the enterprise and have become a rigid demand for most companies. On the one hand, it can greatly improve brand recognition and help consumers quickly remember the company among many competitors; on the other hand, the mascot can attract the attention of the audience in various marketing and promotion activities, shorten the distance with consumers, and strengthen brand affinity. Building a chat system for an enterprise mascot has significant benefits. It can not only enable the mascot to interact with customers at any time to further consolidate the brand image, but also present a unique service experience to customers through personalized chatting, collect customer feedback, and promote the optimization of the company's business.

[0003] Currently, the common construction methods of LLM (Large Language Model) role-playing chat systems are as follows: relying on role profiles and a large amount of knowledge texts, using prompt engineering, and using closed-source LLM represented by GPT-4 to carry out role-playing; or fine-tuning open-source LLM to build a role-playing model. However, these methods have some drawbacks. For example, when calling a closed-source LLM model, it often incurs high costs. Moreover, the closed-source LLM model cannot directly fine-tune professional data. Limited by the context window, it has limited flexibility and is difficult to implement a personalized enterprise IP chat system. When using open-source LLM to build a chat system, although deep customization can be achieved, its best full-parameter fine-tuning has extremely high requirements for computing resources, which is undoubtedly a heavy burden for small and medium-sized enterprises. In contrast, the LoRA fine-tuning with relatively small computing resource requirements achieves far less effect than full-parameter fine-tuning.

[0004] Therefore, in the scenario of limited computing resources, how to adopt an optimization strategy to enable the system to significantly exceed the performance of the LoRA fine-tuning system, while effectively avoiding the high computing cost of the full-parameter fine-tuning system, and finally enabling it to accurately understand complex business scenarios, accurately answer enterprise professional knowledge, efficiently utilize enterprise data, and achieve high-quality interaction with customers, so as to achieve the goals of enhancing the brand image and optimizing the business, is an urgent problem to be solved by those skilled in the art. Summary of the Invention

[0005] In view of this, the present invention provides a method for constructing an enterprise IP chat system based on large model fine-tuning, which can make the constructed dialogue system exceed the effect of simple LoRA fine-tuning in terms of dialogue ability, and at the same time be lower than the full-scale fine-tuning scheme in terms of construction cost.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] A method for constructing an enterprise IP chat system based on large model fine-tuning, comprising the following steps:

[0008] Obtain enterprise knowledge data and casual conversation data, and input them into the first large language model respectively, and correspondingly generate an enterprise knowledge Q&A dataset and a casual conversation Q&A dataset.

[0009] Input the enterprise knowledge Q&A dataset and the casual conversation Q&A dataset into a pre-trained second large language model for fine-tuning. During the fine-tuning process, perform hierarchical configuration on the second large language model, and separately optimize the parameters of each layer through multiple stages.

[0010] Preferably, the enterprise knowledge data is multi-modal data, and the acquisition steps of the multi-modal data include:

[0011] Obtain enterprise voice files and convert them into enterprise voice texts through the Whisper model; obtain enterprise image files and extract key text features through the CogVLM model; fuse the enterprise voice texts, the key text features and the enterprise text data to obtain the enterprise knowledge data.

[0012] Preferably, the steps further include:

[0013] Before fine-tuning, balance the data of the to-be-input enterprise knowledge Q&A dataset and the casual conversation Q&A dataset through data fusion.

[0014] Preferably, the steps of the data fusion include:

[0015] Set a fusion ratio, and sample the enterprise knowledge Q&A dataset and the casual conversation Q&A dataset according to the fusion ratio; splice the sampled data to obtain the final fused data sample.

[0016] Preferably, the first large language model is a closed-source large language model, and the second large language model is an open-source large language model.

[0017] Preferably, the first language model is the GPT-4 model, and the second large language model is ChatGLM-3.

[0018] Preferably, the steps of the fine-tuning include:

[0019] Load the ChatGLM-3 model and add hierarchical tokens. Divide the first l1 layers of tokens into the base layer, the next l2 layers of tokens after the base layer into the adaptation layer, and the last l3 layers into the output layer; set the learning parameters for each training stage; among them, the learning parameters include the first learning parameters corresponding to the first fine-tuning stage and the second learning parameters corresponding to the second fine-tuning stage.

[0020] Freeze the parameters of the base layer and the output layer, and read the first learning parameters to update the parameters of the middle layer under the LoRA fine-tuning framework.

[0021] Freeze the parameters of the trained adaptation layer and unfreeze the parameters of the output layer; read the second learning parameters to update the parameters of the output layer under the full-parameter fine-tuning framework.

[0022] Preferably, the learning parameters include the learning rate, the number of training epochs, the LoRA rank, and the loss function.

[0023] Preferably, the steps further include: constructing an enterprise IP profile, and jointly inputting the enterprise IP profile, the enterprise knowledge data, and the chat conversation data into the first large language model so that the question-and-answer pairs in the generated enterprise knowledge Q&A dataset and the chat conversation Q&A dataset conform to the enterprise IP profile.

[0024] An electronic device includes a processor and a memory. The memory stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the method for constructing an enterprise IP chat system as described above.

[0025] As can be seen from the above technical solutions, compared with the prior art, the present invention discloses a method for constructing an enterprise IP chat system based on large model fine-tuning, which reduces resource consumption through hierarchical fine-tuning and effectively improves the model effect. The present invention uses the role-playing ability of a closed-source model to generate a training dataset and fine-tunes based on an open-source model, which can meet personalized needs at low cost. The present invention uses data fusion to control the dataset ratio, enabling the model to balance learning enterprise knowledge and personalized enterprise styles. The present invention integrates enterprise multi-modal information and uses multi-modal data such as the text materials of the enterprise, the text converted from voice files, and the key text features extracted from image files to train the enterprise mascot chat system. The system can learn richer enterprise data, generate answer content with stronger relevance to the enterprise, fully exploit the enterprise multimedia information resources, improve the utilization degree of the chat system for enterprise information, and enhance the fit between the chat system and the actual business and image of the enterprise. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on the provided drawings.

[0027] Figure 1 Schematic diagram of a method for constructing an enterprise IP chat system based on large model fine-tuning provided by the present invention.

[0028] Figure 2 Schematic diagram of the principle of a method for constructing an enterprise IP chat system based on large model fine-tuning in an embodiment of the present invention. Detailed implementation manners

[0029] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0030] Embodiment 1

[0031] As Figure 1 and Figure 2 show, an embodiment of the present invention discloses a method for constructing an enterprise IP chat system based on large model fine-tuning, including the following steps:

[0032] S1: Obtain enterprise knowledge data and casual conversation data, and input them into the first large language model respectively, and correspondingly generate an enterprise knowledge Q&A dataset and a casual conversation Q&A dataset.

[0033] S2: Input the enterprise knowledge Q&A dataset and the casual conversation Q&A dataset into a pre-trained second large language model for fine-tuning. During the fine-tuning process, configure the second large language model in layers, and separately optimize the parameters of each layer through multiple stages.

[0034] In this embodiment, the present invention can use the first large language model to generate corresponding datasets as training data to train the second large language model, enabling the fine-tuned large language model to learn the capabilities of the first large language model to a certain extent.

[0035] To further implement the above technical solutions, the present invention can be used between closed-source models and open-source models, using the excellent conversation capabilities in the closed-source large model to construct training data and realizing the training of the open-source model.

[0036] Furthermore, the first large language model is the closed-source GPT-4 model, and the second large language model is the open-source ChatGLM-3 model. Hierarchical training is carried out using the LLMLayeredFineTuner: Obtain the question-and-answer dataset and load the pre-trained large language model ChatGLM-3; among them, the question-and-answer dataset includes the enterprise knowledge question-and-answer dataset and the chatting question-and-answer dataset generated by the closed-source GPT-4 model. The ChatGLM-3 model has 28 layers. In the hierarchical configuration, configure the number of layers of each layer in sequence, such as [5, 20, 3]. Divide the first 5 layers of the model into the basic layer, the middle 20 layers into the adaptation layer, and the last 3 layers into the output layer. Before training, it is necessary to configure the configuration parameters of each training stage. The present invention is divided into two training stages, and the two stages separately train the adaptation layer and the output layer in sequence. During the training process, on the basis of hierarchical training, the LoRA fine-tuning framework is used to update the parameters of the adaptation layer, and the full-parameter fine-tuning framework is used to update the parameters of the output layer. Among them, the training parameters that need to be configured include the learning rate, the number of training epochs, the LoRA rank, and the loss function (such as the cross-entropy loss function).

[0037] Specific training steps include:

[0038] S21. Model layering and initialization.

[0039] S211: Mark the first l1 layers of ChatGLM-3 as M1 (basic layer), the middle l2 layers as M2 (adaptation layer), and the last l3 layers as M3 (output layer). Concatenate M1 + M2 + M3 into an overall model, denoted as M o (Combined model).

[0040] S212: According to the freezing strategy in the first-stage training parameters θ1, set the parameters of the basic layer M1 and the output layer M3 to be untrainable (frozen), and retain the trainable state of the adaptation layer M2.

[0041] S22. LoRA training for the adaptation layer.

[0042] S221: Freeze M1 and M3, and only allow M2 to update parameters under the LoRA framework.

[0043] S222: Read θ1, which includes the learning rate (lr1), the number of training epochs (epochs1), the LoRA rank (lora_rank), and the loss function (such as cross-entropy loss), and initialize the LoRA-related modules.

[0044] S223: During the training process, take a batch of data (q, a) from the training dataset QA IP in.

[0045] Forward propagation: Input (q, a) into M o , complete the forward calculation through the LoRA module of M2, output the prediction result and calculate the loss value based on the loss function Calculate the loss value.

[0046] Backward propagation: Based on the loss obtained from forward propagation, calculate the gradient of the weights related to LoRA in M2, and then update the trainable weights of M2 through backward propagation.

[0047] During the parameter update process in the LoRA framework, for the target weight matrix in the adaptation layer M2 (such as the parameters of the attention layer), inject two low-rank matrices (A and B). Among them, A is initialized by Gaussian distribution, B is initialized as a zero matrix, and the rank is much smaller than the dimension of the original matrix. During forward propagation, the input data passes through the original frozen weights and the low-rank bypass branch simultaneously (the output is superimposed as Wx + BAx). During backward propagation, only the gradients of A and B are updated, and the original weights remain fixed. After training is completed, the bypass matrix can be merged with the original weights to avoid inference delay. This method adjusts the model behavior through low-rank approximation, greatly reducing the number of trainable parameters, and freezing M1 and M3 to retain the pre-trained knowledge.

[0048] S224: Repeat step S223 until the number of iterations configured in θ1 is reached or the stop condition is satisfied.

[0049] S23. Full-parameter fine-tuning of the output layer.

[0050] S231: Freeze the base layer M1 and the adaptation layer M2, so that the model is in a state where only the output layer M3 is unfrozen. In this state, all parameters of the output layer M3 can be updated.

[0051] S232: Read the training parameters θ2 of the second stage, which include the learning rate (lr2), the number of training epochs (epochs2), and the loss function (such as mean squared error loss) to prepare the optimizer for full-parameter fine-tuning.

[0052] S233: During the training process, take out a batch of data (q, a) from the training dataset QA IP again.

[0053] Forward propagation: Input (q, a) into M o , obtain the output result through the forward calculation of the model, and calculate the loss value based on the loss function Calculate the loss value.

[0054] Backward propagation: Based on the loss obtained from forward propagation, calculate the gradient of M3, and then update the weights of M3 through backward propagation.

[0055] S234: Repeat step S233 until the number of iterations configured in θ2 is reached or the stop condition is satisfied.

[0056] S24. Output model M o 。

[0057] In one embodiment, the present invention obtains multi-modal enterprise knowledge data, making the fine-tuned model more in line with the actual business of the enterprise. For the acquisition of multi-modal data, it can be collected through the following several ways: (1) With the help of the enterprise ERP system interface, extract the publicly available enterprise data J parsed in JSON format E , such as organizational structure, product information, etc.; (2) Collect the text materials D of the enterprise E , covering enterprise profiles, annual reports, official account tweets, etc.; (3) Collect publicly available enterprise voice files V E , such as recordings of annual meeting speeches, recordings of technical sharing sessions, etc.; (4) Collect enterprise-related image files I E , such as enterprise promotional posters, illustrations for official account articles, etc.

[0058] For the data collected above, it can be processed through the following methods: (1) Use the Whisper-3 model to convert the enterprise voice file V E into enterprise voice text D VE ; (2) Use the CogVLM model (Cognitive Vision Language Model) to extract the key text features D E of the enterprise-related image file I IE 。

[0059] To further implement the above technical solution, when generating training data through the GPT-4 model, it is necessary to make the GPT-4 incorporate a certain answering style when inputting questions or conversations, so as to meet personalized needs in turn.

[0060] For the incorporation of style, the pre-configured character profile with a specific style (i.e., enterprise ip profile) can be used as the input of the GPT-4 model while inputting questions or conversations, and the training data generated by the GPT-4 model can meet the corresponding character style through a specific prompt word template.

[0061] The specific steps include:

[0062] Configure the basic information of the character role (such as role name, nickname, gender, age, etc.), appearance image (such as styling style, physical characteristics, clothing collocation, etc.), personality characteristics (core personality, personality level, personality formation cause), background story (such as character origin, growth experience, connection with the enterprise), skill expertise (such as professional skills, special abilities), catchphrases (such as habitual expressions, emotional expressions) and application scenarios according to the enterprise requirements.

[0063] Generate corresponding prompts according to the configured task role information and build a question library, including self-cognition generation instructions and style information generation instructions. Through self-cognition instructions and style information generation instructions, the generated training data can train ChatGLM-3's self-cognition ability and style imitation ability.

[0064] An example of a self-awareness instruction question bank is as follows:

[0065] ①Basic information:

[0066] May I have your name?

[0067] How do people usually call you (nickname)?

[0068] What is your gender?

[0069] What do you think your gender means to you?

[0070] How old are you this year?

[0071] What are your feelings and thoughts at this age?

[0072] ②Appearance:

[0073] How would you describe your overall style? Cartoon-like, realistic, or something else?

[0074] Tell me about your height and body shape. Are you satisfied with your physical features? Why?

[0075] Tell me about your facial features. Which one are you most satisfied with and which one do you think is the most distinctive?

[0076] Talk about your clothing combinations, colors, styles, and accessories. What do they reflect about your preferences and personality?

[0077] ③Personality characteristics:

[0078] Use a few words to summarize your core personality traits. How do you think these traits were formed?

[0079] Do your emotions and reactions change in different situations? Give an example.

[0080] Looking back on your growing up experience, what things had the greatest impact on the shaping of your character?

[0081] ④Background story:

[0082] Do you still remember how you were born? What was the reason for your birth?

[0083] What important growth points have you experienced along the way? What do these experiences mean to you?

[0084] What role do you play in the enterprise?

[0085] What kind of close connection do you have with the enterprise's business and culture?

[0086] ⑤ Skills and Expertise:

[0087] What is your most proficient professional skill? How did you acquire this skill?

[0088] Besides professional skills, do you have any special abilities? How did you discover and master these special abilities?

[0089] ⑥ Catchphrase:

[0090] What is the sentence you often say? Why is it this sentence?

[0091] What special meaning does it have for you and the enterprise?

[0092] ⑦ Application Scenarios:

[0093] In the online virtual world, where do you like to appear and what do you like to do?

[0094] Coming to the offline real scenario, in which kind of occasion do you think you can play the most role? Why?

[0095] Generate a self - awareness Q&A dataset through the above instructions.

[0096] Furthermore, construct corresponding question banks based on enterprise knowledge and chat conversations respectively, and generate style - less answers through the ChatGLM - 3 model to construct a style - less answer dataset.

[0097] Input the style - less answer dataset into GPT - 4 and polish it according to the style of fictional enterprise characters to obtain an enterprise - style dataset.

[0098] Finally, adopt the above method to form an enterprise knowledge Q&A dataset and a chat conversation Q&A dataset with enterprise style.

[0099] In one embodiment, use the data fuser Mix to allocate and fuse the enterprise knowledge QA pairs QA E-IP and the chat conversation QA pairs QA chat-IP according to a specific fusion ratio ratio (default is 1) to generate a fused training dataset QA IP , which evenly contains QA E-IP and QA chat-IPQA pair data is used for subsequent fine-tuning of ChatGLM-3 to evenly learn both contents and prevent its answers from being overly biased towards one side of the data.

[0100] In this embodiment, the data integrator can confirm the sampling scale of various data according to the configured parameters to achieve the balance of the two types of data. The fusion ratio is set to 1. When sampling, the data integrator will compare the sizes of the data sets, take out the same amount of data from the larger data set among the two types of data and merge it with the smaller data set to obtain balanced training data.

[0101] Next, the technical effects of the present invention will be specifically described:

[0102] 1. Instead of directly using the closed-source model to build the final dialogue system, the present invention uses the closed-source model such as GPT-4 to generate training data, and then uses these data to fine-tune the open-source LLM. This not only utilizes the powerful role-playing ability of the closed-source model but also avoids the high cost of directly calling the closed-source model API. Enterprises do not need to bear a heavy cost burden, having an obvious advantage in cost control. At the same time, it can also be deeply customized based on the open-source model, balancing the cost and customization requirements.

[0103] 2. The present invention solves the balance problem between computing resources and fine-tuning effects. By designing an LLM hierarchical fine-tuner and adopting a hierarchical fine-tuning strategy, the open-source LLM is divided into a basic layer, an adaptation layer, and an output layer. This hierarchical design is based on the differences in functions and data processing at different levels of the language model.

[0104] As the bottom layer of the model, the basic layer undertakes the key task of initially extracting and processing input data, and its parameters determine the way the model understands and processes basic language information. Freezing the parameters of the basic layer is like building a solid foundation for the model, ensuring the stability of the model's bottom-layer architecture during the training process and avoiding deviations in the model's understanding of basic language features caused by frequent changes in the bottom-layer parameters. This guarantees that when processing various inputs, the model can perform preliminary analysis with a stable bottom-layer logic.

[0105] The adaptation layer is in the middle layer of the model. Its role is to further adjust and adapt the data on the basis of the processing of the basic layer to better fit the characteristics of the training data. The LoRA fine-tuning method is adopted. LoRA introduces a small number of trainable parameters through low-rank matrix factorization instead of updating all the parameters in the middle layer, greatly reducing the computational amount. Under limited computing resources, this method enables the middle layer of the model to be specifically optimized according to the characteristics of the training data, effectively improving the model's adaptability to different types of data patterns without consuming too many resources. For example, when processing two different types of data, namely enterprise knowledge Q&A and chat conversations, LoRA fine-tuning can enable the adaptation layer to quickly adapt to the data characteristics and optimize the understanding of different language patterns and semantics.

[0106] The output layer directly determines the final output result of the model. Adopting full-parameter fine-tuning means finely adjusting all the parameters of the output layer. Since the output layer needs to generate answers that highly match the input questions, in the face of complex and diverse user questions, full-parameter fine-tuning can fully consider the semantics, context, and various subtle differences of the questions, making the model output more accurate and flexible. For example, when the user asks a complex question about the usage method of an enterprise's new product in a specific scenario, the output layer after full-parameter fine-tuning can comprehensively analyze various information in the question and give the answer that best meets the user's needs.

[0107] Stratification method selection: When determining the stratification method, the present invention comprehensively considers various factors such as model structure, task requirements, and computing resources. Taking the ChatGLM-3 model as an example, it contains 28 GLM blocks inside. After a large number of experiments and analyses, it is determined that the first 5 layers are set as the basic layer, the middle 20 layers are set as the adaptation layer, and the last 3 layers are set as the output layer. For the basic layer, the selection of the number of frozen parameter layers needs to find a balance between ensuring the stability of the model architecture and retaining a certain degree of flexibility. If the number of layers is too small, the underlying architecture may not be fully stabilized; if the number of layers is too large, it may limit the model's ability to learn new knowledge. The adaptation layer has a relatively large number of layers because it needs to undertake the task of adapting to various data patterns, and more layers can provide a richer parameter adjustment space to adapt to complex training data. Although the output layer has a relatively small number of layers, due to its key role in output accuracy, full-parameter fine-tuning can achieve precise control of the output results within a limited number of layers. In addition, in different enterprise application scenarios, the number of layers of each layer can be flexibly adjusted according to the characteristics of enterprise data, the focus of dialogue tasks, and the available computing resources. For example, if the amount of casual conversation data of the enterprise is large and complex, the number of sub-layers used to process casual conversation data in the adaptation layer can be appropriately increased to better optimize the effect of casual conversation; if the enterprise pays more attention to the accuracy of professional knowledge Q&A or has more available computing resources, the parameter adjustment intensity or the number of layers of the output layer can be appropriately increased to improve the accuracy of professional question answering. This method of flexibly adjusting the stratification method according to the actual situation is the key improvement of the present invention compared with simply freezing some parameter training models in the prior art, which can more efficiently utilize computing resources, improve the fine-tuning effect, and meet the diverse needs of enterprises.

[0108] 3. The present invention integrates enterprise multimodal information and uses multimodal data such as the text materials of the enterprise, the text converted from voice files, and the key text features extracted from image files to train the enterprise mascot chat system. The system can learn richer enterprise data, generate answer content with stronger relevance to the enterprise, fully exploit the enterprise multimedia information resources, improve the utilization degree of the chat system for enterprise information, and enhance the fit between the chat system and the actual business and image of the enterprise.

[0109] 4. Achieve balanced data utilization: The data fusion device can reasonably fuse enterprise knowledge QA pairs and casual conversation QA pairs according to a specific ratio. In the training data link, it avoids the problem that the chat system is overly biased towards one side of the data due to the difference in data volume, and avoids the situation that the answer is rigid and stereotyped due to over-reliance on enterprise knowledge data and cannot meet the user's emotional communication and interaction needs, or the situation that when facing professional consultations, key information is often not given and clichés are generated due to over-biasing towards casual conversation materials. It provides users with a more comprehensive and high-quality service experience, and improves the practicality and user satisfaction of the chat system.

[0110] Embodiment 2

[0111] Based on the same inventive concept, an embodiment of the present invention discloses an electronic device, which includes a processor and a memory. The memory stores a computer program, and when the computer program is executed, it implements the chat system construction method in Embodiment 1.

[0112] Among them, the processor can be composed of integrated circuits in some embodiments. For example, it can be composed of a single packaged integrated circuit, or it can be composed of multiple packaged integrated circuits with the same or different functions, including a combination of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor is the control core of the electronic device, connecting various components of the entire electronic device through various interfaces and circuits, and executing various functions of the electronic device and processing data by running or executing programs or modules stored in the memory and calling data stored in the memory.

[0113] Among them, the memory can be, for example, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the above. More specific examples of the storage medium (a non-exhaustive list) include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, mechanically encoded device, and any suitable combination of the above.

[0114] In this specification, the various embodiments are described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the description in the method part.

[0115] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for constructing an enterprise IP chat system based on large model fine-tuning, characterized in that, It includes the following steps: Obtain enterprise knowledge data and chat conversation data, and input them into the first large language model respectively to generate an enterprise knowledge Q&A dataset and a chat conversation Q&A dataset accordingly; Input the enterprise knowledge Q&A dataset and the chat conversation Q&A dataset into a pre-trained second large language model for fine-tuning. During the fine-tuning process, configure the second large language model in layers and optimize the parameters of each layer separately through multiple stages.

2. The method for constructing an enterprise IP chat system based on large model fine-tuning according to claim 1, wherein, The enterprise knowledge data is multi-modal data, and the acquisition steps of the multi-modal data include: Obtain enterprise text data; Obtain enterprise voice files and convert them into enterprise voice text through the Whisper model; Obtain enterprise image files and extract key text features through the CogVLM model; Fuse the enterprise voice text, the key text features and the enterprise text data to obtain the enterprise knowledge data.

3. A method for constructing an enterprise IP chat system based on large model fine-tuning according to claim 1, characterized in that, The steps also include: Before fine-tuning, balance the data of the to-be-input enterprise knowledge Q&A dataset and the chat conversation Q&A dataset through data fusion.

4. A method for constructing an enterprise IP chat system based on large model fine-tuning according to claim 3, characterized in that, The steps of the data fusion include: Set the fusion ratio and sample the enterprise knowledge Q&A dataset and the chat conversation Q&A dataset according to the fusion ratio; Merge the sampled data to obtain the final fused data sample.

5. A method for constructing an enterprise IP chat system based on large model fine-tuning according to claim 1, characterized in that, The first large language model is a closed-source large language model, and the second large language model is an open-source large language model.

6. A method for constructing an enterprise IP chat system based on large model fine-tuning according to claim 1 or 5, characterized in that, The first language model is the GPT-4 model, and the second large language model is ChatGLM-3.

7. A method for constructing an enterprise IP chat system based on large model fine-tuning according to claim 6, characterized in that, The steps of the fine-tuning include: Load the ChatGLM-3 model and add hierarchical tags. Divide the first l1 layers of tags into the basic layer, the l2 layers after the basic layer are marked as the adaptation layer, and the last l3 layers are divided into the output layer; set the learning parameters for each training stage; Freeze the parameters of the basic layer and the output layer, and read the first learning parameter to update the parameters of the middle layer under the LoRA fine-tuning framework; Freeze the parameters of the trained adaptation layer and unfreeze the parameters of the output layer; read the second learning parameter to update the parameters of the output layer under the full-parameter fine-tuning framework.

8. A method for constructing an enterprise IP chat system based on large model fine-tuning according to claim 7, characterized in that, The learning parameters include the learning rate, the number of training epochs, the LoRA rank and the loss function.

9. A method for constructing an enterprise IP chat system based on large model fine-tuning according to claim 1, characterized in that, The steps also include: construct an enterprise IP profile and input it together with the enterprise knowledge data and the chat conversation data into the first large language model so that the Q&A pairs in the generated enterprise knowledge Q&A dataset and the chat conversation Q&A dataset conform to the enterprise IP profile.

10. An electronic device, characterized in that, It includes a processor and a memory. The memory stores machine-executable instructions that can be executed by the processor. The processor executes the machine-executable instructions to implement the chat system construction method according to any one of claims 1-9.

Citation Information

Patent Citations

  • Large model assembly line parallel training method and system based on gradient sensing parameter freezing

    CN118568499A

  • Intelligent question answering system and question answering method for professional knowledge of tobacco industry based on big language

    CN119088927A

  • Parameter fine tuning method and device for downstream task adaptation of large language model

    CN119294362A

  • Knowledge-based dialogue system for self-learning dialogues and learning method thereof

    US20230134933A1

Cited By

  • Intelligent question and answer method, electronic equipment, medium and product

    CN120448509A