Large model training method and device for cement industry and storage medium
By sub-sample division and standard problem vocabulary setting of cement industry data, combined with attention mechanism and loss function optimization, the problem of low training accuracy caused by the non-single cement plant data is solved, and the efficient application of the model in cement production is achieved.
Patent Information
- Application Number
- CN202510304966.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-07-18
AI Technical Summary
The cement plant data sources are not single and the data volume is small, resulting in low training accuracy of large models.
By dividing the pending data into multiple subsamples, setting up submodels for each business scenario, and training using the vocabulary and calculation parameter table of standard problems, combining attention mechanism and loss function optimization, joint training of submodels is achieved.
It significantly improves the industry adaptability and prediction accuracy of the model, improves the efficiency and quality of the cement production process, and reduces resource waste.
Smart Images

Figure CN120336844A_ABST
Abstract
Description
Technical Field
[0001] This document relates to the field of big data technology, and in particular, to a large model training method, device, and storage medium for the cement industry. Background Art
[0002] In deep learning, a large model refers to a neural network model with a large number of parameters. In the fields of natural language processing and computer vision, these large models have made remarkable progress in automatic language generation, image generation, and decision-making.
[0003] During the training process of a large model, a large amount of data with a single source is required to ensure the computational accuracy of the model.
[0004] However, the data of many cement plants cannot be interconnected, resulting in a non-single data source. The non-single data source also leads to a small amount of data from each data source, thus affecting the training accuracy of the model. Summary of the Invention
[0005] In view of the above solutions, the present application aims to propose a large model training method, device, and storage medium for the cement industry to solve at least one of the above problems.
[0006] In a first aspect, one or more embodiments of the present application provide a large model training method for the cement industry, including:[[]]END]]
[0007] Obtain the data to be processed for the cement industry;
[0008] Divide the data to be processed into multiple sub-samples according to a preset business scenario;
[0009] Train a first sub-model based on the multiple sub-samples, with the sub-samples corresponding to the first sub-model one by one;
[0010] Set a vocabulary table and a calculation parameter table for standard questions in the cement industry for each of the first sub-models;
[0011] Using the data to be processed as input, train a second sub-model based on the vocabulary table of standard questions and the calculation parameter table; and
[0012] Using the data to be processed as input, train the first sub-model and the second sub-model based on the vocabulary table of standard questions and the calculation parameter table.
[0013] Further, after dividing into multiple sub-samples, the method further includes:
[0014] Calculate the number of samples of each of the sub-samples;
[0015] Determine whether the number of samples of each of the sub-samples reaches a preset value; and
[0016] Perform data augmentation on the sub-samples whose sample quantity does not reach the preset value.
[0017] Furthermore, training the first sub-model based on the multiple sub-samples includes:
[0018] Convert the data in the sub-samples into a preset data format, and the data format includes: calculation parameters, domain labels, question labels, and target answers;
[0019] Input the calculation parameters into the first sub-model to obtain a predicted answer;
[0020] Calculate the first cosine similarity between the predicted answer and the question label;
[0021] Calculate the second cosine similarity between the predicted answer and the target answer;
[0022] Calculate the score of the predicted answer according to the first cosine similarity and the second cosine similarity;
[0023] If the score of the predicted answer is not greater than the preset value, perform the backpropagation process of the corresponding parameters to update the parameters of the second pre-trained model; perform the above operations until the score of the predicted answer satisfies being greater than the preset value; and
[0024] If the score of the predicted answer is greater than the preset value, determine that the predicted answer is correct and end the current process.
[0025] Furthermore, the second sub-model uses an attention mechanism.
[0026] Furthermore, taking the data to be processed as the input, training the second sub-model based on the vocabulary of the standard questions and the calculation parameters includes:
[0027] Convert the vocabulary of the standard questions and the calculation parameter table into key vectors; and
[0028] Based on the attention mechanism, the second sub-model outputs the standard questions and calculation parameters corresponding to the first sub-model.
[0029] Furthermore, taking the data to be processed as the input, training the first sub-model and the second sub-model based on the vocabulary of the standard questions and the calculation parameter table includes:
[0030] Taking the data to be processed as the input, based on the vocabulary of the standard questions and the calculation parameters, the second sub-model outputs the standard questions and calculation parameters corresponding to the first sub-model;
[0031] According to the standard problem, the first sub-model obtains corresponding calculation parameters, and the calculation parameters include: domain labels;
[0032] The first sub-model takes the calculation parameters as input and outputs a predicted answer;
[0033] Calculate the loss function value of the predicted answer and the loss function value of the domain label respectively;
[0034] According to the loss function values of the predicted answers and the loss function values of the domain labels of each first sub-model, determine the total loss function value;
[0035] Determine whether the total loss function value meets the preset conditions;
[0036] When the total loss function value does not meet the preset conditions, perform the backpropagation process of the parameters and repeat the above process until the total loss function value meets the preset conditions; and
[0037] When the total loss function value meets the preset conditions, end the current process.
[0038] In a second aspect, one or more embodiments of the present application provide a large model training device for the cement industry, including:
[0039] An acquisition module, configured to acquire data to be processed;
[0040] A data processing module, configured to divide the data to be processed into multiple sub-samples according to a preset business scenario;
[0041] A first training module, which trains a first sub-model based on the multiple sub-samples, and the sub-samples and the first sub-model are in one-to-one correspondence;
[0042] A second training module, configured to set a vocabulary table and a calculation parameter table for standard problems for each first sub-model; take the data to be processed as input, and train a second sub-model based on the vocabulary table of the standard problems and the calculation parameter table; and
[0043] A third training module, configured to take the data to be processed as input, and train the first sub-model and the second sub-model based on the vocabulary table of the standard problems and the calculation parameter table.
[0044] Further, the first training module is configured to convert the data in the subsample into a preset data format, where the data format includes: calculation parameters, domain labels, question labels, and target answers; input the calculation parameters into the first sub-model to obtain a predicted answer; calculate a first cosine similarity between the predicted answer and the question label; calculate a second cosine similarity between the predicted answer and the target answer; calculate a score of the predicted answer according to the first cosine similarity and the second cosine similarity; if the score of the predicted answer is not greater than a preset value, perform a backpropagation process of corresponding parameters to update the parameters of the second pre-trained model; perform the above operations until the score of the predicted answer satisfies being greater than the preset value; and if the score of the predicted answer is greater than the preset value, determine that the predicted answer is correct and end the current process.
[0045] Further, the third training module is configured to use the data to be processed as input, and based on the vocabulary of the standard questions and the calculation parameters, the second sub-model outputs the standard questions and calculation parameters corresponding to the first sub-model; according to the standard questions, the first sub-model obtains corresponding calculation parameters, where the calculation parameters include: domain labels; the first sub-model uses the calculation parameters as input and outputs a predicted answer; calculate a loss function value of the predicted answer and a loss function value of the domain label respectively; determine a total loss function value according to the loss function values of the predicted answers and domain labels of each first sub-model; determine whether the total loss function value satisfies a preset condition; when the total loss function value does not satisfy the preset condition, perform a backpropagation process of parameters and repeat the above process until the total loss function value satisfies the preset condition; and when the total loss function value satisfies the preset condition, end the current process.
[0046] In a third aspect, an embodiment of the present application provides a storage medium for storing computer-executable instructions, characterized in that the computer-executable instructions, when executed, implement the steps of the large model training method for the cement industry according to any one of the first aspects.
[0047] Compared with the prior art, the present application can at least achieve the following technical effects:
[0048] Through the data processing and model training strategies for the cement industry, the industry adaptability and prediction accuracy of the model are significantly improved. Using the preset business scenarios for data partitioning and sub-model training can capture industry characteristics more accurately, and the setting of the vocabulary of standard questions and the calculation parameter table further enhances the model's understanding and processing ability of common problems in the cement industry. In practical applications, this method can effectively improve the efficiency and quality in the cement production process, reduce resource waste, and achieve more accurate production control and decision-making support. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] To more clearly illustrate the technical solutions in one or more embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments described in the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0050] Figure 1 A flowchart of a large model training method for the cement industry provided for one or more embodiments of the present application;
[0051] Figure 2 A schematic structural diagram of a large model training device for the cement industry provided for one or more embodiments of the present application. Detailed implementation manners
[0052] To enable those skilled in the art to better understand the technical solutions in one or more embodiments of the present application, the following will clearly and completely describe the technical solutions in one or more embodiments of the present application in conjunction with the drawings in one or more embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on one or more embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this document.
[0053] The embodiments of the present application provide a large model training method for the cement industry, including the following steps:
[0054] Step 1: Obtain the data to be processed for the cement industry.
[0055] In the embodiments of the present application, the data to be processed includes: cement formula, cement output, cement quality, raw material information, equipment failures, etc. Specifically, obtain the cement formula, cement output, cement quality, raw material information, and equipment failures of different cement plants at different time periods every day from the cement enterprise operation management system.
[0056] Step 2: Divide the data to be processed into multiple sub-samples according to the preset business scenarios.
[0057] In the embodiments of the present application, the business scenarios include: technological processes such as raw meal preparation, clinker calcination, and cement grinding.
[0058] Step 3: Train the first sub-model based on multiple sub-samples, with each sub-sample corresponding to the first sub-model.
[0059] In the embodiments of the present application, in order to improve the adaptability of the model to industry characteristics, a sub-model is set for each business scenario. That is, it is ensured that the overall model can answer questions in various business scenarios.
[0060] Step 4: Set a vocabulary table and a calculation parameter table for standard questions in the cement industry for each first sub-model.
[0061] In the embodiments of the present application, the vocabulary table of standard questions is based on the questions proposed by users, and the calculation parameter table corresponds to the parameters that need to be considered to solve the corresponding questions. Therefore, constructing the vocabulary table of standard questions and the calculation parameter table can improve the selection accuracy of the model, so as to adapt to the situation of a relatively small number of training samples in the cement industry.
[0062] Step 5: Use the data to be processed as input, and train the second sub-model based on the vocabulary table of standard questions and the calculation parameter table.
[0063] In the embodiments of the present application, the second sub-model is used to determine the questions that the user wants to ask and the parameters used to answer these questions according to the data input by the user.
[0064] Step 6: Use the data to be processed as input, and train the first sub-model and the second sub-model based on the vocabulary table of standard questions and the calculation parameter table.
[0065] In the embodiments of the present application, a question-answering system is formed based on the first sub-model and the second sub-model, that is, the second model is responsible for asking questions, and the first sub-model is responsible for answering. By dividing business scenarios, the accuracy of the answers of each first sub-model is ensured. By setting the vocabulary table of standard questions and the calculation parameter table, the problem of a relatively small number of training samples is solved.
[0066] In the embodiments of the present application, after dividing into multiple sub-samples, calculate the sample quantity of each of the sub-samples; determine whether the sample quantity of each of the sub-samples reaches a preset value; perform data augmentation on the sub-samples whose sample quantity does not reach the preset value. For the cement industry, its data is less and the professionalism is stronger. Therefore, after dividing technical scenarios, data augmentation needs to be performed on the data in each scenario to ensure the professionalism of the data.
[0067] In the embodiments of the present application, the specific process of training the first sub-model is as follows:
[0068] Convert the data in the sub-sample into a preset data format, and the data format includes: calculation parameters, domain labels, question labels, and target answers;
[0069] Input the calculation parameters into the first sub-model to obtain a predicted answer;
[0070] Calculate the first cosine similarity between the predicted answer and the question label;
[0071] Calculate the second cosine similarity between the predicted answer and the target answer;
[0072] Calculate the score of the predicted answer based on the first cosine similarity and the second cosine similarity;
[0073] If the score of the predicted answer is not greater than the preset value, perform the backpropagation process of the corresponding parameters to update the parameters of the second pre-trained model; perform the above operations until the score of the predicted answer satisfies being greater than the preset value.
[0074] If the score of the predicted answer is greater than the preset value, determine that the predicted answer is correct and end the current process.
[0075] In the case of a small amount of data, use the matching degree between the question and the predicted answer, and the matching degree between the predicted answer and the standard answer to double-verify the predicted answer to achieve multiple rounds of iteration, thereby improving the accuracy of the calculation result.
[0076] In the embodiments of the present application, the attention mechanism can dynamically adjust the attention area for the input data according to different task requirements. It is just suitable for the application scenario of the second sub-model, so the second sub-model uses the attention mechanism.
[0077] In the embodiments of the present application, the training process of the second sub-model is as follows:
[0078] Convert the vocabulary of the standard question and the calculation parameter table into key vectors. Based on the attention mechanism, the second sub-model outputs the standard question and calculation parameters corresponding to the first sub-model.
[0079] In the embodiments of the present application, after the first sub-model and the second sub-model are trained, joint training is also required. The joint training process is as follows:
[0080] Taking the data to be processed as the input, based on the vocabulary of the standard question and the calculation parameters, the second sub-model outputs the standard question and calculation parameters corresponding to the first sub-model;
[0081] According to the standard question, the first sub-model obtains the corresponding calculation parameters, and the calculation parameters include: domain labels;
[0082] The first sub-model takes the calculation parameters as the input and outputs a predicted answer;
[0083] Calculate the loss function value of the predicted answer and the loss function value of the domain label respectively;
[0084] Determine the total loss function value according to the loss function values of the predicted answers and domain labels of each first sub-model;
[0085] Determine whether the total loss function value meets a preset condition;
[0086] When the total loss function value does not meet the preset condition, perform the backpropagation process of the parameters and repeat the above process until the total loss function value meets the preset condition; and
[0087] When the total loss function value meets the preset condition, end the current process.
[0088] An embodiment of the present application provides a large model training device for the cement industry, as Figure 2 shown, including:
[0089] An acquisition module 201 for acquiring data to be processed;
[0090] A data processing module 202 for dividing the data to be processed into multiple sub-samples according to a preset business scenario;
[0091] A first training module 203 for training a first sub-model based on the multiple sub-samples, with the sub-samples corresponding to the first sub-model one by one;
[0092] A second training module 204 for setting a vocabulary and a calculation parameter table of standard questions for each of the first sub-models; using the data to be processed as input, and training a second sub-model based on the vocabulary of the standard questions and the calculation parameter table; and
[0093] A third training module 205 for using the data to be processed as input, and training the first sub-model and the second sub-model based on the vocabulary of the standard questions and the calculation parameter table.
[0094] In the embodiment of the present application, the first training module is used to convert the data in the sub-sample into a preset data format, and the data format includes: calculation parameters, domain labels, question labels, and target answers; input the calculation parameters into the first sub-model to obtain a predicted answer; calculate a first cosine similarity between the predicted answer and the question label; calculate a second cosine similarity between the predicted answer and the target answer; calculate a score of the predicted answer according to the first cosine similarity and the second cosine similarity; if the score of the predicted answer is not greater than a preset value, perform the backpropagation process of the corresponding parameters to update the parameters of the second pre-trained model; perform the above operations until the score of the predicted answer meets the condition of being greater than the preset value; and if the score of the predicted answer is greater than the preset value, determine that the predicted answer is correct and end the current process.
[0095] In an embodiment of the present application, the third training module is configured to use the data to be processed as input, and based on the vocabulary of the standard questions and the calculation parameters, the second sub-model outputs the standard questions and calculation parameters corresponding to the first sub-model; according to the standard questions, the first sub-model obtains corresponding calculation parameters, and the calculation parameters include: domain labels; the first sub-model uses the calculation parameters as input and outputs a predicted answer; calculates the loss function value of the predicted answer and the loss function value of the domain label respectively; determines the total loss function value according to the loss function values of the predicted answers and the loss function values of the domain labels of each first sub-model; determines whether the total loss function value meets a preset condition; when the total loss function value does not meet the preset condition, performs the backpropagation process of the parameters and repeats the above process until the total loss function value meets the preset condition; and when the total loss function value meets the preset condition, ends the current process.
[0096] An embodiment of the present application provides a storage medium for storing computer-executable instructions, characterized in that the computer-executable instructions, when executed, implement the steps of the large model training method for the cement industry described in any one of the embodiments.
[0097] It should be noted that the embodiment of the storage medium in the present application and the embodiment of the blockchain-based service providing method in the present application are based on the same inventive concept. Therefore, the specific implementation of this embodiment can refer to the corresponding implementation of the blockchain-based service providing method described above, and the repeated parts will not be elaborated.
[0098] The above describes specific embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain implementations, multitasking and parallel processing are also possible or may be advantageous.
[0099] In the 1930s, it was obvious to distinguish whether an improvement to a technology was a hardware improvement (e.g., improvement to circuit structures such as diodes, transistors, switches, etc.) or a software improvement (improvement to method flows). However, with the development of technology, many improvements to method flows today can be regarded as direct improvements to hardware circuit structures. Almost all designers obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement to a method flow cannot be implemented with a hardware entity module. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is an integrated circuit whose logical function is determined by the user programming the device. Designers can program by themselves to "integrate" a digital system on a piece of PLD without having to ask a chip manufacturer to design and produce a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compiler used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called Hardware Description Language (HDL), and there is not only one kind of HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones currently are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be aware that by simply performing a little logical programming on the method flow with the above-mentioned several hardware description languages and programming it into the integrated circuit, it is easy to obtain the hardware circuit that implements the logical method flow.
[0100] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to make the controller implement the same function in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules implementing the method or structures within the hardware component.
[0101] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.
[0102] For the convenience of description, when describing the above devices, they are described separately as various units according to their functions. Of course, when implementing the embodiments of the present application, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0103] Those skilled in the art should understand that one or more embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, one or more embodiments of the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.
[0104] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple blocks.
[0105] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple blocks.
[0106] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple blocks.
[0107] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0108] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. The memory is an example of computer-readable media.
[0109] A computer-readable medium includes permanent and non-permanent, removable and non-removable media and can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information accessible by a computing device. As defined herein, a computer-readable medium does not include transitory computer-readable media, such as modulated data signals and carrier waves.
[0110] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising the element.
[0111] One or more embodiments of the present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. One or more embodiments of the present application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0112] Each embodiment in the present application is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and reference can be made to the relevant parts of the method embodiments for the related content.
[0113] The above are only embodiments of this document and are not intended to limit this document. For those skilled in the art, various modifications and variations can be made to this document. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this document shall be included within the scope of the claims of this document.
Claims
1. A large model training method for the cement industry, characterized in that Including the following steps: Obtain the data to be processed for the cement industry; According to the preset business scenario, divide the data to be processed into multiple sub-samples; Train the first sub-model based on the multiple sub-samples, with one-to-one correspondence between the sub-samples and the first sub-model; Set a vocabulary table and a calculation parameter table for standard questions in the cement industry for each of the first sub-models; Using the data to be processed as input, train the second sub-model based on the vocabulary table of standard questions and the calculation parameter table; And Using the data to be processed as input, train the first sub-model and the second sub-model based on the vocabulary table of standard questions and the calculation parameter table.
2. The method according to claim 1, characterized in that After dividing into multiple sub-samples, the method further includes: Calculate the sample quantity of each of the sub-samples; Determine whether the sample quantity of each of the sub-samples reaches a preset value; and Perform data augmentation on the sub-samples whose sample quantity does not reach the preset value.
3. The method according to claim 1, characterized in that Training the first sub-model based on the multiple sub-samples includes: Convert the data in the sub-samples into a preset data format, where the data format includes: calculation parameters, domain labels, question labels, and target answers; Input the calculation parameters into the first sub-model to obtain a predicted answer; Calculate the first cosine similarity between the predicted answer and the question label; Calculate the second cosine similarity between the predicted answer and the target answer; Calculate the score of the predicted answer based on the first cosine similarity and the second cosine similarity; If the score of the predicted answer is not greater than the preset value, perform the backpropagation process of the corresponding parameters to update the parameters of the second pre-trained model; perform the above operations until the score of the predicted answer satisfies being greater than the preset value; and If the score of the predicted answer is greater than the preset value, determine that the predicted answer is correct and end the current process.
4. The method according to claim 1, characterized in that The second sub-model uses an attention mechanism.
5. The method according to claim 4, characterized in that Using the data to be processed as input, training the second sub-model based on the vocabulary table of standard questions and the calculation parameter table includes: Convert the vocabulary table of standard questions and the calculation parameter table into key vectors; and Based on the attention mechanism, the second sub-model outputs the standard questions and calculation parameters corresponding to the first sub-model.
6. The method according to claim 1, characterized in that Using the data to be processed as input, training the first sub-model and the second sub-model based on the vocabulary table of standard questions and the calculation parameter table includes: Using the data to be processed as input, based on the vocabulary table of standard questions and the calculation parameters, the second sub-model outputs the standard questions and calculation parameters corresponding to the first sub-model; According to the standard questions, the first sub-model obtains the corresponding calculation parameters, where the calculation parameters include: domain labels; The first sub-model uses the calculation parameters as input and outputs a predicted answer; Calculate the loss function value of the predicted answer and the loss function value of the domain label respectively; Determine the total loss function value according to the loss function values of the predicted answers of each of the first sub-models and the loss function value of the domain label; Determine whether the total loss function value meets the preset conditions; When the total loss function value does not meet the preset conditions, perform the backpropagation process of the parameters and repeat the above process until the total loss function value meets the preset conditions; and When the total loss function value meets the preset conditions, end the current process.
7. A large model training device for the cement industry, characterized in that, Including: An acquisition module for acquiring data to be processed; A data processing module for dividing the data to be processed into multiple sub-samples according to a preset business scenario; A first training module for training a first sub-model based on the multiple sub-samples, with the sub-samples corresponding to the first sub-models one by one; A second training module for setting a vocabulary of standard questions and a calculation parameter table for each of the first sub-models; Using the data to be processed as input, training a second sub-model based on the vocabulary of the standard questions and the calculation parameter table; And A third training module for using the data to be processed as input, and training the first sub-model and the second sub-model based on the vocabulary of the standard questions and the calculation parameter table.
8. The apparatus according to claim 7, wherein The first training module is used to convert the data in the sub-sample into a preset data format, and the data format includes: calculation parameters, domain labels, question labels, and target answers; input the calculation parameters into the first sub-model to obtain a predicted answer; calculate the first cosine similarity between the predicted answer and the question label; calculate the second cosine similarity between the predicted answer and the target answer; calculate the score of the predicted answer according to the first cosine similarity and the second cosine similarity; if the score of the predicted answer is not greater than a preset value, perform the backpropagation process of the corresponding parameters to update the parameters of the second pre-trained model; perform the above operations until the score of the predicted answer meets the condition of being greater than the preset value; and if the score of the predicted answer is greater than the preset value, determine that the predicted answer is correct and end the current process.
9. The apparatus according to claim 7, wherein The third training module is used to use the data to be processed as input, and based on the vocabulary of the standard questions and the calculation parameters, the second sub-model outputs the standard questions and calculation parameters corresponding to the first sub-model; According to the standard problem, the first sub-model obtains corresponding calculation parameters, and the calculation parameters include: domain labels; the first sub-model takes the calculation parameters as inputs and outputs predicted answers; calculates the loss function value of the predicted answers and the loss function value of the domain labels respectively; determines the total loss function value according to the loss function values of the predicted answers and the domain labels of each first sub-model; determines whether the total loss function value meets a preset condition; when the total loss function value does not meet the preset condition, performs the backpropagation process of the parameters and repeats the above process until the total loss function value meets the preset condition; and when the total loss function value meets the preset condition, ends the current process.
10. A storage medium for storing computer-executable instructions, characterized in that, The computer-executable instructions, when executed, implement the steps of the large model training method for the cement industry according to any one of claims 1-6.
Citation Information
Patent Citations
Model training method, business risk control method and device
CN114997472A
Data processing method and device, computer equipment, storage medium and program product
CN117764116A
Multi-task processing model training method and device and multi-task processing method
CN118171111A
Multi-model, multi-task trained neural network for analyzing unstructured and semi-structured electronic documents
US20210286989A1