A Construction Method and Device for a Large Model in the Vertical Field of Coal Mines

By constructing a large model of vertical coal mines and combining field and general corpus data, the problem of insufficient knowledge learning in the coal mine professional field is solved, deep learning and precise expression of professional knowledge in the coal mine industry is achieved, the model's understanding and processing capabilities in the coal mine field are improved, and the level of safe production is improved.

CN119848553BActive Publication Date: 2025-07-29BEIJING LONGRUAN TECHNOLOGIES INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510323917.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-07-29
Estimated Expiration
2045-03-19

AI Technical Summary

Technical Problem

The general-language model lacks knowledge learning in the coal mine field, which leads to the model being unable to accurately understand and generate relevant content when dealing with specific terms and workflows of coal mines, which is prone to "illusion".

Method used

By obtaining the domain corpus data and the general corpus data, the domain vocabulary list and the general vocabulary list are constructed separately, and weighted fusion is carried out based on their respective weights to form a fusion vocabulary list. The embedding model is trained using the fusion vocabulary list, the pre-trained original large language model is loaded and incrementally pre-trained to obtain the coal mine vertical field large model.

Benefits of technology

The model's understanding of the professional knowledge of the coal mine industry has been improved, the "illusion" phenomenon has been avoided, and the performance accuracy in tasks such as coal mine safety management, technical operation and mine monitoring has been improved, and more professional and accurate solutions have been provided, which has improved coal mine management and operation efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119848553B_ABST
    Figure CN119848553B_ABST
Patent Text Reader

Abstract

The present invention provides a method and device for constructing a large model in the vertical field of coal mines, relating to the technical field of deep learning. The method includes: obtaining domain corpus data and general corpus data, and respectively constructing a domain word list and a general word list based on the domain corpus data and the general corpus data; performing weighted fusion on each word segment in the general word list and the domain word list based on their respective weights to obtain a fused word list; training an embedding model using the domain corpus data and the general corpus data based on the fused word list; loading a pre-trained original large language model, and replacing the embedding layer of the original large language model with the trained embedding model to obtain an updated large language model; performing incremental pre-training on the updated large language model using the domain corpus data and the general corpus data to obtain a large model in the vertical field of coal mines. The large model in the vertical field of coal mines constructed by the present invention realizes the accurate expression of professional knowledge in the coal mining industry.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of deep learning, and specifically relates to a method and device for constructing a large model in the coal mine vertical field. Background Art

[0002] A coal mine is an open and complex giant system. With the development of technologies such as mine digitization and intelligent mines, mine personnel use digital technologies to monitor and collect data on the mine environment, and quickly make responses and decisions based on the real-time on-site environment, effectively improving the safety production level of coal mines. After the emergence of large language models, the qualitative change brought about by the quantitative change enables large models to have "emergent capabilities" and can handle some complex problems that small models cannot solve.

[0003] Although general large language models have very powerful general capabilities, due to problems such as uneven distribution of training data, overfitting of model parameters, and label noise, general large language models have insufficient learning of professional knowledge in the coal mine field and are extremely prone to the "hallucination" phenomenon. Therefore, general large language models lack professional data for the coal mine industry, resulting in the model being unable to accurately understand and generate relevant content when dealing with coal mine specific terms and work processes. Therefore, how to construct a large model in the coal mine vertical field is an urgent problem to be studied. Summary of the Invention

[0004] The present invention provides a method and device for constructing a large model in the coal mine vertical field, aiming to solve the problems existing in the above background art.

[0005] To solve the above technical problems, the present invention is implemented as follows:

[0006] In a first aspect, the present invention provides a method for constructing a large model in the coal mine vertical field, the method comprising:

[0007] Obtain domain corpus data and general corpus data, and respectively construct a domain vocabulary and a general vocabulary based on the domain corpus data and the general corpus data;

[0008] Perform weighted fusion on each token in the general vocabulary and the domain vocabulary based on their respective weights to obtain a fused vocabulary, where the weights represent the relative frequencies of the tokens appearing in the domain corpus data and the general corpus data;

[0009] Based on the fused vocabulary, use the domain corpus data and the general corpus data to train an embedding model, where the embedding model is used to map tokens to corresponding word vectors based on semantic information;

[0010] Load a pre-trained original large language model, and replace the embedding layer of the original large language model with the trained embedding model to obtain an updated large language model;

[0011] Incrementally pre-train the updated large language model using the domain corpus data and the general corpus data to obtain a large model for the coal mine vertical domain.

[0012] Optionally, the weighted fusion of each token in the general vocabulary and the domain vocabulary based on their respective weights to obtain a fused vocabulary includes:

[0013] In the general vocabulary and the domain vocabulary, calculate the respective word frequencies and weights of each token;

[0014] Based on the weights, linearly combine each token in the general vocabulary and the domain vocabulary to obtain the fused vocabulary.

[0015] Optionally, the calculating the respective word frequencies and weights of each token in the general vocabulary and the domain vocabulary includes:

[0016] For each token in the general vocabulary and the domain vocabulary, determine the first word frequency as the frequency of occurrence of the token in the domain corpus data, or determine the second word frequency as the frequency of occurrence of the token in the general corpus data;

[0017] For each token in the general vocabulary, calculate the relative frequency of the token in the general corpus data according to the first word frequency, and determine the relative frequency of the token in the general corpus data as the first linear term;

[0018] For each token in the domain vocabulary, calculate the relative frequency of the token in the domain corpus data according to the second word frequency, and determine the relative frequency of the token in the general corpus data as the second linear term;

[0019] For the same token in the general vocabulary and the domain vocabulary, determine the logarithm of the ratio of the first word frequency and the second word frequency of the token as the logarithmic term;

[0020] Respectively assign preset weight adjustment parameters to the first linear term, the second linear term, and the logarithmic term, and calculate the weight of each token.

[0021] Optionally, after obtaining the fused vocabulary, the method further includes:

[0022] Check whether there are unmatched tokens in the domain vocabulary, where the unmatched tokens are tokens that exist in the domain vocabulary but do not exist in the general vocabulary;

[0023] In the case of checking the unmatched token, add the token to the fused vocabulary.

[0024] Optionally, training the embedding model using the domain corpus data and the general corpus data based on the fusion vocabulary includes:

[0025] Mix the domain corpus data and the general corpus data according to a preset ratio to obtain training corpus data, where the general corpus data occupies a higher proportion than the domain corpus data;

[0026] Based on the fusion vocabulary, perform word segmentation on the training corpus data to obtain a plurality of training word segments;

[0027] Based on the plurality of training word segments, train the embedding model according to preset training parameters, where the training parameters include batch size, learning rate, and loss function.

[0028] Optionally, the method further includes:

[0029] Obtain the high-dimensional word vectors output by the embedding model, and determine the core semantic information and secondary semantic information represented by the high-dimensional word vectors;

[0030] Perform dimensionality reduction on the high-dimensional word vectors to obtain corresponding low-dimensional word vectors, where the core semantic information represented by the high-dimensional word vectors is stored in the head dimension of the low-dimensional word vectors, and the secondary semantic information represented by the high-dimensional word vectors is stored in the tail dimension of the low-dimensional word vectors;

[0031] Save the low-dimensional word vectors in the predefined data structure of the embedding model.

[0032] Optionally, replacing the embedding layer of the original large language model with the trained embedding model to obtain an updated large language model includes:

[0033] Replace the embedding layer of the original large language model with the trained embedding model to obtain the embedding layer of the updated large language model;

[0034] Initialize the embedding layer of the updated large language model using the word vectors output by the trained embedding model.

[0035] Optionally, incrementally pre-training the updated large language model using the domain corpus data and the general corpus data to obtain a large model for the coal mine vertical domain includes:

[0036] Mix the domain corpus data and the general corpus data according to a preset ratio to obtain training corpus data, where the general corpus data occupies a higher proportion than the domain corpus data;

[0037] Add a low-rank matrix and a bias matrix to the weight matrix of the linear layer of the updated large language model;

[0038] Freeze the other layers of the updated large language model except the embedding layer, and use the training corpus data to train the embedding layer of the updated large language model;

[0039] When the preset iteration condition is reached, unfreeze the other layers, and use the training corpus data to fine-tune and train all layers of the updated large language model to obtain the low-rank adaptation model corresponding to the large language model;

[0040] Merge the low-rank adaptation model with the original large language model to obtain the large model for the coal mine vertical domain.

[0041] Optionally, the obtaining domain corpus data and general corpus data, and respectively constructing a domain vocabulary and a general vocabulary based on the domain corpus data and the general corpus data, includes:

[0042] Organize the pre-obtained coal mine vertical domain dataset into the domain corpus data, and organize the pre-obtained open-source dataset into the general corpus data. The coal mine vertical domain dataset includes equipment monitoring data, mine geological exploration report data, forum data, and industry website data;

[0043] Process the general corpus data into the general vocabulary;

[0044] Process the domain corpus data into the domain vocabulary with the same format as the general vocabulary.

[0045] In a second aspect, the present invention provides a device for constructing a large model for the coal mine vertical domain, the device includes:

[0046] An acquisition module, configured to acquire domain corpus data and general corpus data, and respectively construct a domain vocabulary and a general vocabulary based on the domain corpus data and the general corpus data;

[0047] A fusion module, configured to perform weighted fusion on each word segment in the general vocabulary and the domain vocabulary based on their respective weights to obtain a fusion vocabulary, where the weights represent the relative frequencies of the word segments appearing in the domain corpus data and the general corpus data;

[0048] A first training module, configured to train an embedding model based on the fusion vocabulary by using the domain corpus data and the general corpus data, where the embedding model is used to map a word segment to a corresponding word vector based on semantic information;

[0049] An update module, configured to load a pre-trained original large language model, and replace the trained embedding model into the embedding layer of the original large language model to obtain an updated large language model;

[0050] A second training module, configured to perform incremental pre-training on the updated large language model by using the domain corpus data and the general corpus data to obtain a large model for the coal mine vertical domain.

[0051] Optionally, the fusion module includes:

[0052] A weight calculation sub-module, configured to calculate the word frequency and weight of each token in the general vocabulary and the domain vocabulary;

[0053] A linear combination sub-module, configured to linearly combine each token in the general vocabulary and the domain vocabulary based on the weights to obtain the fused vocabulary.

[0054] Optionally, the weight calculation sub-module includes:

[0055] A word frequency determination unit, configured to, for each token in the general vocabulary and the domain vocabulary, determine the frequency of occurrence of the token in the domain corpus data as the first word frequency, or determine the frequency of occurrence of the token in the general corpus data as the second word frequency;

[0056] A first relative frequency determination unit, configured to, for each token in the general vocabulary, calculate the relative frequency of the token in the general corpus data according to the first word frequency, and determine the relative frequency of the token in the general corpus data as the first linear term;

[0057] A second relative frequency determination unit, configured to, for each token in the domain vocabulary, calculate the relative frequency of the token in the domain corpus data according to the second word frequency, and determine the relative frequency of the token in the general corpus data as the second linear term;

[0058] A logarithm term determination unit, configured to, for the same token in the general vocabulary and the domain vocabulary, determine the logarithm of the ratio of the first word frequency and the second word frequency of the token as the logarithm term;

[0059] A linear combination unit, configured to respectively assign preset weight adjustment parameters to the first linear term, the second linear term, and the logarithm term, and calculate the weight of each token.

[0060] Optionally, the device further includes:

[0061] A checking module for checking whether there are unmatched word segments in the domain vocabulary, where the unmatched word segments are word segments that exist in the domain vocabulary but do not exist in the general vocabulary;

[0062] An adding module for adding the word segment to the fusion vocabulary when the unmatched word segment is detected.

[0063] Optionally, the first training module includes:

[0064] A first mixing sub-module for mixing the domain corpus data and the general corpus data according to a preset ratio to obtain training corpus data, where the general corpus data occupies a higher proportion than the domain corpus data;

[0065] A word segmentation processing sub-module for performing word segmentation processing on the training corpus data based on the fusion vocabulary to obtain a plurality of training word segments;

[0066] A first training sub-module for training the embedding model based on the plurality of training word segments according to preset training parameters, where the training parameters include batch size, learning rate, and loss function.

[0067] Optionally, the device further includes:

[0068] A word vector acquisition module for acquiring the high-dimensional word vectors output by the embedding model and determining the core semantic information and secondary semantic information represented by the high-dimensional word vectors;

[0069] A dimensionality reduction module for performing dimensionality reduction processing on the high-dimensional word vectors to obtain corresponding low-dimensional word vectors, where the core semantic information represented by the high-dimensional word vectors is stored in the head dimension of the low-dimensional word vectors, and the secondary semantic information represented by the high-dimensional word vectors is stored in the tail dimension of the low-dimensional word vectors;

[0070] A word vector storage module for storing the low-dimensional word vectors in a data structure predefined by the embedding model.

[0071] Optionally, the update module includes:

[0072] A replacement sub-module for replacing the embedding layer of the original large language model with the trained embedding model to obtain the embedding layer of the updated large language model;

[0073] An initialization sub-module for initializing the embedding layer of the updated large language model by using the word vectors output by the trained embedding model.

[0074] Optionally, the second training module includes:

[0075] A second mixing sub-module, configured to mix the domain corpus data and the general corpus data according to a preset ratio to obtain training corpus data, wherein the proportion of the general corpus data is higher than that of the domain corpus data;

[0076] A matrix addition sub-module, configured to add a low-rank matrix and a bias matrix to the weight matrix of the linear layer of the updated large language model;

[0077] A freezing sub-module, configured to freeze other layers of the updated large language model except the embedding layer, and use the training corpus data to train the embedding layer of the updated large language model;

[0078] A thawing sub-module, configured to thaw the other layers when a preset iteration condition is reached, and use the training corpus data to fine-tune and train all layers of the updated large language model to obtain a low-rank adaptation model corresponding to the large language model;

[0079] A merging sub-module, configured to merge the low-rank adaptation model with the original large language model to obtain the large model for the coal mine vertical domain.

[0080] Optionally, the obtaining module includes:

[0081] An arrangement sub-module, configured to arrange a pre-obtained coal mine vertical domain data set into the domain corpus data, and arrange a pre-obtained open source data set into the general corpus data, where the coal mine vertical domain data set includes equipment monitoring data, mine geological exploration report data, forum data, and industry website data;

[0082] A first processing sub-module, configured to process the general corpus data into the general vocabulary;

[0083] A second processing sub-module, configured to process the domain corpus data into the domain vocabulary that is unified with the format of the general vocabulary.

[0084] The technical solution provided by the present invention at least brings the following beneficial effects:

[0085] The present invention constructs a large model in the vertical field of coal mines. By combining domain corpus data and general corpus data, it realizes in-depth learning and accurate expression of professional knowledge in the coal mining industry. Through weighted fusion of the domain vocabulary and the general vocabulary, it can effectively improve the model's understanding ability of specific terms and work processes in the coal mining field, avoiding the "hallucination" phenomenon that often occurs when general large language models handle coal mining professional problems. In addition, the embedding model can fully capture domain semantic information during the training process, thereby improving the performance accuracy of the model in tasks such as coal mine safety management, technical operations, and mine monitoring. Finally, the large model in the vertical field of coal mines obtained through incremental pre-training can, while ensuring the generality of the model, provide more professional and accurate coal mine solutions, greatly improving the management, operation efficiency, and work safety level of coal mines, and providing strong support for the intelligent and digital transformation of the coal mining industry. BRIEF DESCRIPTION OF THE DRAWINGS

[0086] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0087] Figure 1 It is a schematic diagram of the steps of a method for constructing a large model in the vertical field of coal mines provided by an embodiment of the present invention;

[0088] Figure 2 It is a block diagram of the structure of a device for constructing a large model in the vertical field of coal mines provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0089] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention belong to the scope of protection of the present invention.

[0090] With the rapid development of digital and intelligent technologies, the coal mining industry is facing unprecedented opportunities and challenges. Traditional coal mine safety management and operation models often rely on manual experience and professional knowledge. However, due to the complexity and diversity of coal mine safety issues, existing technologies are unable to cope with a large number of coal mine safety problems. In addition, although existing large language models have shown strong generality in multiple fields, there are significant deficiencies in knowledge learning and application in the coal mine professional field, resulting in insufficient capabilities of the models in dealing with coal mine specific terms, work processes, and safety management, prone to the "hallucination" phenomenon, and unable to provide accurate and professional solutions.

[0091] To address the above technical deficiencies, the core concept of the present invention is to construct a domain vocabulary and a general vocabulary by obtaining domain corpus data and general corpus data respectively, and perform weighted fusion based on their respective weights to form a fused vocabulary. This can not only effectively improve the model's understanding ability of professional knowledge in the coal mining industry, but also enhance the model's performance in specific fields while ensuring generality. By performing incremental pre-training on the updated large language model, the present invention aims to create a large model for the coal mine vertical domain with high efficient reasoning ability and professional knowledge, providing strong technical support for the safe production and intelligent management of the coal mining industry.

[0092] Figure 1 is a schematic diagram of the steps of a method for constructing a large model for the coal mine vertical domain provided by an embodiment of the present invention. Please refer to Figure 1 , the method includes:

[0093] Step S101, obtain domain corpus data and general corpus data, and construct a domain vocabulary and a general vocabulary based on the domain corpus data and the general corpus data respectively.

[0094] In the coal mine field, there are a large number of professional terms, unique work processes, and strict safety specifications, which are usually not included in the knowledge scope of general large language models. Compared with traditional methods, the present invention has pre-collected a large amount of rich text and data related to the coal mine field, including industrial documents, technical reports, operation manuals, safety records, etc., covering all aspects from mine operation to safety management. Clean and preprocess the collected data, including steps such as removing noise, standardizing text formats, and dealing with missing values, to ensure high-quality input during the subsequent model training stage. During the construction process of the coal mine vertical domain model, integrate this professional knowledge into the incremental pre-training stage to improve the accuracy and effectiveness of the model in understanding and dealing with coal mine field specific problems, significantly enhancing the model's performance in the coal mine field and providing more professional, accurate, and safe solutions. This not only improves the efficiency of coal mine management and operation, but also provides strong support for the industry's safe production and accident prevention.

[0095] The domain corpus data is text data specifically collected for the coal mining industry, including professional terms, industry standards, operating procedures, technical reports, safety management documents, etc. in the coal mining field, which can reflect the specific knowledge and language usage habits in the coal mining field. It can be understood that the core principle and technical framework of the present invention have wide applicability. With the equivalent replacement of the domain corpus data, the present invention can also be applied to other vertical fields, such as medical, financial, legal, manufacturing and other industries. Just collect relevant professional text data for a specific field, construct a corresponding domain vocabulary, and perform data cleaning and preprocessing to effectively train a pre-trained model that meets the requirements of that field.

[0096] The general corpus data refers to extensive text data that is not specific to the coal mining field, covering the language used in daily life, and is sourced from various public resources, such as news articles, social media, Wikipedia, books, etc., reflecting the usage of daily language and people's general expression methods.

[0097] The domain word segmentation is obtained by performing word segmentation on the domain corpus data, and includes commonly used word segmentation in the coal mining field, such as "mine", "coal mining", "safety regulations", "ventilation system", "miner", etc. Similarly, the general word segmentation is obtained by performing word segmentation on the general corpus data, and includes commonly used word segmentation in daily life, applicable to various language scenarios and topics, such as "person", "place", "work", "study", "safety", etc. In an optional implementation manner, step S101 specifically includes steps S1011 - S1013:

[0098] Step S1011, organize the pre-obtained coal mining vertical domain dataset into the domain corpus data, and organize the pre-obtained open-source dataset into the general corpus data. The coal mining vertical domain dataset includes equipment monitoring data, mine geological exploration report data, forum data, and industry website data.

[0099] The coal mining vertical domain dataset is sourced from various data in the coal mining industry, including equipment monitoring data (monitoring data of equipment for mining, transportation, ventilation, etc. in coal mines), mine geological exploration report data (including detailed descriptions of mine geological conditions, such as rock formations, mineral deposits, groundwater levels), forum data (discussion data from relevant forums in the coal mining industry), and industry website data (articles, news, technical reports, etc. from coal mining industry websites). The open-source dataset is not specific to the coal mining field, but is extensive text data sourced from various public resources.

[0100] The quality, diversity, and relevance of coal mine vertical domain datasets and open-source datasets play a crucial role in subsequent training quality and application effects. The processing flow of the embodiments of the present invention for coal mine vertical domain datasets or open-source datasets specifically includes determining data sources, data collection, data cleaning, data organization and arrangement, privacy and security, and data quality assessment.

[0101] Specifically, the steps of determining data sources include determining data sources and obtaining data permissions. First, identify and determine coal mine vertical domain datasets available for the coal mine vertical domain, covering all aspects from mine operation to safety management. Ensure appropriate data usage permissions and authorizations are obtained from suitable data providers to ensure the legality and compliance of data collection. The steps of data collection include using different methods to collect data from different sources. For industry data such as equipment monitoring data and mine geological exploration reports, data is manually collected by experts or personnel in the field. For public open-source data such as forum discussions and industry websites, web crawler tools are used to scrape relevant content from online resources. The data cleaning steps include removing noise and redundancy, handling missing values, and standardization and normalization to ensure the quality, consistency, and comparability of data. Specifically, identify and remove noise and redundant information in the data, such as blank records, duplicate data, incomplete entries, etc., to ensure data quality. Handle missing values in the data by filling default values, interpolation, or deleting records with missing data to ensure the integrity and availability of the data. Standardize the data, such as unifying date formats, units, and standardizing specific terms, to ensure the consistency and comparability of the data. The steps of data organization and arrangement include that the collected data may come from multiple sources and formats and need to be uniformly arranged, including file format conversion, text encoding unification, data structure standardization, etc., for subsequent processing and analysis. To facilitate management and retrieval, an index and directory structure of the data can be selected. This can be achieved through a database management system or a simple folder structure to ensure the orderly storage and quick access of the data. The steps of privacy and security and quality assessment include anonymizing or desensitizing data that may contain personal identity or sensitive information to protect data privacy and security. Establish appropriate access control policies and permission management mechanisms to ensure that only authorized personnel can access and use the dataset. The steps of data quality assessment include conducting quality assessment and verification on the cleaned and arranged data, including inspections in aspects such as data integrity, accuracy, consistency, and availability, to ensure that the dataset meets the expected quality standards.

[0102] Step S1012, process the general corpus data into the general word list.

[0103] The Llama series of models possess very powerful capabilities. By using tools such as Chinese-Alpaca-2 to expand the Chinese vocabulary of the Llama 2 series of models and then performing incremental pre-training, a general vocabulary is generated based on the general corpus data.

[0104] Step S1013, process the domain corpus data into the domain vocabulary that is unified with the format of the general vocabulary.

[0105] In this embodiment, the SentencePiece tool is used to train the domain vocabulary based on the coal mine vertical domain dataset, and then the tokenization results are aligned with the general vocabulary of the Chinese-Alpaca-2 tool. The SentencePiece tool is an unsupervised tokenization tool that converts text data into a standard tokenized form by adaptively learning sub-word units in the text. Applying the SentencePiece tool to the training of the coal mine vertical domain dataset generates a domain vocabulary specifically for the coal mine domain based on the domain corpus data.

[0106] Align the trained domain vocabulary and the general vocabulary to ensure that the domain vocabulary can be in the same format as the general vocabulary, enabling both to share the same tokenization and encoding methods, thereby avoiding issues of inconsistency or conflict during model training. The goal of aligning the tokenization results is to ensure that specific terms in the coal mine domain are appropriately represented in the finally generated large model for the coal mine vertical domain and can be seamlessly combined with the tokenization of the general language during the inference of the large model for the coal mine vertical domain.

[0107] Step S102, perform weighted fusion on each token in the general vocabulary and the domain vocabulary based on their respective weights to obtain a fused vocabulary, where the weights represent the relative frequencies of the tokens in the domain corpus data and the general corpus data.

[0108] In this embodiment, weighted fusion is performed on the tokens in the general vocabulary and the domain vocabulary to generate a fused vocabulary that accurately represents the importance of each token in different corpora. The fused vocabulary not only retains the knowledge unique to the coal mine domain but also maintains the broad applicability of the general language. Specifically, the frequencies of tokens in the domain corpus data and the general corpus data may vary, so the weight of each token is determined by its relative frequencies in the domain corpus data and the general corpus data to ensure that domain-specific tokens receive higher weights in the fused vocabulary, thereby enhancing the model's ability to understand domain-related content. In short, the purpose of weighted fusion is to balance the importance of general tokens and domain tokens, enabling the finally generated fused vocabulary to effectively process general text and accurately identify domain-specific tokens in the coal mine domain, thereby improving the performance of the large model for the coal mine vertical domain in the professional field.

[0109] Step S103: Based on the fused vocabulary, train the embedding model using the domain corpus data and the general corpus data. The embedding model is used to map word segments to corresponding word vectors based on semantic information.

[0110] Use the fused vocabulary to train the embedding model to map word segments to corresponding word vectors and optimize the representation of word vectors through semantic information. Specifically, the fused vocabulary combines word segments in the general corpus data and the domain corpus data and assigns corresponding weights to each word segment, so that the embedding model can learn how to generate corresponding word vectors according to the context information of the word segments during the training process. When training the embedding model, input the text data of the domain corpus data and the general corpus data into the embedding model. The embedding model continuously adjusts the word vectors so that word segments with similar meanings have close vector representations, and then through semantic information, each word segment can be correctly positioned in the semantic space. The goal of this training process is to enable the word vectors to accurately capture the semantic relationships between word segments, which can reflect both the meanings in the general context and the specific knowledge in the coal mine domain. Finally, the trained embedding model will convert word segments (whether general or specific to the coal mine domain) into meaningful word vectors.

[0111] Step S104: Load the pre-trained original large language model and replace the embedding layer of the original large language model with the trained embedding model to obtain an updated large language model.

[0112] First, select and load a general large language model that has undergone large-scale pre-training as the original large language model, that is, the base model. Common base models include BERT (Bidirectional Encoder Representations from Transformers), GPT (Generative Pre-trained Transformer), RoBERTa (Robustly optimized BERT approach), T5 (Text-To-Text Transfer Transformer), etc. For the coal mine vertical domain, in this embodiment, the open-source Llama-2-7b model using the GPT framework is selected as the base model. The original large language model has been trained with a large amount of general data and has good language understanding ability, but its embedding layer (i.e., the part of the word vectors) may not be suitable for the specific word segmentation and context of a particular domain. Therefore, update its embedding layer so that the large language model can better handle the specific terms and language in the coal mine domain. Replace the trained embedding model (that is, the embedding layer corresponding to the fused vocabulary obtained through step S103) into the embedding layer of the original large language model. The role of the embedding layer is to convert the input word segmentation into vector representations, and these word vectors will serve as the basis for the large language model to process text. Through replacement, the new embedding model can provide more domain-specific word vectors for the large language model, enabling the model to more accurately understand the professional terms in the domain.

[0113] Step S105: Use the domain corpus data and the general corpus data to perform incremental pre-training on the updated large language model to obtain a large model for the coal mine vertical domain.

[0114] The basic idea of incremental pre-training is to further train on the existing updated large language model (which has replaced the embedding layer through step S104) instead of training from scratch. The advantage of incremental training is that, while maintaining most of the learned knowledge, the model can be fine-tuned with additional domain corpus data to enhance its performance in the coal mine domain. In this way, a model that performs well in the general context can better adapt to the special needs and knowledge of the coal mine domain. During the incremental pre-training process, the domain corpus data and the general corpus data are used for further training. By using both types of data simultaneously, the large language model can enhance domain knowledge while retaining the ability to process general language. After incremental pre-training, a large model for the coal mine vertical domain that can perform more precisely on coal mine domain data is finally obtained. For example, for texts such as coal mine operation manuals and safety records, the large model for the coal mine vertical domain can understand and generate more professional content. At the same time, since the knowledge of general data is still retained, the large model for the coal mine vertical domain can also maintain a certain degree of generality in tasks involving other domains.

[0115] The present invention constructs a large model for the coal mine vertical domain, combines domain corpus data and general corpus data, and realizes in-depth learning and precise expression of professional knowledge in the coal mine industry. By weighted fusion of the domain vocabulary and the general vocabulary, the understanding ability of the model for specific terms and work processes in the coal mine domain can be effectively improved, avoiding the "hallucination" phenomenon that often occurs in general large language models when dealing with coal mine professional problems. In addition, the embedding model can fully capture domain semantic information during the training process, thereby improving the performance accuracy of the model in tasks such as coal mine safety management, technical operations, and mine monitoring. Finally, the large model for the coal mine vertical domain obtained through incremental pre-training can provide more professional and accurate coal mine solutions while ensuring the generality of the model, greatly improving the efficiency of coal mine management, operation, and safety production levels, and providing strong support for the intelligent and digital transformation of the coal mine industry.

[0116] In an alternative embodiment, the weighted fusion of each tokenization in the general vocabulary and the domain vocabulary based on their respective weights to obtain a fused vocabulary includes:

[0117] Step S1031, calculate the word frequency and weight of each tokenization in the general vocabulary and the domain vocabulary respectively.

[0118] For each segmented word in the general word list and the domain word list, calculate its occurrence frequency in their respective corpora. Words with a high occurrence frequency are more representative in the language. The calculation of weights is based on the relative importance of word frequencies in the entire corpus. The weight of each segmented word reflects its relative importance in the coal mine field and the general context. Through weights, it is possible to highlight those professional terms or segmented words that frequently appear in the domain corpus data and reduce the influence of words that are common in the general corpus but not very important in the domain. The design of weights usually takes into account the impact of domain-specific segmented words on the model, so that in the subsequent word list fusion process, the high weights of domain segmented words can be retained, thereby improving the performance of domain tasks. Through the calculation of word frequencies and weights in this step, it can provide a basis for the word list fusion in the subsequent step S104. During fusion, domain segmented words with high weights will receive more attention and optimization, making the resulting large model for the coal mine vertical domain more accurate in coal mine domain tasks. At the same time, the processing of general segmented words will also be based on their weights to ensure that the model does not lose its general language processing ability due to overemphasis on a certain type of segmented word during the fusion process.

[0119] Step S1032, based on the weights, linearly combine each segmented word in the general word list and the domain word list to obtain the fused word list.

[0120] In this step, the segmented words in the general word list and the domain word list are not simply added together, but are weighted and combined according to the weight of each segmented word, which can effectively highlight important professional terms in the domain while retaining the broad applicability of general vocabulary. When performing linear combination, the weight of each segmented word reflects its importance and relevance in its respective corpus. Based on the weights, ensure that segmented words that frequently appear in the domain corpus and are crucial for the model task occupy a more important position in the fused word list. On the contrary, those that are not very important in the coal mine field but are common in the general context are appropriately reduced in their influence. In an optional implementation manner, calculating the respective word frequencies and weights of each segmented word in the general word list and the domain word list specifically includes the following steps:

[0121] For each segmented word in the general word list and the domain word list, determine the first word frequency as the occurrence frequency of the segmented word in the domain corpus data, or determine the second word frequency as the occurrence frequency of the segmented word in the general corpus data.

[0122] For each segmented word in the general word list and the domain word list calculate their occurrence frequencies in their respective corpora to obtain the first word frequency (domain word frequency) indicating the number of occurrences of the segmented word in the domain corpus, and obtain the second word frequency (general word frequency) indicating the segmented word Number of occurrences in the general corpus.

[0123] For each word segment in the general word list, calculate the relative frequency of the word segment in the general corpus data according to the first word frequency, and determine the relative frequency of the word segment in the general corpus data as the first linear term .

[0124] Calculate the relative frequency of each word segment in the general corpus data, that is, the domain relative frequency . Domain relative frequency As the first linear term, it represents the relative importance of the word segment in the domain. The calculation formula is as follows:

[0125]

[0126] In the formula, is the total occurrence frequency of all word segments in the domain corpus data.

[0127] For each word segment in the domain word list, calculate the relative frequency of the word segment in the domain corpus data according to the second word frequency, and determine the relative frequency of the word segment in the general corpus data as the second linear term .

[0128] Calculate the relative frequency of each word segment in the general corpus data, that is, the general relative frequency . General relative frequency As the second linear term, it represents the relative importance of the word segment in the general corpus. The calculation formula is as follows:

[0129]

[0130] s In the formula, is the total occurrence frequency of all word segments in the general corpus data.

[0131] For the same word segment in the general word list and the domain word list, determine the logarithm of the ratio of the first word frequency and the second word frequency of the word segment as the logarithmic term.

[0132] For the same word segment in the general word list and the domain word list, calculate the logarithm of the ratio of the first word frequency and the second word frequency of the word segment as the logarithmic term. As the weight of the word segment in the fusion word list, highlight the word segments with a relative frequency higher in the domain corpus than in the general corpus and smooth the ratio through the logarithm. The calculation formula is as follows:

[0133]

[0134] Assign preset weight adjustment parameters to the first linear term, the second linear term, and the logarithmic term respectively, and calculate the weight of each word segment.

[0135] Assign preset weight adjustment parameters to each linear term and logarithmic term. The weight adjustment parameters corresponding to the first linear term, the second linear term, and the logarithmic term are , and . The final weight is obtained by weighted summing the above linear term and logarithmic term:

[0136]

[0137] Combine the weight with the original word frequencies (the first word frequency and the second word frequency) through the above formula to obtain a smoother and more effective weight . In this embodiment, in order to make the word segments in the domain vocabulary more important than those in the general vocabulary in a specific task, it is selected that the value of is greater than the value of . It can be appropriately adjusted according to the relative importance of the word segments in the domain vocabulary in the general vocabulary. The value of should be lower than , but can be equivalent to or slightly higher than the value of

[0138] In an alternative embodiment, after obtaining the fused vocabulary, the method further includes: checking whether there are unmatched word segments in the domain vocabulary, where the unmatched word segments are those that exist in the domain vocabulary but do not exist in the general vocabulary; in the case of checking the unmatched word segments, adding the word segment to the fused vocabulary.

[0139] When constructing the fused vocabulary, the domain vocabulary and the general vocabulary are weighted and fused through weight adjustment parameters. Although the vast majority of word segments may have been successfully merged, in some specific cases, some word segments in the domain vocabulary may not appear in the general vocabulary. Unmatched word segments refer to the vocabulary that exists in the domain vocabulary but has no corresponding item in the general vocabulary. It may be due to the particularity of the domain vocabulary or the general vocabulary not covering the professional terms, abbreviations, etc. in this domain. By comparing the domain vocabulary and the general vocabulary, check which words only appear in the domain vocabulary but not in the general vocabulary. These words are the unmatched word segments. For example, the coal mining field may involve terms such as "gas leakage", "mining face", etc., which may not be common in the general corpus data, so they may not appear in the general vocabulary.

[0140] To ensure that the final integrated vocabulary can fully cover domain knowledge without missing any important domain terms, these unmatched word segments are extracted from the domain vocabulary and directly added to the final integrated vocabulary. In this way, the integrated vocabulary contains both the word segments of the general vocabulary and the word segments in the domain vocabulary that are not covered by the general vocabulary. The priority of adding unmatched word segments can be determined according to the importance or frequency of the word segments to ensure that the most common and important word segments are added to the integrated vocabulary first. Optionally, after obtaining the integrated vocabulary, it is also possible to check whether the vocabulary contains terms and common words specific to the coal mine domain to ensure that the required content is covered. According to the verification results and the feedback from domain experts, the aligned vocabulary is optimized to ensure that it covers the terms specific to the coal mine domain and the common words of the general language.

[0141] In an optional implementation manner, training the embedding model by using the domain corpus data and the general corpus data based on the integrated vocabulary includes: mixing the domain corpus data and the general corpus data according to a preset ratio to obtain training corpus data, where the proportion of the general corpus data is higher than that of the domain corpus data; performing word segment processing on the training corpus data based on the integrated vocabulary to obtain a plurality of training word segments; and training the embedding model according to the preset training parameters based on the plurality of training word segments, where the training parameters include batch size, learning rate, and loss function.

[0142] When training the embedding model, by combining the domain corpus data and the general corpus data and mixing the two types of corpus data, it is ensured that the training corpus data can both reflect domain knowledge and retain the broad coverage of the general language. The mixing ratio of the domain corpus data and the general corpus data is preset, and the proportion of the domain corpus data is relatively small, which can be 20% or 30%; the proportion of the general corpus data is relatively large.

[0143] To convert the text data into a form that can be understood by the embedding model, word segment processing is performed on the original text of the training corpus data to obtain the vocabulary form required for training the embedding model. The integrated vocabulary is used for word segment processing to ensure that the word segment results cover domain vocabulary and general vocabulary. Word segmenting is performed on the mixed training corpus data according to the integrated vocabulary. Through word segment processing, a plurality of training word segments are obtained, and each word segment is a unit that the embedding model needs to recognize. The finally obtained training word segments will be used as the input of the embedding model.

[0144] The goal of training an embedding model is to enable the model to learn how to convert the input word segments into high-dimensional word vectors, so that these word vectors can fully represent the semantic information of the vocabulary, thereby making accurate predictions or inferences in subsequent tasks. When training an embedding model, specific training parameters are set to guide the training process of the model, including: batch size, learning rate, and loss function. When training an embedding model, contrastive loss or multiple negatives sampling loss can be used as the loss function to help the embedding model learn the semantic relationships between word embeddings.

[0145] Table 1 shows an example of the training parameter settings of the embedding model. Please refer to Table 1:

[0146] The base model selected is Dmeta-Embedding-base, which is a pre-trained embedding model that can generate embedding vectors of text.

[0147] The train batch size is set to 16, and the amount of training corpus data used for each training is 16 samples.

[0148] The number of epochs is 20. During the training process, the training corpus data will be traversed 20 times (i.e., 20 epochs). Each epoch contains a complete pass of the training data.

[0149] The loss function selects MultipleNegativesSymmetricRankingLoss. The embedding model is trained by comparing the distances between positive and negative samples. The goal is to make the embeddings of positive samples closer and the embeddings of negative samples farther apart.

[0150] The learn rate is set to 2e-05, that is, 0.00002, which represents the step size for each parameter update. The learning rate scheduler selects the WarmupLinearSchedule method. At the beginning of training, the learning rate gradually increases and then decays linearly during the training process. This helps to avoid instability caused by too large a learning rate at the start of training and gradually reduces the learning rate as training progresses, improving the training accuracy.

[0151] The number of warmup steps is set to 5384, which means that at the beginning of training, the learning rate will gradually increase from an initial small value until it reaches the target learning rate of 2e-05.

[0152] The Evaluation steps are set to 2692, which means that during the training process, the embedding model is evaluated every 2692 steps (that is, several small batches in a training process) to ensure that there is no overfitting or performance degradation during the training process.

[0153] Table 1

[0154]

[0155] In an optional embodiment, the method further includes: obtaining a high-dimensional word vector output by the embedding model, and determining the core semantic information and secondary semantic information represented by the high-dimensional word vector; performing dimensionality reduction processing on the high-dimensional word vector to obtain a corresponding low-dimensional word vector, wherein the core semantic information represented by the high-dimensional word vector is stored in the head dimension of the low-dimensional word vector, and the secondary semantic information represented by the high-dimensional word vector is stored in the tail dimension of the low-dimensional word vector; and storing the low-dimensional word vector in a data structure predefined by the embedding model.

[0156] First, obtain high-dimensional word vectors from the embedding model. "High-dimensional" here refers to the multidimensional representation generated by the embedding model for each word. These high-dimensional representations (e.g., 768 or higher) distinguish them from the low-dimensional word vectors described later. In high-dimensional word vectors, the embedding model represents different levels of semantic information. Core semantic information refers to the most important and representative meaning of a word in a specific context, while secondary semantic information refers to possible other meanings of the word or additional contextual information. By analyzing high-dimensional word vectors, we can identify which dimensions primarily correspond to core semantics and which dimensions correspond to secondary semantics.

[0157] The purpose of dimensionality reduction processing on high-dimensional word vectors in this embodiment is to reduce storage space and computational complexity while retaining important semantic information as much as possible. The low-dimensional word vector obtained after dimensionality reduction has a low dimension (such as 256 or 512 dimensions) and is suitable for subsequent large language model training and reasoning. During the dimensionality reduction process, the core semantic information is retained in the head dimension of the low-dimensional word vector, while the secondary semantic information is stored in the tail dimension. Ensure that in the low-dimensional representation, the most important information (core semantics) can be accessed and used first, while the secondary information can be called when needed. This can improve the efficiency and accuracy of the model in specific tasks.

[0158] Finally, save the generated low-dimensional word vectors in the data structure predefined by the embedding model for subsequent use and invocation. The present invention improves the storage process of word vectors, which helps to quickly access and process word vectors and improve the running efficiency of the finally generated large model in the coal mine vertical field.

[0159] Table 2 shows the performance improvement of the trained embedding model. Please refer to Table 2:

[0160] Table 2

[0161]

[0162] The embedding dimension of the original embedding model is 768, and the FAHR index is 93.48. When the embedding dimension of the embedding model is reduced to 256, the FAHR index drops to 93.02. Although the dimension is reduced, the performance hardly decreases, with only a slight decrease of 0.46 percentage points. This indicates that after reducing the embedding dimension, the optimized model still maintains strong performance, and the storage and computing resource consumption are significantly reduced. When the dimension increases to 512, the FAHR value is the same as that in the case of 256 dimensions, still 93.02. That is to say, there is no significant difference in performance between the embedding models with 256 dimensions and 512 dimensions. This result shows that within a specific range of embedding dimensions, the performance of the embedding model tends to be stable, and the embedding dimension can be further optimized to reduce resource consumption without sacrificing performance. When the embedding dimension returns to 768, the FAHR value slightly rebounds to 93.13. Although it is still lower than the original embedding model (93.48), the performance gap is very small, with only a decrease of 0.35 percentage points. This indicates that the optimized embedding model can still maintain good performance in the case of high dimensions.

[0163] In an optional implementation manner, the replacing the trained embedding model into the embedding layer of the original large language model to obtain an updated large language model includes: replacing the embedding layer of the original large language model with the trained embedding model to obtain the embedding layer of the updated large language model; initializing the embedding layer of the updated large language model with the word vectors output by the trained embedding model.

[0164] The embedding layer is the first layer of the original large language model and is used to convert the input text into a vector representation. The embedding layer in the original large language model contains a general vocabulary and embedding method, which is not optimized for the specific coal mine domain. Therefore, the embedding layer originally belonging to the general large language model is replaced with a pre-trained embedding model. In this way, the large language model can obtain word vectors specifically for the coal mine domain, and these word vectors can more accurately represent the unique terms, concepts, and contexts in the coal mine domain. In the updated large language model, the new embedding layer will be initialized with the word vectors output by the trained coal mine domain embedding model. Specifically, these word vectors will be directly applied to the embedding layer of the large language model to replace the content in the original embedding layer. Using the word vectors output by the trained embedding model to initialize the embedding layer of the updated large language model determines how the updated large language model processes the input text.

[0165] In an alternative implementation, the incremental pre-training of the updated large language model using the domain corpus data and the general corpus data to obtain a large model for the coal mine vertical domain includes:

[0166] Step S1051, mixing the domain corpus data and the general corpus data according to a preset ratio to obtain training corpus data, where the general corpus data occupies a higher proportion than the domain corpus data.

[0167] When the embedding model was trained previously, how to obtain the training corpus data has been described in detail. Based on the same or similar process, it will not be elaborated here.

[0168] Step S1052, adding a low-rank matrix and a bias matrix to the weight matrix of the linear layer of the updated large language model.

[0169] In this embodiment, the LoRA (Low-Rank Adaptation) technology is used. By adding a low-rank matrix to the weight matrix of the updated large language model, the large language model is allowed to perform effective adaptive adjustments with fewer parameters. The addition of the low-rank matrix can help the large language model learn new domain knowledge with less computational and storage overhead during incremental learning. The role of the bias matrix is to further adjust the calculation result of the weight matrix and increase the flexibility of the large language model. By appropriately adjusting the weights and biases, the large language model can better adapt to the new data distribution.

[0170] Step S1053, freezing other layers of the updated large language model except the embedding layer, and training the embedding layer of the updated large language model using the training corpus data.

[0171] Freezing the other layers except the embedding layer is to avoid excessive adjustment of the weights of other layers in the large language model during the initial stage of incremental pre-training. At this time, only the embedding layer (i.e., the word vector layer) will be trained, which can ensure that while maintaining the original semantics, the newly added word embeddings focus on learning specific vocabulary and terms in the coal mine field. With the other layers frozen, only the embedding layer will be trained based on the mixed training corpus data, helping the large language model understand the professional terms in the coal mine field and ensuring that the semantic representation of the large language model is closer to the knowledge of the coal mining industry.

[0172] Step S1054, when the preset iteration condition is reached, unfreeze the other layers, and use the training corpus data to fine-tune and train all layers of the updated large language model to obtain the low-rank adaptation model corresponding to the large language model.

[0173] After the training of the embedding layer is completed, unfreeze the other layers in the large language model so that these layers can also participate in the training. By fine-tuning all layers, the large language model can gradually adapt to the special context of the coal mine field while retaining the original general knowledge. The unfreezing operation can be selected after the embedding layer has completed preliminary learning to ensure good model performance during fine-tuning. At this time, the large language model performs a complete fine-tuning training based on the mixed training corpus data. Through this process, the large language model can be further refined on the basis of its original general capabilities to adapt to the grammar, terms, and context of the coal mine field. Fine-tuning can improve the large language model's ability to handle domain-specific tasks. After this step, what is obtained is a large language model adjusted by low-rank adaptation (LoRA), which can effectively utilize the knowledge of the coal mine field while retaining the advantages of the general large language model.

[0174] Step S1055, merge the low-rank adaptation model with the original large language model to obtain the large model for the coal mine vertical domain.

[0175] Finally, merge the low-rank adaptation model after incremental pre-training and fine-tuning with the original large language model. The purpose is to integrate the model results adapted to the domain into a complete large language model so that it can handle both general tasks and tasks in the coal mine field. The merged model has cross-domain knowledge and adaptability, and can perform more professionally and precisely in the coal mine field. After the merger is completed, the large model for the coal mine vertical domain aimed at by the present invention is obtained.

[0176] Figure 2 It is a structural block diagram of a device for constructing a large model for the coal mine vertical domain provided by an embodiment of the present invention. Please refer to Figure 2 , the device includes:

[0177] An acquisition module 201, configured to acquire domain corpus data and general corpus data, and respectively construct a domain word list and a general word list based on the domain corpus data and the general corpus data;

[0178] A fusion module 202, configured to perform weighted fusion on each word segment in the general word list and the domain word list based on their respective weights to obtain a fusion word list, where the weights represent the relative frequencies of the word segments appearing in the domain corpus data and the general corpus data;

[0179] A first training module 203, configured to train an embedding model based on the fusion word list by using the domain corpus data and the general corpus data, where the embedding model is used to map word segments to corresponding word vectors based on semantic information;

[0180] An update module 204, configured to load a pre-trained original large language model, and replace the trained embedding model into the embedding layer of the original large language model to obtain an updated large language model;

[0181] A second training module 205, configured to perform incremental pre-training on the updated large language model by using the domain corpus data and the general corpus data to obtain a large model for the coal mine vertical domain.

[0182] In an alternative embodiment, the fusion module includes:

[0183] A weight calculation sub-module, configured to calculate the word frequency and weight of each word segment in the general word list and the domain word list;

[0184] A linear combination sub-module, configured to perform a linear combination of each word segment in the general word list and the domain word list based on the weights to obtain the fusion word list.

[0185] In an alternative embodiment, the weight calculation sub-module includes:

[0186] A word frequency determination unit, configured to, for each word segment in the general word list and the domain word list, determine the frequency of occurrence of the word segment in the domain corpus data as the first word frequency, or determine the frequency of occurrence of the word segment in the general corpus data as the second word frequency;

[0187] A first relative frequency determination unit, configured to, for each word segment in the general word list, calculate the relative frequency of the word segment in the general corpus data according to the first word frequency, and determine the relative frequency of the word segment in the general corpus data as the first linear term;

[0188] A second relative frequency determination unit, configured to, for each word segment in the domain vocabulary, calculate the relative frequency of the word segment in the domain corpus data according to the second word frequency, and determine the relative frequency of the word segment in the general corpus data as a second linear term;

[0189] A logarithm term determination unit, configured to, for the same word segments in the general vocabulary and the domain vocabulary, determine the logarithm of the ratio of the first word frequency and the second word frequency of the word segment as a logarithm term;

[0190] A linear combination unit, configured to respectively assign preset weight adjustment parameters to the first linear term, the second linear term, and the logarithm term, and calculate the weight of each word segment.

[0191] In an optional implementation manner, the device further includes:

[0192] An inspection module, configured to inspect whether there are unmatched word segments in the domain vocabulary, where the unmatched word segments are word segments that exist in the domain vocabulary but do not exist in the general vocabulary;

[0193] An addition module, configured to, in the case of detecting the unmatched word segments, add the word segments to the fusion vocabulary.

[0194] In an optional implementation manner, the first training module includes:

[0195] A first mixing sub-module, configured to mix the domain corpus data and the general corpus data according to a preset ratio to obtain training corpus data, where the proportion of the general corpus data is higher than that of the domain corpus data;

[0196] A word segment processing sub-module, configured to perform word segment processing on the training corpus data based on the fusion vocabulary to obtain a plurality of training word segments;

[0197] A first training sub-module, configured to train the embedding model based on the plurality of training word segments according to preset training parameters, where the training parameters include batch size, learning rate, and loss function.

[0198] In an optional implementation manner, the device further includes:

[0199] A word vector acquisition module, configured to acquire the high-dimensional word vectors output by the embedding model, and determine the core semantic information and secondary semantic information represented by the high-dimensional word vectors;

[0200] A dimensionality reduction module for performing dimensionality reduction on the high-dimensional word vectors to obtain corresponding low-dimensional word vectors, where the core semantic information represented by the high-dimensional word vectors is preserved in the head dimension of the low-dimensional word vectors, and the secondary semantic information represented by the high-dimensional word vectors is preserved in the tail dimension of the low-dimensional word vectors;

[0201] A word vector storage module for storing the low-dimensional word vectors in a data structure predefined by the embedding model.

[0202] In an alternative embodiment, the update module includes:

[0203] A replacement sub-module for replacing the embedding layer of the original large language model with the trained embedding model to obtain the embedding layer of the updated large language model;

[0204] An initialization sub-module for initializing the embedding layer of the updated large language model using the word vectors output by the trained embedding model.

[0205] In an alternative embodiment, the second training module includes:

[0206] A second mixing sub-module for mixing the domain corpus data and the general corpus data according to a preset ratio to obtain training corpus data, where the general corpus data occupies a higher proportion than the domain corpus data;

[0207] A matrix addition sub-module for adding a low-rank matrix and a bias matrix to the weight matrix of the linear layer of the updated large language model;

[0208] A freezing sub-module for freezing other layers of the updated large language model except the embedding layer and training the embedding layer of the updated large language model using the training corpus data;

[0209] A thawing sub-module for thawing the other layers when a preset iteration condition is reached and performing fine-tuning training on all layers of the updated large language model using the training corpus data to obtain the low-rank adaptation model corresponding to the large language model;

[0210] A merging sub-module for merging the low-rank adaptation model with the original large language model to obtain the large model for the coal mine vertical domain.

[0211] In an alternative embodiment, the acquisition module includes:

[0212] A sorting sub-module, configured to sort the pre-acquired coal mine vertical domain dataset into the domain corpus data, and sort the pre-acquired open-source dataset into the general corpus data. The coal mine vertical domain dataset includes equipment monitoring data, mine geological exploration report data, forum data, and industry website data;

[0213] A first processing sub-module, configured to process the general corpus data into the general vocabulary;

[0214] A second processing sub-module, configured to process the domain corpus data into the domain vocabulary that is unified with the format of the general vocabulary.

[0215] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, devices, electronic devices, and storage media. Therefore, the embodiments of the present invention can take the form of completely hardware embodiments, completely software embodiments, or embodiments combining software and hardware aspects. Moreover, the embodiments of the present invention can take the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.

[0216] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of the methods and devices according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing terminal devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks. These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks. These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, such that a series of operation steps are executed on the computer or other programmable terminal device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable terminal device provide for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1Steps of the functions specified in one or more boxes.

[0217] Although the preferred embodiments of the embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications to these embodiments once they know the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of the embodiments of the present invention.

[0218] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or terminal device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article, or terminal device. Without further limitation, an element defined by the statement "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or terminal device comprising the element.

[0219] The above has introduced in detail a method and device for constructing a large model in the vertical field of coal mines provided by the present invention. Specific examples are used in this text to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, there will be changes in the specific implementation manner and application scope according to the idea of the present invention. In summary, the content of this specification should not be construed as a limitation on the present invention.

Claims

1. A method for constructing a large model in the vertical field of coal mines, characterized in that, The method includes: Obtaining domain corpus data and general corpus data, and respectively constructing a domain vocabulary and a general vocabulary based on the domain corpus data and the general corpus data; Weightedly fusing each token in the general vocabulary and the domain vocabulary based on their respective weights to obtain a fused vocabulary, where the weights are calculated through a weighted combination of a first linear term, a second linear term, and a logarithmic term; wherein, the first linear term is the relative frequency of each token in the general vocabulary in the general corpus data, the second linear term is the relative frequency of each token in the domain vocabulary in the domain corpus data, the logarithmic term is the logarithm of the ratio of the occurrence frequency of each token in the general vocabulary and the domain vocabulary in the domain corpus data to the occurrence frequency in the general corpus data, and the weight adjustment parameter of the first linear term is the highest; Based on the fused vocabulary, training an embedding model using the domain corpus data and the general corpus data, where the embedding model is used to map tokens to corresponding word vectors based on semantic information; Loading a pre-trained original large language model, and replacing the embedding layer of the original large language model with the trained embedding model to obtain an updated large language model; Freezing other layers of the updated large language model except the embedding layer, training the embedding layer of the updated large language model using the domain corpus data and the general corpus data, and in the case of reaching a preset iteration condition, unfreezing the other layers, and fine-tuning and training all layers of the updated large language model using the domain corpus data and the general corpus data to obtain a large model for the coal mine vertical domain.

2. The method according to claim 1, wherein The weightedly fusing each token in the general vocabulary and the domain vocabulary based on their respective weights to obtain a fused vocabulary includes: Calculating the respective word frequencies and weights of each token in the general vocabulary and the domain vocabulary; Based on the weights, linearly combining each token in the general vocabulary and the domain vocabulary to obtain the fused vocabulary.

3. The method according to claim 2, wherein The calculating the respective word frequencies and weights of each token in the general vocabulary and the domain vocabulary includes: For each token in the general vocabulary and the domain vocabulary, determining the first word frequency as the occurrence frequency of the token in the domain corpus data, or determining the second word frequency as the occurrence frequency of the token in the general corpus data; For each token in the general vocabulary, calculating the relative frequency of the token in the general corpus data according to the first word frequency, and determining the relative frequency of the token in the general corpus data as the first linear term; For each token in the domain vocabulary, calculating the relative frequency of the token in the domain corpus data according to the second word frequency, and determining the relative frequency of the token in the domain corpus data as the second linear term; For the same token in the general vocabulary and the domain vocabulary, determining the logarithm of the ratio of the first word frequency and the second word frequency of the token as the logarithmic term; Allocate preset weight adjustment parameters to the first linear term, the second linear term, and the logarithmic term respectively, and calculate the weight of each token.

4. The method according to claim 1 or 2, characterized in that After obtaining the fused vocabulary, the method further includes: Check whether there are unmatched tokens in the domain vocabulary, where the unmatched tokens are tokens that exist in the domain vocabulary but do not exist in the general vocabulary; In the case where the unmatched tokens are detected, add the tokens to the fused vocabulary.

5. The method according to claim 1, wherein The training of the embedding model using the domain corpus data and the general corpus data based on the fused vocabulary includes: Mix the domain corpus data and the general corpus data according to a preset ratio to obtain training corpus data, where the general corpus data occupies a higher proportion than the domain corpus data; Based on the fused vocabulary, perform tokenization on the training corpus data to obtain a plurality of training tokens; Based on the plurality of training tokens, train the embedding model according to preset training parameters, where the training parameters include batch size, learning rate, and loss function.

6. The method according to claim 5, wherein The method further includes: Obtain the high-dimensional word vectors output by the embedding model, and determine the core semantic information and secondary semantic information represented by the high-dimensional word vectors; Perform dimensionality reduction on the high-dimensional word vectors to obtain corresponding low-dimensional word vectors, where the core semantic information represented by the high-dimensional word vectors is stored in the head dimension of the low-dimensional word vectors, and the secondary semantic information represented by the high-dimensional word vectors is stored in the tail dimension of the low-dimensional word vectors; Save the low-dimensional word vectors in a predefined data structure of the embedding model.

7. The method according to claim 1, wherein The replacement of the embedding layer of the original large language model with the trained embedding model to obtain the updated large language model includes: Replace the embedding layer of the original large language model with the trained embedding model to obtain the embedding layer of the updated large language model; Initialize the embedding layer of the updated large language model using the word vectors output by the trained embedding model.

8. The method according to claim 7, wherein The incremental pre-training of the updated large language model using the domain corpus data and the general corpus data to obtain a large model for the coal mine vertical domain includes: Mix the domain corpus data and the general corpus data according to a preset ratio to obtain training corpus data, where the general corpus data occupies a higher proportion than the domain corpus data; Add a low-rank matrix and a bias matrix to the weight matrix of the linear layer of the updated large language model; Freeze other layers of the updated large language model except the embedding layer, and train the embedding layer of the updated large language model using the training corpus data; In the case where a preset iteration condition is reached, unfreeze the other layers, and perform fine-tuning training on all layers of the updated large language model using the training corpus data to obtain a low-rank adaptation model corresponding to the large language model; Merge the low-rank adaptation model with the original large language model to obtain the large model for the coal mine vertical domain.

9. The method according to claim 1, characterized in that Obtaining the domain corpus data and the general corpus data, and respectively constructing a domain vocabulary and a general vocabulary based on the domain corpus data and the general corpus data, includes: Sorting the pre-obtained coal mine vertical domain dataset into the domain corpus data, and sorting the pre-obtained open-source dataset into the general corpus data, where the coal mine vertical domain dataset includes equipment monitoring data, mine geological exploration report data, forum data, and industry website data; Processing the general corpus data into the general vocabulary; Processing the domain corpus data into the domain vocabulary that is unified with the format of the general vocabulary.

10. An apparatus for constructing a large model in the vertical field of coal mines, characterized in that, The device includes: An acquisition module, configured to acquire domain corpus data and general corpus data, and respectively construct a domain vocabulary and a general vocabulary based on the domain corpus data and the general corpus data; A fusion module, configured to perform weighted fusion on each token in the general vocabulary and the domain vocabulary based on their respective weights to obtain a fused vocabulary, where the weights are calculated through a weighted combination of a first linear term, a second linear term, and a logarithmic term; wherein, the first linear term is the relative frequency of each token in the general vocabulary in the general corpus data, the second linear term is the relative frequency of each token in the domain vocabulary in the domain corpus data, the logarithmic term is the logarithm of the ratio of the occurrence frequency of each token in the general vocabulary and the domain vocabulary in the domain corpus data to the occurrence frequency in the general corpus data, and the weight adjustment parameter of the first linear term is the highest; A first training module, configured to train an embedding model based on the fused vocabulary using the domain corpus data and the general corpus data, where the embedding model is used to map tokens to corresponding word vectors based on semantic information; An update module, configured to load a pre-trained original large language model, and replace the embedding layer of the original large language model with the trained embedding model to obtain an updated large language model; A second training module, configured to freeze other layers of the updated large language model except the embedding layer, train the embedding layer of the updated large language model using the domain corpus data and the general corpus data, unfreeze the other layers when reaching a preset iteration condition, and perform fine-tuning training on all layers of the updated large language model using the domain corpus data and the general corpus data to obtain a coal mine vertical domain large model.

Citation Information

Patent Citations

  • Word embedding method, device and equipment for model to carry out financial field task processing

    CN116738983A

  • Training method and device of vertical field large language model and electronic equipment

    CN119047566A