Intelligent manufacturing field large language model construction method fusing field knowledge distillation

By constructing the knowledge graph and knowledge distillation technology in the field of intelligent manufacturing, the problem of insufficient application of general large language models in the field of intelligent manufacturing is solved, efficient understanding and decision-making support for professional terms and complex tasks is achieved, and the domain specificity and generalization capabilities of the model are improved.

CN120338070AInactive Publication Date: 2025-07-18SEC ZHILIAN TECH (JIANGSU) CO LTD
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510427683.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-18
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing general-purpose large language models lack the ability to understand professional terms, knowledge-related reasoning and complex task decision-making in the application of intelligent manufacturing, and it is difficult to effectively integrate structured domain knowledge.

Method used

By collecting intelligent manufacturing domain knowledge and performing structured encoding to build a knowledge graph, selecting a pre-trained large language model as the teacher model, combining knowledge graph embedding technology, using knowledge distillation training method to transfer domain knowledge to lightweight student models, and using a dual-path supervision mechanism to balance generalization and domain specificity.

Benefits of technology

It realizes accurate understanding and decision-making support for professional terms and complex process flows in the field of intelligent manufacturing, improves the domain specificity and generalization capabilities of the model, and adapts to efficient training of small sample task data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120338070A_ABST
    Figure CN120338070A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent manufacturing field big language model construction method fused with field knowledge distillation, and relates to the field of intelligent manufacturing, which collects intelligent manufacturing field knowledge and performs knowledge clearness and structured coding to obtain an intelligent manufacturing field knowledge graph, and synchronously selects a pre-trained big language model as a teacher model. And the intelligent manufacturing domain knowledge graph is combined to obtain an intelligent manufacturing domain teacher model. Then, intelligent manufacturing field task data is collected, and knowledge distillation training is performed on the student model in combination with the intelligent manufacturing field teacher model to obtain an intelligent manufacturing field student model, so that structured field knowledge is effectively integrated, and the structural characteristics and semantic relevance of the field knowledge are concerned; therefore, the semantic similarity of the soft tag and the supervision signal of the hard tag are considered, and generalization and domain specificity are balanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent manufacturing, and more particularly, in the embodiments of this application, it relates to a method for constructing a large language model in the field of intelligent manufacturing that integrates domain knowledge distillation. Background Art

[0002] In recent years, large language models pre-trained on massive general corpora (such as BERT, GPT series) have demonstrated powerful generalization capabilities in natural language processing tasks. However, their application in vertical domains (such as intelligent manufacturing) still faces significant challenges. For example, in the field of intelligent manufacturing, due to the characteristics of high specialization, knowledge intensiveness, and logical complexity in the intelligent manufacturing field, it involves structured and semi-structured knowledge such as equipment parameters, process flows, quality control standards, and industry norms, and there is a large amount of implicit expert experience. General large language models usually lack the ability to explicitly model this domain knowledge, resulting in deficiencies in professional term understanding, knowledge association reasoning, and complex task decision-making.

[0003] Therefore, an optimized solution for constructing a large language model in the field of intelligent manufacturing is desired. Summary of the Invention

[0004] To solve the above technical problems, this application is proposed. Embodiments of this application provide a method for constructing a large language model in the field of intelligent manufacturing that integrates domain knowledge distillation, which collects knowledge in the field of intelligent manufacturing and performs knowledge cleaning and structured encoding to obtain a knowledge graph of the intelligent manufacturing field. Meanwhile, a pre-trained large language model is selected as the teacher model, and a teacher model in the intelligent manufacturing field is obtained by combining it with the knowledge graph of the intelligent manufacturing field. Then, task data in the intelligent manufacturing field is collected, and the student model is trained by knowledge distillation with the teacher model in the intelligent manufacturing field to obtain a student model in the intelligent manufacturing field, so as to effectively integrate structured domain knowledge, and pay attention to the structured characteristics and semantic relevance of domain knowledge, thereby taking into account the semantic similarity of soft labels and the supervision signals of hard labels, and balancing generalization and domain specificity.

[0005] According to one aspect of this application, there is provided a method for constructing a large language model in the field of intelligent manufacturing that integrates domain knowledge distillation, which includes:

[0006] Collect knowledge in the field of intelligent manufacturing;

[0007] Perform knowledge cleaning and structured encoding on the knowledge in the field of intelligent manufacturing to obtain a knowledge graph of the intelligent manufacturing field;

[0008] Select a pre-trained large language model as the teacher model;

[0009] Use knowledge graph embedding technology to fuse the knowledge graph of the intelligent manufacturing field with the model representation space of the teacher model to obtain a teacher model in the intelligent manufacturing field;

[0010] Collect task data in the field of intelligent manufacturing;

[0011] Using the teacher model in the field of intelligent manufacturing and the task data in the field of intelligent manufacturing, perform knowledge distillation training on the student model to obtain the student model in the field of intelligent manufacturing. The student model has a relatively smaller number of parameters compared to the pre-trained large language model. Compared with the prior art, a method for constructing a large language model in the field of intelligent manufacturing that integrates domain knowledge distillation provided by the present application collects knowledge in the field of intelligent manufacturing, performs knowledge cleaning and structured encoding to obtain a knowledge graph in the field of intelligent manufacturing, synchronously selects a pre-trained large language model as the teacher model, and combines the knowledge graph in the field of intelligent manufacturing to obtain the teacher model in the field of intelligent manufacturing. Then, collect task data in the field of intelligent manufacturing, and perform knowledge distillation training on the student model in combination with the teacher model in the field of intelligent manufacturing to obtain the student model in the field of intelligent manufacturing, so as to effectively integrate structured domain knowledge, and pay attention to the structured characteristics and semantic relevance of domain knowledge, so as to take into account the semantic similarity of soft labels and the supervision signal of hard labels, and balance generalization and domain specificity. Brief Description of the Drawings

[0012] By describing the embodiments of the present application in more detail with reference to the accompanying drawings, the above and other objects, features, and advantages of the present application will become more obvious. The drawings are used to provide a further understanding of the embodiments of the present application, and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation to the present application. In the drawings, the same reference numerals generally represent the same components or steps.

[0013] Figure 1 It is a flowchart of a method for constructing a large language model in the field of intelligent manufacturing that integrates domain knowledge distillation according to an embodiment of the present application.

[0014] Figure 2 It is a schematic diagram of data flow of a method for constructing a large language model in the field of intelligent manufacturing that integrates domain knowledge distillation according to an embodiment of the present application.

[0015] Figure 3 It is a flowchart of using the teacher model in the field of intelligent manufacturing and the task data in the field of intelligent manufacturing to perform knowledge distillation training on the student model to obtain the student model in the field of intelligent manufacturing in a method for constructing a large language model in the field of intelligent manufacturing that integrates domain knowledge distillation according to an embodiment of the present application.

[0016] Figure 4 It is a flowchart of calculating the distillation loss function value between the prediction result and the soft label in a method for constructing a large language model in the field of intelligent manufacturing that integrates domain knowledge distillation according to an embodiment of the present application.

[0017] Figure 5It is a flowchart for enhancing the features of the distilled semantic difference coding vector to obtain the distilled semantic enhanced coding vector in the method for constructing a large language model in the intelligent manufacturing field by fusing domain knowledge distillation according to an embodiment of the present application.

[0018] Figure 6 It is a flowchart for calculating the distillation semantic difference feature phase reshaping gain operator of each local phase coding vector of the distillation semantic difference feature by extracting the statistical number of effective components of the local phase coding vector of the distillation semantic difference feature based on each distillation semantic difference feature in the method for constructing a large language model in the intelligent manufacturing field by fusing domain knowledge distillation according to an embodiment of the present application. Detailed implementation manners

[0019] Various exemplary embodiments, features and aspects of the present application will be described in detail below with reference to the accompanying drawings. The same reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings are not necessarily drawn to scale unless otherwise specified.

[0020] In recent years, although large language models such as BERT and GPT series have shown powerful capabilities in natural language processing tasks, they still face challenges when applied in vertical fields such as intelligent manufacturing. These fields are highly specialized, knowledge-intensive and logically complex, involving specific knowledge such as technical terms, process flows and quality control, as well as a large amount of implicit expert experience. Due to the lack of explicit modeling of such domain knowledge, general large language models perform poorly in understanding technical terms, reasoning about knowledge associations and making complex decisions.

[0021] It should be understood that the integration of domain knowledge is the fundamental issue in constructing domain large language models and the core goal of knowledge distillation technology in domain applications. Although traditional domain adaptation methods (such as fine-tuning) can alleviate some problems, they rely on a large amount of labeled data and are difficult to effectively integrate structured domain knowledge. Knowledge distillation technology provides an efficient way for domain adaptation by transferring the knowledge of the teacher model to a lightweight student model. However, existing knowledge distillation methods mostly focus on the transfer of output layer probability distributions in general tasks and ignore the structured characteristics and semantic relevance of domain knowledge. How to deeply integrate the domain knowledge graph with the representation space of the language model and achieve efficient knowledge transfer through distillation has become a key issue in constructing large language models in the field of intelligent manufacturing. At the same time, knowledge graph embedding technology provides a solution for the vector representation of domain knowledge. By aligning the entities, relationships in the knowledge graph with the semantic space of the pre-trained language model, the implicit understanding of domain concepts and their associations by the model can be enhanced. However, there is still a semantic gap in the fusion of the heterogeneous representation spaces of the knowledge graph and the language model. In addition, the task data in the field of intelligent manufacturing usually has the characteristics of small samples and long-tail distribution, requiring the distillation process to take into account both the semantic similarity of soft labels and the supervision signals of hard labels to balance generalization and domain specificity.

[0022] To address the above technical problems, the present application proposes a method for constructing a large language model in the field of intelligent manufacturing that integrates domain knowledge distillation. Figure 1 FIG. is a flowchart of a method for constructing a large language model in the field of intelligent manufacturing that integrates domain knowledge distillation according to an embodiment of the present application. Figure 2 FIG. is a schematic diagram of data flow of a method for constructing a large language model in the field of intelligent manufacturing that integrates domain knowledge distillation according to an embodiment of the present application. As Figure 1 and Figure 2 shown, the method for constructing a large language model in the field of intelligent manufacturing that integrates domain knowledge distillation according to an embodiment of the present application includes: S110, collecting domain knowledge in the field of intelligent manufacturing; S120, performing knowledge cleaning and structured encoding on the domain knowledge in the field of intelligent manufacturing to obtain a domain knowledge graph in the field of intelligent manufacturing; S130, selecting a pre-trained large language model as the teacher model; S140, integrating the domain knowledge graph in the field of intelligent manufacturing with the model representation space of the teacher model using knowledge graph embedding technology to obtain a teacher model in the field of intelligent manufacturing; S150, collecting task data in the field of intelligent manufacturing; S160, using the teacher model in the field of intelligent manufacturing and the task data in the field of intelligent manufacturing to perform knowledge distillation training on the student model to obtain a student model in the field of intelligent manufacturing, and the student model has a relatively smaller number of parameters compared to the pre-trained large language model.

[0023] In an embodiment of the present application, in step S110, knowledge in the field of intelligent manufacturing is collected, including: collecting knowledge in the field of intelligent manufacturing from multiple data sources, where the multiple data sources include domain documents, expert knowledge, knowledge bases or ontology libraries, or structured databases. It should be understood that domain documents are an important way to obtain the latest research progress and technological trends in the field of intelligent manufacturing. These documents may cover academic papers, industry reports, technical manuals, etc. They detail the theoretical research results and practical application cases in this field. By deeply analyzing these materials, the key concepts, process flows, and their interrelationships of intelligent manufacturing can be extracted. Among them, expert knowledge, as an important source, reflects the profound insights and unique understandings of professionals with years of experience in a specific field. This tacit knowledge often includes practical operation skills, troubleshooting methods, and best practice strategies that are not easily obtained through literature. In addition, knowledge bases or ontology libraries are also indispensable data sources. They usually store a large amount of domain knowledge in a structured manner, including entities (such as equipment, materials), attributes (such as dimensions, performance indicators), and relationships (such as causal relationships, association rules). By mining this structured information, a detailed knowledge network can be established, making the logical connections between various knowledge points clearer, thus helping to improve the understanding ability and reasoning level of the model. Finally, structured databases provide a large amount of real-time and historical data on the manufacturing process, such as production plans, quality control records, equipment operating status, etc. These data can not only reflect various situations in actual production but also reveal potential problems and optimization spaces. The knowledge graph constructed based on these data can better reflect the complexity and diversity of the real world and provide rich materials for the training and verification of the model. By integrating information from various channels such as domain documents, expert knowledge, knowledge bases or ontology libraries, and structured databases, not only can the coverage and depth of the constructed knowledge graph in the field of intelligent manufacturing be ensured, but also the complementarity and integration of different forms of knowledge can be promoted, thereby laying a solid foundation for the development of efficient and accurate large language models.

[0024] In the method for constructing a large language model in the field of intelligent manufacturing by fusing domain knowledge distillation, in step S120, the knowledge in the field of intelligent manufacturing is clarified and structurally encoded to obtain a knowledge graph of the field of intelligent manufacturing. It should be understood that converting scattered and diverse domain knowledge into a systematic and structured representation form facilitates subsequent processing and application. First of all, through knowledge clarification, redundant information and noise in the original data can be effectively identified and removed, ensuring that the final knowledge graph has high relevance and accuracy. In this process, advanced natural language processing techniques need to be adopted to deeply analyze text materials, including but not limited to methods such as named entity recognition, relationship extraction, and semantic parsing, so as to accurately extract key information about equipment parameters, process flows, quality control standards, etc. and convert it into a machine-readable form. Further, the goal of structural encoding is to organize the cleaned knowledge fragments into a logically rigorous and hierarchical knowledge system. This step requires precise modeling of the relationships between each knowledge point to form a complex network composed of nodes (representing entities or concepts) and edges (representing relationships or connections). For example, in the field of intelligent manufacturing, these internal connections can be captured by defining the dependencies between different manufacturing processes, the constraints between raw materials and finished product specifications, etc. This structured expression method can not only intuitively display the relationships between elements, but also support deeper knowledge reasoning and decision-making support.

[0025] In the method for constructing a large language model in the field of intelligent manufacturing by fusing domain knowledge distillation, in step S130, a pre-trained large language model is selected as the teacher model. It should be understood that through the training of a large-scale corpus, the pre-trained large language model has learned rich language structures and patterns, and this knowledge is crucial for understanding and processing text information. Among them, pre-trained large language models such as BERT and GPT, based on a deep neural network architecture, have undergone a pre-training stage with a large amount of general-purpose corpus and possess powerful natural language understanding capabilities, capable of capturing complex semantic relationships at the lexical, sentence, and even discourse levels. Specifically, these models adopt self-supervised learning methods and are trained on a large-scale dataset without manual annotation. By predicting the probability distribution of masked words or the next word, etc., they automatically learn the internal laws and feature representations of the input text. Specifically, in a specific embodiment of the present application, when selecting a pre-trained large language model as the teacher model, a comprehensive evaluation of various existing pre-trained large language models on the market needs to be carried out, including but not limited to BERT, GPT series, T5, etc. These models have received wide attention due to their excellent natural language processing capabilities. The evaluation process should be carried out based on several key factors. First is the model scale and architecture characteristics. Different pre-trained large language models have significant differences in the number of parameters, the number of layers, and the number of neurons in each layer, which directly affect the expressive ability and computational efficiency of the model. And, an ideal pre-trained corpus should cover a wide range of topics, from news reports, literary works to scientific papers, patent documents, etc., to ensure that the model can learn comprehensive language structures and usages. In addition, the time span of the corpus needs to be concerned, because as time goes by, language usage habits and social and cultural backgrounds will change, and the latest corpus helps the model better understand the current language environment. To achieve this goal, a pre-trained large language model that supports incremental learning or continuous update mechanism can be selected. They allow users to further train the model based on the newly obtained data, so as to maintain the relevance and accuracy of the model. Among them, injecting the pre-processed domain knowledge into the teacher model can enhance its domain knowledge representation ability. Common knowledge injection methods include Fine-tuning, Adapter, KnowledgeGraphEmbedding, and PromptEngineering. Among them, Fine-tuning uses a domain knowledge corpus to continuously pre-train or fine-tune the pre-trained model so that it learns domain language patterns and knowledge. Adapter inserts an Adapter module into the pre-trained model and only fine-tunes the Adapter parameters while keeping the pre-trained model parameters unchanged to achieve lightweight injection of knowledge.KnowledgeGraphEmbedding (Knowledge Graph Embedding) integrates the entity and relationship representations of a knowledge graph into the model representation space, enhancing the model's understanding and reasoning capabilities regarding the knowledge graph. PromptEngineering (Prompt Engineering), on the other hand, designs prompts containing domain knowledge to guide the teacher model to utilize domain knowledge during generation or prediction.

[0026] In the above method for constructing a large language model in the field of intelligent manufacturing that integrates domain knowledge distillation, in step S140, the knowledge graph of the intelligent manufacturing field is integrated into the model representation space of the teacher model using knowledge graph embedding technology to obtain the teacher model for the intelligent manufacturing field. It should be understood that when integrating the knowledge graph of the intelligent manufacturing field into the representation space of the pre-trained large language model through knowledge graph embedding technology, in fact, a bridge is being created to connect the general language model with the professional knowledge system within a specific domain. This process is not just a simple superposition or splicing, but a deep integration and interaction. On the one hand, the large language model provides powerful text understanding and generation capabilities. It can capture subtle meaning changes, context dependencies, and potential semantic connections in language. On the other hand, the embedded knowledge graph supplements precise descriptions of important concepts and their interactions within a specific domain. After the two are combined, the resulting teacher model for the intelligent manufacturing field can not only handle general language tasks like traditional language models but also provide more accurate and in-depth understanding and answers to questions involving professional terms, complex process flows, and industry standards. Specifically, in a specific embodiment of the present application, first, a suitable knowledge graph embedding algorithm needs to be selected, such as TransE, DistMult, or RotatE, etc. These methods learn vector representations of entities and relationships in different ways. Once the vector representations of entities and relationships are obtained, the next step is to integrate them into the representation space of the teacher model. This usually involves adjusting the architecture of the teacher model or introducing new components to facilitate handling this newly added structured information. For example, by designing a special input layer or intermediate layer, the embedding vectors of the knowledge graph can be directly input into the model as additional features, or the model can automatically learn how to balance information from text data and the knowledge graph through attention mechanisms and other means. In addition, a joint training method can also be considered, that is, using both text data and knowledge graph data to update the model parameters to achieve effective integration of the two types of information.

[0027] In the method for constructing a large language model in the field of intelligent manufacturing that integrates domain knowledge distillation, in step S150, task data in the field of intelligent manufacturing is collected. It should be understood that collecting task data in the field of intelligent manufacturing can provide sufficient samples to train the model and enable it to have the ability to understand specific tasks. For example, during fault diagnosis, the model needs to identify possible problems based on data of the operating state of the equipment and give corresponding solutions. This requires a large number of historical fault records and corresponding solutions as training samples so that the model can learn the characteristics under different fault modes and their associated solution strategies. By analyzing this data, the model can gradually build an understanding of the working principle of the equipment and common faults, thereby improving the accuracy of its prediction and diagnosis. At the same time, considering the small sample characteristics and long-tailed distribution phenomenon of data in the field of intelligent manufacturing, that is, some types of events or situations occur relatively rarely, but their occurrence may have a significant impact on the performance of the entire system, it is particularly important to collect diverse task data. This not only helps to improve the model's performance in high-frequency common scenarios but also enhances its ability to handle rare but important situations. For example, although the probability of a specific type of product defect occurring is low, once it occurs, it may cause serious economic losses. Therefore, collecting data containing such rare events can make the model more robust when facing similar challenges. Specifically, in a specific embodiment of the present application, the task data in the field of intelligent manufacturing usually comes from multiple channels, including sensor networks on the production line, enterprise resource planning (ERP) systems, manufacturing execution systems (MES), historical databases, and technical documents, etc. For the sensor network, various sensors deployed on production equipment can be used to collect real-time information on the operating state of the equipment, such as physical parameters like temperature, pressure, and vibration frequency. These data can reflect the current working condition of the equipment and provide a basis for predictive maintenance. At the same time, by integrating the data in the ERP and MES systems, information at the management level such as production plan arrangements, bill of materials, and order details can be obtained, as well as information such as detailed process flow descriptions, quality standard definitions, and past operation records. This is of great significance for understanding the overall layout of the production process and optimizing resource allocation.

[0028] Figure 3 The flowchart shows the knowledge distillation training of the student model using the intelligent manufacturing domain teacher model and the intelligent manufacturing domain task data in the method for constructing a large language model in the field of intelligent manufacturing that integrates domain knowledge distillation according to an embodiment of the present application to obtain the intelligent manufacturing domain student model. As Figure 3As shown, in the embodiment of the present application, step S160, using a teacher model in the field of intelligent manufacturing and task data in the field of intelligent manufacturing, performing knowledge distillation training on the student model to obtain a student model in the field of intelligent manufacturing, includes: S161, inputting the task data in the field of intelligent manufacturing into the teacher model in the field of intelligent manufacturing to obtain soft labels; S162, inputting the task data in the field of intelligent manufacturing into the student model to obtain prediction results; S163, calculating the value of the hard label loss function between the prediction result and the true label; S164, calculating the value of the distillation loss function between the prediction result and the soft label; S165, using the weighted sum of the distillation loss function value and the hard label loss function value as the loss function value, training the student model to obtain a student model in the field of intelligent manufacturing.

[0029] That is, in the technical solution of this application, the efficient transfer of domain knowledge from the enhanced teacher model to the lightweight student model is achieved through the knowledge distillation framework, while solving the contradiction between the structural sparsity of knowledge in the intelligent manufacturing field and the characteristics of small samples of task data. In terms of the technical path, first, the domain knowledge graph is deeply integrated with the semantic space of the general pre-trained large language model through the knowledge graph embedding technology to construct a teacher model with domain cognitive ability. In this process, structured knowledge such as equipment parameters and process specifications is mapped into vectors aligned with the representation space of the language model, forming a domain teacher model with multi-granularity semantic enhancement. In the distillation stage, a dual-path supervision mechanism is adopted: on the one hand, the cross-entropy loss of the true label is used to ensure the basic performance of the task, and on the other hand, a distillation loss function based on semantic differences is designed to focus on capturing the correlation differences between the soft labels output by the teacher model and the prediction results of the student model in the deep semantic space. Specifically, the soft label and the prediction result are mapped to a shared vector space through semantic embedding encoding, the position-differential vector is extracted to reveal the fine-grained semantic deviation, and then the key difference features are strengthened by the feature enhancement module, and finally decoded into a distillation loss value. This mechanism enables the lightweight student model to not only inherit the domain knowledge reasoning ability of the teacher model but also maintain sensitivity to domain terms and logical relationships under limited labeled data, achieving a balance between knowledge transfer efficiency and model compression effect. Specifically, in a specific embodiment of this application, first, a language model with a simpler structure or fewer parameters is selected as the student model. The structure of the student model can be the same as or different from that of the teacher model, but usually a lighter model is selected. Then, the enhanced teacher model is used to process the domain task data to generate soft labels (i.e., the probability distribution output by the teacher model) or extract the intermediate layer features of the teacher model. The soft label contains richer knowledge information of the teacher model, such as the similarity between categories and the confidence of prediction. Next, the domain task data and the soft label provided by the teacher model are used to train the student model. Finally, a suitable loss function is designed, usually combining two losses, namely HardLabelLoss and DistilationLoss. Among them, HardLabelLoss is the loss between the prediction result of the student model and the true label (such as cross-entropy loss). DistilationLoss is the loss between the prediction result of the student model and the soft label of the teacher model (such as KL divergence, MSE loss), and the goal is to make the output distribution of the student model as close as possible to the output classification of the teacher model.

[0030] Specifically, in step S161, the task data in the field of intelligent manufacturing is input into the teacher model in the field of intelligent manufacturing to obtain soft labels. It should be understood that soft labels are different from hard labels (true labels). They represent the probability distribution prediction of the teacher model for the input data, rather than a single definite answer. This probability distribution not only contains the most likely answer but also provides information about other possibilities, thus retaining more semantic information and uncertainty. When the task data in the field of intelligent manufacturing is input into the teacher model, based on the knowledge graph embedding it has learned and the language understanding and generation capabilities obtained during the pre-training process, the model can conduct in-depth understanding and analysis of each input instance. For example, when dealing with the task of equipment fault diagnosis, the teacher model can utilize its knowledge of historical fault records, equipment parameters, and process flows to infer potential problems and their solutions under the current equipment state. In this way, the teacher model not only identifies specific types of fault patterns but also attempts to understand the operating mechanism of the entire system and gives a soft label containing multiple possible fault causes and their corresponding probabilities based on these understandings. Such an output form is richer and more detailed compared to only providing the most likely fault type (hard label) because it takes into account various possible situations and their probabilities, helping to more comprehensively describe the problem space. Specifically, to enable the teacher model to effectively process the input data, it is necessary to consider how to map this data into the knowledge representation space that the teacher model has learned. This means making full use of the embedding layer or encoder inside the teacher model to convert the input task data into a high-dimensional vector form to capture its deep semantic information. In this process, considering that the task data in the field of intelligent manufacturing is often highly professional and complex, specially designed feature engineering methods may be required to enhance the expressiveness of the data. For example, domain-specific word embeddings can be introduced or relationship embedding techniques based on knowledge graphs can be used, so that each input instance not only contains surface information but also implies the complex logical relationships and context behind it. This not only improves the understanding accuracy of the teacher model but also helps to generate more accurate soft labels. Once the data preprocessing and feature mapping are completed, the prepared task data in the field of intelligent manufacturing can be formally input into the teacher model. At this stage, the teacher model will conduct a comprehensive and in-depth analysis of each input instance according to its trained parameters and architecture and output a probability distribution as a soft label. This process involves calculations at multiple levels. First, the input data is passed layer by layer to different levels of the teacher model through the forward propagation algorithm. Each layer will transform and abstract the information received at that layer, gradually refining higher-level feature representations.For deep neural network architectures like Transformer, this feature extraction process particularly relies on the self-attention mechanism, which allows the model to dynamically adjust the weight distribution between different parts, thereby capturing more precisely the key information and its interrelationships in the input data.

[0031] Specifically, in step S162, the task data in the field of intelligent manufacturing is input into the student model to obtain a prediction result. It should be understood that the student model aims to inherit the knowledge and capabilities of the teacher model, but usually has fewer parameters to facilitate efficient operation in resource-constrained environments. Therefore, by inputting real task data in the field of intelligent manufacturing into the student model and observing the output prediction result, it is possible to intuitively verify whether the student model has successfully absorbed the essence of the knowledge transmitted by the teacher model. Further, by comparing the prediction result of the student model with the soft labels provided by the teacher model, it is possible to effectively identify the deficiencies of the student model in which aspects, and thus make targeted improvements. As the deep understanding and analysis result of the input data by the teacher model, the soft labels not only reflect the most likely answers, but also contain the evaluation of other possibilities and their relative importance. Therefore, when there are significant differences between the prediction result of the student model and the soft labels, this often means that the student model may have deviations in the understanding of some key concepts, or its learning strategy fails to fully capture the core knowledge of the teacher model. By systematically analyzing these differences, it is possible to more specifically adjust the architecture design, hyperparameter settings, and training algorithms of the student model in order to narrow the gap between the two. In the embodiment of the present application, in step S163, the hard label loss function value between the prediction result and the true label is calculated, including: calculating the cross-entropy loss function value between the prediction result and the true label as the hard label loss function value. It should be understood that the hard label refers to the exact classification or numerical answer corresponding to each input instance in the dataset, which represents the most ideal prediction result. However, in practical applications, due to factors such as data noise, uneven sample distribution, and model itself limitations, the prediction results generated by the model often deviate from these ideal targets to a certain extent. In order to narrow this gap and ensure that the model can accurately capture the potential patterns in the data, a measure is needed to evaluate the quality of the prediction results, and the hard label loss function is an effective evaluation index designed for this purpose. Specifically, the role of the hard label loss function is to provide a quantifiable error metric, enabling the model to automatically adjust its internal parameters during the training process according to this metric to minimize the gap between the prediction result and the true label. The most common forms of the hard label loss function include cross-entropy loss (for classification tasks) and mean squared error loss (for regression tasks). Further, the process of calculating the hard label loss function value actually establishes a feedback mechanism, enabling the model to continuously correct its own deviations in each iteration of training. In the field of deep learning, especially for complex architectures like Transformer, the model usually contains a huge number of parameters, and it is obviously infeasible to directly manually adjust these parameters. Therefore, an optimization strategy based on the gradient descent algorithm is needed, and the hard label loss function provides the necessary direction guidance for this process.At this time, by comparing the production plan predicted by the model with the actual situation (i.e., the true label), the corresponding hard label loss value can be calculated. This loss value is then backpropagated to each layer of the network, using the chain rule to calculate the gradients of the parameters of each layer with respect to the loss function, and updating the parameter values accordingly, so that the prediction result of the model in the next round of training is closer to the actual situation. This is iterated repeatedly until the model converges or reaches the preset stop condition.

[0032] Figure 4 A flowchart for calculating the distillation loss function value between the prediction result and the soft label in the method for constructing a large language model in the field of intelligent manufacturing that integrates domain knowledge distillation according to an embodiment of the present application. As Figure 4 shown, in the embodiment of the present application, step S164, calculating the distillation loss function value between the prediction result and the soft label, includes: S1641, performing semantic embedding encoding on the prediction result to obtain a prediction result semantic embedding encoding vector; S1642, performing semantic embedding encoding on the soft label to obtain a soft label semantic embedding encoding vector; S1643, calculating the position-wise difference vector between the prediction result semantic embedding encoding vector and the soft label semantic embedding encoding vector to obtain a distillation semantic difference encoding vector; S1644, performing feature enhancement on the distillation semantic difference encoding vector to obtain a distillation semantic enhancement encoding vector; S1645, performing feature decoding on the distillation semantic enhancement encoding vector to obtain the distillation loss function value.

[0033] Specifically, steps S1641 and S1642, semantic embedding encoding is performed on the prediction results to obtain the semantic embedding encoding vector of the prediction results, and semantic embedding encoding is performed on the soft labels to obtain the semantic embedding encoding vector of the soft labels. It should be understood that in the technical solution of the present application, in the process of calculating the distillation loss function value between the prediction results and the soft labels, first, in the process of distillation of knowledge in the intelligent manufacturing field, there are limitations in directly comparing the output layer probability distribution of the student model prediction results and the teacher model soft labels. Specifically, since the implicit association of domain knowledge (such as the logical constraints of equipment parameters and process flows, and the mapping relationship between industry specifications and quality control standards) is often implicit in the semantic space, it is difficult to fully convey the reasoning logic of structured knowledge by relying solely on probabilistic similarity. For example, the soft labels output by the general model may not be able to distinguish the fine-grained semantic differences between professional terms such as "processing accuracy 0.01mm" and "tolerance range ±0.02mm", and these differences are crucial to intelligent manufacturing decision-making tasks. Therefore, in the technical solution of the present application, the prediction results are semantically embedded to obtain the prediction result semantic embedding coding vector, and the soft labels are semantically embedded to obtain the soft label semantic embedding coding vector. By processing the prediction results and soft labels in a semantic embedding coding manner to map them to a shared semantic vector space, it is possible to decouple the surface probability distribution and the deep semantic association, thereby more accurately quantifying the deviations between the two in the understanding of domain knowledge. Through the vector alignment of the semantic space, the potential differences between the teacher model and the student model in the representation of domain concepts are revealed. Specifically, semantic embedding coding can convert the domain knowledge implicit in the soft labels (such as the nonlinear association between process parameters, the causal relationship between fault codes and equipment status) into a high-dimensional vector representation, so that the student model can not only imitate the output behavior of the teacher model, but also learn its internal semantic reasoning mode. For example, in the task of equipment fault diagnosis, the soft label may contain complex temporal dependencies between the fault code and the sensor data, and the semantic embedding coding can explicitly map such relationships into geometric constraints in the vector space, thereby enhancing the student model's modeling ability for implicit domain logic.

[0034] Specifically, step S1643, calculate the position difference vector between the semantic embedding coding vector of the prediction result and the semantic embedding coding vector of the soft label to obtain the distilled semantic difference coding vector. It should be understood that since domain knowledge (such as the constraint relationship between equipment parameters and process specifications, the mapping logic between industry standards and quality control indicators) is usually embedded in the semantic space in a multi-dimensional and hierarchical manner, relying only on the global similarity measurement will ignore the differences in key local features. Therefore, the position difference vector between the semantic embedding coding vector of the prediction result and the semantic embedding coding vector of the soft label is further calculated to obtain the distilled semantic difference coding vector. The calculation of the position difference vector can compare the difference between the semantic embedding coding vector of the prediction result and the semantic embedding coding vector of the soft label dimension by dimension, so as to accurately locate the deviation between the two in the representation of specific domain concepts, which can identify the semantic offset between the prediction result and the soft label in key domain concepts (such as industry abbreviations, equipment models), and avoid the semantic ambiguity problem caused by the sparsity of domain terms in general distillation. In particular, through the difference analysis of fine-grained semantics, the learning of the domain knowledge structure by the student model can be strengthened. Specifically, the position difference operation can decouple the complex domain relationships implied in the soft labels (such as the temporal dependence of fault codes and sensor data, and the synergy between process parameters) into the difference components of each dimension in the vector space. This dimension-by-dimension difference analysis can reveal the weak links of the student model in domain knowledge reasoning and provide targeted signals for subsequent feature enhancement.

[0035] Figure 5 This is a flowchart of performing feature enhancement on the distilled semantic difference coding vector to obtain the distilled semantic enhancement coding vector in the method for constructing a large language model in the intelligent manufacturing field by integrating domain knowledge distillation according to an embodiment of the present application. Figure 5As shown, in the embodiment of the present application, step S1644, feature enhancement is performed on the distilled semantic difference encoding vector to obtain a distilled semantic enhanced encoding vector, including: S1644-1, feature phase reconstruction and information squeezing are performed on the distilled semantic difference encoding vector to obtain a set of distilled semantic difference feature extraction local phase encoding vectors; S1644-2, calculate the distilled semantic difference feature effective component statistics of each distilled semantic difference feature extraction local phase encoding vector in the set of distilled semantic difference feature extraction local phase encoding vectors; S1644-3, based on the distilled semantic difference feature effective component statistics of each distilled semantic difference feature extraction local phase encoding vector, calculate the distilled semantic difference feature phase reshaping gain operator of each distilled semantic difference feature local phase encoding vector; S1644-4, based on the distilled semantic difference feature phase reshaping gain operator of each distilled semantic difference feature local phase encoding vector, perform feature phase significance reshaping on the set of distilled semantic difference feature local phase encoding vectors to obtain a distilled semantic enhanced encoding vector. It should be understood that due to the structured characteristics of domain knowledge (such as the non-linear coupling between process parameters and the multi-dimensional association between equipment status and fault codes), the distilled semantic difference encoding vector contains both deviation signals of key domain logics and may also be mixed with irrelevant fluctuations of general semantics. For example, when predicting equipment maintenance strategies, the differences between the teacher model and the student model in the correlation dimension of "bearing wear rate" and "lubrication cycle" may be masked by other irrelevant features (such as the context generalization deviation of general terms), and directly using the distilled semantic difference encoding vector for loss calculation may face challenges of information redundancy and noise interference. Therefore, in the technical solution of the present application, further feature enhancement is performed on the distilled semantic difference encoding vector to obtain a distilled semantic enhanced encoding vector. The feature enhancement method here is through local phase reconstruction and information squeezing, so as to decouple the effective components and redundant noises in the distilled semantic difference encoding vector, thereby focusing on the core contradiction of domain knowledge transfer.

[0036] Specifically, the feature enhancement method here captures the local structure of the distilled semantic difference coding vector through one-dimensional convolutional encoding (such as the co-variation trend between the tolerance range and machining accuracy in process specifications, and the temporal dependence between sensor data and abnormal codes in fault diagnosis). Then, it eliminates dimensional redundancy in the information squeezing stage (such as suppressing the smooth transition features of general semantics) and retains the difference components strongly related to the domain task (such as the calculation logic deviation of dimensional tolerance under the "ISO2768-mK" standard). For example, in a quality inspection task, the teacher model may establish a strong association between the "surface roughness Ra value" and "grinding parameters" in a specific dimension based on knowledge graph embedding. If the student model fails to fully learn this relationship, the difference signal in the corresponding dimension needs to be amplified through the reshaping gain operator of the distilled semantic difference feature phase to achieve targeted knowledge transfer. In terms of execution effect, this module significantly improves the transfer efficiency of domain-structured knowledge in the distillation process. On the one hand, phase reconstruction and information squeezing can identify the hidden domain logic faults in the distilled semantic difference coding vector (such as the process jump rule deviation in the process flow diagram), avoiding detail loss caused by global feature smoothing. On the other hand, the calculation of the reshaping gain operator of the distilled semantic difference feature can adaptively enhance the key phase differences (such as the dynamic balance relationship between the equipment OEE (Overall Equipment Effectiveness) and production rhythm), while maintaining the translational invariance of the feature space. This enables the student model to accurately inherit the reasoning ability of the teacher model for implicit knowledge (such as the abnormal threshold setting rule in expert experience) under small-sample data. For example, in a production capacity optimization task, it can accurately model the non-linear relationship between equipment utilization rate and energy consumption curve, or maintain the causal constraint between material heat treatment parameters and mechanical property indicators in process planning, ultimately achieving a double breakthrough in lightweight deployment and domain task accuracy.

[0037] In an embodiment of the present application, in step S1644-1, feature phase reconstruction and information squeezing are performed on the distilled semantic difference coding vector to obtain a set of distilled semantic difference feature extraction local phase coding vectors, including: S1644-11, performing feature phase reconstruction based on one-dimensional convolutional encoding on the distilled semantic difference coding vector to obtain a set of distilled semantic difference feature local phase coding vectors; S1644-12, performing information squeezing on each distilled semantic difference feature local phase coding vector in the set of distilled semantic difference feature local phase coding vectors to obtain a set of distilled semantic difference feature extraction local phase coding vectors.

[0038] Specifically, in step S1644-11, performing feature phase reconstruction based on one-dimensional convolutional encoding on the distilled semantic difference coding vector to obtain a set of distilled semantic difference feature local phase coding vectors, which is represented by the distilled semantic feature phase reconstruction formula as:

[0039] Conv l×1(X) = {x1, x2,..., x i ,..., x n}

[0040] where X is the distilled semantic difference coding vector, Conv l×1 is one-dimensional convolutional coding processing, l is the feature phase reconstruction step length, and x1, x2, x i , x n are the 1st, 2nd, i-th, and n-th distilled semantic difference feature local phase coding vectors in the set of distilled semantic difference feature local phase coding vectors, respectively. It should be understood that the distilled semantic difference coding vector carries the semantic deviation information between the prediction result of the student model and the soft label of the teacher model. Due to the structured characteristics of domain knowledge (such as the constraint relationship between device parameters and process flows, the mapping logic between industry standards and quality indicators), there are multi-dimensional local association patterns within the distilled semantic difference coding vector. Directly adopting global feature matching is difficult to effectively capture the relative relationships between adjacent feature dimensions (such as the co-variation trend of process parameters, the temporal dependence between fault codes and sensor data), and these local patterns often imply the core reasoning logic of domain knowledge. Therefore, it is necessary to analyze the internal structure of the distilled semantic difference coding vector through feature phase reconstruction technology to provide a fine-grained analysis basis for subsequent difference feature enhancement. Through the one-dimensional convolutional coding mechanism, multi-scale local structure features can be extracted from the distilled semantic difference coding vector to achieve refined modeling of domain knowledge differences. Specifically, using the sliding window feature of the convolutional kernel, the relative phase relationships between adjacent dimensions of the distilled semantic difference coding vector (such as feature mutation edges, smooth change trends) are scanned at multiple levels, and the implicit domain logic structure (such as the tolerance range calculation rule in process specifications, the association pattern between device status and exception codes) is mapped to the set of distilled semantic difference feature local phase coding vectors. Through the collaborative action of multiple groups of convolutional kernels (different sizes or weights), local feature expressions covering different receptive fields are constructed to form a multi-angle feature space that can represent the structured differences of domain knowledge, providing an intermediate representation rich in semantic associations for subsequent information squeezing and phase reshaping. In this way, the output set of distilled semantic difference feature local phase coding vectors not only retains the spatial distribution characteristics of the original difference vector, but also highlights the fault information of key domain logics (such as the deviation of process jump rules in process flow diagrams) through the local structure enhancement mechanism, enabling the subsequent feature enhancement module to specifically strengthen the semantic difference components strongly related to intelligent manufacturing tasks, and ultimately improving the transfer efficiency of domain structured knowledge in the knowledge distillation process. Specifically, in step S1644-12, information squeezing is performed on each distilled semantic difference feature local phase coding vector in the set of distilled semantic difference feature local phase coding vectors to obtain the set of distilled semantic difference feature squeezed local phase coding vectors, which is expressed by the distilled semantic information squeezing formula as:

[0041]

[0042] where, ‖·‖ is the one-norm of a vector, and v i is the i-th distilled semantic difference feature extraction local phase encoding vector in the set of distilled semantic difference feature extraction local phase encoding vectors. It should be understood that after the feature phase reconstruction based on one-dimensional convolution, although the set of distilled semantic difference feature local phase encoding vectors contains the structured difference information of domain knowledge (such as the co-variation pattern of process parameters, the temporal correlation between equipment status and fault codes), due to the multi-scale characteristics of the convolution operation and the high-dimensional complexity of the feature space, there may be redundant dimensions (such as different convolution kernel responses representing the same process constraint repeatedly) and information mixing (such as the coupling of noise signals and key domain logics) inside. This redundancy and mixing will lead to an increase in the computational load of the subsequent feature enhancement module, and at the same time dilute the significance of key domain difference signals (such as tolerance calculation rule deviations, implicit associations between material properties and processing parameters). Therefore, it is necessary to perform feature rectification on the distilled semantic difference feature local phase encoding vectors through information squeezing technology to eliminate redundant interference and strengthen the representation ability of core difference features. Through the non-linear information screening mechanism, key difference components highly relevant to intelligent manufacturing tasks can be extracted from multi-scale local phase features (such as the non-linear coupling feature between machining accuracy and equipment vibration frequency, the mapping deviation between industry standard terms and quality control indicators). Specifically, by designing a dynamic weight adjustment strategy based on feature norms (such as using the normalized square function to suppress low-energy feature dimensions), the information density of the distilled semantic difference feature local phase encoding vectors is redistributed, retaining the core dimensions representing the structured differences of domain knowledge (such as the correlation feature between bearing wear rate and lubrication period), while weakening general semantic noise (such as irrelevant context generalization deviations) and redundant components (such as redundant expressions of repeated process constraints). This process compresses the feature space from the original high-dimensional representation to a low-dimensional dense space, realizing the explicit decoupling and semantic focusing of domain difference features. After performing the information squeezing operation, the obtained distilled semantic difference feature extraction local phase encoding vectors exhibit stronger domain specificity and information purity. Specifically, through a non-linear squeezing function (such as norm-based feature energy recalibration), selective enhancement and suppression of local phase encoding are performed, effectively separating the key phase fluctuations (such as the phase synchronization deviation between production rhythm and energy consumption curve) and noise interference (such as random fluctuations introduced by sensor acquisition errors) in the dynamic balance relationship of equipment OEE (Overall Equipment Effectiveness). This enables the subsequent feature reshaping module to focus on strengthening the difference patterns highly relevant to domain tasks (such as process jump rule conflicts in process flow diagrams), avoiding the problem of gradient dispersion caused by redundant dimensions.

[0043] Specifically, in step S1644-2, calculate the count of the effective components of the distilled semantic difference features for each local phase encoding vector in the set of distilled semantic difference feature extraction local phase encoding vectors, which is represented by the formula for calculating the count of the effective components of the distilled semantic difference features:

[0044]

[0045] Where is the feature value at the j-th position in the i-th local phase encoding vector for distilled semantic difference feature extraction, count i represents the effective component count, ε is a trainable preset threshold, en i is the count of the effective components of the distilled semantic difference features corresponding to v i . It should be understood that by establishing a feature effectiveness evaluation system, the effective components strongly related to the intelligent manufacturing tasks (such as the co-variation pattern of process parameters, the mapping deviation of industry standard clauses) in the local phase encoding vectors for distilled semantic difference feature extraction can be identified through a statistical threshold determination mechanism (such as ε threshold screening based on the difference values of adjacent feature dimensions). Specifically, by calculating the number of dimensions that meet the difference conditions within the local phase encoding vectors for distilled semantic difference feature extraction, the count of the effective components of the distilled semantic difference features can objectively reflect the significance of the local phase encoding in the structured difference representation of domain knowledge (such as the number of mutation features in the correlation between bearing wear rate and lubrication period). This process transforms the abstract semantic differences of the local phase encoding vectors for distilled semantic difference feature extraction into measurable spatial distribution indicators, ensuring the interpretability and domain adaptability of the subsequent calculation of the phase reshaping gain operator. After calculating the count of the effective components of the distilled semantic difference features, high-value difference features (such as the non-linear coupling anomaly points between material hardness and cutting speed) can be accurately located, and low-contribution components (such as the smooth features generated by the generalization deviation of general terms) can be suppressed. Figure 6 is a flowchart for calculating the phase reshaping gain operator of the distilled semantic difference features for each local phase encoding vector of the distilled semantic difference features based on the count of the effective components of the distilled semantic difference features in the method for constructing a large language model in the intelligent manufacturing field that fuses domain knowledge distillation according to an embodiment of the present application. As Figure 6As shown, in the embodiment of the present application, in step S1644-3, based on the statistical number of effective components of the distilled semantic difference features of the local phase encoding vectors extracted from each distilled semantic difference feature, calculate the distilled semantic difference feature phase reshaping gain operator for each local phase encoding vector of the distilled semantic difference feature, including: S1644-31, based on the statistical number of effective components of the distilled semantic difference features of the local phase encoding vectors extracted from each distilled semantic difference feature, determine the suppression factor corresponding to each local phase encoding vector of the distilled semantic difference feature; S1644-32, based on the suppression factor corresponding to each local phase encoding vector of the distilled semantic difference feature, calculate the distilled semantic difference feature phase reshaping gain operator for each local phase encoding vector of the distilled semantic difference feature.

[0046] In the embodiment of the present application, in step S1644-32, based on the suppression factor corresponding to each local phase encoding vector of the distilled semantic difference feature, calculate the distilled semantic difference feature phase reshaping gain operator for each local phase encoding vector of the distilled semantic difference feature, including: S1644-321, based on the suppression factor corresponding to each local phase encoding vector of the distilled semantic difference feature, calculate the initial distilled semantic difference feature phase reshaping gain operator for each local phase encoding vector of the distilled semantic difference feature; S1644-322, perform feature phase spread missing correction on the initial distilled semantic difference feature phase reshaping gain operator to obtain the distilled semantic difference feature phase reshaping gain operator.

[0047] Specifically, step S1644-3 is represented by the formula for the distilled semantic difference feature phase reshaping gain operator as:

[0048]

[0049] where n is the number of vectors in the set of local phase encoding vectors of the distilled semantic difference feature, θ i is the polar angle corresponding to v i λ i is the suppression factor corresponding to v i π represents the circumference ratio, arctan represents the arctangent function, e i is the initial distilled semantic difference feature phase reshaping gain operator, e(v i ) is the initial distilled semantic difference feature phase reshaping gain operator corresponding to v i K i is the holomorphic flatness factor of the distilled semantic difference space, T i is the metric representation factor of the flatness decomposition of the distilled semantic difference, e i' is the distillation semantic difference feature phase reshaping gain operator. It should be understood that after calculating the number of effective components of the distillation semantic difference features, although the local phase encoding vectors extracted from each distillation semantic difference feature have quantified their information contribution degrees (such as the dimensional mutation quantity of process parameter collaboration deviation, the feature significance of fault diagnosis rule conflicts), static statistical indicators cannot directly adapt to the structured characteristics of knowledge in the field of intelligent manufacturing (such as the multi-dimensional dynamic coupling relationship between equipment status and process specifications). Since domain knowledge transfer needs to achieve high-fidelity semantic alignment in the feature space (such as the causal constraint mapping between material properties and processing parameters), it is necessary to transform the number of effective components of the distillation semantic difference features into a dynamic weight regulation signal through the non-linear transformation mechanism of the gain operator to solve the spatial geometric mismatch problem between the feature phase distribution and domain logic (such as the lack of flatness constraint under the holomorphic structure). Constructing a gain calculation model based on domain knowledge guidance can realize the adaptive reshaping of the feature phase space through the synergistic effect of the suppression factor and the gain operator.

[0050] And, based on the number of effective components en of the distillation semantic difference features i Calculate the suppression factor λ i , and quantify the redundancy degree of the local phase encoding vector v extracted from the distillation semantic difference features in the feature space (such as the inefficient representation of repeated process rules), that is, regard each local phase encoding vector v extracted from the distillation semantic difference features i as a set v of local phase encoding vectors extracted from the distillation semantic difference features i (i = 1~n) is based on the generator of the effective components. Then, the generator is used as the subspace generation vector pointing to the set space. Let θ i = en i / ∑ i en i=1~n en i , and it is also expected that the polar angle representation θ i satisfies the direction symmetry, so that the set space maintains the translational invariance of the effective components.

[0051] Furthermore, combined with the holomorphic structure metric decomposition, that is:

[0052]

[0053] Then, enhance the effectiveness of the local phase encoding as the flatness decomposition metric representation based on the holomorphic structure to construct the single-mode coupling representation as a gauge field, that is:

[0054]

[0055] Next, through the distillation semantic difference feature phase reshaping gain operator e iThe enhanced validity representation shows that its separate mode is the highest-weight state excited by flatness in the spatial holomorphic structure. Thus, the distillation semantic difference feature phase reshaping gain operator e can be updated i , namely:

[0056]

[0057] In this way, it is possible to adaptively strengthen the phase difference patterns strongly related to intelligent manufacturing tasks (such as the non-linear correlation characteristics between bearing wear rate and lubrication period), while suppressing noise and redundant components (such as the random fluctuations introduced by sensor acquisition errors). For example, in the tolerance calculation scenario, the gain operator dynamically amplifies the phase correlation dimension between "dimensional tolerance under ISO2768-mK standard" and "processing accuracy deviation" through flatness decomposition metric, and eliminates the gradient deviation caused by the missing feature dispersion through the correction term, enabling the student model to accurately capture the structural differences between process specifications and measured data (such as the phase shift between the tolerance zone boundary and the actual size). This mechanism ensures the efficient transfer of domain logic during the knowledge distillation process (such as multi-sensor data fusion rules in fault diagnosis), and ultimately realizes the robust inference ability of the lightweight model for implicit expert experience (such as abnormal threshold setting strategies).

[0058] Specifically, in step S1644-4, based on the distillation semantic difference feature phase reshaping gain operator of each distillation semantic difference feature local phase encoding vector, the set of distillation semantic difference feature local phase encoding vectors is subjected to feature phase significance reshaping to obtain distillation semantic enhanced encoding vectors, which is expressed by the feature phase significance reshaping formula as:

[0059]

[0060] where exp is the natural exponential function value with base e, a i is the distillation semantic difference feature phase reshaping gain weight, v cIt is a distilled semantic enhanced encoding vector. It should be understood that by dynamically allocating weights to local phase encoding through the distilled semantic difference feature phase reshaping gain operator, it is possible to achieve the directional enhancement of domain-specific difference features, aligning the geometric distribution of the distilled semantic difference feature local phase encoding vector with the structured logic of domain knowledge. Specifically, based on the weight regulation of each distilled semantic difference feature local phase encoding vector by the distilled semantic difference feature phase reshaping gain operator (such as the dynamic scaling factor based on the effective ingredient statistics), the spatial topological structure of the reconstructed feature vector is highlighted, emphasizing the phase difference patterns strongly related to intelligent manufacturing tasks (such as the correlation features between machining accuracy mutations and equipment vibrations), while suppressing redundant dimensions (such as the inefficient representation of repeated process constraints) and noise components (such as the random fluctuations introduced by sensor acquisition errors). In this way, a distilled semantic enhanced encoding vector with domain knowledge orientation can be formed, enabling the feature space to explicitly encode the implicit mapping rules between equipment states and process specifications (such as the non-linear relationship between bearing wear rate and lubrication cycle). After performing feature phase saliency reshaping, the distilled semantic enhanced encoding vector exhibits a high-resolution representation ability for domain-structured differences. For example, in the process optimization scenario, the reshaped feature vector can accurately capture the collaborative deviation of "cutting speed" and "tool life" in a specific phase dimension, amplifying the key dimension difference signal (such as the processing parameter adjustment requirements caused by material hardness mutations) through the distilled semantic difference feature phase reshaping gain operator and weakening the interference of irrelevant features (such as the noise impact of environmental temperature fluctuations). This operation enables the student model to focus on the core domain logic (such as the dimensional tolerance calculation rules under ISO standards and the fault inference path of multi-sensor data fusion) during the knowledge distillation process, significantly improving its transfer efficiency and inference robustness for implicit knowledge (such as the abnormal threshold setting strategy in expert experience), and ultimately achieving the high-precision decision-making ability of the lightweight model in complex intelligent manufacturing tasks.

[0061] In an embodiment of the present application, in step S1645, decoding the features of the distilled semantic enhanced encoding vector to obtain a distilled loss function value includes: passing the distilled semantic enhanced encoding vector through a loss value calculator based on a decoder to obtain the distilled loss function value. It should be understood that since feature enhancement has performed high-dimensional abstraction on the distilled semantics through phase reconstruction and information squeezing (such as the geometric constraints of non-linear coupling between process parameters and the implicit association between fault modes and sensor data), there is no directly differentiable mapping relationship between its encoding form and the task objective. Therefore, further, the features of the distilled semantic enhanced encoding vector are decoded to obtain the distilled loss function value. The core role of feature decoding is to establish a bridge between the distilled semantic enhanced space and the loss function space, converting the structured differences of domain knowledge into an optimizable objective function to obtain the distilled loss function value.

[0062] Specifically, in step S165, the weighted sum of the distillation loss function value and the hard-label loss function value is used as the loss function value to train the student model to obtain a student model in the field of intelligent manufacturing. It should be understood that in complex task scenarios in the field of intelligent manufacturing, such as fault diagnosis and process optimization, relying solely on hard labels may not be able to comprehensively reflect the deep patterns and potential correlations contained in the data. For example, in the equipment fault classification task, although the true label can clearly indicate the specific fault type of a certain device, it cannot convey the potential possibilities and relative importance of other related fault types. The soft labels provided by the teacher model, on the other hand, contain the predicted probabilities of all possible fault types, which not only helps the student model to understand the input data more comprehensively but also provides more dimensional information support for the subsequent decision-making process. Although soft labels can provide rich semantic information, after all, it is an indirect form of supervision, and its accuracy depends on the performance of the teacher model. Therefore, in practical applications, simply using the distillation loss function to guide the training of the student model may have certain limitations, especially when the teacher model is not absolutely perfect or there are certain biases. At this time, it is particularly important to combine the hard-label loss function to form a comprehensive loss function. The hard-label loss function directly reflects the difference between the model output and the actual target, providing a clear and definite learning goal for the student model and helping the model establish a basic understanding of basic concepts and logical relationships. This direct supervision method based on hard labels can effectively prevent the student model from deviating from the correct learning path and ensure its sufficient accuracy and reliability when dealing with key tasks. Adding the distillation loss function value and the hard-label loss function value with appropriate weights to form the final loss function value is essentially to balance the relationship between two different types of supervision signals. Specifically, in a specific embodiment of the present application, the method of combining the distillation loss function value and the hard-label loss function value is adopted, and an appropriate weight parameter is introduced to balance the relationship between the two. Among them, the distillation loss function uses forms such as Kullback-Leibler divergence or mean squared error to measure the difference between the output of the student model and the output of the teacher model. On the other hand, the hard-label loss function directly measures the gap between the prediction result of the student model and the true label. To construct the comprehensive loss function, the weights of the two loss terms need to be determined. Generally speaking, in the initial stage of training, since the student model has not established a sufficient basic knowledge framework, the model relies more on the rich information provided by the teacher model; as the training process progresses, more attention is paid to the hard-label loss, prompting the student model to be closer to the true data distribution and improving its ability to solve problems independently. Once the overall loss function is defined, it can be applied to the training process of the student model. In each iteration, first, a batch of input data is fed into the student model to obtain the corresponding prediction results. Then, the distillation loss value and the hard-label loss value for this batch of data are calculated, and then the total loss value is obtained.Subsequently, the gradients of the parameters of each layer with respect to the total loss function are calculated using the backpropagation algorithm, and the model parameters are updated accordingly.

[0063] In summary, the method for constructing a large language model in the field of intelligent manufacturing that integrates domain knowledge distillation according to the embodiments of the present application is elucidated. It collects the knowledge in the field of intelligent manufacturing and performs knowledge cleaning and structured encoding to obtain the knowledge graph in the field of intelligent manufacturing. Meanwhile, a pre-trained large language model is selected as the teacher model, and the teacher model in the field of intelligent manufacturing is obtained by combining with the knowledge graph in the field of intelligent manufacturing. Then, the task data in the field of intelligent manufacturing is collected, and the student model is trained by knowledge distillation in combination with the teacher model in the field of intelligent manufacturing to obtain the student model in the field of intelligent manufacturing, so as to effectively integrate the structured domain knowledge, and pay attention to the structured characteristics and semantic relevance of the domain knowledge, thus taking into account the semantic similarity of the soft labels and the supervision signals of the hard labels, and balancing the generalization and domain specificity.

Claims

1. A method for constructing a large language model in the field of intelligent manufacturing that integrates domain knowledge distillation, characterized in that, Including: Collecting knowledge in the field of intelligent manufacturing; Performing knowledge clarification and structured encoding on the knowledge in the field of intelligent manufacturing to obtain a knowledge graph of the field of intelligent manufacturing; Selecting a pre-trained large language model as the teacher model; Fusing the model representation space of the teacher model with the knowledge graph of the field of intelligent manufacturing by using knowledge graph embedding technology to obtain a teacher model for the field of intelligent manufacturing; Collecting task data in the field of intelligent manufacturing; Using the teacher model for the field of intelligent manufacturing and the task data in the field of intelligent manufacturing to perform knowledge distillation training on the student model to obtain a student model for the field of intelligent manufacturing, where the student model has a relatively smaller number of parameters compared to the pre-trained large language model.

2. The method for constructing a large language model in the field of intelligent manufacturing that fuses domain knowledge distillation according to claim 1, wherein, Collecting knowledge in the field of intelligent manufacturing, including: collecting the knowledge in the field of intelligent manufacturing from multiple data sources, where the multiple data sources include domain documents, expert knowledge, knowledge bases or ontologies, or structured databases.

3. The method for constructing a large language model in the field of intelligent manufacturing that integrates domain knowledge distillation according to claim 1, wherein, Using the teacher model for the field of intelligent manufacturing and the task data in the field of intelligent manufacturing to perform knowledge distillation training on the student model to obtain a student model for the field of intelligent manufacturing, including: Inputting the task data in the field of intelligent manufacturing into the teacher model for the field of intelligent manufacturing to obtain soft labels; Inputting the task data in the field of intelligent manufacturing into the student model to obtain a prediction result; Calculating the value of the hard label loss function between the prediction result and the true label; Calculating the value of the distillation loss function between the prediction result and the soft labels; Using the weighted sum of the value of the distillation loss function and the value of the hard label loss function as the value of the loss function to train the student model to obtain the student model for the field of intelligent manufacturing.

4. The method for constructing a large language model in the field of intelligent manufacturing by fusing domain knowledge distillation according to claim 3, wherein, Calculating the value of the hard label loss function between the prediction result and the true label, including: calculating the value of the cross-entropy loss function between the prediction result and the true label as the value of the hard label loss function.

5. The method for constructing a large language model in the field of intelligent manufacturing that integrates domain knowledge distillation according to claim 3, characterized in that, Calculating the value of the distillation loss function between the prediction result and the soft labels, including: Performing semantic embedding encoding on the prediction result to obtain a prediction result semantic embedding encoding vector; Performing semantic embedding encoding on the soft labels to obtain a soft label semantic embedding encoding vector; Calculating the position-wise difference vector between the prediction result semantic embedding encoding vector and the soft label semantic embedding encoding vector to obtain a distillation semantic difference encoding vector; Performing feature enhancement on the distillation semantic difference encoding vector to obtain a distillation semantic enhancement encoding vector; Performing feature decoding on the distillation semantic enhancement encoding vector to obtain the value of the distillation loss function.

6. The method for constructing a large language model in the field of intelligent manufacturing that integrates domain knowledge distillation according to claim 5, wherein, Performing feature enhancement on the distillation semantic difference encoding vector to obtain a distillation semantic enhancement encoding vector, including: Performing feature phase reconstruction and information squeezing on the distillation semantic difference encoding vector to obtain a set of distillation semantic difference feature extraction local phase encoding vectors; Calculating the effective component statistics of the distillation semantic difference features for each distillation semantic difference feature extraction local phase encoding vector in the set of distillation semantic difference feature extraction local phase encoding vectors; Extract the statistical number of effective components of the distilled semantic difference features for extracting the local phase encoding vectors based on the respective distilled semantic difference features, and calculate the distilled semantic difference feature phase reshaping gain operator for each local phase encoding vector of the distilled semantic difference features; Based on the distilled semantic difference feature phase reshaping gain operator for each local phase encoding vector of the distilled semantic difference features, perform feature phase significance reshaping on the set of local phase encoding vectors of the distilled semantic difference features to obtain the distilled semantic enhanced encoding vectors.

7. The method for constructing a large language model in the field of intelligent manufacturing by fusing domain knowledge distillation according to claim 6, wherein, Perform feature phase reconstruction and information squeezing on the distilled semantic difference encoding vectors to obtain a set of local phase encoding vectors for extracting the distilled semantic difference features, including: Perform feature phase reconstruction based on one-dimensional convolutional encoding on the distilled semantic difference encoding vectors to obtain a set of local phase encoding vectors of the distilled semantic difference features; Perform information squeezing on each local phase encoding vector in the set of local phase encoding vectors of the distilled semantic difference features to obtain the set of local phase encoding vectors for extracting the distilled semantic difference features.

8. The method for constructing a large language model in the field of intelligent manufacturing by fusing domain knowledge distillation according to claim 7, characterized in that, Extract the statistical number of effective components of the distilled semantic difference features for extracting the local phase encoding vectors based on the respective distilled semantic difference features, and calculate the distilled semantic difference feature phase reshaping gain operator for each local phase encoding vector of the distilled semantic difference features, including: Based on the statistical number of effective components of the distilled semantic difference features for extracting the local phase encoding vectors, determine the suppression factors corresponding to the local phase encoding vectors for extracting the respective distilled semantic difference features; Based on the suppression factors corresponding to the local phase encoding vectors for extracting the respective distilled semantic difference features, calculate the distilled semantic difference feature phase reshaping gain operator for each local phase encoding vector of the distilled semantic difference features.

9. The method for constructing a large language model in the field of intelligent manufacturing by fusing domain knowledge distillation according to claim 8, wherein Based on the suppression factors corresponding to the local phase encoding vectors for extracting the respective distilled semantic difference features, calculate the distilled semantic difference feature phase reshaping gain operator for each local phase encoding vector of the distilled semantic difference features, including: Based on the suppression factors corresponding to the local phase encoding vectors for extracting the respective distilled semantic difference features, calculate the initial distilled semantic difference feature phase reshaping gain operator for each local phase encoding vector of the distilled semantic difference features; Perform feature phase scatter loss correction on the initial distilled semantic difference feature phase reshaping gain operator to obtain the distilled semantic difference feature phase reshaping gain operator.

10. The method for constructing a large language model in the field of intelligent manufacturing that fuses domain knowledge distillation according to claim 9, characterized in that, Perform feature decoding on the distilled semantic enhanced encoding vectors to obtain the distilled loss function value, including: passing the distilled semantic enhanced encoding vectors through a loss value calculator based on a decoder to obtain the distilled loss function value.

Citation Information

Cited By

  • Low-altitude multi-agent cooperative control method and system

    CN120523098A

  • Method and device for extracting legal clause information of license of open source software

    CN120952011A

  • Model distillation method and device based on knowledge base

    CN121009966A

  • Power equipment identification method based on deep learning

    CN121412758A

  • A power equipment identification method based on deep learning

    CN121412758B