Method and device for finely adjusting large language model applied to generic semiconductor field

By combining knowledge graphs from the semiconductor field with SPARQL query statements, and employing prefix fine-tuning and low-rank adaptive techniques to fine-tune large language models, the problems of lack of specificity and insufficient corpus in the application of large language models in the semiconductor field are solved, achieving higher accuracy and reliability.

CN120806191APending Publication Date: 2025-10-17CLP JIUTIAN INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411109458.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-08-13
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing large-scale language models lack specificity in the application of the semiconductor industry, leading to frequent misunderstandings and errors. Furthermore, the corpus is scarce and incomplete, making it difficult to meet the professional terminology and usage standards of the semiconductor industry.

Method used

By combining knowledge graphs from the semiconductor field, utilizing SPARQL query statements and pre-trained corpora, and employing prefix fine-tuning and low-rank adaptive techniques to fine-tune the large language model, natural language answers that conform to semiconductor field standards are generated.

Benefits of technology

It improves the performance and practicality of large language models in the general semiconductor field, enabling accurate understanding and handling of specific problems, supporting applications such as equipment maintenance and fault diagnosis, and reducing erroneous output.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120806191A_ABST
    Figure CN120806191A_ABST
Patent Text Reader

Abstract

The invention discloses a method and device for finely adjusting a large language model applied to the field of generic semiconductors, and belongs to the technical field of artificial intelligence. The method for finely adjusting the large language model applied to the generic semiconductor field comprises the steps of obtaining a first pre-training corpus based on a first SPARQL query statement and a first model in a knowledge graph of the generic semiconductor field; querying a knowledge graph of the generic semiconductor field through a second SPARQL query statement to obtain a second pre-training corpus; and performing fine tuning on the target large language model based on the first pre-training corpus and the second pre-training corpus by adopting a prefix fine tuning and low-rank adaptive technology. According to the method and the device for finely tuning the large language model applied to the generic semiconductor field, the knowledge graph of the generic semiconductor field is utilized, the text of the generic semiconductor field is combined, and the large language model capable of accurately understanding and processing specific problems in the generic semiconductor field is obtained by combining two fine tuning technologies.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of artificial intelligence, and particularly relates to a method and device for fine-tuning a large language model applied to the field of general semiconductors. BACKGROUND

[0002] At present, large language models have shown excellent processing and analysis capabilities in general language tasks. These models can understand and generate complex text content by utilizing a wide range of general-purpose corpora for pre-training. However, due to the lack of industry-specific training data, these models often misinterpret and make mistakes in specific domain question-answering processes, a phenomenon known as "hallucination", which limits their practical application effectiveness. To solve this problem, specific industries such as finance and medicine have begun to fine-tune models by combining industry-specific public or private corpora. This fine-tuning method based on industry-specific corpora significantly improves the accuracy and reliability of models in specific domains, thus enabling large language model technology to be applied in industry applications.

[0003] With the significant progress in natural language processing technology in recent years, large language models have become a key technical tool in multiple knowledge domains, demonstrating strong capabilities in handling complex language tasks. However, general large language models still have limitations in dealing with specific domain problems. These models often lack specificity due to the broadness of training data, especially when facing professional or technical content, often resulting in insufficient information accuracy, a phenomenon commonly referred to as "hallucination" in the academic community.

[0004] Specifically, these large language models, without specific domain training, may generate content that does not meet the requirements of practical applications, especially in industrial applications that require high reliability, where this deficiency is particularly evident. Although the introduction of specific domain corpora for pre-training can alleviate this problem to some extent, such as applications in the fields of finance, education, law, and medicine, which have proven the effectiveness of this method, the related research and application in the field of general semiconductors is relatively lagging behind.

[0005] At present, the general semiconductor industry lacks efficient language models that can fully utilize domain-specific knowledge when facing complex manufacturing processes and equipment maintenance. In addition, the specific text corpus in this field is more scarce and incomplete compared to other fields, which further exacerbates the limitations of large language models in this field.

[0006] However, in the semiconductor industry, the fine-tuning scheme of such models encounters several key problems and limitations. First, compared with other fields, the corpus in the semiconductor field is more scarce and incomplete, which significantly limits the training effect and application range of the model in this field. Second, due to the complexity of manufacturing process control and equipment maintenance scenarios, the "hallucination" phenomenon of the model is difficult to accept. In addition, the semiconductor field has strict specifications for terminology and language, and existing public datasets often cannot meet these specific needs. In order to fully utilize the advantages brought by large language model technology, it is necessary to establish a fine-tuning large model scheme suitable for the general semiconductor field. SUMMARY

[0007] The present application aims to at least solve one of the technical problems existing in the prior art. To this end, the present application proposes a method and device for fine-tuning a large language model applied to the general semiconductor field, which can make the large language model applicable to the general semiconductor field through fine-tuning.

[0008] In a first aspect, the present application provides a method for fine-tuning a large language model applied to the general semiconductor field, comprising:

[0009] Based on a first SPARQL query statement in a knowledge graph of the general semiconductor field and a first model, a first pre-training corpus is obtained; the first pre-training corpus includes multiple groups of natural language questions after prefix processing and second SPARQL query statements corresponding to the natural language questions;

[0010] Based on querying the knowledge graph of the general semiconductor field through the second SPARQL query statement, a second pre-training corpus is obtained; the second pre-training corpus includes multiple groups of natural language questions and natural language answers corresponding to the second SPARQL query statement; the natural language answers correspond to the query results obtained by querying the knowledge graph of the general semiconductor field through the second SPARQL query statement;

[0011] Using prefix fine-tuning and low-rank adaptive technology, the target large language model is fine-tuned based on the first pre-training corpus and the second pre-training corpus.

[0012] According to the method for fine-tuning a large language model applied to the general semiconductor field of the present application, by utilizing the knowledge graph of the general semiconductor field, the rich text and knowledge graph of the general semiconductor field are combined, and through advanced fine-tuning technology, the performance and practicality of the large language model in the general semiconductor field are improved, so that a large language model capable of accurately understanding and processing problems specific to the general semiconductor field can be obtained, providing more accurate and reliable support in device maintenance and other applications in the general semiconductor field.

[0013] According to one embodiment of the present application, the first SPARQL query statement and the first model based on the knowledge graph in the field of general semiconductor are used to obtain a first pre-training corpus, which includes:

[0014] The first SPARQL query statement is converted into a natural language question corresponding to the first SPARQL query statement, and a third corpus is obtained; the third corpus includes a plurality of groups of the first SPARQL query statement and the natural language question corresponding to the first SPARQL query statement.

[0015] The third corpus is used as a template instruction, and entities in the knowledge graph in the field of general semiconductor are traversed by the first model to obtain a plurality of corpus pairs; each corpus pair includes a second SPARQL query statement and a natural language question corresponding to the second SPARQL query statement.

[0016] Each corpus pair is preprocessed to obtain the first pre-training corpus.

[0017] According to one embodiment of the present application, the second pre-training corpus is obtained by querying the knowledge graph in the field of general semiconductor through the second SPARQL query statement, which includes:

[0018] Each second SPARQL query statement is used to query the knowledge graph in the field of general semiconductor to obtain a target triple.

[0019] The second pre-training corpus is obtained based on the natural language answer corresponding to each second SPARQL query statement and the target triple.

[0020] According to one embodiment of the present application, the second pre-training corpus is obtained based on the natural language answer corresponding to each second SPARQL query statement and the target triple, which includes:

[0021] A first sub-corpus is obtained based on the natural language answer corresponding to each second SPARQL query statement and the target triple; the first sub-corpus includes a plurality of groups of the natural language question corresponding to the second SPARQL query statement, the target triple, and the natural language answer corresponding to the target triple.

[0022] A second sub-corpus is obtained; the second sub-corpus includes a plurality of groups of randomly selected natural language questions, randomly selected target triples, and target answers; the target answer is used to indicate no answer.

[0023] The first sub-corpus and the second sub-corpus are merged to obtain the second pre-training corpus.

[0024] According to one embodiment of the present application, the prefix fine-tuning and low-rank adaptive technology is used to fine-tune the target large language model based on the first pre-training corpus and the second pre-training corpus, which includes:

[0025] The first pre-training corpus and the second pre-training corpus are mixed to obtain a merged corpus, and a source label indicating the source is added to the corpus in the merged corpus;

[0026] According to the source label, a corresponding prefix encoder is selected to generate a prefix embedding vector;

[0027] Through the low-rank adaptive technology, the target large language model is fine-tuned based on the added prefix embedding vector and the merged corpus.

[0028] According to one embodiment of the present application, the prefix embedding vector is generated according to the source label by selecting a corresponding prefix encoder, which includes:

[0029] In the case of freezing the parameters of the target large language model, a plurality of prefix embedding vectors are generated and concatenated to the front end of the embedding vectors of each layer of the target large language model;

[0030] After each round of training, the vector at the front end of the embedding vectors of each layer of the target large language model is extracted as the fine-tuned prefix embedding vector.

[0031] In a second aspect, the present application provides a device for fine-tuning a large language model applied to the field of general semiconductors, which comprises:

[0032] The first acquisition module is configured to acquire a first pre-training corpus based on a first SPARQL query statement and a first model in a knowledge graph of the field of general semiconductors; the first pre-training corpus includes a plurality of groups of natural language questions processed by prefixes and second SPARQL query statements corresponding to the natural language questions;

[0033] The second acquisition module is configured to acquire a second pre-training corpus based on querying the knowledge graph of the field of general semiconductors by using the second SPARQL query statement; the second pre-training corpus includes a plurality of groups of natural language questions and natural language answers corresponding to the second SPARQL query statement; the natural language answers correspond to the query results obtained by querying the knowledge graph of the field of general semiconductors by using the second SPARQL query statement;

[0034] The fine-tuning module is configured to fine-tune a target large language model based on the first pre-training corpus and the second pre-training corpus by using a prefix fine-tuning and low-rank adaptive technology.

[0035] The device for fine-tuning a large language model applied to the field of general semiconductors according to the present application can accurately understand and process problems specific to the field of general semiconductors by combining rich texts and knowledge graphs in the field of general semiconductors through advanced fine-tuning technology, thereby providing more accurate and reliable support in device maintenance and other applications in the field of general semiconductors.

[0036] In a third aspect, the present application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method for fine-tuning a large language model applied to the field of general semiconductors according to the first aspect.

[0037] In a fourth aspect, the present application provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program is executable by a processor to implement the method for fine-tuning a large language model applied to the field of general semiconductors according to the first aspect.

[0038] In a fifth aspect, the present application provides a chip, comprising a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is configured to run a program or an instruction to implement the method for fine-tuning a large language model applied to the field of general semiconductors according to the first aspect.

[0039] In a sixth aspect, the present application provides a computer program product comprising a computer program, wherein the computer program is executable by a processor to implement the method for fine-tuning a large language model applied to the field of general semiconductors according to the first aspect.

[0040] Additional aspects and advantages of the present application will be in part apparent and in part pointed out hereinafter. BRIEF DESCRIPTION OF DRAWINGS

[0041] The above and / or additional aspects and advantages of the present application will become apparent and be readily appreciated from the following description, including the appended drawings, wherein:

[0042] Figure 1 is a flowchart of the method for fine-tuning a large language model applied to the field of general semiconductors according to an embodiment of the present application;

[0043] Figure 2 is a structural schematic diagram of the device for fine-tuning a large language model applied to the field of general semiconductors according to an embodiment of the present application;

[0044] Figure 3 is a structural schematic diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0045] The technical solutions in the embodiments of the present application will be clearly described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art are within the scope of protection of the present application.

[0046] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of a kind and do not limit the number of objects, for example, the first object can be one or more. In addition, "and / or" in the specification and claims indicates at least one of the connected objects, and the character " / ", generally indicates that the objects before and after are in a "or" relationship.

[0047] The method for fine-tuning a large language model in the field of general semiconductors, the device for fine-tuning a large language model in the field of general semiconductors, the electronic device and the readable storage medium provided by the embodiments of the present application will be described in detail below with reference to the drawings and through specific embodiments and application scenarios.

[0048] The method for fine-tuning a large language model in the field of general semiconductors can be applied to a terminal, and can be executed by hardware or software in the terminal.

[0049] The terminal includes, but is not limited to, a portable communication device such as a mobile phone or a tablet computer having a touch-sensitive surface (e.g., a touchscreen display and / or a touchpad). It should also be understood that in some embodiments, the terminal can not be a portable communication device, but a desktop computer having a touch-sensitive surface (e.g., a touchscreen display and / or a touchpad).

[0050] In each of the following embodiments, a terminal including a display and a touch-sensitive surface is described. However, it should be understood that the terminal can include one or more other physical user interface devices such as physical keyboards, mice, and joysticks.

[0051] The method for fine-tuning a large language model applied to the field of general semiconductors provided by the embodiments of the present application can be executed by an electronic device or a functional module or functional entity of the electronic device capable of implementing the method for fine-tuning a large language model applied to the field of general semiconductors. The electronic device mentioned in the embodiments of the present application includes but is not limited to a mobile phone, a tablet computer, a computer, a camera, a wearable device, and the like. The method for fine-tuning a large language model applied to the field of general semiconductors provided by the embodiments of the present application will be described below by taking an electronic device as an example.

[0052] As shown in Figure 1 The method for fine-tuning a large language model applied to the field of general semiconductors includes steps 110, 120, and 130.

[0053] The method for fine-tuning a large language model applied to the field of general semiconductors provided by the embodiments of the present application aims to fine-tune a large language model so that it can be applied to the field of general semiconductors to perform production process control, fault diagnosis and maintenance, and the like in the field of general semiconductors.

[0054] In the field of general semiconductors, production process control can be performed by applying a large language model fine-tuned by the embodiments of the present application to a semiconductor production line to analyze production data in real time, predict and optimize the production process, and reduce the defect rate and improve the yield.

[0055] In the field of general semiconductors, fault diagnosis and maintenance can be performed by a large language model fine-tuned by the embodiments of the present application to achieve fault diagnosis and maintenance guidance for semiconductor equipment, and through natural language processing technology, maintenance personnel can quickly identify problems and obtain solutions.

[0056] Step 110: obtaining a first pre-training corpus based on a first SPARQL query statement and a first model in a knowledge graph of the field of general semiconductors; the first pre-training corpus includes multiple groups of natural language questions processed by a prefix and a second SPARQL query statement corresponding to the natural language questions.

[0057] In actual execution, the first pre-training corpus can be established first. The first pre-training corpus can be a pre-training corpus for converting natural language to a SPARQL query statement. The first pre-training corpus is a pre-training corpus for converting natural language to a logical form (expressed as a SPARQL query statement).

[0058] Step 110 involves starting from the annotation of field entities to build a first pre-training corpus capable of converting extensive natural language tables to SPARQL query statements for training a model to perform effective semantic parsing.

[0059] The SPARQL language is a query language in the form of logical expressions for a knowledge graph database. A SPARQL query statement can be composed of multiple SPO triples and their logical association statements, and is used to perform knowledge retrieval of a knowledge graph database. The knowledge graph database supports parsing and execution of such query statements. For example, the query statement “(AND(JOIN[manufacturers,manufacturer,wafer][12-inch wafer])(JOIN(R[wafer,wafer size,12-inch]))” is used to retrieve manufacturers capable of producing 12-inch wafers.

[0060] The knowledge graph of the general semiconductor field can adopt a Resource Description Framework (RDF) structure composed of triples of subject, predicate, and object (SPO) to describe the knowledge content of the general semiconductor field. For example, the triple (wafer, wafersize, 12-inch) on the “wafer” entity can be described as “wafer with a 12-inch size”.

[0061] The SPARQL query statement corresponding to the natural language question can be generated using the general semiconductor field knowledge graph database.

[0062] In some embodiments, the SPARQL query statement corresponding to the natural language question can be obtained through a target dataset, thereby generating a corpus composed of a large number of corpora corresponding the SPARQL query statement to the corresponding natural language question. In some embodiments, the above-mentioned target dataset can be a MetaQA, WebQuestionsSP, or CWQ dataset, etc. For example, the CWQ dataset collects a large number of SPARQL query statements for Internet search and their corresponding natural language questions. Through entity ID replacement in data preprocessing, the SPARQL query statement can be corresponded to the corresponding natural language question.

[0063] In some embodiments, the above-mentioned target dataset can be enhanced with knowledge of the general semiconductor field to make up for the lack of specific knowledge of the semiconductor field based on the construction of the general field knowledge database.

[0064] It should be noted that the second SPARQL query statement is obtained by expanding the first SPARQL query statement in the general semiconductor field knowledge graph using the general semiconductor field knowledge graph and the first model. The first model can be a large language model. The specific first model used is not limited by the embodiments of the present application.

[0065] Step 120, based on querying the knowledge graph in the field of generic semiconductor by the second SPARQL query statement, a second pre-training corpus is obtained; the second pre-training corpus includes a plurality of sets of natural language questions and natural language answers corresponding to the second SPARQL query statement; the natural language answers correspond to the query results obtained by querying the knowledge graph in the field of generic semiconductor by the second SPARQL query statement.

[0066] In actual execution, with the help of the existing context reasoning ability of the large language model, combined with the natural language question, a set of natural language answers can be obtained. In order to better conform to the language habits and norms of the generic semiconductor field, a data set for fine-tuning of the language large model is further constructed, which can be the second pre-training corpus.

[0067] After obtaining the logical form query statement, a corpus can be further constructed to map the query results of the knowledge graph in the field of generic semiconductor to the correct natural language answers. The second SPARQL query statement of step 110 can be used to construct a mapping from the knowledge graph query triple to the natural language answer. Through step 120, the fine-tuned large language model can generate accurate and contextually appropriate answers based on the query results.

[0068] It should be noted that constructing a domain-specific pre-training corpus based on a knowledge graph (KG) is one of the key steps to achieve efficient natural language understanding. In the field of generic semiconductors, by integrating structured databases within the industry (a knowledge graph covering generic semiconductor knowledge), a mapping relationship between natural language questions and SPARQL query statements is generated, as well as a conversion from query result triples to natural language answers, forming two high-quality pre-training corpora.

[0069] It should be noted that using the knowledge graph in the field of generic semiconductors can improve the accuracy and practicality of the fine-tuned large language model. Given the specific terminology and language norms in the field of generic semiconductors, and the "hallucination" phenomenon of large language models, by integrating the large language model with the existing knowledge graph and fine-tuning training, the fine-tuned large language model can generate answers based on information in the knowledge graph as much as possible. This approach can effectively reduce incorrect answers caused by the "hallucination" phenomenon, thereby improving accuracy in the field.

[0070] It should be noted that the first pre-training corpus and the second pre-training corpus not only enable the large language model to perform effective semantic analysis, but also generate natural language answers that conform to the standards of the generic semiconductor field.

[0071] Step 130, fine-tuning the target large language model based on the first pre-training corpus and the second pre-training corpus using prefix fine-tuning and low-rank adaptive techniques.

[0072] In actual implementation, the target large language model can be fine-tuned by using the prefix fine-tuning and LoRA technology to obtain a fine-tuned large language model dedicated to the field of general semiconductors. It can be understood that the fine-tuned large language model dedicated to the field of general semiconductors can be applied to the field of general semiconductors.

[0073] In some embodiments, the target large language model can be any open-source large language model, such as the GLM130B model, etc. The input of the target large language model is text data. According to different tasks performed by the target large language model, the output of the target large language model can include at least one of query results, judgment results, and answers to questions (which can be in the form of text and / or voice, etc.).

[0074] Prefix fine-tuning (Prefix-Tuning) and LoRA (Low-Rank Adaptation) are two efficient large language model fine-tuning techniques that adjust a small number of model parameters to adapt to specific tasks, reducing resource consumption while maintaining model flexibility. Prefix fine-tuning adds learnable "prefix" parameters before each decoding layer without directly modifying the parameters of the pre-trained model. These prefix parameters act as conditional information to guide the model to output task-related results. LoRA, on the other hand, achieves fine-tuning by applying low-rank updates to weight matrices. This method modifies parameters in a small and structured manner rather than retraining the entire network to adapt the model.

[0075] GLM-130B is an open-source bilingual (Chinese and English) bidirectional dense large language model with 130 billion parameters. It is based on the general language model (GLM) architecture and is suitable for a variety of language processing tasks. The pre-training corpus prepared in steps 110 and 120 can be used in conjunction with prefix fine-tuning and LoRA technology to fine-tune this model to meet the specific needs of the field of general semiconductors.

[0076] During fine-tuning, LoRA technology is applied to add fine-tuning residual parameters ΔW = A x B to the weight matrix W of the large language model, where the ranks of matrices A and B are much smaller than W. During training, the W parameters remain fixed, and only A and B are allowed to update, thereby minimizing the total number of parameters during training. During the forward propagation of the network, the original input x is processed through the network layer fine-tuned by LoRA technology, and the output changes from Wx to Wx + (A x B)x. The fine-tuning process based on LoRA technology can be as shown in Figure 2 .

[0077] It should be noted that step 130 adopts the bilingual corpus model fine-tuning of the joint action of LoRA technology and prefix fine-tuning. LoRA technology and prefix fine-tuning are two advanced model fine-tuning methods, and through their joint action, efficient fine-tuning of large language models can be achieved. The LoRA layer is shared in the training of two pre-training corpora (i.e., the first pre-training corpus and the second pre-training corpus), effectively passing on the usage specifications and background knowledge of generic semiconductor field terms. Prefix fine-tuning is performed separately for each training, enabling the large language model to adapt to different specific tasks within the generic semiconductor field. This fast-switching fine-tuning prefix approach aims to minimize model load while ensuring the accuracy of question answering, enhancing the model's applicability and generalization ability in different tasks.

[0078] The parameter optimization of the large language model using LoRA technology and prefix fine-tuning technology enables the fine-tuned large language model to effectively utilize the training of two corpora in the same field and adapt to the needs of two specific tasks in the generic semiconductor field. The parameter optimization of the large language model using LoRA technology and prefix fine-tuning technology can achieve the purpose of reducing the parameter amount while maintaining the complexity of the model. The parameter optimization of the large language model using LoRA technology and prefix fine-tuning technology can make the fine-tuned large language model more lightweight, improve the running efficiency and response speed, and optimize the performance.

[0079] The fine-tuned language model obtained through the above steps has significantly improved performance in semantic understanding, logical form generation, and correct answer generation in the generic semiconductor field. The model can more accurately process professional knowledge and data queries, effectively supporting professional applications in the semiconductor industry such as equipment maintenance and fault diagnosis, improving work efficiency and safety.

[0080] The method of fine-tuning the large language model applied to the generic semiconductor field provided by the embodiments of the present application combines rich generic semiconductor field text and knowledge graphs through the use of a knowledge graph in the generic semiconductor field, improves the performance and practicality of the large language model in the generic semiconductor field through advanced fine-tuning technology, and thus obtains a large language model that can accurately understand and process problems specific to the generic semiconductor field, providing more accurate and reliable support in applications such as equipment maintenance in the generic semiconductor field.

[0081] In some embodiments, based on the first SPARQL query statement and the first model in the knowledge graph of the generic semiconductor field, the first pre-training corpus is obtained, including: converting the first SPARQL query statement into a natural language question corresponding to the first SPARQL query statement, and obtaining a third corpus; the third corpus includes multiple groups of first SPARQL query statements and natural language questions corresponding to the first SPARQL query statements.

[0082] In actual execution, the first SPARQL query statement can be a SPARQL query statement obtained by manual annotation on an existing SPARQL query statement in a semiconductor industry knowledge graph database. Through manual annotation, it can be ensured that these first SPARQL query statements can correctly retrieve information.

[0083] Through manual annotation of the text, the general semiconductor field-specific language specification can be imparted to the large language model, and the practicability of the large language model in the field is enhanced.

[0084] In some embodiments, a large model distillation technology can be used to convert the first SPARQL query statement into a natural language question through a complex large model interface, thereby obtaining a small sample of a batch of dedicated "first SPARQL query statement-natural language question". The above small sample can constitute a third corpus.

[0085] In some embodiments, the "first SPARQL query statement-natural language question" converted through the complex large model interface can be manually reviewed and modified first.

[0086] The third corpus is used as a template instruction to traverse entities in a knowledge graph in the general semiconductor field through the first model, and a plurality of corpus pairs are obtained; each corpus pair includes a second SPARQL query statement and a natural language question corresponding to the second SPARQL query statement.

[0087] In actual execution, each entity in an existing general semiconductor knowledge graph can be traversed, and the above small sample is used as a template instruction to guide the first model to generate a semiconductor field "second SPARQL query statement-natural language question" corpus pair with consistent format.

[0088] In some embodiments, the above corpus pairs can also be screened to screen out corpus pairs corresponding to SPARQL statement pairs that can obtain retrieval results, thereby creating a domain-specific (in the embodiments of the present application, the general semiconductor field) dataset for complex language model knowledge distillation.

[0089] Each corpus pair is preprocessed to obtain a first pre-training corpus.

[0090] In actual execution, the "second SPARQL query statement-natural language question" corpus pair generated in the above step can be further preprocessed to construct a pre-training corpus for natural language to logical form conversion, that is, a first pre-training corpus. In some embodiments, the number of corpus in the first pre-training corpus can reach several hundred or several thousand, for example, about 1000 data, etc.

[0091] In the first pre-training corpus, the second SPARQL query statement r is taken as the target statement, and the natural language question q after prefix processing is taken as the input statement. The prefix processing here involves adding the prefix text "Please give the SPARQL query statement corresponding to the following natural language question" to the natural language question to clearly indicate the task target of the large language model.

[0092] According to the method for fine-tuning the large language model applied to the field of general semiconductors provided by the embodiments of the present application, by utilizing the knowledge graph in the field of general semiconductors, the rich texts and knowledge graphs in the field of general semiconductors are combined to form a high-quality pre-training corpus, so that the large language model can not only perform effective semantic analysis, but also generate natural language answers conforming to the standards in the field of general semiconductors.

[0093] In some embodiments, the second pre-training corpus is obtained based on querying the knowledge graph in the field of general semiconductors by the second SPARQL query statement, including: querying the knowledge graph in the field of general semiconductors by each second SPARQL query statement to obtain a target triple.

[0094] In actual execution, the knowledge graph in the field of general semiconductors can be searched by each second SPARQL query statement to obtain an SPO triple answer. The SPO triple answer is the target triple, i.e., the target triple corresponding to the second SPARQL query statement.

[0095] The second pre-training corpus is further constructed using the second SPARQL query statement. The SPO triple answer r (i.e., the target triple r) is obtained by searching the knowledge graph in the field of general semiconductors using the above query statement.

[0096] The second pre-training corpus is obtained based on the natural language answer corresponding to each second SPARQL query statement and the target triple.

[0097] In actual execution, each target triple r can be manually annotated or automatically annotated to annotate the corresponding natural language answer a. By manually annotating the text, the large language model can be taught the specific terminology specifications in the field of general semiconductors, and the practicality of the large language model in this field can be enhanced.

[0098] The corpus in the second pre-training corpus can be represented as (r|q, a). Each corpus can include a natural language question and a target triple as training input text, and a natural language answer as output text. The second pre-training corpus can be used to convert triples in the knowledge graph into natural language answers. In some embodiments, the number of corpora in the second pre-training corpus can reach several hundred or several thousand, for example, about 1000 data, etc.

[0099] The method for fine-tuning a large language model in the field of general semiconductors provided by the embodiments of the present application can combine rich texts and knowledge graphs in the field of general semiconductors to form a high-quality pre-training corpus, so that the large language model can not only perform effective semantic analysis, but also generate natural language answers that meet the standards in the field of general semiconductors.

[0100] In some embodiments, based on the natural language answer corresponding to each second SPARQL query statement and the target triple, the second pre-training corpus is obtained, including: based on the natural language answer corresponding to each second SPARQL query statement and the target triple, a first sub-corpus is obtained; the first sub-corpus includes a plurality of groups of natural language questions corresponding to the second SPARQL query statements, target triples and natural language answers corresponding to the target triples

[0101] In actual execution, the first sub-corpus can be composed of real natural language questions, target triples obtained by querying, and natural language answers converted from the target triples. The process of obtaining the first sub-corpus can be referred to in the foregoing embodiments, which will not be described here.

[0102] A second sub-corpus is obtained; the second sub-corpus includes a plurality of groups of randomly selected natural language questions, randomly selected target triples and target answers; the target answer is used to indicate no answer.

[0103] In actual execution, the second sub-corpus can also be constructed. The second sub-corpus can be a pseudo corpus. A triple rf randomly queried from the knowledge graph in the field of general semiconductors is combined with an irrelevant random natural language question qf to form a pseudo input (rf|qf). In some embodiments, a fixed text format can be used for the output text af, such as “I am sorry, I cannot query the answer to the question from the knowledge graph”. The target answer is the output text af.

[0104] It can be understood that the randomly selected natural language question is irrelevant to the randomly selected target triple, that is, the randomly selected target triple is different from the target triple obtained by searching the knowledge graph in the field of general semiconductors through the second SPARQL query statement corresponding to the randomly selected natural language question.

[0105] By constructing the pseudo corpus, the scenario in which the knowledge graph in the field of general semiconductors cannot provide an answer can be simulated, the large language model can be prevented from incorrectly outputting answers due to the “hallucination” phenomenon, and the “hallucination” phenomenon of the large language model can be reduced.

[0106] The first sub-corpus and the second sub-corpus are merged to obtain the second pre-training corpus.

[0107] In actual execution, the first sub-corpus and the second sub-corpus can be merged to obtain the second pre-training corpus.

[0108] According to the method for fine-tuning the large language model applied to the field of general semiconductors provided in the embodiments of the present application, the correlation between the query result and the input question is automatically explored by using the language processing capability of the large language model through the construction of the pseudo corpus, and then it is judged whether the answer of "I don't know" should be output, which can further improve the accuracy of the output answer of the fine-tuned large language model, improve the anti-interference capability of the fine-tuned large language model to irrelevant corpus, and effectively reduce the false output.

[0109] In some embodiments, the first pre-training corpus and the second pre-training corpus are used to fine-tune the target large language model based on prefix fine-tuning and low-rank adaptive technology, including: mixing the first pre-training corpus and the second pre-training corpus to obtain a merged corpus, and adding a source label for indicating the source to the corpus in the merged corpus.

[0110] In actual execution, in the actual fine-tuning process, the first pre-training corpus and the second pre-training corpus can be trained at the same time. The first pre-training corpus and the second pre-training corpus can be mixed first, and the source labels are marked in them to distinguish them.

[0111] According to the source label, a corresponding prefix encoder is selected to generate a prefix embedding vector.

[0112] In actual execution, since the first pre-training corpus and the second pre-training corpus involve different language tasks in the same knowledge field, the same LoRA fine-tuning parameter matrix A and B are used in the fine-tuning process, and different prefix fine-tuning encoders p and p' are used to distinguish specific tasks.

[0113] During training, the corresponding prefix encoder p or p' is selected according to the source label, and the two tasks share a set of LoRA parameters.

[0114] Through the low-rank adaptive technology, the target large language model is fine-tuned based on the prefix embedding vector and the merged corpus.

[0115] In actual execution, the LoRA technology can be used to fine-tune the target large language model based on the prefix embedding vector and the merged corpus obtained by mixing the first pre-training corpus and the second pre-training corpus.

[0116] It can be understood that the large language model applicable to the field of general semiconductors can be obtained under the condition of a preset number or model convergence. Through experiments, the large language model applicable to the field of general semiconductors can be obtained after 700 batches of fine-tuning training.

[0117] According to the method for fine-tuning the large language model applied to the field of general semiconductors provided in the embodiments of the present application, through the parameter optimization of the large language model combined with the LoRA technology and the prefix fine-tuning technology, the fine-tuned large language model can effectively utilize the training of two corpora in the same field, can adapt to the requirements of two specific tasks in the field of general semiconductors, can reduce the parameter amount while maintaining the complexity of the model, can make the fine-tuned large language model more lightweight, improve the running efficiency and response speed, and optimize the performance.

[0118] In some embodiments, according to the source label, a corresponding prefix encoder is selected to generate a prefix embedding vector, including: generating a plurality of prefix embedding vectors and concatenating them to the front end of the embedding vectors of each layer of the target large language model in the case of freezing the parameters of the target large language model; after each round of training is completed, the vectors at the front end of the embedding vectors of each layer of the target large language model are extracted as the fine-tuned prefix embedding vectors.

[0119] In actual execution, in the training based on the corpora in the first pre-training corpus and the second pre-training corpus respectively, a plurality of prefix embedding vectors e1, e2, …, en can be generated using a learnable encoder network after the parameters of the target large language model, and these vectors are concatenated to the front end of the embedding vectors of each layer of the large model respectively. After the training is completed, the vectors embedded at the front end can be extracted and corresponded as the fine-tuned prefixes of the first pre-training corpus and the second pre-training corpus respectively, that is, the fine-tuned prefix embedding vectors.

[0120] According to the method for fine-tuning the large language model applied to the field of general semiconductors provided in the embodiments of the present application, through the parameter optimization of the large language model combined with the LoRA technology and the prefix fine-tuning technology, the fine-tuned large language model can effectively utilize the training of two corpora in the same field, can adapt to the requirements of two specific tasks in the field of general semiconductors, can reduce the parameter amount while maintaining the complexity of the model, can make the fine-tuned large language model more lightweight, improve the running efficiency and response speed, and optimize the performance.

[0121] The method for fine-tuning the large language model applied to the field of general semiconductors provided in the embodiments of the present application can be executed by the device for fine-tuning the large language model applied to the field of general semiconductors.

[0122] The embodiments of the present application also provide a device for fine-tuning the large language model applied to the field of general semiconductors.

[0123] AsFigure 2 As shown, the device for fine-tuning a large language model applied to the field of general semiconductors includes a first acquisition module 210, a second acquisition module 220, and a fine-tuning module 230.

[0124] The first acquisition module 210 is configured to acquire a first pre-training corpus based on a first SPARQL query statement and a first model in a knowledge graph in the field of general semiconductors; the first pre-training corpus includes a plurality of groups of natural language questions after prefix processing and second SPARQL query statements corresponding to the natural language questions.

[0125] The second acquisition module 220 is configured to acquire a second pre-training corpus by querying the knowledge graph in the field of general semiconductors based on the second SPARQL query statement; the second pre-training corpus includes a plurality of groups of natural language questions and natural language answers corresponding to the second SPARQL query statements; the natural language answers correspond to query results obtained by querying the knowledge graph in the field of general semiconductors based on the second SPARQL query statement.

[0126] The fine-tuning module 230 is configured to fine-tune a target large language model based on the first pre-training corpus and the second pre-training corpus by using prefix fine-tuning and low-rank adaptive technology.

[0127] The device for fine-tuning a large language model applied to the field of general semiconductors provided by the embodiments of the present application combines rich texts and knowledge graphs in the field of general semiconductors by using the knowledge graph in the field of general semiconductors, improves the performance and practicality of the large language model in the field of general semiconductors by using advanced fine-tuning technology, and thus obtains a large language model that can accurately understand and process problems specific to the field of general semiconductors, thereby providing more accurate and reliable support in device maintenance and other applications in the field of general semiconductors.

[0128] In some embodiments, the first acquisition module 210 can be specifically configured to:

[0129] convert the first SPARQL query statement into a natural language question corresponding to the first SPARQL query statement to acquire a third corpus; the third corpus includes a plurality of groups of the first SPARQL query statement and the natural language question corresponding to the first SPARQL query statement;

[0130] use the third corpus as a template instruction to traverse entities in the knowledge graph in the field of general semiconductors by using the first model to acquire a plurality of corpus pairs; each corpus pair includes a second SPARQL query statement and a natural language question corresponding to the second SPARQL query statement;

[0131] perform preprocessing on each corpus pair to acquire the first pre-training corpus.

[0132] In some embodiments, the second obtaining module 220 can include:

[0133] The query unit is configured to query the knowledge graph in the field of general semiconductors by each second SPARQL query statement to obtain target triples, respectively.

[0134] The obtaining unit is configured to obtain a second pre-training corpus based on each second SPARQL query statement and a natural language answer corresponding to the target triples.

[0135] In some embodiments, the obtaining unit can be specifically configured to:

[0136] obtain a first sub-corpus based on each second SPARQL query statement and a natural language answer corresponding to the target triples; the first sub-corpus includes multiple groups of natural language questions corresponding to the second SPARQL query statements, target triples, and natural language answers corresponding to the target triples;

[0137] obtain a second sub-corpus; the second sub-corpus includes multiple groups of randomly selected natural language questions, randomly selected target triples, and target answers; the target answers are used to indicate no answer;

[0138] merge the first sub-corpus and the second sub-corpus to obtain the second pre-training corpus.

[0139] In some embodiments, the fine-tuning module 230 can include:

[0140] The mixing unit is configured to mix the first pre-training corpus and the second pre-training corpus to obtain a merged corpus, and add a source label for indicating a source to the corpus in the merged corpus;

[0141] The prefix unit is configured to select a corresponding prefix encoder to generate a prefix embedding vector according to the source label;

[0142] The fine-tuning unit is configured to fine-tune the target large language model based on the added prefix embedding vector and the merged corpus by using a low-rank adaptive technique.

[0143] In some embodiments, the prefix unit can be specifically configured to:

[0144] generate multiple prefix embedding vectors and concatenate them to the front end of the embedding vectors of each layer of the target large language model under the condition that the parameters of the target large language model are frozen;

[0145] After each round of training is completed, the vectors at the front end of the embedding vectors of each layer of the target large language model are extracted as the fine-tuned prefix embedding vectors.

[0146] The device for fine-tuning a large language model in the generic semiconductor field in the embodiments of the present application can be an electronic device or a component in an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other device other than a terminal. For example, the electronic device can be a mobile phone, a tablet computer, a notebook computer, a palm computer, a vehicle-mounted electronic device, a mobile Internet device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), and can also be a server, a network attached storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, and the like. The embodiments of the present application are not limited in this regard.

[0147] The device for fine-tuning a large language model in the generic semiconductor field in the embodiments of the present application can be a device with an operating system. The operating system can be an Android operating system, an iOS operating system, or other possible operating systems, and the embodiments of the present application are not limited in this regard.

[0148] The device for fine-tuning a large language model in the generic semiconductor field provided in the embodiments of the present application can implement the method embodiments Figure 1 The method embodiments implement various processes, and to avoid repetition, the details are not described here.

[0149] In some embodiments, as shown in Figure 3 The embodiments of the present application also provide an electronic device 300, which includes a processor 310, a memory 320, and a computer program stored in the memory 320 and executable on the processor 310. When the processor 310 executes the program, it implements various processes of the above-mentioned method embodiments for fine-tuning a large language model in the generic semiconductor field and achieves the same technical effects. To avoid repetition, the details are not described here.

[0150] It should be noted that the electronic device in the embodiments of the present application includes the mobile electronic device and the non-mobile electronic device described above.

[0151] The embodiment of the present application also provides a non-transitory computer-readable storage medium, which stores a computer program. The computer program is executed by a processor to implement each process of the method for fine-tuning a large language model applied to the field of general semiconductors and achieve the same technical effects. To avoid repetition, details are not described herein.

[0152] The processor is the processor in the electronic device in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer readable only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0153] The embodiment of the present application also provides a computer program product, which includes a computer program. The computer program is executed by a processor to implement the method for fine-tuning a large language model applied to the field of general semiconductors.

[0154] The processor is the processor in the electronic device in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer readable only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0155] The embodiment of the present application also provides a chip, which includes a processor and a communication interface. The communication interface is coupled with the processor. The processor is configured to run a program or an instruction to implement each process of the method for fine-tuning a large language model applied to the field of general semiconductors and achieve the same technical effects. To avoid repetition, details are not described herein.

[0156] It should be understood that the chip mentioned in the embodiment of the present application can also be referred to as a system-level chip, a system chip, a chip system or a system-on-chip, etc.

[0157] It should be noted that, in this document, the terms "comprising", "containing" or any other variant thereof are intended to cover non-exclusive inclusion, so that processes, methods, articles or devices that include a series of elements not only include those elements, but also include other elements not explicitly listed or inherent to such processes, methods, articles or devices. Without more limitations, the element defined by the statement "comprising a" does not exclude the presence of additional identical elements in the process, method, article or device that includes the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to the order of performing the functions as shown or discussed, but can also include performing the functions in a substantially simultaneous manner or in a reverse order, for example, the described method can be performed in an order different from the described order, and various steps can be added, omitted or combined. In addition, the features described with reference to certain examples can be combined in other examples.

[0158] Through the above description of the embodiments, those skilled in the art can clearly understand that the above-mentioned example methods can be realized by means of software and necessary general hardware platforms, and of course, can also be realized by hardware, but in many cases, the former is a better embodiment. Based on such understanding, the technical solutions of the present application can be embodied in the form of a computer software product in essence or in the form of a part of the prior art that makes a contribution. The computer software product is stored in a storage medium (such as a ROM / RAM, a magnetic disc, an optical disc), and includes a plurality of instructions for causing a terminal (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in various embodiments of the present application.

[0159] The embodiments of the present application are described above in combination with the accompanying drawings, but the present application is not limited to the above-mentioned specific embodiments, and the above-mentioned specific embodiments are only illustrative, but not restrictive. Those skilled in the art can make many forms under the inspiration of the present application without departing from the scope of the present application and the scope protected by the claims.

[0160] In the description of the present specification, the description of the terms "one embodiment", "some embodiments", "an illustrative embodiment", "an example", "a specific example", or "some examples" means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In the present specification, the illustrative description of the above terms does not necessarily mean the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0161] Although the embodiments of the present application have been shown and described, those skilled in the art can understand that various changes, modifications, replacements and variations can be made to the embodiments without departing from the principles and purposes of the present application, and the scope of the present application is defined by the claims and their equivalents.

Claims

1. A method for fine-tuning a large language model applied to the pan-semiconductor field, characterized in that: include: Based on a first SPARQL query statement and a first model in a knowledge graph of the pan-semiconductor field, a first pre-trained corpus is obtained; The first pre-training corpus includes a plurality of sets of prefix-processed natural language questions and second SPARQL query statements corresponding to the natural language questions; Based on querying the knowledge graph of the pan-semiconductor field through the second SPARQL query statement, obtaining a second pre-trained corpus; the second pre-trained corpus includes multiple groups of the natural language questions and natural language answers corresponding to the second SPARQL query statement; The natural language answer corresponds to a query result obtained by querying the knowledge graph of the pan-semiconductor field through the second SPARQL query statement; Prefix fine-tuning and low-rank adaptation techniques are used to fine-tune the target large language model based on the first pre-training corpus and the second pre-training corpus.

2. The method for fine-tuning a large language model applied to the pan-semiconductor field according to claim 1, characterized in that: The step of obtaining a first pre-trained corpus based on a first SPARQL query statement and a first model in a knowledge graph in the pan-semiconductor field includes: Converting the first SPARQL query statement into a natural language question corresponding to the first SPARQL query statement to obtain a third corpus; the third corpus includes multiple groups of the first SPARQL query statements and the natural language questions corresponding to the first SPARQL query statements; Using the third corpus as a template instruction, traversing entities in the knowledge graph of the pan-semiconductor field through the first model to obtain multiple corpus pairs; each of the corpus pairs includes a second SPARQL query statement and a natural language question corresponding to the second SPARQL query statement; Preprocess each of the corpus pairs to obtain the first pre-training corpus.

3. The method for fine-tuning a large language model applied to the pan-semiconductor field according to claim 1, characterized in that: The acquiring of a second pre-trained corpus based on querying the knowledge graph of the pan-semiconductor field through the second SPARQL query statement includes: Using each of the second SPARQL query statements, query the knowledge graph of the pan-semiconductor field to obtain target triples; The second pre-training corpus is obtained based on each second SPARQL query statement and a natural language answer corresponding to the target triple.

4. The method for fine-tuning a large language model applied to the pan-semiconductor field according to claim 3, characterized in that: The acquiring of the second pre-trained corpus based on each second SPARQL query statement and the natural language answer corresponding to the target triple includes: Based on each of the second SPARQL query statements and the natural language answers corresponding to the target triples, a first sub-corpus is obtained; the first sub-corpus includes multiple sets of natural language questions corresponding to the second SPARQL query statements, the target triples, and the natural language answers corresponding to the target triples; Obtaining a second sub-corpus; the second sub-corpus includes a plurality of randomly selected natural language questions, randomly selected target triples, and target answers; the target answer is used to indicate that there is no answer; The first sub-corpus and the second sub-corpus are merged to obtain the second pre-training corpus.

5. The method for fine-tuning a large language model applied to the pan-semiconductor field according to any one of claims 1 to 4, characterized in that: The method of fine-tuning the target large language model based on the first pre-training corpus and the second pre-training corpus by using prefix fine-tuning and low-rank adaptation technology includes: Mixing the first pre-trained corpus and the second pre-trained corpus to obtain a merged corpus, and adding a source tag for indicating the source to the corpus in the merged corpus; According to the source label, selecting a corresponding prefix encoder to generate a prefix embedding vector; The target large language model is fine-tuned based on the added prefix embedding vector and the merged corpus through a low-rank adaptation technique.

6. The method for fine-tuning a large language model applied to the pan-semiconductor field according to claim 5, characterized in that: The step of selecting a corresponding prefix encoder to generate a prefix embedding vector according to the source label includes: While freezing the parameters of the target large language model, generating a plurality of prefix embedding vectors and connecting them in series to the front end of the embedding vectors of each layer of the target large language model; After each round of training, the front-end vectors of the embedding vectors of each layer of the target large language model are extracted as the fine-tuned prefix embedding vectors.

7. A device for fine-tuning a large language model applied to the pan-semiconductor field, characterized in that: include: A first acquisition module is configured to acquire a first pre-trained corpus based on a first SPARQL query statement and a first model in a knowledge graph of a pan-semiconductor field; The first pre-training corpus includes a plurality of sets of prefix-processed natural language questions and second SPARQL query statements corresponding to the natural language questions; A second acquisition module is configured to acquire a second pre-trained corpus based on querying the knowledge graph of the pan-semiconductor field using the second SPARQL query statement; the second pre-trained corpus includes multiple sets of the natural language questions and natural language answers corresponding to the second SPARQL query statement; The natural language answer corresponds to a query result obtained by querying the knowledge graph of the pan-semiconductor field through the second SPARQL query statement; A fine-tuning module is used to fine-tune the target large language model based on the first pre-training corpus and the second pre-training corpus by using prefix fine-tuning and low-rank adaptation technology.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, it implements the method for fine-tuning a large language model applied to the pan-semiconductor field as described in any one of claims 1-6.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for fine-tuning a large language model applied to the pan-semiconductor field as described in any one of claims 1 to 6 is implemented.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method for fine-tuning a large language model applied to the pan-semiconductor field as described in any one of claims 1 to 6 is implemented.