Automobile industry knowledge base construction method and system based on large language model
Through a method based on a large language model, automobile industry data is collected and marked from the Internet, and a multi-task model is built for fine-tuning, solving the problems of high time and cost and poor information timeliness in the existing automobile industry knowledge base construction methods, and achieving efficient and accurate knowledge base construction and information retrieval.
Patent Information
- Application Number
- CN202411926960.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-25
- Publication Date
- 2025-05-13
AI Technical Summary
The existing automotive industry knowledge base construction method requires a lot of time and labor costs, cannot effectively meet the timeliness of knowledge and information, and cannot make full use of the massive sparse industry data and knowledge in the Internet.
Using a method based on a large language model, raw unstructured data from a wide range of sources is collected from the Internet, and manual labeling costs are reduced through small proportion sampling and large language model annotation, and multi-task model is built for fine-tuning, so as to realize vectorized data processing and real-time updates.
It has realized the efficient construction of the automotive industry knowledge base, improved the accuracy of the information retrieval of the knowledge base, provided accurate and detailed information answers, promoted knowledge dissemination and sharing, and supported smart decision-making.
Smart Images

Figure CN119990271A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of vector knowledge base construction, and in particular to a method and system for constructing an automobile industry knowledge base based on a large language model. Background Art
[0002] It is still a challenge to screen and organize the industry knowledge and build an industry knowledge base to support the retrieval-augmented generation (RAG) of large language models. Existing knowledge base construction methods usually collect and organize industry-related text data manually, and then process the data content-independent data (including text segmentation, combination, vectorization, etc.) and finally store it in a vector database for storage.
[0003] However, the above solutions require a lot of time and manpower costs, cannot meet the timeliness requirements of knowledge information, and cannot fully utilize the massive sparse industry data and knowledge existing on the Internet. Therefore, the existing methods of building industry knowledge bases have problems such as large workload, untimely information updates and narrow knowledge coverage.
[0004] Therefore, those skilled in the art provide a method and system for constructing an automotive industry knowledge base based on a large language model to solve the problems raised in the above background technology. Summary of the invention
[0005] The purpose of the present invention is to provide a method and system for constructing an automobile industry knowledge base based on a large language model to solve the problems raised in the above background technology.
[0006] To achieve the above object, the present invention provides the following technical solutions:
[0007] A method and system for constructing an automobile industry knowledge base based on a large language model, comprising the following steps:
[0008] Step S101: Collect a large amount of original unstructured and unlabeled data from a wide range of sources such as public encyclopedias, news and finance, knowledge question-and-answer websites, and automobile forums through the Internet, including source files of types such as text data, documents, and reports, and store the data in an object storage system, where each source data corresponds to an object unique identifier and related metadata;
[0009] Step S102: Sampling the original data in small proportion, wherein the sampling standard is that the amount of data from different sources is not less than 100, and then the labeling personnel label the sampled data;
[0010] Step S103, in order to reduce the cost of manual labeling, a large language model is used to label the sample data, which specifically includes the following contents:
[0011] (1) Sampling the original data of the automobile industry obtained in step S101, wherein the sampling ratio is not higher than 1%;
[0012] (2) constructing sample data, where the source of the sample data is the manually annotated data in S102;
[0013] (3) To ensure the quality of the generated annotation information, the present invention selects a large language model with a large number of parameters and strong reasoning ability, and uses the sample data in content (2) as the initial prompt word, and uses the large language model to reason and annotate the data in (1). The content and format of the annotation are the same as those in step S102;
[0014] Step S104, using the data annotated in step S102 and step S103, according to the multiple annotated label categories in step S102, construct training data sets for different tasks, including text classification data sets, text keyword extraction data sets, text hierarchical and normalized data sets, etc.;
[0015] Step S105, using the multiple training data sets constructed in step S104, peft fine-tuning training is performed on the model with small parameter quantity, general performance, but less required resources, and faster reasoning speed, and a multi-task model of MoE (Mixture of Experts) architecture is constructed. The peft fine-tuning method used in the present invention is LoRA fine-tuning, which is characterized by comprising the following contents:
[0016] When the weight matrix of the pre-trained model and its update have low intrinsic dimensions, they are approximated by low-rank decomposition, that is, for the weight matrix W0∈R of the pre-trained model d×k , initialize the low-rank matrix A∈R with a random Gaussian distribution r×k and a low-rank matrix B∈R initialized with a zero matrix d×r To express its update ΔW, it can be expressed by the following formula, namely:
[0017] W0+ΔW=W0+BA
[0018] Among them, the dimensions of B and A are much smaller than the dimension of W0. During the fine-tuning process, the weight W0 of the fixed and trained model is unchanged, and the trained low-rank matrices A and B are injected into each layer of the Transformer architecture to fine-tune the Self-Attention part. This fine-tuning method can greatly reduce the amount of training parameters;
[0019] After the model fine-tuning is completed, the pre-trained weights of the model are merged with the fine-tuned weights, and the merged model file is saved;
[0020] Step S106, using the MoE model fine-tuned in step S105, annotating the unsampled original data according to different task types, wherein the annotated content and format of the data are the same as those in step S102;
[0021] Step S107, vectorizing the labeled hierarchical data, which is characterized by including the following contents:
[0022] The text sequence related to the automotive industry is input into the text2vec model, and then the text sequence is divided into multiple windows. A context vector is generated for each window, and the generated context vectors are weighted averaged to obtain the vector representation of the text sequence object related to the automotive industry.
[0023] Step S108, writing the vectorized data of the automobile industry obtained in step S107 into the vector knowledge base;
[0024] Step S109: To ensure the timeliness of data information, a real-time update strategy is adopted. Based on the Flink task and the Kafka data pipeline, real-time streaming collection and extraction processing is performed on data from a wide range of sources. Then, real-time annotation and text vectorization are performed through the process of steps S106, S107, and S108, and the data is added or updated to the automotive industry vector knowledge base.
[0025] Furthermore, the method further comprises step S110, a knowledge retrieval module, which is used to query according to the user's natural language, and is characterized in that it comprises the following contents:
[0026] (1) Vectorize the user's question and compare the similarity in the vector knowledge base, and use the L2 distance to calculate the five knowledge blocks with the highest similarity;
[0027] (2) Only text segments with content higher than the threshold set by the user are taken. If any, they are concatenated using natural language. The concatenated content is then input into the large language model for inference, and the model's inference content is returned to the user. Otherwise, the question is directly input into the large language model for question and answering.
[0028] Preferably, the annotation information in step S102 includes the following content:
[0029] (1) An overview of the data content;
[0030] (2) Based on whether the data is completely related to the information or knowledge of the automotive industry, possibly related to the automotive industry or has a certain relationship with the automotive industry, or has no relationship with the automotive industry, the data is labeled with three levels of relevance: relevant, possibly related, or irrelevant;
[0031] (3) If the degree of relevance of the data to the automotive industry is related as in content (2), it is necessary to further annotate the data in the fine-grained subcategories under the automotive industry category, including automobile brand, automobile operation, automobile maintenance, automobile insurance, road traffic regulations, etc.;
[0032] (4) The categories are marked in the format required by the Chain of Thought (CoT) technology, and the basis for the classification is further marked;
[0033] (5) If the relevance of the data to the automotive industry is as described in (2), the data should be further annotated with a list of keywords in the automotive industry;
[0034] (6) If the degree of relevance of the data to the automotive industry is as described in content (2), the data needs to be further organized and layered, and the expression needs to be rewritten in a standardized manner.
[0035] Compared with the prior art, the present invention has the following beneficial effects:
[0036] Based on the large language model, the present invention efficiently extracts and integrates the knowledge of the automotive industry from massive text data, achieving the comprehensiveness and diversity of knowledge; secondly, the system improves the accuracy of information retrieval in the existing automotive industry knowledge base, and the large language model question-answering service based on the system can provide users with accurate and detailed information answers. In addition, the semantic understanding and generation capabilities of the large language model help to explain and express automotive industry knowledge with higher coherence and accuracy; building an automotive industry knowledge base not only promotes the dissemination and sharing of knowledge in the automotive industry, but also provides decision makers with information in professional fields to support smart decision-making. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 It is a flow chart of a method and system for constructing an automobile industry knowledge base based on a large language model of the present invention. DETAILED DESCRIPTION
[0038] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0039] See also Figure 1 In an embodiment of the present invention, a method and system for constructing an automobile industry knowledge base based on a large language model, a method and system for constructing an automobile industry knowledge base based on a large language model, comprises the following steps:
[0040] Step S101: Collect a large amount of original unstructured and unlabeled data from a wide range of sources such as public encyclopedias, news and finance, knowledge question-and-answer websites, and automobile forums through the Internet, including source files of types such as text data, documents, and reports, and store the data in an object storage system, where each source data corresponds to an object unique identifier and related metadata;
[0041] Step S102: Sampling the original data in small proportion, wherein the sampling standard is that the amount of data from different sources is not less than 100, and then the labeling personnel label the sampled data;
[0042] Step S103, in order to reduce the cost of manual labeling, a large language model is used to label the sample data, which specifically includes the following contents:
[0043] (1) Sampling the original data of the automobile industry obtained in step S101, wherein the sampling ratio is not higher than 1%;
[0044] (2) constructing sample data, where the source of the sample data is the manually annotated data in S102;
[0045] (3) To ensure the quality of the generated annotation information, the present invention selects a large language model with a large number of parameters and strong reasoning ability, and uses the sample data in content (2) as the initial prompt word, and uses the large language model to reason and annotate the data in (1). The content and format of the annotation are the same as those in step S102;
[0046] Step S104, using the data annotated in step S102 and step S103, according to the multiple annotated label categories in step S102, construct training data sets for different tasks, including text classification data sets, text keyword extraction data sets, text hierarchical and normalized data sets, etc.;
[0047] Step S105, using the multiple training data sets constructed in step S104, peft fine-tuning training is performed on the model with small parameter quantity, general performance, but less required resources, and faster reasoning speed, and a multi-task model of MoE (Mixture of Experts) architecture is constructed. The peft fine-tuning method used in the present invention is LoRA fine-tuning, which is characterized by comprising the following contents:
[0048] When the weight matrix of the pre-trained model and its update have low intrinsic dimensions, they are approximated by low-rank decomposition, that is, for the weight matrix W0∈R of the pre-trained model d×k , initialize the low-rank matrix A∈R with a random Gaussian distribution r×k and a low-rank matrix B∈R initialized with a zero matrix d×r To express its update ΔW, it can be expressed by the following formula, namely:
[0049] W0+ΔW=W0+BA
[0050] Among them, the dimensions of B and A are much smaller than the dimension of W0. During the fine-tuning process, the weight W0 of the fixed and trained model is unchanged, and the trained low-rank matrices A and B are injected into each layer of the Transformer architecture to fine-tune the Self-Attention part. This fine-tuning method can greatly reduce the amount of training parameters;
[0051] After the model fine-tuning is completed, the pre-trained weights of the model are merged with the fine-tuned weights, and the merged model file is saved;
[0052] Step S106, using the MoE model fine-tuned in step S105, annotating the unsampled original data according to different task types, wherein the annotated content and format of the data are the same as those in step S102;
[0053] Step S107, vectorizing the labeled hierarchical data, which is characterized by including the following contents:
[0054] The text sequence related to the automotive industry is input into the text2vec model, and then the text sequence is divided into multiple windows. A context vector is generated for each window, and the generated context vectors are weighted averaged to obtain the vector representation of the text sequence object related to the automotive industry.
[0055] Step S108, writing the vectorized data of the automobile industry obtained in step S107 into the vector knowledge base;
[0056] Step S109: To ensure the timeliness of data information, a real-time update strategy is adopted. Based on the Flink task and the Kafka data pipeline, real-time streaming collection and extraction of data from a wide range of sources are performed. Then, real-time annotation and text vectorization are performed through the process of steps S106, S107, and S108, and the data is added or updated to the automotive industry vector knowledge base.
[0057] By adopting the above technical solutions, based on the large language model, the knowledge of the automotive industry can be efficiently extracted and integrated from massive text data, achieving the comprehensiveness and diversity of knowledge; secondly, the system improves the accuracy of information retrieval in the existing automotive industry knowledge base, and the large language model question-answering service based on the system can provide users with accurate and detailed information answers. In addition, the semantic understanding and generation capabilities of the large language model help to explain and express automotive industry knowledge with higher coherence and accuracy; building an automotive industry knowledge base not only promotes the dissemination and sharing of knowledge in the automotive industry, but also provides decision makers with information in professional fields to support smart decision-making.
[0058] Furthermore, the method further comprises step S110, a knowledge retrieval module, which is used to query according to the user's natural language, and is characterized in that it comprises the following contents:
[0059] (1) Vectorize the user's question and compare the similarity in the vector knowledge base, and use the L2 distance to calculate the five knowledge blocks with the highest similarity;
[0060] (2) Only text segments with content higher than the threshold set by the user are taken. If any, they are concatenated using natural language. The concatenated content is then input into the large language model for inference, and the model's inference content is returned to the user. Otherwise, the question is directly input into the large language model for question and answering.
[0061] In this embodiment, the annotation information in step S102 includes the following content:
[0062] (1) An overview of the data content;
[0063] (2) Based on whether the data is completely related to the information or knowledge of the automotive industry, possibly related to the automotive industry or has a certain relationship with the automotive industry, or has no relationship with the automotive industry, the data is labeled with three levels of relevance: relevant, possibly related, or irrelevant;
[0064] (3) If the degree of relevance of the data to the automotive industry is related as in content (2), it is necessary to further annotate the data in the fine-grained subcategories under the automotive industry category, including automobile brand, automobile operation, automobile maintenance, automobile insurance, road traffic regulations, etc.;
[0065] (4) The categories are marked in the format required by the Chain of Thought (CoT) technology, and the basis for the classification is further marked;
[0066] (5) If the relevance of the data to the automotive industry is as described in (2), the data should be further annotated with a list of keywords in the automotive industry;
[0067] (6) If the degree of relevance of the data to the automotive industry is as described in content (2), the data needs to be further organized and layered, and the expression needs to be rewritten in a standardized manner.
[0068] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical scheme and inventive concept of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.
Claims
1. A method and system for constructing an automobile industry knowledge base based on a large language model, characterized in that: The steps include: Step S101: Collect a large amount of original unstructured and unlabeled data from a wide range of sources such as public encyclopedias, news and finance, knowledge question-and-answer websites, and automobile forums on the Internet, including text data, documents, and source files of report type, and store the data in an object storage system, where each source data corresponds to an object unique identifier and related metadata; Step S102: Sampling the original data in small proportion, wherein the sampling standard is that the amount of data from different sources is not less than 100, and then the labeling personnel label the sampled data; Step S103, using a large language model to complete the labeling of sample data, specifically includes the following contents: (1) Sampling the original data of the automobile industry obtained in step S101, wherein the sampling ratio is not higher than 1%; (2) constructing sample data, where the source of the sample data is the manually annotated data in S102; (3) To ensure the quality of the generated annotation information, the present invention selects a large language model with a large number of parameters and strong reasoning ability, and uses the sample data in content (2) as the initial prompt word, and uses the large language model to reason and annotate the data in (1). The content and format of the annotation are the same as those in step S102; Step S104, using the data annotated in step S102 and step S103, according to the multiple annotated label categories in step S102, construct training data sets for different tasks, including a text classification data set, a text keyword extraction data set, and a text hierarchical and normalized data set; Step S105, using the multiple training data sets constructed in step S104, peft fine-tuning training is performed on the model with small parameter quantity, general performance, but less required resources, and faster reasoning speed, and a multi-task model of MoE architecture is constructed. The peft fine-tuning method used in the present invention is LoRA fine-tuning, which is characterized by comprising the following contents: When the weight matrix of the pre-trained model and its update have low intrinsic dimensions, they are approximated by low-rank decomposition, that is, for the weight matrix W0∈R of the pre-trained model d×k , initialize the low-rank matrix A∈R with a random Gaussian distribution r×k and a low-rank matrix B∈R initialized with a zero matrix d×r To express its update ΔW, it can be expressed by the following formula, namely: W0+ΔW=W0+BA Among them, the dimensions of B and A are much smaller than the dimension of W0. During the fine-tuning process, the weight W0 of the fixed and trained model is unchanged, and the trained low-rank matrices A and B are injected into each layer of the Transformer architecture to fine-tune the Self-Attention part. This fine-tuning method can greatly reduce the amount of training parameters; After the model fine-tuning is completed, the pre-trained weights of the model are merged with the fine-tuned weights, and the merged model file is saved; Step S106, using the MoE model fine-tuned in step S105, annotating the unsampled original data according to different task types, wherein the annotated content and format of the data are the same as those in step S102; Step S107, vectorizing the labeled hierarchical data, which is characterized by including the following contents: The text sequence related to the automotive industry is input into the text2vec model, and then the text sequence is divided into multiple windows. A context vector is generated for each window, and the generated context vectors are weighted averaged to obtain the vector representation of the text sequence object related to the automotive industry. Step S108, writing the vectorized data of the automobile industry obtained in step S107 into the vector knowledge base; Step S109: To ensure the timeliness of data information, a real-time update strategy is adopted. Based on the Flink task and the Kafka data pipeline, real-time streaming collection and extraction of data from a wide range of sources are performed. Then, through the process of steps S106, S107, and S108, real-time annotation and text vectorization are performed, and the data is added or updated to the automotive industry vector knowledge base.
2. The method and system for constructing an automobile industry knowledge base based on a large language model according to claim 1, characterized in that: The method further includes step S110, a knowledge retrieval module, which is used to query according to the user's natural language, and is characterized by comprising the following contents: (1) Vectorize the user's question and compare the similarity in the vector knowledge base, and use the L2 distance to calculate the five knowledge blocks with the highest similarity; (2) Only text segments with content higher than the threshold set by the user are taken. If any, they are concatenated using natural language. The concatenated content is then input into the large language model for inference, and the model's inference content is returned to the user. Otherwise, the question is directly input into the large language model for question and answering.
3. The method and system for constructing an automobile industry knowledge base based on a large language model according to claim 1, characterized in that: The annotation information in step S102 includes the following contents: (1) An overview of the data content; (2) Based on whether the data is completely related to the information or knowledge of the automotive industry, possibly related to the automotive industry or has a certain relationship with the automotive industry, or has no relationship with the automotive industry, the data is labeled with three levels of relevance: relevant, possibly related, or irrelevant; (3) If the degree of relevance of the data to the automotive industry is related as in content (2), it is necessary to further annotate the data in the fine-grained subcategories under the automotive industry category, including automobile brand, automobile operation, automobile maintenance, automobile insurance, and road traffic regulations; (4) The categories are marked in the format required by the thinking chain technology, and the basis for the classification is further marked; (5) If the relevance of the data to the automotive industry is as described in (2), the data should be further annotated with a list of keywords in the automotive industry; (6) If the degree of relevance of the data to the automotive industry is as described in content (2), the data needs to be further organized and layered, and the expression needs to be rewritten in a standardized manner.
Citation Information
Cited By
Automobile part name matching method, device and equipment based on large model
CN121542410A
An automobile accessory name matching method, device and equipment based on a large model
CN121542410B