Key technology identification and IPC classification method and system based on large language model
Through incremental pre-training and supervised fine-tuning technology, combined with domain data processing and multi-agent collaborative reasoning framework, the efficiency and accuracy problems of key technology identification and IPC classification in existing technologies are solved, and efficient and accurate identification and classification in specific fields are achieved.
Patent Information
- Application Number
- CN202510305882.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-09-23
AI Technical Summary
Existing methods have problems in key technology identification and IPC classification, such as being time-consuming and labor-intensive, insufficient semantic understanding, lack of domain knowledge, and insufficient adaptation of general models to specific fields, resulting in incomplete or inaccurate identification.
Through incremental pre-training and supervised fine-tuning technology, combined with domain data for data preprocessing, a structured domain corpus is constructed, the byte pair encoding algorithm is used to generate BPE merging rules, low-rank adaptation technology is used to inject domain knowledge into large language models, and a multi-agent collaborative reasoning framework is introduced to improve the model's recognition and IPC classification capabilities in specific fields.
It significantly improves the accuracy and efficiency of the model's recognition and IPC classification in specific fields, reduces training resource consumption, and provides a more efficient and accurate key technology identification and IPC classification solution.
Smart Images

Figure CN120687865A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of natural language processing, and in particular to a method and system for key technology identification and IPC classification based on a large language model, which is used for key technology identification and international patent classification IPC identification in any specified field. Background Art
[0002] In today's era of increasingly fierce global scientific and technological competition, identifying and mastering key core technologies is of great significance. Key core technologies are the core support for a country's scientific and technological strength and industrial competitiveness, and they directly determine a country's position and voice in the global industrial chain. For example, in the field of semiconductor chips, high-end chip manufacturing technology is a key core technology.
[0003] Identifying key core technologies can help countries and businesses clarify R&D priorities and resource allocation. By precisely identifying these technologies, superior resources can be concentrated on key breakthroughs, avoiding blind and fragmented technological deployment. Furthermore, identifying key core technologies facilitates preemptive intellectual property protection, building technological barriers, and preventing technologies from being easily copied and surpassed. Furthermore, breakthroughs in key core technologies often drive the clustered development of related industries, forming new economic growth points and promoting the transformation and upgrading of economic structures. Therefore, only by accurately identifying and mastering key core technologies can we seize the initiative in global scientific and technological competition. However, despite extensive research and exploration into key core technology identification in academia and industry, existing methods still face numerous challenges in terms of timeliness, universality, and accuracy.
[0004] Traditional methods for identifying key technologies rely primarily on qualitative methods such as expert evaluation and literature analysis, as well as quantitative methods based on big data processing and machine learning algorithms. While these methods can reveal the development trends and potential impact of technologies to a certain extent, they often suffer from issues such as being time-consuming and labor-intensive, and lacking semantic understanding. For example, the "Delphi method" relies on the judgment and summary of domain experts, which is subjective and costly. Quantitative methods based on models such as LDA / Bert can be used to extract information from large amounts of patent documents, but due to the "semantic gap" they present, their precision and accuracy still need to be improved. Furthermore, their scalability is poor, and given a new domain, the entire process of data processing and model training and inference needs to be repeated, which is time-consuming.
[0005] With the rapid development of large-scale AI models, a number of general-purpose models, such as GPT-4 and Tongyi Qianwen, have emerged, demonstrating impressive capabilities in a wide range of language understanding and generation. However, these general-purpose models still have significant limitations when handling specific tasks in specialized fields, particularly in understanding the meaning of IPC classifications and identifying key technical points. These models often lack in-depth optimization for specific fields, resulting in incomplete or inaccurate identification of key technologies, which in turn impacts the efficiency and accuracy of patent classification.
[0006] Therefore, there is an urgent need for a method that can effectively combine domain knowledge and improve the performance of general models in domain key technology identification and IPC classification tasks. Summary of the Invention
[0007] The purpose of the invention is to address the above technical problems and propose a method and system for key technology identification and IPC classification based on a large language model. Through incremental pre-training and supervised fine-tuning technology, the performance of the general model in key technology identification and IPC classification tasks in specific fields is improved, and the problems of existing models in semantic understanding deviation, insufficient task adaptation and lack of domain knowledge are solved. The limitations of existing technology identification methods in terms of timeliness and universality are significantly improved. The purpose of the invention can be achieved through the following technical solutions: The present invention provides a method for key technology identification and IPC classification based on a large language model, comprising the following steps: Step S1: Collect domain data and perform data preprocessing to build a structured domain corpus. The domain corpus includes public patent text data and IPC classifications, multiple IPC classification comparison tables, and general text corpus. Step S2: Based on the Llama-Factory framework, the byte pair encoding algorithm is used to perform word segmentation training on the domain corpus, and the corresponding BPE merging rules are generated. The BPE merging rules are merged with the original BPE merging rules of the Qwen2.5 base model to reconstruct the hybrid word segmenter and generate the corresponding domain vocabulary. At the same time, the domain data is converted into the corresponding pre-training data corpus. Using the pre-training data corpus and the domain vocabulary, the Qwen2.5 base model is incrementally pre-trained by adopting low-rank adaptation technology to inject domain data. Step S3: Construct a supervised fine-tuning training dataset for the specified task based on the domain data, and perform data preprocessing to convert it into a format suitable for SFT unit input. Then, supervised fine-tune the SFT unit through low-rank adaptation technology to make the SFT unit reach the expected performance standard. Step S4: Based on the proposed key technology identification and IPC classification problems, they are assigned to the corresponding query agent or reasoning agent through the task diversion mechanism, and the SFT unit is called to process the problem and output the answer.
[0008] Furthermore, step S1 includes collecting multi-source heterogeneous domain data and performing data preprocessing, including data cleaning, format conversion and semantic analysis, and integrating the preprocessed domain data to form a unified domain corpus.
[0009] Furthermore, in step S2, the domain corpus is trained for word segmentation using a byte pair encoding algorithm to generate corresponding BPE merging rules, which are then merged with the original BPE merging rules of the Qwen2.5 base model to reconstruct the hybrid word segmenter and generate the corresponding domain vocabulary, including: Extract the original vocabulary and BPE merging rules from the Qwen2.5 base model, and retain the subword units and their corresponding identifiers in the original Qwen2.5 base model; Use the byte pair encoding algorithm to perform word segmentation training on the domain corpus in step S1, iteratively count and merge high-frequency character pairs in the corpus to form BPE merging rules for specific fields; The formed BPE merging rule is superimposed on the original BPE merging rule, and the new domain vocabulary is appended to the end of the original vocabulary for fusion, and a unique identifier is assigned to form the corresponding domain vocabulary; Based on the fused domain vocabulary and BPE merging rules, a corresponding hybrid word segmenter is constructed to process and analyze general and domain-specific text data.
[0010] Furthermore, in step S2, the domain data is converted into corresponding pre-training data corpus, including: Based on the classified general text corpus, categories associated with specific fields are extracted and used as pre-training data corpus; Convert patent text data into natural text and combine it with the extracted patent number information to assemble the patent information into a complete pre-training data corpus; the patent number information includes the patent number, invention name and IPC classification; Based on the hierarchical relationship information in multiple IPC classification comparison tables, the hierarchical relationship information is converted into descriptive language that can be parsed by a large language model and used as pre-training data corpus.
[0011] Furthermore, in step S2, the Qwen2.5 base model is incrementally pre-trained using the pre-training data corpus and the domain vocabulary by adopting a low-rank adaptation technique to inject domain data, including: The pre-training data corpus is divided into a training set and a validation set according to a preset ratio. The training set is used for incremental pre-training of the Qwen2.5 base model, and the validation set is used to monitor the performance of the incremental pre-training process. The pre-trained weights of the Qwen2.5 base model are loaded, and the model parameters are fine-tuned using low-rank adaptation techniques. The Qwen2.5 base model is configured by inserting a trainable rank factorization matrix into the query and key matrices of the Transformer layer, with the rank dimension set to 64, and the number of new parameters is controlled within 0.3% of the original model. At the same time, more than 90% of the underlying parameters in the Qwen2.5 base model are frozen to retain its general language understanding capabilities, and only the LoRA adapter is trained; Pre-training data corpus was introduced according to preset standards, and small batch training and gradient accumulation techniques were used to balance memory usage and efficiency. At the same time, the cosine annealing strategy was used to adjust the learning rate to ensure the convergence of the Qwen2.5 base model. Based on the fully trained Qwen2.5 base model, the LoRA adapter weights and pre-trained weights are merged to export the complete model file.
[0012] Furthermore, in step S3, based on the proposed key technology identification and IPC classification problems, the tasks are assigned to the corresponding query agent or reasoning agent through the task offloading mechanism, and the SFT unit is called to process the problem and output the answer, including: Based on the collected domain data, a supervised fine-tuning training set for key technology identification and IPC classification is constructed, and data enhancement is performed. For structured domain data, prompt words are set and a large language model is used to convert IPC classification questions into question-answering format. For unstructured domain data, technical features in patent text data are extracted and converted into key technology identification tasks. Enable the supervised fine-tuning mode in the SFT unit and use low-rank adaptation technology to fine-tune the model parameters. Based on the Llama-Factory framework, the weighted multi-task loss function is extended to implement: in, is the cross entropy loss function, which is used to optimize the IPC classification task; It is the KL divergence loss, which constrains the similarity of probability distribution in feature extraction tasks; is the L2 regularization term; Based on the trained SFT unit, the LoRA adapter weights and the Qwen2.5 base model weights are merged to export the complete model file.
[0013] Furthermore, in step S4, based on the proposed key technology identification and IPC classification problems, the tasks are assigned to the corresponding query agent or reasoning agent through the task offloading mechanism, and the SFT unit is called to process the problem and output the answer, including: Based on the trained SFT unit, a corresponding multi-agent collaborative reasoning unit is constructed. The multi-agent collaborative reasoning unit includes a query agent and a reasoning agent. Combine the Bert-based classification prediction model to analyze and judge the input questions; When the question is simple, it is assigned to the query agent after keyword matching and semantic analysis, and the SFT unit is called to output the answer; When the problem is a complex problem, the reasoning agent breaks down the complex problem into several sub-problems, and calls the SFT unit to identify the corresponding key technologies and match the corresponding IPC classification.
[0014] Based on the same inventive concept, the present invention provides a system for key technology identification and IPC classification based on a large language model, which adopts the above-mentioned method for key technology identification and IPC classification based on a large language model, including: The data acquisition module is used to collect domain data and perform data preprocessing to build a structured domain corpus. The domain corpus includes public patent text data and IPC classifications, multiple IPC classification comparison tables, and general text corpus. The incremental pre-training module is used to perform word segmentation training on the domain corpus using the byte pair encoding algorithm based on the Llama-Factory framework, generate corresponding BPE merging rules, merge the BPE merging rules with the original BPE merging rules of the Qwen2.5 base model to reconstruct the hybrid word segmenter and generate the corresponding domain vocabulary; at the same time, the domain data is converted into the corresponding pre-training data corpus; using the pre-training data corpus and the domain vocabulary, the Qwen2.5 base model is incrementally pre-trained by adopting low-rank adaptation technology to inject domain data; The supervised fine-tuning module is used to construct a supervised fine-tuning training dataset for a specified task based on domain data, and perform data preprocessing to convert it into a format suitable for SFT unit input. The SFT unit is supervised and fine-tuned through low-rank adaptation technology to ensure that the SFT unit reaches the expected performance standard. Result output module: Based on the proposed key technology identification and IPC classification problems, the task is assigned to the corresponding query agent or reasoning agent through the task diversion mechanism, and the problem is processed and the answer is output by calling the SFT unit.
[0015] Furthermore, the incremental pre-training module includes, The pre-training data corpus processing unit is used to extract categories associated with specific fields based on classified general text corpus and use them as pre-training data corpus; convert patent text data into natural text and combine it with the extracted patent number information to assemble a complete pre-training data corpus.
[0016] Furthermore, the patent number information includes the patent number, invention name and IPC classification; based on the hierarchical relationship information in multiple IPC classification comparison relationship tables, the hierarchical relationship information is converted into a descriptive language parsed by a large language model and used as a pre-training data corpus.
[0017] Compared with the prior art, the present invention has at least one of the following technical effects: The present invention effectively integrates and processes patent text data and IPC classifications, multiple IPC classification comparison tables and general text corpora, uses pre-trained data corpora and domain vocabulary, and adopts low-rank adaptation technology to incrementally pre-train the Qwen2.5 base model to inject domain data, so that the model has a deeper and more accurate understanding of specific fields, especially IPC classifications, and overcomes the problem of IPC meaning recognition errors in other general large models. While consuming low training resources, it greatly improves the basic capabilities of the model in professional fields.
[0018] The multi-agent collaborative reasoning unit improves the reasoning ability and accuracy in specific fields through the division of labor and cooperation among agents, enhances the robustness of the overall system, effectively solves the problem of insufficient task adaptability of existing general large models, and provides a more efficient and accurate solution for patent classification and key technology identification.
[0019] Based on learning and training of large-scale general corpus and patent data, this method has outstanding universality in the problem of key technology identification. It can efficiently output the key core technologies and IPC classification numbers in any specified field without the need to repeat tasks such as expert interviews, data labeling, model training and fitting in the specified field like existing methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments: Figure 1 This is a flowchart of the method steps for key technology identification and IPC classification based on a large language model of the present invention; Figure 2 This is an example flow chart of key technologies in the field of new energy vehicle batteries and their IPC classification in an embodiment of the present invention. DETAILED DESCRIPTION
[0021] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0022] First embodiment The present invention provides a method for key technology identification and IPC classification based on a large language model. Through incremental pre-training and supervised fine-tuning technology, the performance of the general model in key technology identification and IPC classification tasks in specific fields is improved, and the problems of existing models in semantic understanding deviation, task adaptation deficiency and lack of domain knowledge are solved. In addition, the limitations of existing technology identification methods in terms of timeliness and universality are significantly improved. Figure 1 As shown, the steps include: Step S1: Collect domain data and perform data preprocessing to build a structured domain corpus. The domain corpus includes public patent text data and IPC classifications, multiple IPC classification comparison tables, and general text corpus. Step S2: Based on the Llama-Factory framework, the byte pair encoding algorithm is used to perform word segmentation training on the domain corpus, and the corresponding BPE merging rules are generated. The BPE merging rules are merged with the original BPE merging rules of the Qwen2.5 base model to reconstruct the hybrid word segmenter and generate the corresponding domain vocabulary. At the same time, the domain data is converted into the corresponding pre-training data corpus. Using the pre-training data corpus and the domain vocabulary, the Qwen2.5 base model is incrementally pre-trained by adopting low-rank adaptation technology to inject domain data. Step S3: Construct a supervised fine-tuning training dataset for the specified task based on the domain data, and perform data preprocessing to convert it into a format suitable for SFT unit input. Then, supervised fine-tune the SFT unit through low-rank adaptation technology to make the SFT unit reach the expected performance standard. Step S4: Based on the proposed key technology identification and IPC classification problems, they are assigned to the corresponding query agent or reasoning agent through the task diversion mechanism, and the SFT unit is called to process the problem and output the answer.
[0023] Furthermore, step S1 includes collecting multi-source heterogeneous domain data and performing data preprocessing, including data cleaning, format conversion and semantic analysis, and integrating the preprocessed domain data to form a unified domain corpus.
[0024] Furthermore, in step S2, the domain corpus is trained for word segmentation using a byte pair encoding algorithm to generate corresponding BPE merging rules, which are then merged with the original BPE merging rules of the Qwen2.5 base model to reconstruct the hybrid word segmenter and generate the corresponding domain vocabulary, including: Extract the original vocabulary and BPE merging rules from the Qwen2.5 base model, and retain the subword units and their corresponding identifiers in the original Qwen2.5 base model; Use the byte pair encoding algorithm to perform word segmentation training on the domain corpus in step S1, iteratively count and merge high-frequency character pairs in the corpus to form BPE merging rules for specific fields; The formed BPE merging rule is superimposed on the original BPE merging rule, and the new domain vocabulary is appended to the end of the original vocabulary for fusion, and a unique identifier is assigned to form the corresponding domain vocabulary; Based on the fused domain vocabulary and BPE merging rules, a corresponding hybrid word segmenter is constructed to process and analyze general and domain-specific text data.
[0025] Furthermore, in step S2, the domain data is converted into corresponding pre-training data corpus, including: Based on the classified general text corpus, categories associated with specific fields are extracted and used as pre-training data corpus; Convert patent text data into natural text and combine it with the extracted patent number information to assemble the patent information into a complete pre-training data corpus; the patent number information includes the patent number, invention name and IPC classification; Based on the hierarchical relationship information in multiple IPC classification comparison tables, the hierarchical relationship information is converted into descriptive language that can be parsed by a large language model and used as pre-training data corpus.
[0026] Furthermore, in step S2, the Qwen2.5 base model is incrementally pre-trained using the pre-training data corpus and the domain vocabulary by adopting a low-rank adaptation technique to inject domain data, including: The pre-training data corpus is divided into a training set and a validation set according to a preset ratio. The training set is used for incremental pre-training of the Qwen2.5 base model, and the validation set is used to monitor the performance of the incremental pre-training process. The pre-trained weights of the Qwen2.5 base model are loaded, and the model parameters are fine-tuned using low-rank adaptation techniques. The Qwen2.5 base model is configured by inserting a trainable rank factorization matrix into the query and key matrices of the Transformer layer, with the rank dimension set to 64, and the number of new parameters is controlled within 0.3% of the original model. At the same time, more than 90% of the underlying parameters in the Qwen2.5 base model are frozen to retain its general language understanding capabilities, and only the LoRA adapter is trained; Pre-training data corpus was introduced according to preset standards, and small batch training and gradient accumulation techniques were used to balance memory usage and efficiency. At the same time, the cosine annealing strategy was used to adjust the learning rate to ensure the convergence of the Qwen2.5 base model. Based on the fully trained Qwen2.5 base model, the LoRA adapter weights and pre-trained weights are merged to export the complete model file.
[0027] Furthermore, in step S3, based on the proposed key technology identification and IPC classification problems, the tasks are assigned to the corresponding query agent or reasoning agent through the task offloading mechanism, and the SFT unit is called to process the problem and output the answer, including: Based on the collected domain data, a supervised fine-tuning training set for key technology identification and IPC classification is constructed, and data enhancement is performed. For structured domain data, prompt words are set and a large language model is used to convert IPC classification questions into question-answering format. For unstructured domain data, technical features in patent text data are extracted and converted into key technology identification tasks. Enable the supervised fine-tuning mode in the SFT unit and use low-rank adaptation technology to fine-tune the model parameters. Based on the Llama-Factory framework, the weighted multi-task loss function is extended to implement: Among them, is the cross entropy loss function, which is used to optimize the IPC classification task; is the KL divergence loss, which constrains the similarity of probability distribution in the technical feature extraction task; is the L2 regularization term; Based on the trained SFT unit, the LoRA adapter weights and the Qwen2.5 base model weights are merged to export the complete model file.
[0028] Furthermore, in step S4, based on the proposed key technology identification and IPC classification problems, the tasks are assigned to the corresponding query agent or reasoning agent through the task offloading mechanism, and the SFT unit is called to process the problem and output the answer, including: Based on the trained SFT unit, a corresponding multi-agent collaborative reasoning unit is constructed. The multi-agent collaborative reasoning unit includes a query agent and a reasoning agent. Combine the Bert-based classification prediction model to analyze and judge the input questions; When the question is simple, it is assigned to the query agent after keyword matching and semantic analysis, and the SFT unit is called to output the answer; When the problem is a complex problem, the reasoning agent breaks down the complex problem into several sub-problems, and calls the SFT unit to identify the corresponding key technologies and match the corresponding IPC classification.
[0029] Second embodiment Based on the same inventive concept, the present invention also provides a system for key technology identification and IPC classification based on a large language model, which adopts the above-mentioned method for key technology identification and IPC classification based on a large language model, including: The data acquisition module is used to collect domain data and perform data preprocessing to build a structured domain corpus. The domain corpus includes public patent text data and IPC classifications, multiple IPC classification comparison tables, and general text corpus. The incremental pre-training module is used to perform word segmentation training on the domain corpus using the byte pair encoding algorithm based on the Llama-Factory framework, generate corresponding BPE merging rules, merge the BPE merging rules with the original BPE merging rules of the Qwen2.5 base model to reconstruct the hybrid word segmenter and generate the corresponding domain vocabulary; at the same time, the domain data is converted into the corresponding pre-training data corpus; using the pre-training data corpus and the domain vocabulary, the Qwen2.5 base model is incrementally pre-trained by adopting low-rank adaptation technology to inject domain data; The supervised fine-tuning module is used to construct a supervised fine-tuning training dataset for a specified task based on domain data, and perform data preprocessing to convert it into a format suitable for SFT unit input. The SFT unit is supervised and fine-tuned through low-rank adaptation technology to ensure that the SFT unit reaches the expected performance standard. Result output module: Based on the proposed key technology identification and IPC classification problems, the task is assigned to the corresponding query agent or reasoning agent through the task diversion mechanism, and the problem is processed and the answer is output by calling the SFT unit.
[0030] Furthermore, the incremental pre-training module includes, The pre-training data corpus processing unit is used to extract categories associated with specific fields based on classified general text corpus and use them as pre-training data corpus; convert patent text data into natural text and combine it with the extracted patent number information to assemble a complete pre-training data corpus.
[0031] Furthermore, the patent number information includes the patent number, invention name and IPC classification; based on the hierarchical relationship information in multiple IPC classification comparison relationship tables, the hierarchical relationship information is converted into a descriptive language parsed by a large language model and used as a pre-training data corpus.
[0032] Third embodiment In order to further verify the effectiveness and adaptability of this technology in practical applications, the third embodiment combines patent data from a specific industry and focuses on experiments on key technology identification and IPC classification tasks in specific fields. This embodiment demonstrates the advantages of the present invention in processing data in specific fields and improving model accuracy and efficiency by adopting different field data sets and specific tasks. At the same time, this embodiment also compares the performance differences between traditional methods and this technology when processing similar tasks, further verifying the effectiveness of the present invention. The specific scheme is as follows: This paper proposes a method for identifying key technologies in a field and classifying them into international patents (IPC) based on incremental pre-training and supervised fine-tuning. The core solution includes the following three key steps: (1) Parameter-efficient knowledge injection mechanism. In order to enhance the model's ability to understand domain-specific terminology, especially IPC classification, the present invention adopts the following method for knowledge injection: Expanding the domain dictionary, or domain corpus: Build and expand a domain corpus containing IPC-specific terminology. This domain corpus not only covers a wide range of professional terms but also contains rich technical background information, aiming to improve the model's basic semantic understanding capabilities, enabling it to better capture and parse professional technical terms.
[0033] Low-rank adaptation of LoRA incremental pre-training: Based on existing Chinese base models, such as Qwen2.5, we use the low-rank adaptation of LoRA for incremental pre-training. This approach efficiently introduces new domain knowledge without affecting the overall architecture of the original model, reducing the consumption of training resources. In this way, the model can significantly improve its performance in specific domains while maintaining its original advantages.
[0034] (2) Supervised fine-tuning (SFT) strategy for scene adaptation. In order to make the model more accurately adapt to IPC classification and other key technology identification tasks, this paper introduces a domain-oriented supervised fine-tuning strategy: We built a domain-specific dataset, known as a supervised fine-tuning training dataset, for IPC classification and key technology identification tasks, encompassing a rich set of conversational question-answering scenarios. This supervised fine-tuning training dataset not only covers explanations of basic concepts but also delves into problem analysis in complex application scenarios, ensuring the model can fully address technical consulting needs at all levels. This provides a learning resource for the model.
[0035] LoRA-based SFT fine-tuning: Based on the high-quality supervised fine-tuning training dataset described above, the model is fine-tuned using a low-rank adaptation method. This effectively mitigates the catastrophic forgetting problem that can occur with traditional fine-tuning, while significantly reducing computational cost and time. This approach allows the model to efficiently adjust parameters to better adapt to new tasks while maintaining the stability of the original architecture.
[0036] (3) Multi-agent collaborative reasoning framework To further enhance the model's reasoning capabilities and accuracy in specific domains, this paper introduces a multi-agent collaborative reasoning framework, namely a multi-agent collaborative reasoning unit. Multiple agents are designed, each focusing on a specific task or domain knowledge, and performing routing based on classification predictions. Through division of labor and collaboration, the agents complement and support each other, thereby improving the accuracy and robustness of the overall system.
[0037] The specific implementation method is carried out in the following steps: Step 1: Systematically collect domain data To enhance the model's understanding of domain-specific terminology, especially IPC classifications and related technical terms, we systematically collected and processed heterogeneous data from multiple sources. The following are the specific steps: Step 1.1: Data Collection Clarify the list of major data sources involved in this study to ensure comprehensive and representative field data. Major data sources include but are not limited to the following categories: ① Public patent text dataset This large-scale collection of Chinese patent documents, obtained through web crawling and authorized batch downloading from platforms such as Google Patents and Dawei, covers patent applications and authorization texts across multiple technology areas. These patent documents not only include detailed technical solution descriptions, claims, and specifications, but also include complete International Patent Classification (IPC) information and technical background descriptions. The text in the dataset has undergone rigorous cleaning and formatting to ensure data integrity and consistency. By introducing this dataset, the model can better grasp the usage standards of professional terminology and IPC-related knowledge in patent documents, laying a solid foundation for subsequent key technology identification and IPC classification matching tasks. It also filters based on time and patent status to eliminate outdated or low-quality patent documents. This patent dataset contains approximately 1.34 million patents and nearly 10 billion words.
[0038] ②Key core technology related data sets The integration of multiple structured classification information obtained from the websites of relevant authoritative institutions, such as the official website of the State Intellectual Property Office, mainly includes three core components: (1) the International Patent Classification (IPC) 2024.01, which provides a complete patent classification system and its hierarchical relationship; (2) the industrial technology classification and IPC classification comparison table, which establishes a mapping bridge between the field of technological innovation and the patent classification system; (3) the International Patent Classification and National Economic Industry Classification (GB / T 4754-2022) comparison table, which realizes the precise association between patent technology fields and industrial economic sectors. These comparison tables have been strictly verified and manually annotated by domain experts to ensure the accuracy and reliability of classification mapping. By introducing this dataset, we have constructed a multi-dimensional technology classification knowledge system, injecting a comprehensive knowledge encyclopedia from patent technology to industrial application into the model. This dataset not only supports the model to understand the semantic connotation of patent classification, but also enhances its industrial correlation analysis ability in the key technology identification task, which is the prerequisite for the subsequent key technology identification and IPC accurate classification in specific fields.
[0039] ③ Chinese open source general text corpus In addition to the aforementioned domain data, to avoid the loss of general capabilities of large models, we also collected Chinese general corpus data, such as WUDAO2.0, which contains approximately 60 million samples and covers general text corpora in many fields such as encyclopedias, blogs, technology, and education.
[0040] Step 1.2: Data cleaning and format conversion The acquired data is cleaned and formatted to ensure data consistency and usability.
[0041] Data cleaning: Remove duplicates, invalid characters, and noisy data to ensure data integrity and consistency. Consider deleting or filling in reasonable default values for data entries with a high number of missing values. Also, remove patent data that is outdated or has expired.
[0042] Format conversion: Data from different sources is uniformly converted into standardized natural language sentence formats. For example, since crawlers crawl webpage source code, which contains a lot of redundant HTML-formatted information, tools are used to extract only unstructured text such as technical solution summaries, claims, and specifications, and convert them into standard paragraph or sentence formats.
[0043] Semantic parsing: Use natural language processing techniques to extract key information and generate high-quality domain corpora. For example, named entity recognition (NER) can be used to extract technical terms and classification codes from patent documents to form structured metadata.
[0044] Step 1.3: Build the domain corpus The cleaned and format-converted data is integrated into a high-quality domain corpus for subsequent model training.
[0045] Data integration: Merge cleaned data from various data sources to form a unified domain corpus. Ensure that the corpus covers a wide range of technical fields and contains rich technical terminology and classification information.
[0046] Quality Check: Domain experts conduct quality checks on the integrated corpus to ensure data accuracy and completeness. Further data cleaning and correction are performed if necessary.
[0047] Step 2: Domain Knowledge Injection and Incremental Pre-training After completing the systematic collection, cleaning, format conversion, and construction of the domain corpus in step 1, this step will carry out domain knowledge injection based on the Llama-Factory pre-training framework. The specific operations are as follows: Step 2.1: Build a domain vocabulary ①Extract the original vocabulary and word segmentation rules Extract the original vocabulary file and BPE merging rule file from the Qwen2.5 base model, and retain the subword units and their corresponding identifiers in their original word segmentation strategy.
[0048] ② New BPE word segmenter for training domain corpus Using the Byte-Pair Encoding (BPE) algorithm, perform word segmentation training on the domain corpus constructed in Step 1. By iteratively counting high-frequency character pairs in the corpus, a new BPE merging rule set tailored to the specific domain is generated. During this process, the integrity of domain terminology (such as IPC classification numbers and technical keywords) is prioritized to prevent incorrect segmentation of specialized terms.
[0049] ③Integrating new and old vocabulary The newly generated domain BPE merge rules are sequentially superimposed with the merge rules of the original model to ensure the priority of the original common vocabulary. The newly added domain vocabulary is merged to the end of the original vocabulary in an appended manner and assigned a unique token ID (incremented from the maximum ID value in the original vocabulary) to avoid identifier conflicts with the original vocabulary.
[0050] ④Reconstruct the hybrid word segmenter Based on the merged vocabulary and merging rules, the BPE word segmenter was rebuilt to support domain expansion. A dynamic mapping mechanism enabled seamless compatibility between the newly added domain vocabulary and the existing general vocabulary. This process preserved the original model's word segmentation capabilities for general text while significantly enhancing the recognition accuracy of domain-specific terms (such as IPC classification level descriptions and technical feature phrases).
[0051] Step 2.2: Generate pre-training data corpus Each type of data corpus is converted into a pre-training corpus format (json format) suitable for the Llama-Factory pre-training framework, and is divided into three categories according to the data format and processed separately.
[0052] Category 1: Wudao Dataset The Wudao dataset has been categorized into categories such as "encyclopedia," "blog," "entertainment," and "automobiles." Only technology-related content, such as "encyclopedia," "blog," and "automobiles," will be extracted. Determine the categories and the number of categories to use.
[0053] Category 2: Google patent data After the format conversion of the patent crawler data in step 1, the crawled unstructured HTML data has been converted into natural text D and recombined with the extracted patent number and other information. The following is an example: Natural Text D: Abstract:\nThe present invention discloses a thermoelectric module, comprising an N-type thermoelectric material and a P-type thermoelectric material; the N-type thermoelectric material and the P-type thermoelectric material are in the form of filaments, alternately connected head to tail to form a "W"-shaped thermoelectric wire string; the thermoelectric wire strings are arranged at intervals in the longitudinal direction and connected to each other at the ends; the good shape of the thermoelectric wire strings and the more solid foundation are conducive to promoting the application and development of thermoelectric technology.\n\nDescription:\n
[001] Technical field\n
[002] The present invention relates to the field of thermoelectric technology, and in particular to a thermoelectric material module and a method for manufacturing the same.\n
[003] Background technology\n
[004] Energy is the foundation of modern life and development. Electricity is the most widely used and most studied energy and has become a research hotspot in materials science.
[005] Using thermoelectric modules to generate electricity is a new type of power generation method that does not require transmission components and is noiseless and waste-free. Like the use of secondary energy sources such as solar energy, wind energy, and hydropower, it is environmentally friendly. Furthermore, the material is reliable and has a long service life, which is a very practical requirement, leading to the rapid development of thermoelectric materials. Since the 1970s, due to the development of Freon refrigeration technology, research on thermoelectric refrigeration and thermoelectric materials has been neglected and has almost come to a standstill. Since the 1990s, the environmental damage caused by Freon has become widely recognized. The extracted information is as follows: { "Patent Publication Number":"CN101170157A", "Invention Name": "A thermoelectric module and its manufacturing method", "IPC Classification (Public)":"H01L35 / 28;H01L35 / 32;H01L35 / 34" } Assemble the natural text and extracted information, mainly assembling the invention name, IPC classification number and natural text, such as: {"id":"a259c0abd","text":"Patent name: A thermoelectric module and its manufacturing method. The IPC classification numbers involved in this patent include: H01L35 / 28; H01L35 / 32; H01L35 / 34. Abstract: The present invention discloses a thermoelectric module, including N-type thermoelectric material and P-type thermoelectric material; the N-type thermoelectric material and P-type thermoelectric material are in filament form, alternate with each other and connected head to tail in sequence to form a "W"-shaped thermoelectric wire string; the thermoelectric wire strings are arranged at intervals in the longitudinal direction and connected to each other at the ends; the good shape of the thermoelectric wire strings and the more solid foundation are conducive to promoting the application and development of thermoelectric technology. Description:
[001] Technical field\n
[002] The present invention relates to the field of thermoelectric technology, and in particular to a thermoelectric material module and its manufacturing method.\n
[003] Background technology\n
[004] Energy is the foundation of modern life and development. Electric energy is the most widely used and the most studied energy and has become a research hotspot in materials science.\n
[005] Using thermoelectric modules to generate electricity is a new type of power generation method that does not require the use of transmission components, operates silently and produces no waste. Like the use of secondary energy sources such as solar energy, wind energy, and hydropower, it is environmentally friendly. In addition, this material has reliable performance and a long service life, which is a precise requirement and has led to the rapid development of thermoelectric materials. Since the 1970s, due to the development of Freon refrigeration technology, research on thermoelectric refrigeration and thermoelectric materials has been neglected and almost came to a standstill. Since the 1990s, the destructive effects of Freon on the environment have become widely recognized...","source":"Google_Patent","patent_id":"CN101170157A" Category 3: IPC classification number and other structured data Convert key core technology-related data sets, such as the international patent IPC classification table, the industrial technology classification and IPC classification comparison table, and other structured data into natural text.
[0054] Take the international patent IPC classification table data as an example: [A] Necessities of human life [A01] Agriculture; forestry; animal husbandry; hunting; trapping; fishing [A01B] Land preparation for agriculture or forestry; parts, components or accessories of agricultural machinery or implements in general [A01B1 / 00] Hand Tools [A01B1 / 02] shovel; shovel [A01B1 / 04] Toothed [A01B1 / 06] Hoe; Manual Cultivator [A01B1 / 08] With single blade [A01B1 / 10] With double or multiple blades [A01B1 / 12] Toothed blade [A01B1 / 14] With teeth only [A01B1 / 16] Weeding tools [A01B1 / 18] Pliers [A01B1 / 20] Combination of different types of hand tools [A01B1 / 22] Blades or similar working parts fixed to handles; replaceable or adjustable blades [A01B1 / 24] For treating grass or lawn [A01B11 / 00] Ploughs with vibrating, digging or perforating working parts [A01B13 / 00] Ploughs or similar implements for special purposes [A01B13 / 02] For ridging or trimming, such as with symmetrically arranged mouldboards [A01B13 / 04] For use in vineyards, orchards, etc. [A01B13 / 06] Device for preventing damage to grapevines and other crops [A01B13 / 08] For soil preparation [A01B13 / 10] Special equipment for lifting the subsoil layer [A01B13 / 12] Device for spreading subsoil on the ground surface [A01B13 / 14] For cultivating two or more layers of soil [A01B13 / 16] Machinery for controlling soil erosion, such as trough diggers and trenchers [A01B15 / 00] Components, working parts or parts of ploughs [A01B15 / 02] Coulters; fixed coulters [A01B15 / 04] Plowshare ... The International Patent Classification (IPC) is a hierarchical structure, starting with the major class [A] and then breaking it down into smaller classes, such as [A01B1 / 00] Hand tools, and further subclasses, such as [A01B1 / 02] Spades. Each level has a corresponding description.
[0055] The above structured data is converted into natural text suitable for pre-training of large language models. The goal is that the converted text needs to contain rich semantic information and have a clear structure, and the hierarchical relationship should be described in natural language so that the model can fully understand and learn this classification knowledge.
[0056] During the conversion process, the following principles need to be followed: Preserve semantic integrity: Express the complete path of each classification node in natural language. For example, "Under the category of necessities for human life, agriculture, forestry, animal husbandry, hunting, trapping, and fishing belong to subcategory A01. Within subcategory A01, land preparation for agriculture or forestry; components, parts, or accessories of general agricultural machinery or farm tools are classified as group A01B." This method can clearly display hierarchical relationships.
[0057] Avoid ambiguity: Use clear conjunctions and terms such as "belongs to", "divided into", "includes", etc. to clarify hierarchical relationships.
[0058] Information density: While maintaining the fluency of natural language, include as much information as possible, such as classification number, name, superior-subordinate relationship, etc., to enhance the training effect of the model.
[0059] Here is an example of converting the International Patent Classification table into natural text suitable for pre-training a large model: Category A - Necessities of life This broad category covers technical fields directly related to basic human needs, including basic industries such as agriculture, forestry and animal husbandry.
[0060] Subclass A01 - Agriculture; forestry; animal husbandry; hunting; trapping; fishing As the first subcategory of Category A, A01 focuses on the entire agricultural production process technology, including full-chain innovation from soil treatment to crop protection.
[0061] Group A01B - Land preparation for agriculture or forestry Under the A01 subcategory system, A01B specifically deals with land cultivation technology, including both manual agricultural tool design and mechanized land preparation equipment.
[0062] - Group A01B1 / 00 - Hand tools Define the basic agricultural tool type, which is further subdivided into multiple specific tool forms: - A01B1 / 02: Spades, shovels and other flat digging tools - A01B1 / 04: Improved hand tool with teeth - A01B1 / 06: Hoes and manual cultivators for soil loosening - A01B1 / 08: Special design of single blade structure - A01B1 / 10: Dual or multi-blade combination tool - A01B1 / 12: Blade tool with tooth structure - A01B1 / 14: Special tools with pure tooth structure (such as rakes) - A01B1 / 16: Targeted weed removal tool - A01B1 / 18: Pliers-type grabbing implements - A01B1 / 20: Multifunctional combination hand tool - A01B1 / 22: Replaceable parts related to handle connection technology - A01B1 / 24: Special tools for lawn maintenance - Group A01B13 - Special Plough This category covers farming equipment for special scenarios: - A01B13 / 02: Ridging equipment with symmetrical mouldboard - A01B13 / 04: Terrain-adaptable implement for vineyards / orchards - A01B13 / 06: Crop protection devices (e.g., designs to prevent vine damage) - A01B13 / 08: Deep Soil Treatment Technology - A01B13 / 10: Special machinery for subsoil renovation - A01B13 / 12: Soil layering treatment device - A01B13 / 14: Multi-layer soil compound farming system - A01B13 / 16: Soil and water conservation equipment (such as water tank excavators) - Component Classification A01B15 - Plow Parts Define the components of tillage machinery in detail: - A01B15 / 02: Coulter and fixing device - A01B15 / 04: Ploughshare (core component for turning over the soil) This presentation method not only maintains machine-readable structural features but also organizes knowledge in a natural human language manner, which helps large models establish hierarchical cognition of the classification system and understand the technical relationships between various patent fields.
[0063] Key core technologies and IPC classification mapping tables are converted using the above solution.
[0064] Step 2.3: Incremental pre-training After completing the construction of the domain vocabulary and the generation of the pre-training data corpus, this step uses the incremental pre-training technology based on the Llama-Factory framework to inject domain knowledge into the Qwen2.5 base model. The specific implementation steps are as follows: Step 2.3.1: Data partitioning and mixing strategy Dataset partitioning Split the pre-training data corpus generated in step 2.2 into a training set and a validation set at a ratio of 99:1. The training set is used for incremental pre-training of the model, and the validation set is used to monitor model performance during training to prevent overfitting.
[0065] ①Data hybrid strategy To achieve a balance between the model's general language understanding capabilities and domain knowledge, a progressive domain weighting scheme is employed. For example, in the initial stages of training, domain data (e.g., patent texts and IPC classification data) is mixed with general corpus (e.g., the WUDAO 2.0 dataset) in a ratio of 3:7. This phase aims to preserve the model's understanding of general text while initially introducing domain knowledge. In the mid-term, the ratio is gradually adjusted to 5:5, enhancing the model's adaptability to domain terminology and semantics. Ultimately, the ratio is increased to 8:2, with domain data taking the lead, further strengthening the model's performance in identifying key technologies in specific domains and in IPC classification tasks.
[0066] ② Adversarial sample enhancement To improve the robustness of the model, we randomly insert 5%-10% of non-technical text snippets (such as news and social media content) into the training data. These adversarial examples effectively prevent the model from overfitting to the domain data and enhance its generalization ability in diverse scenarios.
[0067] Step 2.3.2 Incremental pre-training implementation ①Model initialization and parameter configuration Load the pre-trained weights of the Qwen2.5-7B base model and use low-rank adaptation (LoRA) technology for efficient parameter fine-tuning. The specific configuration is: LoRA target module: insert a trainable rank decomposition matrix into the query and key matrices of the Transformer layer, set the rank dimension to 64, and control the number of new parameters to within 0.3% of the original model.
[0068] Freeze Parameters: Freeze more than 90% of the underlying parameters in the model, retaining its general language understanding capabilities and training only the LoRA adapter.
[0069] ② Training process design Dynamic course learning: Data is input step by step from simple to difficult according to the complexity of the text. In the early stage, the focus is on patent content such as patent text data. In the middle stage, international IPC classification data and national economic industry data are introduced. In the later stage, key industry technologies and their corresponding IPC data are fed in.
[0070] Batch training and gradient accumulation: Set the per-device training batch size to 4 and the number of gradient accumulation steps to 8 to balance video memory usage and training efficiency.
[0071] Learning rate scheduling: A cosine annealing learning rate scheduling strategy is used. The initial learning rate is set to 1e-4 and gradually decays to 1e-5 during training to ensure stable model convergence.
[0072] Training Rounds and Evaluation: Considering training costs, we set a total of three training rounds. We recorded the loss every 500 training steps and saved a model checkpoint every 1000 steps. We evaluated the model's perplexity for domain terms and the implicit accuracy of IPC classification on the validation set. If the model did not improve after three consecutive evaluations, we terminated training early.
[0073] ③Training execution Multi-node distributed training was launched based on the Llama-Factory framework, utilizing the parallel computing resources of eight NVIDIA V100 GPUs. The DeepSpeed optimizer (ZERO-3 configuration) was used to achieve efficient use of video memory, enabling FP16 precision, accelerating training and reducing video memory usage.
[0074] ④Model preservation and verification After training is complete, the LoRA adapter weights are combined with the original model weights and exported as a complete model file, which is saved to a specified directory. A training loss curve and validation set performance report are also generated for subsequent analysis. Performance verification is divided into two aspects: first, evaluating its ability to understand Chinese semantics, such as performance on benchmarks like CMMLU and C-EVAL; second, given some IPCs, observing whether the model's output of the meaning is correct.
[0075] Step 3: Supervised Fine-tuning: After completing incremental pre-training, this step constructs a domain-specific dataset and uses low-rank adaptation (LoRA) technology to perform supervised fine-tuning (SFT) on the model to accurately adapt it to key technology identification and IPC classification tasks. The specific implementation process is as follows: Step 3.1: Construct SFT training dataset ①Data integration and processing Based on the domain data collected in step 1 (IPC classification table, key technology descriptions, patent texts, etc.), a task-oriented supervised fine-tuning training set is constructed, which includes the following two types of core data: ②Structured data conversion: Design prompt words and use large language models such as ChatGPT to convert International Patent Classification (IPC) data into question-answer pairs to strengthen the model's understanding of classification logic. For example, the following example: Example 1: { "instruction":"Please determine which category the following technology belongs to according to the IPC classification rules: 'A crop growth monitoring system based on lidar, involving three-dimensional point cloud data processing and machine learning algorithms'", "input":"", "output": "This technology belongs to G01S17 / 88 (target detection and recognition through lidar) and G06N20 / 00 (machine learning model application)", "source":"IPC Classification Table" } Example 2: { "instruction":"What does B29K69 / 00 technically refer to?", "input":null, "output":" In the International Patent Classification (IPC) system, B29K69 / 00 is a specific subclass of class B29K, which indicates that B29K is used in conjunction with subclasses B29B, B29C or B29D and relates to molding materials or to materials used for reinforcement, fillers or preforms, such as inserts. B29K69 / 00 specifically refers to polycarbonate-related Specific explanation: 1.B: represents the first major category of IPC classification, namely operation; transportation.
[0076] 2.B29: represents the processing of plastics; generally the processing of materials in a plastic state.
[0077] 3. B29K: Further subdivided to indicate materials used in conjunction with subclasses B29B, B29C, or B29D, relating to molding materials or to materials used for reinforcement, fillers, or preforms, such as inserts.
[0078] 4.69 / : indicates polycarbonate 5. / 00: Indicates that the subcategory has no further subdivision. ", "source":"IPC Classification Table" } The same method is used to compare the key technology classification data - industrial technology classification and IPC classification relationship table, such as: { "instruction":"What are the core technologies of packaging substrates and their IPC classifications?" "input": null, "output":"Technical content: including packaging substrates, packaging carriers, etc.\nIPC classification:\nH01L33: Semiconductor devices specially adapted for light emission, such as infrared emission, having at least one potential transition barrier or surface barrier; methods or apparatus specific to the manufacture or processing of these semiconductor devices or their components; components of these semiconductor devices (H01L 51 / 50 preferred; devices consisting of multiple components formed in or on a common substrate enter H01L 27 / 00; coupling of light guides and optoelectronic elements enters G02B 6 / 42; semiconductor lasers enter H01S 5 / 00; electroluminescent light sources themselves enter H05B 33 / 00)\nH01L: Semiconductor devices; solid-state devices not included in other categories\nH01L31: Semiconductor devices sensitive to infrared radiation, light, electromagnetic radiation of shorter wavelength, or corpuscular radiation, and specially adapted for converting such radiation into electrical energy or for controlling electrical energy by means of such radiation; methods or apparatus specially adapted for the manufacture or processing of such semiconductor devices or components thereof; components thereof (H01L 51 / 42 preferred); devices consisting of a plurality of solid-state components, other than radiation-sensitive elements, formed in or on a common substrate in combination with one or more electric light sources H01L 27 / 00; energy harvesting devices for covering roof surfaces E04D 13 / 18; heat generation from solar energy F24J 2 / 00; semiconductor monitors for measuring X-radiation, gamma radiation, corpuscular radiation or cosmic radiation G0\n-H01L31 / 048: encapsulated or encased" "source":"Comparison table between industrial technology classification and IPC classification" } ③ Unstructured data conversion: Extract technical features and classification labels from patent texts to build key technology identification task data. For example: { "instruction":"Extract key technical features from the following patent abstracts:", The present invention discloses a graphene composite electrode material. A single-layer graphene is grown on a copper substrate by chemical vapor deposition, and then composited with a transition metal sulfide heterojunction..." "output":"1. Preparation of single-layer graphene by chemical vapor deposition; 2. Transition metal sulfide heterojunction composite technology; 3. Copper substrate interface optimization process", "source":"patent"} ④Data enhancement and quality control In terms of data augmentation and quality control, semantic equivalence expansion techniques are first used to replace synonyms (e.g., replacing "laser radar" with "LiDAR") and restructure sentences (e.g., converting active to passive voice) in technical descriptions of question-answer pairs. This generates semantically consistent yet diverse examples, effectively expanding the training set and improving the model's generalization. Secondly, for the IPC classification task, a negative sample generation strategy is employed to randomly replace the correct classification labels in 20% of the examples with similar but incorrect options (e.g., replacing "H01L35 / 28" with "H01L35 / 38"). This enhances the model's interference tolerance and classification robustness in complex scenarios. Finally, to ensure high data quality, domain experts manually review 10% of the sampled data, focusing on verifying the completeness of the extracted technical features and the accuracy of the classification labels. The data error rate is strictly controlled below 2%, providing a reliable data foundation for subsequent model training.
[0079] Step 3.2: LoRA-based supervised fine-tuning Step 3.2.1: Model initialization and parameter configuration Enable supervised fine-tuning (SFT) mode, specify the fine-tuning type as low-rank adaptation (LoRA), and set the adaptation target to all trainable modules of the model (including the query, key, and value matrices of the attention layer and the feedforward neural network parameters) to achieve full-level knowledge injection.
[0080] Based on the Llama-Factory framework, the weighted multi-task loss function is extended to implement: in, is the cross entropy loss function, which is used to optimize the IPC classification task; It is the KL divergence loss, which constrains the similarity of probability distribution in feature extraction tasks; is an L2 regularization term to prevent overfitting. This composite loss is dynamically calculated and back-propagated through the framework's customized callback interface.
[0081] Step 3.2.2: Supervised fine-tuning training ①Dynamic course learning mechanism Gradually increase the complexity of tasks in three stages: Phase 1 (first 1.5 epochs): Focuses on basic classification tasks. Input data is mainly IPC category recognition (such as "the technical meaning of A01B1 / 24"), accounting for 80% of the tasks. The length of technical description text in batch data is limited to ≤2048 tokens.
[0082] Phase 2 (Intermediate 1.0 Epoch): Introducing fine-grained classification of key technology identification (such as "Extract key technical features from the following patent abstracts:") and joint technical feature extraction tasks, and expanding the text length to ≤4096 tokens.
[0083] Phase 3 (last 0.5 epochs): Multi-hop reasoning tasks are enabled (e.g., "What are the core technologies of packaging substrates and their IPC classifications?"). The maximum length of the input text is maintained at 4096 tokens, covering the entire context window.
[0084] ② Training implementation The maximum number of training samples in a single cycle is limited to 50,000, 16 multi-process parallel data loading is enabled, and the input text is dynamically truncated or padded to 4096.
[0085] Set the single device batch size to 1, set the equivalent batch size to 8, and combine the 4-card parallel training to form a global batch size of 32. Use the cosine annealing strategy, and set the initial learning rate to , enable BF16 floating point format calculations.
[0086] Step 3.2.3: Model output and verification After training is complete, the LoRA adapter weights are merged and exported with the base model weights to form a complete model file that can be directly deployed. The merging process is implemented using the merge_lora tool built into the framework to ensure that inference efficiency is not compromised.
[0087] Construct a test dataset containing 2,000 samples that were not used in training to verify the model's performance in the following dimensions: Fine-grained classification accuracy: The classification accuracy of IPC groups (such as "A01B1 / 24") reached 98.2% Technical Relevance Reasoning: In cross-domain technical consulting tasks (such as "Key Technologies in the New Energy Vehicle Industry"), the consistency between the output results and expert judgment reaches 80%.
[0088] Step 4: Multi-agent collaborative reasoning Based on the SFT (supervised fine-tuning) model trained in step 3, a multi-agent collaborative reasoning framework is constructed to achieve accurate response through task offloading and collaboration mechanisms. The specific implementation process is as follows: Step 4.1: Agent division of labor and routing mechanism First, we design two agents, which can be extended to more agents in the future according to the needs of the question-answering task: ① Basic query agent (Agent_Base) Responsible for handling simple and straightforward questions, such as explaining the meaning of a specific IPC classification number, such as "the technical meaning of A01B1 / 24," or directly describing common key technologies, such as "semiconductor key technologies." This agent can directly call upon the knowledge learned by the SFT model for output.
[0089] For the problems handled by the basic agent, a general prompt word template is designed to directly guide the output of relevant knowledge of the SFT model saved in step 3. The general prompt word template is as follows: Regarding the question: {question content}, please give a concise and accurate answer based on professional knowledge and logic. ②Inference Agent (Agent_Inf) The system is triggered when faced with complex, compound questions, such as "What are the core technologies for package substrates and their IPC classifications?" This agent breaks down the complex question into multiple sub-questions, each of which is answered by invoking the SFT model saved in step 3, and then reasoning step by step. First, it reasoned about the specific content of "Core technologies for package substrates," then matched each identified core technology to the corresponding IPC classification number. The execution steps are as follows: Technical decomposition: extracting the core entities and task objectives in the problem Feature recognition: calling key technology recognition submodule Classification mapping: parallel matching of IPC classifications for each technology point Results integration: generating final structured output Step 4.2: Collaborative Reasoning Process Construct a problem recognition module and analyze the input problem in combination with the BERT-based classification prediction model: When the problem belongs to the simple and direct category, it is directly assigned to the basic intelligent agent after keyword matching and semantic analysis; if the problem involves multi-step reasoning, complex concept understanding or requires knowledge integration, such as complex problems such as "core technology and its classification", the problem will be transferred to the reasoning problem intelligent agent for processing, and the basic intelligent agent will be called as needed.
[0090] The following is an example of a multi-agent collaborative reasoning workflow for complex problems: Input question: core technology of packaging substrate and its IPC classification → Agent_Inf triggers decomposition: a) Subtask 1: Identify key technologies for "package substrate" → Output {substrate material selection technology, substrate structure design technology, manufacturing process control technology, reliability testing and evaluation technology} b) Subtask 2: IPC classification of each technology → Parallel call Agent_Base: - Substrate material selection technology → HH01L23 / 12, H01L23 / 15, C08G73 / 06 - Substrate structure design technology → HH01L23 / 498, H05K1 / 02, H05K3 / 46 - Manufacturing process control technology → H01L21 / 48, H05K3 / 06, H05K3 / 10 - Reliability testing and evaluation technology → G01R31 / 28, H01L21 / 66, H01L23 / 58 Results integration: Generate technology-classification mapping table Step 4.3: Output Example The following is an example of a specific process for the question "Key technologies in the field of new energy vehicle batteries and their IPC classification" Figure 2 As shown: After entering the reasoning agent Agent_Inf, the following important steps are implemented: Extract Entities Design a prompt word template and use the SFT model to provide the core entity and task objectives in the question. The corresponding prompt word design for this example is as follows: Identify the core entities in the question: "Key technologies in the field of new energy vehicle batteries and their IPC classification?" For the input question: "Key technologies in the field of new energy vehicle batteries and their IPC classification?", the extracted core entity is "new energy vehicle battery field".
[0091] Technology Identification Similarly, a technical recognition prompt word template is designed, the SFT model is called, and the key technologies corresponding to the core entities are given. The prompt word design corresponding to this example is as follows: Please analyze and list the core technologies in the field of new energy vehicle batteries in detail. The analysis process must be organized and the reasons for each step of reasoning must be explained.
[0092] After calling the SFT model, you can get the response: The field of new energy vehicle batteries is one of the key development directions of the global automotive industry. Its core technologies mainly include the following aspects: 1. High energy density battery materials Cause Analysis: Demand for increased range: As consumers' requirements for the range of electric vehicles continue to increase, high-energy-density battery materials have become the key to improving the range of electric vehicles.
[0093] Cost control: By increasing energy density, vehicle range can be extended without increasing battery weight, thereby reducing overall costs.
[0094] 2. Fast charging technology Cause Analysis: Charging infrastructure limitations: Fast charging technology can shorten charging time, reduce user waiting time, and help overcome the problem of insufficient charging facilities.
[0095] User experience optimization: For users who need to travel frequently, the ability to quickly replenish power will greatly improve the user experience.
[0096] 3. Security Technology Cause Analysis: Battery thermal runaway problem: Safety is one of the most core issues of electric vehicles, especially in cases of battery overcharging, short circuit, etc., which can easily cause fire or explosion.
[0097] Regulatory requirements and market trust: As countries pay more and more attention to the safety of electric vehicles, improving battery safety has become an industry consensus.
[0098] 4. Environmentally friendly recycling technology Cause Analysis: Resource recycling: Lithium-ion batteries contain a variety of valuable metal elements, such as lithium, cobalt, nickel, etc. The development of efficient recycling technology is conducive to resource recycling.
[0099] Environmental considerations: Proper disposal of used batteries reduces environmental pollution and complies with the concept of sustainable development.
[0100] 5. Intelligent Management System Cause Analysis: Performance optimization: Real-time monitoring of battery status through the intelligent management system can effectively extend battery life while ensuring optimal working efficiency.
[0101] Fault warning: Discover potential problems in advance to avoid safety accidents caused by battery failure. ... IPC classification mapping For each of the key core technologies mentioned above, the basic agent (Agent_Base) is called to match the IPC classification number. For example, the corresponding prompt words for "high energy density battery materials" in this example are as follows: Question: Give the IPC classification number corresponding to the technology "high energy density battery materials". Please give a concise and accurate answer based on professional knowledge and logic. The following results are obtained: Key technologies: positive electrode materials (such as lithium iron phosphate, ternary materials), negative electrode materials (such as graphite, silicon-carbon composite materials), electrolyte, diaphragm, etc.
[0103] Corresponding IPC classification number: H01M 4 / 00: Electrode materials H01M 10 / 0525: Lithium-ion battery H01M 10 / 056: Electrolyte H01M 10 / 0562: Solid Electrolyte H01M 50 / 00: Battery separator Others such as "fast charging technology" and "safety technology" obtained results in turn.
[0104] Summary Output Combine the results of each technology output, summarize the content by calling the SFT model, and finally get the output: 1. Battery material technology Key technologies: positive electrode materials (such as lithium iron phosphate, ternary materials), negative electrode materials (such as graphite, silicon-carbon composite materials), electrolyte, diaphragm, etc.
[0105] IPC classification: H01M 4 / 00: Electrode materials H01M 10 / 0525: Lithium-ion battery H01M 10 / 056: Electrolyte H01M 10 / 0562: Solid Electrolyte H01M 50 / 00: Battery separator 2. Battery Management System (BMS) Key technologies: battery state estimation (SOC, SOH), thermal management, balancing control, fault diagnosis, etc.
[0106] IPC classification: H01M 10 / 42: Battery Management H01M 10 / 48: Battery status monitoring H01M 10 / 615: Thermal Management H01M 10 / 625: Temperature control H01M 10 / 63: Battery balancing 3. Battery manufacturing process Key technologies: electrode coating, electrode cutting, battery assembly, packaging, etc.
[0107] IPC classification: H01M 4 / 04: Electrode Manufacturing H01M 10 / 04: Battery Assembly H01M 10 / 058: Battery Manufacturing Process H01M 50 / 00: Battery package 4. Battery safety technology Key technologies: overcharge / over-discharge protection, short circuit protection, thermal runaway protection, etc.
[0108] IPC classification: H01M 10 / 42: Battery protection H01M 10 / 44: Charging Control H01M 10 / 52: Battery Safety H01M 50 / 00: Battery safety design 5. Battery recycling and reuse Key technologies: battery disassembly, material recycling, cascade utilization, etc.
[0109] IPC classification: H01M 10 / 54: Battery Recycling H01M 10 / 60: Battery Recycling H01M 6 / 52: Waste Battery Disposal Similarly, the above process is used to solve the problem "Key technologies in the rare earth field and their IPC classification", and the output is: 1. Rare earth mining and extraction technology Key technologies: rare earth mining, beneficiation, leaching, separation and purification, etc.
[0110] IPC classification: C22B 59 / 00: Extraction of rare earth metals C22B 3 / 00: Hydrometallurgical extraction C22B 1 / 00: Ore pretreatment C22B 3 / 04: Acid leaching C22B 3 / 26: Solvent Extraction 2. Rare earth separation and purification technology Key technologies: solvent extraction, ion exchange, chromatographic separation, etc.
[0111] IPC classification: C22B 3 / 26: Solvent Extraction C22B 3 / 40: Ion Exchange C22B 3 / 42: Chromatographic separation C22B 59 / 00: Separation of rare earth metals 3. Rare earth functional material preparation technology Key technologies: rare earth permanent magnet materials, rare earth luminescent materials, rare earth catalytic materials, etc.
[0112] IPC classification: H01F 1 / 00: Magnetic materials (such as NdFeB permanent magnets) C09K 11 / 00: Luminescent materials (such as phosphors) B01J 23 / 00: Catalyst Materials C22C 28 / 00: Rare earth alloy 4. Rare earth recovery and reuse technology Key technologies: recycling of waste rare earth materials, utilization of rare earth secondary resources, etc.
[0113] IPC classification: C22B 7 / 00: Scrap Metal Recycling C22B 59 / 00: Rare Earth Metal Recovery C22B 3 / 00: Hydrometallurgical Recovery C22B 3 / 26: Solvent Extraction Recovery 5. Rare earth application technology Key technologies: Application of rare earths in new energy, electronic information, aerospace and other fields.
[0114] IPC classification: H01M 4 / 00: Application of rare earths in batteries H01F 1 / 00: Application of rare earth in magnetic materials C09K 11 / 00: Application of rare earths in luminescent materials B01J 23 / 00: Application of rare earths in catalysts The model that has undergone incremental pre-training and supervised fine-tuning can instantly respond to key core technology identification and IPC classification in any specified field, fully demonstrating the universality of this method for key core technology identification queries in different fields.
[0115] The protection scope of the present invention is obviously not limited to the above specific embodiments, but also includes the following equivalent changes or replacements to achieve the present invention: 1. Base model selection: Qwen2.5 is selected as the base model in this invention, but other models with similar architecture and performance (such as GPT-4, GLM, etc.) can also be used to implement the technical solution of this invention; 2. Supervised fine-tuning strategy: In addition to using LoRA technology for supervised fine-tuning, other parameter efficient fine-tuning technologies (such as Adapter and Prefix-tuning) can also be used to achieve similar effects.
[0116] Although the present invention has been disclosed above in terms of preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art may make possible changes and modifications to the technical solutions of the present invention by using the methods and technical contents disclosed above without departing from the spirit and scope of the present invention. Therefore, any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solutions of the present invention shall fall within the scope of protection of the technical solutions of the present invention.
Claims
1. A method for key technology identification and IPC classification based on a large language model, characterized in that the steps include: Step S1: collecting domain data and performing data preprocessing to construct a structured domain corpus, wherein the domain corpus includes public patent text data and IPC classifications, multiple IPC classification comparison tables, and general text corpus; Step S2: Based on the Llama-Factory framework, a byte pair encoding algorithm is used to perform word segmentation training on the domain corpus to generate corresponding BPE merging rules, and the BPE merging rules are merged with the original BPE merging rules of the Qwen2.5 base model to reconstruct the hybrid word segmenter and generate the corresponding domain vocabulary; at the same time, the domain data is converted into the corresponding pre-training data corpus; Using the pre-training data corpus and the domain vocabulary, incrementally pre-training the Qwen2.5 base model by adopting a low-rank adaptation technique to inject the domain data; Step S3: constructing a supervised fine-tuning training dataset for a specified task based on the domain data, performing data preprocessing to convert the data into a format suitable for SFT unit input, and performing supervised fine-tuning on the SFT unit through the low-rank adaptation technology to enable the SFT unit to achieve the expected performance standard; Step S4: Based on the proposed key technology identification and IPC classification problems, they are assigned to the corresponding query agent or reasoning agent through the task diversion mechanism, and the SFT unit is called to process the problems and output answers.
2. The method for key technology identification and IPC classification based on a large language model according to claim 1, characterized in that: Step S1 includes collecting multi-source heterogeneous domain data and performing data preprocessing, including data cleaning, format conversion and semantic analysis, and integrating the domain data after the data preprocessing to form a unified domain corpus.
3. The method for key technology identification and IPC classification based on a large language model according to claim 2, characterized in that: In step S2, the domain corpus is trained with a byte pair encoding algorithm to generate corresponding BPE merging rules, which are then merged with the original BPE merging rules of the Qwen2.5 base model to reconstruct a hybrid word segmenter and generate a corresponding domain vocabulary, including: Extracting the original vocabulary and the BPE merging rules from the Qwen2.5 base model, retaining the subword units and their corresponding identifiers in the original Qwen2.5 base model; The byte pair encoding algorithm is used to perform the word segmentation training on the domain corpus in step S1, and high-frequency character pairs in the corpus are iteratively counted and merged to form the BPE merging rules for a specific domain; The formed BPE merging rule is superimposed on the original BPE merging rule, and the new domain vocabulary is appended to the end of the original vocabulary for fusion, and a unique identifier is assigned to form the corresponding domain vocabulary; Based on the fused domain vocabulary and the BPE merging rules, the corresponding hybrid word segmenter is constructed for processing and analyzing general and domain-specific text data.
4. The method for key technology identification and IPC classification based on a large language model according to claim 3, characterized in that: In step S2, the domain data is converted into corresponding pre-training data corpus, including: Based on the classified general text corpus, extracting categories associated with the specific field and using them as the pre-training data corpus; Converting the patent text data into natural text, and combining the extracted patent number information to assemble the patent information into the complete pre-training data corpus; wherein the patent number information includes the patent number, invention name and the IPC classification; Based on the hierarchical relationship information in the multiple IPC classification comparison relationship tables, the hierarchical relationship information is converted into a descriptive language parsed by a large language model and used as the pre-training data corpus.
5. The method for key technology identification and IPC classification based on a large language model according to claim 4 is characterized in that: In step S2, using the pre-training data corpus and the domain vocabulary, the Qwen2.5 base model is incrementally pre-trained by adopting a low-rank adaptation technique to inject the domain data, including: Dividing the pre-training data corpus into a training set and a validation set according to a preset ratio, wherein the training set is used for the incremental pre-training of the Qwen2.5 base model, and the validation set is used to monitor the performance of the incremental pre-training process; Load the pre-trained weights of the Qwen2.5 base model and fine-tune the model parameters using the low-rank adaptation technique. The Qwen2.5 base model is configured by inserting a trainable rank factorization matrix into the query and bond matrices of the Transformer layer, with the rank dimension set to 64, and the amount of new parameters is controlled within 0.3% of the original model. At the same time, more than 90% of the underlying parameters of the Qwen2.5 base model are frozen to retain its general language understanding capabilities and only train the LoRA adapter; The pre-training data corpus is introduced according to the preset standards, and small batch training and gradient accumulation techniques are used to balance video memory usage and efficiency. At the same time, the learning rate is adjusted using the cosine annealing strategy to ensure the convergence of the Qwen2.5 base model; Based on the fully trained Qwen2.5 base model, the LoRA adapter weights and the pre-trained weights are merged to export the complete model file.
6. The method for key technology identification and IPC classification based on a large language model according to claim 5, characterized in that: In step S3, based on the key technology identification and IPC classification questions, the tasks are assigned to the corresponding query agent or reasoning agent through the task offloading mechanism, and the SFT unit is called to process the questions and output answers, including: Based on the collected domain data, a supervised fine-tuning training set for key technology identification and IPC classification is constructed, and data enhancement is performed. For structured domain data, prompt words are set and the large language model is called to convert the IPC classification questions into question-answer format. For unstructured domain data, technical features in the patent text data are extracted and converted into the key technology identification task. The supervised fine-tuning mode in the SFT unit is enabled, and the low-rank adaptation technology is used to fine-tune the model parameters; wherein, based on the Llama-Factory framework, a weighted multi-task loss function is extended to implement: in, is a cross entropy loss function used to optimize the IPC classification task; KL divergence loss constrains the similarity of probability distribution in the feature extraction task; is the L2 regularization term; Based on the SFT unit after training, the LoRA adapter weights and the Qwen2.5 base model weights are merged to export the complete model file.
7. The method for key technology identification and IPC classification based on a large language model according to claim 6, characterized in that: In step S4, based on the key technology identification and IPC classification questions, the tasks are assigned to the corresponding query agent or reasoning agent through the task offloading mechanism, and the SFT unit is called to process the questions and output answers, including: Based on the fully trained SFT unit, a corresponding multi-agent collaborative reasoning unit is constructed, wherein the multi-agent collaborative reasoning unit includes the query agent and the reasoning agent; Analyze and judge the input problem by combining the classification prediction model based on Bert; When the question is a simple question, it is assigned to the query agent after keyword matching and semantic analysis, and the SFT unit is called to output the answer; When the problem is a complex problem, the reasoning agent breaks down the complex problem into several sub-problems, and calls the SFT unit to identify the corresponding key technology and match the corresponding IPC classification.
8. A system for key technology identification and IPC classification based on a large language model, using the method for key technology identification and IPC classification based on a large language model according to any one of claims 1 to 7, characterized in that: include: A data acquisition module is used to collect domain data and perform data preprocessing to build a structured domain corpus. The domain corpus includes public patent text data and IPC classifications, multiple IPC classification comparison tables, and general text corpus. An incremental pre-training module is used to perform word segmentation training on the domain corpus using a byte pair encoding algorithm based on the Llama-Factory framework, generate corresponding BPE merging rules, merge the BPE merging rules with the original BPE merging rules of the Qwen2.5 base model, and reconstruct the hybrid word segmenter and generate the corresponding domain vocabulary; at the same time, the domain data is converted into the corresponding pre-training data corpus; Using the pre-training data corpus and the domain vocabulary, incrementally pre-training the Qwen2.5 base model by adopting a low-rank adaptation technique to inject the domain data; A supervised fine-tuning module is used to construct a supervised fine-tuning training dataset for a specified task based on the domain data, perform data preprocessing to convert the data into a format suitable for SFT unit input, and perform supervised fine-tuning on the SFT unit using the low-rank adaptation technology to enable the SFT unit to achieve the expected performance standard; Result output module: Based on the proposed key technology identification and IPC classification problems, the task is assigned to the corresponding query agent or reasoning agent through the task diversion mechanism, and the SFT unit is called to process the problem and output the answer.
9. The system for key technology identification and IPC classification based on a large language model according to claim 8, characterized in that: The incremental pre-training module includes: A pre-training data corpus processing unit, configured to extract categories associated with a specific field based on the classified general text corpus, and use the categories as the pre-training data corpus; The patent text data is converted into natural text and combined with the extracted patent number information to assemble the complete pre-training data corpus.
10. The system for key technology identification and IPC classification based on a large language model according to claim 8, characterized in that: The patent number information includes the patent number, the invention name and the IPC classification; based on the hierarchical relationship information in multiple IPC classification comparison relationship tables, the hierarchical relationship information is converted into a descriptive language parsed by a large language model and used as the pre-training data corpus.
Citation Information
Patent Citations
A heat electric module and its making method
CN101170157A
Cited By
Method and system for improving task reasoning speed
CN121880560A
Intellectual property multi-branch task processing method and device
CN122064771A