Method and system for identifying material attribute information by using large language model

Through the large language model combining word segmentation algorithm and rule model to identify material attribute information, the problems of low recognition efficiency and unstable accuracy in traditional material recognition technology are solved, and more efficient and accurate material recognition is achieved, supporting the transparency and synergy of supply chain management.

CN120067288APending Publication Date: 2025-05-30CHINA COAL DEVELOPMENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411930127.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-25
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Traditional material identification technology has problems such as low identification efficiency, unstable accuracy, high labor costs, difficulty in adapting to material diversity and limited data processing capabilities.

Method used

A large language model is used to combine word segmentation algorithms and rule models to determine whether manual verification is required by extracting material information, predicting attribute extraction results, analyzing material types, determining the recognition accuracy and reliability coefficients.

Benefits of technology

It improves the accuracy and efficiency of material identification, reduces communication costs, enhances the transparency and coordination of supply chain management, and supports more accurate management decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067288A_ABST
    Figure CN120067288A_ABST
Patent Text Reader

Abstract

The invention provides a method and system for identifying material attribute information by using a large language model, and belongs to the technical field of data processing, and the method specifically comprises the following steps: extracting material information of a material, and determining extraction results of the material information in different types of attributes by using a word segmentation algorithm and an identification result of a rule model, and determining prediction extraction results of the material information in different types of attributes by using an identification result of the large language model, and determining an artificial verification proportion of the material type corresponding to the material according to a deviation condition of the prediction extraction results corresponding to the material type and a deviation condition of the extraction results, and whether the material needs to be subjected to manual verification processing is determined by utilizing the manual verification proportion and the reliability coefficient of the material, so that the reliability of management of the material attribute information is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data processing, and particularly relates to a method and system for identifying material attribute information by using a large language model. Background Art

[0002] In the contemporary industrial and logistics management fields, the material identification link is crucial for ensuring the efficiency and accuracy of the process. However, traditional material identification technologies face many challenges, which seriously limit the operation efficiency and cost control of enterprises. The obvious defects of traditional material identification technologies include low identification efficiency, unstable accuracy, high labor costs, insufficient adaptability to material diversity, and limitations in data processing capabilities.

[0003] In view of the above technical problems, the present invention provides a method and system for identifying material attribute information by using a large language model. Summary of the Invention

[0004] The object of the present invention is to provide a method for identifying material attribute information by using a large language model.

[0005] To solve the above technical problems, the first aspect of the present invention provides a method for identifying material attribute information by using a large language model, which specifically includes:

[0006] S1 Extract the material information of the material, and use the recognition results of the word segmentation algorithm and the rule model to determine the extraction results of the material information in different types of attributes, and use the recognition results of the large language model to determine the predicted extraction results of the material information in different types of attributes;

[0007] S2 Based on the predicted extraction results and extraction results of different types of attributes, determine the material type of the material, and according to the analysis results of the material type, determine the deviation situation of the predicted extraction results corresponding to the material type, and when the recognition accuracy rate of the material type meets the requirements by using the deviation situation, proceed to the next step;

[0008] S3 Determine the deviation situation between the predicted extraction results and extraction results of the material in different types of attributes, and combine the matching situation between the predicted extraction results in different types of attributes, and when the reliability coefficient of the predicted extraction results meets the requirements, proceed to the next step;

[0009] S4 According to the deviation situation of the predicted extraction results corresponding to the material type and the deviation situation of the extraction results, determine the manual verification ratio of the material type corresponding to the material, and use the manual verification ratio and the reliability coefficient of the material to determine whether the material needs to be manually verified.

[0010] The beneficial effects of the present invention are as follows:

[0011] By identifying the material attribute information, large language models can provide a unified material data standard for all parties in the supply chain, enhancing the transparency and collaboration of supply chain management. This helps reduce communication costs and improve the response speed and flexibility of the supply chain.

[0012] The material attribute information identified by large language models can serve as an important input for data analysis. Through the mining and analysis of this data, enterprises can gain in-depth insights into the consumption patterns, demand trends, and potential risk points in the supply chain, providing strong support for management decision-making.

[0013] The technology of large language models identifying material attribute information has broad application prospects and important value in aspects such as material management, production process optimization, and supply chain management. Through the application of this technology, enterprises can achieve automated management and efficient utilization of material information, improve production efficiency and product quality, and reduce operating costs and risks.

[0014] A further technical solution is that the attributes include material model, name, brand, specification, and size.

[0015] A further technical solution is that the large language model is built by fusing the NER model with the large model.

[0016] A further technical solution is that the method for determining the recognition result of the word segmentation algorithm is as follows:

[0017] Construct a feature library:

[0018] Construct a feature library by using the material brand data collected from various sources as the basis. Each brand name is used as an independent feature item and stored as a database table, file, or data structure in memory;

[0019] Word segmentation algorithm recognition:

[0020] Use the word segmentation algorithm to segment the text, extract key words, and use the large language model to preliminarily recognize the preprocessed key words to obtain a candidate set of possible brand names;

[0021] Match the candidate set of brand names obtained from the preliminary recognition with the brand names in the feature library, and determine the recognition result of the word segmentation algorithm according to the matching result of the feature library.

[0022] A further technical solution is that the sources include official databases, industry standards, and market research reports.

[0023] A further technical solution is that the specific steps for constructing the feature library are as follows:

[0024] Collect material brand data from various sources, clean and organize the collected material brand data, remove duplicates, errors, or irrelevant items, and format the collected material brand data;

[0025] Construct the sorted material brand data into a feature library, take each brand name as an independent feature item, and use the feature library constructed by the feature items to store it as a database table, file, or data structure in memory.

[0026] A further technical solution lies in determining whether the material needs to be manually verified, specifically including:

[0027] Use the manual verification ratio of the materials of the material type and the quantity of the materials of the material type to determine the manual verification quantity of the materials of the material type;

[0028] According to the manual verification quantity, determine the verification target of the materials of the material type in descending order of the reliability coefficient of the material;

[0029] Use the verification target to determine whether the material needs to be manually verified.

[0030] A further technical solution lies in that when the material belongs to the verification target, it is determined that the material needs to be manually verified.

[0031] On the other hand, the present application provides a system for identifying material attribute information using a large language model, which is applied to the above method for identifying material attribute information using a large language model, specifically including:

[0032] Material information parsing module, manual verification module;

[0033] Among them, the material information parsing module is responsible for extracting the material information of the material, and using the recognition results of the word segmentation algorithm and the rule model to determine the extraction results of the material information in different types of attributes, and using the recognition results of the large language model to determine the predicted extraction results of the material information in different types of attributes;

[0034] The manual verification module is responsible for manually verifying the material.

[0035] Other features and advantages will be described in the subsequent description, and, in part, will become apparent from the description, or will be understood by implementing the present invention. The objectives and other advantages of the present invention are achieved and obtained by the structures specifically pointed out in the description and the drawings.

[0036] To make the above objectives, features, and advantages of the present invention more obvious and understandable, the following specific preferred embodiments are given, and in conjunction with the accompanying drawings, the detailed description is as follows. Brief Description of the Drawings

[0037] By referring to the accompanying drawings and describing in detail its exemplary embodiments, the above and other features and advantages of the present invention will become more apparent.

[0038] Figure 1 It is a flowchart of a method for identifying material attribute information using a large language model according to Embodiment 1.

[0039] Figure 2 It is a flowchart of a method for determining the recognition result of a word segmentation algorithm.

[0040] Figure 3 It is a flowchart of the specific steps for constructing a feature library.

[0041] Figure 4 It is a flowchart of the specific steps for developing a rule engine. Detailed Embodiments

[0042] Now, the exemplary embodiments will be described more fully with reference to the accompanying drawings. However, the exemplary embodiments can be implemented in various forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this invention will be thorough and complete, and will fully convey the concept of the exemplary embodiments to those skilled in the art. Like reference numerals in the figures denote like or similar structures, and thus their detailed descriptions will be omitted.

[0043] The terms "a", "an", "the", and "said" are used to denote the presence of one or more elements / components / etc.; the terms "comprising" and "having" are used to mean an open inclusion and mean that there may be additional elements / components / etc. in addition to the listed elements / components / etc.

[0044] First, collect material information. This includes recording the basic information of the material, such as name, model, specification, material, manufacturer, etc. This information is usually obtained by manual entry or by scanning barcodes, QR codes, etc. on the material. Subsequently, organize the collected material information into a document or spreadsheet for subsequent query and management.

[0045] Secondly, material classification and coding. Classify the materials according to factors such as the attributes, uses, and characteristics of the materials to optimize organization and management. At the same time, assign a unique code to each material for quick retrieval and identification in the system. The coding system may include letters, numbers, or a combination of them to reflect information such as the classification and attributes of the materials.

[0046] Next, conduct material identification. Based on the coding and classification information of the materials, make corresponding labels and paste them on the materials or their packages. The labels usually contain key information such as the name, coding, quantity, production date, and expiration date of the materials. In addition, to improve the identification efficiency, barcodes or QR codes can also be generated for the materials and printed on the labels. In this way, the material information can be quickly read through scanning devices.

[0047] Then, conduct material identification and verification. In the processes of material receipt, issue, and inventory count, etc., the staff manually identify and verify according to the information on the labels or barcodes / QR codes. With the development of technology, more and more enterprises begin to adopt automatic identification technologies (such as barcode scanners, RFID readers, etc.) to improve the identification efficiency and accuracy. These devices can automatically read the information on the material labels and transmit it to the management system for processing.

[0048] After that, conduct material information management. Establish a material information management system and store the collected material information in the database. The system should support operations such as querying, modifying, and deleting material information so as to keep track of the inventory status and flow status of the materials in real time. Through the analysis of material information, the consumption patterns, demand trends, and potential risk points in the supply chain of the materials can be understood. This information provides strong support for the decision-making of the enterprise's procurement plan, production plan, inventory management, etc.

[0049] Finally, conduct material traceability and quality management. To ensure the quality and safety of the materials, a material traceability system needs to be established. Through the traceability system, information such as the source, production batch, production date, and inspection records of the materials can be traced, so that the cause can be quickly located and measures can be taken when quality problems occur. Conduct quality inspections before material receipt and before use to ensure that the materials meet the specified quality standards. The inspection results should be recorded and associated with the material information for subsequent query and traceability.

[0050] To sum up, the workflow of traditional material identification includes multiple links such as material information collection, classification and coding, identification, identification and verification, information management, and traceability and quality management. These links are interrelated and mutually supportive, jointly constituting a complete system for material identification and management.

[0051] In the contemporary industrial and logistics management fields, the material identification link is crucial for ensuring the efficiency and accuracy of the process. However, traditional material identification technologies face many challenges, which seriously limit the operation efficiency and cost control of enterprises. The obvious defects of traditional material identification technologies include low identification efficiency, unstable accuracy, high labor costs, insufficient adaptability to material diversity, and limitations in data processing capabilities. Specifically, the deficiencies of traditional material identification technologies are mainly manifested in the following aspects:

[0052] 1. Low recognition efficiency

[0053] High manual dependency: Traditional material recognition technologies often rely heavily on manual operations, such as manually checking and recording material information. In the case of a large variety and quantity of materials, this method is inefficient and difficult to meet the requirements of rapid recognition.

[0054] Slow processing speed: Manually recognizing materials requires one-by-one checking and comparison, and its processing speed is far slower than that of automated recognition systems. In emergency or high-frequency material recognition scenarios, traditional methods may not be able to respond in time.

[0055] 2. Low recognition accuracy

[0056] Interference of human factors: The manual recognition process is easily affected by factors such as subjective judgment, fatigue, and distraction of attention, resulting in deviations or errors in the recognition results.

[0057] Inconsistent recognition standards: Different personnel may have different recognition standards for materials, resulting in inconsistent recognition results and affecting subsequent data statistics and analysis.

[0058] 3. High cost

[0059] Labor cost: A large amount of manual participation increases the labor cost of enterprises. Especially in the context of the continuous rise of labor costs, this problem is particularly prominent.

[0060] Time cost: The inefficient recognition process occupies a large amount of time resources and reduces the operational efficiency of enterprises.

[0061] 4. Difficulty in adapting to material diversity

[0062] Diverse material forms: In the case of diverse material forms and complex packaging, traditional recognition technologies may be difficult to accurately identify material information.

[0063] 5. Limited data processing ability

[0064] Data storage and query: Traditional material recognition technologies usually rely on paper documents or simple spreadsheets to store and query material information. In the case of an increasing amount of data, the query efficiency is low and errors are prone to occur.

[0065] Data analysis ability: Due to the lack of advanced data processing and analysis tools, traditional technologies are difficult to deeply mine and analyze material data, thus limiting the exertion of data value.

[0066] In summary, traditional material recognition technologies have defects such as low recognition efficiency, low accuracy, high cost, difficulty in adapting to material diversity, and limited data processing ability.

[0067] With the rapid progress of artificial intelligence technology, large language models have become a research hotspot in the field of natural language processing (NLP). These models are pre-trained on vast text datasets, thereby acquiring profound semantic understanding and context processing capabilities. They can handle various complex natural language processing tasks, such as text generation, question-answering systems, and language translation, fully demonstrating their excellent language processing capabilities. By integrating these artificial intelligence technologies, intelligent material coding governance and management can be achieved, significantly improving work efficiency and enhancing the data quality and accuracy of material coding.

[0068] Traditional material identification methods rely on manual input or barcode scanning. These methods are not only inefficient but also prone to errors. Large language models, through natural language processing technology, can automatically identify and parse key attribute information in material descriptions, significantly improving the accuracy and efficiency of identification.

[0069] By leveraging large language models, enterprises can achieve automated management and updating of material information. When new materials enter the system, the model can automatically extract their attribute information and store it in the database without manual intervention, which greatly reduces the workload of material management personnel and the possibility of human errors. Accurate material attribute information helps enterprises formulate more reasonable procurement plans and inventory strategies.

[0070] By identifying material attributes through large language models, enterprises can better understand the demand situation, inventory status, and supplier supply capabilities of materials, thereby making more accurate and timely procurement decisions, reducing inventory costs, and improving capital turnover.

[0071] During the production process, accurate material attribute information is crucial for the formulation and execution of production plans. Large language models can quickly identify and match the materials required for production, reducing production delays and waste caused by material errors or shortages. At the same time, the model can also automatically adjust production parameters and process flows based on material attribute information, improving production efficiency and product quality.

[0072] Supply chain management involves collaborative cooperation among multiple links and multiple participants. By identifying material attribute information, large language models can provide a unified material data standard for all parties in the supply chain, enhancing the transparency and collaboration of supply chain management. This helps reduce communication costs and improve the response speed and flexibility of the supply chain.

[0073] The material attribute information identified by large language models can serve as an important input for data analysis. By mining and analyzing this data, enterprises can gain in-depth insights into the consumption patterns, demand trends, and potential risk points in the supply chain, providing strong support for management decisions.

[0074] The technology of large language models for identifying material attribute information has broad application prospects and important value in aspects such as material management, production process optimization, and supply chain management. Through the application of this technology, enterprises can achieve automated management and efficient utilization of material information, improve production efficiency and product quality, and reduce operating costs and risks.

[0075] 1. Data Preparation and Preprocessing:

[0076] The sources of material attribute data are diverse, of uneven quality, and have complex descriptions. To ensure that the large language model can accurately identify materials, this patent proposes a brand-new data preparation and preprocessing process, covering all aspects from data collection to quality improvement, with the characteristics of high automation and intelligence.

[0077] ● Multi-source Data Acquisition and Integration: This patent designs a multi-level data acquisition mechanism covering a wide range of data sources. First, internal data mainly comes from enterprise management systems such as ERP, PLM, and SCM, covering material descriptions and historical records in aspects such as enterprise production, design, and procurement.

[0078] ● Intelligent Data Cleaning and Standardization: To address the noise and inconsistencies in the data, this patent proposes a set of intelligent data cleaning and standardization mechanisms. During the cleaning process, deep learning algorithms are used for automatic spelling correction, synonym recognition, and detection and resolution of semantic conflicts. For the common unit difference problems in material descriptions, the system realizes automatic unit conversion and standardization processing through semantic analysis and context understanding. In addition, a multi-round feedback mechanism is combined during the data cleaning process, which can gradually optimize the cleaning strategy according to historical data and user feedback to ensure that the cleaned data reaches high precision and high consistency.

[0079] ● Data Quality Improvement and Enhancement: The quality of data directly affects the training effect of the large language model. To ensure the high quality of training data, this patent introduces a variety of data enhancement technologies. First, a data synthesis method based on generative adversarial networks (GANs) is adopted to generate diverse material descriptions, especially for data expansion of rare materials and new materials. Secondly, combined with data enhancement techniques, through methods such as synonym replacement, random noise injection, and context perturbation, the diversity of data and the robustness of the model are further improved. In addition, this patent also designs a knowledge fusion method based on an expert system to combine the experience knowledge of domain experts with the data, enhancing the semantic depth and industry adaptability of the data.

[0080] 2. Selection of Large Language Model and Application of Pre-trained Model:

[0081] To ensure that the large language model can exhibit the best performance in the field of material recognition, this patent adopts a variety of innovative technologies during model selection and pre-training, fully leveraging the potential of the model and specifically optimizing it for material recognition tasks.

[0082] ● Multi-task learning and model collaborative optimization: During the model selection process, this patent adopts a multi-task learning (MTL) framework to integrate multiple tasks related to material recognition into the same model for training. Through shared parameters and feature learning, different tasks can promote each other, thereby improving the overall performance of the model. To further optimize the model, this patent designs a dynamic weight adjustment mechanism that can automatically adjust the weight allocation according to the performance of each task during training, ensuring that the model always maintains the best state in a multi-task environment. In addition, the system also introduces a model architecture based on multi-scale feature extraction, which can capture semantic information at different levels and enhance the model's ability to understand complex material descriptions.

[0083] ● Industry-specific pre-training and adaptive semantic mapping: To improve the model's performance in a specific industry, this patent designs an adaptive industry corpus expansion and pre-training strategy. By collecting and constructing an industry-specific corpus, the model can deeply learn the specialized terms, expressions, and semantic relationships in the material field during the pre-training stage. Combining with the adaptive semantic mapping algorithm, the system can quickly switch between different fields and perform accurate semantic mapping. In addition, the system also introduces few-shot learning technology, enabling the model to still maintain high recognition accuracy and adaptability in the absence of a large amount of labeled data.

[0084] 3. Model fine-tuning and optimization:

[0085] After completing the pre-training, this patent designs a comprehensive model fine-tuning and optimization process, combining the most advanced technical means to enable the model to exhibit excellent performance and high adaptability in material attribute recognition tasks.

[0086] ● Fine-grained fine-tuning based on contrastive learning: During the model fine-tuning stage, this patent introduces contrastive learning technology, enabling the model to further optimize its discrimination ability by identifying similar and different material descriptions. This technology not only enhances the model's perception of subtle differences but also improves its accuracy in processing similar material descriptions. Combining with data assimilation technology, the system can maintain data consistency and high quality while introducing new data, ensuring that the model can quickly adapt and maintain efficient recognition when facing new material descriptions.

[0087] ● Cross - domain Transfer Learning and Knowledge Distillation: To enhance the generalization ability and cross - domain adaptability of the model, this patent proposes an optimization strategy based on Transfer Learning and Knowledge Distillation. By transferring knowledge from other domains to the material recognition task, the model can quickly adapt and demonstrate excellent performance in new domains. The knowledge distillation technique transfers the knowledge of complex models to smaller sub - models, enabling the sub - models to maintain high performance while reducing computational resource consumption. In addition, the introduction of transfer learning also enables the model to quickly respond to and adapt to changing material description requirements.

[0088] 4. Model Evaluation and Improvement:

[0089] To ensure the efficiency and stability of the model in practical applications, this patent designs a comprehensive and detailed model evaluation and automated improvement mechanism, enabling the model to continuously optimize and improve in various complex scenarios.

[0090] ● Multi - dimensional Performance Evaluation and Comparative Analysis: During the model evaluation process, the system adopts multi - dimensional performance evaluation metrics, covering aspects such as recognition accuracy, computational efficiency, resource consumption, response speed, etc. Through comparative analysis with baseline models and industry standards, the system can accurately identify the advantages and disadvantages of the model and formulate optimization strategies accordingly. An adaptive evaluation mechanism based on Meta - Learning is introduced during the evaluation process, which can dynamically adjust the weights of evaluation metrics according to different task and data characteristics to ensure the comprehensiveness and accuracy of evaluation results.

[0091] ● Adaptive Model Update and Intelligent Feedback Loop: During the model improvement process, this patent introduces an adaptive model update mechanism that can automatically adjust the parameters and structure of the model according to task requirements and changes in the external environment. Combined with an intelligent feedback loop, the system can, after each iteration, make targeted adjustments and optimizations to the model based on user feedback and evaluation results to ensure that the model always maintains the best performance in a changing environment.

[0092] 5. Deployment and Application:

[0093] This patent adopts a series of leading technologies and strategies in model deployment and practical applications to ensure the efficiency, adaptability, and long - term reliability of the model.

[0094] ● Intelligent Online Learning: The system design of this patent supports multi - task dynamic adaptation, can intelligently switch between different material recognition tasks, and adjust the calculation strategy according to task complexity. By introducing an online learning mechanism, the system can update the model's knowledge base in real - time to ensure its long - term stable operation in different application scenarios.

[0095] Automated Performance Monitoring: The system integrates advanced automated performance monitoring and autonomous optimization update technologies, capable of real-time detecting the running status of the model and the recognition accuracy. When it detects a decline in recognition performance or encounters new material descriptions, the system will automatically trigger the optimization update program to keep the model in the best state at all times.

[0096] Example 1

[0097] As Figure 1 shown, a method for identifying material attribute information using a large language model specifically includes:

[0098] S1 Extract the material information of the material, and use the recognition results of the word segmentation algorithm and the rule model to determine the extraction results of the material information in different types of attributes, and use the recognition results of the large language model to determine the predicted extraction results of the material information in different types of attributes;

[0099] The large model algorithm involved in the present invention is mainly optimized and upgraded based on the RoBERTa decoder of the Transformer model architecture. In the field of natural language processing (NLP), since the Transformer model was proposed in 2017, it has become one of the key technologies driving the development of this field. The original Transformer architecture, namely "Attention Is All You Need", first revealed the great potential of the self-attention mechanism in processing sequence data, especially outstanding in machine translation tasks. This model realizes the in-depth understanding of the input text and the generation of the translated text through the cooperation of the encoder and the decoder. In this process, the encoder plays a crucial role, responsible for extracting rich semantic information from the input text and transmitting this information in a continuous representation form to the decoder, and then generating the translated text in the target language.

[0100] Over time, models based on the original Transformer encoder module have emerged one after another, among which the most prominent ones are BERT and RoBERTa. These models have significantly improved the performance of NLP tasks through large-scale pre-training and fine-tuning. BERT uses masked language modeling and next sentence prediction tasks to pre-train on a large amount of text data, so as to be able to capture the deep context information of the input text. As an improved version of BERT, RoBERTa further improves the performance and efficiency of the model by means of increasing the amount of training data and optimizing the training process. These Transformer-based models have not only made breakthroughs in language understanding, but also demonstrated strong generalization ability in various downstream NLP tasks, bringing revolutionary progress to the field of natural language understanding.

[0101] 1. Original Transformer

[0102] The original Transformer architecture ("Attention Is All You Need", 2017) was developed for English-French and English-German language translation. It uses both an encoder and a decoder. The input text (i.e., the sentence to be translated) is first tokenized into individual word tokens, and then these tokens are encoded through an embedding layer. After that, it enters the encoder part. Next, a positional encoding vector is added to each embedded word. Then, these embeddings pass through a multi-head self-attention layer. After the multi-head attention layer, there is a residual and layer normalization (Add&normalize), which performs a layer of normalization operation and adds the original embeddings through a skip connection (also known as a residual connection or shortcut connection). Finally, after entering the "fully connected layer" (a small multi-layer perceptron composed of two fully connected layers with a non-linear activation function between them), the output is again "residual and layer normalized", and then the output is passed to the multi-head self-attention layer of the decoder module. The overall structure of the decoder part in the above figure is very similar to that of the encoder part, and the key difference is their input and output content. The encoder receives the input text to be translated, while the decoder is responsible for generating the translated text.

[0103] 2. Encoder

[0104] The encoder part in the original Transformer architecture is responsible for understanding and extracting relevant information from the input text. It outputs a continuous representation (embedding) of the input text, which is then passed to the decoder. Finally, the decoder generates the translated text (target language) based on the continuous representation received from the encoder.

[0105] Over the years, various encoder-only architectures have been developed based on the encoder module in the original Transformer model. Two of the most representative examples are BERT (Bidirectional Encoder Representations from Transformers for language understanding, 2018) and RoBERTa (Robustly Optimized BERT Pretraining Approach, 2018).

[0106] BERT (Bidirectional Encoder Representations from Transformers) is an encoder-only architecture based on the Transformer encoder module. It uses masked language modeling and next sentence prediction tasks to be pre-trained on a large text corpus.

[0107] Illustration of the masked language modeling pre-training objective used in BERT-style Transformers.

[0108] The main idea of masked language modeling is to randomly mask (or replace) some word tokens in the input sequence and train the model to predict the original masked tokens based on the context.

[0109] In addition to the masked language modeling pre-training task shown in the above figure, the next sentence prediction task requires the model to predict whether the sentence order of two randomly permuted sentences in the original document is correct. The masked language and next sentence pre-training objectives enable BERT to learn a large amount of context representations of the input text, and then these representations can be fine-tuned for various downstream tasks such as sentiment analysis, question answering, and named entity recognition.

[0110] RoBERTa (Robustly optimized BERT approach) is an optimized version of BERT. It maintains the same overall architecture as BERT but makes some training and optimization improvements, such as a larger batch size, more training data, and removing the next sentence prediction task. These improvements enable RoBERTa to have better performance and can handle various natural language understanding tasks better than BERT.

[0111] Although RoBERTa has made significant progress in the field of natural language processing, to achieve efficient identification and extraction of specific domain data such as material properties, more professional technologies and methods need to be combined. By applying large language models and multi-model fusion technologies, we can further improve the processing ability of this type of data. Next, we combine intelligent large language models with rule bases and rule engines to achieve standardized identification and extraction of material property data and ensure the accuracy of the data.

[0112] Apply large language models and multi-model fusion technologies to efficiently identify material property data using a rule base and a rule engine. Through the integration of intelligent large language model recognition, tokenization algorithm extraction, and rule engine data extraction, achieve the standard identification and extraction of material property data, and combine manual review to ensure the accuracy of the data.

[0113] The entire process of material property data identification, including model training, data extraction, application of feature libraries and tokenization algorithms, and construction of rule bases and rule engines, finally achieves the standardized identification and extraction of material property data and ensures high accuracy of the data.

[0114] Material Recognition Model Training

[0115] In the process of material attribute data recognition and extraction, an artificial intelligence model is applied, and through multi-model fusion technology, with the minimum number of labeled samples for training, an efficient material recognition effect is achieved. This method not only gives full play to the generalization and recognition ability of large models, but also, with the assistance of large models for sample annotation, optimizes the model training set, thereby improving the efficiency of the annotation work.

[0116] The optimization process of the model training set is as follows:

[0117] 1. Prepare training samples

[0118] Collect data: First, a large amount of labeled text data needs to be collected. This data should contain various entity types and clearly label the boundaries and categories of the entities.

[0119] Data cleaning: Clean the collected data to remove noise, invalid or duplicate samples to ensure data quality.

[0120] Data partitioning: Partition the cleaned data into a training set, a validation set, and a test set. Usually, the training set is used for model training, the validation set is used for model hyperparameter tuning and early stopping, and the test set is used for final evaluation of model performance.

[0121] 2. Initialize the parameter model

[0122] Select the RoBERTa model: Select an appropriate version of the RoBERTa model according to the task requirements, such as RoBERTa-base or RoBERTa-large. These models have different parameter scales and performances.

[0123] Initialize the task-specific layer: Add task-specific layers on top of the RoBERTa encoder, such as a fully connected layer or a CRF layer, for decoding (prediction) in the NER task. The parameters of these layers need to be randomly initialized.

[0124] Set hyperparameters: Set hyperparameters during model training, such as the learning rate, batch size, number of training epochs, etc.

[0125] 3. Start training

[0126] Load data: Load the training set data into memory to prepare for model training.

[0127] Forward propagation: Input the training data into the RoBERTa encoder to obtain the vector representation of the text. Then, pass these vector representations to the task-specific layer for decoding to obtain the NER prediction results.

[0128] Calculate loss: Calculate the loss value based on the prediction results and the actual annotations. Commonly used loss functions include cross-entropy loss.

[0129] Backpropagation: Update the model parameters through optimization algorithms such as gradient descent to minimize the loss value.

[0130] 4. Model parameter adjustment

[0131] Validation set evaluation: After each training round, the validation set is used to evaluate the model performance. This usually involves calculating indicators such as accuracy, recall, and F1 score on the validation set.

[0132] Hyperparameter tuning: Adjust hyperparameters such as learning rate decay, increase / decrease batch size, etc. based on the performance on the validation set to improve model performance.

[0133] Early stopping strategy: If the performance on the validation set does not improve significantly over multiple consecutive rounds, training is stopped early to prevent overfitting.

[0134] 5. Set verification indicators

[0135] Clarify the evaluation criteria: Before training begins, clarify the evaluation criteria for the NER task, such as accuracy, recall, F1 score, etc. These indicators will be used to evaluate the performance of the model on the validation set and test set.

[0136] Multi-indicator balance: Considering the particularity of the NER task, it may be necessary to balance indicators such as accuracy, recall, and F1 score to obtain a more comprehensive evaluation result.

[0137] 6. Verify training results

[0138] Test set evaluation: After model training is completed, the final performance of the model is evaluated using the test set. The test set should be completely independent of the training set and validation set to ensure the fairness of the evaluation results.

[0139] Result analysis: Analyze the evaluation results on the test set to understand the performance of the model in terms of different entity types, different text types, etc.

[0140] Model optimization: Based on the evaluation results and result analysis on the test set, the model is further optimized and adjusted to improve its performance in practical applications.

[0141] Material data recognition-dual model fusion algorithm

[0142] By adopting the technology of integrating NER model and large model, the automatic extraction of material-related attribute information is realized. In the NER task, RoBERTa is usually used as a feature extractor, and its output is used as the input for the subsequent task-specific layers. The process of the NER task is roughly as follows:

[0143] Text preprocessing: Perform preprocessing operations such as word segmentation and stop word removal on the original text so that the model can process it better.

[0144] Encoding: Input the preprocessed text into the RoBERTa model, and obtain the vector representation of the text through multiple layers of Transformer encoders. These vector representations contain rich context information and language features.

[0145] Decoding (prediction): Pass the output of RoBERTa to task-specific layers (such as fully connected layers, CRF layers, etc.) for decoding and prediction. In the NER task, these layers will predict whether each word in the text belongs to a certain entity category and the boundaries of the entity based on the vector representation provided by RoBERTa.

[0146] Post-processing: Perform post-processing on the prediction results, such as merging adjacent identical entities and removing impossible entity categories, to obtain the final NER results.

[0147] When processing the material title "Stanley adjustable wrench | 8 inches British | STMT94551-8-23", through the above steps, the system can successfully identify and extract key attributes such as material model, name, brand, etc., and clean the data in combination with manual review. Through the collaborative work of the dual models, the comprehensive extraction accuracy has been increased to more than 90%, ensuring high recognition rate and accuracy of the data.

[0148] Material data recognition - feature library + word segmentation algorithm

[0149] Given the possible probabilistic limitations of the large language model in the recognition process, for the "enumerable" attribute values in material attributes, such as material brand data, its value range is relatively fixed. For the recognition of such material attribute data, a feature library combined with a word segmentation algorithm is used to optimize the recognition accuracy.

[0150] The feature library, as the core element of machine learning and data analysis, is a collection of data features (data values) that have been systematically collected and sorted in advance.

[0151] The word segmentation algorithm is one of the commonly used AI algorithms in natural language processing. It depends on the word library to identify the text content and extracts key words through algorithm logic.

[0152] The optimization process of material attribute recognition is as follows:

[0153] 1. Feature Library Construction

[0154] Step 1.1 Data Collection

[0155] First, collect material brand data from various reliable sources (such as official databases, industry standards, market research reports, etc.).

[0156] Ensure that the collected data is extensive and accurate.

[0157] Step 1.2 Data Cleaning and Sorting

[0158] Clean and sort the collected data, removing duplicate, incorrect, or irrelevant items.

[0159] Format the data to ensure the consistency and standardization of each brand name.

[0160] Step 1.3 Feature Library Construction

[0161] Construct the sorted brand data into a feature library, with each brand name as an independent feature item.

[0162] The feature library can be stored as a database table, file, or data structure in memory for fast retrieval and matching.

[0163] 2. Selection and Implementation of Word Segmentation Algorithm

[0164] Step 2.1 Selection of Word Segmentation Algorithm

[0165] Select a word segmentation algorithm suitable for natural language processing tasks, such as rule-based word segmentation, statistics-based word segmentation, or deep learning word segmentation algorithm.

[0166] Considering the particularity of material brand names (such as may contain non-standard words, abbreviations, etc.), it may be necessary to customize the word segmentation algorithm or word library.

[0167] Step 2.2 Customized Word Library

[0168] Customize or expand the word library of the word segmentation algorithm according to the brand names in the feature library.

[0169] Ensure that the word library contains all important brand names so that the word segmentation algorithm can accurately identify them.

[0170] Step 2.3 Implementation of Word Segmentation Algorithm

[0171] Implement the word segmentation algorithm and integrate it into the material attribute recognition system.

[0172] When processing text, first use the word segmentation algorithm to segment the text and extract key words.

[0173] 3. Recognition in Combination with Large Language Model

[0174] Step 3.1 Text Preprocessing

[0175] Preprocess the text of the material attributes to be recognized, including removing irrelevant characters, unifying the encoding format, etc.

[0176] Step 3.2 Preliminary Recognition

[0177] Use a large language model to perform preliminary recognition on the preprocessed text to obtain a candidate set of possible brand names.

[0178] This step may produce some misidentifications or omissions because of the probabilistic limitations of the large language model.

[0179] Step 3.3 Feature Library Matching

[0180] Match the candidate set obtained from the preliminary recognition with the brand names in the feature library.

[0181] For the brand names that match successfully, confirm their accuracy and extract relevant information.

[0182] For the candidates that do not match successfully, further analysis or manual verification can be carried out.

[0183] Step 3.4 Result Optimization

[0184] According to the results of the feature library matching, correct and optimize the results of the preliminary recognition.

[0185] Confidence scores, context information, or other auxiliary information can be considered to improve the accuracy of recognition.

[0186] 4. Performance Evaluation and Iteration

[0187] Step 4.1 Evaluate the Recognition Effect

[0188] Use the test set to evaluate the performance of the optimized recognition system, and calculate metrics such as accuracy, recall rate, and F1 score.

[0189] Step 4.2 Iterative Optimization

[0190] Adjust the parameters of the feature library, word segmentation algorithm, or large language model according to the evaluation results.

[0191] Repeat the above steps to continuously optimize the performance of the recognition system.

[0192] Material Data Recognition - Rule Library + Rule Engine

[0193] For the case where some material attribute data follows specific rules, such as the wrench size usually following the composition form of "number" plus "inch", the constructed rule library is combined with the rule engine to identify and extract the attribute values. This solution will significantly improve the recognition efficiency and accuracy of such material data.

[0194] Rule library definition: A database used to define data recognition rules, where each rule corresponds to a data processing method.

[0195] Rule engine: Technical personnel write the processing method for data strings according to the defined rules to find the parts that match the rules from the strings.

[0196] Material attribute data recognition and extraction process

[0197] 1. Rule library construction

[0198] Step 1.1 Rule definition

[0199] Identify the material attributes to be recognized and the rules they follow. For example, the rule for wrench size is "number + inch".

[0200] Define the specific format and scope of the rules. For example, the number can be from 1 to 999, and inch is the unit and always follows the number.

[0201] Step 1.2 Rule encoding

[0202] Encode the defined rules into a format that can be understood by the computer and store them in the rule library.

[0203] The rule library can be a database table, a configuration file, or a data structure in memory, facilitating access by the rule engine.

[0204] 2. Rule engine development

[0205] Step 2.1 Engine design

[0206] Design the architecture of the rule engine, including the input interface, rule matching module, processing module, and output interface.

[0207] Determine how the rule engine receives the data string to be processed, how to traverse the rule library for matching, and how to handle the case of successful matching.

[0208] Step 2.2 Rule implementation

[0209] Write the corresponding processing logic according to the rules in the rule library.

[0210] Use technical means such as regular expressions and string processing functions to identify the parts that match the rules from the input data string.

[0211] Step 2.3 Integration Testing

[0212] Conduct unit testing and integration testing on the rule engine to ensure that it can correctly identify and process material attribute data that complies with the rules.

[0213] 3. Combine with large language model preprocessing

[0214] Step 3.1 Data Preprocessing

[0215] Before sending the material attribute data into the rule engine, first use the large language model for preprocessing.

[0216] The large language model can help correct spelling mistakes, identify abbreviations or aliases, and may provide additional context information, thus improving the accuracy of subsequent rule matching.

[0217] Step 3.2 Preliminary Identification

[0218] The large language model can conduct preliminary identification on the data, marking the areas or keywords that may contain material attribute information.

[0219] These preliminary identification results can be used as the input of the rule engine, reducing invalid matches and improving efficiency.

[0220] 4. Rule Engine Execution and Result Processing

[0221] Step 4.1 Rule Matching

[0222] Send the preprocessed data string into the rule engine to perform rule matching.

[0223] The rule engine traverses the rule base, finds the rules that match the data string, and executes the corresponding processing logic.

[0224] Step 4.2 Result Extraction

[0225] Extract the material attribute values from the data string that matches successfully, such as the specific size of the wrench.

[0226] Format and store the extraction results for subsequent use or display.

[0227] Step 4.3 Verification and Feedback

[0228] Verify the extraction results to ensure their accuracy and integrity.

[0229] If errors or omissions are found, they can be fed back to the rule base or the rule engine for adjustment and optimization.

[0230] 5. Performance Evaluation and Iteration

[0231] Step 5.1 Performance Evaluation

[0232] Evaluate the recognition efficiency and accuracy of the rule engine using the test set.

[0233] Calculate key metrics such as recognition speed, accuracy, and recall.

[0234] Step 5.2 Iterative Optimization

[0235] Adjust and optimize the rule base, rule engine, or large language model according to the evaluation results.

[0236] Repeat the above process to continuously improve the recognition efficiency and accuracy of material attribute data.

[0237] Recognition Data Fusion and Confirmation

[0238] Through the fusion of data recognized by the intelligent large language model, extracted by the word segmentation algorithm, and extracted by the rule engine, the standard recognition and extraction of material attribute data are finally realized. At the same time, the accuracy of the data is ensured by combining the manual review ability.

[0239] Using large language models and multi-model fusion technologies, significant advantages and effects are demonstrated in the field of material attribute data recognition compared to traditional recognition methods:

[0240] High Efficiency and Accuracy:

[0241] Large Language Model: Trained on a large amount of text data, the large language model can understand and generate natural language text, demonstrating excellent language processing capabilities and generalization performance. During the process of recognizing material attribute data, the model can accurately identify key information such as material names, specifications, and uses, thereby improving the recognition accuracy.

[0242] Multi-Model Fusion: Combining the advantages of multiple models, such as large language models, word segmentation algorithms, and rule engines, can achieve more comprehensive data extraction and more accurate attribute recognition. The multi-model fusion technology can make full use of the strengths of different models, make up for the deficiencies of single models, and thus improve the overall recognition efficiency and accuracy.

[0243] Intelligence and Automation:

[0244] Intelligent Recognition: The large language model can automatically understand the meaning of material attribute data and generate corresponding recognition results, completing most of the recognition work without manual intervention. This significantly improves the intelligence level of material attribute data recognition.

[0245] Automated Process: From data input to result output, the entire recognition process can be highly automated, reducing the time and cost of manual operations.

[0246] Scalability and Flexibility:

[0247] Rule Library and Rule Engine: By constructing a rule library and a rule engine, the recognition rules and extraction logic of material attribute data can be flexibly defined. When the material type or attributes change, only the rule library needs to be updated to adapt to the new recognition requirements, without the need for large-scale modification of the system.

[0248] Word Segmentation Algorithm: The word segmentation algorithm can be customized and optimized according to different language habits and rules to meet the recognition requirements of material attribute data in different fields.

[0249] Data Quality Assurance:

[0250] Manual Review: Combining the manual review process can further ensure the accuracy of material attribute data. Manual review can promptly detect and correct errors or omissions that may occur in the automatic recognition process.

[0251] Data Standardization: By standardizing the recognition and extraction process, the consistency and comparability of material attribute data can be ensured, providing strong support for subsequent data analysis and applications.

[0252] Improve Recognition Efficiency: The automated and intelligent recognition process significantly shortens the recognition time of material attribute data and improves work efficiency.

[0253] Reduce Error Rate: The combination of multi-model fusion technology and the manual review process effectively reduces the error rate in the recognition process and improves data accuracy.

[0254] Optimize Resource Allocation: Through the automated recognition process, the dependence on human resources can be reduced, resource allocation can be optimized, and enterprise operating costs can be lowered.

[0255] Promote Digital Transformation: The intelligent recognition and extraction of material attribute data is an important part of promoting the digital transformation of enterprises. Through the application of this technology, enterprises can manage and utilize material data resources more efficiently and enhance their overall competitiveness.

[0256] In summary, using large language models and multi-model fusion technology for material attribute data recognition has significant advantages and effects compared to traditional methods. With the continuous development and improvement of technology, this method will be widely applied and promoted in more fields.

[0257] Furthermore, the attributes include material model, name, brand, format, and size.

[0258] Specifically, the large language model is built by fusing the NER model with the large model.

[0259] Specifically, as Figure 2 shown, the method for determining the recognition result of the word segmentation algorithm is:

[0260] Build a feature library:

[0261] Construct a feature library by using the material brand data collected from various sources as the basis. Each brand name is used as an independent feature item and stored as a data structure in a database table, file, or memory.

[0262] Tokenization algorithm recognition:

[0263] Use a tokenization algorithm to perform tokenization on the text, extract key words, and use a large language model to perform preliminary recognition on the preprocessed key words to obtain a candidate set of possible brand names.

[0264] Match the candidate set of brand names obtained from the preliminary recognition with the brand names in the feature library, and determine the recognition result of the tokenization algorithm according to the matching result of the feature library.

[0265] Optionally, the sources include official databases, industry standards, and market research reports.

[0266] It should be noted that as Figure 3 shown, the specific steps for building the feature library are as follows:

[0267] Collect material brand data from various sources, clean and organize the collected material brand data, remove duplicate items, error items, or irrelevant items, and format the collected material brand data.

[0268] Construct the sorted material brand data into a feature library. Each brand name is used as an independent feature item, and the feature library constructed by the feature items is stored as a data structure in a database table, file, or memory.

[0269] It should be noted that the method for determining the recognition result of the rule model is:

[0270] Rule library construction:

[0271] Clarify the material attributes to be recognized and the rules they follow, encode the defined rules into a computer-readable format, store them in the rule library, and construct them through a database table, configuration file, or data structure in memory.

[0272] Rule engine development:

[0273] Design the architecture of the rule engine and write the corresponding processing logic according to the rules in the rule library.

[0274] Combined with large language model preprocessing:

[0275] Before sending the material information into the rule engine, it is first preprocessed using a large language model to preliminarily identify the data and mark the areas or keywords that may contain material attribute information.

[0276] Rule engine execution and result processing:

[0277] Send the preprocessed data string into the rule engine, perform rule matching, and extract the material attribute values from the successfully matched data strings.

[0278] Furthermore, the architecture of the rule engine includes an input interface, a rule matching module, a processing module, and an output interface.

[0279] Furthermore, as Figure 4 shown, the specific steps for developing the rule engine are:

[0280] Design the architecture of the rule engine to determine how the rule engine receives the data string to be processed, how to traverse the rule library for matching, and how to handle successful matches.

[0281] According to the rules in the rule library, write the corresponding processing logic, and use technical means such as regular expressions and string processing functions to identify the parts that match the rules from the input data string.

[0282] Conduct unit tests and integration tests on the rule engine to obtain the development processing results of the rule engine.

[0283] Specifically, the specific steps for constructing the large language model are:

[0284] Text preprocessing: Perform word segmentation and stop word removal preprocessing operations on the material information.

[0285] Encoding: Input the preprocessed text into the RoBERTa model to obtain the vector representation of the text through multiple layers of Transformer encoders.

[0286] Decoding: Pass the output of RoBERTa to the task-specific layer for decoding and prediction to obtain the prediction result.

[0287] Post-processing: Post-process the prediction result, merge adjacent identical entities, remove impossible entity categories to obtain the final NER result, and use it as the output result of the large language model.

[0288] Furthermore, the specific layer includes a fully connected layer and a CRF layer.

[0289] In addition, it should be noted that the material type of the material is determined according to the matching situation between the prediction extraction results of different types of attributes in the material and the material type.

[0290] Based on the prediction extraction results and extraction results of different types of attributes, S2 determines the material type of the material. According to the analysis result of the material type, it determines the deviation situation of the prediction extraction result corresponding to the material type, and when the recognition accuracy rate of the material type meets the requirements by using the deviation situation, it proceeds to the next step;

[0291] Further, determining that the recognition accuracy rate of the material type meets the requirements specifically includes:

[0292] Based on the deviation situation of the prediction extraction result corresponding to the material type, it determines the material corresponding to the material type and uses it as the historical matching material;

[0293] According to the deviation situation of the prediction extraction result of the historical matching material, it conducts historical matching materials with deviation in the prediction extraction result and uses them as deviation materials;

[0294] Based on the quantity of the deviation materials, it determines whether the recognition accuracy rate of the material type meets the requirements.

[0295] Optionally, when the quantity of the deviation materials is greater than the preset quantity of deviation materials, it is determined that the recognition accuracy rate of the material type does not meet the requirements.

[0296] Specifically, when the recognition accuracy rate of the material type does not meet the requirements, the material is subjected to manual verification processing.

[0297] It should be noted that determining that the recognition accuracy rate of the material type meets the requirements specifically includes:

[0298] Based on the deviation situation of the prediction extraction result corresponding to the material type, it determines the material corresponding to the material type and uses it as the historical matching material;

[0299] According to the deviation situation of the prediction extraction result of the historical matching material, it conducts historical matching materials with deviation in the prediction extraction result and uses them as deviation materials. When the quantity of the deviation materials does not meet the requirements, it is determined that the recognition accuracy rate of the material type does not meet the requirements;

[0300] When the quantity of the deviation materials meets the requirements:

[0301] Based on the deviation situation of the prediction extraction results of different deviation materials, it determines the prediction extraction results with deviation for different deviation materials and uses them as deviation extraction results. When the total quantity of the deviation extraction results of different deviation materials does not meet the requirements, it is determined that the recognition accuracy rate of the material type does not meet the requirements;

[0302] When the total quantity of the deviation extraction results of different deviation materials meets the requirements:

[0303] According to the deviation situation of the prediction extraction results of different deviation materials, determine the number of prediction extraction results with deviations for different types of attributes. When there are attributes for which the number of prediction extraction results with deviations does not meet the requirements, it is determined that the recognition accuracy rate of the material type does not meet the requirements;

[0304] When there are no attributes for which the number of prediction extraction results with deviations does not meet the requirements:

[0305] Based on the deviation situation of the prediction extraction results of different deviation materials, when it is determined that there are no deviation materials with a deviation amount not meeting the requirements, it is determined that the recognition accuracy rate of the material type meets the requirements;

[0306] When there are deviation materials with a deviation amount not meeting the requirements:

[0307] Based on the deviation situation of the prediction extraction results of different deviation materials, determine the material deviation coefficients of different deviation materials, determine the comprehensive deviation coefficient according to the material deviation coefficients of different deviation materials, and based on the comprehensive deviation coefficient, determine whether the recognition accuracy rate of the material type meets the requirements.

[0308] S3 Determine the deviation situation between the prediction extraction results and the extraction results of the material for different types of attributes, and in combination with the matching situation between the prediction extraction results for different types of attributes, when it is determined that the reliability coefficient of the prediction extraction results meets the requirements, proceed to the next step;

[0309] Furthermore, the matching situation between the prediction extraction results is determined according to the matching situation of different prediction extraction results for the same type of material. Specifically, the prediction extraction results that cannot be for the same type of material are used as the matching deviation extraction results;

[0310] Specifically, determining that the reliability coefficient of the prediction extraction results meets the requirements specifically includes:

[0311] Based on the deviation situation between the prediction extraction results and the extraction results of the material for different types of attributes, determine the attributes with deviations and use them as deviation attributes, and determine the attribute deviation coefficient according to the proportion of the number of the deviation attributes;

[0312] According to the matching situation between the prediction extraction results for different types of attributes, use the prediction extraction results that cannot be for the same type of material as the matching deviation extraction results, and determine the matching deviation coefficient using the proportion of the number of the matching deviation extraction results;

[0313] Based on the average value of the matching deviation coefficient and the attribute deviation coefficient, determine the comprehensive deviation coefficient, and use the comprehensive deviation coefficient to determine whether the reliability coefficient of the prediction extraction results meets the requirements.

[0314] Further, when the comprehensive deviation coefficient is greater than the preset deviation coefficient threshold, it is determined that the reliability coefficient of the predicted extraction result does not meet the requirements.

[0315] In addition, it should be noted that when the reliability coefficient of the predicted extraction result does not meet the requirements, the material is manually verified.

[0316] S4 determines the manual verification ratio of the material type corresponding to the material according to the deviation situation of the predicted extraction result corresponding to the material type and the deviation situation of the extraction result, and determines whether the material needs to be manually verified by using the manual verification ratio and the reliability coefficient of the material.

[0317] Further, the method for determining the manual verification ratio of the material type corresponding to the material is as follows:

[0318] Based on the deviation situation of the predicted extraction result corresponding to the material type and the deviation situation of the extraction result, determine the proportion of the number of deviations in the predicted extraction result and the proportion of the number of deviations in the extraction result for different types of attributes in the material type;

[0319] Determine the attribute deviation coefficient of different types of attributes according to the proportion of the number of deviations in the predicted extraction result and the proportion of the number of deviations in the extraction result;

[0320] Take the average of the attribute deviation coefficients of different types of attributes as the average deviation coefficient, and use the average deviation coefficient to determine the manual verification ratio of the material type corresponding to the material.

[0321] Specifically, using the average deviation coefficient to determine the manual verification ratio of the material type corresponding to the material specifically includes:

[0322] Based on the average deviation coefficient, determine the preset verification ratio under the average deviation coefficient;

[0323] Use the preset verification ratio to determine the manual verification ratio of the material type corresponding to the material.

[0324] In another embodiment, the method for determining the manual verification ratio of the material type corresponding to the material is as follows:

[0325] Take the material corresponding to the material type as the matching material, and based on the deviation situation of the predicted extraction result and the deviation situation of the extraction result in the matching material, determine the matching materials with deviations, and use them as the deviation matching materials. When the number of the deviation matching materials is greater than the preset number of matching materials, use the preset verification ratio to determine the manual verification ratio of the material type corresponding to the material;

[0326] When the quantity of the materials matching the deviation is not greater than the preset quantity of matching materials:

[0327] Determine the attributes with deviations based on the proportion of the quantity of different types of attributes in the predicted extraction results with deviations in the material type and the proportion of the quantity of extraction results with deviations. When the quantity of attributes with deviations does not meet the requirements, the manual verification proportion of the material type corresponding to the material is determined using the preset verification proportion;

[0328] When the quantity of attributes with deviations meets the requirements:

[0329] Determine the attribute deviation coefficients of different types of attributes based on the proportion of the quantity of different types of attributes in the predicted extraction results with deviations in the material type and the proportion of the quantity of extraction results with deviations. When there are attributes with attribute deviation coefficients that do not meet the requirements, the manual verification proportion of the material type corresponding to the material is determined using the preset verification proportion;

[0330] When there are no attributes with attribute deviation coefficients that do not meet the requirements:

[0331] Based on the attribute deviation coefficients of different types of attributes, determine the attributes with attribute deviation coefficients within the preset deviation coefficient range. When the quantity of attributes with attribute deviation coefficients within the preset deviation coefficient range does not meet the requirements, the manual verification proportion of the material type corresponding to the material is determined using the preset verification proportion;

[0332] When the quantity of attributes with attribute deviation coefficients within the preset deviation coefficient range meets the requirements:

[0333] Take the average of the attribute deviation coefficients of different types of attributes as the average deviation coefficient, and use the average deviation coefficient to determine the manual verification proportion of the material type corresponding to the material.

[0334] Further, determining whether the material needs to be manually verified specifically includes:

[0335] Use the manual verification proportion of the materials of the material type and the quantity of the materials of the material type to determine the manual verification quantity of the materials of the material type;

[0336] According to the manual verification quantity, determine the verification targets of the materials of the material type in descending order of the reliability coefficient of the materials;

[0337] Use the verification targets to determine whether the material needs to be manually verified.

[0338] A further technical solution is that when the material belongs to the verification target, it is determined that the material needs to be manually verified.

[0339] Embodiment 2

[0340] On the other hand, the present application provides a system for identifying material attribute information using a large language model, which is applied to the above method for identifying material attribute information using a large language model, and specifically includes:

[0341] A material information parsing module and a manual verification module;

[0342] Among them, the material information parsing module is responsible for extracting the material information of the material, and using the recognition results of the word segmentation algorithm and the rule model to determine the extraction results of the material information in different types of attributes, and using the recognition results of the large language model to determine the predicted extraction results of the material information in different types of attributes;

[0343] The manual verification module is responsible for manually verifying the material.

[0344] In the description of this specification, the description of terms such as "one embodiment" and "one preferred embodiment" means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the embodiments of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or instance. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0345] The above are only the preferred embodiments of the embodiments of the present invention, and are not used to limit the embodiments of the present invention. For those skilled in the art, the embodiments of the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the embodiments of the present invention shall be included in the protection scope of the embodiments of the present invention.

Claims

1. A method for identifying material attribute information using a large language model, characterized in that: Specifically include: Extract material information of the material, and determine the extraction results of different types of attributes of the material information using the word segmentation algorithm and the recognition results of the rule model, and determine the predicted extraction results of the different types of attributes of the material information using the recognition results of the large language model; Based on the prediction and extraction results of different types of attributes and the extraction results, the material type of the material is determined, and according to the analysis results of the material type, the deviation of the prediction and extraction results corresponding to the material type is determined, and when the recognition accuracy of the material type is determined by using the deviation to meet the requirements, the next step is entered; Determine the deviation between the predicted extraction results and the extraction results of the material in different types of attributes, and combine the matching between the predicted extraction results of different types of attributes to determine that the reliability coefficient of the predicted extraction result meets the requirements, and then proceed to the next step; According to the deviation of the predicted extraction result corresponding to the material type and the deviation of the extraction result, the manual verification ratio of the material type corresponding to the material is determined, and the manual verification ratio and the reliability coefficient of the material are used to determine whether the material needs to be manually verified.

2. The method for identifying material attribute information using a large language model according to claim 1, characterized in that: The attributes include material model, name, brand, standard, and size.

3. The method for identifying material attribute information using a large language model according to claim 1, characterized in that: The large language model is constructed by fusing the NER model with the large model.

4. The method for identifying material attribute information using a large language model according to claim 1, characterized in that: The method for determining the recognition result of the word segmentation algorithm is: Build the feature library: Collect material brand data from various sources as the basis to build a feature library, with each brand name as an independent feature item, and store it as a database table, file or in-memory data structure; Word segmentation algorithm recognition: Use the word segmentation algorithm to segment the text, extract key words, and use the large language model to perform preliminary recognition on the pre-processed key words to obtain possible brand name candidate sets; The brand name candidate set obtained by preliminary recognition is matched with the brand names in the feature library, and the recognition result of the word segmentation algorithm is determined according to the result of the feature library matching.

5. The method for identifying material attribute information using a large language model as claimed in claim 4, characterized in that: The sources mentioned include official databases, industry standards, and market research reports.

6. The method for identifying material attribute information using a large language model according to claim 1, characterized in that: The specific steps of constructing the feature library are: Collect material brand data from various sources, clean and organize the collected material brand data, remove duplicates, erroneous items or irrelevant items, and format the collected material brand data; The sorted material brand data is constructed into a feature library, and each brand name is used as an independent feature item. The feature library constructed using the feature items is stored as a database table, file, or data structure in memory.

7. The method for identifying material attribute information using a large language model according to claim 1, characterized in that: The method for determining the recognition result of the rule model is: Rule base construction: Clarify the material attributes that need to be identified and the rules they follow, encode the defined rules into a computer-understandable format, store them in a rule base, and build them through database tables, configuration files, or in-memory data structures. Rules Engine Development: Design the architecture of the rule engine and write the corresponding processing logic according to the rules in the rule base; Combined with large language model preprocessing: Before feeding the material information into the rule engine, a large language model is used for preprocessing to perform preliminary identification of the data and mark out areas or keywords that may contain material attribute information; Rule engine execution and result processing: The preprocessed data string is sent to the rule engine to perform rule matching, and the material attribute value is extracted from the successfully matched data string.

8. The method for identifying material attribute information using a large language model as claimed in claim 7, characterized in that: The architecture of the rule engine includes an input interface, a rule matching module, a processing module and an output interface.

9. The method for identifying material attribute information using a large language model according to claim 7, characterized in that: The specific steps of rule engine development are: Design the architecture of the rule engine to determine how the rule engine receives the data string to be processed, how to traverse the rule base for matching, and how to handle successful matching situations; According to the rules in the rule base, write the corresponding processing logic, use regular expressions and string processing functions to identify the parts that match the rules from the input data string; Perform unit testing and integration testing on the rule engine to obtain the development and processing results of the rule engine.

10. A system for identifying material attribute information using a large language model, applied to a method for identifying material attribute information using a large language model as claimed in any one of claims 1 to 9, characterized in that: Specifically include: Material information analysis module, manual verification module; The material information parsing module is responsible for extracting the material information of the material, and using the word segmentation algorithm and the recognition results of the rule model to determine the extraction results of the material information in different types of attributes, and using the recognition results of the large language model to determine the predicted extraction results of the material information in different types of attributes; The manual verification module is responsible for performing manual verification on the materials.