Material main data optimization and standardization processing method and system based on multi-dimensional matching rule
Through the material master data optimization method based on multi-dimensional matching rules, problems such as duplicate codes, wrong codes, and missing entries in material master data management are solved, the accuracy and structural rationality of the data are improved, and the operating efficiency and management costs of the enterprise are reduced.
Patent Information
- Application Number
- CN202510755180.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-09-26
AI Technical Summary
Problems such as duplicate codes, incorrect codes, missing entries, incorrect or missing attributes exist in material master data management, which affect the company's operating efficiency and management costs.
A material master data optimization and standardization processing method based on multi-dimensional matching rules is adopted, including data preparation, standard library construction, matching rule definition, data matching, revision and verification, etc., using programming language and natural language processing technology to improve matching accuracy.
The problems of duplicate codes, wrong codes and missing data were solved to ensure data accuracy, deduplication rate reached 10%, the rationality of data structure was improved, data optimization and cleaning were completed, and the final data reached 800,000 items.
Smart Images

Figure CN120705193A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data management technology, and in particular to a material master data optimization and standardization processing method and system based on multi-dimensional matching rules. Background Art
[0002] In materials management, the quality of material master data is directly related to a company's operational efficiency and management costs. However, as a company's business grows, the volume of purchased material master data continues to increase, making management more difficult. This leads to issues such as duplicate codes, incorrect codes, missing entries, and incorrect or missing attributes, which seriously impact the quality of master data and, in turn, the company's operational efficiency.
[0003] Therefore, a material master data optimization and standardization processing method and system based on multi-dimensional matching rules are needed to solve the above technical problems. Summary of the Invention
[0004] The present invention provides a material master data optimization and standardization processing method and system based on multi-dimensional matching rules, aiming to solve the above technical problems.
[0005] The present invention is implemented as follows: a material master data optimization and standardization processing method based on multi-dimensional matching rules, comprising the following steps:
[0006] S1. Data preparation: Collect and organize material master data to ensure that the data has fields that match the standard library;
[0007] S2. Standard library construction: Build a standard library containing standard data based on national standards and industry standards documents, and regularly update and maintain the standard library;
[0008] S3. Define matching rules: Based on business needs and data characteristics, define multi-dimensional matching rules including exact matching and fuzzy matching;
[0009] S4. Data matching: Use programming language to read material master data and match it with the standard library according to the defined matching rules;
[0010] S5. Data revision: When a matching entry is found in the standard library, relevant information of the entry is extracted to revise deficiencies in the material master data;
[0011] S6. Data Verification and Backup: Verify the revised data to ensure its accuracy and completeness, and perform data backup;
[0012] S7. Cycle optimization: Regularly check and update the standard library, monitor changes in material master data, and make corresponding revisions and maintenance.
[0013] Preferably, the construction of the standard library in step S2 further includes the following sub-steps:
[0014] S2.1 Document collection and organization: Collect all relevant national and industry standard documents and organize them to ensure their completeness and accuracy;
[0015] S2.2 Document parsing and data extraction: Parse electronic documents, extract required data elements, and perform cleaning and standardization;
[0016] S2.3 Data import and storage: Import the cleaned and standardized data into the database and establish a reasonable database table structure and relationships.
[0017] Preferably, the multi-dimensional matching rules in step S3 include exact matching based on a single field, fuzzy matching based on multiple fields, regular expression matching based on attribute fields, and text similarity matching based on material description fields.
[0018] Preferably, the data matching in step S4 further includes preprocessing the material description field using natural language processing technology to improve matching accuracy.
[0019] Preferably, the data revision in step S5 includes updating field values, adding new fields or performing other necessary operations to ensure that the material master data remains consistent with the data in the standard library.
[0020] The present invention also proposes a material master data optimization and standardization processing system based on multi-dimensional matching rules, comprising:
[0021] Data preparation module, used to collect and organize material master data;
[0022] Standard library construction module, used to build standard libraries based on national standards and industry standards documents, and to perform regular updates and maintenance;
[0023] Matching rule definition module, used to define multi-dimensional matching rules;
[0024] Data matching module, used to match material master data with the standard library according to defined matching rules;
[0025] The data revision module is used to extract matching entry information from the standard library and revise deficiencies in the material master data;
[0026] Data verification and backup module, used to verify the revised data, ensure the accuracy and integrity of the data, and perform data backup;
[0027] The cycle optimization module is used to regularly check and update the standard library, monitor changes in material master data, and make corresponding revisions and maintenance.
[0028] Preferably, the standard library construction module also includes a document collection and organization unit, a document parsing and data extraction unit, and a data import and storage unit.
[0029] Preferably, the multi-dimensional matching rules defined by the matching rule definition module include exact matching based on a single field, fuzzy matching based on multiple fields, regular expression matching based on attribute fields, and text similarity matching based on material description fields.
[0030] Preferably, the data matching module further includes a natural language processing unit for preprocessing the material description field to improve matching accuracy.
[0031] Preferably, the data revision module includes a field updating unit, a field adding unit or other necessary operation units to ensure that the material master data remains consistent with the data in the standard library.
[0032] Compared with related technologies, the material master data optimization and standardization processing method and system based on multi-dimensional matching rules provided by the present invention have the following beneficial effects:
[0033] The material master data optimization and standardization processing method and system based on multi-dimensional matching rules proposed in the present invention focus on solving the problems of duplicate codes, wrong codes, and missing data, ensuring data accuracy, and removing duplicate data to 10% of the overall data; completing the material architecture optimization, solving the coding technology attribute problem, and improving the rationality of the data structure, removing duplicate data to 10% of the overall data; achieving comprehensive optimization and cleaning of the data, ensuring that the final data reaches 800,000 items. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 A schematic diagram of the structure of a material master data optimization and standardization processing method and system based on multi-dimensional matching rules provided by the present invention;
[0035] Figure 2 This is an architectural diagram of the technical solution of the present invention;
[0036] Figure 3 This is a diagram showing the objectives of the three stages of the present invention;
[0037] Figure 4 A flowchart of database performance optimization in the present invention;
[0038] Figure 5 This is an example diagram of the matching rules in the present invention;
[0039] Figure 6 This is an example diagram of converting non-text format into text format using OCR (Optical Character Recognition) technology;
[0040] Figure 7 This is the standard document preprocessing flow chart. DETAILED DESCRIPTION
[0041] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0042] Please refer to Figure 1 , Figure 1 The flowchart of the material master data optimization and standardization processing method based on multi-dimensional matching rules provided by the present invention. The material master data optimization and standardization processing method based on multi-dimensional matching rules includes the following steps:
[0043] S1. Data preparation: Collect and organize material master data to ensure that the data has fields that match the standard library;
[0044] S2. Standard library construction: Based on national standards and industry standards, a standard library containing standard data is constructed, and the standard library is regularly updated and maintained. The specific steps include:
[0045] S2.1 Document collection and organization: Collect all relevant national and industry standard documents and organize them to ensure their completeness and accuracy;
[0046] S2.2 Document parsing and data extraction: Parse electronic documents, extract required data elements, and perform cleaning and standardization;
[0047] S2.3 Data import and storage: Import the cleaned and standardized data into the database and establish a reasonable database table structure and relationships;
[0048] S3. Define matching rules: Based on business needs and data characteristics, define multi-dimensional matching rules, including exact matching and fuzzy matching. Multi-dimensional matching rules include exact matching based on a single field, fuzzy matching based on multiple fields, regular expression matching based on attribute fields, and text similarity matching based on material description fields.
[0049] S4. Data matching: Use a programming language to read the material master data and match it with the standard library according to the defined matching rules. The data matching in this step also includes pre-processing the material description fields using natural language processing technology to improve matching accuracy.
[0050] S5. Data revision: When a matching entry is found in the standard library, relevant information about the entry is extracted to correct deficiencies in the material master data. Data revision in this step includes updating field values, adding new fields, or performing other necessary operations to ensure that the material master data is consistent with the data in the standard library.
[0051] S6. Data Verification and Backup: Verify the revised data to ensure its accuracy and completeness, and perform data backup;
[0052] S7. Cycle optimization: Regularly check and update the standard library, monitor changes in material master data, and make corresponding revisions and maintenance.
[0053] The following will specifically describe the material master data optimization and standardization processing method based on multi-dimensional matching rules proposed in the present invention.
[0054] 1. The objectives of the present invention are divided into three stages (such as Figure 3 shown):
[0055] Phase 1: Focus on solving duplicate codes, wrong codes, and missing data to ensure data accuracy. Duplicate data should reach 10% of the total data.
[0056] Phase 2: Complete the optimization of material structure, solve the problem of coding technology attributes, and improve the rationality of data structure. Deduplication of data reaches 10% of the total data.
[0057] Phase 3: Complete the rectification of the above issues, comprehensively optimize and clean the data, and ensure that the final data reaches 800,000 items.
[0058] 2. The overall solution architecture of the present invention:
[0059] The overall solution of the present invention is from the collection of standard documents to the processing and storage, and finally the matching and revision of material master data. The overall architecture diagram is as follows Figure 2 shown.
[0060] The overall architecture for building a standards library can be summarized as follows: First, based on business needs and data characteristics, determine the goals and scope of the standards library, including the types and number of standards to be covered, as well as the sources and formats of the data. Once these standards documents are collected, they are digitized and converted to digital form. It's worth noting that if the standards documents are only available in paper form, they can be digitized through scanning. Non-text documents like PDFs and JPGs can be converted to electronic text using optical character recognition (OCR) technology.
[0061] Secondly, the standard library architecture is designed, including database design, data model, table structure, and indexing strategy, to ensure efficient data storage and retrieval. For unstructured data, if it consists of large sections of natural language, NLP (natural language processing) technology can be used to extract key text and convert it into structured fields. This meets database storage requirements and completes the conversion of unstructured text into structured text.
[0062] Next, we developed data import and processing programs to integrate standard data from various sources into the standard library, performing necessary cleansing, conversion, and validation to ensure data accuracy and consistency. Furthermore, we established an update and maintenance mechanism for the standard library, including regular updates to standard data, fixing data errors, and adding new features to maintain the library's timeliness and usability. Finally, we provided access interfaces and documentation support for the standard library to facilitate user query and use of the data. The entire solution architecture aims to build an efficient, reliable, and easy-to-use standard library to provide strong data support for the business.
[0063] After building the standard library, we match it against the material master data. This matching can be done directly using key fields, such as material specifications and models, or by using fuzzy queries and regular expressions on attribute fields, or by using word vector similarity matching on the text in the material description field. After the matching is complete, we perform a search and comparison of the field content in each material master data to revise and supplement the attribute fields.
[0064] 3. Build a standard library
[0065] 3.1 Document collection and organization
[0066] First, collect all relevant national and industry standards documents. Organize them to ensure their completeness and accuracy. If the documents are in paper form, you may want to scan them into electronic form for easier processing.
[0067] If the document comes from an official Internet portal, you can use crawler technology to crawl the document data and obtain the latest data updated in real time.
[0068] 3.2 Document Parsing and Data Extraction
[0069] Parse electronic documents and identify their structure (e.g., chapters, titles, paragraphs, etc.). Extract required data elements, such as standard number, standard name, scope of application, technical requirements, etc. For complex documents, natural language processing (NLP) technology may be needed to assist in parsing and extracting data. This is shown in the following table:
[0070]
[0071] If it is a non-text document such as an image document or PDF, OCR text recognition technology is required to identify this part of the document content as electronic text data to facilitate subsequent document parsing and extraction.
[0072] Clean the extracted data to remove duplicate, erroneous, or irrelevant information. Standardize the data, such as by unifying data formats, units, and terminology. Categorize and organize the data according to the specific requirements of national and industry standards.
[0073] For example, material data can be filed in categories, subcategories, and leaf categories, and different fields can be used to distinguish and classify them, forming data with a logical relationship of inclusion and sub-inclusion.
[0074] Design the database structure based on the document structure and data requirements of national and industry standards. Determine which tables, fields, and relationships are needed, and define data types, constraints, and indexes. Design a reasonable database table structure and relationships to facilitate subsequent data storage and query. This is shown in the following table (material master data original table entry table):
[0075]
[0076]
[0077] 3.3 Data import and storage
[0078] Import cleaned and standardized data into the database. Store the data in appropriate tables and fields according to the database structure. Ensure data accuracy and completeness, and perform necessary verification and testing.
[0079] Design reasonable query statements according to needs to obtain the required information from the database. Optimize database performance, such as creating indexes, adjusting query statements, etc., to improve query efficiency and accuracy. Back up and maintain the database regularly to ensure data security and recoverability (such as Figure 4 shown).
[0080] 3.4 Development and Implementation
[0081] If you need to integrate the database with other systems or applications, you'll need to develop corresponding interfaces. Based on the interface requirements, design and implement the data exchange format and protocol. Test and verify the interface to ensure its stability and reliability.
[0082] API path: / api / material-master
[0083] HTTP method: POST
[0084] Request header:
[0085] Content-Type:application / json
[0086] Authorization:Bearer<token>
[0087] The data return format is as follows:
[0088]
[0089]
[0090]
[0091]
[0092] With the continuous updates and changes to national and industry standards, the database data needs to be regularly updated and maintained. New documents are parsed and data extracted, and then updated into the database. Old documents are reviewed and cleaned to ensure the accuracy and completeness of the data in the database.
[0093] 4Data matching
[0094] 4.1 Data Preparation
[0095] Ensure that the existing data has been organized and has fields that can be matched with the standard library (such as model specifications, standard numbers, etc.). If there are problems with the format or structure of the existing data, appropriate data cleaning and preprocessing are required.
[0096] 4.2 Define matching rules
[0097] Define appropriate matching rules based on business needs and data characteristics. This can be an exact match based on a single field or a fuzzy match based on multiple fields. For exact matching, use the = operator in SQL statements; for fuzzy matching, use the LIKE operator with wildcard characters (such as % and _).
[0098] Use a programming language (such as Python, Java, etc.) to write a program that can read existing data and match it with the standard library according to the defined matching rules. The program should be able to traverse each record of the existing data and find the matching entry in the standard library (such as Figure 5 shown).
[0099] When a matching entry is found in the standards library, the program should be able to extract relevant information about the entry (such as the standard name, technical requirements, etc.). The program should then use this information to correct deficiencies in the existing data. This may include updating field values, adding new fields, or performing other necessary operations.
[0100] 4.3 Development and Implementation
[0101] Before revising data, it is recommended to test a small amount of data to ensure that the matching and revision logic are correct. During the testing process, you can manually check whether the revised data meets expectations, or write automated test scripts to verify the accuracy of the data.
[0102] Once the matching and revision logic has been verified, batch processing can be performed on the entire existing dataset. Depending on the size and complexity of the dataset, you may want to consider using parallel processing, batch processing, or other optimization techniques to improve processing efficiency.
[0103] Before revising data, always back up the existing data. This ensures that you can restore it to its original state in the event of an error or unexpected situation. After the revision is completed, it is recommended to back up the data again for future reference or restoration.
[0104] As national and industry standards are updated and changed, the data in the standard library needs to be regularly checked and updated. At the same time, it is also necessary to monitor changes in existing data and perform corresponding revisions and maintenance work as needed.
[0105] 5. Natural Language Extraction
[0106] 5.1 Document Preprocessing
[0107] First, the standard document is preprocessed (such as Figure 7 As shown in the figure), including noise removal, text cleaning (such as removing HTML tags, special characters, extra spaces, etc.), text sentence segmentation, word segmentation, etc. If the document is in PDF, image or other non-text format, you may need to use OCR (Optical Character Recognition) technology (such as Figure 6 ), convert it to text format.
[0108] Named entity recognition in NLP technology is used to identify key entities in documents, such as names of people, places, organizations, standard numbers, etc. These entities are very important for subsequent data structuring and querying.
[0109] Relationship extraction technology is used to identify the relationships between entities in the document, such as the specific content and scope of application corresponding to a certain standard number. This helps organize the data into a structured form for subsequent query and analysis.
[0110] Semantic annotation techniques (such as part-of-speech tagging and semantic role tagging) are used to further understand the meaning and context of the text. This helps to extract and transform text data more accurately.
[0111] Based on the structure and content of standard documents, a set of templates and rules are developed to guide the NLP model in extracting key information from the text. These templates and rules can be adjusted and optimized according to the different types and characteristics of documents.
[0112] 5.2 Model Training and Evaluation
[0113] Use labeled datasets (if available) to train NLP models to improve their accuracy in identifying and extracting key information. Evaluate the performance of the models and make adjustments and optimizations based on the evaluation results.
[0114] 5.3 Development and Implementation
[0115] Convert the key information and relationships extracted by the NLP model into a structured data format (such as JSON, XML, database table, etc.). Store this data in a database or other storage system for subsequent query and analysis.
[0116] As standard documents are updated and changed, NLP models and data structures are regularly reviewed and updated. The entire processing flow is continuously optimized and improved based on user feedback and business needs.
[0117] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the scope of protection of the invention. Obviously, the embodiments described are only some embodiments of the present invention, rather than all embodiments. Based on these embodiments, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in this field can still combine, add, delete or make other adjustments to the features in the various embodiments of the present invention according to the circumstances without conflict, without making creative work, so as to obtain different other technical solutions that do not deviate from the concept of the present invention in essence, and these technical solutions also fall within the scope of protection of the present invention.< / token>
Claims
1. A material master data optimization and standardization processing method based on multi-dimensional matching rules, characterized in that: The following steps are involved: S1. Data preparation: Collect and organize material master data to ensure that the data has fields that match the standard library; S2. Standard library construction: Build a standard library containing standard data based on national standards and industry standards documents, and regularly update and maintain the standard library; S3. Define matching rules: Based on business needs and data characteristics, define multi-dimensional matching rules including exact matching and fuzzy matching; S4. Data matching: Use programming language to read material master data and match it with the standard library according to the defined matching rules; S5. Data revision: When a matching entry is found in the standard library, relevant information of the entry is extracted to revise deficiencies in the material master data; S6. Data Verification and Backup: Verify the revised data to ensure its accuracy and completeness, and perform data backup; S7. Cycle optimization: Regularly check and update the standard library, monitor changes in material master data, and make corresponding revisions and maintenance.
2. The material master data optimization and standardization processing method based on multi-dimensional matching rules according to claim 1 is characterized in that: The construction of the standard library in step S2 further includes the following sub-steps: S2.1 Document collection and organization: Collect all relevant national and industry standard documents and organize them to ensure their completeness and accuracy; S2.2 Document parsing and data extraction: Parse electronic documents, extract required data elements, and perform cleaning and standardization; S2.3 Data import and storage: Import the cleaned and standardized data into the database and establish a reasonable database table structure and relationships.
3. The material master data optimization and standardization processing method based on multi-dimensional matching rules according to claim 1 is characterized in that: The multi-dimensional matching rules in step S3 include exact matching based on a single field, fuzzy matching based on multiple fields, regular expression matching based on attribute fields, and text similarity matching based on material description fields.
4. The material master data optimization and standardization processing method based on multi-dimensional matching rules according to claim 1 is characterized in that: The data matching in step S4 also includes pre-processing the material description field using natural language processing technology to improve the matching accuracy.
5. The material master data optimization and standardization processing method based on multi-dimensional matching rules according to claim 1 is characterized in that: The data revision in step S5 includes updating field values, adding new fields, or performing other necessary operations to ensure that the material master data is consistent with the data in the standard library.