BOM material information automatic coding method based on AI
Through the AI-based automatic encoding method of material information, the problems of inefficient and difficult to ensure accuracy in traditional manual encoding are solved, and the intelligence and refinement of material management are realized, coding efficiency and data accuracy are improved, and the production and procurement processes are optimized.
Patent Information
- Application Number
- CN202510155676.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-02-12
AI Technical Summary
Traditional manual coded material information has problems such as inefficiency, difficulty in ensuring accuracy and lack of unified standards, resulting in inefficient bill of materials management and affecting the coordination of production processes and supply chains.
The automatic encoding method of material information based on AI is adopted, and the data structure of the bill of materials is unified, and the bit number and material usage is used to identify the bit number and material usage, filter the single bit number, perform data comparison and cleaning, and generate a unique model number to realize the automatic encoding of material information.
It improves the accuracy and standardization of material data, greatly improves coding efficiency, realizes the intelligence and refinement of material management, breaks information barriers, enhances coordination, optimizes production and procurement processes, and reduces operating costs.
Smart Images

Figure CN120011321A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of material informatization. More specifically, the present invention relates to an AI-based automatic coding method for BOM material information. Background Art
[0002] In many industries, including manufacturing and the electronics and information technology sector, the bill of materials (BOM), a core document describing product structure, details all the material information required. Efficient management of material information, particularly accurate and unified coding, is key to optimizing production processes, controlling costs, and facilitating supply chain collaboration. Traditional material information coding methods rely heavily on manual labor and present numerous challenges. Manual coding suffers from the following insurmountable drawbacks: first, low efficiency, which impacts the progress of procurement, production, and inventory processes; second, uncertainty about accuracy, which triggers chain reactions across subsequent procurement, production, and inventory management processes; and third, the lack of unified standards, which prevents the sharing of material coding information. Therefore, the development of a method that can automatically, accurately, and efficiently code material information in a BOM is urgent. With the rapid development of artificial intelligence (AI) technology and its increasing application across various industries, it presents new opportunities for solving material coding problems. With its powerful data analysis, pattern recognition, and predictive capabilities, AI can theoretically automate the coding of material information, potentially overcoming the limitations of traditional manual coding. However, in actual material management, numerous factors present significant challenges for AI-based automated material information coding.
[0003] First, in the actual material management process, the BOM formats used by different departments or business processes within an enterprise vary widely. In R&D, to conveniently record product design ideas, the BOM may focus on technical specifications and design parameters, using a more flexible format. In contrast, in production, to facilitate production process planning and material distribution, the BOM focuses more on the production process and assembly sequence, using a different format. In electronics manufacturing companies, R&D may list materials by functional module, while production may arrange them according to production line sequence. This results in a chaotic BOM data structure and makes it difficult to directly convert it into a unified standard BOM. Suppliers also provide a wide variety of BOM formats, ranging from Excel spreadsheets to PDF documents, and some even provide them in custom electronic document formats. This requires companies to spend a significant amount of time and effort on format conversion and data sorting when integrating BOM data.
[0004] Secondly, data errors in bills of materials are extremely common. Misspellings of material names, unclear specifications and models, and inaccurate packaging are common. For example, an automobile manufacturer might misspell the name of a key component in a bill of materials for procuring engine parts, leading to serious problems in subsequent supplier communications and quality traceability. These erroneous data renders the bill of materials unusable as a standard bill of materials, as standard bills require accurate data and can negatively impact the entire production process.
[0005] Finally, there is a serious problem of missing information in the bill of materials. Some bills of materials may lack key material group information, making it difficult for companies to clearly distinguish material categories, which is not conducive to material classification management and inventory counting. Some bills of materials may not record material location numbers, resulting in the inability to accurately determine the material installation location during the production and assembly process, affecting production efficiency and product quality. Material usage information may also be inaccurate or missing, which will cause great difficulties in procurement planning and inventory management, making it difficult to convert bills of materials into standard lists and unable to provide reliable data support for the company's production operations.
[0006] With the development of intelligent manufacturing, the automation and informatization of production processes are increasing, and the demand for standardized bills of materials is also increasing. Automated production lines require accurate standard bills of materials for automated distribution and assembly. The development of the Industrial Internet also requires efficient data sharing and collaboration between companies, suppliers, and partners based on standard bills of materials. However, the aforementioned issues with current bills of materials pose challenges to the automated coding of material information, seriously hindering enterprises' digital transformation and intelligent development. Summary of the Invention
[0007] An object of the present invention is to solve at least the above problems and to provide at least the advantages which will be described hereinafter.
[0008] Another purpose of the present invention is to provide an AI-based automatic coding method for BOM material information, which can improve the accuracy and standardization of material data, greatly improve coding efficiency, realize intelligent and refined material management, break down internal and external information barriers of the enterprise, enhance synergy, optimize production, procurement and other processes, and provide strong support for enterprises to reduce costs and increase efficiency.
[0009] In order to achieve these purposes and other advantages according to the present invention, a method for automatic coding of BOM material information based on AI is provided, comprising: S1. Unify the data structure of the bill of materials (BOM) prepared at any stage of a product and store it as a collection of multiple individual BOM data items. Each individual BOM data item includes at least the material name, material group, material specification model, tag number, material packaging form, and material quantity. S2. Use AI to identify the position number and material quantity in the bill of materials; S3. Filter out each piece of BOM data with unique tag data, and extract four pieces of data from each piece of BOM data: material name, material specification model, tag number, and material packaging form; compare the four pieces of data from the single BOM data with data in a pre-built standard database; if the four pieces of data from the single BOM data completely match an entry in the standard database, replace the single piece of material data with the corresponding entry in the standard database, but retain the tag number and corresponding material usage data in the single piece of material data, to form a single piece of material standard data; otherwise, clean the single piece of BOM data and compare it with the data in the standard database again; if the four pieces of data from the cleaned single BOM data completely match an entry in the standard database, replace the single piece of material data with the corresponding entry in the standard database, but retain the tag number and corresponding material usage data in the single piece of material data, to form a single piece of material standard data; S4. Use AI to generate a unique model code based on the material specification model, position number, and material packaging form in a single piece of material standard data, and encode the material in the form of material group plus unique model code; Among them, the method for cleaning a single bill of materials data is to correct spelling errors in material names, material specifications and models, and material packaging forms according to a pre-built standard vocabulary in the field of electronic devices.
[0010] Preferably, the single BOM data that still cannot be completely matched with the data in the standard database after cleaning in step S3 is further processed as follows: A1. Using the pre-built knowledge graph of electronic components, find the key material parameter data from the single bill of materials data. The key material parameter data includes component type, packaging form and performance parameters. A2. Unify the expression of similar key parameters and convert the material key parameter data into low-dimensional digital vectors through a vector conversion model; A3. After assigning weights to different key parameter data, all key parameter vector data are weighted and combined to generate a comprehensive vector for the single BOM data; A4. Calculate the cosine distance between the comprehensive vector of the single BOM data item and the comprehensive vector of each entry in the standard database using a cosine distance algorithm. If the obtained cosine distance exceeds a set matching threshold, replace the single BOM data item with the corresponding entry in the standard database, but retain the material usage data in the single BOM data item to form the single BOM data item.
[0011] Preferably, if the number of entries in A4 whose cosine distance exceeds the matching threshold is greater than 1, then: The plurality of entries whose cosine distance exceeds the matching threshold are sorted in ascending order according to the cosine distance, and the entry with the smallest cosine distance is selected as the corresponding entry in the standard database.
[0012] Preferably, if the number of entries in A4 whose cosine distance exceeds the matching threshold is greater than 1, then: Sort multiple entries whose cosine distance exceeds the matching threshold in ascending order of cosine distance, select the two entries with the smallest and second smallest cosine distances, and further calculate the difference in cosine distances corresponding to the two entries. If the difference is not greater than 0.05, select the entry with the smallest cosine distance as the corresponding entry in the standard database. Otherwise, further count the number of material key parameter data of the two entries, and select the entry with the larger number of material key parameter data as the corresponding entry in the standard database.
[0013] Preferably, A4 further includes the following steps: B1. Establish a historical matching database and calculate the frequency with which each entry in the standard database was selected during the historical matching process; B2. When the number of entries whose cosine distance exceeds the matching threshold is greater than 1, the entry with the highest selection frequency in the historical matching process is selected as the corresponding entry in the standard database.
[0014] Preferably, step S1 also includes identifying the format of the bill of materials prepared at any stage of the product and selecting corresponding processing tools to perform preliminary inspection of the data in the bill of materials, and returning the bill of materials that is missing any data such as material name, material group, material specification model, position number, material packaging form and material usage to the provider of the bill of materials to supplement the missing information.
[0015] Preferably, in step S1, unifying the data structure of the bill of materials prepared at any stage of the product includes the following steps: Identify the format of the bill of materials developed at any stage of the product, including: C1. Identify file format A by file extension and file format B by file header; C2. Utilize multiple rules in the file format judgment rule library to judge file format A and file format B respectively. Quantify the judgment results of various rules into scores, assign weights to each score, and form a comprehensive score for file format A and a comprehensive score for file format B based on a weighted combination of the scores and weights. Select the file format with the larger comprehensive score as the file format for the bill of materials. According to the identified file format, select the corresponding processing tool to extract various data in the bill of materials and generate a single bill of materials data with a unified data structure.
[0016] Preferably, the unified data structure in step S1 includes at least the following fields: The material name field is used to clearly identify the specific name of the material, and its data type is string; The material group field is used to group materials according to the classification criteria of function, source or production stage. Its data type is string; The material specification and model field is used to record the technical parameters and model specifications of the material. Its data type is string. The position number field is used to identify the specific position number of the material in the product structure, and its data type is a string; The material packaging type field is used to record the material packaging type, and its data type is a string; and The material usage information field is used to record the quantity of the material, and its data type is integer.
[0017] Preferably, the method for screening the unique bit number data in the bill of materials in step S3 includes: reading a piece of bill of materials data, comparing its bit number with the recorded bit number, and updating the frequency information; if the frequency information of the bit number is greater than 1, it is determined to be a non-unique bit number.
[0018] Preferably, noise data needs to be removed before cleaning a single BOM data. The method for removing noise data includes: D1. Use a single BOM data item as the core data source, determine its material name information, and extract the key parameter data of the material; D2. The material product design drawings, production process documents, and material specifications provided by the supplier are then used as data sources. After standardizing the data formats of different data sources, key material parameter data from one or more sources is extracted. D3, combining the data extracted in steps D1 and D2 to form a comprehensive feature vector to represent the material; D4. Use data mining algorithms to mine frequent item sets and association rules between different features in the feature vector; D5. When reading a single BOM data item to be denoised, first determine its material name information, use the material name information to match the single BOM data item with the same material name in the single BOM data item, then extract the frequent item sets and association rules between the corresponding different features, and perform denoising operations on the single BOM data item to be denoised based on these frequent item sets and association rules.
[0019] The present invention has at least the following beneficial effects: First, the AI-based automatic BOM material information encoding method provided by the present invention unifies the bill of materials data structure and clarifies the recording specifications for key information such as material names and material groups. This not only ensures data integrity and consistency from the source, but also simplifies the data processing for subsequent encoding work. Through a comparison and cleaning mechanism with a standard database, spelling errors are corrected based on a standard vocabulary in the field of electronic devices, ensuring that material data complies with industry standards and effectively avoiding information misunderstandings and incorrect transmissions caused by non-standard data. Secondly, the AI-based BOM material information automatic coding method provided by the present invention runs through the entire process of enterprise material management, from data collection, processing, standardization to coding of the bill of materials. In the production process, accurate material coding helps to accurately formulate production plans and distribute materials; in the procurement process, it facilitates efficient communication with suppliers; in the inventory management process, it is conducive to inventory counting and control, comprehensively improving the refinement level of enterprise material management, optimizing enterprise operation processes, reducing operating costs, and enhancing enterprise competitiveness. Third, the AI-based BOM material information automatic coding method provided by the present invention targets material data that cannot be directly matched with the standard database. It uses the electronic component field knowledge graph, vector conversion model and cosine distance algorithm technology to find the key parameter data of the material, convert it into a low-dimensional digital vector, and calculate the cosine distance to achieve intelligent matching of complex material data. When faced with new electronic components, even if their data format is special, the corresponding entry in the standard database can be efficiently found, significantly improving the applicability and flexibility of material coding. Fourthly, in the AI-based BOM material information automatic coding method provided by the present invention, when the cosine distance matching result is not unique, a variety of strategies are adopted for processing, such as selecting the minimum entry by cosine distance sorting, or further calculating the cosine distance difference between the minimum and second-minimum entries, combining the amount of key material parameter data, and even considering factors such as historical matching frequency. These strategies optimize matching decisions from multiple dimensions, adapt to various scenarios, ensure the selection of the most practical standard entry, and improve the accuracy and reliability of material coding; Fifth, the AI-based BOM material information automatic coding method provided by the present invention discloses a process for identifying and processing the format of the bill of materials. It identifies the file format through the file extension and file header, quantifies the judgment result using a format judgment rule library, and uses a weighted combination to form a comprehensive score to determine the format. Then, the corresponding tool is selected to extract data. This makes the material information automatic coding method adaptable to bills of materials in various formats, ensuring accurate data extraction from bills of materials from different sources, providing effective data support for subsequent standardization and coding work, and ensuring the consistency and efficiency of the material management process. Sixth, the AI-based BOM material information automatic encoding method provided by the present invention discloses the use of data mining algorithms to remove noise data. It is supported by multi-source data such as material product design drawings and production process documents. By extracting key parameter data, merging them to form a comprehensive feature vector, and mining frequent item sets and association rules, it can accurately identify and remove noise in a single bill of materials data; ensure the purity of the encoded data, and lay a solid foundation for the accuracy of subsequent encoding.
[0020] Other advantages, objectives and features of the present invention will be reflected in part from the following description and will be understood by those skilled in the art through study and practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 This is a flow chart of the AI-based automatic coding method for BOM material information described in one technical solution of the present invention. DETAILED DESCRIPTION
[0022] The present invention will be described in further detail below in conjunction with the accompanying drawings so that those skilled in the art can implement the invention with reference to the description.
[0023] It should be understood that terms such as “having”, “including” and “comprising” used herein do not preclude the existence or addition of one or more other elements or combinations thereof.
[0024] like Figure 1 As shown, the present invention provides an AI-based BOM material information automatic encoding method, comprising: S1. Unify the data structure of the bill of materials (BOM) prepared at any stage of a product and store it as a collection of multiple individual BOM data items. Each individual BOM data item includes at least the material name, material group, material specification model, tag number, material packaging form, and material quantity. S2. Use AI to identify the position number and material quantity in the bill of materials; S3. Filter out each piece of BOM data with unique tag data, and extract four pieces of data from each piece of BOM data: material name, material specification model, tag number, and material packaging form; compare the four pieces of data from the single BOM data with data in a pre-built standard database; if the four pieces of data from the single BOM data completely match an entry in the standard database, replace the single piece of material data with the corresponding entry in the standard database, but retain the tag number and corresponding material usage data in the single piece of material data, to form a single piece of material standard data; otherwise, clean the single piece of BOM data and compare it with the data in the standard database again; if the four pieces of data from the cleaned single BOM data completely match an entry in the standard database, replace the single piece of material data with the corresponding entry in the standard database, but retain the tag number and corresponding material usage data in the single piece of material data, to form a single piece of material standard data; S4. Use AI to generate a unique model code based on the material specification model, position number, and material packaging form in a single piece of material standard data, and encode the material in the form of material group plus unique model code; Among them, the method for cleaning a single bill of materials data is to correct spelling errors in material names, material specifications and models, and material packaging forms according to a pre-built standard vocabulary in the field of electronic devices.
[0025] In the above technical solution, a specific implementation method of step S1 is as follows: use a format recognition tool written in Python to perform format analysis on the bill of materials formulated at any stage of the product, such as research and development, production, quality inspection, inventory, etc., and determine the file format of the bill of materials by identifying file extensions, such as .xlsx, .csv, .pdf, etc., and specific identification information in the file header. For different identified file formats, call the corresponding data extraction plug-in. If it is an Excel file (.xlsx), use the pandas library to read the data, and extract information such as material name, material group, material specification model, position number, material powder packaging form, and material usage according to the pre-set column name mapping relationship; if it is a PDF file, use OCR (optical character recognition) technology in combination with regular expressions to match key information, and perform structured processing after extraction. The extracted material information is further stored as a collection of multiple single bill of materials data according to a unified data structure, where each single bill of materials data is stored in the following form: material name: specific name; material group: the group it belongs to; material specification model: detailed specifications; bit number: unique number; material packaging form: packaging type; material usage: specific quantity.
[0026] In the above technical solution, a specific implementation of step S2 is as follows: using the BERT model as an AI tool, further adding a relational reasoning module. In the above technical solution, a specific implementation of step S3 is as follows: writing a Python screening program, which uses a loop structure to traverse a single bill of materials data set. During the traversal process, the bit numbers that have appeared and their number of occurrences are recorded through a hash table. For each bill of materials data, its bit number is extracted and the hash table is queried. If the bit number does not appear in the hash table, it is added to the hash table and the number of occurrences is recorded as 1; if it has appeared, its number of occurrences is updated; and each single bill of materials data with a bit number data occurrence of 1 is screened out. For the single bill of materials data that has been screened, four data are extracted: material name, material specification model, bit number, and material packaging form. Using SQL query statements, these four data are compared with a standard database pre-built in a MySQL database. Each record in the standard database contains complete material information and has been reviewed by industry experts and verified by historical data. If the four data items in a single BOM record completely match an entry in the standard database, the material name, material group, material specification model, position number, material packaging form, and other data of the single material record are replaced with the corresponding entry in the standard database, while the material usage data in the original single material record is retained, forming a single piece of standard material data. If a complete match is not found, the cleaning process begins. Based on a pre-built standard vocabulary for the electronic device field, which is stored in JSON format and contains correct and incorrect spellings of common material names, material specifications, and material packaging forms, a string matching algorithm, such as the KMP (Knuth-Morris-Pratt) algorithm, is used to match the material name, material specification model, and material packaging form one by one. If a spelling error is found, it is corrected according to the correct expression in the vocabulary. After correction, the SQL query statement is used again to compare with the standard database. If the match is successful, the single piece of standard material data is formed in the above manner.
[0027] In the above technical solution, a specific implementation method of step S4 is as follows: based on the material category, such as capacitors, resistors, inductors, integrated circuits, etc., a unique material group coding rule is pre-set, for example, capacitors are represented by "C", resistors are represented by "R", inductors are represented by "L", and integrated circuits are represented by "IC". AI is used to generate a unique model code based on the material specification model, bit number, and material packaging form in a single piece of material standard data. The rules are as follows: For the material specification model part: first condense the key information of the material specification model. Taking capacitors as an example, first extract the key digital part of the capacitance value and the withstand voltage value, discard the unit information (because the material group has been indicated as capacitance, the unit can be defaulted), and perform digital processing. For example, if the capacitance value is 10μF, take 10, and if the withstand voltage value is 50V, take 50, and then splice the two into 1050. If the material specification model contains multiple parameters, the key parameters can be selected for processing according to importance or fixed order to ensure that this part of the code reflects the material characteristics to a certain extent and the number of bits is 4. For the bit number portion, which typically consists of letters and numbers, only the alphabetical portion of the bit number is encoded and converted to numbers (e.g., A=01, B=02, C=03, ..., Z=26). This ensures that this portion of the code identifies the bit number and has a digit length of two. For the material packaging type, a mapping table is established between the material packaging type and the digital code. For example, "SMD" is encoded as "01," "Plug-in" is encoded as "02," and "BGA package" is encoded as "03." If there are many different packaging types, layered or compressed encoding can be used to ensure a 2-digit code length. Finally, the material is encoded using the material group-unique model code format to ensure uniqueness and traceability.
[0028] Another specific embodiment of the above technical solution is as follows: During the product development phase at an electronics manufacturer, the design team developed detailed bills of materials (BOMs). These lists were stored in various formats, such as Excel spreadsheets and PDF documents. First, in step S1, the company's technical staff used data processing software to unify the BOM data structure. By identifying key information from the various file formats, they extracted information such as material name, group, specification, tag, packaging, and quantity, and stored it as a collection of multiple individual BOM data items.
[0029] Then proceed to step S2, using AI tools to identify the position numbers and material quantities in the bill of materials, fix the material quantities with the position numbers, and ensure that the material quantities are not replaced when replacing the entries with those in the standard database.
[0030] Then proceed to step S3, where technicians use a screening program to screen out each single BOM data with a unique bit number data. Taking capacitor materials as an example, when the screening program reads a BOM data of a capacitor with a bit number of "R005", the program compares it with the recorded bit number and finds that the bit number does not appear repeatedly and is a unique bit number. Then extract the four data of the material name, material specification model, bit number and material packaging form from the single BOM data and compare them with the pre-built standard database. During the comparison process, it is found that the material name of the capacitor "electrolytic capacitor" is different from the description of "electrolytic capacitor" in the standard database and cannot be fully matched. Therefore, according to the pre-built standard vocabulary in the field of electronic devices, the material name is corrected and "electrolytic capacitor" is changed to "electrolytic capacitor". Compare with the standard database again and successfully match the corresponding entry. Replace the single material data with the corresponding entry in the standard database, while retaining the bit number and original material usage data.
[0031] After completing the above steps, enter step S4. The technicians use the coding algorithm to encode the materials in a single piece of material standard data. Taking this capacitor as an example, according to the coding rules, the material specification model part extracts the capacitance value as 10μF and takes 10, and the withstand voltage value as 50V and takes 50, and then splices the two into 1050. For the bit number part, the letter part is converted into a value of 18, and for the material packaging part, it is converted into a number of 04. The material group is "C", and the material group and the unique model code are connected with "-", so the code is C-10501804.
[0032] In the above technical solution, the present invention has at least the following technical effects: 1. Through step S1, the data structure of the bill of materials is unified, making the originally chaotic material information orderly and standardized. This facilitates data sharing and collaboration among departments within the enterprise and reduces communication costs and errors caused by inconsistent data formats. In step S2, material usage is bound to reference numbers to ensure the accuracy of material usage in the BOM. In step S3, it is compared with the standard database and cleaned according to the standard vocabulary to further ensure the accuracy and standardization of material information, improve the quality of material information, and provide a reliable data foundation for subsequent production, procurement, and other links. 2. The screening and comparison mechanism in step S3 can quickly and accurately identify BOM data that needs to be standardized and promptly correct and replace it. This greatly shortens the time for collating and standardizing material information, improves the efficiency of material management, and greatly accelerates the progress of new product development and production. 3. In step S4, the materials are coded according to their various characteristics so that the codes are closely associated with the key information of the materials. This not only ensures the accuracy of the codes, but also enables the detailed information of the materials to be quickly traced through the codes in subsequent production and inventory management processes. Once quality problems or changes in production requirements occur, the relevant materials can be quickly located and corresponding measures can be taken, thereby improving the company's ability to control materials and the flexibility to deal with problems.
[0033] In another embodiment of the present invention, a single piece of BOM data that still cannot be completely matched with the data in the standard database after cleaning in step S3 is further processed as follows: A1. Using the pre-built knowledge graph of electronic components, find the key material parameter data from the single bill of materials data. The key material parameter data includes component type, packaging form and performance parameters. A2. Unify the expression of similar key parameters and convert the material key parameter data into low-dimensional digital vectors through a vector conversion model; A3. After assigning weights to different key parameter data, all key parameter vector data are weighted and combined to generate a comprehensive vector for the single BOM data; A4. Calculate the cosine distance between the comprehensive vector of the single BOM data item and the comprehensive vector of each entry in the standard database using a cosine distance algorithm. If the obtained cosine distance exceeds a set matching threshold, replace the single BOM data item with the corresponding entry in the standard database, but retain the material usage data in the single BOM data item to form the single BOM data item.
[0034] In the above technical solution, the present invention further optimizes the method for processing single BOM data that still cannot completely match the data in the standard database after cleaning. A specific implementation method thereof is as follows: Consider a single BOM data item (using capacitors as an example) that still cannot be fully matched with the standard database after cleaning. The BOM data item is as follows: Material name: capacitor; Material group: electronic components; Material specification model: 10μF 50V; Position number: R005; Material packaging form: SMD; Material quantity: 5.
[0035] Use pre-built knowledge graphs in the field of electronic components (such as Corning's knowledge graph) to find the key material parameter data of the capacitor. These key parameter data include but are not limited to: Component type: extracted from the material name, which is a capacitor; packaging form: SMD; performance parameters: extracted from the material specification model, which is 10uF 50V, indicating a capacitance value of 10uF and a rated voltage of 50V respectively.
[0036] To ensure that key parameter data from different sources or formats can be processed uniformly, the representation of these relationship parameters needs to be standardized. For example, the component type "capacitor" is uniformly represented as "capacitor" (this can be further refined to more specific types, such as electrolytic capacitors and ceramic capacitors, as needed. To simplify the process, this is assumed here without further refinement); the SMD package type is uniformly represented as "SMD"; and the performance parameter "10μF 50V" is uniformly represented as "10μF50V." The standardized material key parameter data is converted into low-dimensional digital vectors using a vector conversion model (e.g., Word2Vec). For example, the capacitor is converted into the vector [1,0,0]; the SMD is converted into the vector [0,1,0]; and the 10μF 50V is converted into the vector [0,0,1]. Weights are further assigned to the different key parameter data, for example: the component type has a weight of 0.4; the package type has a weight of 0.3; and the performance parameter has a weight of 0.3. The comprehensive vector of this single BOM data item is 0.4×[1,0,0]+0.3×[0,1,0]+0.3×[0,0,1]=[0.4,0.3,0.3]. Assume that the comprehensive vector of item 1 in the standard database after conversion is [0.5,0.3,0.2], and the comprehensive vector of item 2 after conversion is [0.3,0.4,0.3].
[0037] The cosine distance formula (for example, the cosine_similarity function in the scipy library) is further used to calculate the cosine distance between the single BOM data and item 1 and item 2. Assuming that the cosine distance between the single BOM data and item 1 is 0.8, and the cosine distance between the single BOM data and item 2 is 0.6, and the matching threshold is set to 0.7, the single material data is replaced with item 1, but the material usage data in the single material data is retained to form a single material standard data.
[0038] The above technical solution uses a pre-built knowledge graph in the field of electronic components to identify key material parameter data from a single bill of materials data, covering multiple dimensions such as component type, packaging form, and performance parameters. This overcomes the limitations of matching based solely on material name, specification model, part number, and packaging form. This approach can provide a deeper understanding of the essential characteristics of the material and avoid matching failures caused by inconsistent surface information. For example, for a capacitor, not only its simple name and specifications are considered, but also the specific type of capacitor (such as electrolytic capacitor, ceramic capacitor, etc.) and more detailed performance parameters (such as the capacitor's capacitance, withstand voltage value, error range, temperature characteristics, etc.) are extracted, making the description of the material more complete and improving the matching accuracy with entries in the standard database. The key material parameter data is converted into a low-dimensional digital vector and weighted combination is performed. The similarity with each entry in the standard database is then calculated using the cosine distance algorithm, achieving a quantitative assessment of the similarity between materials. Compared to traditional text matching methods, this method, based on vectors and cosine distance, can more accurately measure the similarity between materials. This is particularly true for materials with similar but not identical descriptions, enabling a more accurate identification of the closest matches. For example, if different manufacturers describe the same type of capacitor with subtle differences (such as "10uF±5% 50V" and "10uF 50V"), vector transformation and cosine distance calculation can better determine their similarity, rather than simply assuming they are different materials.
[0039] In another embodiment of the present invention, if the number of entries in A4 whose cosine distance exceeds the matching threshold is greater than 1, then: Sort multiple entries whose cosine distance exceeds the matching threshold by cosine distance from small to large, and select the entry with the smallest cosine distance as the corresponding entry in the standard database. Sort multiple entries whose cosine distance exceeds the matching threshold by cosine distance from small to large, and select the entry with the smallest cosine distance as the corresponding entry in the standard database. This strategy ensures that when multiple similar entries exist, the most similar one is selected first, reducing mismatches caused by inaccurate information and ensuring that the final single material standard data is as close to the actual situation as possible. In another embodiment of the present invention, if the number of entries in A4 whose cosine distance exceeds the matching threshold is greater than 1, then: Sort the multiple entries whose cosine distance exceeds the matching threshold in order of cosine distance from small to large, select the two entries with the smallest and second smallest cosine distances, and further calculate the difference in cosine distances corresponding to the two entries. If the difference is not greater than 0.05, select the entry with the smallest cosine distance as the corresponding entry in the standard database. Otherwise, further count the number of material key parameter data of the two entries, and select the entry with the larger number of material key parameter data as the corresponding entry in the standard database. Select the entry with the larger number of material key parameter data as the corresponding entry in the standard database. This strategy comprehensively considers the similarity between entries and the richness of key parameters, avoids misselection due to small cosine distance differences, and takes into account the integrity of key parameter information, thereby improving the rationality of selecting the most matching entry. In another embodiment of the present invention, A4 further includes the following steps: B1. Establish a historical matching database and calculate the frequency with which each entry in the standard database was selected during the historical matching process; B2. When the cosine distance of an entry exceeding the matching threshold is greater than 1, the entry most frequently selected in the historical matching process is selected as the corresponding entry in the standard database. This technical solution leverages historical data experience. For some common material matching scenarios, based on past matching experience, it can more accurately select appropriate entries. This is particularly applicable to materials with ambiguous descriptions or that appear frequently in different projects, improving the reliability and stability of matching.
[0040] In another embodiment of the present invention, step S1 also includes identifying the format of the bill of materials prepared at any stage of the product and selecting corresponding processing tools to perform preliminary inspection of the data in the bill of materials, and returning the bill of materials that is missing any data such as material name, material group, material specification model, position number, material packaging form and material usage to the provider of the bill of materials to supplement the missing information.
[0041] The above technical solution can systematically check the integrity of the bill of materials by identifying the format of the bill of materials prepared at any stage of the product and using the corresponding processing tools to perform preliminary inspections on the data in the bill of materials. For files containing multiple individual bill of materials data, this technical solution can ensure that each data contains key information, such as material name, material group, material specification model, position number, material packaging form, and material usage. Selecting the appropriate processing tool to parse the bill of materials can convert bills of materials from different sources and formats into a unified data structure. This helps to eliminate data processing difficulties caused by format differences, allowing subsequent operations (such as screening, matching, encoding, etc.) to be performed based on a consistent data format, improving the consistency and standardization of the entire material information processing process.
[0042] In another embodiment of the present invention, in step S1, unifying the data structure of the bill of materials prepared at any stage of the product includes the following steps: Identify the format of the bill of materials developed at any stage of the product, including: C1. Identify file format A by file extension and file format B by file header; C2. Utilize multiple rules in the file format judgment rule library to judge file format A and file format B respectively. Quantify the judgment results of various rules into scores, assign weights to each score, and form a comprehensive score for file format A and a comprehensive score for file format B based on a weighted combination of the scores and weights. Select the file format with the larger comprehensive score as the file format for the bill of materials. According to the identified file format, select the corresponding processing tool to extract various data in the bill of materials and generate a single bill of materials data with a unified data structure.
[0043] In the above technical solution, the present invention further optimizes the method of unifying the data structure of the bill of materials. When the format of the bill of materials is identified, a specific implementation method is as follows: Assume that the extension of the read bill of materials file is ".csv" (file format A), and the file header contains special identification information and is identified as file format B.
[0044] The file format judgment rule library contains multiple rules. The rules for file format A (.csv) are as follows: Check if the file is comma-delimited. This rule has a score of 30 and a weight of 0.3. Check whether the file contains a table header. This rule has a score of 40 and a weight of 0.5. Check whether the file has a specific suffix. The score of this rule is 20 points and the weight is 0.2.
[0045] The rules for file format B are as follows: Check whether the file header contains the "ProductBOM" logo. The score of this rule is 50 points and the weight is 0.6; Check whether the date format in the file header conforms to "YYYY-MM-DD". The score of this rule is 30 points and the weight is 0.4.
[0046] For file format A (.csv), after inspection, the file is comma-delimited, scoring 30 points; it contains a table header, scoring 40 points; but the suffix information does not meet the requirements, scoring 0 points. The overall score = 0.3 × 30 + 0.5 × 40 + 0.2 × 0 = 29 points. For file format B, after inspection, the file header contains the "ProductBOM" identifier, scoring 50 points; the date format meets the requirements, scoring 30 points. The overall score = 0.6 × 50 + 0.4 × 30 = 42 points.
[0047] In summary, since the comprehensive score of file format B is greater than the comprehensive score of file format A, we select file format B as the file format of the bill of materials.
[0048] The above technical solution uses file extensions to identify file format A and file headers to identify file format B. These two formats are then judged separately based on multiple rules in a file format judgment rule library. The results of each rule are quantified into scores, which are then combined with weights to form a comprehensive score, accurately determining the true format of the bill of materials. This multi-dimensional format determination approach avoids the one-sidedness and inaccuracy of relying solely on a single factor (such as the file extension) for file format judgment. Identifying and detecting formats at the initial stage of data processing can identify and resolve potential formatting issues at the source, preventing errors caused by formatting issues in subsequent data processing steps (such as screening, cleaning, and encoding).
[0049] In another embodiment of the present invention, the unified data structure in step S1 includes the following fields: The material name field is used to clearly identify the specific name of the material, and its data type is string; The material group field is used to group materials according to the classification criteria of function, source or production stage. Its data type is string; The material specification and model field is used to record the technical parameters and model specifications of the material. Its data type is string. The position number field is used to identify the specific position number of the material in the product structure, and its data type is a string; The material packaging type field is used to record the material packaging type, and its data type is a string; and The material usage information field is used to record the quantity of the material, and its data type is integer.
[0050] The above technical solution clearly categorizes and presents the information in the bill of materials by defining a unified data structure, including fields for material name, material group, material specification and model, tag number, material packaging form, and material usage information. By assigning a data type to each field, the solution ensures data consistency and standardization, avoids errors caused by inconsistent data types, and improves the data quality and reliability of the entire material information system.
[0051] In another embodiment of the present invention, the method for screening BOM data for unique tag number data in step S3 includes: reading a BOM data item, comparing its tag number with recorded tag numbers, and updating frequency information. If the tag number's frequency information is greater than 1, the item is determined to be non-unique. This accurately screens for unique tag numbers, reduces the risk of tag number confusion, and helps improve the efficiency and specificity of material information screening. Ensuring the tag number uniqueness of BOM data involved in subsequent processing can improve the quality of the entire material information processing. Because subsequent operations (such as comparing a single BOM data item with a standard database or encoding) are based on accurate tag numbers, mismatches or inaccurate encoding results caused by tag number issues can be reduced.
[0052] In another embodiment of the present invention, noise data needs to be removed before cleaning a single BOM data. The method for removing noise data includes: D1. Use a single BOM data item as the core data source, determine its material name information, and extract the key parameter data of the material; D2. The material product design drawings, production process documents, and material specifications provided by the supplier are then used as data sources. After standardizing the data formats of different data sources, key material parameter data from one or more sources is extracted. D3, combining the data extracted in steps D1 and D2 to form a comprehensive feature vector to represent the material; D4. Use data mining algorithms to mine frequent item sets and association rules between different features in the feature vector; D5. When reading a single BOM data item to be denoised, first determine its material name information, use the material name information to match the single BOM data item with the same material name in the single BOM data item, then extract the frequent item sets and association rules between the corresponding different features, and perform denoising operations on the single BOM data item to be denoised based on these frequent item sets and association rules.
[0053] The above solution further optimizes the method of removing noise from a single BOM data. A specific implementation method is as follows: The bill of materials data for a capacitor is as follows: Material name: capacitor; Material group: electronic components; Material specification model: 10μF 50V; Position number: R005; Material packaging form: SMD; Material quantity: 5.
[0054] The above single BOM data is used as the core data source, and its material name information is determined to be "capacitor". The key parameter data of the material is extracted, including capacitance value: 10μF; rated voltage 50V; tolerance ±10%.
[0055] Other data sources for this material include the product design drawings, production process documents, and the material specifications provided by the supplier. Key parameters extracted from the product design drawings (CAD files were parsed using a CAD file parsing tool) include length: 5mm; width: 3mm; and pin pitch: 2mm. Key parameters extracted from the production process documents (PDF files were parsed using a PDF parsing tool) include operating temperature range: -20°C to 85°C. Key parameters extracted from the material specifications provided by the supplier (text files were parsed using a text parsing tool) include lifespan: 1000 hours; maximum ripple current: 1A.
[0056] The extracted information is integrated into a comprehensive feature vector represented as [capacitor, 10μF, 50V, ±10%, SMD, R005, 5, 5mm, 3mm, 2mm, -20°C~85°C, 1000 hours, 1A]. A data mining algorithm (such as the Apriori algorithm) is then used to mine frequent itemsets and association rules between different features in the feature vector. First, use the create_transaction_list function to convert the comprehensive feature vector into a transaction list. Then, use pandas to convert the transaction list into a data frame for subsequent use in the Apriori algorithm. Set the apriori function to set the minimum support to 0.1, use column names as item set elements, and identify frequent itemsets.
[0057] When reading another single bill of materials data to be denoised, the following is shown: Material name: capacitor; Material group: electronic components; Material specification model: 10μF 50V; Position number: R005; Material packaging form: SMD; Material quantity: 5.
[0058] Based on the frequent itemsets and association rules previously mined, we can find that information commonly associated with "capacitor" includes tolerance, size, operating temperature range, lifespan, and maximum ripple current. Therefore, if this information is missing from the data to be denoised, we can supplement or modify it based on the frequent itemsets and association rules to make it more complete and accurate, and obtain the following result: Material name: Capacitor; Material group: Electronic components; Material specification model: 10μF 50V; Position number: R005; Material package type: SMD; Material quantity: 5; Length: 5mm (supplement); Width: 3mm (supplement); Pin spacing: 2mm (supplement); Operating temperature range: -20℃~85℃ (supplement); Lifespan: 1000 hours (supplement); Maximum ripple current: 1A (supplement).
[0059] The above technical solution uses a single BOM data item as the core data source and combines it with multiple data sources such as product design drawings, production process documents, and material descriptions provided by suppliers to extract key parameter data for materials, thereby achieving multi-dimensional collection of material information. Merging information from different data sources and synthesizing it into a comprehensive feature vector helps to unify the representation of scattered material information and condense information from multiple dimensions into a single vector, providing a comprehensive and structured data format for subsequent data processing. By mining frequent item sets and association rules, targeted denoising operations can be performed to improve the consistency and reliability of quantity processing. After denoising, matching errors or encoding errors caused by noisy data can be avoided, improving the reliability and success rate of subsequent processing steps. For example, for "capacitor" BOMs provided by multiple different suppliers, the same denoising logic can be used, so that the final processed data has the same data structure and information integrity, avoiding data inconsistencies caused by different information representations from different suppliers.
[0060] The number of devices and processing scale described here are used to simplify the description of the present invention. Applications, modifications and variations of the AI-based BOM material information automatic encoding method of the present invention are obvious to those skilled in the art.
[0061] Although the embodiments of the present invention have been disclosed above, they are not limited to the applications listed in the description and implementation methods. They can be fully applied to various fields suitable for the present invention. For those familiar with the art, additional modifications can be easily implemented. Therefore, without departing from the general concept defined by the claims and the scope of equivalents, the present invention is not limited to the specific details and illustrations shown and described herein.
Claims
1. The AI-based BOM material information automatic coding method is characterized by: include: S1. The data structure of the bill of materials prepared at any stage of the unified product is stored as a collection of multiple single bill of materials data. Each single bill of materials data at least includes the material name, material group, material specification model, position number, material packaging form and material usage; S2. Use AI to identify the position number and material usage in the bill of materials; S3. Filter out each single BOM data with unique position number data, and extract four data of material name, material specification model, position number and material packaging form from each single BOM data, and compare the four data of the single BOM data with the data in the pre-built standard database. If the four data of the single BOM data completely match an entry in the standard database, replace the single material data with the corresponding entry in the standard database but retain the position number and corresponding material usage data in the single material data to form a single material standard data; otherwise, clean the single BOM data and compare it with the data in the standard database again. If the four data of the cleaned single BOM data completely match an entry in the standard database, replace the single material data with the corresponding entry in the standard database but retain the position number and corresponding material usage data in the single material data; S4. Use AI to generate a unique model code based on the material specification model, position number and material packaging form in a single piece of material standard data, and encode the material in the form of material group plus unique model code; Among them, the method for cleaning a single BOM data is: correct the spelling errors of material names, material specifications and material packaging forms according to the pre-built standard vocabulary in the field of electronic devices.
2. The AI-based BOM material information automatic coding method according to claim 1, characterized in that: The single BOM data that cannot be completely matched with the data in the standard database after cleaning in step S3 is further processed as follows: A1. Using the pre-built knowledge graph in the field of electronic components, find the key material parameter data from the single BOM data. The key material parameter data includes component type, packaging form and performance parameters. A2. Unify the expression form of similar key parameters, and convert the key parameter data of the materials into low-dimensional digital vectors through a vector conversion model; A3. After assigning weights to different key parameter data, all key parameter vector data are weighted and combined to generate a comprehensive vector of the single BOM data; A4. Calculate the cosine distance between the comprehensive vector of the single material list data and the comprehensive vector of each entry in the standard database through the cosine distance algorithm. If the obtained cosine distance exceeds the set matching threshold, replace the single material data with the corresponding entry in the standard database, but retain the material usage data in the single material data to form a single material standard data.
3. The AI-based BOM material information automatic coding method according to claim 2, characterized in that: If the number of entries in A4 whose cosine distance exceeds the matching threshold is greater than 1, then: The multiple entries whose cosine distance exceeds the matching threshold are sorted in ascending order according to the cosine distance, and the entry with the smallest cosine distance is selected as the corresponding entry in the standard database.
4. The AI-based BOM material information automatic coding method according to claim 2, characterized in that: If the number of entries in A4 whose cosine distance exceeds the matching threshold is greater than 1, then: Sort multiple entries whose cosine distance exceeds the matching threshold from small to large according to the cosine distance, select the two entries with the smallest and second smallest cosine distances, and further calculate the difference in cosine distances corresponding to the two entries. If the difference is not greater than 0.05, select the entry with the smallest cosine distance as the corresponding entry in the standard database. Otherwise, further count the number of material key parameter data of the two entries, and select the entry with the larger number of material key parameter data as the corresponding entry in the standard database.
5. The AI-based BOM material information automatic coding method according to claim 2, characterized in that: A4 also includes the following steps: B1. Establish a historical matching database and calculate the frequency of each entry in the standard database being selected during the historical matching process; B2. When the number of entries whose cosine distance exceeds the matching threshold is greater than 1, the entry with the highest frequency selected in the historical matching process is selected as the corresponding entry in the standard database.
6. The AI-based BOM material information automatic coding method according to claim 1, characterized in that: Step S1 also includes identifying the format of the bill of materials prepared at any stage of the product and selecting the corresponding processing tools to perform preliminary inspection of the data in the bill of materials, and returning the bill of materials that is missing any data such as material name, material group, material specification model, position number, material packaging form and material usage to the provider of the bill of materials to supplement the missing information.
7. The AI-based BOM material information automatic coding method according to claim 6, characterized in that: In step S1, unifying the data structure of the bill of materials prepared at any stage of the product includes the following steps: Identify the format of the bill of materials developed at any stage of the product, including: C1. Identify file format A through file extension and file format B through file header; C2. Use multiple rules in the file format judgment rule library to judge file format A and file format B respectively, quantify the judgment results of various rules into scores, and assign weights to each score. According to the weighted combination of the score and the weight, the comprehensive score of file format A and the comprehensive score of file format B are formed, and the file format with the larger comprehensive score is selected as the file format of the bill of materials; According to the identified file format, select the corresponding processing tool to extract various data in the bill of materials and generate a single bill of materials data with a unified data structure.
8. The AI-based BOM material information automatic coding method according to claim 7, characterized in that: The unified data structure in step S1 includes at least the following fields: The material name field is used to clearly identify the specific name of the material, and its data type is string; The material group field is used to group materials according to the classification criteria of function, source or production stage. Its data type is string; The material specification model field is used to record the technical parameters and model specifications of the material, and its data type is a string; The position number field is used to identify the specific position number of the material in the product structure, and its data type is a string; The material packaging type field is used to record the material packaging type, and its data type is a string; and The material usage information field is used to record the quantity of the material, and its data type is integer.
9. The AI-based BOM material information automatic coding method according to claim 1, characterized in that: The method for screening the bit number data in the bill of materials for uniqueness in step S3 includes: reading a piece of bill of materials data, comparing its bit number with the recorded bit number, and updating the frequency information; if the frequency information of the bit number is greater than 1, it is determined to be a non-unique bit number.
10. The AI-based BOM material information automatic coding method according to claim 1, characterized in that: Before cleaning a single BOM data, noise data needs to be removed. The method of removing noise data includes: D1. Take a single BOM data as the core data source, determine its material name information, and extract the key parameter data of the material; D2. The material product design drawings, production process documents and material descriptions provided by suppliers are used as data sources. After unified processing based on the differences in data formats of different data sources, key material parameter data from one or more sources are extracted. D3, combining the data extracted in steps D1 and D2 to form a comprehensive feature vector to represent the material; D4. Use data mining algorithms to mine frequent item sets and association rules between different features in the feature vector; D5. When reading a single BOM data to be denoised, first determine its material name information, use the material name information to match the single BOM data with the same material name in the previous single BOM data, then extract the frequent item sets and association rules between the corresponding different features, and perform denoising operations on the single BOM data to be denoised based on these frequent item sets and association rules.
Citation Information
Patent Citations
Enterprise material cleaning service system and data cleaning method thereof
CN114328495A
Data cleaning method and device based on artificial intelligence, equipment and medium
CN118295991A
Fabricated building quality element list extraction method and system, terminal and medium
CN118761387A
Building industry material classification and attribute extraction method based on large language model
CN119378553A
Artificial intelligence system and method for processing multilevel bills of materials
US20150262124A1
Cited By
Electronic material intelligent coding method and system based on dynamic parameter mapping
CN121093201A
Material main data intelligent treatment method and system for electric power equipment manufacturing
CN121684813A
Server and Method for Automatic Matching of BOM and Process Based on Drawing Recognition Using AI Agent
KR103027383B1