Material BOM rapid comparison method and device, electronic equipment and storage medium

Through the rapid material BOM comparison method of data cleaning, hash indexing and graph theory clustering analysis, the problems of low processing efficiency and difficulty in identifying subtle differences in the existing technology are solved, and efficient and accurate material BOM comparison is achieved.

CN120045958AInactive Publication Date: 2025-05-27SHENZHEN QIANHENG ELECTRONICS CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510241997.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-05-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing fast material BOM comparison methods are inefficient in the face of complex integrated circuit design, are susceptible to human factors, and are difficult to identify subtle differences, especially when multi-dimensional and complex BOM data, the computing resource and time bottlenecks are obvious.

Method used

By obtaining the original BOM data for data cleaning, building a hash index to generate a hash feature vector group, calculating the difference matrix and performing graph theory clustering analysis, extracting key paths, generating a differential impact sorting table, and finally encapsulated into a material BOM comparison report.

Benefits of technology

It realizes rapid and accurate identification of differences between materials, reduces artificial deviations and omissions, improves the reliability and accuracy of comparison results, and supports efficient processing of multi-dimensional complex BOM data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045958A_ABST
    Figure CN120045958A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data analysis, and provides a material BOM rapid comparison method and device, electronic equipment and a storage medium. The method comprises the steps of obtaining original BOM data, performing data cleaning processing on the original BOM data to obtain structured BOM data, performing Hash index construction on the structured BOM data to obtain a Hash feature vector group, and further performing difference matrix calculation on the Hash feature vector group to obtain a difference identification matrix; performing graph theory clustering analysis on the difference identification matrix to obtain a difference classification topological graph, performing key path extraction on the difference classification topological graph to obtain a difference influence degree sorting table, and finally performing report structured packaging on the difference influence degree sorting table to obtain a material BOM comparison report. According to the method, through data cleaning, hash indexing, difference matrix calculation and graph theory clustering analysis, sorting is performed according to the influence degree while the difference is quickly recognized, and the accuracy and speed of BOM comparison are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of data analysis, and in particular, to a method, device, electronic device, and storage medium for quickly comparing material BOMs. Background Art

[0002] With the continuous improvement of the complexity of integrated circuit design, the version management and comparison of material BOMs have become increasingly important. Especially between different versions, different design stages, or suppliers, material BOMs often need to be compared to identify differences, ensure the accuracy of the design, and optimize the procurement and production processes.

[0003] The existing methods for quickly comparing material BOMs mainly include two categories: rule-based comparison and algorithm-based comparison. Rule-based comparison methods usually rely on manually set rules and standards, and identify differences by comparing item information in BOM data of each version, such as material names, models, specifications, etc. They are suitable for situations with simple structures and small data volumes. However, when faced with complex integrated circuit designs, the processing efficiency is low, and it is easily affected by human factors, resulting in inaccurate comparison results or omission of important differences. The existing algorithm-based comparison methods usually rely on pre-set rules and feature extraction methods, and are prone to ignoring relatively subtle differences in the BOM. For example, when there are ambiguities or inconsistencies in material names and models, it may lead to errors in the comparison results. Moreover, the existing methods have limited processing capabilities for complex BOM data. Especially when faced with multi-dimensional and hierarchically complex BOM data, it is easy to encounter bottlenecks in computing resources and time, resulting in a decline in processing efficiency. Summary of the Invention

[0004] In view of this, the present application provides a method, device, electronic device, and storage medium for quickly comparing material BOMs to solve the problem of quickly comparing multi-dimensional complex material BOM data.

[0005] The first aspect of the present application provides a method for quickly comparing material BOMs, the method comprising: Obtaining original BOM data and performing data cleaning processing on the original BOM data to obtain structured BOM data; Constructing a hash index for the structured BOM data to obtain a hash feature vector group; Calculating a difference matrix for the hash feature vector group to obtain a difference identification matrix; Performing graph theory clustering analysis on the difference identification matrix to obtain a difference classification topology graph; Extracting a critical path from the difference classification topology graph to obtain a difference impact degree ranking table; Performing report structured encapsulation on the difference impact degree ranking table to obtain a material BOM comparison report.

[0006] In an alternative embodiment, the construction of the hash index for the structured BOM data to obtain the hash feature vector group includes: Concatenate each material parameter combination in the structured BOM data to obtain a string sequence for each material; Perform a hash calculation on the string sequence to obtain a hash value for each material; Perform a storage structure conversion on the hash values to obtain the hash feature vector group.

[0007] In an alternative embodiment, the calculation of the difference matrix for the hash feature vector group to obtain the difference identification matrix includes: Perform a bitwise exclusive OR calculation on every two hash values in the hash feature vector group to obtain the difference value between each pair of hash values; Perform a calculation of the number of different bits on the difference value according to a preset PopCount operation to obtain the difference measure between each pair of materials; Perform a difference matrix calculation on the difference measure to obtain the difference matrix between every two materials; Perform a non-matching item screening according to a preset matching threshold and the difference matrix to obtain the difference identification matrix.

[0008] In an alternative embodiment, the graph theory clustering analysis of the difference identification matrix to obtain the difference classification topology graph includes: Construct a graph according to the difference identification matrix to obtain a weighted graph; Perform a modularity optimization process on the weighted graph to obtain the clustering category to which each material belongs; Perform a structural analysis on the weighted graph according to the clustering category to obtain the difference classification topology graph.

[0009] In an alternative embodiment, the modularity optimization process of the weighted graph to obtain the clustering category to which each material belongs includes: Step S41: Analyze the adjacency relationship of each node in the weighted graph to obtain the connection strength between the nodes; Step S42: Obtain the similarity measure between materials according to the edge weights and the connection strength in the weighted graph; Step S43: Perform a local optimization process on the weighted graph according to the similarity measure to obtain a weighted graph to be measured; Step S44: Calculate the modularity of the weighted graph to be measured to obtain a modularity value to be measured; Step S45: Obtain the modularity value of the weighted graph, and calculate the difference between the modularity value and the to-be-tested modularity value to obtain a modularity difference value; Step S46: Update the clustering structure of the weighted graph according to the to-be-tested weighted graph; Repeat steps S41 to S46 until the modularity difference value is equal to or less than a preset difference value threshold, and then obtain the clustering category to which each material belongs according to the updated weighted graph.

[0010] In an optional embodiment, the obtaining the differential classification topology graph by performing structural analysis on the weighted graph according to the clustering category includes: Perform structural analysis on the weighted graph according to the clustering category to obtain the coordinates and edge weights of the clustering category; Map the clustering category according to the coordinates to obtain a topological structure; Adjust the topological structure according to the edge weights to obtain a primary differential classification topology graph; Perform visualization processing on the primary differential classification topology graph according to a preset visualization method to obtain the differential classification topology graph.

[0011] In an optional embodiment, the obtaining the differential impact degree ranking table by performing critical path extraction on the differential classification topology graph includes: Perform traversal analysis on the differential classification topology graph to obtain differential paths; Calculate the impact degree of each differential path to obtain the impact degree value of each differential path; Sort the differential paths according to the impact degree value to obtain the impact degree ranking of the differential paths; Filter the differential paths according to a preset sequence threshold and the impact degree ranking to obtain a set of critical paths; Generate the differential impact degree ranking table according to the set of critical paths.

[0012] The second aspect of the present application provides a device for quickly comparing material BOMs, and the device includes: A data cleaning module, configured to obtain original BOM data and perform data cleaning processing on the original BOM data to obtain structured BOM data; A hash index module, configured to construct a hash index for the structured BOM data to obtain a group of hash feature vectors; A difference identification module, configured to calculate a difference matrix for the group of hash feature vectors to obtain a difference identification matrix; A difference classification module, configured to perform graph theory clustering analysis on the difference identification matrix to obtain a difference classification topology graph; An influence ranking module, configured to extract a critical path from the difference classification topology graph to obtain a difference influence degree ranking table; A comparison report module, configured to perform report structured encapsulation on the difference influence degree ranking table to obtain a material BOM comparison report.

[0013] A third aspect of the present application provides an electronic device, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the material BOM fast comparison method described above are implemented.

[0014] A fourth aspect of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the material BOM fast comparison method described above are implemented.

[0015] In summary, the present application at least includes the following beneficial technical effects: 1. Calculate the difference matrix through the hash feature vector group, and apply the PopCount operation to quantify the difference measure between materials, so as to quickly identify the differences between materials and avoid the cumbersome manual comparison.

[0016] 2. Use graph theory clustering analysis to classify the difference data, and perform clustering processing on the weighted graph through the modularity optimization algorithm, so that the differences between different materials can be clearly organized and identified.

[0017] 3. Through the critical path extraction method, rank the difference influence degree, help quickly find the most critical difference part, provide a basis for subsequent decision-making, and reduce human biases and omissions.

[0018] 4. By filtering non-matching items and setting a matching threshold, errors and irrelevant data can be effectively removed, the accuracy of the comparison can be improved, interference from irrelevant information can be avoided, and the reliability of the comparison result can be enhanced. Description of the Drawings

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0020] Figure 1 is a flowchart of a material BOM fast comparison method provided by an embodiment of the present application; Figure 2 It is a functional module diagram of a material BOM rapid comparison device provided by an embodiment of the present application; Figure 3 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Specific embodiments

[0021] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0022] As Figure 1 shown, it is a flowchart of a material BOM rapid comparison method provided by an embodiment of the present application. The material BOM rapid comparison method provided by an embodiment of the present application includes the following steps.

[0023] Step S1: Obtain the original BOM data and perform data cleaning processing on the original BOM data to obtain structured BOM data.

[0024] It should be understood that the Bill of Materials (BOM) is a basic document in processes such as electronic engineering, manufacturing, and assembly, listing all the parts and raw materials required for product production. Through the API interface or the import function of the system, according to the type of the material management system, select to obtain the original BOM data from an ERP system (such as SAP, Oracle, etc.), an electronic design automation tool (EDA tool, such as Altium, Cadence, etc.) or an Excel file. Usually, the ERP system will contain comprehensive material information, such as material number, description, supplier, inventory, price, etc.; the EDA tool contains material data related to circuit design, such as the model number, reference designator, electrical characteristics, etc. of components. Excel is usually used for data storage of exports or manual inputs.

[0025] Since the data formats from different sources may be inconsistent (for example, the units are not unified, the formats of material numbers are not unified, the date formats are inconsistent, etc.), in order to ensure the accuracy and reliability of subsequent processing, it is necessary to unify the data formats. For the unification of units, the units can be parsed according to regular expressions and converted into a unified standard. For example, "10μF" of capacitance value should be converted to the standard format "10uF". The date formats in different systems may be different, and all dates need to be unified into a specific format (for example, YYYY-MM-DD).

[0026] At the same time, different data sources may use different field names, such as: material number, model number, quantity, part number, etc. To make the data consistent, it is necessary to unify the naming of these fields. Usually, a mapping table is used to unify the field names in different data sources to ensure that each field has a standard name throughout the BOM data.

[0027] There may be missing values in the original BOM data, such as missing information about the material description, quantity, supplier, etc. To ensure the integrity of the data and the smooth progress of subsequent processing, the missing values need to be processed. Depending on the type and context of the missing values, the following methods can be used for filling: 1. Filling based on the mean or median: For numerical data (such as quantity), it can be filled according to the mean or median of similar materials. If the quantity of a certain material is missing, the quantity of other materials in the same category can be referred to for filling.

[0028] 2. Filling based on rules: For missing descriptive fields (such as material description, model number), it can be filled by rule inference. For example, if a certain material lacks a description, it can be inferred based on its model number or other known information.

[0029] 3. Marking missing: If the missing quantity exceeds a certain threshold (such as 10%), it is marked as abnormal data and not filled.

[0030] After completing the basic data formatting, the original BOM data needs to be cleaned to detect and correct abnormal data. Exemplarily, some data such as material quantity, price, etc. may have logical errors (for example, the quantity is negative). These abnormalities can be automatically detected by a rule engine and marked or corrected. If there are duplicate material items in the original BOM data (for example, the same material number and model number appear multiple times), they should be de-duplicated. It can be judged whether it is a duplicate record according to fields such as material number, model number, etc. If so, their quantities are merged.

[0031] At the same time, during data cleaning, it is also necessary to perform consistency verification on the data. For example, check whether the material number exists in the system and whether it matches other fields (such as model number, supplier), or check whether the material quantity is reasonable. For inconsistent data, it can be corrected through rule verification or manual review.

[0032] After data cleaning, the original BOM data will become more standardized, complete, and consistent. Next, the data needs to be organized into structured BOM data for subsequent processing. Structured data generally refers to converting data into the form of a standard table or a database table, so that each data item of the material has a clear field. Design a suitable table structure according to the requirements of subsequent use. Each BOM record should include but not be limited to: material number, model, parameters, quantity, item number, description, supplier, etc. By filling the cleaned data into the standard table, structured data is generated. Depending on the size of the data volume, the BOM data can be stored as an Excel table or imported into a database. For small projects, Excel can store it conveniently and quickly; while for large-scale material data, it can be chosen to store it in a relational database (e.g., MySQL, PostgreSQL) or a non-relational database (e.g., MongoDB) to support more efficient querying and analysis.

[0033] Finally, the generated structured BOM data can be exported in a standard format, such as CSV, Excel, or stored in the form of a table in the database. The structured data format can facilitate subsequent processing, such as calculation, comparison, screening, report generation, etc.

[0034] Step S2: Construct a hash index for the structured BOM data to obtain a group of hash feature vectors.

[0035] Among them, hash index construction refers to generating a hash value according to the characteristics of each material in the structured BOM data, so as to quickly identify the material according to the hash value. In the structured BOM data, each material usually contains multiple characteristics, such as material number, model, description, quantity, unit, item number, etc. To construct a hash index, first, the multiple characteristics of each material need to be concatenated into a string in a predetermined order. For example, for a capacitor material, the concatenation method of "material number + model + parameters + quantity + unit" can be adopted. By concatenating all the key characteristics, it can be ensured that each material has a unique string representation, ensuring that the data of different materials are distinguished in the hash calculation and reducing the probability of conflicts.

[0036] After completing the concatenation of the material parameters, perform a hash calculation on the string sequence of each material to obtain its hash value. Hash calculation is to map the concatenated string to a hash value of a fixed length through a hash function. Among them, the hash value is the unique identifier of the material, which is used for efficient querying and comparison in subsequent processing. The hash algorithm has high efficiency and a low collision probability. Through the hash algorithm, a string of any length can be mapped to a hash value of a fixed length (e.g., 128 bits, 256 bits, etc.). The embodiment of the present application adopts the FNV-1a hash algorithm shown as follows: Among them, is the ASCII code sequence of the input string, that is, the concatenated material feature string is converted. is the prime constant used by the FNV algorithm, usually 16777619. is the length of the string. is the ASCII code value of the character in the string.

[0037] After obtaining the hash value of the material, store the hash value of each material and its corresponding index in a structured hash feature vector group. The hash feature vector group not only contains hash values, but may also contain other metadata related to the material (for example, material number, model, quantity, etc.). The hash feature vector group is used to support fast query and difference analysis. The hash feature vector group usually adopts the form of a list or a dictionary (which can be stored as a table in a database). Each hash value is used as an index, and other information of the material (such as model, quantity, etc.) is stored as additional metadata in the hash table. When using a hash table for storage, the key is the hash value of the material, and the value is other metadata of the material (such as material number, description, etc.). This way can complete the query operation in constant time; when using an array for storage, the index is the hash value, and the detailed information of the material is stored. Although the storage is relatively simple, more processing may be required during query. The hash feature vector group provides the function of fast retrieval and comparison. By using the hash value of each material as an identifier, it supports fast lookup through the hash value, and can greatly reduce memory consumption and calculation time, especially when processing large-scale material BOM data. Exemplarily, to determine whether two materials are the same, only need to compare their hash values. If the hash values are the same, it can be considered that they are the same, thus avoiding the traditional one-by-one comparison.

[0038] Step S3, perform a difference matrix calculation on the hash feature vector group to obtain a difference identification matrix.

[0039] Extract the hash values of each pair of materials from the hash feature vector group. Each material has a unique hash value in the hash feature vector group, representing the feature information of the material. Perform the following bitwise exclusive OR operation on every two hash values: Among them, is the difference value between the hash values of material A and material B. and They are the hash values of Material A and Material B respectively. is the bitwise XOR operator. The bitwise XOR operator is used to compare whether the bits at the same position of two binary numbers are the same. If the values of the two numbers at a certain bit are different, the result of that bit is 1; if they are the same, the result is 0.

[0040] After the difference value (i.e., binary number) obtained by bitwise XOR calculation, the number of 1s in the difference value is calculated through the PopCount operation. The PopCount operation calculates the number of 1s in a binary number, reflecting how many bits are different between the materials in terms of hash values. The larger the value of the difference metric, the greater the difference between the two materials.

[0041] After calculating the difference metrics for each pair of materials through the PopCount operation, a difference matrix needs to be constructed to provide a quantitative data basis for matching screening and difference classification. The difference matrix is a two-dimensional matrix. In the difference matrix, each row represents the difference metric between one material and other materials. Specifically, the elements of the matrix represent the difference metric between Material A and Material B. By calculating the difference metrics for each pair of materials, the entire matrix is obtained. The size of the matrix is N×N, where N is the total number of materials. The specific formula is as follows: Among them, is the difference metric between Material A and Material B, calculated by PopCount. is the total number of materials. is a difference matrix, and the elements in the difference matrix are calculated through the difference metric.

[0042] After obtaining the difference matrix, non-matching items in the difference matrix are screened according to a preset matching threshold, thereby generating a difference identification matrix. The difference identification matrix is used to identify which materials are matching and which are non-matching. Among them, the matching threshold is a value that needs to be set according to specific application requirements, representing the maximum allowable value of the difference metric. If the difference metric between two materials is less than this threshold, they are considered to be matching; otherwise, they are considered to be non-matching. According to the value of the difference matrix, if < matching threshold, it is considered that Material A and Material B are matching, and the corresponding element in the difference identification matrix is set to 0; if ≥ matching threshold, it is set to 1, indicating that these two materials do not match. The calculation formula of the difference identification matrix is specifically as follows: Among them, , The difference metric obtained through PopCount. Through the matching threshold screening process, the generated difference identification matrix can clearly identify which materials are matched and which are not, providing an effective data basis for subsequent clustering analysis and difference identification.

[0043] Step S4: Perform graph theory clustering analysis on the difference identification matrix to obtain a difference classification topology graph.

[0044] Among them, graph theory clustering analysis is an important tool for identifying the relationships between materials. By converting the difference identification matrix into a weighted graph and performing modularity optimization on this graph, materials can be divided into different clustering categories. By optimizing the structure of the graph, the difference classification between materials is revealed, and a clear difference classification topology graph is formed.

[0045] In the difference identification matrix, if = 1, it indicates that there is a significant difference between material A and material B, otherwise it is 0. In the weighted graph, will be converted into the edge weight of the graph. Specifically, the 1 or 0 in the difference identification matrix may be further mapped to a continuous value through a certain formula to represent the magnitude of the difference. Obtain the edges of all material pairs from the difference matrix and assign edge weights to form an undirected graph (i.e., a weighted graph). By converting the difference identification matrix into a weighted graph, graph theory tools can be effectively used for further analysis and optimization of materials, such as clustering analysis, difference detection, etc.

[0046] In the weighted graph, each material or material cluster is represented by a node, and the relationship between nodes is represented by the edge weight. Calculate the connection strength between material nodes through adjacency relationship analysis to analyze the similarity between materials. That is, there is an edge between material node i and material node j, and the weight of the edge reflects the similarity or difference degree between these two nodes (materials). Adjacency relationship analysis is to calculate their connection strength through the edge weights of material nodes. There is an edge between material node i and material node j, and the weight of the edge reflects the similarity or difference degree between these two nodes (materials). Adjacency relationship analysis is to calculate their connection strength through the edge weights of material nodes. The specific connection strength formula is as follows: Among them, is the connection strength between material node and material node , which is equal to the edge weight between them. The edge weight is calculated according to the similarity or difference degree of materials, usually obtained based on the difference identification matrix.

[0047] Based on the connection strength of each edge and the edge weight, the similarity measure between each pair of materials can be calculated. The similarity measure reflects the similarity between materials. The similarity measure is a direct reflection of the edge weight. The larger the edge weight, the larger the similarity measure. The cosine similarity or Pearson correlation coefficient can be used to measure the similarity between materials. The specific similarity measure formula is as follows: where, is the similarity measure between material node and material node . is the edge weight between material node and node , representing the similarity between them. and are the degrees of material node and node , representing the number of their respective connected edges.

[0048] Based on the similarity measure, local optimization is performed on the weighted graph. The local optimization process clusters similar material nodes together according to the similarity between materials, thus forming a more reasonable clustering structure. The optimization process usually uses a greedy algorithm to gradually adjust the connection relationship between material nodes until optimization. Local optimization improves the quality of the clustering result by considering the similarity of materials and merging similar material nodes, ensuring that the final clustering categories can accurately reflect the differences between materials.

[0049] After one round of local optimization, the structure of the weighted graph to be measured changes, and it is necessary to calculate the modularity of the weighted graph to be measured to evaluate whether the optimized graph structure is appropriate. Modularity is a measure of the quality of graph partitioning and reflects whether the distribution of nodes in the clustering is reasonable. The specific formula is as follows: where, is the edge weight between material node and node . and are the degrees of material node and node , representing the number of their respective connected edges. is the total number of edges in the graph. is an indicator function. If node and node belong to the same cluster, the value is 1, otherwise it is 0. The higher the modularity value, the closer the clustering relationship between nodes and the more reasonable the partitioning of the weighted graph to be measured.

[0050] After obtaining the modularity value of the weighted graph to be measured, the modularity value of the weighted graph before optimization (hereinafter collectively referred to as the original weighted graph) is obtained by accessing the database, and the difference in the modularity value between the original weighted graph and the weighted graph to be measured (i.e., the modularity difference value) is used to determine whether the optimization process has achieved the expected effect. If the modularity difference value is small, it indicates that there is no significant change in the clustering structure of the optimized weighted graph; if the modularity difference value is large, it indicates that the optimization process has introduced significant changes.

[0051] It should be understood that the modularity difference value gradually becomes smaller after multiple rounds of local optimization processing (i.e., the modularity difference value gradually converges), thus indicating that the structure of the weighted graph is gradually approaching perfection. After obtaining the modularity difference value, the modularity difference value is compared with a preset difference value threshold to determine whether the clustering structure of the weighted graph has reached the expected appropriate degree. At the same time, the clustering structure of the weighted graph is updated according to the weighted graph to be measured. If the modularity difference value is less than or equal to the difference value threshold, it can be determined that the current optimized clustering structure is the final structure, and the clustering category of the weighted graph is updated. If the modularity difference value is greater than the difference value threshold, local optimization processing is continued until the preset convergence condition is met.

[0052] The clustering structure is optimized by performing modularity optimization on the weighted graph. Each material is assigned to an appropriate clustering category in the weighted graph, and finally an optimized clustering category is generated.

[0053] Based on the clustering category, a structural analysis of the weighted graph is performed to construct a clear differential classification topology graph, which is helpful for the differential analysis of the material BOM. By constructing the differential classification topology graph, the differences, similarities, and relationships between clustering categories among materials can be more intuitively displayed.

[0054] According to the clustering category of each material in the weighted graph, the coordinates of each clustering category on the weighted graph are determined through structural analysis, and the edge weights between clustering categories are calculated. The relative positions and connection relationships of each clustering category in the differential classification topology graph can be determined according to the similarity and difference degrees of the materials.

[0055] The coordinates of the clustering category are usually calculated by a force-directed layout algorithm or other graph layout algorithms. By optimizing the distances between nodes and the layout of edges, the nodes of clustering categories with high similarity are close to each other, while the nodes of clustering categories with large differences are far from each other. In the embodiments of the present application, by optimizing the x and y coordinates of each clustering category node, the coordinates of the nodes of clustering categories with high similarity are close, so as to form a tight clustering structure in the differential classification topology graph. The specific implementation can use the force-directed layout algorithm to adjust the node positions iteratively until the overall structure of the differential classification topology graph is stable. The specific formula is as follows: Among them, is the total energy of the graph, representing the layout quality of the graph. is the clustering category and the clustering category the connection strength between them. is the node and node distance. is the ideal distance, usually related to their similarity. The more similar the nodes are, the closer the distance should be. By minimizing the total energy , optimize the positions of the clustering category nodes in the weighted graph (i.e., the coordinates of the clustering categories), ensuring that the nodes of similar clustering categories are closer and the nodes of clustering categories with larger differences are farther apart.

[0056] Place the coordinate points of each clustering category at the corresponding positions in the differential classification topology graph. Through the coordinates of the clustering categories, the clustering category nodes occupy clear positions in the differential classification topology graph, and the edges between the nodes are connected according to the similarity and difference degrees. During the mapping process, the magnitude of the edge weight determines the thickness or color of the edge. When the edge weight is larger, it indicates that the difference between the materials is smaller, and the connected edge can be thinner; on the contrary, when the edge weight is smaller, it means that the difference between the materials is larger, and the connected edge can be thicker.

[0057] Through mapping, ensure that the position of each clustering category in the differential classification topology graph is reasonable, so that all clustering categories form a reasonable topological structure in the differential classification topology graph, and the differences between the materials are shown through the weights of the edges.

[0058] After obtaining the preliminary topological structure generation, further adjust the topological structure according to the edge weights to optimize the connections between the clustering category nodes, making the relationship between the clustering categories more in line with the difference degree of the materials. The edge weight reflects the similarity or difference degree between the materials. Therefore, when constructing the differential classification topology graph, it must be adjusted according to these weights. Specifically, for the clustering category nodes with higher similarity (i.e., with smaller edge weights), the connected edges should be adjusted to make their distances closer; while for the clustering category nodes with larger differences (i.e., larger edge weights), the distance between them should be increased. The nodes and edges in the topological graph need to be optimized and adjusted according to the difference degree between the materials, so that the relationship of each clustering category can more accurately reflect the actual differences of the materials.

[0059] Finally, according to the characteristics of the primary differential classification topology graph, select an appropriate graph layout algorithm (e.g., force-directed layout, hierarchical layout, etc.) so that the similarity and difference degrees between clustering categories can be clearly presented in the differential classification topology graph. The layout algorithm will adjust according to the edge weights and positions between nodes to optimize the final visualization effect. In the differential classification topology graph, the thickness or color of the edges will be adjusted according to the edge weights. The larger the edge weight, the thinner the edge, indicating a material clustering with smaller differences; the smaller the edge weight, the thicker the edge, indicating a material clustering with larger differences. Add identifiers to each clustering category node, such as the material number, name, or label of the clustering category, to help analysts understand the specific content of each clustering category. By clearly displaying the differential classification topology graph, the differences between materials and the relationships between clustering categories are intuitively presented, thus providing strong support for subsequent analysis, decision-making, report generation, etc.

[0060] Step S5: Extract the critical path from the differential classification topology graph to obtain a differential impact degree ranking table.

[0061] It should be understood that a differential path is an edge or a set of edges between two different clustering categories in the differential classification topology graph. The differential path shows the connection and differences between material clusterings. In the differential path, the weight of the edge reflects the similarity or difference between materials. By using a traversal algorithm (e.g., depth-first search algorithm or breadth-first search algorithm) to traverse all nodes (i.e., the clustering categories on both sides of the path) in the differential classification topology graph and check the connections between each pair of clustering categories, all paths that meet the differential conditions (i.e., differential paths) are recorded. The differential path can be a directly connected edge or a path connected through intermediate nodes. For each path, collect its nodes and the weights of the edges.

[0062] After obtaining the differential paths, it is necessary to calculate the impact degree value of each differential path. The impact degree value is an important quantitative evaluation of the differential path in the final comparison result. The impact degree value of each differential path is directly related to the difference degree and connection strength of the material nodes on the path. Specifically, the calculation of the impact degree value can be obtained by weighted summation of the weights of all edges on the differential path, or by multiplying the weights of each edge. The calculation of the impact degree value reflects the importance of the path, that is, the degree of influence of this path on the clustering result. The impact degree calculation formula of the embodiment of the present application is as follows: Where, is the impact degree value of the differential path. is the weight of the th edge on the path, that is, the similarity or difference degree between materials. is the number of edges on the difference path. By calculating the product of the edge weights on the difference path, the overall influence degree value of the difference path can be obtained. Each edge on the path contributes to the influence degree, and the edges with larger weights have a greater impact on the influence degree.

[0063] After obtaining the influence degree values of each difference path, all the difference paths are sorted from high to low according to the influence degree values by common quicksort algorithm or mergesort algorithm. The paths with higher influence degree values indicate that the paths have a greater impact on the clustering result. By sorting the difference paths, it can be determined which differences need to be focused on the most.

[0064] Furthermore, according to a preset sequence threshold, the difference paths in the influence degree sorting are screened to obtain a key path set. Among them, the key paths are the paths with higher rankings in the influence degree sorting, which are used to indicate that the difference paths have an important impact on the material BOM comparison result. Specifically, according to the preset sequence threshold, the difference paths with sorting numbers greater than or equal to the sequence threshold are selected from the influence degree sorting, and the sequence threshold can be adjusted according to the actual situation. For example, the top 20% of the difference paths in the influence degree sorting are selected. Finally, the selected key paths are summarized to form a key path set.

[0065] Finally, according to each path in the key path set and its corresponding influence degree value, they are sorted from high to low according to the influence degree value and listed in the difference influence degree sorting table. Each row in the difference influence degree sorting table should include information such as the path number, the material nodes of the path, and the influence degree value of the path.

[0066] Exemplarily, the following difference influence degree sorting table is obtained according to the key path Through the difference influence degree sorting table, the influence degrees of each material difference can be visually viewed.

[0067] Step S6: Structurally encapsulate the difference influence degree sorting table to obtain a material BOM comparison report.

[0068] To integrate the analysis results into a structured, visual, and easily interpretable material BOM comparison report, it is necessary to encapsulate the above-obtained difference impact degree ranking table into the report structure. Specifically, convert the difference impact degree ranking table into XML format. Among them, XML is a format commonly used for data representation, which is convenient for organizing, storing, and sharing data. By converting the difference impact degree ranking table into XML format, it can ensure that the data has good scalability and easy parsing. Among them, the XML format file includes, but is not limited to, the root node used as the container for all information and the sub-nodes created according to each difference path in the difference impact degree ranking table. Among them, each sub-node nests the detailed information about the path, such as the code, parameters, difference metrics, clustering categories, etc. of the material.

[0069] After obtaining the XML format encapsulation, it is necessary to further convert the XML data into an HTML report to display the comparison results in a graphical way. The HTML report can intuitively display the comparison results and provide a more understandable interface for users. The XML format difference impact degree ranking table is converted into HTML format through XSLT. Specifically, an XSLT style sheet is pre-written according to XSLT to define how to convert the XML format into HTML format. The style sheet specifies how to display the information of each difference path, including but not limited to the path number, material information, impact degree value, etc. At the same time, the mapping rules of XSLT are preset according to XSLT, which are used to stipulate that the sub-nodes in XML are mapped to a row in the HTML table, and each field corresponds to a column in the table. Finally, according to the mapping rules of XSLT and the XSLT style sheet, the data recorded on XML is converted into HTML format. In order to make the report more readable, the generated HTML report is beautified through CSS styles, so as to add colors, fonts, table formats, etc. in the report to highlight key information, such as paths with larger difference degrees, paths with higher impact degrees, etc.

[0070] After converting the difference impact degree ranking table into an HTML format report, the HTML file is output in the form of a PDF file (i.e., the material BOM comparison report) for reference by decision-makers, engineers, or other relevant personnel. The material BOM comparison report contains the most important material differences, helping users understand which material differences have greater impacts, thus providing a basis for subsequent optimization, adjustment, and correction.

[0071] This application is applied to the field of data analysis technology. By obtaining the original BOM data and performing data cleaning on the original BOM data to obtain structured BOM data, and constructing a hash index for the structured BOM data to obtain a hash feature vector group, further calculating a difference matrix for the hash feature vector group to obtain a difference identification matrix, thereby performing graph theory clustering analysis on the difference identification matrix to obtain a difference classification topology graph, and then extracting the critical path from the difference classification topology graph to obtain a difference impact degree ranking table, and finally performing report structured encapsulation on the difference impact degree ranking table to obtain a material BOM comparison report. This application quickly identifies differences and sorts them according to the impact degree through data cleaning, hash indexing, difference matrix calculation, and graph theory clustering analysis, improving the accuracy and speed of BOM comparison.

[0072] As Figure 2 shown, it is a functional module diagram of a material BOM rapid comparison device provided by an embodiment of this application.

[0073] In some embodiments, the material BOM rapid comparison device 2 may include multiple functional modules composed of computer program segments. The computer programs of each program segment in the material BOM rapid comparison device 2 can be stored in the memory of the server and executed by at least one processor to execute (see details in Figure 1 description) the functions of the material BOM rapid comparison method.

[0074] In this embodiment, according to the functions it executes, the material BOM rapid comparison device 2 can be divided into multiple functional modules. The functional modules may include: a data cleaning module 21, a hash indexing module 22, a difference identification module 23, a difference classification module 24, an impact ranking module 25, and a comparison report module 26. The module referred to in the present invention means a series of computer program segments that can be executed by at least one processor and can complete fixed functions, and are stored in the memory. In this embodiment, the functions of each module will be described in detail in subsequent embodiments.

[0075] The data cleaning module 21 is used to obtain the original BOM data and perform data cleaning on the original BOM data to obtain structured BOM data.

[0076] The hash indexing module 22 is used to construct a hash index for the structured BOM data to obtain a hash feature vector group.

[0077] In an optional implementation manner, the hash indexing module 22 is specifically used for: Concatenating each material parameter combination in the structured BOM data to obtain a string sequence for each material; Perform a hash calculation on the string sequence to obtain the hash value of each material; Perform a storage structure conversion on the hash value to obtain a hash feature vector group.

[0078] The difference identification module 23 is used to perform a difference matrix calculation on the hash feature vector group to obtain a difference identification matrix.

[0079] In an optional implementation manner, the difference identification module 23 is specifically used for: Perform a bitwise exclusive OR calculation on every two hash values in the hash feature vector group to obtain the difference value between each pair of hash values; Perform a calculation on the number of different bits of the difference value according to a preset PopCount operation to obtain the difference measure between each pair of materials; Perform a difference matrix calculation on the difference measure to obtain a difference matrix between every two materials; Perform a non-matching item screening according to a preset matching threshold and the difference matrix to obtain the difference identification matrix.

[0080] The difference classification module 24 is used to perform a graph theory clustering analysis on the difference identification matrix to obtain a difference classification topology graph.

[0081] In an optional implementation manner, the difference classification module 24 is specifically used for: Perform a graph construction according to the difference identification matrix to obtain a weighted graph; Perform a modularity optimization process on the weighted graph to obtain the clustering category to which each material belongs; Perform a structure analysis on the weighted graph according to the clustering category to obtain the difference classification topology graph.

[0082] In an optional implementation manner, the difference classification module 24 is further used for: Step S41: Perform an adjacency relationship analysis on each node in the weighted graph to obtain the connection strength between the nodes; Step S42: Obtain the similarity measure between the materials according to the edge weight and the connection strength in the weighted graph; Step S43: Perform a local optimization process on the weighted graph according to the similarity measure to obtain a weighted graph to be measured; Step S44: Perform a modularity calculation on the weighted graph to be measured to obtain a measured modularity value; Step S45: Obtain the modularity value of the weighted graph, and calculate the difference between the modularity value and the measured modularity value to obtain a modularity difference value; Step S46: Update the clustering structure of the weighted graph according to the weighted graph to be measured; Repeat the execution of the steps S41 to S46 until the modularity difference value is equal to or less than a preset difference value threshold, and obtain the clustering category to which each material belongs according to the updated weighted graph.

[0083] In an optional embodiment, the difference classification module 24 is further configured to: Perform a structural analysis on the weighted graph according to the clustering category to obtain the coordinates and edge weights of the clustering category; Map the clustering category according to the coordinates to obtain a topological structure; Adjust the topological structure according to the edge weights to obtain a primary difference classification topological graph; Perform a visualization process on the primary difference classification topological graph according to a preset visualization method to obtain the difference classification topological graph.

[0084] The impact sorting module 25 is configured to extract a critical path from the difference classification topological graph to obtain a difference impact degree sorting table.

[0085] In an optional embodiment, the impact sorting module 25 is specifically configured to: Perform a traversal analysis on the difference classification topological graph to obtain a difference path; Calculate the impact degree of the difference path to obtain the impact degree value of each difference path; Sort the difference paths according to the impact degree value to obtain the impact degree sorting of the difference paths; Filter the difference paths according to a preset sequence threshold and the impact degree sorting to obtain a critical path set; Generate the difference impact degree sorting table according to the critical path set.

[0086] The comparison report module 26 is configured to perform a report structure encapsulation on the difference impact degree sorting table to obtain a material BOM comparison report.

[0087] It should be understood that the various change modes and specific embodiments in the method provided in the above embodiments are equally applicable to the material BOM rapid comparison device in this embodiment. Through the foregoing detailed description of the material BOM rapid comparison method, those skilled in the art can clearly know the implementation method of the material BOM rapid comparison device in this embodiment. For the sake of brevity of the specification, it will not be described in detail here.

[0088] As Figure 3 shown, it is a schematic structural diagram of an electronic device provided by an embodiment of the present application.

[0089] In a preferred embodiment of the present invention, the electronic device 3 may include, but is not limited to: a memory 31, at least one processor 32, and at least one communication bus 33.

[0090] Those skilled in the art should understand that Figure 3 the structure of the illustrated electronic device 3 does not constitute a limitation on the embodiments of the present invention. The electronic device 3 may further include more or fewer other hardware or software than shown, or different component arrangements.

[0091] In some embodiments, the electronic device 3 is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions. Its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits, programmable gate arrays, digital signal processors, and embedded devices, etc.

[0092] It should be noted that the electronic device 3 is only an example. Other existing or future electronic products that can be adapted to this application should also be included within the protection scope of this application and are hereby incorporated by reference.

[0093] In some embodiments, a computer program is stored in the memory 31. When the computer program is executed by the at least one processor 32, all or part of the steps in the material BOM rapid comparison method as described are implemented. The memory 31 includes read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM), or other optical disc memories, magnetic disk memories, tape memories, or any other computer-readable medium capable of carrying or storing data. Further, the computer-readable storage medium mainly includes a storage program area and a storage data area. Among them, the storage program area may store an operating system, application programs required for at least one function, etc.

[0094] In some embodiments, the at least one processor 32 is the control core (Control Unit) of the electronic device 3, connecting various components of the entire electronic device 3 through various interfaces and circuits, and by running or executing programs or modules stored in the memory 31, and calling data stored in the memory 31, to perform various functions of the electronic device 3 and process data. For example, when the at least one processor 32 executes the computer program stored in the memory 31, all or part of the steps of the method for quickly comparing material BOMs described in the embodiments of the present application are implemented; or all or part of the functions of the device for quickly comparing material BOMs are implemented. The at least one processor 32 may be composed of integrated circuits. For example, it may be composed of a single packaged integrated circuit, or may be composed of multiple integrated circuits with the same or different functions packaged, including a combination of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips, etc.

[0095] In some embodiments, the at least one communication bus 33 is configured to implement connection communication between the memory 31 and the at least one processor 32, etc. Although not shown, the electronic device 3 may further include a power source (such as a battery) for supplying power to each component. Preferably, the power source can be logically connected to the at least one processor 32 through a power management device, so as to implement functions such as management of charging, discharging, and power consumption management through the power management device. The power source may also include any components such as one or more DC or AC power sources, a recharge device, a power failure detection circuit, a power converter or inverter, a power status indicator, etc. The electronic device 3 may also include various sensors, a Bluetooth module, a Wi-Fi module, etc., which will not be elaborated here.

[0096] The integrated units implemented in the form of software function modules as described above can be stored in a computer-readable storage medium. The above-mentioned software function modules are stored in a storage medium, including several instructions for causing an electronic device (which may be a personal computer, an electronic device, or a network device, etc.) or a processor to execute part of the methods described in the various embodiments of the present application.

[0097] In several embodiments provided by the present application, it should be understood that the disclosed device and method can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation.

[0098] The module described as a separation component may or may not be physically separated. The component shown as a module may or may not be a physical unit, and it may be located in one place or distributed across multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0099] The above are all preferred embodiments of this application. The protection scope of this application is not limited thereby. Therefore, all equivalent changes made according to the structure, shape, and principle of this application shall be covered within the protection scope of this application.

Claims

1. A material BOM quick comparison method, characterized in that: The method comprises: Acquire original BOM data and perform data cleaning on the original BOM data to obtain structured BOM data; Performing hash index construction on the structured BOM data to obtain a hash feature vector group; Performing difference matrix calculation on the hash feature vector group to obtain a difference identification matrix; Performing graph theory cluster analysis on the difference identification matrix to obtain a difference classification topology map; Extracting key paths from the difference classification topology graph to obtain a difference impact ranking table; The difference impact ranking table is packaged in a report structure to obtain a material BOM comparison report.

2. The material BOM rapid comparison method according to claim 1 is characterized in that: The step of constructing a hash index on the structured BOM data to obtain a hash feature vector group includes: Concatenate each material parameter combination in the structured BOM data to obtain a character string sequence for each material; Performing hash calculation on the string sequence to obtain a hash value for each material; The hash value is subjected to storage structure conversion to obtain a hash feature vector group.

3. The material BOM rapid comparison method according to claim 1 is characterized in that: The performing difference matrix calculation on the hash feature vector group to obtain a difference identification matrix comprises: Performing bitwise XOR calculation on every two hash values ​​in the hash feature vector group to obtain a difference value between each pair of hash values; Counting the number of different digits of the difference value according to a preset PopCount operation to obtain a difference measure between each pair of materials; Performing difference matrix calculation on the difference measure to obtain a difference matrix between every two materials; Non-matching items are screened according to a preset matching threshold and the difference matrix to obtain the difference identification matrix.

4. The material BOM rapid comparison method according to claim 1 is characterized in that: The performing graph theory cluster analysis on the difference identification matrix to obtain a difference classification topology map comprises: Performing graph construction according to the difference identification matrix to obtain a weighted graph; Performing modularity optimization processing on the weighted graph to obtain the cluster category to which each material belongs; The weighted graph is structurally analyzed according to the clustering categories to obtain the difference classification topology graph.

5. The material BOM rapid comparison method according to claim 4 is characterized in that: The performing modularity optimization processing on the weighted graph to obtain the cluster category to which each material belongs includes: Step S41, performing adjacency analysis on each node in the weighted graph to obtain the connection strength between the nodes; Step S42: obtaining a similarity measure between materials according to the edge weights in the weighted graph and the connection strength; Step S43: performing local optimization processing on the weighted graph according to the similarity metric to obtain a weighted graph to be tested; Step S44, performing modularity calculation on the weighted graph to be tested to obtain a modularity value to be tested; Step S45, obtaining the modularity value of the weighted graph, and calculating the difference between the modularity value and the modularity value to be measured to obtain a modularity difference value; Step S46, updating the clustering structure of the weighted graph according to the weighted graph to be tested; The step S41 to the step S46 are repeatedly executed until the modularity difference value is equal to / less than a preset difference value threshold, and the cluster category to which each material belongs is obtained according to the updated weighted graph.

6. The material BOM rapid comparison method according to claim 5 is characterized in that: The performing structural analysis on the weighted graph according to the clustering category to obtain the difference classification topology graph comprises: Performing structural analysis on the weighted graph according to the clustering categories to obtain coordinates and edge weights of the clustering categories; Mapping the cluster categories according to the coordinates to obtain a topological structure; Adjust the topological structure according to the edge weights to obtain a primary difference classification topological graph; The primary difference classification topology map is visualized according to a preset visualization method to obtain the difference classification topology map.

7. The material BOM rapid comparison method according to claim 1 is characterized in that: The extracting of the key path from the difference classification topology graph to obtain the difference impact ranking table comprises: Performing traversal analysis on the difference classification topology graph to obtain difference paths; Calculating the influence of the difference paths to obtain the influence value of each difference path; Sorting the difference paths according to the influence values ​​to obtain an influence ranking of the difference paths; The difference paths are screened according to a preset sequence threshold and the impact ranking to obtain a critical path set; The difference impact ranking table is generated according to the critical path set.

8. A material BOM quick comparison device, characterized in that: The device comprises: A data cleaning module is used to obtain original BOM data and perform data cleaning on the original BOM data to obtain structured BOM data; A hash index module, used for constructing a hash index on the structured BOM data to obtain a hash feature vector group; A difference identification module, used for performing difference matrix calculation on the hash feature vector group to obtain a difference identification matrix; A difference classification module is used to perform graph theory cluster analysis on the difference identification matrix to obtain a difference classification topology map; An impact ranking module is used to extract key paths from the difference classification topology map to obtain a difference impact ranking table; The comparison report module is used to perform report structured packaging on the difference impact ranking table to obtain a material BOM comparison report.

9. An electronic device, characterized in that: The electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the material BOM rapid comparison method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the material BOM rapid comparison method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Data comparison method and device, equipment and storage medium

    CN120849409A