Coal machine spare part similarity identification and coding unification system and method and storage medium
Patent Information
- Application Number
- CN202511612994.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-06
- Publication Date
- 2026-09-15
- Estimated Expiration
- 2045-11-06
AI Technical Summary
人工整理依赖经验,效率低且易出错,无法应对大规模数据
[0015] The present invention relates to a unified system for similarity identification and coding of coal mining machinery spare parts. This invention utilizes large models and deep learning to extract and fuse multi-dimensional features from multi-source spare parts data. Based on the comprehensive feature analysis formed by the fusion of multi-dimensional features, the similarity between spare parts data is analyzed, breaking through the dependence on the original coding format of spare parts and improving the accuracy and robustness of cross-system spare parts data matching.
Smart Images

Figure CN121434738B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of coal mining equipment supply chain information management and data governance technology; specifically, this invention relates to a unified system, method and storage medium for coal mining equipment spare parts similarity identification and coding. Background Technology
[0002] With the acceleration of digitalization and intelligentization in the coal mining industry, enterprises have successively launched business systems such as ERP, WMS, and MRO, realizing information management of spare parts procurement, inventory, and allocation. However, due to differences in system construction time, manufacturer standards, and data entry habits, the names, codes, models, and specifications of the same spare parts are often inconsistent in different systems, resulting in problems such as "one item, multiple names" and "one item, multiple codes," leading to redundant inventory data, difficulties in inventory counting, and low efficiency in allocation and procurement.
[0003] Existing technologies mainly employ manual sorting or simple rule-based matching methods. Manual sorting relies on experience, is inefficient and error-prone, and cannot handle large-scale data. While rule-based matching can handle format differences, it lacks flexibility and semantic understanding capabilities. Summary of the Invention
[0004] In view of this, the present invention provides a unified system, method and storage medium for similarity identification and coding of coal mining machinery spare parts, thereby solving or at least alleviating one or more of the above-mentioned problems and other problems existing in the prior art.
[0005] To achieve the aforementioned objectives, a first aspect of the present invention provides a unified system for similarity identification and coding of coal mining machinery spare parts, wherein the system comprises: The data acquisition module is used to extract raw data of coal mining machine spare parts from multiple business systems. The raw data includes the name, code, description of the coal mining machine spare parts, and information about the equipment where the spare parts are located. The cleaning and preprocessing module is used to perform uniform formatting on the raw data and output preprocessed data. A feature extraction and matching engine is used to extract multi-dimensional features from the preprocessed data. The multi-dimensional features include text similarity features, semantic features, structural features, and attribute features. The semantic features are extracted through a large model. The similarity calculation and clustering module is used to calculate the comprehensive similarity between different coal mining machinery spare parts data through a deep fusion network based on the multi-dimensional features, and to automatically cluster them according to a set threshold, classifying spare parts with a comprehensive similarity higher than the threshold into one category. The master data generation and system integration module is used to generate unified spare parts master data for each cluster and synchronize the spare parts master data to the multiple business systems.
[0006] Optionally, in the system described above, the system further includes a word segmentation and industry dictionary module, used to perform word segmentation on the preprocessed data and, in conjunction with a coal mining machinery industry-specific dictionary, to perform specialized identification on the preprocessed data. This specialized identification includes recognizing industry terminology and identifying the model, specifications, and units of measurement of spare parts or equipment. The word segmentation and industry dictionary module is also used to calculate the text similarity features between spare parts data based on the results of the word segmentation processing and the specialization recognition. The text similarity features include the edit distance or Jaccard similarity between strings.
[0007] Optionally, in the system described above, the system further includes a coding rule normalization module, used to learn the spare parts coding rules of different business systems using a large model, converting spare parts data containing spare parts codes into a unified semantic vector, and using the semantic vector as the semantic feature of the spare parts. The large model is a finely tuned large language model based on a corpus of coal mining machinery, which includes coal mining machinery manuals, industry standards, equipment manuals, and after-sales records.
[0008] Optionally, in the system described above, the system further includes an attribute information parsing module for extracting the attribute features from the preprocessed data, the attribute features including specifications, material, compatible models, and functional uses.
[0009] In the system described above, optionally, the feature extraction and matching engine integrates the text similarity features, the semantic features, the structural features, and the attribute features to form a comprehensive feature, wherein the structural features include the bill of materials, location, and function of the equipment where the spare parts are located.
[0010] In the system described above, optionally, the similarity calculation and clustering module calculates the comprehensive similarity through a deep fusion network in the following way: in, This is a comprehensive feature vector of different spare parts data. This represents vector concatenation. Here are the network parameters, and ReLU is the linear rectified function. The Sigmoid activation function outputs... This represents the overall similarity score; The deep fusion network is trained using a binary cross-entropy loss function: in, For the average binary cross-entropy loss, The total number of training samples, For the actual matching labels of spare parts pairs, Indicates a pair of truly similar spare parts. Indicates dissimilar pairs.
[0011] Optionally, in the system described above, the system further includes a human feedback and self-learning module, which is used to manually review the clustering results when the similarity calculation and clustering module performs clustering based on similarity, and to use the results of the human review for the self-optimization of the system.
[0012] To achieve the aforementioned objective, a second aspect of the present invention provides a method for similarity identification and unified coding of coal mining machinery spare parts based on a system as described in any of the first aspects above, wherein the method includes: The raw data of coal mining machinery spare parts is extracted from multiple business systems. The raw data includes the name, code, description of the coal mining machinery spare parts and information about the equipment where the spare parts are located. The original data is formatted uniformly to obtain preprocessed data; Multi-dimensional feature extraction is performed on the preprocessed data, including text similarity features, semantic features, structural features, and attribute features; Based on the multi-dimensional features, the comprehensive similarity between different coal mining machinery spare parts is calculated through a deep fusion network, and automatic clustering is performed according to a set threshold, classifying spare parts with a comprehensive similarity higher than the threshold into one category; Generate unified spare parts master data for each cluster, and synchronize the spare parts master data to the multiple business systems.
[0013] In the method described above, optionally, the multi-dimensional feature extraction includes: The preprocessed data is segmented into words and then combined with a coal mining machinery industry-specific dictionary to perform specialized identification. This specialized identification includes identifying industry terms and identifying the model, specifications, and units of measurement of spare parts or equipment. Based on the results of the word segmentation and the specialized recognition, the text similarity features between spare parts data are calculated, including the edit distance or Jaccard similarity between strings. By using a large model to learn the spare parts coding rules of different business systems, spare parts data containing spare parts codes are converted into a unified semantic vector. Extract the attribute characteristics of the spare parts, including specifications, material, compatible models and functions; Extract the structural features of the spare parts, including the bill of materials, location, and function of the equipment in which the spare parts are located; The text similarity features, semantic vectors, structural features, and attribute features are integrated to form a comprehensive feature.
[0014] To achieve the foregoing objectives, a third aspect of the present invention provides a computer-readable storage medium having computer-executable instructions or a computer program that, when executed by a processor, implement the method as described in any of the preceding second aspects.
[0015] The present invention relates to a unified system for similarity identification and coding of coal mining machinery spare parts. This invention utilizes large models and deep learning to extract and fuse multi-dimensional features from multi-source spare parts data. Based on the comprehensive feature analysis formed by the fusion of multi-dimensional features, the similarity between spare parts data is analyzed, breaking through the dependence on the original coding format of spare parts and improving the accuracy and robustness of cross-system spare parts data matching.
[0016] The present invention further provides a method for similarity identification and unified coding of coal mining machinery spare parts based on the above system, and therefore the method also has the above advantages.
[0017] The present invention further provides a computer-readable storage medium for implementing the above-described method, and thus the computer-readable storage medium also has the above-described advantages. Attached Figure Description
[0018] The disclosure of this invention will become more apparent from the accompanying drawings. It should be understood that these drawings are for illustrative purposes only and are not intended to limit the scope of protection of this invention. In the drawings: Figure 1 This is a schematic block diagram of an embodiment of the coal mining machinery spare parts similarity identification and coding unified system of the present invention; Figure 2 This is a schematic block diagram of another embodiment of the coal mining machinery spare parts similarity identification and coding unified system of the present invention. Detailed Implementation
[0019] Referring to the accompanying drawings and specific embodiments, the structure, composition, features, and advantages of the coal mining machinery spare parts similarity identification and coding unified system, method, and storage medium of the present invention will be described by way of example below. However, all descriptions should not be construed as limiting the present invention in any way.
[0020] Furthermore, for any single technical feature described or implied in the embodiments mentioned herein, or any single technical feature shown or implied in the various figures, the present invention still allows for any combination or deletion of these technical features (or their equivalents) without any technical obstacle, and thus these further embodiments according to the present invention should also be considered within the scope of this description.
[0021] Figure 1 This is a schematic block diagram of an embodiment of the coal mining machinery spare parts similarity identification and coding unified system of the present invention.
[0022] like Figure 1 As shown, in this embodiment, the system includes a data acquisition module, a cleaning and preprocessing module, a feature extraction and matching engine, a similarity calculation and clustering module, and a master data generation and system integration module. The data acquisition module is communicatively connected to the cleaning and preprocessing module, which in turn is communicatively connected to the feature extraction and matching engine. The feature extraction and matching engine is communicatively connected to the similarity calculation and clustering module, which is also communicatively connected to the master data generation and system integration module. Furthermore, the data acquisition module and the master data generation and system integration module are communicatively connected to multiple business systems.
[0023] The data acquisition module is used to collect spare parts data for coal mining machinery through data interfaces of multiple business systems within the coal mining enterprise. These systems include ERP (Enterprise Resource Planning), WMS (Warehouse Management System), and MRO (Maintenance, Repair & Operations). Spare parts data includes name, code, BOM (Bill of Materials), equipment location, function, and specifications.
[0024] The cleaning and preprocessing module is used to standardize the format of spare parts data. Because different systems have significant differences in input rules and habits, this module receives the collected spare parts data and cleans and preprocesses it, including standardizing capitalization for both Chinese and English characters, removing meaningless symbols, eliminating redundant spaces, and converting between simplified and traditional Chinese characters. The purpose of this module is to eliminate differences caused by formatting variations, making subsequent analysis more accurate.
[0025] The feature extraction and matching engine not only analyzes the text itself in spare parts data, but also uses various methods to extract different types of features, including text similarity features, semantic features, structural features, and attribute features. By fusing these multiple features and subsequently inputting these feature vectors into a deep learning model for joint judgment, matching accuracy can be improved, and incomplete or non-standardized data can be effectively identified.
[0026] Optionally, text similarity features can be represented by calculating edit distance, Jaccard similarity, etc. Edit distance refers to the minimum number of edits required to transform one string into another; allowed edit operations include replacing, inserting, and deleting a character. Jaccard similarity is a metric used to measure the similarity between two sets; it represents similarity by calculating the ratio of the intersection to the union of the two sets, with a higher value indicating higher similarity.
[0027] Optionally, semantic features can be used to convert spare parts descriptions into vectors using large-scale pre-trained language models (such as BERT), and semantic similarity can be determined to identify spare parts with different descriptive text but the same meaning. BERT stands for Bidirectional Encoder Representations from Transformers, and it is a pre-trained model proposed by Google AI Research.
[0028] Optionally, structural features may include the relationship between components and the whole machine, the location of the equipment, and its function and purpose, as analyzed from the BOM table; attribute features may include the results of extracting and comparing information such as specifications, materials, and compatible models.
[0029] The similarity calculation and clustering module is used to integrate the above features, input them into the deep learning model for fusion judgment, calculate the comprehensive similarity between different spare parts data, and automatically cluster them according to the set threshold, grouping the same or highly similar spare parts into the same category, reducing problems such as "one item, multiple names" and "one item, multiple codes", and significantly improving the efficiency of data standardization.
[0030] Optionally, the similarity calculation and clustering module can support semi-automatic manual confirmation, and submit spare parts that are uncertain whether they should be classified into the same category for manual review and judgment.
[0031] The master data generation and system integration module is used to generate unified spare parts master data for each cluster, assigning a unique standard code, standardized name and standardized attributes, and synchronizing it to business systems such as ERP, supply chain, and warehousing through interfaces to achieve data consistency and standardized management across the entire enterprise, reduce inventory redundancy, improve supply chain collaboration efficiency, and update in real time to avoid business interruptions caused by data delays or inconsistencies.
[0032] Figure 2 This is a schematic block diagram of another embodiment of the coal mining machinery spare parts similarity identification and coding unified system of the present invention.
[0033] like Figure 2As shown, in this embodiment, in addition to the modules described in the previous embodiments, the system also includes a word segmentation and industry dictionary module, an encoding rule normalization module, an attribute information parsing module, and a human feedback and self-learning module. The word segmentation and industry dictionary module, the encoding rule normalization module, and the attribute information parsing module are all communicatively connected to the cleaning and preprocessing module, and also to the feature extraction and matching engine; the human feedback and self-learning module is communicatively connected to the similarity calculation and clustering module, and the master data generation and system integration module.
[0034] like Figure 2 As shown, after the spare parts data is processed by the cleaning and preprocessing module, it is processed in parallel through three links: the word segmentation and industry dictionary module, the encoding rule normalization module, and the attribute information parsing module.
[0035] The word segmentation and industry dictionary module is used to segment the cleaned spare parts names, breaking down a complete sentence or text into several words with independent meanings, and then combining this with a specialized dictionary for the coal mining machinery industry for professional recognition. This embodiment constructs an industry-specific thesaurus covering the coal mining equipment field, including commonly used component terms, model rules, and specifications. During word segmentation, the system prioritizes matching to the industry thesaurus, ensuring accurate recognition of industry-specific expressions, abbreviations, and model formats, reducing missegmentation and mismatches. For example, it can recognize industry terms such as "cutting motor" and "traction motor," and distinguish different models, specifications, and units of measurement. This ensures the accuracy of subsequent feature extraction, avoiding incorrect segmentation or misjudgment, and enabling the system to recognize complex situations such as coal mining industry-specific naming habits, synonyms, abbreviations, and model variations.
[0036] The encoding rule normalization module uses a large model to learn the encoding rules of different systems, transforming complex and diverse encodings into unified semantic variables. The large model can be fine-tuned to suit the coal industry, performing semantic parsing on spare part names and codes from different source systems. It automatically identifies naming patterns and encoding generation logic, abstracting various heterogeneous codes into unified semantic variables, thereby achieving unified encoding matching across systems and manufacturers. This allows the system to handle complex encoding variations such as different character formats, symbol differences, and positional transformations, overcoming the limitations of traditional reliance on original encoding patterns. Even with completely different spare part encoding rules, it can understand their true correspondences, improving the accuracy and robustness of cross-system matching of spare part data.
[0037] The attribute information parsing module is used to extract attribute features from spare parts data, including features such as material, compatible model, and function.
[0038] The outputs of the three links mentioned above are fused in the feature extraction and matching engine. The feature extraction and matching engine combine text similarity, semantic similarity, BOM structure, and attribute features to form a comprehensive feature. Among them, text similarity can be calculated based on the word segmentation results of the word segmentation and industry dictionary modules, while semantic similarity is calculated based on the semantic variables transformed by the encoding rule normalization module.
[0039] The manual feedback and self-learning module is used to review the output results of the similarity calculation and clustering modules and provide feedback on the review results for model optimization. This module receives the results of automatic spare parts grouping and high-similarity merging from the similarity calculation and clustering modules. For spare parts matches with similarity close to the threshold but uncertain, manual review is submitted to ensure the reliability of the results. The results of manual review can be dynamically fed back to the deep learning models of other modules in the system as new training samples, achieving online incremental learning. Through multiple iterations, the system continuously corrects the similarity weights and feature selection strategies. This self-learning mechanism enables the system to have adaptive optimization capabilities, allowing it to continuously optimize during practical applications, constantly improving the accuracy of spare parts matching and gradually reducing the proportion of manual intervention.
[0040] This invention also provides a unified method for similarity identification and coding of coal mining machinery spare parts. In one embodiment, the method includes steps such as collecting spare parts data, data preprocessing, extracting and fusing multi-dimensional features, similarity clustering, and unified coding of spare parts of the same type.
[0041] In the step of collecting spare parts data, raw data of coal mining machine spare parts is extracted from multiple business systems, including the name, code, specifications, material, compatible machine model, and functional description of the spare parts. In this embodiment, similarity identification is performed on the multi-source spare parts data of the 1860-WD coal mining machine. The spare parts data comes from three systems: ERP, WMS, and MRO, and a total of 6,842 records are collected.
[0042] In the data preprocessing step, the raw spare parts data extracted in the previous step is standardized, including unifying Chinese and English, standardizing units, cleaning up symbols, and mapping to a thesaurus. In this embodiment, 6580 valid data entries were retained after data cleaning. This step eliminates differences caused by writing format, making subsequent analysis more accurate.
[0043] Optionally, after performing the aforementioned preprocessing on the spare parts data, word segmentation and terminology normalization are then performed using a coal mining machinery industry dictionary. This identifies and unifies terms with different characters but the same meaning in multi-source spare parts data, ensuring the accuracy of subsequent feature extraction and the ability to recognize complex situations such as naming conventions, synonyms, and abbreviations unique to the coal mining industry. The coal mining machinery industry dictionary, as an industry thesaurus covering the coal mining equipment field, includes commonly used component terms, model rules, and specifications. During word segmentation, the industry thesaurus is prioritized for matching, thereby ensuring the correct identification of industry-specific expressions, abbreviations, and model formats, reducing mis-segmentation and mis-matching.
[0044] In the step of extracting and fusing multi-dimensional features, multi-dimensional features are extracted from the spare parts data after the above processing, and then fused. In this embodiment, multi-dimensional features include semantic features, as well as structured features such as structural features and attribute features.
[0045] For example, a large language model is used to convert spare parts data into semantic vectors. The large language model employs a base model and undergoes lightweight fine-tuning based on a corpus from the coal mining machinery domain, which includes coal mining machine manuals, industry standards, equipment manuals, and after-sales records. The base model is a technical paradigm of large-scale modeling, obtaining general knowledge representations through pre-training on massive amounts of unlabeled data. In this embodiment, Qwen3-32B (a dense structured large language model developed by Alibaba's Tongyi Qianwen team) is used as the base model. Each spare parts data... The semantic vector is obtained by encoding with this large language model: in, This represents the semantic embedding function of the model. This indicates that the semantic embedding generated by the model can identify complex device component expressions, such as "traction gear assembly" and "reduction transmission component." Semantic embedding algorithms are used to capture semantic information in text, allowing the degree of semantic similarity between different texts to be measured by distance in vector space.
[0046] This step utilizes a large-scale model fine-tuned across the industry to perform semantic analysis on spare part names and codes from different source systems. It automatically identifies naming patterns and code generation logic, abstracting various heterogeneous codes into unified semantic variables, thereby achieving standardized code matching across systems and vendors. Therefore, this method can handle complex coding variations such as different character formats, symbol differences, and positional transformations, overcoming the limitations of traditional reliance on original coding patterns. Even if spare part coding rules are completely different, it can understand their true correspondence, improving the accuracy and robustness of cross-system matching of spare part data.
[0047] For example, the extracted structured features include a specification parameter vector. Attribute vectors and BOM level embedding The attribute vector includes one-hot encoding of material, function, and location. Finally, the multi-modal features are combined to form a comprehensive feature. : in," " indicates a vector concatenation operation.
[0048] In the similarity clustering step, based on the aforementioned comprehensive features, a deep fusion network is used to calculate the comprehensive similarity between different spare parts data, and automatic clustering is performed according to a set threshold, grouping spare parts with a comprehensive similarity higher than the threshold into one category. By jointly judging the comprehensive features fused from the above multiple features, matching accuracy can be improved, and incomplete or non-standardized data can be effectively identified.
[0049] For example, any two spare parts data can be calculated using a deep fusion network. The similarity method is as follows: in, Spare parts data The corresponding comprehensive feature vector, This represents vector concatenation. The network parameters of this deep fusion network ( This is the weight matrix. (where is the bias vector), and ReLU is the linear rectified function. The Sigmoid activation function outputs... This represents the similarity score. The network learns the coupling relationship between semantic and structured features to achieve matching and identification of spare parts from different sources.
[0050] Optionally, the deep fusion network model is trained using the following binary cross-entropy loss function: in, For the average binary cross-entropy loss, The total number of training samples, For the actual matching labels of spare parts pairs, Indicates a pair of truly similar spare parts. Indicates dissimilar pairs.
[0051] For example, after obtaining the similarity score, based on the similarity matrix Using threshold Filter: when At that time, output candidate matching spare parts pairs This allows for manual verification, or it can be automatically marked as a similar spare part. In this embodiment, the threshold... The default value is 0.87.
[0052] This step combines the coding normalization results with multidimensional similarity calculation to automatically cluster identical or highly similar spare parts into the same group, and supports semi-automatic marking and confirmation of incomplete matches, thereby reducing problems such as "one item, multiple names" and "one item, multiple codes", and significantly improving data standardization efficiency.
[0053] In the process of unifying the coding of spare parts of the same type, spare parts in the same cluster are uniformly coded, and a spare part master data entry containing a unified code, a normalized name, and standardized attributes is generated for each cluster. The spare part master data entry is then synchronized to business systems such as ERP, WMS, and MRO through system interfaces to achieve cross-system data consistency and real-time updates, avoid business interruptions caused by data delays or inconsistencies, achieve data consistency and standardized management across the entire enterprise, reduce spare parts inventory redundancy, and improve supply chain collaboration efficiency.
[0054] To evaluate the effectiveness of this method, this embodiment compares it with a traditional rule-based classification and matching method on the same spare parts dataset. The traditional rule-based classification and matching method is based on conventional logic in the field, such as complete matching of field keywords, encoding prefix rules, and specification numerical threshold judgment. The results show that the above-mentioned coal mining machinery spare parts similarity identification and encoding unification method in this embodiment can accurately identify spare parts items in the dataset that are described differently but are actually the same, successfully clustering and merging originally scattered duplicate data, effectively improving accuracy compared to traditional rule-based classification and matching. Table 1 below shows the accuracy comparison results of different algorithms: Table 1 Comparison of accuracy of different algorithms In Table 1, "Matching method based on semantic understanding" refers to the above-mentioned unified method for similarity identification and encoding of coal mining machinery spare parts in this embodiment.
[0055] This invention also provides a computer-readable storage medium on which computer-executable instructions or computer programs can be stored. When the computer-executable instructions or computer programs are executed by a processor, they implement the steps of the coal mining machinery spare parts similarity identification and encoding unification method as described in any of the foregoing embodiments.
[0056] Specifically, the process described above for implementing the method for similarity identification and unified coding of coal mining machinery spare parts can be implemented as computer-executable instructions or computer software programs. For example, embodiments of the present invention may include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the above-described method. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device. When the computer program is executed by a processing device, it performs the functions defined in the methods of embodiments of the present invention.
[0057] It should be noted that the computer-readable medium described above in this invention can be a computer-readable signal medium, a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this invention, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0058] The modules described in the embodiments of the present invention can be implemented in software or hardware. The names of the modules are not, in some cases, limiting the scope of the module itself.
[0059] Some embodiments of this invention address issues such as inconsistent naming, chaotic coding rules, and lack of standardization in multi-source spare parts data within the coal mining equipment supply chain. By combining industry knowledge bases with algorithms such as large-scale pre-trained language models, semantic similarity calculation, and multi-feature deep learning, it automatically processes large-scale data in batches, performing spare parts similarity matching and data unification. This achieves high-precision unification and systematic management of spare parts data. Compared to manual methods and simple rule-based matching methods, it effectively improves processing efficiency, semantic understanding, accuracy, and cross-system generalization, making it particularly suitable for accurately identifying spare parts with significant textual differences. Furthermore, it possesses a self-learning mechanism to continuously optimize spare parts similarity matching. The generated unified spare parts master data can automatically interface with existing business systems, achieving multi-system data consistency, meeting the needs of standardized and efficient management of coal mining equipment supply chain data, reducing inventory redundancy, and improving supply chain collaboration efficiency.
[0060] The technical scope of this invention is not limited to the contents of the above specification. Those skilled in the art can make various modifications and variations to the above embodiments without departing from the technical concept of this invention, and all such modifications and variations should fall within the scope of this invention.
Claims
1. A unified system for similarity identification and coding of coal mining machinery spare parts, characterized in that, The system includes: The data acquisition module is used to extract raw data of coal mining machine spare parts from multiple business systems. The raw data includes the name, code, description of the coal mining machine spare parts, and information about the equipment where the spare parts are located. The cleaning and preprocessing module is used to perform uniform formatting on the raw data and output preprocessed data. A feature extraction and matching engine is used to extract multi-dimensional features from the preprocessed data. These multi-dimensional features include text similarity features, semantic features, structural features, and attribute features. The semantic features are extracted using a large model. Specifically, the text similarity features include edit distance or Jaccard similarity between strings; the structural features include the bill of materials, location, and function of the equipment containing the spare part; and the attribute features include specifications, material, compatible models, and function. The semantic features are obtained by learning the spare part coding rules of different business systems using a large model, converting the spare part data containing the spare part code into a unified semantic vector, and using this semantic vector as the semantic feature of the spare part. The similarity calculation and clustering module is used to calculate the comprehensive similarity between different coal mining machinery spare parts data through a deep fusion network based on the multi-dimensional features, and to automatically cluster them according to a set threshold, classifying spare parts with a comprehensive similarity higher than the threshold into one category. The master data generation and system integration module is used to generate unified spare parts master data for each cluster and synchronize the spare parts master data to the multiple business systems. The similarity calculation and clustering module calculates the comprehensive similarity through a deep fusion network in the following way: in, This is a comprehensive feature vector of different spare parts data. This represents vector concatenation. Here are the network parameters, and ReLU is the linear rectified function. The Sigmoid activation function outputs... This represents the overall similarity score.
2. The system as described in claim 1, characterized in that, The system also includes a word segmentation and industry dictionary module, used to segment the preprocessed data and, in conjunction with a coal mining machinery industry-specific dictionary, to perform specialized identification on the preprocessed data. This specialized identification includes recognizing industry terminology, as well as identifying the model, specifications, and units of measurement of spare parts or equipment. The word segmentation and industry dictionary module is also used to calculate the text similarity features between spare parts data based on the results of the word segmentation processing and the specialization recognition.
3. The system as described in claim 2, characterized in that, The large model is a finely tuned large language model based on a corpus of coal mining machinery, which includes coal mining machinery manuals, industry standards, equipment manuals, and after-sales records.
4. The system as described in claim 1, characterized in that, The feature extraction and matching engine integrates the text similarity features, semantic features, structural features, and attribute features to form a comprehensive feature.
5. The system as described in claim 1, characterized in that, The deep fusion network is trained using a binary cross-entropy loss function: in, For the average binary cross-entropy loss, The total number of training samples, For the actual matching labels of spare parts pairs, Indicates a pair of truly similar spare parts. Indicates dissimilar pairs.
6. The system as described in claim 1, characterized in that, The system also includes a human feedback and self-learning module, which is used to manually review the clustering results when the similarity calculation and clustering module performs clustering based on similarity, and to use the results of the human review for the self-optimization of the system.
7. A method for similarity identification and unified coding of coal mining machinery spare parts based on the system described in any one of claims 1-6, characterized in that, The method includes: The raw data of coal mining machinery spare parts is extracted from multiple business systems. The raw data includes the name, code, description of the coal mining machinery spare parts and information about the equipment where the spare parts are located. The original data is formatted uniformly to obtain preprocessed data; Multi-dimensional feature extraction is performed on the preprocessed data, including text similarity features, semantic features, structural features, and attribute features; Based on the multi-dimensional features, the comprehensive similarity between different coal mining machinery spare parts is calculated through a deep fusion network, and automatic clustering is performed according to a set threshold, classifying spare parts with a comprehensive similarity higher than the threshold into one category; Generate unified spare parts master data for each cluster, and synchronize the spare parts master data to the multiple business systems.
8. The method as described in claim 7, characterized in that, The multi-dimensional feature extraction includes: The preprocessed data is segmented into words and then combined with a coal mining machinery industry-specific dictionary to perform specialized identification. This specialized identification includes identifying industry terms and identifying the model, specifications, and units of measurement of spare parts or equipment. Based on the results of the word segmentation and the specialized recognition, the text similarity features between spare parts data are calculated, including the edit distance or Jaccard similarity between strings. The spare parts coding rules of different business systems are learned by using a large model, and the spare parts data containing spare parts codes are converted into a unified semantic vector, which is then used as the semantic feature of the spare parts. Extract the attribute characteristics of the spare parts, including specifications, material, compatible models and functions; Extract the structural features of the spare parts, including the bill of materials, location, and function of the equipment in which the spare parts are located; The text similarity features, semantic features, structural features, and attribute features are integrated to form a comprehensive feature.
9. A computer-readable storage medium having computer-executable instructions or a computer program that, when executed by a processor, implements the method as claimed in claim 7 or 8.
Citation Information
Patent Citations
Data grading and classifying method and device
CN119538118A
Ship multi-source data integration method based on Word2Vec model
CN119576917A