Material similarity calculation method and system based on high-dimensional features and multi-modal features

By fusing multimodal features and dynamically assigning weights to structured data, text descriptions, and visual data, a material knowledge graph is constructed, solving the material matching problem in supply chain material management and enabling efficient material similarity calculation and rapid response in emergency substitution scenarios.

CN120873636APending Publication Date: 2025-10-31AVIC GOLD NETWORK (BEIJING) TECHNOLOGY CO LTD

Patent Information

Application Number
CN202511383395.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-26
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing technologies in supply chain material management suffer from problems such as difficulty in matching materials across suppliers, low utilization of unstructured data, and slow response in emergency replacement scenarios. In particular, they are insufficient in multimodal feature fusion and large-scale data processing efficiency, and cannot effectively solve the problems of accuracy and robustness in similarity calculation of key production materials.

Method used

By preprocessing structured data, text descriptions, and visual data, multimodal features are extracted, and a dynamic weight allocation strategy is used for weighted fusion to construct a material knowledge graph. Then, the HNSW graph index is used to calculate similarity, thereby achieving accurate matching between materials.

Benefits of technology

It enables more comprehensive material similarity calculation, improves matching accuracy and data utilization, enhances the efficiency of large-scale data processing, reduces manual intervention, and improves response speed and work efficiency in emergency replacement scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120873636A_ABST
    Figure CN120873636A_ABST
Patent Text Reader

Abstract

The invention discloses a material similarity calculation method and system based on high-dimensional features and multi-modal features, and belongs to the technical field of supply chain management. The method comprises the following steps: carrying out multi-modal feature extraction on preprocessed structured data, text description and visual data, and respectively generating a structured vector, a text vector and a visual vector; synchronizing the structured data through metadata, entity names and time, and establishing association between the structured data and text description and visual data; carrying out weighted fusion on the structured vector, the text vector and the visual vector by adopting a dynamic weight distribution strategy to generate a unified material feature vector; according to the multi-modal features, identifying entities and extracting relationships, and constructing and storing a material knowledge graph; and calculating the similarity, and retrieving and outputting a matching result. The system comprises a multi-modal feature extraction module, a cross-modal association module, a feature fusion module, a knowledge graph module and a similarity calculation and retrieval module. According to the method, the weight optimization of each modal is dynamically adjusted, and the matching precision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of supply chain management technology, and in particular to a method and system for calculating material similarity based on high-dimensional features and multimodal features. Background Technology

[0002] In supply chain material management, there are prominent problems such as difficulty in matching materials across suppliers, low utilization of unstructured data, and slow response in emergency replacement scenarios. Existing technologies are insufficient in terms of multimodal feature fusion, industry adaptability, and efficiency in large-scale data processing.

[0003] Some existing material similarity calculation methods do not effectively combine unstructured data (such as CAD drawings, 3D models, quality inspection reports, etc.) and structured data (such as physical properties, geometric parameters, document attributes, etc.), and some have low industry adaptability, requiring further focus on technologies (such as few-shot learning).

[0004] Critical production materials refer to materials that have a critical impact on a company's production operations, supply chain stability, or core competitiveness. They are often characterized by high value, high risk, strong supply dependence, or close ties to core business. Core components of aero engines (such as turbine blades), high-end chips, and specialty chemical raw materials are all examples of critical production materials. Shortages or mismatches of these materials can lead to significant losses such as production line shutdowns and project delays. Therefore, in similarity calculations, it is essential to prioritize ensuring the consistency of structured parameters (such as dimensional accuracy and material certification).

[0005] For example, patent document CN119691151A, entitled "Material Similarity Retrieval Method and System," discloses the following method: Material information is obtained based on a material data table, including material name, model / specification, and long description. The obtained material information is preprocessed to obtain preprocessed material information. The preprocessed material name is input into a material name similarity model to obtain the material name similarity between the target material and existing materials in the current material pool. The preprocessed model / specification is input into a model / specification similarity model to obtain the model / specification similarity between the target material and existing materials in the current material pool. The preprocessed long description is input into a long description similarity model to obtain the long description similarity between the target material and existing materials in the current material pool. Based on the obtained material name similarity, model / specification similarity, and long description similarity, the similarity between the target material and existing materials in the material pool is calculated using configured preset weights and adjustment parameters. Although it can calculate the similarity of material name, model specification, and long description separately, and obtain the total similarity between the target material and existing materials by configuring preset weights and adjusting parameters, it is limited to the processing of single-modal text data, does not take into account unstructured data, and lacks dynamic weight allocation and high-dimensional feature fusion strategies for cross-modal attention mechanisms. It cannot solve the problems of low utilization of unstructured data and cross-modal semantic alignment. In scenarios with heterogeneous material descriptions and complex matching of multiple attributes across suppliers, its accuracy and robustness are insufficient.

[0006] For example, patent document CN117974016A, entitled "A Rapid Comparison and Analysis Method, Device, and Medium for Definable BOM Dimensions," discloses the following: Based on the business requirements of the bill of materials (BOM), the definition dimensions of the BOM are determined; attributes are added to the defined dimensions and pre-stored material dimensions to obtain dedicated material dimensions; multi-level data structure partitioning is performed on the dedicated material dimensions to obtain a multi-level BOM model; through the multi-level BOM model, multi-source material data is integrated and processed according to the relevant hierarchical structure to obtain integrated material data; based on the enterprise IDP framework, the integrated material data is compared and analyzed to determine the difference marker data; and the difference marker data is visualized. While this can solve problems such as difference identification, data consistency, version control, data traceability, and decision support, and help improve production efficiency and quality, and reduce costs and risks, it focuses on BOM dimension comparison, does not involve multi-modal feature fusion, and the data types processed are limited to BOM-related dimensions, ignoring the issue of material description differences across suppliers in the supply chain; it does not consider efficient processing of large-scale data, and its performance is limited when the data volume is large.

[0007] Therefore, it is crucial to effectively integrate multimodal features, efficiently process large-scale data, and construct knowledge graphs to achieve accurate material similarity calculations. Summary of the Invention

[0008] In view of this, the main objective of the present invention is to provide a material similarity calculation method and system based on high-dimensional features and multimodal features, in order to at least partially solve the above-mentioned technical problems.

[0009] To achieve the above objectives, as a first aspect of the present invention, a material similarity calculation method based on high-dimensional features and multimodal features is proposed, comprising: S1: Preprocess the raw data, which includes structured data, text descriptions, and visual data; S2: Perform multimodal feature extraction on the preprocessed structured data, text descriptions, and visual data to generate structured vectors, text vectors, and visual vectors, respectively; S3: Establish a connection between structured data and text descriptions and visual data through metadata, entity names, and time synchronization; S4: Employ a dynamic weighting strategy to weight and fuse structured vectors, text vectors, and visual vectors to generate a unified material feature vector; S5: Identify entities and extract relationships based on multimodal features, and construct and store a material knowledge graph; S6: Calculate the similarity between materials based on the material feature vectors, retrieve and output the matching results through the HNSW graph index.

[0010] The preprocessing includes: Perform missing value imputation, outlier filtering, and numerical normalization on structured data; The text description is processed by removing special symbols, using a unified encoding format of UTF-8, and then segmented into word sequences using a word segmentation tool. Perform resolution unification, grayscale conversion, or color enhancement on visual data.

[0011] The multimodal feature extraction includes: Map non-numerical data in structured data to low-dimensional vectors, and normalize numerical data in structured data into vectors; The text description is segmented and encoded to generate multiple semantic vectors, and keywords are given additional weights. Extract multidimensional global features from material images, perform edge detection to construct local feature maps, align material images with text descriptions, and generate multidimensional visual vectors.

[0012] Step S3 includes: Visual data and text descriptions are stored in the form of files and associated with structured data through metadata; Associating the same entity across different modalities using entity names; Multimodal features are synchronized in time to ensure that the time reference of multimodal data is consistent.

[0013] The feature fusion described in step S4 includes: S41: Dynamic weight allocation: Set structured vector weights based on material type. Text vector weights Visual vector weights ,and + ; S42: Fusion Vector The calculation formula is: in, For structured vectors, For text vectors, Visual vectors; S43: Dimensionality Reduction Optimization: The fusion vector is reduced in dimension using the PCA algorithm while retaining variance information.

[0014] Step S5, the knowledge graph construction, includes: S51: Map entities in multimodal features to nodes in a knowledge graph; S52: Extract entity relationships from text descriptions and visual relationships from material images using a pre-trained model, and verify the consistency between entity relationships and visual relationships through a bidirectional attention mechanism; S53: Determine the node type and relation type, store the entity nodes, relation types and fusion vectors in the graph database, and configure vector indexes for entity nodes to accelerate retrieval.

[0015] Step S6, calculating the similarity between materials, includes: S61: The similarity between the target material and the candidate material is calculated using the weighted Euclidean distance formula; S62: Cosine normalization of similarity, the formula is: Where q represents the fusion vector of the target material, and t represents the fusion vector of the candidate material. The Euclidean norm of the fusion vector of the target material is represented. The Euclidean norm of the fusion vector of candidate materials; S63: Retrieve the top k items with the highest similarity through the HNSW graph index, where k represents a positive integer set by the user.

[0016] It also includes step S7: adjusting the weights of each modality feature based on the user's click feedback on the search results.

[0017] The materials include industrial parts, raw materials or finished products, the visual data includes material appearance images, CAD drawings or 3D point cloud data, and the structured data includes material codes, specifications, supplier information, prices and inventory status.

[0018] The dynamic weight allocation strategy is as follows: the weight of the structured vector is 0.5-0.8, the weight of the text vector is 0.1-0.3, and the weight of the visual vector is 0.05-0.2; the sum of the weights of the structured vector, the text vector, and the visual vector is 1; the weights are dynamically adjusted according to the material type, wherein the weight of the structured vector for key production materials is increased to 0.7-0.8, and the weight of the visual vector for standardized materials is increased to 0.15-0.2.

[0019] As a second aspect of the present invention, a material similarity calculation system based on high-dimensional features and multimodal features is also proposed, comprising: The multimodal feature extraction module is used to extract features from the structured data, text descriptions, and visual data of materials, and generate structured vectors, text vectors, and visual vectors, respectively. The cross-modal association module is used to establish associations between structured data and text descriptions and visual data through metadata association, entity name association, and time synchronization; The feature fusion module is used to perform weighted fusion of the structured vector, text vector and visual vector using a dynamic weight allocation strategy to generate a unified material feature vector; The knowledge graph module is used to construct a knowledge graph of material entities and relationships. The knowledge graph includes entity nodes, relationship types, and vector indexes. The similarity calculation and retrieval module is used to calculate the similarity between materials based on their feature vectors, and retrieves and outputs matching results through the HNSW graph index.

[0020] The multimodal feature extraction module includes: The structured data processing unit normalizes material codes, specifications, supplier information, prices, and inventory, and maps them into low-dimensional structured vectors through the TransE model. The text processing unit uses BERT or BGE-M3 models to encode the technical parameters, usage descriptions, and procurement contract terms of the materials, generating text vectors. The vision processing unit uses ResNet or CLIP models to extract features from material appearance images, CAD drawings, and 3D point cloud data, generating visual vectors.

[0021] The cross-modal association module associates text descriptions with the same material entities in visual data through entity name matching, and performs time synchronization of multimodal features to ensure that the time reference of multimodal data is consistent.

[0022] The knowledge graph module includes: The entity recognition unit identifies material entities in cross-modal data through similarity calculation; The relation extraction unit uses a pre-trained language model to extract entity relations from text, a Detectron2 model to extract visual relations from images, and a bidirectional attention mechanism to verify the rationality of the relations. The graph database storage unit is used to store entity nodes, relation types, and vector indexes. The entity nodes contain multimodal fusion vectors of materials.

[0023] The similarity calculation and retrieval module calculates the similarity between materials based on an improved weighted Euclidean distance, using the following formula:

[0024] Among them, This represents the weight of the i-th feature. This represents the i-th dimension feature of the target material. Let represent the i-th feature of the material to be matched, and n represent the dimension of the feature vector; The HNSW graph index is used for approximate nearest neighbor retrieval, and the top k matching results are output.

[0025] Based on the above technical solution, it can be seen that the material similarity calculation method and system based on high-dimensional features and multimodal features of the present invention has at least one of the following beneficial effects compared with the prior art: 1. By combining structured data, text descriptions, and visual features, a more comprehensive material similarity calculation is achieved; the dynamic adjustment of the weights of each modality enables the algorithm to be optimized according to different business needs and material types, thereby improving matching accuracy.

[0026] 2. The use of similarity matrix block computation and parallel computing framework effectively improves the efficiency of large-scale data processing; the adoption of vector retrieval technology significantly reduces storage costs and speeds up querying, ensuring real-time response.

[0027] 3. It has achieved effective association between text, images, and structured data, enhancing the overall utilization of data; by constructing a knowledge graph, it not only helps to better understand the relationships between materials, but also supports rapid retrieval and decision-making.

[0028] 4. Solutions were provided to address the challenges of material matching difficulties and slow response times in emergency substitution scenarios within the supply chain, reducing manual intervention and improving work efficiency and accuracy. Attached Figure Description

[0029] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0030] Figure 1 This is a flowchart illustrating the multimodal feature extraction module of the material similarity calculation system based on high-dimensional features and multimodal features of the present invention. Figure 2 This is a schematic diagram of the similarity calculation and retrieval module of the material similarity calculation system based on high-dimensional features and multimodal features of the present invention; Figure 3 This is a flowchart illustrating an embodiment of the material similarity calculation method based on high-dimensional features and multimodal features of the present invention. Figure 4 This is a flowchart illustrating Embodiment 2 of the material similarity calculation method based on high-dimensional features and multimodal features of the present invention. Detailed Implementation

[0031] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to specific embodiments and accompanying drawings.

[0032] The terminology used in this invention is for the purpose of describing particular embodiments only and is not intended to limit the embodiments of the invention. The singular forms “a,” “the,” and “the” as used in the embodiments of the invention and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.

[0033] The inventors discovered that existing technologies suffer from several problems in supply chain material management, including difficulties in matching materials across suppliers, low utilization of unstructured data, slow response to emergency replacement scenarios, insufficient consideration of capability differences at each level when optimizing joint economic batches for multi-product, multi-level supply chains, difficulty in freely configuring BOM dimensions, and challenges in comparing and querying multiple versions with poor identifiability. Through in-depth research, they found that by integrating high-dimensional features and multimodal features to calculate material similarity, using hybrid intelligent algorithms to optimize joint economic batches, constructing a multi-level BOM model with definable dimensions, and combining it with an enterprise IDP framework for comparative analysis, it is possible to improve the accuracy and efficiency of material matching, reduce the overall cost of the supply chain, and enhance the flexibility and identifiability of BOM version comparison analysis.

[0034] Therefore, as Figure 3As shown, the inventors have proposed a material similarity calculation method based on high-dimensional features and multimodal features, including: S1: Preprocess the raw data, which includes structured data, text descriptions, and visual data; S2: Perform multimodal feature extraction on the preprocessed structured data, text descriptions, and visual data to generate structured vectors, text vectors, and visual vectors, respectively; S3: Establish a connection between structured data and text descriptions and visual data through metadata, entity names, and time synchronization; S4: Employ a dynamic weighting strategy to weight and fuse structured vectors, text vectors, and visual vectors to generate a unified material feature vector; S5: Identify entities and extract relationships based on multimodal features, and construct and store a material knowledge graph; S6: Calculate the similarity between materials based on the material feature vectors, retrieve and output the matching results through the HNSW graph index.

[0035] Structured data contains fundamental information about materials, such as material codes, specifications (size, weight, material, etc.), supplier information, price, and inventory status. These features can be converted into vector representations through numerical normalization. Converting size parameters (length, width, height) of different magnitudes and dimensions into vectors within a unified range (usually [0, 1]) eliminates the impact of dimensional differences on subsequent similarity calculations. For example, the original values ​​of 10cm (length) and 2cm (height) differ significantly; directly using them for calculations would overemphasize the influence of length on similarity. Normalization ensures a more reasonable weighting of features across dimensions. For instance, material dimensions (e.g., 10cm × 5cm × 2cm) can be normalized to a vector of [0.8, 0.4, 0.1], and supplier IDs can be mapped to low-dimensional vectors using the TransE model. Due to their explicit numerical relationships, structured features are typically assigned higher weights (e.g., 0.7) in similarity calculations, especially for key parameters (e.g., price, delivery time).

[0036] Normalization presupposes that the maximum and minimum dimensions of this type of material are clearly defined. These maximum and minimum values ​​are typically derived from historical data or industry standards. Assuming that in a supply chain scenario, the size range of this type of material (such as small mechanical parts) is: Length: Minimum 1cm, maximum 12.5cm (corresponding to the common length range of this type of part in the business). Width: Minimum 1cm, maximum 12.5cm; Height: Minimum 0.5cm, maximum 20cm (the height range is wider because it includes sheet-like parts).

[0037] The following normalization formula is used. ,in, Indicates the original size value. This represents the minimum value of that dimension. This represents the maximum value of that dimension. This represents the normalized value.

[0038] Substituting the length 10cm into the formula, we get... Substituting the width of 5cm into the formula, we get... Rounded to 0.4; based on a height of 2cm, substituting into the formula, we get... That is, the vector after normalization based on the material size is [0.8, 0.4, 0.1].

[0039] Textual descriptions include material technical parameters, usage specifications, and procurement contract terms. These long texts can be segmented and encoded using a text embedding model to generate multiple vectors. For example, BERT can generate a 768-dimensional semantic vector for a description such as "304 stainless steel, corrosion-resistant, suitable for the food industry." Textual features are particularly important in determining the similarity of material functions and uses, and are typically assigned a moderate weight (e.g., 0.2).

[0040] Visual data includes material appearance images, packaging designs, 3D point cloud data, etc. These features can be extracted using CLIP or ViT models for visual embedding, or analyzed using VL models combined with natural language processing. For example, the CLIP model can align material images with text descriptions to generate 512-dimensional visual vectors. Visual features have unique advantages in judging the similarity of material appearance and packaging, and are usually assigned low weights (e.g., 0.1), but their weights can be increased in certain specific scenarios (e.g., alternative material selection).

[0041] The present invention will be further illustrated below through specific embodiments. It should be noted that the following embodiments are merely illustrative and not intended to limit the scope of the invention. All other embodiments obtained by those skilled in the art based on the embodiments shown below without inventive effort are within the scope of protection of the embodiments of the present invention.

[0042] Example 1 In this embodiment, further, Figure 3 This is a flowchart illustrating the first embodiment of the material similarity calculation method based on high-dimensional features and multimodal features of the present invention. In this embodiment, the similarity calculation of "bolt" material is used as an example to explain the method flow in detail: like Figure 1 As shown, the structured vector representation of the bolt is: ; Text vector Generate a 768-dimensional vector using BERT. Among them, the weight of keywords such as "304 stainless steel" and "food machinery" increased by 10%; Visual vectors Image features are extracted using ResNet50 to generate 2048-dimensional visual vectors. .

[0043] Structured data and text / images are linked by the material code "BOLT001", and text descriptions and image annotations are linked by the entity name "bolt". Data synchronization is achieved based on timestamps (such as 2025-01-01 14:30).

[0044] Dynamic weights =0.7, fusion vector The covariance is calculated using the following formula, and then the PCA algorithm is used to maximize the retention of variance information in the data: , Where n represents the total number of data points. This represents the standardized data matrix. express The transpose of the matrix, Indicates the preceding A matrix composed of eigenvectors This represents a low-dimensional data matrix, and the eigenvector retains 80% or 95% of the data features as needed.

[0045] This is because multimodal feature fusion generates high-dimensional vectors (e.g., 500-dimensional structured vectors + 768-dimensional text vectors + 2048-dimensional visual vectors, resulting in a total dimension of 3268). Directly using these vectors for similarity calculations would lead to problems such as high computational cost (distance calculation for 3268-dimensional vectors is time-consuming) and noise interference (redundant dimensions may introduce irrelevant information). Therefore, PCA is needed to project the high-dimensional vectors into a low-dimensional space (e.g., 512-dimensional), while retaining 80%-95% of the variance information (i.e., the main features of the original data).

[0046] In this embodiment, it is assumed that the multimodal fusion vectors of 10,000 materials have been obtained, forming a 10,000×3268 matrix (each row represents the 3268-dimensional features of a material, and each column represents a feature dimension). The 10,000 materials represent the total number of materials in this embodiment, i.e., the overall data.

[0047] PCA is sensitive to data scale, so the matrix X is first standardized to eliminate the influence of different dimensional units.

[0048] First, calculate the mean of each dimension (column). and standard deviation For each element Standardize (the j-th dimension feature of the i-th material): The standardized data matrix is ​​represented as Ensure that the mean of each dimension is 0 and the standard deviation is 1.

[0049] The covariance matrix is ​​used to describe the correlation between different dimensions, and the formula is: Where n represents the total number of materials, in this embodiment n=10000, and C represents a 3268×3268 symmetric matrix, with elements... This indicates the correlation between the j-th and k-th dimensions of the feature; a larger value indicates a higher correlation.

[0050] The eigenvalues ​​and eigenvectors of the covariance matrix C are the two most important outputs of PCA. The eigenvectors represent the "principal component directions" (the axes of the new low-dimensional coordinate system), and each eigenvector corresponds to a "principal component". The eigenvalues ​​represent the variance information contained in the principal component (the larger the eigenvalue, the more original information the principal component retains).

[0051] After solving the problem using eigenvalue decomposition (or SVD singular value decomposition), 3268 eigenvalues ​​were obtained. and the corresponding feature vector Each feature vector has 3268 dimensions.

[0052] Based on the principle of maximizing the retained variance, select the top k principal components whose cumulative variance contribution rate reaches the target value (e.g., 95%): Calculate the variance contribution rate of each principal component: ; Calculate the cumulative variance contribution rate: When this value is ≥95%, selection stops.

[0053] The cumulative variance contribution rate of the first 512 principal components was calculated to be 95.3%, meaning these 512 principal components retained 95.3% of the information from the original 3268-dimensional data. Therefore, the eigenvectors corresponding to these 512 principal components were selected. The projection matrix W (3268×512 dimensions) is formed.

[0054] The standardized high-dimensional matrix Multiplying by the projection matrix W yields a low-dimensional matrix. In this embodiment, Y represents a 10000×512 matrix, and each row represents a 512-dimensional low-dimensional feature of a material.

[0055] The vector dimension was reduced from 3268 to 512, decreasing the computational load of subsequent similarity calculations by 84%. This dimensionality reduction lowers computational complexity, significantly improving retrieval speed and meeting the real-time matching needs of the supply chain, especially for urgent alternative material queries requiring sub-second response times. 95% variance preservation means that the low-dimensional vector still contains key information from the original features, such as key material specifications, core semantics, and significant visual features, avoiding the loss of accuracy in similarity calculations caused by dimensionality reduction. For example, core features such as "3mm diameter" and "stainless steel material" of the bolt are still retained, ensuring matching accuracy. High-dimensional vectors may contain redundant dimensions (such as repeated adjectives in text descriptions or background noise in images). PCA, by selecting principal components, can filter out this redundant information, allowing the low-dimensional vector to focus more on the key features distinguishing material similarity (such as bolt size, material, and head shape).

[0056] Preprocessing includes filling missing values, filtering outliers, and normalizing the structured data; the normalization uses the min-max standardization method to map the data to the [0,1] interval; in this embodiment, the specifications of the bolt (diameter 3mm, length 5mm, material 304 stainless steel) are normalized to [0.3, 0.5, 0.8] by min-max. Special symbols are removed from the text description, the unified encoding format is UTF-8, and it is segmented into word sequences using a word segmentation tool; in this embodiment, the key sentence "304 stainless steel cross head bolt, suitable for food machinery, corrosion resistant" is retained after cleaning. The visual data is processed by unifying resolution, grayscale, or color enhancement; in this embodiment, the bolt image is adjusted to 224×224 pixels and grayscale processing is performed.

[0057] In this embodiment, the model used to extract image features (such as CLIP) is a pre-trained model. These models have fixed the input image size to 224×224 pixels (or 224×224 is its compatible standard size) during the training phase. For example, the visual encoder of the CLIP model also supports 224×224 pixels as standard input, ensuring that the convolution kernel sliding, pooling operations, etc. during image feature extraction match the model weights.

[0058] If the input image size is incorrect (e.g., too large or too small), it needs to be forcibly adjusted by stretching, cropping, etc., which may cause image distortion (e.g., the length-to-width ratio of the bolt is distorted), thus affecting the accuracy of visual features (e.g., misjudging the shape of the bolt head).

[0059] The reason for adjusting the bolt image to 224×224 pixels is to consider both feature preservation and computational cost. If the size is too small (e.g., 32×32 pixels), image details (e.g., bolt thread texture, head cross groove) will be lost, resulting in incomplete visual feature extraction and affecting subsequent similarity calculation (e.g., unable to distinguish between "cross groove bolt" and "slotted bolt"). If the size is too large (e.g., 1024×1024 pixels), although more details are preserved, the computational load will increase significantly. The number of image pixels increased from 224×224≈50,000 pixels to 1024×1024≈1,000,000 pixels, and the computational cost of the convolutional layer increased quadratically. Although the dimensions of the generated visual vectors (such as 2048 dimensions in ResNet) remain unchanged, the feature extraction time will increase several times, which cannot meet the needs of "real-time retrieval" in supply chain scenarios (such as the need for second-level response for matching emergency alternative materials).

[0060] The 224×224 pixel resolution preserves the key visual features of the bolt (such as size proportions, head shape, and thread details) while controlling computational complexity to ensure efficient system operation.

[0061] In this embodiment, further, as Figure 1 As shown, multimodal feature extraction includes: Map non-numerical data in structured data to low-dimensional vectors, and normalize numerical data in structured data into vectors; The text description is segmented and encoded using BERT or BGE-M3 models to generate multiple semantic vectors, and additional weights are assigned to keywords. The ResNet model is used to extract multidimensional global features from material images, perform edge detection to construct local feature maps, align material images with text descriptions, and generate multidimensional visual vectors.

[0062] Step S3 includes: Visual data and text descriptions are stored in the form of files and associated with structured data through metadata; Associating the same entity across different modalities using entity names; Multimodal features are synchronized in time to ensure that the time reference of multimodal data is consistent.

[0063] Step S4, feature fusion, includes: S41: Dynamic weight allocation: Set structured vector weights based on material type. Text vector weights Visual vector weights ,and + ; S42: Fusion Vector The calculation formula is: in, For structured vectors, For text vectors, Visual vectors; S43: Dimensionality Reduction Optimization: The fusion vector is reduced in dimension using the PCA algorithm while retaining variance information.

[0064] Step S5, knowledge graph construction, includes: S51: Map entities (such as "M3 bolt" and "304 stainless steel") in multimodal features to nodes in a knowledge graph; In this embodiment, extract entities related to materials. Text description: From "M3 cross head bolt, material is 304 stainless steel, suitable for food machinery", identify entities "M3 bolt" (material), "304 stainless steel" (material), and "food machinery" (application scenario); Image data: From the appearance image of the bolt, identify entities "bolt head" and "thread" (material component) through a target detection model (such as YOLO); Structured data: Extract entities "M3" (size standard) and "cross head" (head type) from the specification parameter table.

[0065] The extracted entities are standardized to avoid ambiguity: "M3 bolts" and "M3×5 bolts" are uniformly classified as "M3 specification bolts" to avoid redundancy in nodes due to differences in details; The terms "304 stainless steel" and "304 stainless steel" will be merged into "304 stainless steel" to unify the naming rules.

[0066] Each entity corresponds to a node in the knowledge graph. The node attributes include basic attributes and multimodal features, among which: Basic attributes include entity name (such as "M3 specification bolt") and type label (such as "material" or "material type"). Multimodal features include the fusion vector (512-dimensional) generated by associated S4 and references to the original features (such as the corresponding text fragment ID and image URL).

[0067] S52: Entity relationships are extracted from text descriptions using a pre-trained model, and visual relationships are extracted from material images. The consistency between entity relationships and visual relationships is verified using a bidirectional attention mechanism. In this embodiment, a pre-trained language model (such as BERT or ERNIE) is used to extract entity relationships from text descriptions. Enter the text: "M3 cross head bolt, made of 304 stainless steel, suitable for food machinery"; Extraction relationship: (M3 specification bolt, material is 304 stainless steel); (M3 specification bolt, suitable for food machinery). Visual relationships are extracted from bolt images using object detection and relationship recognition models (such as Detectron2 and VisualBERT): Input image: High-resolution image of an M3 bolt (including head, thread, and material texture); Identify entity pairs: (bolt head, including cross groove), (bolt, made of, metal material); Combined with the cross-modal association (entity name matching) in step S3, "metal material" is mapped to "stainless steel 304", resulting in the visual relationship: (M3 bolt, made of, stainless steel 304).

[0068] To ensure that textual and visual relationships do not conflict, cross-validation is performed using a bidirectional attention mechanism: On the one hand, the semantic similarity between the textual relation “(M3 bolt, material is, stainless steel 304)” and the visual relation “(M3 bolt, made of..., stainless steel 304)” is calculated by encoding relational phrases through a pre-trained model and calculating cosine similarity. On the other hand, the visual features (such as metallic luster and texture) of "stainless steel 304" in the image are checked from visual to text to see if they match the semantic features (such as "corrosion resistant" and "food grade") of the text description "stainless steel 304". If the similarity is ≥0.8 (threshold), the relationship is determined to be consistent and the relationship is retained; otherwise, it is marked as conflict (such as the text description is "iron" but the image is identified as "stainless steel", which requires manual verification).

[0069] S53: Determine the node type and relation type, store the entity nodes, relation types and fusion vectors in the graph database, and configure vector indexes for entity nodes to accelerate retrieval.

[0070] Define node types and relationship types based on the characteristics of supply chain materials: Node type: "Materials" (e.g., "M3 bolts"); Material (e.g., "304 stainless steel"); “Application scenarios” (e.g., “food machinery”); “Components” (such as “bolt heads”).

[0071] Relationship type: "Material is" (material → material); "Adapted to" (material → application scenario); "Includes" (materials → components); "Made from" (material → material).

[0072] In this embodiment, the Neo4j graph database is used to store the structured graph, implementing node storage, where each node contains an ID, name, type label, and fusion vector (512 dimensions); and relation storage, where each relation contains a starting node ID, a target node ID, relation type, and confidence level (e.g., when the text and visual relations are consistent, the confidence level is 0.9).

[0073] To accelerate the vector similarity retrieval in step S6, the fused vectors of nodes are synchronized to the vector database. In this embodiment, Milvus is used and an index is created, employing either IVF_FLAT (inverted index + Flat search) or HNSW (hierarchical approximate nearest neighbor) index. For 512-dimensional vectors, the number of cluster centers is set to 1024 (IVF_FLAT) or the search depth to 64 (HNSW) to ensure a retrieval latency ≤100ms. Vectors in the vector database are associated with nodes in Neo4j through unique IDs. During retrieval, similar vectors are first found through the vector index, and then the complete node attributes and relationships are obtained from Neo4j through the IDs.

[0074] Taking "M3 stainless steel bolts" as an example, the final knowledge graph fragment includes: Nodes: "M3 bolt", "304 stainless steel", "food machinery", "bolt head"; Relationship: (M3 bolt - material - 304 stainless steel), (M3 bolt - suitable for - food machinery), (M3 bolt - includes - bolt head). When a user queries "bolts with the same material as M3 stainless steel bolts", the "304 stainless steel" node can be quickly located through the graph, and then all materials with "material as" that node can be retrieved in reverse. The result with the highest similarity can be returned by combining the vector index.

[0075] The above process, through the fusion of multimodal entities and relationships, solves the limitation of traditional knowledge graphs relying on a single text data source. At the same time, the configuration of vector indexes lays the foundation for efficient retrieval in the next step S6, ultimately achieving accurate matching of semantic association and feature similarity.

[0076] Step S6, calculating the similarity between materials, includes: S61: The similarity between the target material and the candidate material is calculated using the weighted Euclidean distance formula; S62: Cosine normalization of similarity, the formula is: Where q represents the fusion vector of the target material and t represents the fusion vector of the candidate material; S63: Retrieve the top k items with the highest similarity through the HNSW graph index, where k represents a positive integer set by the user.

[0077] The target material is "M3×5 cross head bolt" to be matched; the weighted Euclidean distance D=0.12 between the target material and the candidate material "M3×6 cross head bolt" is calculated; the top 5 similarity materials are returned through the HNSW index, among which "M3×6 cross head bolt" is the best match.

[0078] Example 2 In this embodiment, as Figure 4 As shown, the method also includes step S7: adjusting the weights of each modal feature based on user click feedback on the search results. This dynamic optimization based on actual user preferences aims to make multimodal feature fusion more aligned with actual business needs by dynamically adjusting based on user feedback.

[0079] like Figure 2 As shown, when a user submits a query request containing multimodal information, in this embodiment, the request includes the text "Screw GB / T818-2016" and an image of a screw. The server first extracts semantic vectors from the text using BGE-M3 and enhances key parameters, extracts global and local features from the image using ResNet and Canny, and normalizes or encodes implicit structured parameters, including the M3×5 size, using word embedding. Then, dynamic weighted fusion is performed, with the initial weights being 0.5 for text, 0.3 for the image, and 0.2 for the structured parameters. After PCA dimensionality reduction to 128 dimensions, HNSW hierarchical indexing is performed with L=5 levels. First, coarse-grained localization is performed at the top level, and then fine-grained matching is performed at the bottom level using a mixture of cosine similarity and weighted Euclidean distance. In this embodiment, the Top-K result is "M3×5 cross-head screw". If the user provides a matching result, the server dynamically adjusts the modal weights, such as increasing the text weight, to achieve multimodal input to feature fusion, perform efficient retrieval, and then provide feedback optimization, achieving accurate and efficient matching of cross-heterogeneous descriptions of industrial materials.

[0080] Materials include industrial parts, raw materials or finished products; visual data includes material appearance images, CAD drawings or 3D point cloud data; and structured data includes material codes, specifications, supplier information, prices and inventory status.

[0081] The dynamic weight allocation strategy is as follows: the weight of the structured vector is 0.5-0.8, the weight of the text vector is 0.1-0.3, and the weight of the visual vector is 0.05-0.2; the sum of the weights of the structured vector, the text vector, and the visual vector is 1; the weights are dynamically adjusted according to the material type, wherein the weight of the structured vector for key production materials is increased to 0.7-0.8, and the weight of the visual vector for standardized materials is increased to 0.15-0.2.

[0082] Structured data of critical production materials (such as specifications, supplier information, price, and delivery cycle) has a greater impact on similarity calculations, and therefore receives higher weight (e.g., 0.7-0.8) in multimodal feature fusion. This is because key parameters of such materials (such as material compatibility, delivery cycle, and supplier qualifications) directly determine whether they can meet production requirements; the matching priority of function and performance is far higher than secondary features such as appearance. Multimodal feature fusion is a crucial step in material similarity algorithms. A weighted fusion strategy is adopted, dynamically adjusting the weights of each modality based on material type and business needs. For example, for critical production materials, the weight of structured attributes can be increased (e.g., 0.8); for standardized materials, the weight of visual features can be increased (e.g., 0.3). This dynamic weighting mechanism optimizes similarity calculation results across different scenarios.

[0083] Standardized materials refer to materials with unified industry standards or general specifications, high substitutability, and wide supply channels. They are often of lower value, with stable demand, and their core characteristics are consistency in appearance and general parameters. Common bolts, nuts, standard resistors, and general packaging materials are all examples of standardized materials. Industry standards (such as GB / T and ISO) clearly define the specifications of these materials, and products from different suppliers have minimal functional differences. Therefore, similarity can be quickly determined through visual characteristics (such as appearance images), meeting the need for efficient matching.

[0084] Because the visual features of standardized materials (such as appearance images and packaging designs) have a greater impact on similarity calculations, the weight of visual vectors is increased (e.g., 0.15-0.2) in multimodal feature fusion. This is because the function and specifications of such materials are highly standardized, and the consistency of appearance and packaging is more likely to serve as a basis for rapid matching (e.g., bolt head shape, packaging specifications).

[0085] In supply chain material management applications, material similarity algorithms need to process large-scale data. Therefore, a similarity matrix block calculation is adopted to achieve an efficient processing strategy for large-scale material data (such as 1 million-level materials).

[0086] Calculating the global similarity matrix of 1 million materials directly would present two core problems: an explosion in computational load and redundant storage.

[0087] Because the matrix contains Each element (each element representing the similarity between two materials) takes approximately 1 ns to calculate, resulting in a total processing time of about 115 days, which is completely insufficient to meet real-time requirements. Storing this matrix requires approximately 4 TB of space (in float32 type, with each element being 4 bytes), far exceeding the storage capacity of a conventional server.

[0088] Therefore, block computing, by "breaking down a large matrix into smaller parts," can solve the aforementioned problems by decomposing a large matrix into sub-matrices that can be processed in parallel. By dividing data into blocks according to material categories or suppliers, and combining them with parallel computing frameworks, computational efficiency can be effectively improved. For example, if materials are divided into 100 categories, each containing 10,000 materials, then 100 10,000-fold similarity sub-matrices can be computed in parallel, and finally merged into a complete similarity matrix.

[0089] First, the data is divided into blocks according to the category of the materials. That is, based on the category attribute of the materials (in this embodiment, the material type is used as the dividing basis), 1 million materials are divided into 100 categories, with 10,000 materials in each category.

[0090] Category 1: Electronic components (resistors, capacitors, etc., mostly key production materials); Category 2: Standard mechanical parts (bolts, nuts, etc., mostly standardized materials); Category 3: Packaging materials (cardboard boxes, tape, etc., standardized materials); ... (100 categories in total).

[0091] The similarity calculation for similar materials is more demanding (e.g., when looking for substitutes for resistors, it is more likely to be matched within the "electronic components" category). Different categories of materials have different feature weights (e.g., key production materials emphasize structural features, while standardized materials emphasize visual features), and weight strategies can be applied in a targeted manner after segmentation.

[0092] Next, using parallel computing frameworks such as Spark or Kubernetes, the 100 categories are distributed across 100 computing nodes (or 100 CPU cores), with each node independently computing its own submatrix for each category. Node 1 calculates a 10000×10000 submatrix for the "Electronic Components" category (storing the similarity between all materials within this category); Node 2 calculates the 10000×10000 submatrix of the "Mechanical Standard Parts" category; ... (All nodes start computing simultaneously).

[0093] Elements of each submatrix To calculate the similarity between the i-th material and the j-th material within a category, the feature weights of that category are used: For critical production materials (such as electronic components): use structured feature weights of 0.8, text weights of 0.15, and visual weights of 0.05, and calculate based on the improved weighted Euclidean distance; For standardized materials (such as bolts): use a structured feature weight of 0.5, text weight of 0.2, and visual weight of 0.3, and calculate according to the above formula.

[0094] Prioritize retaining similar submatrices: After calculating 100 nodes, 100 submatrices of 10000×10000 are obtained, which are then merged into a block diagonal matrix, meaning that there are submatrices only on the diagonal, and the off-diagonal regions are 0 or not calculated for the time being; Supplement cross-class similarity as needed: If users need cross-class matching, such as finding temporary conductors to replace resistors, they can temporarily start cross-class submatrix calculation, such as the submatrix of electronic components and conductor material categories, but it is not calculated by default to save resources.

[0095] The effective computational cost of the global matrix is ​​from Reduce to This represents a 99% reduction, with storage requirements decreasing from 4TB to 40GB, and the total computation time being significantly shortened through parallel computing on 100 nodes.

[0096] The submatrix calculation for critical production materials (such as electronic components) focuses on structured features, while the submatrix calculation for standardized materials (such as bolts) focuses on visual features. After segmentation, weights can be applied independently, avoiding the accuracy loss caused by globally uniform weights. More than 90% of the substitute material matching needs in the supply chain occur within the same type of material (such as the substitute material for resistors is still resistors). After segmentation, submatrix calculations of the same type are prioritized, ignoring low-value cross-category calculations to improve efficiency. When adding new materials, it is only necessary to add the corresponding category and recalculate the submatrix of that category (10000×10001), without reconstructing the global matrix, adapting to the dynamic changes in the supply chain.

[0097] Material similarity calculation systems based on high-dimensional features and multimodal features include: The multimodal feature extraction module is used to extract features from the structured data, text descriptions, and visual data of materials, and generate structured vectors, text vectors, and visual vectors, respectively. The cross-modal association module is used to establish associations between structured data and text descriptions and visual data through metadata association, entity name association, and time synchronization; The feature fusion module is used to perform weighted fusion of structured vectors, text vectors and visual vectors using a dynamic weight allocation strategy to generate a unified material feature vector. The knowledge graph module is used to build a knowledge graph of material entities and relationships. The knowledge graph includes entity nodes, relationship types, and vector indexes. The similarity calculation and retrieval module is used to calculate the similarity between materials based on their feature vectors, and retrieves and outputs matching results through the HNSW graph index.

[0098] The multimodal feature extraction module includes: The structured data processing unit normalizes structured data such as material specifications and supplier information, and maps it into low-dimensional structured vectors through the TransE model. The text processing unit uses BERT or BGE-M3 models to encode text descriptions such as the technical parameters and usage instructions of the materials, and generates text vectors. The vision processing unit uses ResNet or CLIP models to extract features from material images and 3D point cloud data, generating visual vectors.

[0099] The cross-modal association module associates text descriptions with the same material entities in visual data through entity name matching, and performs time synchronization of multimodal features to ensure that the time base of multimodal data is consistent.

[0100] By overcoming the barrier of large description differences through multimodal features and semantic association, "screws GB / T818-2016" and "M3×5 cross head screws" from different suppliers can be accurately matched; the problem of low utilization of unstructured data is solved, and visual data such as CAD drawings, 3D models, and quality inspection reports, as well as text descriptions, are incorporated into similarity calculations through feature extraction and association, improving the comprehensiveness of matching; the combination of efficient retrieval and semantic filtering reduces the time for substitute material matching from several hours of manual work to seconds in the system, greatly reducing the error rate.

[0101] The knowledge graph module includes: The entity recognition unit identifies material entities in cross-modal data through similarity calculation; The relation extraction unit uses a pre-trained language model to extract entity relations from text, a Detectron2 model to extract visual relations from images, and a bidirectional attention mechanism to verify the rationality of the relations; and The graph database storage unit is used to store entity nodes, relation types, and vector indexes. Entity nodes contain multimodal fusion vectors of materials.

[0102] The similarity calculation and retrieval module calculates the similarity between materials based on an improved weighted Euclidean distance, using the following formula: Among them, This represents the weight of the i-th feature. This represents the i-th dimension feature of the target material. Let represent the i-th feature of the material to be matched, and n represent the dimension of the feature vector; The HNSW graph index is used for approximate nearest neighbor retrieval, and the top k matching results are output.

[0103] This invention integrates multimodal features of structured attributes, text descriptions, and visual features, and employs a dynamic weight allocation strategy, similarity matrix block calculation, and index retrieval to construct a knowledge graph to enhance data association and retrieval performance, effectively improving the accuracy and efficiency of material similarity calculation.

[0104] The foregoing has described specific embodiments of the present invention. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired results. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0105] In the description of the embodiments of the present invention, the terms "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the present invention. In the embodiments of the present invention, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in a suitable manner in any one or more embodiments or examples. Furthermore, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in the embodiments of the present invention, as well as the features of the different embodiments or examples.

[0106] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of embodiments of the present invention, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0107] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing custom logic functions or processes, and the scope of preferred embodiments of the invention includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as should be understood by those skilled in the art to which embodiments of the invention pertain.

[0108] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A material similarity calculation method based on high-dimensional features and multimodal features, characterized in that, include: S1: Preprocess the raw data, which includes structured data, text descriptions, and visual data; S2: Perform multimodal feature extraction on the preprocessed structured data, text descriptions, and visual data to generate structured vectors, text vectors, and visual vectors, respectively; S3: Establish a connection between structured data and text descriptions and visual data through metadata, entity names, and time synchronization; S4: Employ a dynamic weighting strategy to weight and fuse structured vectors, text vectors, and visual vectors to generate a unified material feature vector; S5: Identify entities and extract relationships based on multimodal features, and construct and store a material knowledge graph; S6: Calculate the similarity between materials based on the material feature vectors, retrieve and output the matching results through the HNSW graph index.

2. The material similarity calculation method based on high-dimensional features and multimodal features according to claim 1, characterized in that, The preprocessing includes: Perform missing value imputation, outlier filtering, and numerical normalization on structured data; The text description is processed by removing special symbols, using a unified encoding format of UTF-8, and then segmented into word sequences using a word segmentation tool. Perform resolution unification, grayscale conversion, or color enhancement on visual data.

3. The material similarity calculation method based on high-dimensional features and multimodal features according to claim 1, characterized in that, The multimodal feature extraction includes: Map non-numerical data in structured data to low-dimensional vectors, and normalize numerical data in structured data into vectors; The text description is segmented and encoded to generate multiple semantic vectors, and keywords are given additional weights. Extract multidimensional global features from material images, perform edge detection to construct local feature maps, align material images with text descriptions, and generate multidimensional visual vectors.

4. The material similarity calculation method based on high-dimensional features and multimodal features according to claim 1, characterized in that, Step S3 includes: Visual data and text descriptions are stored in the form of files and associated with structured data through metadata; Associating the same entity across different modalities using entity names; Multimodal features are synchronized in time to ensure that the time reference of multimodal data is consistent.

5. The material similarity calculation method based on high-dimensional features and multimodal features according to claim 1, characterized in that, S4 include: S41: Dynamic weight allocation: Set structured vector weights based on material type. Text vector weights Visual vector weights ,and ; S42: Fusion Vector The calculation formula is: in, For structured vectors, For text vectors, Visual vectors; S43: Dimensionality Reduction Optimization: The fusion vector is reduced in dimension using the PCA algorithm while retaining variance information.

6. The material similarity calculation method based on high-dimensional features and multimodal features according to claim 1, characterized in that, S5 include: S51: Map entities in multimodal features to nodes in a knowledge graph; S52: Extract entity relationships from text descriptions and visual relationships from material images using a pre-trained model, and verify the consistency between entity relationships and visual relationships through a bidirectional attention mechanism; S53: Determine the node type and relation type, store the entity nodes, relation types and fusion vectors in the graph database, and configure vector indexes for entity nodes to accelerate retrieval.

7. The material similarity calculation method based on high-dimensional features and multimodal features according to claim 1, characterized in that, S6 include: S61: The similarity between the target material and the candidate material is calculated using the weighted Euclidean distance formula; S62: Cosine normalization of similarity, the formula is: Where q represents the fusion vector of the target material, and t represents the fusion vector of the candidate material. The Euclidean norm of the fusion vector of the target material is represented. The Euclidean norm of the fusion vector of candidate materials; S63: Retrieve the top k items with the highest similarity through the HNSW graph index, where k represents a positive integer set by the user.

8. The material similarity calculation method based on high-dimensional features and multimodal features according to claim 1, characterized in that, It also includes S7: adjusting the weights of each modal feature based on user click feedback on the search results.

9. The material similarity calculation method based on high-dimensional features and multimodal features according to claim 1, characterized in that, The materials include industrial parts, raw materials or finished products, the visual data includes material appearance images, CAD drawings or 3D point cloud data, and the structured data includes material codes, specifications, supplier information, prices and inventory status.

10. The material similarity calculation method based on high-dimensional features and multimodal features according to claim 1, characterized in that, The dynamic weight allocation strategy is as follows: the weight of the structured vector is 0.5-0.8, the weight of the text vector is 0.1-0.3, and the weight of the visual vector is 0.05-0.2; the sum of the weights of the structured vector, the text vector, and the visual vector is 1; the weights are dynamically adjusted according to the material type, wherein the weight of the structured vector for key production materials is increased to 0.7-0.8, and the weight of the visual vector for standardized materials is increased to 0.15-0.

2.

11. A material similarity calculation system based on high-dimensional features and multimodal features, employing the material similarity calculation method based on high-dimensional features and multimodal features as described in any one of claims 1 to 10, characterized in that the system... include: The multimodal feature extraction module is used to extract features from the structured data, text descriptions, and visual data of materials, and generate structured vectors, text vectors, and visual vectors, respectively. The cross-modal association module is used to establish associations between structured data and text descriptions and visual data through metadata association, entity name association, and time synchronization; The feature fusion module is used to perform weighted fusion of the structured vector, text vector and visual vector using a dynamic weight allocation strategy to generate a unified material feature vector; The knowledge graph module is used to construct a knowledge graph of material entities and relationships. The knowledge graph includes entity nodes, relationship types, and vector indexes. The similarity calculation and retrieval module is used to calculate the similarity between materials based on their feature vectors, and retrieves and outputs matching results through the HNSW graph index.

12. The material similarity calculation system based on high-dimensional features and multimodal features according to claim 11, characterized in that, The multimodal feature extraction module includes: The structured data processing unit normalizes material codes, specifications, supplier information, prices, and inventory status, and maps them into low-dimensional structured vectors through the TransE model. The text processing unit uses BERT or BGE-M3 models to encode the technical parameters, usage descriptions, and procurement contract terms of the materials, generating text vectors. The vision processing unit uses ResNet or CLIP models to extract features from material appearance images, CAD drawings, and 3D point cloud data, generating visual vectors.

13. The material similarity calculation system based on high-dimensional features and multimodal features according to claim 11, characterized in that, The cross-modal association module associates text descriptions with the same material entities in visual data through entity name matching, and performs time synchronization of multimodal features to ensure that the time reference of multimodal data is consistent.

14. The material similarity calculation system based on high-dimensional features and multimodal features according to claim 11, characterized in that, The knowledge graph module includes: The entity recognition unit identifies material entities in cross-modal data through similarity calculation; The relation extraction unit uses a pre-trained language model to extract entity relations from text, a Detectron2 model to extract visual relations from images, and a bidirectional attention mechanism to verify the rationality of the relations. The graph database storage unit is used to store entity nodes, relation types, and vector indexes. The entity nodes contain multimodal fusion vectors of materials.

15. The material similarity calculation system based on high-dimensional features and multimodal features according to claim 11, characterized in that, The similarity calculation and retrieval module calculates the similarity between materials based on an improved weighted Euclidean distance, using the following formula: Among them, This represents the weight of the i-th feature. This represents the i-th dimension feature of the target material. Let represent the i-th feature of the material to be matched, and n represent the dimension of the feature vector.

Citation Information

Patent Citations

  • Rapid comparative analysis method and device capable of defining BOM dimension and medium

    CN117974016A

  • Material similarity retrieval method and system

    CN119691151A

  • Patent retrieval method and system based on multi-modal attention map

    CN115617956A

  • Cross-modal knowledge graph construction method and device

    CN119443224A

  • Enterprise-level simulation knowledge graph construction method based on multi-modal data integration

    CN120336547A

Cited By

  • Aircraft cargo name identification method and device and electronic equipment

    CN121118898A

  • Investment project duplicate checking method and system, computer and storage medium

    CN121257510A

  • Contactor electrical life multi-stage prediction method based on clustering analysis and sequential network

    CN121279151A

  • Coal machine spare part similarity identification and coding unification system and method and storage medium

    CN121434738A

  • Aero-engine fault data processing method and system

    CN121636985A