Multi-modal intelligent industrial design recommendation method and system

Through multimodal data acquisition and feature fusion, combined with convolutional neural networks and natural language processing models, the problems of low efficiency and poor dynamic adaptability of traditional industrial design methods are solved, and more efficient and accurate recommendations of industrial design solutions are achieved.

CN120180376AInactive Publication Date: 2025-06-20BEIJING CENTURY STAR APPL TECH RES CENT
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510667929.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-06-20
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the prior art, industrial design methods rely on experience and intuition, are inefficient and are susceptible to subjective factors, and are difficult to respond to design parameter adjustments or environmental changes in real time, and have poor dynamic adaptability.

Method used

The multimodal intelligent industrial design recommendation method is adopted, and the visual and text data is extracted and fused through multimodal data acquisition, feature extraction and feature fusion, and the visual and text data is extracted and fused by convolutional neural networks and natural language processing models to generate a unified feature representation, thereby recommending industrial design solutions with high matching degree, and dynamically sorting based on structured and physical data.

Benefits of technology

It improves the matching degree between the design plan and actual production needs, enhances the dynamic adaptability of the system, can respond to design parameters and environmental changes in real time, and improves the efficiency and accuracy of industrial design.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120180376A_ABST
    Figure CN120180376A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent recommendation, and discloses a multi-modal intelligent industrial design recommendation method and system, and the system comprises a multi-modal data collection module, a multi-modal data feature extraction module, a multi-modal data feature fusion module, and an industrial design scheme recommendation module. Feature extraction is carried out through multi-modal data, feature extraction is carried out on visual modal data through a convolutional neural network, and feature extraction is carried out on text modal data through a natural language processing model; weight distribution is carried out on the visual modal data features and the text modal data features, feature fusion is carried out on the visual modal data features and the text modal data features, and unified feature representation is generated; corresponding feature extraction modes are adopted for different modal data, multi-source data such as vision, texts and sensors are fused, the problem of single information dimension is solved, and it is guaranteed that recommendation of design schemes meets actual production requirements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent recommendation, and more particularly to a multi-modal intelligent industrial design recommendation method and system. Background Art

[0002] With the development of technology, industrial design is more and more widely used in various industries. However, traditional industrial design methods often rely on designers' experience and intuition, with low efficiency and being easily affected by subjective factors. In the prior art, integrating artificial intelligence into industrial design well solves the problems existing in traditional industrial design methods.

[0003] However, there are still some problems in the prior art: 1. The information dimension is single, and multi-source data such as vision, text, and sensors cannot be integrated, resulting in a low matching degree between the recommended design scheme and the actual production requirements; 2. Traditional algorithms are difficult to respond in real time to the adjustment of design parameters or changes in the environment, and the dynamic adaptability is poor.

[0004] In view of this, the present invention provides a multi-modal intelligent industrial design recommendation method and system, which analyzes based on multi-modal data to improve the matching degree between the recommended design scheme and the actual production requirements, and uses advanced algorithms to improve the dynamic adaptability. Summary of the Invention

[0005] In order to overcome the above-mentioned defects of the prior art, the present invention provides a multi-modal intelligent industrial design recommendation method and system to solve the problems existing in the above background art.

[0006] The present invention provides the following technical solutions: A multi-modal intelligent industrial design recommendation method includes the following steps: Step S1: Multi-modal data collection: Collect multi-modal data of the target object, and achieve the temporal alignment of structured modal data and physical modal data through a timestamp synchronization algorithm; Step S2: Multi-modal data feature extraction: Extract features from the multi-modal data in Step S1, extract features from visual modal data through a convolutional neural network, and extract features from text modal data through a natural language processing model; Step S3: Multi-modal data feature fusion: Assign weights to the visual modal data features and text modal data features, and perform feature fusion on the visual modal data features and text modal data features to generate a unified feature representation; Step S4: Obtain the industrial design scheme set of the target object, obtain the matching degree of each industrial design recommendation scheme in the industrial design scheme set according to the features fused in Step S3, and obtain the objective function of each industrial design recommendation scheme based on the structured modal data and physical modal data, and output the top ones with higher rankings after sorting One industrial design solution is used as the final recommended industrial design solution.

[0007] Preferably, the target object is the product that needs industrial design; the multi-modal data includes visual modal data, text modal data, structured modal data, and physical modal data. The visual modal data includes the design drawings and 3D models of the target object; the text modal data includes the design requirement text and design specification text; the structured modal data is the production data of the target object; the physical modal data is the sensor data of the target object during the production process.

[0008] Preferably, the timestamp synchronization algorithm in step S1 includes: Step S11: Use a high-precision time reference source to unify the clocks of each acquisition device; the acquisition device is the device for acquiring structured modal data and physical modal data; Step S12: Ensure the timing consistency of data acquisition through redundant time synchronization paths; Step S13: Perform sliding window segmentation and interpolation alignment on asynchronous data.

[0009] Preferably, the sliding window segmentation in step S13 is expressed by the formula: ; where represents the data set within the th sliding window, represents the th data, represents the acquisition time of the th data, represents the total number of sliding windows, represents the data acquisition reference time period, represents the window tolerance threshold; The interpolation alignment uses cubic spline interpolation, and the formula is expressed as: ; where represents the data after interpolation alignment at time point , represents the nearest neighbor timestamp of the missing data point, , , , are all spline coefficients.

[0010] Preferably, the convolutional neural network in step S2 includes an input layer, a convolutional layer, a pooling layer, a loss function, and an output layer; The input layer is used to input visual modality data; the convolutional layer has 5 layers of depthwise separable convolutions; the pooling layer uses the max pooling method with a window set to 2×2 and a stride of 2; the output layer outputs the extracted visual modality data features; The loss function is designed as:

[0011] where, represents the loss function, represents the predicted feature vector of the th sample data, represents the true feature vector of the th sample data, represents the total amount of sample data, represents the L2 regularization coefficient, represents the trainable parameters; The formula for convolution in the convolutional layer is expressed as: ; where, represents the feature value output at the th position in the th channel, represents the weight from the input channel to the output channel at the position , represents the pixel value of the input image at the position in the input channel ; represents the bias of the th channel; represents the height of the convolutional kernel, represents the width of the convolutional kernel; represents the total number of channels.

[0012] Preferably, the feature extraction of the text modality data by the natural language processing model in step S2 specifically includes: Step S21: Generate word embeddings: Use a pre-trained model to extract the embedding vectors of each keyword in the text modality data, and the keywords are identified from the text using the self-attention mechanism in the BERT model; Step S22: Obtain the semantic weight of each keyword; Step S23: Perform weighted summation based on the semantic weights to obtain the features of the text modality data.

[0013] Preferably, the specific method for obtaining the semantic weight of each keyword in step S22 is: ; Among them, represents the semantic weight of the th keyword, represents the term frequency-inverse document frequency value of the th keyword, represents the position attenuation factor of the th keyword, , represents the term frequency-inverse document frequency value weight factor, represents the position weight factor, represents the total number of keywords, .

[0014] Preferably, the feature of the text modality data obtained by weighted summation based on the semantic weight in step S23 is expressed by the formula: ; where represents the feature of the text modality data, represents the embedding vector of the th keyword; The self-attention mechanism in step S21 is expressed by the formula: ; where is the self-attention mechanism representation, represents the query matrix, represents the key matrix, represents the value matrix, represents the scaling factor, represents the transpose of the key matrix, is the transpose representation of the matrix.

[0015] Preferably, the feature fusion of the visual modality data feature and the text modality data feature in step S3 is expressed by the formula: ; where represents the fused feature, represents the visual modality data feature, and are the corresponding scale factors respectively, satisfying , ; The specific method for obtaining the matching degree of each industrial design recommended solution in the industrial design solution set is: ; where represents the matching degree of the th industrial design solution in the industrial design solution set, represents the fused feature, represents the feature of the th industrial design solution in the industrial design solution set, It is a norm representation.

[0016] A multi-modal intelligent industrial design recommendation system, including a multi-modal data acquisition module, a multi-modal data feature extraction module, a multi-modal data feature fusion module, and an industrial design scheme recommendation module; The multi-modal data acquisition module is used to collect multi-modal data of the target object, and realize the timing alignment of structured modal data and physical modal data through a timestamp synchronization algorithm; at the same time, preprocess the collected multi-modal data; The multi-modal data feature extraction module is used to extract features from the multi-modal data in the multi-modal data acquisition module, extract features from visual modal data through a convolutional neural network, and extract features from text modal data through a natural language processing model; The multi-modal data feature fusion module is used to assign weights to visual modal data features and text modal data features, fuse visual modal data features and text modal data features, and generate a unified feature representation; The industrial design scheme recommendation module is used to obtain the industrial design scheme set of the target object, obtain the matching degree of each industrial design recommendation scheme in the industrial design scheme set according to the features fused by the multi-modal data feature fusion module, and obtain the objective function of each industrial design recommendation scheme based on the structured modal data and the physical modal data, and output the top several industrial design schemes as the final industrial design recommendation schemes after sorting.

[0017] The technical effects and advantages of the present invention: By setting step S2 and step S3, the present invention is beneficial to extract features from multi-modal data, extract features from visual modal data through a convolutional neural network, and extract features from text modal data through a natural language processing model; assign weights to visual modal data features and text modal data features, fuse visual modal data features and text modal data features, and generate a unified feature representation; adopt corresponding feature extraction methods for different modal data, fuse multi-source data such as vision, text, and sensors, solve the problem of single information dimension, ensure that the recommendation of the design scheme meets the actual production requirements, and utilize the self-attention mechanism and the neural network model to respond in real time to the adjustment of design parameters or changes in the environment, improve the dynamic adaptability, and at the same time improve the matching degree between the recommendation of the design scheme and the actual production requirements. Description of the Drawings

[0018] Figure 1 It is a flow chart of the multi-modal intelligent industrial design recommendation method of the present invention.

[0019] Figure 2This is the structural diagram of the multi-modal intelligent industrial design recommendation system of the present invention. Detailed implementation manners

[0020] Next, the technical solutions in the present invention will be clearly and completely described in conjunction with the accompanying drawings in the present invention. In addition, the forms of the various structures described in the following embodiments are merely examples, and a multi-modal intelligent industrial design recommendation method and system according to the present invention are not limited to the various structures described in the following embodiments. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0021] As Figure 1 shown, the present invention provides a multi-modal intelligent industrial design recommendation method, including the following steps: Step S1: Multi-modal data collection: Collect multi-modal data of the target object, and achieve the time-series alignment of the structured modal data and the physical modal data through the timestamp synchronization algorithm; at the same time, preprocess the collected multi-modal data; the target object is the product that needs to be industrially designed; the multi-modal data includes visual modal data, text modal data, structured modal data, and physical modal data. The visual modal data includes, but is not limited to, the design drawings and 3D models of the target object; the text modal data includes, but is not limited to, the design requirement text and the design specification text. The design requirement text is the requirement content of the user when industrially designing the target object, and the design specification text is the design specification when industrially designing the target object; the structured modal data is the production data of the target object, that is, the production efficiency and production cost of the target object during the production process, etc.; the physical modal data is the sensor data of the target object during the production process, that is, the production line temperature and pressure during the production process, etc.; both the structured modal data and the physical modal data are historical data of the target object and can be obtained from the database through big data technology; the preprocessing includes, but is not limited to, cleaning, denoising, and removing outliers from the data, etc. Step S2: Multi-modal data feature extraction: Extract features from the multi-modal data in Step S1, extract features from the visual modal data through a convolutional neural network, and extract features from the text modal data through a natural language processing model. Step S3: Multi-modal data feature fusion: Assign weights to the visual modal data features and the text modal data features, and perform feature fusion on the visual modal data features and the text modal data features to generate a unified feature representation. Step S4: Obtain the industrial design solution set of the target object. According to the features after fusion in Step S3, obtain the matching degrees of each industrial design recommended solution in the industrial design solution set, and obtain the objective function of each industrial design recommended solution based on the structured modal data and physical modal data. After sorting, output the top industrial design solutions as the final industrial design recommended solutions; you can select them by yourself from the recommended industrial design solutions. The value of can be set by yourself. If three are needed to be recommended, then . In this embodiment, no specific limitation is imposed on this specific value.

[0022] In this embodiment, it should be specifically noted that the timestamp synchronization algorithm in Step S1 includes: Step S11: Use a high-precision time reference source to unify the clocks of each acquisition device; the acquisition device is a device for acquiring structured modal data and physical modal data, including but not limited to various sensors; Step S12: Ensure the timing consistency of data acquisition through redundant time synchronization paths; Step S13: Perform sliding window segmentation and interpolation alignment on asynchronous data.

[0023] In this embodiment, it should be specifically noted that the sliding window segmentation in Step S13 is represented by the formula: ; where represents the data set within the th sliding window, represents the rd data, represents the acquisition time of the th data, represents the total number of sliding windows, represents the data acquisition reference time period. In this embodiment, is selected; represents the window tolerance threshold for compensating the clock deviation of the acquisition device. In this embodiment, is selected; The interpolation alignment uses cubic spline interpolation to compensate for the data at missing time points, and the formula is represented as:

[0024] where represents the data after interpolation alignment at time point , represents the nearest neighbor timestamp of the missing data point. During the interpolation process, the timestamps near the missing data point are needed to determine the position of the interpolation point; , , , are all spline coefficients, which describe the shape and trend of the interpolation function. These coefficients are obtained by solving the least squares method using adjacent data points.

[0025] In this embodiment, it should be specifically noted that the convolutional neural network in step S2 includes an input layer, a convolutional layer, a pooling layer, a loss function, and an output layer; The input layer is used to input visual modality data; the convolutional layer has 5 layers of depthwise separable convolutions, and the number of filters in each layer is 32, 64, 128, 256, and 512 respectively; the pooling layer uses the max pooling method, with a window set to 2×2 and a stride of 2; the output layer outputs the extracted visual modality data features; automatically extracts key features such as geometric shapes, topological relationships, and dimensional ratios from industrial design drawings to ensure that the feature vectors can accurately represent the geometric attributes of the target object, align with the manually labeled features, and provide structured data support for subsequent recommendations; The loss function is designed as: ; where represents the loss function, represents the predicted feature vector of the th sample data, represents the true feature vector of the th sample data, represents the total amount of sample data, represents the L2 regularization coefficient, , which is used to penalize values with overly large model weights to prevent overfitting; represents the trainable parameters, which are used to control the model complexity, can be any one or more of weights or biases; the sample data is the data used to train the convolutional neural network. All products undergoing industrial design are used as samples, and the visual modality data of the samples is used as the sample data of the convolutional neural network. The true feature vector of the sample data is the true feature of the sample, which is obtained by exporting it from professional design software and then annotated by professionals in the field; The formula for convolution in the convolutional layer is expressed as: ; where represents the feature value output at the th position in the th channel, represents the weight from the input channel to the output channel at the th position, Indicates the position of the input image In the input channel The pixel value of; Indicates the Bias of the th channel; Indicates the height of the convolutional kernel, Indicates the width of the convolutional kernel; Indicates the total number of channels. Exemplarily, if it is an RGB image, then .

[0026] In this embodiment, it should be specifically noted that the purpose of extracting features from the text modal data through the natural language processing model in step S2 is to convert the user requirement text into a structured semantic vector and capture implicit design constraints; specifically including: Step S21: Generate word embeddings: Use a pre-trained model to extract the embedding vectors of each keyword in the text modal data, and the keywords are identified using the self-attention mechanism in the BERT model; the pre-trained model selects the BERT model, and the BERT model is a prior art, and this embodiment will not elaborate on it too much; Step S22: Obtain the semantic weight of each keyword; Step S23: Perform weighted summation based on the semantic weights to obtain the features of the text modal data.

[0027] In this embodiment, it should be specifically noted that the specific method for obtaining the semantic weight of each keyword in step S22 is: ; Wherein, Indicates the Semantic weight of the th keyword, Indicates the th term frequency-inverse document frequency value of the keyword, which is used to quantify the importance of the word in the text, Indicates the th position decay factor of the keyword, , which is used to capture the influence of word order on semantic priority, and the first word , and the position is reduced by 0.02 for each subsequent shift, Indicates the term frequency-inverse document frequency value weight factor, which is used to amplify the influence of high-discrimination keywords. In this embodiment, is selected, Indicates the position weight factor, which is used to enhance the weight of the first word in the sentence. In this embodiment, is selected; Indicates the total number of keywords, ; The feature of the text modal data obtained by performing weighted summation based on the semantic weights in step S23 is expressed by the formula: ; Among them, represents the feature of the text modality data, represents the th keyword embedding vector.

[0028] In this embodiment, it should be specifically noted that the self-attention mechanism in step S21 is expressed by the formula: ; Among them, is the self-attention mechanism representation, represents the query matrix, characterizing the semantic focus that needs to be concerned currently, represents the key matrix, providing the semantic keys to be matched, and calculating the similarity with ; represents the value matrix, storing the actual semantic information, and generating a context vector after weighted summation, represents the scaling factor, used to prevent gradient explosion, represents the transpose of the key matrix, is the transpose representation of the matrix.

[0029] In this embodiment, it should be specifically noted that the feature fusion of the visual modality data feature and the text modality data feature in step S3 is expressed by the formula: ; Among them, represents the fused feature, represents the visual modality data feature, and are the corresponding proportionality factors respectively, satisfying , ; In this embodiment, no specific limitations are imposed on the specific values of and , which can be specifically set by those skilled in the art according to the actual situation. If the design of the target object pays more attention to the presentation of objective drawings and the content of the customer's needs is the focus, then can be set. If the design of the target object pays more attention to the content of the customer's needs and the content of the customer's drawings is the focus for reference, then can be set. If both are equally important, then can be set.

[0030] In this embodiment, it should be specifically noted that the acquisition method of the industrial design scheme set of the target object is: retrieving the historical industrial design schemes of the target object from the database, and at the same time obtaining all the new achievable industrial design schemes based on the modification by those skilled in the art, and summarizing all the industrial design schemes to obtain the industrial design scheme set of the target object; The specific method for obtaining the matching degree of each industrial design recommended solution in the industrial design solution set is as follows: ; where represents the matching degree of the -th industrial design solution in the industrial design solution set, represents the fused features, represents the features of the -th industrial design solution in the industrial design solution set, is a norm representation used to calculate the norm of a vector; The larger the value of , the higher the matching degree, indicating that the -th industrial design solution in the industrial design solution set has a higher similarity to the fused features of the target object, that is, using this design solution can meet the industrial design requirements of the target object; , representing the total number of industrial design solutions in the industrial design solution set.

[0031] In this embodiment, it should be specifically noted that the objective function of each industrial design recommended solution can be an objective function such as a cost constraint function or a performance constraint function, and the corresponding objective function can be set according to the actual needs of those skilled in the art. The cost constraint function is a function that limits and constrains the cost value, and the performance constraint function is a function that limits and constrains the product performance; in this embodiment, the cost constraint function is selected as the objective function; the formula is expressed as: ; where represents the objective function, represents the cost limit value; select the industrial design solutions that satisfy the objective function and sort them; the cost limit value is the maximum acceptable cost of the target object, which can be set by those skilled in the art according to the product situation of the target object. This embodiment does not specifically limit this specific value. Different target objects have different costs and different requirements for costs; Combine the matching degree with the objective function to form the final recommended score. The specific method is as follows: Sort the industrial design solutions according to the matching degree, from the highest matching degree to the lowest matching degree, to form the first list; sort the industrial design solutions according to the production cost, from the lowest cost to the highest cost, to form the second list; the method for obtaining the production cost of the industrial design solution is: obtain the structured modal data and physical modal data corresponding to each industrial design solution, so as to obtain the production cost of the industrial design solution; since only the industrial design solutions that satisfy the objective function are sorted, the industrial design solutions that do not satisfy the objective function are simultaneously excluded from the sorting of the matching degree to ensure the consistency of the industrial design solutions in the first list and the second list; If there are An industrial design scheme, the arrangement serial number of the industrial design scheme in the sequence list and the sum of its corresponding recommended scores are , for example, if there are 5 industrial design schemes in total, the recommended score of the industrial design scheme ranked first in the first sequence list satisfies , that is, the recommended score of the industrial design scheme ranked first is 5 points, and so on, the recommended score of the industrial design scheme ranked second is 4 points, the recommended score of the industrial design scheme ranked third is 3 points, the recommended score of the industrial design scheme ranked fourth is 2 points, and the recommended score of the industrial design scheme ranked fifth is 1 point; the calculation method of the recommended scores of the second sequence list and the first sequence list is the same; Combining the recommended scores of the first sequence list and the second sequence list, the final recommended score is expressed as: ; Among them, represents the final recommended score of the th industrial design scheme, represents the recommended score of the th industrial design scheme in the first sequence list, represents the recommended score of the th industrial design scheme in the second sequence list; and are the corresponding weight coefficients respectively, and satisfy , and The specific values of can be set by those skilled in the art according to the actual situation. In this embodiment, is selected, ; ; Sort the industrial design schemes according to the final recommended scores, sort from the highest recommended score to the lowest recommended score, and output the top industrial design schemes as the final industrial design recommended schemes; .

[0032] As Figure 2 shown, the present invention provides a multi-modal intelligent industrial design recommendation system, including a multi-modal data acquisition module, a multi-modal data feature extraction module, a multi-modal data feature fusion module, and an industrial design scheme recommendation module; The multi-modal data acquisition module is used to collect multi-modal data of the target object, and realize the temporal alignment of structured modal data and physical modal data through a timestamp synchronization algorithm; at the same time, preprocess the collected multi-modal data; The multi-modal data feature extraction module is used to extract features from the multi-modal data in the multi-modal data acquisition module, extract features from the visual modal data through a convolutional neural network, and extract features from the text modal data through a natural language processing model; The multi-modal data feature fusion module is used to assign weights to the visual modal data features and the text modal data features, fuse the visual modal data features and the text modal data features, and generate a unified feature representation; The industrial design solution recommendation module is used to obtain the industrial design solution set of the target object, obtain the matching degree of each industrial design recommendation solution in the industrial design solution set according to the features fused by the multi-modal data feature fusion module, and obtain the objective function of each industrial design recommendation solution based on the structured modal data and the physical modal data, and output the top several industrial design solutions as the final industrial design recommendation solutions.

[0033] Finally: The above are only the preferred embodiments of the present invention and are not used to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

[0034] The above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in this application, and all should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

Claims

1. A multimodal intelligent industrial design recommendation method, characterized in that: It includes the following steps: Step S1: Multimodal data acquisition: Acquire the multimodal data of the target object, and achieve the temporal alignment of the structured modal data and the physical modal data through the timestamp synchronization algorithm; Step S2: Multimodal data feature extraction: Extract features from the multimodal data in Step S1, extract features from the visual modal data through a convolutional neural network, and extract features from the text modal data through a natural language processing model; Step S3: Multimodal data feature fusion: Assign weights to the visual modal data features and the text modal data features, and perform feature fusion on the visual modal data features and the text modal data features to generate a unified feature representation; Step S4: Obtain the industrial design solution set of the target object. According to the features after fusion in step S3, obtain the matching degree of each industrial design recommended solution in the industrial design solution set, and obtain the objective function of each industrial design recommended solution based on the structured modal data and physical modal data. After sorting, output the top industrial design solutions with the highest ranking as the final industrial design recommended solutions.

2. The multimodal intelligent industrial design recommendation method according to claim 1, characterized in that: The target object is the product that needs to be industrially designed; the multimodal data includes visual modal data, text modal data, structured modal data, and physical modal data. The visual modal data includes the design drawings and 3D models of the target object; the text modal data includes design requirement texts and design specification texts; the structured modal data is the production data of the target object; the physical modal data is the sensor data of the target object during the production process.

3. The multimodal intelligent industrial design recommendation method according to claim 2, characterized in that: The timestamp synchronization algorithm in Step S1 includes: Step S11: Use a high-precision time reference source to unify the clocks of each acquisition device; the acquisition device is the device for acquiring structured modal data and physical modal data; Step S12: Ensure the temporal consistency of data acquisition through redundant time synchronization paths; Step S13: Perform sliding window segmentation and interpolation alignment on asynchronous data.

4. The multimodal intelligent industrial design recommendation method according to claim 3, characterized in that: The sliding window segmentation in Step S13 is represented by the formula: ; Among them, represents the data set within the th sliding window, represents the th data, represents the acquisition time of the th data, represents the total number of sliding windows, represents the data acquisition reference time period, represents the window tolerance threshold; The interpolation alignment uses cubic spline interpolation, and the formula is represented as: ; Among them, represents a time point the data after interpolation alignment, represents the nearest neighbor timestamp of the missing data point, and and and are all spline coefficients.

5. The multimodal intelligent industrial design recommendation method according to claim 4, characterized in that: The convolutional neural network in Step S2 includes an input layer, a convolutional layer, a pooling layer, a loss function, and an output layer; The input layer is used to input visual modal data; the convolutional layer has 5 layers of depthwise separable convolution; the pooling layer uses the maximum pooling method, the window is set to 2×2, and the stride is 2; the output layer outputs the extracted visual modal data features; The loss function is designed as: ; Among them, represents the loss function, represents the predicted feature vector of the th sample data, represents the true feature vector of the th sample data, represents the total amount of sample data, represents the L2 regularization coefficient, represents the trainable parameter; The formula for convolution in the convolutional layer is represented as: ; Among them, represents the eigenvalue output by the th channel at position ; represents the weight from the input channel to the output channel at position ; represents the pixel value of the input image at position in the input channel ; represents the bias of the th channel; represents the height of the convolution kernel, represents the width of the convolution kernel; represents the total number of channels.

6. The multimodal intelligent industrial design recommendation method according to claim 5, characterized in that: The specific process of extracting features from the text modal data through a natural language processing model in Step S2 includes: Step S21: Generate word embeddings: Use a pre-trained model to extract the embedding vectors of each keyword in the text modal data. The keyword uses the self-attention mechanism in the BERT model to identify the text; Step S22: Obtain the semantic weight of each keyword; Step S23: Perform weighted summation based on the semantic weight to obtain the features of the text modal data.

7. The multimodal intelligent industrial design recommendation method according to claim 6, characterized in that: The specific method for obtaining the semantic weight of each keyword in Step S22 is: ; Among them, represents the semantic weight of the th keyword, represents the th keyword's term frequency - inverse document frequency value, represents the th keyword's position attenuation factor, , represents the term frequency - inverse document frequency value weight factor, represents the position weight factor, represents the total number of keywords, .

8. A multimodal intelligent industrial design recommendation method according to claim 7, characterized in that: The formula for performing weighted summation based on the semantic weight to obtain the features of the text modal data in Step S23 is represented as: ; wherein, represents the feature of the text modal data, represents the embedding vector of the The self-attention mechanism in step S21 is expressed by the formula: ; Among them, is the self-attention mechanism representation, represents the query matrix, represents the key matrix, represents the value matrix, represents the scaling factor, represents the transpose of the key matrix, is the transpose representation of the matrix.

9. A multimodal intelligent industrial design recommendation method according to claim 8, characterized in that: The formula for feature fusion of the visual modal data features and the text modal data features in Step S3 is represented as: ; wherein, represents the fused feature, represents the visual modality data feature, and are the corresponding scale factors respectively, satisfying , ; The specific method for obtaining the matching degree of each industrial design recommendation plan in the industrial design plan set is: ; wherein, represents the matching degree of the th industrial design solution in the industrial design solution set, represents the fused features, represents the features of the th industrial design solution in the industrial design solution set, is represented by a norm.

10. A multimodal intelligent industrial design recommendation system for using a multimodal intelligent industrial design recommendation method according to any one of claims 1-9, characterized in that: It includes a multi-modal data acquisition module, a multi-modal data feature extraction module, a multi-modal data feature fusion module, and an industrial design solution recommendation module; The multi-modal data acquisition module is used to collect multi-modal data of a target object, and achieve temporal alignment of structured modal data and physical modal data through a timestamp synchronization algorithm; meanwhile, preprocess the collected multi-modal data; The multi-modal data feature extraction module is used to extract features from the multi-modal data in the multi-modal data acquisition module, extract features from visual modal data through a convolutional neural network, and extract features from text modal data through a natural language processing model; The multi-modal data feature fusion module is used to assign weights to visual modal data features and text modal data features, fuse visual modal data features and text modal data features, and generate a unified feature representation; The industrial design solution recommendation module is used to obtain the industrial design solution set of the target object, obtain the matching degree of each industrial design recommendation solution in the industrial design solution set according to the features fused by the multi-modal data feature fusion module, and obtain the objective function of each industrial design recommendation solution based on the structured modal data and the physical modal data. After sorting, the top industrial design solutions are output as the final industrial design recommendation solutions.

Citation Information

Patent Citations

  • Hard bus prefabrication production line control method and system based on BIM model

    CN118584925A

  • Large-scale engineering equipment design intention extraction method based on multi-modal data fusion

    CN119068302A

  • Cross-modal feature alignment and fusion method based on comparative learning

    CN119272024A

  • Performance optimization method and device of intelligent manufacturing chip based on historical data

    CN119761771A

  • An intelligent recommendation system and method for precision shaft design

    CN119782625A