Method for solving supply chain information island problem based on multi-mode Ai
Through multimodal AI technology, multiple data formats are collected and integrated in all links of the supply chain, the problems of poor data circulation, incompatible formats and incomplete decision-making are solved, rapid response and accurate decision-making of the supply chain are achieved, and the overall efficiency of the supply chain is improved.
Patent Information
- Application Number
- CN202510366839.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-08
AI Technical Summary
There are problems in the existing supply chain management with poor data circulation, incompatible information formats, lack of comprehensiveness in decision-making and slow response, resulting in increased operating costs of enterprises, disconnected production plans and missing market opportunities.
Multimodal AI technology is adopted to collect text, images, speech and sensor data through the acquisition equipment at various links of the supply chain, use word vector models, CNN models, MFCC feature extraction and neural networks to perform feature extraction and encoding, and information fusion is carried out based on attention mechanisms, and finally intelligent decision-making is made in the deep neural network of Transformer architecture.
It realizes the barrier-free circulation of data, unified representation of information formats and rapid decision-making, improves the response speed of the supply chain and the accuracy of decision-making, and can quickly adapt to market changes and optimize production and logistics operations.
Smart Images

Figure CN120277611A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the cross - technical field of AI and supply chain management technology, and specifically to a method for solving the problem of information silos in the supply chain based on multimodal AI. Background Art
[0002] In today's digital age, supply chain management is crucial for the operation and development of enterprises. However, the existing methods for solving the problem of information silos in the supply chain have many defects and deficiencies, severely restricting the improvement of the overall efficiency of the supply chain.
[0003] 1. Poor data circulation: The supply chain covers multiple links such as procurement, production, logistics, and sales, and each link often uses independent information systems. The procurement department relies on the supplier management system to record contracts and purchase orders; the production link uses the production execution system to monitor production progress; logistics uses the transportation management system to track the transportation of goods; and sales processes orders and customer feedback through the customer relationship management system. There is a lack of effective integration between these systems, resulting in difficulty in real - time data sharing. For example, information on logistics transportation delays cannot be promptly transmitted to the production department, causing the production plan to be out of sync with the actual transportation progress, which may then lead to production stagnation or inventory backlog, increasing the operating costs of the enterprise.
[0004] 2. Incompatible information formats: The data formats generated in each link of the supply chain are rich and diverse. Procurement contracts are mostly text files; the output of production equipment sensors is structured data; the images of transportation vehicles in the logistics link are in multimedia format; and customer feedback at the sales end may be text, voice, or image. Integrating and analyzing data in different formats poses great challenges. Traditional methods require a large amount of manpower for format conversion and pre - processing, and data loss or errors are extremely likely to occur during the process, affecting the accuracy and availability of the data.
[0005] 3. Lack of comprehensiveness in decision - making: Since information is scattered in various independent systems, it is difficult for decision - makers to obtain comprehensive and accurate information. Procurement decisions may only be based on historical procurement data and supplier quotes, without fully considering changes in production requirements, fluctuations in logistics costs, and customer feedback. For example, customers at the sales end put forward new requirements for product quality, but due to the ineffective transmission of information, the procurement department still purchases raw materials according to the original standards, ultimately affecting product quality and customer satisfaction and damaging the market competitiveness of the enterprise.
[0006] 4. Slow response speed: In the face of market changes or emergencies, the delay in information transmission and processing makes the supply chain respond slowly. For example, when market demand suddenly increases, the sales department cannot promptly transmit information to the production and procurement departments, or the production department cannot adjust the production plan in a timely manner because it has not obtained information on blocked logistics transportation in time, resulting in the supply chain's inability to quickly adapt to changes, missing market opportunities or causing customer loss, bringing economic losses to the enterprise.
[0007] Based on this, a method for solving the problem of information silos in the supply chain based on multimodal AI is now provided, which can eliminate the drawbacks existing in the existing devices. Summary of the Invention
[0008] The purpose of the present invention is to provide a method for solving the problem of information silos in the supply chain based on multimodal AI, so as to solve the problems of the shortcomings of the modern product in the background technology.
[0009] To achieve the above object, the present invention provides the following technical solutions:
[0010] A method for solving the problem of information silos in the supply chain based on multimodal AI, comprising the following steps:
[0011] Step 1: Supply chain multimodal information collection and integration:
[0012] Using the collection devices distributed in each link of the supply chain procurement, production, logistics, and sales, collect multimodal information such as text, images, voices, and sensor data; and classify and store the collected multimodal information according to the source and modality;
[0013] Step 2: Feature extraction and encoding:
[0014] For text information, use the word vector model to map the words in the text into word vectors, and then encode them through the RNN network to obtain the feature vectors representing the text information;
[0015] For image information, use the pre-trained CNN model for feature extraction, and through convolution, pooling, and fully connected layer operations, obtain the feature vectors of a fixed length;
[0016] For voice information, use a specific library to extract the MFCC features of the voice record, and then use the neural network for further encoding to obtain the feature vectors representing the voice information;
[0017] For sensor data information: first normalize the sensor data, and then perform feature extraction and encoding through a custom neural network layer to obtain the feature vectors representing the sensor data;
[0018] Step 3: Supply chain island information fusion:
[0019] Based on the attention mechanism, calculate the cosine similarity between different modality feature vectors, use the Softmax function to convert the correlation into attention weights, and perform weighted summation on different modality feature vectors to achieve information fusion;
[0020] Step 4: Intelligent decision-making and collaborative application:
[0021] The fused feature vectors are input into a pre-trained deep neural network model based on the Transformer architecture for intelligent decision-making. According to the decision results output by the model, corresponding collaborative operations are automatically triggered.
[0022] Based on the above technical solutions, the present invention also provides the following alternative technical solutions:
[0023] In an alternative solution: the word vector model in step two is a Word2Vec model;
[0024] The pre-trained CNN model in step two is a VGG16 model under the TensorFlow framework;
[0025] The specific library in step two is the Librosa library. The MFCC features of the voice recording are extracted using the Librosa library. Let the voice signal be s(t), and the signal of the nth frame after frame processing is Sn(m), m = 0, 1,..., M - 1, where M is the frame length. After windowing, we get Calculate the power spectrum After filtering through the Mel filter bank, we get H is the number of Mel filters. Finally, after discrete cosine transform, the MFCC coefficients are obtained L is the number of MFCC coefficients. After obtaining the MFCC features, a neural network is built using PyTorch for further encoding, and finally, feature vectors representing voice information are obtained;
[0026] In step two, the MinMaxScaler is used to normalize the sensor data, and the fully connected layer is used as a custom neural network layer for feature extraction and encoding.
[0027] In an alternative solution: in step one, in the procurement link, text contracts provided by suppliers, product sample images, and voice records of the communication between both parties are collected;
[0028] In the production link, sensor data of equipment operation, production process monitoring images, and voice reports of operators on the production situation are collected;
[0029] In the logistics link, position and status sensor data of transport vehicles, road condition images, and voice communications between drivers and the dispatching center are collected;
[0030] In the sales link, text evaluations feedback by customers, product usage scenario images, and voice communication records between customer service and customers are collected.
[0031] In an alternative solution: in step four, when training the deep neural network model based on the Transformer architecture, the cross-entropy loss function is adopted, and the model parameters are adjusted through an optimization algorithm. The optimization algorithm is the stochastic gradient descent method.
[0032] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0033] 1. Through the multi-modal AI technology, the present invention collects comprehensive multi-modal information, including text, images, voice, and sensor data, etc., through acquisition devices deployed in various links of the supply chain. It penetrates into each link to capture various types of information in real time, and then integrates the collected information into a unified platform, breaking the isolated state of data in each link and realizing unobstructed data circulation.
[0034] 2. By using word vector models (such as Word2Vec) and RNN networks (such as LSTM), the present invention converts text into feature vectors that retain semantic and context information, uses pre-trained CNN models (such as VGG16) to extract key features of images, adopts MFCC feature extraction combined with neural network encoding to process voice, and performs normalization and neural network processing on sensor data to generate feature vectors, converting information in various formats into a unified feature representation, solving the problem of incompatible information formats and facilitating subsequent fusion and analysis.
[0035] 3. Through the multi-modal information fusion module of multi-modal AI based on the attention mechanism, the present invention fuses different modal feature vectors, measures the correlation between feature vectors by calculating the cosine similarity, and uses the Softmax function to convert it into attention weights to highlight key information. When making procurement decisions, it fuses multi-modal information such as procurement contract texts, supplier product images, communication voice records, as well as production requirements, logistics costs, customer feedback, etc., enabling decision-makers to comprehensively understand the situation and make more accurate decisions.
[0036] 4. Through the multi-modal AI technology, the present invention processes and analyzes multi-source information in real time, and quickly makes decisions with the help of a deep neural network model based on the Transformer architecture. Once changes occur in market demand, logistics transportation, or production links, the model can quickly analyze and fuse the information, automatically triggering corresponding collaborative operations. When market demand increases, the model will quickly adjust the production plan, arrange raw material procurement, and optimize logistics distribution according to multi-modal information of sales, production, procurement, and logistics, improving the response speed of the supply chain and better adapting to market changes. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 It is a schematic diagram of the step flow of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0038] In order to make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments.
[0039] In one embodiment, as Figure 1As shown, a method for solving the problem of supply chain information islands based on multimodal Ai includes the following steps:
[0040] Step 1: Supply chain multimodal information collection and integration:
[0041] Utilize collection equipment distributed in various links of supply chain procurement, production, logistics, and sales to collect multimodal information such as text, images, voice, sensor data, etc.; and classify and store the collected multimodal information according to source and modality.
[0042] Step 2: Feature extraction and encoding:
[0043] For text information, the word vector model is used to map the words in the text into word vectors, which are then encoded through the RNN network to obtain feature vectors representing the text information;
[0044] For image information, a pre-trained CNN model is used for feature extraction. After convolution, pooling and full connection layer operations, a fixed-length feature vector is obtained.
[0045] For speech information, the MFCC features of the speech recording are extracted with the help of a specific library, and then further encoded using a neural network to obtain a feature vector representing the speech information;
[0046] For sensor data information: the sensor data is first normalized, and then feature extraction and encoding are performed through a custom neural network layer to obtain a feature vector representing the sensor data.
[0047] Step 3: Supply chain island information integration:
[0048] Based on the attention mechanism, the cosine similarity between feature vectors of different modalities is calculated, and the Softmax function is used to convert the correlation into attention weight. The weighted summation of feature vectors of different modalities is performed to achieve information fusion.
[0049] Step 4: Intelligent decision-making and collaborative application:
[0050] The fused feature vector is input into a pre-trained deep neural network model based on the Transformer architecture for intelligent decision-making, and the corresponding collaborative operations are automatically triggered based on the decision results output by the model.
[0051] The above embodiment discloses a method for solving the problem of supply chain information islands based on multimodal Ai, and its specific working process is as follows:
[0052] (1) Supply chain multimodal information collection and integration steps:
[0053] Collect multimodal information using collection devices distributed across various links in the supply chain. In the procurement link, collect the text contracts provided by suppliers, product sample images, and voice records of the communication between both parties; in the production link, collect sensor data of equipment operation, production process monitoring images, and voice reports of operators on production conditions; in the logistics link, collect position and status sensor data of transport vehicles, road condition images, and voice communications between drivers and the dispatching center; in the sales link, collect text evaluations from customers, product usage scenario images, and voice communication records between customer service and customers, etc.
[0054] Store the collected information in the corresponding database tables according to the links such as procurement, production, logistics, and sales, as well as modalities such as text, images, voice, and sensor data, and establish indexes for quick query and invocation.
[0055] Preliminarily sort out the collected multimodal information, classify and store it according to the source and modality, and prepare for subsequent processing.
[0056] (2) Feature extraction and encoding steps:
[0057] Text information processing: For text information such as text contracts and customer feedback, use a word vector model (such as Word2Vec) to map the words in the text to word vectors. Assume the text sequence is S = [w1, w2,..., wn], where wi represents the i-th word in the text, and n is the total number of words in the text. After being processed by the Word2Vec model, each word wi is mapped to a word vector Then encode it through an RNN network (taking LSTM as an example). Finally, obtain a feature vector that can represent the entire text information.
[0058] Use the NLTK library in Python to preprocess the text information, including operations such as word segmentation, stop word removal, and stemming, and then combine it with a pre-trained Word2Vec model to generate word vectors, and encode them through an LSTM network built by Keras. For example, for the procurement contract text, after preprocessing, obtain the vector representation of each word through the Word2Vec model, and then process it through the LSTM network to generate a feature vector representing the key information of the contract.
[0059] (3) Image information processing: For product sample images, production process monitoring images, etc., use a pre-trained CNN model (such as VGG16) under the TensorFlow framework to extract features. Let the input image be X, after convolutional layer operations, the convolutional kernel is W, and the convolution operation formula is where Y is the convolution output result, M and N are the sizes of the convolutional kernel, and b is the bias. After a series of convolution and pooling operations, flatten the feature map through the fully connected layer and further process it to obtain a feature vector F with a fixed length.
[0060] Use the VGG16 model under the TensorFlow framework to extract features from images. Process the input images according to the convolution operation formula. After convolution, pooling, and fully connected layer operations, obtain the feature vectors of the images. For example, for production process monitoring images, after being processed by the VGG16 model, feature vectors reflecting the production status are obtained.
[0061] (4) Speech information processing: Extract MFCC features of speech records with the help of the Librosa library. Let the speech signal be s(t), and the nth frame signal after frame processing is Sn(m), where m = 0, 1,..., M - 1, M is the frame length. After windowing, obtain Calculate the power spectrum After filtering through the Mel filter bank, obtain H is the number of Mel filters. Finally, obtain the MFCC coefficients through discrete cosine transform L is the number of MFCC coefficients. After obtaining the MFCC features, use PyTorch to build a neural network (take MLP as an example) for further encoding, and finally obtain the feature vectors that can represent the speech information.
[0062] Process the speech with the help of the Librosa library according to the MFCC feature extraction formula, and use the MLP network built by PyTorch to further encode the MFCC features. For example, for the voice communication between the driver and the dispatching center, after MFCC feature extraction and MLP network encoding, vectors representing the key features of the speech content are obtained.
[0063] (5) Sensor data processing: For sensor data such as equipment operation and vehicle position, first use MinMaxScaler for normalization processing to map the data to the [0, 1] interval. Then, perform feature extraction and encoding through a custom neural network layer (take the fully connected layer as an example). Let the weight matrix of the fully connected layer, then output the feature vector.
[0064] Use MinMaxScaler in the Scikit-learn library to normalize the sensor data, and then perform feature extraction and encoding through a custom fully connected layer neural network. For example, for sensor data such as the temperature and pressure of equipment operation, after normalization and neural network processing, feature vectors that can represent the equipment operation status are obtained
[0065] (6) Steps for information fusion of supply chain islands:
[0066] Based on the attention mechanism, the feature vectors of different modalities are fused. First, calculate the cosine similarity between the feature vectors of different modalities. For the text feature vector and the image feature vector, this is used to measure the correlation between them.
[0067] Then, use the Softmax function to transform the correlation into attention weights. Set the correlation vector between the feature vectors of different modalities. After being processed by the Softmax function, an attention weight vector is obtained, where the value of each element is between 0 and 1, and the sum of all elements is 1, highlighting the modality information with a higher relevance to the current supply chain business scenario.
[0068] Finally, perform a weighted sum of the feature vectors of different modalities to achieve information fusion. Let the feature vectors of text, image, voice, and sensor data, and the fused feature vector comprehensively reflect the actual situation of each link in the supply chain.
[0069] (7) Intelligent decision-making and collaborative application steps:
[0070] Input the fused feature vector into a pre-trained deep neural network model based on the Transformer architecture for intelligent decision-making. During the training process, use the cross-entropy loss function, where is the number of samples, is the number of decision categories (such as categories like adjusting the procurement plan, optimizing the production process, adjusting logistics distribution, responding to customer demands, etc.), is the true label (0 or 1) of the sample belonging to category, and is the probability that the model predicts the sample belongs to category. Continuously adjust the parameters of the model through an optimization algorithm (such as the stochastic gradient descent method) so that the model can accurately make decisions based on the input multi-modal fusion information.
[0071] According to the decision result output by the model, automatically trigger the corresponding collaborative operation. For example, if the model determines that the customer demand has changed, based on the analysis result of the fusion information, automatically adjust the procurement plan, notify the production department to adjust the production arrangement, and coordinate with the logistics department to optimize the distribution plan to achieve the collaborative operation of each link in the supply chain and break the information silos.
[0072] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method for solving the problem of information silos in the supply chain based on multimodal AI, characterized in that, The following steps are involved: Step 1: Supply chain multimodal information collection and integration: Utilize collection equipment distributed in various links of supply chain procurement, production, logistics, and sales to collect multimodal information such as text, images, voice, and sensor data; and classify and store the collected multimodal information according to source and modality; Step 2: Feature extraction and encoding: For text information, the word vector model is used to map the words in the text into word vectors, which are then encoded through the RNN network to obtain feature vectors representing the text information; For image information, a pre-trained CNN model is used for feature extraction. After convolution, pooling and full connection layer operations, a fixed-length feature vector is obtained. For speech information, the MFCC features of the speech recording are extracted with the help of a specific library, and then further encoded using a neural network to obtain a feature vector representing the speech information; For sensor data information: the sensor data is first normalized, and then feature extracted and encoded through a custom neural network layer to obtain a feature vector representing the sensor data; Step 3: Supply chain island information integration: Based on the attention mechanism, the cosine similarity between feature vectors of different modalities is calculated, and the correlation is converted into attention weight using the Softmax function. The feature vectors of different modalities are weighted summed to perform information fusion. Step 4: Intelligent decision-making and collaborative application: The fused feature vector is input into a pre-trained deep neural network model based on the Transformer architecture for intelligent decision-making, and the corresponding collaborative operations are automatically triggered based on the decision results output by the model.
2. According to claim 1, a method for solving the problem of supply chain information islands based on multimodal Ai is characterized in that: The word vector model in step 2 is a Word2Vec model; The pre-trained CNN model in step 2 is a VGG16 model under the TensorFlow framework; In the second step, the specific library is the Librosa library. The MFCC features of the voice record are extracted using the Librosa library. Let the voice signal be s(t). After frame division, the nth frame signal is Sn(m), where m = 0, 1,..., M - 1 and M is the frame length. After windowing, we get m = 0, 1,..., M - 1, and calculate the power spectrum k = 0, 1,..., M - 1; then, after filtering through the Mel filter bank, we get H is the number of Mel filters. Finally, through discrete cosine transform, we obtain the MFCC coefficients L is the number of MFCC coefficients. After obtaining the MFCC features, a neural network is built using PyTorch for further encoding, and finally, a feature vector representing the voice information is obtained; In the step 2, the sensor data is normalized using MinMaxScaler, and the fully connected layer is used as a custom neural network layer for feature extraction and encoding.
3. A method for solving the problem of supply chain information silos based on multimodal AI according to claim 1, characterized in that, In the step 1, the procurement process collects text contracts, product sample images, and voice records of communication between the two parties provided by the supplier; The production process collects sensor data of equipment operation, production process monitoring images, and operators’ voice reports on production conditions; The logistics link collects the location and status sensor data of the transport vehicles, road condition images, and voice communications between the driver and the dispatch center; During the sales process, text evaluations of customer feedback, images of product usage scenarios, and voice communication records between customer service and customers are collected.
4. A method for solving the problem of information silos in the supply chain based on multimodal AI according to claim 1, characterized in that, In the step 4, when training the deep neural network model based on the Transformer architecture, a cross entropy loss function is used, and the model parameters are adjusted by an optimization algorithm, and the optimization algorithm is a stochastic gradient descent method.