Deep multi-modal fusion decision-making method based on building supplier selection

By collecting and integrating diverse data and dynamically adjusting weights using an attention mechanism, the problem of human bias and data limitations in traditional construction supplier selection methods has been solved, enabling intelligent and rapid supplier evaluation.

CN121524898APending Publication Date: 2026-02-13SHANGHAI MUZI CLOUD DATA TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511156701.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-19
Publication Date
2026-02-13

AI Technical Summary

Technical Problem

Traditional methods for selecting construction suppliers rely on manual assessments or single-dimensional data, which suffer from subjective biases and difficulty in fully leveraging the value of the data, thus failing to provide effective support for accurate decision-making.

Method used

By establishing a real-time connection between the supplier's enterprise resource planning system, the Ministry of Housing and Urban-Rural Development's regulatory platform, and third-party databases, diverse data is collected, semantic analysis, feature extraction, and fusion are performed, and weights are dynamically adjusted using an attention mechanism to achieve intelligent evaluation.

Benefits of technology

It achieves comprehensive integration and intelligent evaluation of multi-dimensional data, reduces manual intervention, shortens the evaluation cycle, and provides a reliable basis for supplier selection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121524898A_ABST
    Figure CN121524898A_ABST
Patent Text Reader

Abstract

The invention relates to a deep multi-modal fusion decision-making method for building supplier selection. The method comprises the following steps: establishing real-time connection with a supplier enterprise resource planning system, a residential building department supervision platform and a third-party engineering evaluation database through a standardized interface, and collecting multivariate data; performing semantic analysis on the text data to obtain a core semantic information table, performing standardized verification on the numerical data to obtain an abnormal data table, extracting building construction multi-modal features from the image data, and mapping the three types of information into a unified vector through a feature fusion model; intelligently matching the unified vector with a preset weight template, dynamically adjusting the weight through an attention mechanism, and performing fine adjustment and weighted calculation in combination with real-time information to obtain a preliminary evaluation score; and performing multi-dimensional cross validation on the preliminary evaluation score, and outputting a multi-modal fused supplier evaluation report after deviation is corrected in combination with historical data. And efficient integration and intelligent evaluation of the multi-dimensional data of the building suppliers are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the interdisciplinary field of construction engineering management and artificial intelligence, and specifically relates to a deep multimodal fusion decision-making method for selecting construction suppliers. Background Technology

[0002] In the field of construction project management, supplier selection directly impacts project quality, cost, schedule, and safety, making it a core element of construction supply chain management. Traditional supplier selection methods often rely on manual assessments or single-dimensional data, which have significant limitations: manual assessments are heavily influenced by subjective experience, prone to decision-making biases, and struggle to handle massive amounts of information; single data dimensions cannot comprehensively reflect a supplier's overall capabilities. With the advancement of digital transformation in the construction industry, supplier data is exhibiting multimodal characteristics, but existing technologies lack effective multimodal data fusion mechanisms, hindering the full extraction of data value and failing to provide strong support for accurate decision-making. Therefore, there is an urgent need for a decision-making method that can integrate diverse data and achieve intelligent assessment. Summary of the Invention

[0003] To address the aforementioned problems in the existing technology, this invention provides a deep multimodal fusion decision-making method for selecting building suppliers; The objective of this invention can be achieved through the following technical solutions: S1: Establish real-time connections with supplier enterprise resource planning systems, the Ministry of Housing and Urban-Rural Development's regulatory platform, and third-party engineering evaluation databases through standardized application programming interfaces to collect multi-data information from construction suppliers; S2: Extract the text data from the multivariate data and perform semantic parsing to obtain a core semantic information table; perform standardization verification on the numerical data in the multivariate data using a standardization model to obtain an anomaly data table; extract the multimodal features of building construction in key areas from the image data of the multivariate data based on a target recognition model; perform feature fusion using a feature fusion model to map the core semantic information table, the anomaly data table, and the multimodal features of building construction into a unified vector; S3: The unified vector is intelligently matched with the preset weight template of the attention mechanism in the decision fusion model through the intelligent evaluation module. The weights are dynamically adjusted according to the importance of features through the attention mechanism, and then fine-tuned in combination with real-time information. The features are weighted and calculated based on the adjusted template to obtain the preliminary evaluation score. S4: By performing text semantic consistency verification and numerical fluctuation rationality analysis on the preliminary evaluation scores, multi-dimensional cross-validation is achieved. Combined with historical evaluation data, deviations are corrected, and the decision fusion model is driven to output a multi-modal fusion supplier evaluation report.

[0004] Specifically, the registered capital reflects the supplier's financial strength and risk resistance capability. A range of values ​​is set, and the average registered capital of similar companies in the industry is used as a benchmark during standardization. The ratio of the deviation value to the standard deviation is calculated. The construction period of past projects is in days, and the project scale needs to be distinguished. The engineering quality score adopts a percentage system, usually based on the industry average score. The fluctuation range of material prices is calculated on a monthly basis and obtained through a formula.

[0005] Specifically, the attention mechanism dynamically adjusts the weights based on the importance of features. In actual matching, if a supplier's multimodal features of construction show high engineering quality, the weight of that feature is increased. If another supplier's abnormal data table shows anomalies, the weight of its related features is reduced. At the same time, a second fine-tuning is performed in combination with real-time information, and the adjusted weight template is finally obtained for weighted calculation of features to obtain a preliminary evaluation score.

[0006] Specifically, the deep multimodal fusion decision-making method for selecting construction suppliers includes a data acquisition unit, a data processing and fusion unit, a preliminary evaluation unit, and a verification and output unit. The data acquisition unit establishes real-time connections with the supplier's enterprise resource planning system, the Ministry of Housing and Urban-Rural Development's regulatory platform, and a third-party engineering evaluation database through a standardized application programming interface (API), thereby collecting multimodal data of construction suppliers. The data processing and fusion unit classifies the multimodal data, extracts text data for semantic parsing to obtain a core semantic information table, performs standardized verification on numerical data using a standardized model to obtain an anomaly data table, and extracts key regions from image data based on a target recognition model. The multimodal characteristics of building construction are then mapped into a unified vector through a feature fusion model. The preliminary evaluation unit, with the help of an intelligent evaluation module, intelligently matches the unified vector with the preset weight template of the attention mechanism in the decision fusion model. The weights are dynamically adjusted according to the importance of features through the attention mechanism, and a second fine-tuning is performed based on real-time information. The features are weighted and calculated according to the adjusted template to obtain the preliminary evaluation score. The verification output unit performs multi-dimensional cross-validation on the preliminary evaluation score, including text semantic consistency verification and numerical fluctuation rationality analysis. It corrects deviations by combining historical evaluation data, thereby driving the decision fusion model to output a multimodal fusion supplier evaluation report.

[0007] Specifically, the target recognition unit includes an image preprocessing model, a feature extraction model, and a target localization and classification model. The image preprocessing model performs image analysis and processing on the input construction site image. The specific processing methods include: using a median filtering algorithm to remove salt-and-pepper noise from the image, adjusting the image brightness and contrast through gamma correction, uniformly cropping and scaling the image, and standardizing the RGB color space to obtain a standardized image. Based on the standardized image output by the image preprocessing model, the feature extraction model extracts the surface texture features of building materials, the geometric contour features of construction equipment, and the color features of safety facilities, and outputs a high-dimensional feature map. The target localization and classification model adopts an improved Faster R-CNN architecture, generates ROIs through a region proposal network, where the ROI is a rectangular region that covers the minimum bounding rectangle of key targets including building materials, equipment, and safety facilities in the image, with the coordinates originating from the upper left corner of the image, and outputs the coordinates. At the same time, it classifies the targets within the ROI and outputs category labels.

[0008] Specifically, the Faster R-CNN architecture includes: a feature extraction layer, a region proposal network layer, a ROI pooling layer, and a classification and regression layer. The feature extraction layer employs a ResNet-50-based convolutional neural network to perform layer-by-layer convolution and pooling operations on the pre-processed construction images, outputting a high-dimensional feature map. The region proposal network layer, based on the high-dimensional feature map output by the feature extraction layer, generates ROIs through a sliding window, predicting the target probability and bounding box offset for each candidate region. The ROI pooling layer uniformly maps candidate regions of different sizes generated by the region proposal network layer to a fixed-size feature map. The classification and regression layer processes the fixed-size feature map output by the ROI pooling layer, outputting the target category and classification confidence, while simultaneously correcting the candidate region coordinates through bounding box regression, finally outputting the coordinates.

[0009] As a preferred technical solution of the present invention, the intelligent evaluation module in S3 includes a multi-source data preprocessing model, a cross-modal feature extraction model, a two-stage fusion modeling model, a dynamic weight evaluation model, and a multi-dimensional verification and correction model.

[0010] Specifically, the multi-source data preprocessing model is used to receive the collected multi-source data, perform word segmentation, entity recognition and redundant information filtering on text data, conduct preliminary screening of outliers and unit unification processing on numerical data, and perform noise reduction, scaling and format standardization on image data. Through these operations, the original data is transformed into structured text fragments, standardized numerical matrices and preprocessed image sets.

[0011] Specifically, the cross-modal feature extraction model is used to extract high-value features from different types of preprocessed data. For text data, a bidirectional encoder representation from transformers is adopted for semantic encoding to capture key semantics; for numerical data, a gradient boosting tree is used to screen the metrics related to strong supplier fulfillment capabilities; for image data, a real-time object detection module is utilized to detect key objects and extract visual features, finally forming a multi-modal feature library. Among them, the specific calculation method of the real-time object detection module includes: According to the high-dimensional feature map extracted by the convolutional network in object detection, a feature tensor of size H×W×C, which is unfolded into a matrix form , where N = H×W is the total number of spatial pixels, H and W are spatial dimensions, and C is the number of channels. Each column represents the feature vector of a channel; Perform singular value decomposition on X to obtain , where X is the matrix form after unfolding the high-dimensional feature map extracted by the convolutional network in object detection. Among them, U is the left singular matrix. According to the singular value decay rate, the first k dominant singular vectors are selected to form the matrix (k << N, usually k = 5 - 20, balancing accuracy and efficiency), to capture the main mode of the features; From the column vectors of U k , select the index corresponding to the element with the largest absolute value in each column to form the sample point set I = {i1, i2,..., i k} The sample points are the most representative positions in the feature space; Construct the selection matrix , where P[m,n] = 1 if n = i m , otherwise 0, which is used to extract the sample point features from X . At the same time, calculate the interpolation coefficient matrix ; For the original high-dimensional feature X, interpolate through the discrete empirical interpolation method to obtain the reduced-order feature . This reduced-order feature retains the core information, and the dimension is reduced from N×C to k×C, which can be directly input into the detection layer for object localization and classification calculation. On the premise that the detection accuracy loss is less than 5%, the feature processing time-consuming is reduced to k / N of the original; According to the dynamic objects in the video stream, re-execute steps 1 - 5 every 10 frames, update the sample point set I and the interpolation matrix C based on the new frame features, ensure that the reduced-order features can adapt to the target position changes, and maintain the robustness of real-time detection; Specifically, the construction progress feature is used to quantitatively assess the matching degree between the actual progress of the project and the planned construction period, and the progress index is mapped by the state changes of the region in the image; the building material quality feature reflects the quality compliance of building materials, and is generated based on the comparison results of material state and process standards in the image; the safety facility feature is used to assess the safety protection measures at the construction site, and is generated by identifying the existence status and standardization of use of safety facilities; the environmental impact feature is used to measure the degree of impact of construction activities on the surrounding ecology and residents' lives, and is generated based on the implementation effect of environmental protection measures in the image.

[0012] Specifically, the two-stage fusion model is used to address the problem of data heterogeneity and achieve effective fusion of multimodal features. First, text, numerical, and image features are dimensionally aligned and concatenated to form an initial fusion vector. Then, an attention mechanism is introduced to assign higher weights to the core dimensions. Finally, an autoencoder maps the vector to a unified feature vector, providing core input for subsequent evaluation calculations. Specifically, the dynamic weight evaluation model is used to calculate evaluation scores based on a unified vector and achieve intelligent weight adjustment. It incorporates templates for different industry scenarios, selects an appropriate template through vector similarity calculation, dynamically adjusts weights based on real-time data, and then generates preliminary evaluation results using a weighted summation formula, thus bridging feature fusion and cross-validation. Specifically, the multi-dimensional validation correction model is used to correct evaluation biases through cross-validation and output a final report. It compares the semantic consistency between text descriptions and image features, analyzes recent evaluation score fluctuations, fine-tunes the bias of the model trained on historical data, and finally integrates various information to generate a multi-modal evaluation report including visual charts, ensuring the reliability of the evaluation results.

[0013] The beneficial effects of this invention are as follows: (1) By setting up deep multimodal fusion technology, multi-dimensional data can be effectively integrated, breaking through the limitations of traditional single data evaluation. It can combine the supplier's qualification text description, engineering quality numerical score and safety and progress characteristics in the construction site image to more comprehensively depict the supplier's capabilities. At the same time, the multi-dimensional cross-validation mechanism further corrects the data deviation, making the evaluation results more in line with the actual business scenario and providing a reliable basis for selecting high-quality suppliers.

[0014] (2) By setting up an automated processing flow for real-time collection of multi-data, target recognition model, and feature fusion model, the traditional manual data processing and analysis process is replaced, significantly shortening the evaluation cycle. In addition, the intelligent matching method of dynamically adjusting weights through the attention mechanism can adaptively optimize the evaluation logic based on real-time information and business needs, reduce manual intervention, and realize end-to-end intelligent decision-making from data collection to evaluation report output, meeting the needs of the construction industry to respond quickly to market changes. Attached Figure Description

[0015] To facilitate understanding by those skilled in the art, the present invention will be further described below with reference to the accompanying drawings.

[0016] Figure 1 This is a flowchart illustrating a deep multimodal fusion decision-making method for selecting building suppliers according to the present invention. Detailed Implementation

[0017] To further illustrate the technical means and effects of the present invention in achieving its intended purpose, the following detailed description of the specific implementation methods, structures, features, and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided.

[0018] Please see Figure 1 A deep multimodal fusion decision-making method for building supplier selection S1: Establish real-time connections with supplier enterprise resource planning systems, the Ministry of Housing and Urban-Rural Development's regulatory platform, and third-party engineering evaluation databases through standardized application programming interfaces to collect multi-data information from construction suppliers; S2: Extract the text data from the multivariate data and perform semantic parsing to obtain a core semantic information table; perform standardization verification on the numerical data in the multivariate data using a standardization model to obtain an anomaly data table; extract the multimodal features of building construction in key areas from the image data of the multivariate data based on a target recognition model; perform feature fusion using a feature fusion model to map the core semantic information table, the anomaly data table, and the multimodal features of building construction into a unified vector; S3: The unified vector is intelligently matched with the preset weight template of the attention mechanism in the decision fusion model through the intelligent evaluation module. The weights are dynamically adjusted according to the importance of features through the attention mechanism, and then fine-tuned in combination with real-time information. The features are weighted and calculated based on the adjusted template to obtain the preliminary evaluation score. S4: By performing text semantic consistency verification and numerical fluctuation rationality analysis on the preliminary evaluation scores, multi-dimensional cross-validation is achieved. Combined with historical evaluation data, deviations are corrected, and the decision fusion model is driven to output a multi-modal fusion supplier evaluation report.

[0019] Specifically, the registered capital reflects the supplier's financial strength and risk resistance capability. A range of values ​​is set, and the average registered capital of similar companies in the industry is used as a benchmark during standardization. The ratio of the deviation value to the standard deviation is calculated. The construction period of past projects is in days, and the project scale needs to be distinguished. The engineering quality score adopts a percentage system, usually based on the industry average score. The fluctuation range of material prices is calculated on a monthly basis and obtained through a formula.

[0020] Specifically, the attention mechanism dynamically adjusts the weights based on the importance of features. In actual matching, if a supplier's multimodal features of construction show high engineering quality, the weight of that feature is increased. If another supplier's abnormal data table shows anomalies, the weight of its related features is reduced. At the same time, a second fine-tuning is performed in combination with real-time information, and the adjusted weight template is finally obtained for weighted calculation of features to obtain a preliminary evaluation score.

[0021] Specifically, the deep multimodal fusion decision-making method for selecting construction suppliers includes a data acquisition unit, a data processing and fusion unit, a preliminary evaluation unit, and a verification and output unit. The data acquisition unit establishes real-time connections with the supplier's enterprise resource planning system, the Ministry of Housing and Urban-Rural Development's regulatory platform, and a third-party engineering evaluation database through a standardized application programming interface (API), thereby collecting multimodal data of construction suppliers. The data processing and fusion unit classifies the multimodal data, extracts text data for semantic parsing to obtain a core semantic information table, performs standardized verification on numerical data using a standardized model to obtain an anomaly data table, and extracts key regions from image data based on a target recognition model. The multimodal characteristics of building construction are then mapped into a unified vector through a feature fusion model. The preliminary evaluation unit, with the help of an intelligent evaluation module, intelligently matches the unified vector with the preset weight template of the attention mechanism in the decision fusion model. The weights are dynamically adjusted according to the importance of features through the attention mechanism, and a second fine-tuning is performed based on real-time information. The features are weighted and calculated according to the adjusted template to obtain the preliminary evaluation score. The verification output unit performs multi-dimensional cross-validation on the preliminary evaluation score, including text semantic consistency verification and numerical fluctuation rationality analysis. It corrects deviations by combining historical evaluation data, thereby driving the decision fusion model to output a multimodal fusion supplier evaluation report.

[0022] Specifically, the target recognition unit includes an image preprocessing model, a feature extraction model, and a target localization and classification model. The image preprocessing model performs image analysis and processing on the input construction site image. The specific processing methods include: using a median filtering algorithm to remove salt-and-pepper noise from the image, adjusting the image brightness and contrast through gamma correction, uniformly cropping and scaling the image, and standardizing the RGB color space to obtain a standardized image. Based on the standardized image output by the image preprocessing model, the feature extraction model extracts the surface texture features of building materials, the geometric contour features of construction equipment, and the color features of safety facilities, and outputs a high-dimensional feature map. The target localization and classification model adopts an improved Faster R-CNN architecture, generates ROIs through a region proposal network, where the ROI is a rectangular region that covers the minimum bounding rectangle of key targets including building materials, equipment, and safety facilities in the image, with the coordinates originating from the upper left corner of the image, and outputs the coordinates. At the same time, it classifies the targets within the ROI and outputs category labels.

[0023] Specifically, the Faster R-CNN architecture includes: a feature extraction layer, a region proposal network layer, a ROI pooling layer, and a classification and regression layer. The feature extraction layer employs a ResNet-50-based convolutional neural network to perform layer-by-layer convolution and pooling operations on the pre-processed construction images, outputting a high-dimensional feature map. The region proposal network layer, based on the high-dimensional feature map output by the feature extraction layer, generates ROIs through a sliding window, predicting the target probability and bounding box offset for each candidate region. The ROI pooling layer uniformly maps candidate regions of different sizes generated by the region proposal network layer to a fixed-size feature map. The classification and regression layer processes the fixed-size feature map output by the ROI pooling layer, outputting the target category and classification confidence, while simultaneously correcting the candidate region coordinates through bounding box regression, finally outputting the coordinates.

[0024] As a preferred technical solution of the present invention, the intelligent evaluation module in S3 includes a multi-source data preprocessing model, a cross-modal feature extraction model, a two-stage fusion modeling model, a dynamic weight evaluation model, and a multi-dimensional verification and correction model.

[0025] Specifically, the multi-source data preprocessing model is used to receive the collected multi-source data, perform word segmentation, entity recognition and redundant information filtering on text data, conduct preliminary screening of outliers and unit unification processing on numerical data, and perform noise reduction, scaling and format standardization on image data. Through these operations, the original data is transformed into structured text fragments, standardized numerical matrices and preprocessed image sets.

[0026] In this embodiment, the data acquisition unit connects to three types of data sources through a standardized API: Supplier A's ERP system provides financial statements for the past three years and progress reports of projects under construction; the Ministry of Housing and Urban-Rural Development's regulatory platform returns its qualification certificates and safety accident records for the past five years; and a third-party database provides project review video frames and peer evaluation reports, synchronizing the data to the system in real time.

[0027] In this embodiment, specifically, the registered capital reflects the financial strength and risk resistance ability of the supplier, and its value range is usually 5 million yuan to 1 billion yuan. When performing standardization processing, the average registered capital of similar enterprises in the industry is used as the benchmark, and the ratio of the deviation value to the standard deviation is calculated. For example, if a supplier's registered capital is 200 million yuan, the industry average is 150 million yuan, and the standard deviation is 80 million yuan, then the standardized value is 0.625; the construction period of past projects is in days, and it is necessary to distinguish the project scale. For projects with an area of less than 100,000 square meters, the reasonable construction period is 300 - 500 days. After standardization, the performance efficiency of projects of different scales can be directly compared. During verification, the range verification is used to set "planned construction period ± 20%" as the reasonable interval, and if it exceeds, it is marked as abnormal; the engineering quality score uses a 100-point system, usually with the industry average score of 75 points as the benchmark, and the standardized value is (actual score - average score) / industry standard deviation. During logical verification, it is cross-judged in combination with the "number of high-quality engineering awards"; the fluctuation range of material prices is calculated on a monthly basis, and the formula is "(average price this month - average price last month) / average price last month × 100%", and the reasonable interval is set to ±5% (if it exceeds, it will affect cost stability). After standardization, it can be logically verified with the supplier's price commitment. If the committed fluctuation does not exceed 3% but the actual reaches 6%, it is determined as abnormal.

[0028] Specifically, the cross-modal feature extraction model is used to extract high-value features from different types of preprocessed data. For text data, a bidirectional encoder representation from the transformer is used for semantic encoding to capture key semantics; for numerical data, a gradient boosting tree is used to screen indicators strongly related to the supplier's performance ability; for image data, a real-time object detection module is used to detect key objects and extract visual features, and finally a multi-modal feature library is formed. Among them, the specific calculation method of the real-time object detection module includes: According to the high-dimensional feature map extracted by the convolutional network in object detection, a feature tensor with the size of H×W×C, which is unfolded into a matrix form , where N = H×W is the total number of spatial pixels, H and W are spatial dimensions, and C is the number of channels. Each column represents a feature vector of a channel; Perform singular value decomposition on X to obtain , where X is the matrix form after unfolding the high-dimensional feature map extracted by the convolutional network in object detection, and U is the left singular matrix. According to the singular value decay rate, the first k dominant singular vectors are selected to form a matrix (k << N, usually k = 5 - 20, balancing accuracy and efficiency), to capture the main modality of the features; From the column vectors of U k , select the index corresponding to the element with the largest absolute value in each column to form a sample point set I = {i1, i2,..., i k}, and the sample points are the most representative positions in the feature space; Constructing the selection matrix Where P[m,n]=1, if n=i m Otherwise, it is 0, used to extract sample point features from X. Simultaneously calculate the interpolation coefficient matrix. ; For the original high-dimensional features X, reduced-order features are obtained by interpolation using discrete empirical interpolation methods. This reduced-order feature retains the core information, and the dimension is reduced from N×C to k×C. It can be directly input into the detection layer for target localization and classification calculation. While ensuring that the detection accuracy loss is less than 5%, the feature processing time is reduced to k / N of the original. Based on the dynamic targets in the video stream, steps 1-5 are re-executed every 10 frames to update the sample point set I and interpolation matrix C based on the features of the new frame, ensuring that the reduced-order features can adapt to changes in the target position and maintain the robustness of real-time detection. In this embodiment, text data is parsed using the BERT model to extract the core semantics of "Grade A Professional Contractor for Steel Structure Engineering" and form an information table; numerical data, such as annual turnover of 520 million yuan and accident rate of 0.3 times / million man-hours, are standardized and marked as abnormal if "3 times of project overdue"; construction site images are processed using a real-time target detection algorithm to extract multimodal features such as "tower crane verticality deviation of 1.2°" and "rebar binding qualification rate of 98%", and then processed by the DEIM algorithm to reduce the order; the feature fusion model maps the three types of data into a 128-dimensional unified vector, in which the feature weight of "engineering quality" is increased by 15%.

[0029] Specifically, the two-stage fusion model is used to address the problem of data heterogeneity and achieve effective fusion of multimodal features. First, text, numerical, and image features are dimensionally aligned and concatenated to form an initial fusion vector. Then, an attention mechanism is introduced to assign higher weights to the core dimensions. Finally, an autoencoder maps the vector to a unified feature vector, providing core input for subsequent evaluation calculations. Specifically, the dynamic weight evaluation model is used to calculate evaluation scores based on a unified vector and achieve intelligent weight adjustment. It incorporates templates for different industry scenarios, selects an appropriate template through vector similarity calculation, dynamically adjusts weights based on real-time data, and then generates preliminary evaluation results using a weighted summation formula, thus bridging feature fusion and cross-validation.

[0030] In this embodiment, the intelligent evaluation module calls the preset template of "large public buildings" (engineering quality weight 30%). Because Supplier A's "accident rate is better than 90% of the industry average", the weight of safety features is dynamically increased to 20%. Combined with real-time steel price increase information, the weight of "material reserve capacity" is temporarily increased by 5%, and the weighted calculation yields an initial score of 89.6 points. Specifically, the multi-dimensional validation correction model is used to correct evaluation biases through cross-validation and output a final report. It compares the semantic consistency between text descriptions and image features, analyzes recent evaluation score fluctuations, fine-tunes the bias of the model trained on historical data, and finally integrates various information to generate a multi-modal evaluation report including visual charts, ensuring the reliability of the evaluation results.

[0031] In this embodiment, the verification unit found a semantic contradiction between "self-reported zero accidents" and "1 minor accident" on the regulatory platform, and corrected the safety feature weight to 18%. By comparing historical data, it found that the score fluctuation was within a reasonable range (standard deviation 4.2), and finally output an evaluation report with a radar chart, recommending supplier A as the first choice.

[0032] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.

Claims

1. A deep multimodal fusion decision-making method for building supplier selection, characterized in that, include: S1: Establish real-time connections with supplier enterprise resource planning systems, the Ministry of Housing and Urban-Rural Development's regulatory platform, and third-party engineering evaluation databases through standardized application programming interfaces to collect multi-data information from construction suppliers; S2: Extract the text data from the multivariate data and perform semantic parsing to obtain a core semantic information table; perform standardization verification on the numerical data in the multivariate data using a standardization model to obtain an anomaly data table; extract multimodal features of building construction in key areas from the image data of the multivariate data based on a target recognition model; Feature fusion is performed through a feature fusion model, mapping the core semantic information table, the abnormal data table, and the multimodal features of building construction into a unified vector. S3: The unified vector is intelligently matched with the preset weight template of the attention mechanism in the decision fusion model through the intelligent evaluation module. The weights are dynamically adjusted according to the importance of features through the attention mechanism, and then fine-tuned in combination with real-time information. The features are weighted and calculated based on the adjusted template to obtain the preliminary evaluation score. S4: By performing multi-dimensional cross-validation on the preliminary evaluation scores and correcting deviations by combining historical evaluation data, the decision fusion model is driven to output a multimodal fusion supplier evaluation report.

2. The method according to claim 1, characterized in that, The standardized model obtains numerical data related to the selection of construction suppliers from different data dimensions. Specifically, it includes the supplier's registered capital, the construction period of past projects, the engineering quality score, and the fluctuation range of material prices. The standardized model transforms the data into a standard normal distribution including the mean and standard deviation. The verification process sets reasonable ranges for each numerical indicator through range verification. Data that exceeds the range is marked as abnormal, and finally an abnormal data table is formed.

3. The method according to claim 1, characterized in that, The construction multimodal features are based on the target recognition unit extracting the registered capital, the construction period of past projects, the engineering quality score, and the material price fluctuation range from the multi-dimensional data. The target recognition unit is based on a real-time target detection algorithm to locate the physical space or equipment cluster directly associated with the core evaluation indicators in the image; the key features of the construction multimodal features are extracted, and finally the construction multimodal feature table is obtained.

4. The method according to claim 1, characterized in that, The core evaluation indicators include: financial strength indicators, project management efficiency indicators, engineering quality indicators, and cost control indicators, which reflect the supplier's comprehensive performance capabilities. The financial strength indicator corresponds to the registered capital, the project management efficiency indicator corresponds to the construction period of past projects, the engineering quality indicator corresponds to the engineering quality score, and the cost control indicator corresponds to the fluctuation range of material prices.

5. The method according to claim 1, characterized in that, The intelligent evaluation module evaluates the data by adjusting the weights appropriately. It matches the preset weight template of the attention mechanism in the unified vector and decision fusion model with the preset weight template, which is constructed based on historical high-quality supplier evaluation cases and industry expert experience.

6. The method according to claim 1, characterized in that, The multi-dimensional cross-validation is used to assess the accuracy of the results. The specific method is as follows: based on text semantic consistency verification and numerical fluctuation rationality verification, combined with historical evaluation data, the deviation is corrected. The current preliminary evaluation score is compared with the historical evaluation data of similar suppliers. If the weight settings of features in this evaluation are found to be different from those in historical data, the weights are adjusted according to the historical data. Finally, the decision fusion model is driven to output a multimodal fusion supplier evaluation report.

7. The method according to claim 3, characterized in that, The target recognition unit includes an image preprocessing model, a feature extraction model, and a target localization and classification model. The image preprocessing model performs image analysis and processing on the input construction site image. Specific processing methods include: using a median filtering algorithm to remove salt-and-pepper noise from the image; adjusting image brightness and contrast through gamma correction; uniformly cropping and scaling the image; and standardizing the RGB color space to obtain a standardized image. Based on the standardized image output by the image preprocessing model, the feature extraction model extracts surface texture features of building materials, geometric contour features of construction equipment, and color features of safety facilities, outputting a high-dimensional feature map. The target localization and classification model uses an improved Faster R-CNN architecture, generating Region of Interest (ROIs) through a Region Proposal Network. Each ROI is a rectangular region covering the minimum bounding rectangle of key targets in the image, including building materials, equipment, and safety facilities. The coordinates are based on the top-left corner of the image, and the model outputs coordinates. Simultaneously, it classifies targets within the ROI and outputs category labels.

8. The method according to claim 4, characterized in that, The dual-stage weighted collaborative fusion strategy is divided into two fusion stages: the first fusion stage is used to perform word embedding processing on the core semantic information of the text data, convert it into vector form, quantize the abnormal label information in the abnormal data table into binary vectors, and perform dimensionality reduction processing on the multimodal features of building construction. The second fusion stage automatically identifies the weights of different features in decision-making through an attention mechanism and an autoencoder model. The autoencoder maps vectors from different sources to the same dimensional space to generate the unified vector table.

9. The method according to claim 6, characterized in that, The text semantic consistency check is used for the text information involved in the evaluation process, including whether the supplier's self-stated advantages are consistent with the third-party evaluation, and whether the text description of the core semantic information table and the preliminary evaluation score are consistent. If a contradiction is found, the relevant data will be re-checked and the evaluation score will be corrected. The numerical fluctuation reasonableness check is used for the preliminary evaluation score and related numerical indicators. If the fluctuation of the supplier's evaluation score exceeds the threshold, the data processing and fusion process will be traced back to find the cause of the anomaly and make corrections.

10. The method according to claim 7, characterized in that, The Faster R-CNN architecture includes: a feature extraction layer, a region proposal network layer, a ROI pooling layer, and a classification and regression layer. The feature extraction layer uses a ResNet-50-based convolutional neural network to perform layer-by-layer convolution and pooling operations on the pre-processed construction images, outputting a high-dimensional feature map. The region proposal network layer, based on the high-dimensional feature map output by the feature extraction layer, generates ROIs through a sliding window, predicting the target probability and bounding box offset for each candidate region. The ROI pooling layer uniformly maps candidate regions of different sizes generated by the region proposal network layer to a fixed-size feature map. The classification and regression layer processes the fixed-size feature map output by the ROI pooling layer, outputting the target category and classification confidence, and simultaneously corrects the candidate region coordinates through bounding box regression, finally outputting the coordinates.