Enterprise complex decision auxiliary method and system based on multi-modal cognitive intelligent large model

By adopting multimodal cognitive intelligent big model and input data enhancement strategies in complex enterprise decision-making assistance systems, the shortcomings of existing systems in data quality, knowledge acquisition and model interpretability are solved, and more efficient and intelligent decision-making assistance effects are achieved.

CN120163290APending Publication Date: 2025-06-17SHANDONG FUTURE NETWORK TECHNOLOGY DEVELOPMENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510248574.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

Existing enterprise complex decision-making auxiliary methods and systems have shortcomings in data quality, knowledge acquisition and model interpretability, and it is difficult to effectively deal with complex decision-making scenarios with dynamic changes.

Method used

Adopting a complex decision-making assistance method based on multimodal cognitive intelligent big model, we optimize the training and deployment of big models to generate decision suggestions through multimodal data analysis, innovative big model architecture and input data enhancement strategies based on historical data.

Benefits of technology

It significantly improves the rationality and intelligence of model decisions, enhances the robustness and adaptability of the model, and improves the accuracy and efficiency of decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163290A_ABST
    Figure CN120163290A_ABST
Patent Text Reader

Abstract

The invention provides an enterprise complex decision auxiliary method and system based on a multi-modal cognitive intelligent large model, and belongs to the technical field of artificial intelligence. The method comprises the steps of collecting multi-modal enterprise complex decision auxiliary data, and performing data cleaning and data labeling to obtain a training data set; constructing an enterprise complex decision-making auxiliary large model; inputting the training data set into the enterprise complex decision auxiliary large model, and optimizing large model training to obtain a trained enterprise complex decision auxiliary large model; and deploying the trained large model to a server, and generating decision suggestions in combination with an input data enhancement strategy. According to the method, through joint processing of the bimodal data, the reasonability and the intelligent degree of model decision can be remarkably improved, and the model prediction accuracy in different scenes is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an enterprise complex decision-making assistance method and system based on a multi-modal cognitive intelligence large model, belonging to the field of artificial intelligence technology. Background Art

[0002] In today's highly competitive and complex business environment, it is of great significance to develop enterprise complex decision-making assistance methods and systems. From the perspective of the enterprise's own operation, it can significantly improve the decision-making efficiency and accuracy. In the past, enterprise decision-making relied on manual collection and analysis of data, which was time-consuming and error-prone. However, this system can quickly process massive amounts of data with advanced algorithms and models, and promptly provide comprehensive and accurate decision-making suggestions for the enterprise. At the same time, the system can identify potential risks in advance by deeply mining and analyzing multi-dimensional data. For example, when making investment decisions, it comprehensively evaluates risks based on multiple factors, helping the enterprise avoid losses caused by blind decisions. From the perspective of resource utilization, this system analyzes the resource requirements and return expectations of each business segment according to the enterprise's strategic goals and actual situations, enabling resources to be accurately invested in high-value areas, improving resource utilization efficiency, and thus enhancing the overall efficiency of the enterprise. From the long-term development of the enterprise, the data analysis provided by the system can help the enterprise discover new market opportunities and potential demands, and promote the enterprise to develop new products and explore new business fields.

[0003] The existing enterprise complex decision-making assistance methods and systems include the following methods: (1) Decision-making assistance method based on an expert system: Store the knowledge and experience of experts in the knowledge base in the form of rules. When the enterprise faces a decision-making problem, the system uses an inference engine to search for matching rules in the knowledge base according to the input problem and conditions, and then obtains decision-making suggestions. For example, in the risk assessment decision of an enterprise, the judgment rules of risk management experts for different risk scenarios are entered into the knowledge base, and the system can give a risk level assessment and response strategy based on the current risk factors faced by the enterprise, and use expert wisdom to assist the enterprise in making scientific decisions; (2) Decision-making assistance method based on data mining: This method mines potential patterns, trends, and correlation information from a large amount of enterprise data. It first collects and organizes the internal business data, customer data, financial data, etc. of the enterprise, and then uses data mining algorithms to provide strong data support for enterprise decision-making; (3)Machine learning-based decision-making assistance method: Use machine learning algorithms to train historical data and build a prediction model. Common machine learning algorithms include linear regression, logistic regression, support vector machines, neural networks, etc. Taking enterprise sales prediction as an example, use historical sales data, market trend data, promotion activity data, etc. as training data to train a linear regression model or a neural network model. After the model learns the patterns in the data, it can predict future sales situations. Enterprises can formulate production plans, inventory management strategies, etc. based on the prediction results to achieve intelligent decision-making assistance.

[0004] However, the above methods have the following deficiencies and disadvantages in decision-making assistance: The decision-making assistance method based on data mining relies on a large amount of high-quality data. If the data is missing, incorrect, or incomplete, the mining results may have a large deviation, and it is difficult to handle complex decision-making scenarios with dynamic changes. The decision-making assistance method based on expert systems has bottlenecks in the knowledge acquisition process. It takes time and effort to sort out and transform expert experience into rules, and the system has poor flexibility. It is difficult to provide effective suggestions when there are no corresponding rules in the knowledge base for new problems. The model training cost of the machine learning-based decision-making assistance method is high, requiring a large amount of computing resources and time; and the model interpretability is weak. It is difficult for enterprises to understand the decision-making basis of the model, resulting in limited applications in some scenarios with high requirements for decision-making transparency. Summary of the Invention

[0005] The object of the present invention is to provide an enterprise complex decision-making assistance method and system based on a multi-modal cognitive intelligence large model, aiming to overcome the deficiencies of traditional methods through methods such as multi-modal data analysis, innovative large model architecture, and input data enhancement strategy based on historical data.

[0006] To achieve the above object, the present invention is realized through the following technical solutions: An enterprise complex decision-making assistance method based on a multi-modal cognitive intelligence large model, including: Collect multi-modal enterprise complex decision-making assistance data, perform data cleaning and data annotation to obtain a training data set, Each group of data in the training data set includes problem description data and corresponding analysis results and solution labels, where the problem description data includes text data and time series data; Build an enterprise complex decision-making assistance large model including a word embedding layer, a text data encoder module, a time series data encoder module, a feature alignment unit, and a decoder module; Input the training data set into the enterprise complex decision-making assistance large model, and use the cross-entropy loss function, dynamic learning rate adjustment, and sample weight self-adaptive method to optimize the large model training to obtain a trained enterprise complex decision-making assistance large model; Deploy the trained large model to the server and generate decision-making suggestions in combination with the input data augmentation strategy; the input data augmentation strategy is as follows: store the historical Q&A data of the large model through a data experience pool, calculate the similarity between the new question and the historical data, and use the similar historical answers as the input augmented text data.

[0007] Preferably, the data processing method of the enterprise complex decision-making assistance large model is as follows: For the input text data, the feature vector of the text data is calculated through the word embedding layer and the learnable parameter matrix; a text data encoder with a stacked structure is used to perform deep feature extraction on the feature vector; the output of the upper encoder of the text data encoder is used as the input of the lower encoder, and the number of the text data encoders is pieces; For the input time series data, a time series data encoder with a stacked structure is directly used to expand the feature extraction, and the number of stacked layers of the time series data encoder is the same as that of the text data encoder; the output of the upper encoder of the time series data encoder is used as the input of the lower encoder; The output features of the text data encoder module and the time series data encoder module at the same layer are jointly used as the input of the feature alignment unit at this layer, and the number of the feature alignment units is the same as that of the text data encoder; The output of the feature alignment unit is used as the input of the decoder, and the number of decoders is the same as that of the text data encoder. The input of the th layer decoder includes the output of the th layer decoder and the output of the th layer feature alignment unit; The output of the th layer decoder is processed by the softmax function to obtain the prediction result of the large model.

[0008] Preferably, the feature vector of the input text data is calculated through the word embedding layer and the learnable parameter matrix, and the specific method is as follows: , , where represents the word embedding vector of the text data, represents the text data, represents the position vector of the text data, represents the embedding layer operation function, represents the feature vector of the text data; The position vector is realized through the learnable parameter matrix , and the size of this matrix is , Indicates the number of words contained in the input text data, represents the dimension of the position vector; the position vector of each word in the input text data is the parameter matrix The index corresponding to the word in is the row vector of, where .

[0009] Preferably, the text data encoder module sequentially includes: two fully connected layers, a Relu activation function, a multi-head attention mechanism, a BN normalization layer, a Dropout layer, and a Simgoid activation function; The feature vector of the text data After passing through two fully connected layers in sequence, the Relu activation function is used to activate the features to obtain preliminary features ; The preliminary features are sent to the multi-head attention mechanism for processing, and the specific method is as follows: The preliminary features are mapped to three different matrix feature spaces , , where, , , respectively represent , , matrix calculation functions, , , respectively represent the calculated under the , , matrix of the th attention head, and the total number of heads included in the multi-head attention mechanism is , Calculate the single attention head feature output factor : where, represents the softmax function, is the normalization factor, represents the matrix transpose of.

[0010] Based on the calculated feature output factors of each attention head calculate the feature output of each attention head , and the formula is as follows: , The feature outputs of each attention head in the multi-head attention mechanism are integrated through a fully connected layer to obtain a feature output ; The feature output is successively passed through a BN normalization layer, a Dropout layer, and a Sigmoid activation function to further process the feature and obtain the feature vector after being processed by the encoder unit .

[0011] Preferably, the time series data encoder module successively includes: three LSTM layers, a multi-head attention mechanism, two LSTM layers, and a Relu activation function; The input time series data is directly subjected to feature extraction through three LSTM layers in sequence to obtain multi-level features of the time series data ; The features are processed through a multi-head attention mechanism to obtain the processed features ; Then, the features are processed and extracted through two LSTM layers, and after passing through the Relu activation function, the construction of the time series data encoder module is completed, and the features output by the time series data encoder are obtained .

[0012] Preferably, the feature alignment unit processes features as follows: The features , are normalized using a BN normalization layer, and activated using an activation function. When aligning features, an adaptive adjustment scheme is used to determine the number of dimensions after feature alignment : Among them, represents the dimension of the input feature vector , represents the dimension of the input feature vector , is a dimension adjustment parameter, represents the floor operation function; The number of dimensions of the aligned features is obtained, and then a dimension adjustment mapping is performed through a fully connected layer to unify the dimensions of the two features to ; The data features of the two modalities are fused using a fully connected layer to complete feature alignment.

[0013] Preferably, the decoder includes a feed-forward neural network, a multi-head attention mechanism, a BN normalization layer, a Dropout layer, a BN normalization layer, and a Dropout layer; the feed-forward neural network includes two fully-connected layers and a mish activation function.

[0014] Preferably, the specific method of the input data augmentation strategy is as follows: Construct a data experience pool and store the original user input data in the data experience pool , the corresponding output results of the large model , data identifiers, and timestamps. The original user input data is the problem description data input to the large model; Mix the new problem description data and the original problem description data existing in the experience pool . The problem description data includes both text data and time series data . Use the Word2Vec word embedding layer to calculate the text data to obtain its corresponding word embedding vector . Based on the word embedding vectors of each data, use the K-means clustering algorithm to divide all the data in the experience pool into X data clusters to complete the preliminary data classification, and record the data cluster where the new problem description data is located. Calculate the cosine similarity between the word embedding vector of the text data in the data cluster where the new problem description data is located and the word embedding vectors of other text data in the data cluster as well as and the Euclidean distance between the time series data of and the time series data of other data in the data cluster. The formulas are as follows: , , where represents the word embedding vector of the text data of the new problem description data, represents the word embedding vectors of other text data in the data cluster, represents the dot product calculation function, represents the norm calculation function, represents the th element in the time series data of the new problem description data, represents the th element in the time series data of other data in the data cluster, , represents the normalization operation function; Based on the calculated cosine similarity and Euclidean distance, add them to obtain the feature similarity score ; Based on the similarity scores between the new input problem description data and each piece of problem description data in the experience pool, a threshold is set. The adjustable range of is (1.5 - 1.75). The larger the threshold is set, the higher the standard for similarity determination, indicating that only new problem description texts with higher data similarity to the experience pool data can perform the input data enhancement operation at this time. For data with similarity scores greater than the threshold, the corresponding large model response data is added to the problem description text. When a new question generates a Q&A each time, the data ( , ) is automatically stored in the experience pool, and according to the time decay strategy, based on the timestamp information contained in each piece of data, the data in the experience pool that exceeds the save time threshold is deleted; specifically, the adjustable range of the save time threshold is (10 - 14 months). According to the timestamp information of each piece of data in the experience pool, the data whose existence time exceeds this time threshold will be automatically deleted. Therefore, the larger the time threshold, the slower the update frequency of the experience pool data, and vice versa, the faster the update frequency of the experience pool data.

[0015] Preferably, the loss function of the enterprise complex decision-making assistance large model is as follows: , where, represents the number of samples, represents the sequence length, represents the conditional probability of the th data of the th data of the st text data sample and the rd data in the th policy solution predicted by the data model; The dynamic learning rate adjustment is specifically as follows: In the initial training stage of the model, the initial value of the learning rate is set high. As the training progresses step by step, calculate the loss value of the model after each training, and use this loss value for dynamic optimization and update of the learning rate: , where, represents the learning rate after each update, represents the learning rate decay factor, represents the loss function value factor, represents the Euler's constant; The specific method of the sample weight adaptation is as follows: Before the start of training, assign a weight with a value of to each group of samples, where is the number of training sample groups in the dataset. After the first round of training, adjust the sample weights according to the loss value of each group of samples: wherein, represents the weight adjustment factor, represents the sample weight after each update, represents the initial sample weight; after the weights are updated, use the samples with updated weights to recalculate the parameter gradients in the model and update the model parameters. When calculating the gradients, multiply the gradients obtained from each sample by its corresponding sample weight .

[0016] An enterprise complex decision-making assistance system based on a multi-modal cognitive intelligence large model, comprising: Data acquisition and processing module: used to acquire multi-modal enterprise complex decision-making assistance data, perform data cleaning and data annotation to obtain a training dataset; Model architecture design module: used to construct an enterprise complex decision-making assistance large model including a word embedding layer, a text data encoder module, a time series data encoder module, a feature alignment unit and a decoder module; Model training module: adopts a cross-entropy loss function, dynamic learning rate adjustment and sample weight self-adaptation method to optimize the training of the large model to obtain a trained enterprise complex decision-making assistance large model; Model deployment and application module: deploy the trained large model to a server and generate decision-making suggestions in combination with an input data enhancement strategy; Input data enhancement module: store the historical Q&A data of the large model through a data experience pool, calculate the similarity between the new question and the historical data, and use the similar historical answers as input to enhance the text data.

[0017] The advantages of the present invention are as follows: (1) Multi-modal data feature analysis: The present invention jointly uses text data and time series data related to enterprise development (such as monthly sales data, monthly product sales volume data) as input data for the large model, emphasizes the importance of time series data in making enterprise complex decisions, and can significantly improve the rationality and intelligence level of model decisions through the joint processing of bimodal data; (2) Innovative enterprise complex decision-making assistance large model architecture design: In the model architecture designed in the present invention, a text data encoder module and a time series data encoder module are designed, and a feature alignment unit is designed to flatten the dimensions of the two modal data. Using this multi-modal data encoder structure and feature alignment unit ensures that the model can understand the mutual relationships between complex data, enhances the model's context learning ability, and thus can propose relatively complex strategies, enhancing the robustness of the model; (3) Input data enhancement strategy based on historical data: A data reserve experience pool containing historical Q&A data is designed, and a data similarity calculation method is designed to obtain the similarity between the new question description data and the historical question description data, and the solution of the historical question is used as an additional input to the large model at this time according to the similarity score, improving the integrity of the input data. It helps the model better use the existing knowledge reserve for reasoning and judgment, improves the adaptability and generalization ability of the model in different scenarios, and ensures the model prediction accuracy in different scenarios. Description of the Drawings

[0018] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation to the present invention.

[0019] Figure 1 It is a schematic diagram of the method flow structure of the present invention.

[0020] Figure 2 It is a schematic diagram of the text data encoder module.

[0021] Figure 3 It is a schematic diagram of the time series data encoder module.

[0022] Figure 4 It is a schematic diagram of the decoder module.

[0023] Figure 5 Schematic diagram of the enterprise complex decision-making assistance large model architecture.

[0024] Figure 6 It is a comparison chart of the prediction accuracy rate of the large model of the present invention.

[0025] Figure 7 It is a comparison chart of the decision accuracy rate of the large model of the present invention.

[0026] Figure 8 It is a comparison chart of the model performance with and without strategies of the present invention.

[0027] Figure 9 Overall technical route flow chart. Detailed implementation manners

[0028] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0029] An enterprise complex decision-making assistance method based on a multi-modal cognitive intelligence large model includes the following main steps: First, collect data from enterprise internal, public data, and third-party data platforms, and perform cleaning and annotation. In the model architecture design, set up text and time-series data encoder processing modules and a feature fusion unit, and complete the overall architecture in combination with a decoder. Construct an input data enhancement strategy based on historical data, and add historical answers to similar questions as input through a data experience pool and similarity calculation. At the same time, design a loss function, learning rate adaptive optimization, and sample weight strengthening training strategies to optimize model training. Finally, deploy a server system within the enterprise, deploy the trained model therein and test it to assist enterprise operation.

[0030] As Figure 1 shown, specifically including: S1: Construction of an enterprise complex decision-making assistance data set, collect multi-modal enterprise complex decision-making assistance data, perform data cleaning and data annotation to obtain a training data set. Each group of data in the training data set includes problem description data and corresponding analysis results and solution labels, where the problem description data includes text data and time-series data.

[0031] S2: Design of the enterprise complex decision-making assistance model architecture, construct an enterprise complex decision-making assistance large model including a word embedding layer, a text data encoder module, a time-series data encoder module, a feature alignment unit, and a decoder module.

[0032] S3: Construct a model adaptive optimization training strategy, input the training data set into the enterprise complex decision-making assistance large model, and optimize the training of the large model by using a cross-entropy loss function, dynamic learning rate adjustment, and sample weight self-adaptation method to obtain a trained enterprise complex decision-making assistance large model.

[0033] S4: Construction of an input data enhancement strategy based on historical data and actual deployment and use of the model. Deploy the trained large model to the server, and generate decision-making suggestions in combination with the input data enhancement strategy; the input data enhancement strategy is as follows: Store the historical question-and-answer data of the large model through a data experience pool, calculate the similarity between the new question and the historical data, and use the similar historical answers as input to enhance the text data.

[0034] As a refinement of the above embodiment, step S1 is to ensure the integrity of data collection, and collect data from within the enterprise, public data, and third-party data platforms. And to further improve the data quality, data cleaning operations and detailed data annotation operations are performed on the data to obtain an enterprise complex decision-making assistance data set for subsequent model training; specifically including: S1-1: Data collection scenario design: To achieve the diversity and comprehensiveness of data collection, data collection is carried out under three different scenarios. First, data is collected within the enterprise to clarify the key information of enterprise operations, enabling the model to capture the unique laws and patterns of the enterprise, so as to provide decision-making support more in line with the actual situation of the enterprise; secondly, data is collected from public data outside the enterprise, and web crawler technology is used to search for data related to enterprise decisions on websites related to the enterprise by using keyword search. Finally, data is supplemented through a third-party data platform to improve the integrity of the data.

[0035] S1-2: Initial data collection: Data collection is carried out under the data collection scenarios described in S1-1. In the present invention, the problem description data includes two different modalities of corresponding problem description text data and time-series data related to the enterprise. When collecting the data set, various types of enterprise problem description text data and the corresponding time-series description data are collected. To ensure the superiority of the large model decision-making and deal with various levels of enterprise problems, the present invention widely collects the cause analysis and solutions of various types of enterprise problems.

[0036] S1-3: Data cleaning: Denoising processing is performed on the text data collected in the process of S1-2, and invalid characters such as Html tags, special characters, and garbled characters are deleted, and all letters in the text are unified into uppercase. For sequence data, the outlier data is deleted. In addition, when there are missing values in the sequence data but have little impact on the overall data, the data containing missing values is also deleted.

[0037] S1-4: Data annotation: For the data after the data cleaning operation in S1-3, professional staff are organized to perform data annotation on it to provide more accurate training data for model training. After the data annotation operation, each group of training data includes problem description data , problem description data which includes text data and time-series data . For each group of problem description data, a label information is assigned to it, and this label information is the analysis result and solution of this problem .

[0038] For example: 1. Analyze the reasons for the decline in the third-quarter revenue in combination with the monthly revenue situation of last year. This set of data is the problem description text data. 2. The monthly revenue data of the enterprise last year. This data is time-series data with a monthly time interval. 3. Labels: Reasons for the decline in the enterprise's third-quarter revenue analyzed from various perspectives; Another example: 1. The power consumption of the company has been too high recently. What are the reasons for this and some power-saving solutions are proposed. This set of data is the problem description text data. 2. The daily average power consumption data of the enterprise. This data is time-series data with a daily time interval. 3. Labels: Reasons for the excessive power consumption of the enterprise and feasible power-saving solutions.

[0039] After the data annotation is completed, a standard enterprise complex decision-making assistance data set is obtained. Each set of data in the data set includes problem description data as well as analysis results and solution data . The problem description data includes text data and time-series data ; , .

[0040] As a refinement of the above embodiment, step S2 is to ensure the processing ability of the model for multi-modal data. A text data encoder processing module and a time-series data encoder processing module are designed in the model architecture, and a feature fusion unit is designed to achieve effective fusion of bimodal data. Subsequently, the overall architecture design of the enterprise complex decision-making assistance large model is completed in combination with the designed decoder structure.

[0041] The data processing process of the large model specifically includes: For the input text data , after calculating the feature vector of the text data using the word embedding layer and the learnable parameter matrix, a stacked text data encoder is then used to perform deep feature extraction on the feature vector. In this embodiment, the stacking layer number is 12.

[0042] For the input time-series data , a stacked time-series data encoder is directly used to perform feature extraction. Similarly, to maintain the matching of data at different levels, the stacking layer number of the time-series data encoder is also 12.

[0043] In the above encoder structure, the output of the upper-level encoder is used as the input of the lower-level encoder. The decoder architecture used in the present invention also includes 12 decoder modules.

[0044] It should be noted that the counting method of the encoder is from top to bottom, and the counting method of the decoder is from bottom to top. For the For the output of the text data encoder and the timing data encoder of the layer, after being processed by the feature alignment unit, it serves as one of the inputs to the decoder module of the 13th - layer. For the decoder module, the output of the upper - layer decoder also serves as one of the inputs to the lower - layer decoder. After being processed by 12 encoders and then processed by the softmax activation function, the prediction result of the large model is obtained.

[0045] The following specifically describes various data - processing flows S2 - 1: Text data processing process: 1) Construction of text - data feature vectors: To facilitate the subsequent model's feature - analysis and processing of text data, at the beginning of the model, the word - embedding vector and the position vector of the text data need to be calculated, and the two vector data are added to obtain the feature vector corresponding to the text data. Specifically, the text data is sent to the Word2Vec embedding layer to calculate the corresponding word - embedding vector . In this embodiment, the position vector is implemented through a learnable parameter matrix . This parameter matrix is a matrix of size . represents the number of words contained in the input text data, represents the dimension of the position vector. The position vector of each word in the input text is the row vector corresponding to the index of this word in the parameter matrix , where .

[0046] During model training, the parameter matrix will update the value of the parameter matrix according to the update rule of the model optimizer, so that the parameter matrix can learn the position - vector representation most suitable for the current task. Compared with the fixed - position - encoding method, using this learnable position - encoding method enables the model to learn a position vector closer to the actual situation according to the specific task and data to enhance the overall performance of the model. Then, the feature vector is obtained as: , , where, represents the embedding - layer operation function; 2) Design of the text data encoder module: The architecture diagram of the encoder module for text data is as shown in Figure 2 and is described in words as follows: After obtaining the feature vector corresponding to the text data an encoder unit is constructed to further process the feature vector In the encoder unit, first, there are two interconnected fully connected layers, and then the Relu activation function is used to activate the features to obtain the preliminary features : , where represents the fully connected layer operation function, represents the Relu activation function. Subsequently, the features are sent into the multi-head attention mechanism for further processing.

[0047] The construction method of the multi-head attention mechanism is as follows: (1) Based on the feature it is mapped to three different matrix feature spaces to improve the model's understanding ability of complex models: , where , , respectively represent , , matrix calculation functions, , , respectively represent the th attention head's , , matrices. The total number of heads in the multi-head attention mechanism is , .

[0048] Calculate the single attention head feature output factor : where represents the softmax function, is the normalization factor, represents the transpose of the matrix .

[0049] Based on the calculated feature output factors for each attention head, calculate the feature output and integrate the feature outputs of each attention head in the multi-head attention mechanism through a fully connected layer to obtain a feature output : , , , wherein, represents the operation function of the multi-head attention mechanism

[0050] Subsequently, the BN normalization layer, Dropout layer, and Sigmoid activation function are used to further process the feature to obtain the feature vector after being processed by the encoder unit : .

[0051] S2-2: Design of the time series data encoder module: Its model schematic diagram is as Figure 3 shown, and the text description is as follows In the time series data encoder module, for the data three interconnected LSTM layers are designed to directly extract features, and multi-level features of the time series data are collected through this stacked LSTM network layer structure to obtain a richer feature representation , wherein, represents the operation function of the LSTM network layer

[0052] Afterwards, the multi-head attention mechanism described in S2-1 is used to further process the feature to obtain the processed feature : , Subsequently, two interconnected LSTM layers are designed to further process and extract features from the feature , and then through the final ReLU activation function, the construction of the time series data encoder module is completed, and the feature output by the time series data encoder is obtained : , S2-3: Design of the feature alignment unit: For the feature obtained in S2-1 and the feature obtained in S2-2, since the two types of features have different dimensions, in order to facilitate the subsequent model to fuse the two features, a feature alignment unit is designed in the present invention to align the feature and the feature Adjust the dimensions to be consistent. In the feature alignment unit, first use the BN normalization layer to normalize the features to accelerate model convergence. Subsequently, use the Relu activation function to activate the features. When aligning the features, use an adaptive adjustment scheme to determine the number of dimensions after feature alignment. : Among them, represents the input feature vector 's dimension, represents the input feature vector 's dimension, is the dimension adjustment parameter, represents the floor operation function; Calculate the number of dimensions of the aligned features After that, perform dimension adjustment mapping through a fully connected layer to unify the dimensions of the two features to . Subsequently, use another fully connected layer to fuse the data features of the two modalities, thus completing the construction of the feature alignment unit.

[0053] S2-4: Decoder module design: The schematic diagram of the decoder module model is as shown in Figure 4 shown, and the text description is as follows: In the decoder module, first, it contains a feed-forward neural network composed of two fully connected layers and the mish activation function. Then, it is the multi-head attention mechanism described in S2-1. Then, it is a BN normalization layer and a Dropout layer to enhance the generalization ability of the model. Finally, perform activation processing through the Relu activation function to complete the construction of the decoder module.

[0054] As a refinement of the above embodiment, in step S3, a model adaptive optimization training strategy is constructed: a loss function for model training is designed, and a learning rate adaptive optimization adjustment strategy and a sample weight strengthening training strategy during model training are constructed to optimize the training process of the model and improve the overall training quality of the model.

[0055] S4-1: Loss function construction: In the present invention, a loss function used by the model is constructed based on cross-entropy.

[0056] , Among them, represents the number of samples, represents the sequence length, represents based on the th text data sample's th data and the th time series data sample's The conditional probability of the th policy solution obtained from the th data in a data model prediction.

[0057] S4-2 Learning rate dynamic optimization strategy construction: In the initial training stage of the model, the learning rate is set to a relatively large initial value of 0.01 to accelerate the model convergence speed. As the training progresses step by step, calculate the loss value of the model after each training , and use this loss value to perform dynamic optimization update of the learning rate: , where, represents the learning rate after each update, represents the learning rate decay factor, represents the loss function value factor, represents the Euler's constant. According to this learning rate dynamic optimization strategy, during the model training process, the learning rate is dynamically and flexibly adjusted in real time according to the training loss, making the model training more efficient and stable. To prevent the learning rate from decaying too much and causing the model training to stagnate, a lower limit of the learning rate is set in the present invention, . When the learning rate decays to the lower limit value, it will no longer continue to decay and maintain this value for subsequent training. This can ensure that the model still has a certain ability to update parameters in the later stage of training.

[0058] S4-3: Sample weight adaptive adjustment strategy: In this embodiment, in addition to adaptively adjusting the learning rate of training, the training weight of each sample is also adaptively updated during the training stage. Before the training starts, since the dataset contains a total of A groups of training samples, a weight with a value of is assigned to each group of samples. After the first round of training, the larger the loss value corresponding to each group of samples after training, the higher the training difficulty of this group of samples. If the loss value is smaller, it means that this group of samples is easy to be learned by the model: where, represents the weight adjustment factor, represents the sample weight after each update. After the weight is updated, use the samples with updated weights to recalculate the parameter gradients in the model and update the model parameters. When calculating the gradients, the gradients calculated for each sample are multiplied by their corresponding sample weights , so that the model parameters are more biased towards difficult samples. Thereby enhancing the generalization performance of the model and enhancing the accuracy of the model in actual prediction; S4-4: Model Training: Use the complete dataset constructed in S1 and combine it with the model training strategies specified in S4-2 and S4-3 to train the model. During the training process, use the Stochastic Gradient Descent optimizer to update the model parameters. Continuously repeat the above training process until the model training is completed, and obtain the trained enterprise complex decision-making assistance large model.

[0059] As a refinement of the above embodiment, S4 constructs and actually deploys the model based on the input data augmentation strategy of historical data: Deploy the server system required for the operation of the large model within the enterprise, deploy the trained enterprise complex decision-making assistance large model to the server, and adjust and test the software system to ensure that the model can run normally. The response plan obtained from the large model can help relevant personnel in the enterprise carry out enterprise operations efficiently.

[0060] S5-1 Hardware Deployment: To ensure that the model can work properly, deploy a server with a high-performance CPU, a large-capacity memory, and a high-speed storage device in the enterprise to ensure that the server can run this enterprise complex decision-making assistance large model efficiently and stably; S5-2 Model Deployment: Deploy the trained enterprise complex decision-making assistance large model to the enterprise's terminal server system, set the input and output data interfaces of the model, and dock with the existing business systems of the enterprise. After the deployment is completed, debug and test the model to ensure that the model can make predictions normally; S5-3 Historical Information Retrieval: Use the historical data augmentation strategy proposed in S3 to retrieve data on various issues related to enterprise development, including data in two modalities: text data and time series data, to obtain decision-making plans for similar historical events; S5-4 Input Data Augmentation: Combine the decision-making plans of similar historical events retrieved in S5-3 with the current text description text data to obtain enhanced text data, and use this enhanced text data and time series data as the new model input data for the large model; S5-5 Decision-making Obtaining and Use: Send the data enhanced in S5-4 into the enterprise complex decision-making assistance large model to obtain the output result of the large model, and use the predicted enterprise development strategy suggestions to provide forward-looking information based on data and model analysis for enterprise development, helping enterprise managers to have a more comprehensive perspective when formulating strategic decisions.

[0061] In this embodiment, a data experience pool for storing historical Q&A data is constructed, and an input data similarity calculation method is designed. When the input question description data and text data are the same as the historical query data saved in the experience pool, the answer data of the historical query will be used as the input result of the large model together, so as to further improve the integrity of the input data and achieve the purpose of optimizing the model performance.

[0062] Specifically, it includes: In this embodiment, when the large model is actually used, its problem description data and the corresponding large model answer data will be stored in the historical data experience pool. Based on the data stored in the experience pool, when the new problem description data has a high similarity with the problem description data existing in the experience pool, at this time, the historical answer data of the large model will be automatically used as the input text data of the large model, so as to prompt the generation of this model according to the historical data and further enhance the context understanding ability of the model. The specific description is as follows: S5-4-1 Data experience pool construction: Each storage content in the data experience pool contains two core fields, and ; among them, is the text and time series data input by the original user. is the output result of the large model for this problem. In addition to the above core fields, a data identifier is added to each storage content for data indexing, and a timestamp is also added to record the generation time of the data. All storage contents in the experience pool are recorded in a database based on FAISS.

[0063] S5-4-2 Clustering of historical data based on text questions: To calculate the similarity between the newly input new problem description data and the data stored in the experience pool, the present invention designs a similarity calculation unit based on the K-means clustering algorithm and the Euclidean distance. First, the newly asked new problem description data and the problem description data existing in the experience pool are mixed. Since there is text data and time series data in the problem description data, the Word2Vec word embedding layer is used to calculate the text data to obtain its corresponding word embedding vector . Based on the word embedding vectors of each data, the K-means clustering algorithm is designed to divide all the data in the experience pool into X data clusters to complete the preliminary data classification. And record the data cluster where the new problem description data is located. Then calculate the cosine similarity between the word embedding vectors of the text data in the data cluster where the new problem description data is located and the word embedding vectors of other text data in the data cluster and the Euclidean distance between the time series data .

[0064] , , Among them, represents the word embedding vector of the new problem description data text data, represents the word embedding vector of other text data in the data cluster, represents the dot product calculation function, represents the norm calculation function, represents the th element in the time series data of the new problem description data, represents the th element in the time series data of other data in the data cluster, , represents the normalization operation function.

[0065] Calculate the feature similarity score based on the calculated cosine similarity and Euclidean distance : .

[0066] S5-4-3 Input data enhancement implementation: Based on the calculated new input problem description data and the similarity score of each problem description data in the experience pool, set the adjustable range of the threshold to (1.5 - 1.75). The larger the threshold is set, the higher the standard for similarity determination is, indicating that only the new problem description text with a higher data similarity to the experience pool data at this time can perform the input data enhancement operation. For the data with a similarity score greater than the threshold, add the corresponding large model response data to the problem description text to enhance the data comprehensiveness of the input data and optimize the overall performance of the large model; S5-4-4 Experience pool dynamic update: When a new question and answer is generated each time, save the data ( , ) into the experience pool automatically. In the present invention, a time decay strategy is used to delete the data in the experience pool that exceeds the save time threshold based on the timestamp Tic information contained in each data; specifically, the adjustable range of the save time threshold is (10 - 14 months). According to the timestamp information of each data in the experience pool, the data whose existence time exceeds this time threshold will be automatically deleted. Therefore, the larger the time threshold is, the slower the update frequency of the experience pool data is, and vice versa, the faster the update frequency of the experience pool data is.

[0067] Model performance comparison and testing: The proposed algorithm of the invention was compared and tested with two existing enterprise complex decision-making assistance algorithms, Qyt and Thg. When conducting the test, two different performance indicators of the model were evaluated. The first indicator was the decision accuracy rate, which was calculated as the number of correct decisions given by the large model divided by the total number of decisions; the other performance indicator was the prediction accuracy rate, which was calculated as the number of correct predictions divided by the total number of predictions when the large model was performing prediction tasks.

[0068] When conducting the prediction accuracy rate test of the large model, it was tested for 4 different prediction problems, and the results are as Figure 6 shown. The method proposed by the present invention has the highest prediction accuracy rate in all 4 prediction tasks.

[0069] When conducting the decision accuracy rate test of the large model, it was tested for 4 different decision problems, and the results are as Figure 7 shown. The method proposed by the present invention shows the highest decision accuracy rate in all 4 decision tasks.

[0070] At the same time, for the historical data input enhancement strategy method proposed by the present invention, to prove the effectiveness of this strategy, the performance of the model without using this strategy and with using this strategy was compared. The models with and without this strategy were respectively tested in four aspects: market prediction, market decision-making, financial prediction, and financial decision-making, as Figure 8 shown. The results show that when this strategy is not used in the model, the overall prediction accuracy of the model will decline to varying degrees.

[0071] An enterprise complex decision-making assistance system based on a multi-modal cognitive intelligence large model, comprising: Data collection and processing module: used to collect multi-modal enterprise complex decision-making assistance data, perform data cleaning and data annotation to obtain a training data set; Model architecture design module: used to construct an enterprise complex decision-making assistance large model including a word embedding layer, a text data encoder module, a time series data encoder module, a feature alignment unit, and a decoder module; Model training module: adopts a cross-entropy loss function, dynamic learning rate adjustment, and sample weight self-adaptive method to optimize the training of the large model to obtain a trained enterprise complex decision-making assistance large model; Model deployment and application module: deploys the trained large model to a server, and generates decision-making suggestions in combination with the input data enhancement strategy; Input data enhancement module: stores the historical Q&A data of the large model through a data experience pool, calculates the similarity between the new question and the historical data, and uses the similar historical answers as input to enhance the text data.

[0072] Finally, it should be noted that the above are only preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A complex enterprise decision-making assistance method based on a multimodal cognitive intelligence model, characterized in that: include: Collect multi-modal enterprise complex decision-making auxiliary data, perform data cleaning and data labeling, and obtain training data sets. Each set of data in the training data set includes problem description data and corresponding analysis results and solution labels, wherein the problem description data includes text data and time series data; Construct a large enterprise complex decision-making support model including a word embedding layer, a text data encoder module, a time series data encoder module, a feature alignment unit, and a decoder module; The training data set is input into the enterprise complex decision-making support large model, and the cross entropy loss function, dynamic learning rate adjustment and sample weight adaptation method are used to optimize the large model training to obtain the trained enterprise complex decision-making support large model; The trained large model is deployed to the server, and decision recommendations are generated in combination with the input data enhancement strategy; the input data enhancement strategy is as follows: the historical question and answer data of the large model is stored in the data experience pool, the similarity between new questions and historical data is calculated, and similar historical answers are used as input enhanced text data.

2. According to claim 1, the enterprise complex decision-making assistance method based on multimodal cognitive intelligence big model is characterized by: The data processing method of the enterprise complex decision-making auxiliary large model is as follows: The feature vector of the input text data is calculated through the word embedding layer and the learnable parameter matrix; a text data encoder with a stacked structure is used to perform deep feature extraction on the feature vector; the output of the upper encoder is used as the input of the lower encoder, and the number of the text data encoders is indivual; A stacked structured time series data encoder directly extracts features from the input time series data, wherein the number of stacked layers of the time series data encoder is consistent with that of the text data encoder; the output of the upper encoder of the time series data encoder is used as the input of the lower encoder; The output features of the text data encoder module and the time series data encoder module of the same layer are used as the input of the feature alignment unit of this layer, and the number of the feature alignment units is consistent with that of the text data encoder; The output of the feature alignment unit is used as the decoder input, and the number of decoders is consistent with the text data encoder. The input of the layer decoder includes Layer decoder output and Layer feature alignment unit output; No. The output of the layer decoder is processed by the softmax function to obtain the prediction result of the large model.

3. The enterprise complex decision-making assistance method based on multimodal cognitive intelligence big model according to claim 2 is characterized in that: The input text data is calculated through the word embedding layer and the learnable parameter matrix to obtain the feature vector of the text data, specifically in the following way: , , in, word embedding vectors representing text data, Represents text data, represents the position vector of the text data, express Embedding layer operation function, A feature vector representing text data; The position vector Through the learnable parameter matrix Implementation, the size of the matrix is , Indicates the number of characters contained in the input text data. Represents the dimension of the position vector; the position vector of each word in the input text data is the parameter matrix The index corresponding to this word in is A row vector of .

4. The enterprise complex decision-making assistance method based on multimodal cognitive intelligence big model according to claim 1 or 2 is characterized in that: The text data encoder module includes: two fully connected layers, a Relu activation function, a multi-head attention mechanism, a BN normalization layer, a Dropout layer and a Simgoid activation function; Feature vector of text data After passing through two fully connected layers in sequence, the Relu activation function is used to activate the features to obtain preliminary features. ; The initial features It is sent to the multi-head attention mechanism for processing. The specific method is as follows: The initial features Mapping to three different matrix feature spaces , , in, , , Respectively , , Matrix calculation functions, , , Respectively represent the calculated Attention , , Matrix, the total number of heads included in the multi-head attention mechanism is , , Calculate the single attention head feature output factor : in, represents the softmax function, is the normalization factor, Representation Matrix The transpose of . Based on the calculated feature output factor of each attention head Calculate the feature output of each attention head , the formula is as follows: , The feature output of each attention head in the multi-head attention mechanism is obtained through the fully connected layer Integrate and get feature output ; Output the features The features are processed sequentially through the BN normalization layer, the Dropout layer and the Simgoid activation function. Further processing is performed to obtain the feature vector processed by the encoder unit .

5. The enterprise complex decision-making assistance method based on multimodal cognitive intelligence big model according to claim 4 is characterized in that: The time series data encoder module includes: three LSTM layers, a multi-head attention mechanism, two LSTM layers and a Relu activation function; The input time series data is directly extracted through three LSTM layers in sequence to obtain multi-level features of the time series data. ; Through the multi-head attention mechanism Processed features ; Then pass two LSTM layers to the features Feature processing and extraction, after Relu activation function, complete the construction of the time series data encoder module, and obtain the features output by the time series data encoder .

6. The enterprise complex decision-making assistance method based on multimodal cognitive intelligence big model according to claim 5 is characterized in that: The feature alignment unit feature processing method is as follows: The characteristics , Use the BN normalization layer to normalize the features. The activation function activates the features. When aligning the features, an adaptive adjustment scheme is used to determine the number of dimensions after feature alignment. : in, Represents the input feature vector The dimension of Represents the input feature vector The dimension of is the dimension adjustment parameter, Represents the floor operation function ; Get the number of feature dimensions after alignment , and then a fully connected layer is used to adjust the dimension mapping so that the dimensions of the two features are unified to ; A fully connected layer is used to fuse the data features of the two modalities to complete feature alignment.

7. The enterprise complex decision-making assistance method based on multimodal cognitive intelligence big model according to claim 6 is characterized in that: The decoder includes a feedforward neural network, a multi-head attention mechanism, a BN normalization layer, a Dropout layer, and a BN normalization layer and a Dropout layer; the feedforward neural network includes two fully connected layers and a mish activation function.

8. The enterprise complex decision-making assistance method based on multimodal cognitive intelligence big model according to claim 1 is characterized in that: The specific method of the input data enhancement strategy is as follows: Build a data experience pool and store the original user input data in the data experience pool , the corresponding output results of the large model ,data identifier and timestamp, the original user input data is the problem description data input into the large model; Describe the new problem data and the original problem description data in the experience pool Mixed, the problem description data includes text data and time series data , using the Word2Vec word embedding layer to embed text data Calculate and get the corresponding word embedding vector Based on the word embedding vector of each data, the K-means clustering algorithm is used to divide all the data in the experience pool into X data clusters to complete the preliminary data classification, and the new problem description data is recorded The data cluster in which the new problem description data is calculated In the data cluster The cosine similarity between the word embedding vector of the text data and the word embedding vectors of other text data in the data cluster as well as The Euclidean distance between the time series data and other time series data in the data cluster , the formula is as follows: , , in, Represents the word embedding vector of the new problem description data text data, Represents the word embedding vector of other text data in the data cluster, represents the dot product calculation function, represents the modulus length calculation function, Indicates the first elements, Indicates the first elements, , represents the normalization function; The feature similarity score is obtained by adding the calculated cosine similarity and Euclidean distance. ; Set a threshold based on the similarity score between the newly input question description data and each question description data in the experience pool. For data with similarity scores greater than the threshold, the corresponding large model responds to the data Add to the problem description text; Every time a new question is generated, the data ( , ) is automatically stored in the experience pool, and according to the time decay strategy, the data in the experience pool that exceeds the storage time threshold is deleted based on the timestamp information contained in each data.

9. The enterprise complex decision-making assistance method based on multimodal cognitive intelligence big model according to claim 1 is characterized in that: The loss function of the enterprise complex decision-making auxiliary model is as follows: , in, represents the sample size, represents the sequence length, Indicates that based on The first Data and The first time series data sample The first Strategy plan No. The conditional probability of data; The dynamic learning rate adjustment is as follows: In the initial training phase of the model, the learning rate The initial value is set high, and as the training progresses, the loss value of the model after each training is calculated. , and use this loss value to dynamically optimize and update the learning rate: , in, represents the learning rate after each update, represents the learning rate decay factor, represents the loss function value factor, represents Euler's constant; The specific method of sample weight adaptation is as follows: Before training begins, each group of samples is assigned a value of The weight of is the number of training sample groups in the dataset. After the first round of training, the sample weights are adjusted according to the loss value of each group of samples: , in, represents the weight adjustment factor, represents the sample weight after each update, Represents the initial sample weight; after the weight is updated, the sample with the updated weight is used to recalculate the parameter gradient in the model and update the model parameters. When calculating the gradient, the gradient calculated for each sample is multiplied by its corresponding sample weight .

10. An enterprise complex decision support system based on a multimodal cognitive intelligence model, characterized in that: include: Data collection and processing module: used to collect multi-modal enterprise complex decision-making auxiliary data, perform data cleaning and data labeling, and obtain training data sets; Model architecture design module: used to build a large enterprise complex decision-making support model including a word embedding layer, a text data encoder module, a time series data encoder module, a feature alignment unit, and a decoder module; Model training module: Use cross entropy loss function, dynamic learning rate adjustment and sample weight adaptive method to optimize large model training and obtain a trained enterprise complex decision-making auxiliary large model; Model deployment and application module: deploy the trained large model to the server and generate decision recommendations based on input data enhancement strategies; Input data enhancement module: Store the historical question and answer data of the large model through the data experience pool, calculate the similarity between new questions and historical data, and use similar historical answers as input to enhance text data.