Software system development method based on multi-modal AI large model

CN120029605AInactive Publication Date: 2025-05-23四川参盘供应链科技有限公司

Patent Information

Application Number
CN202510502727.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-05-23
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional software system development methods have shortcomings in processing multimodal data and realizing intelligence, and cannot effectively integrate and utilize multiple modal data, such as images, voice and video, resulting in low data utilization.

Method used

Using a software system development method based on multimodal AI big model, multimodal AI big model with convolutional neural network, recurrent neural network and Transformer architecture is used to uniformly characterize and construct multimodal data, feature extraction and convolution operations are performed to realize the fusion of multimodal data, and jointly fine-tune the multimodal AI big model.

Benefits of technology

It improves the degree of integration of multimodal data, improves the intelligence level of software systems, and can make more efficient use of multimodal data, enhancing data analysis and decision-making support capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029605A_ABST
    Figure CN120029605A_ABST
Patent Text Reader

Abstract

The invention relates to a software system development method based on a multi-modal AI large model, which belongs to the technical field of multi-source information integration and comprises the following steps: S1, constructing the multi-modal AI large model; s2, multi-modal data unified representation construction: constructing a heterogeneous data fusion engine, and performing space-time alignment processing on the fused heterogeneous data to generate a multi-modal tensor sequence; s3, software system function development: based on the multi-modal AI large model, developing an analysis function of the software system; s4, establishing a mixed training optimization mechanism: carrying out combined fine tuning on the multi-modal AI large model; s5, establishing a credibility guarantee system: generating a credibility evaluation report of the multi-modal AI large model; the method has the advantages that feature extraction is carried out on the fused tensor sequence, convolution operation is carried out on the tensor sequence through the multi-dimensional convolution kernel so as to extract the combined feature of image and text data, combined fine adjustment is carried out on the multi-modal AI large model, the fusion degree of the multi-modal data is improved, and the intelligent level of a software system is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of multi-source information integration, and in particular relates to a software system development method based on a multimodal AI large model. Background Art

[0002] In today's digital age, software systems have been widely used in various fields, from applications in daily life to complex industrial control systems. The functional requirements of software systems are rapidly advancing towards high intelligence, high diversity, and the ability to process complex multi-source data. However, traditional software system development methods have many limitations that are difficult to overcome when dealing with these new requirements.

[0003] In terms of data processing, traditional software system development methods have extremely limited capabilities for fusion processing of multi-source heterogeneous data. With the popularization of technologies such as the Internet of Things and big data, the data types that software systems need to process are no longer limited to single text or numerical values. For example, a large number of multi-modal data such as images, voice, and video have emerged. The massive emergence of these multi-modal data has caused traditional software system development methods to lack an effective multi-modal data fusion mechanism. They often process different modal data separately and cannot associate or fuse different modal data, resulting in low data utilization.

[0004] The existing software system development methods can no longer adapt to the development trend of multimodal data fusion. Therefore, there is an urgent need for a new development method that can make full use of the rich information of multimodal data and improve the intelligence level of software systems. Summary of the invention

[0005] The present invention mainly solves the shortcomings of existing software system development methods in processing multimodal data and realizing intelligence. The present invention provides a software system development method based on a multimodal AI big model, performs feature extraction on the fused tensor sequence, uses a multidimensional convolution kernel to perform a convolution operation on the tensor sequence to extract the joint features of image and text data, and jointly fine-tunes the multimodal AI big model, thereby improving the fusion degree of multimodal data and enhancing the intelligence level of the software system.

[0006] In order to achieve the above object, the present invention is implemented by the following technical solutions:

[0007] A software system development method based on a multimodal AI large model includes the following steps:

[0008] S1. Build a multimodal AI big model: Build a multimodal AI big model with a convolutional neural network (CNN) for image data processing, a recurrent neural network (RNN) for speech data processing, and a Transformer architecture for text data processing;

[0009] S2. Construction of unified representation of multimodal data: Build a heterogeneous data fusion engine, perform spatiotemporal alignment on the fused heterogeneous data, and generate a multimodal tensor sequence;

[0010] S3. Development of software system functions: Based on the multimodal AI big model, develop the analysis function of the software system; among them, develop the data analysis function of the software system, conduct comprehensive analysis of the multimodal AI big model, mine potential data information, and provide support for the decision-making of the software system;

[0011] S4. Establish a hybrid training optimization mechanism: jointly fine-tune multimodal AI models;

[0012] S5. Establish a trusted assurance system: Generate a credibility assessment report for large multimodal AI models.

[0013] Optionally, in step S1, features are extracted by performing convolution operation on the convolution kernel and the input image. Assume that the input image is , the convolution kernel is , the output feature map is , for a two-dimensional image, the formula for the convolution operation is:

[0014] ;

[0015] in, are the coordinates in the output feature map, are the coordinates in the convolution kernel, and are the height and width of the convolution kernel respectively.

[0016] Optionally, in step S1, the RNN processes continuous speech input, and the loop is calculated as follows: for each time step , perform the following calculations: Calculate the current time step The hidden state , by inputting and the hidden state at the previous time step Perform linear transformation and pass activation function get ,in, and is the weight matrix, It is the bias vector and can also calculate the output as needed ;

[0017] Choose a suitable loss function according to the specific task. Taking classification task as an example, the cross entropy loss function is commonly used:

[0018] ;

[0019] in, is the sample size; is the number of categories; It is a sample Belongs to category The true label of is the model prediction sample Belongs to category probability;

[0020] For backpropagation: calculate the gradient through the time backpropagation algorithm (BPTT) to update the weight matrix and the bias vector .

[0021] Optionally, in step S1, TF calculation: Chinese The number of occurrences is ,document The total number of words is , then the word In the documentation The word frequency in for: ;in, Words In the documentation The number of times it appears in For Documentation Total number of words;

[0022] IDF Calculation: Document Collection There are documents containing the word The number of documents is , then the word Inverse document frequency for: ;

[0023] TF-IDF value: word In the documentation The TF-IDF value in is the product of term frequency and inverse document frequency. .

[0024] Optionally, in step S2, the heterogeneous data fusion engine uses the CRF modeling sequence labeling method, the goal is to When , predict the optimal label sequence ,in, Indicates Entity labels for locations;

[0025] Entity fusion is specifically a rule-based weighted fusion: the core of entity fusion is to calculate the and Similarity , determine whether they are the same entity;

[0026] The similarities of multiple attributes are weighted and combined to obtain the comprehensive similarity: ;in, It is similarity of attributes (e.g. document name similarity), is the total number of attribute similarities; is the attribute weight, satisfying , usually determined through manual experience or machine learning.

[0027] Optionally, in step S3, the development of the analysis function mainly adopts the tensor product fusion algorithm to analyze the multimodal fusion: the data of different modes are represented as tensors, and the image data is represented as a three-dimensional tensor, where the dimensions correspond to the height, width and number of channels of the image respectively; the text data will be represented as a two-dimensional tensor, where one dimension represents the length of the text and the other dimension represents the dimension of the word vector. There are two modes of data, represented as tensors respectively. and , The shape is , The shape is ;

[0028] Tensor product operation: Perform a tensor product operation on two tensors. The operation rule of the tensor product is that for two tensors and , their tensor product Elements ,in, is the tensor product Each element of the matrix in For tensors Each element of the matrix in For tensors The tensor product of the elements of the matrix The shape will become , is the size of the first dimension, is the size of the second dimension, is the size of the third dimension, is the size of the fourth dimension, is the size of the fifth dimension.

[0029] Optional, for 4D tensors: is a The matrix of is a The tensor product of is a The matrix, tensor product of Each element in The calculation formula is:

[0030] ;

[0031] Among them, i=1, 2,…m; j=1, 2,…n; k=1, 2,…,p; o=1, 2,…,q.

[0032] Optionally, for multidimensional tensors, assume is a The tensor of is a tensor, then their tensor product is a A multidimensional tensor is a first-order matrix.

[0033] Optionally, in step S4, the multimodal AI large model is jointly fine-tuned, and the steps adopted are: determining the fine-tuning target, setting the fine-tuning parameters, selecting the optimizer, and back-propagation and parameter updating.

[0034] Optionally, in step S5, the credibility assessment report of the multimodal AI large model generates intra-modal consistency and intra-modal alignment.

[0035] Beneficial effects of the present invention:

[0036] The present invention obtains respective extracted features through a convolutional neural network (CNN) for image data processing, a recurrent neural network (RNN) for speech data processing, and a Transformer architecture for text data processing, constructs a heterogeneous data fusion engine with respective extracted features through a unified representation of multimodal data, performs spatiotemporal alignment processing on the fused heterogeneous data, generates a multimodal tensor sequence, performs feature extraction on the fused tensor sequence, performs convolution operation on the tensor sequence using a multidimensional convolution kernel to extract joint features of image and text data, and jointly fine-tunes the multimodal AI large model, thereby improving the fusion degree of multimodal data and enhancing the intelligence level of the software system. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0038] Figure 1 It is a schematic diagram of the system structure of the present invention;

[0039] Figure 2 It is a schematic diagram of the workflow of the present invention. DETAILED DESCRIPTION

[0040] The embodiments of the present application are described in detail below with reference to the accompanying drawings.

[0041] like Figure 1 As shown, a software system based on a multimodal AI big model is provided. Usually, the software system of the multimodal AI big model includes a data layer, a model layer, an application layer and a control layer.

[0042] The role of the data layer: Responsible for collecting, organizing and preprocessing multimodal data, including various types of data such as images, text and audio. These data are the basis for model training. The quality and diversity of the data directly affect the performance and generalization ability of the model.

[0043] Example: In an intelligent security software system based on a multimodal AI large model, the data layer collects video image data captured by surveillance cameras, text data of card swipe records from the access control system, and audio data from alarm devices.

[0044] The role of the model layer: It is the core part of the software system, including a large multimodal AI model. The model conducts large-scale training on the data provided by the data layer, learns the association and feature representation between different modal data, and can understand and generate multimodal information to achieve various intelligent tasks, such as image recognition, speech recognition, and natural language processing.

[0045] Example: Take OpenAI’s GPT-4 model as an example. After being trained with a large amount of multimodal data of text and images, it can generate related images based on the input text description, or provide a detailed text description of a given image.

[0046] Application layer role: Apply the capabilities of the model layer to specific business scenarios to provide users with various services and functions. By calling the data of the model layer, multi-modal interactive applications such as intelligent customer service, intelligent security monitoring, and intelligent driving assistance can be realized.

[0047] Example: In an intelligent customer service scenario, the application layer receives text or voice questions input by users, calls the relevant models of the model layer to understand and analyze them, and then generates corresponding answers and feeds them back to the user in the form of text or voice.

[0048] Function of the control layer: coordinate and control the entire software system, including model training, scheduling of the inference process, and communication and collaboration between modules. According to the system's operating status and user needs, resources are reasonably allocated to ensure efficient and stable operation of the system.

[0049] Example: During the training of a large multimodal AI model, the control layer will automatically allocate resources based on the size of the data and the complexity of the model, such as deciding how many GPUs to use to improve training efficiency. At the same time, during the model inference phase, the control layer will reasonably arrange the model calling sequence based on the priority of user requests and the occupancy of system resources to ensure timely response to user requests.

[0050] like Figure 2 As shown, a software system based on a multimodal AI big model, the present invention provides a software system development method based on a multimodal AI big model, comprising the following steps:

[0051] S1. Build a multimodal AI big model: Build a multimodal AI big model with a convolutional neural network (CNN) for image data processing, a recurrent neural network (RNN) for speech data processing, and a Transformer architecture for text data processing;

[0052] Regarding CNN, in the convolution layer, features are extracted by convolution operation between the convolution kernel and the input image. Assume that the input image is , the convolution kernel is , the output feature map is , for a two-dimensional image, the formula for the convolution operation is:

[0053] ;

[0054] in, are the coordinates in the output feature map, are the coordinates in the convolution kernel, and are the height and width of the convolution kernel respectively;

[0055] Specifically, assuming the input image is a The matrix, convolution kernel is a The matrix of the output feature map is is a For the first matrix on the output feature map elements, and the calculation formula is:

[0056] ;

[0057] Define a convolution function to implement the convolution operation. The convolution function accepts the input image and convolution kernel parameters and returns the result of the convolution operation. The input image and convolution kernel are tested to obtain the result of the convolution operation.

[0058] The RNN algorithm is as follows:

[0059] RNN processes continuous speech input and realizes continuous speech recognition. Specifically, in the recurrent neural network for speech data processing, the input and initialization are as follows: the speech data (usually a vector sequence after feature extraction) is input into the RNN in sequence. At each time step In this case, the input vector is , and initialize the hidden state to , usually initialize the hidden state Set to an all-zero vector.

[0060] The loop is calculated as: For each time step , perform the following calculations: Calculate the current time step The hidden state of , by inputting and the hidden state at the previous time step Perform linear transformation and pass activation function get ,in, and is the weight matrix, is the bias vector. You can also calculate the output as needed .

[0061] The specific loop hidden state is calculated as: ;in: is the current time step The hidden state of is the current time step The input vector of is the previous time step The hidden state of is the weight matrix input to the hidden layer; is the weight matrix from hidden layer to hidden layer; is the bias vector of the hidden layer; It is an activation function. Commonly used activation functions are Sigmoid function or Tanh function.

[0062] Cycle output calculation as needed ,Right now: ,in, and is the weight matrix for the output and the bias vector Specifically, is the current time step The output vector of is the weight matrix from the hidden layer to the output layer; is the bias vector of the output layer.

[0063] Loss calculation and back propagation: Calculate the loss function according to the task objective (such as cross entropy loss for classification tasks, mean square error loss for regression tasks, etc.) , and then the gradient is calculated and the parameters of the network are updated through the back propagation through time algorithm (BPTT). In BPTT, the error is back propagated along the time series to calculate the gradient of each time step.

[0064] Choose a suitable loss function according to the specific task. Taking classification task as an example, the cross entropy loss function is commonly used:

[0065] ;

[0066] in, is the sample size; is the number of categories; It is a sample Belongs to category The true label of (the true label is 0 or 1); is the model prediction sample Belongs to category probability.

[0067] For backpropagation: calculate the gradient through the time backpropagation algorithm (BPTT) to update the weight matrix and the bias vector , for the weight matrix and the bias vector , their respective gradients are calculated as follows:

[0068] ; ;

[0069] in, is the length of the time series, is the current time step, and the gradient descent optimization algorithm is used to update the parameters:

[0070] ; ;

[0071] in, is the learning rate, which is used to control the step size of parameter updates.

[0072] RNN can process data efficiently. It can read the sequence element by element and update the hidden state based on the current input and the previous hidden state to capture the long-term dependencies in the sequence. The hidden state in RNN can be seen as a summary of the network's past information. At each time step, RNN calculates a new hidden state based on the current input and the hidden state of the previous time step. The hidden state can accumulate and transmit information in the sequence, allowing RNN to model the historical information of the sequence. The cyclic structure of RNN is an important feature that distinguishes it from other neural networks. The cyclic connection allows information to propagate cyclically in the network, that is, the hidden state of the current time step depends not only on the current input, but also on the hidden state of the previous time step. This cyclic mechanism enables RNN to process sequence data with long-term dependencies because information can be transmitted and accumulated between multiple time steps. During the training process, RNN uses the back-propagation algorithm to calculate the gradient of the loss function with respect to the network parameters, and updates the parameters based on the gradient. Through a large amount of training data, RNN can automatically discover various patterns in the sequence. Once the training is completed, RNN can perform tasks such as prediction, classification or generation of new and unseen sequence data. The trained RNN can generate subsequent text content based on a given partial text, or perform speech recognition on a speech signal and convert it into text.

[0073] About the TF-IDF algorithm of Transformer architecture:

[0074] Term frequency (TF): refers to the frequency of a word appearing in a document. It is calculated by dividing the number of times the word appears in the document by the total number of words in the document. The higher the term frequency, the more important the word is in the document.

[0075] Inverse Document Frequency (IDF): It measures the rarity of a word in the entire document collection. It is calculated by dividing the total number of documents by the number of documents containing the word, and then taking the logarithm. If a word appears in many documents, its inverse document frequency is low, indicating that it is a common word; conversely, if a word appears in only a few documents, its inverse document frequency is high, indicating that it is a discriminative word.

[0076] TF calculation: Assume that in a document Chinese The number of occurrences is ,document The total number of words is , then the word In the documentation The word frequency in for: ;

[0077] Words In the documentation The number of times it appears in For Documentation The total number of words.

[0078] Example: Documentation : "the cat sat on the mat, the cat was happy", total number of words is 8, and the word "the" appears 2 times, then: .

[0079] IDF calculation: Assuming a document collection There are documents containing the word The number of documents is , then the word Inverse document frequency for: , the denominator is added by 1 to prevent the denominator from being 0, and the denominator is added by 1 to avoid =0 when infinity appears (i.e., words that have never appeared may be set to 0 in practice). is the total number of documents in the document collection; To contain words The number of documents that contain the term (at least once).

[0080] Example: Total number of documents =1000, the number of documents containing the word "the" ,but: .

[0081] TF-IDF value: word In the documentation The TF-IDF value in is the product of term frequency and inverse document frequency, that is, , .

[0082] By calculating the TF-IDF value of each word in the text, representative feature words are extracted to classify the text. The essence of TF-IDF is that the more frequently a word (i.e., the word "the") appears in a document and the rarer it is in the entire collection, the higher its importance. From the above examples, we can see that high-frequency common words (such as "the") are the reason why TF-IDF can effectively extract text features.

[0083] S2. Construction of unified representation of multimodal data: Build a heterogeneous data fusion engine, perform spatiotemporal alignment on the fused heterogeneous data, and generate a multimodal tensor sequence;

[0084] Entity recognition and fusion algorithms are used to build a heterogeneous data fusion engine, such as identifying and fusing information from different sources.

[0085] Entity recognition is specifically about entity recognition. It often uses CRF modeling sequence labeling. The goal is to When , predict the optimal label sequence ,in, Indicates The entity label of a location (such as "a certain document name", "a certain data image" or "a certain voice").

[0086] Conditional probability formula:

[0087] ;in, is the normalization factor, also called the partition function, which ensures that the sum of probabilities is 1; , where the sum is over all possible label sequences carried out.

[0088] is the transfer feature function, describing the label arrive The transfer relationship (e.g. the probability of a document name being followed by a suffix); is the index of the characteristic function, ranging from 1 to , Represents the number of transition feature functions, which is usually a binary function that takes the value 1 when a specific transition condition is met, otherwise it takes the value 0. In the part-of-speech tagging task, It can represent features that are transferred from one part-of-speech tag to another, such as "transferring from document name to data image".

[0089] is the state characteristic function, describing the position Tags With observed value (e.g., a suffix ending with "doc or docx" may be a document name);

[0090] is the index of the characteristic function, ranging from 1 to , Represents the number of state feature functions. Similarly, it is usually also a binary function. In the named entity recognition task, It can represent the suffix feature of the label "docx" when the current word ends with the suffix "doc".

[0091] is the weight of the feature function, obtained through maximum likelihood estimation training, The weight of the transfer feature function measures the importance of the corresponding transfer feature function. These weights are learned through training data, and the training goal is to maximize the log-likelihood of the training data.

[0092] Entity fusion is specifically a rule-based weighted fusion: the core of entity fusion is to calculate the and Similarity , determine whether they are the same entity.

[0093] The similarities of multiple attributes are weighted and combined to obtain the comprehensive similarity: ;in, It is similarity of attributes (e.g. document name similarity), is the total number of attribute similarities; is the attribute weight, satisfying , usually determined through manual experience or machine learning (e.g. logistic regression).

[0094] In addition, the spatiotemporal alignment process uses a spatiotemporal convolutional network (STCN), specifically: assuming that the input spatiotemporal data is a four-dimensional tensor ,in, Represents the time dimension (e.g., the number of frames in a video). For video data, it represents the number of frames in the video. A video clip contains 10 frames of images. =100, Indicates the number of channels (such as the RGB channels of an image). For color images, there are usually three RGB channels. =3; if it is a grayscale image, then =1, Represents the height in the spatial dimension, that is, the number of pixels in the vertical direction of an image or video frame. The height of an image is 480 pixels. =480, is the width in the spatial dimension, that is, the number of pixels in the horizontal direction of an image or video frame. The width of an image is 640 pixels, then =640.

[0095] For spatial convolution, we use a size of The spatial convolution kernel , is the size of the spatial convolution kernel, usually an odd number such as 3, 5, 7, etc. Represents the size of the convolution kernel in the spatial dimension. =3, the convolution kernel is a 3×3 matrix, which slides on the spatial dimension of the image to perform convolution operations, each time multiplying and summing with the 3×3 image area to extract spatial features, and performing convolution operations on the spatial dimension. The spatial convolution output at for:

[0096] ; Where t represents the time dimension, c represents the channel dimension, The spatial dimension convolution kernels, The spatial dimension convolution kernel.

[0097] t is the index of the time dimension, which specifies which frame of the input data in the time series is currently being processed.

[0098] c is the index of the channel dimension, which is used to distinguish different channels. For example, in an RGB image, c=1 represents the red channel, c=2 represents the green channel, and c=3 represents the blue channel.

[0099] is the index on the spatial dimension, Indicates the height position of an image or video frame. Indicates the position in the width direction. A specific pixel point in space can be determined. Indicates the time t, channel c, spatial position The output value after spatial convolution.

[0100] For temporal convolution, we use a size of Temporal convolution kernel , is the size of the temporal convolution kernel, indicating the number of frames covered by the convolution kernel in the time dimension, =5 means that the temporal convolution kernel will consider the information of the current frame and the two frames before and after (a total of 5 frames) to perform convolution operations, so as to capture the feature changes in the time series of the video and perform convolution operations in the time dimension. The temporal convolution output at for: ;

[0101] Here The meaning is the same as the corresponding parameter in the spatial convolution output. Indicates that at time t, channel c, spatial position The output value after time convolution at A temporal convolution kernel is obtained by further performing a convolution operation in the time dimension based on the result of the spatial convolution.

[0102] When combining spatial and temporal convolution, the results of spatial convolution and temporal convolution are combined. Assume that spatial convolution is performed first and then temporal convolution. , then the output of the spatiotemporal convolution is: ; It is the final output of combining spatial convolution and temporal convolution. It integrates the feature information in both spatial and temporal dimensions for subsequent activation function processing, pooling operations, or input of other network layers.

[0103] Multimodal tensor sequence: When multimodal data is organized in the form of tensors and arranged in a certain order, a multimodal tensor sequence is formed. For example, in video analysis, a video can be regarded as a sequence consisting of a series of image frames (each image frame is a tensor), and each time point may also have corresponding audio data (also represented in the form of tensors) and some text description information (such as subtitles, which can also be converted into tensor representation). These tensors of different modes are combined in chronological order to form a multimodal tensor sequence.

[0104] S3. Development of software system functions: Based on the multimodal AI big model, develop the analysis function of the software system; among them, develop the data analysis function of the software system, conduct comprehensive analysis of the multimodal AI big model, mine potential data information, and provide support for the decision-making of the software system;

[0105] The development of the analysis function is mainly to analyze multimodal fusion, using the tensor product fusion algorithm for analysis:

[0106] Data representation: Data of different modalities are represented as tensors. Image data is represented as a three-dimensional tensor, where the dimensions correspond to the height, width, and number of channels of the image. Text data is represented as a two-dimensional tensor, where one dimension represents the length of the text and the other dimension represents the dimension of the word vector. Suppose there are two modal data, represented as tensors and , The shape is , The shape is .

[0107] Tensor product operation: Perform a tensor product operation on two tensors. The rule for the tensor product operation is that for two tensors and , their tensor product The elements of ,in, is the tensor product Each element of the matrix in For tensors Each element of the matrix in For tensors The tensor product of the elements of the matrix The shape becomes .

[0108] is the size of the first dimension, indicating that in the first dimension, the tensor has elements or components, which can be imagined as arranged along the first dimension A "slice".

[0109] is the size of the second dimension, indicating that each "slice" is further divided into That is, for every element in the first dimension, there is There are related elements corresponding to each "slice", forming A "block".

[0110] is the size of the third dimension, indicating that in the third dimension, each The "block" contains smaller elements, that is, at each position defined by the first two dimensions, there is Elements are arranged to form a "small piece"; , and is a tensor Dimensions in .

[0111] is the size of the fourth dimension, indicating that in the fourth dimension, there are Different "layers", each layer contains all the elements determined by the first three dimensions, forming A "layer block".

[0112] is the size of the fifth dimension, and each "layer block" is further divided into "element blocks", that is, for each "layer block" determined by the first four dimensions, there are related "element blocks", forming "Element blocks". and is a tensor Dimensions in .

[0113] In general, Complete description of the tensor The structure of the tensor can be determined by the size of these five dimensions. The total number of elements in is Each element can be indexed in these five dimensions to determine the position of each element in the tensor, thus realizing the tensor product. Deformation .

[0114] Specifically, four-dimensional tensor (matrix): Let is a The matrix of is a The tensor product of is a The matrix, tensor product of Elements in The calculation formula is:

[0115] ;

[0116] Among them, i=1, 2,…m; j=1, 2,…n; k=1, 2,…,p; o=1, 2,…,q.

[0117] Specifically, In and is an indicator, Representation Matrix The rows in Representation Matrix Columns in . Representative Matrix Middle Line elements of the column. Similarly, In and is an indicator, Representation Matrix OK, Representation Matrix Columns. Representative Matrix Middle Line Elements of a column.

[0118] Represented by a two-dimensional tensor and After performing the tensor product operation, we get a four-dimensional tensor The elements in , which describe the four-dimensional tensor In a specific location The components at , in this way, the four-dimensional tensor can be fully represented All elements in .

[0119] For example: If , ,but ;

[0120] Further, based on the four-dimensional tensor , Multidimensional tensor (matrix): For a general multidimensional tensor, assume is a The tensor of is a tensor, then their tensor product is a dimensional tensor is a first-order matrix. First-order matrices are also called row vectors or column vectors. They do not have the concept of rows and columns, so they do not have inverse matrices. For first-order matrices, they are only one element and there is no inverse in the sense of multiplication, because any matrix multiplied with it will change its dimension. The concept of the inverse matrix of a first-order matrix does not exist, it depends on the rows, columns and dimensions of the matrix.

[0121] For example: for a three-dimensional tensor (Shape is ) and a two-dimensional tensor (Shape is ), their tensor product is a five-dimensional tensor with shape .

[0122] Through feature extraction, the fused tensor For feature extraction, convolutional neural networks (CNN) and recurrent neural networks (RNN) can be used to extract features from tensors. For example, a three-dimensional convolution kernel can be used to extract features from tensors. Convolution operations are performed to extract joint features of image and text data.

[0123] For the 3D convolution kernel tensor The convolution operation calculation process:

[0124] Assume the three-dimensional convolution kernel is , the size is , indicating the length of each of the three dimensions; the input tensor is , the size is , indicating the sizes of the three dimensions, the input tensor can be regarded as a three-dimensional space node containing multiple data points; the output tensor is .

[0125] In a 3D convolution, the convolution kernel slides over the three dimensions of the input tensor, and for the output tensor Each element in , and its calculation formula is:

[0126] ;

[0127] in, , and Represent the output tensors respectively Indexes in three dimensions, , and Represent the convolution kernels Indexes in three dimensions. The formula indicates that the output tensor is A certain position The value of is determined by the convolution kernel With the input tensor Obtained by weighted summing of the areas around the corresponding positions.

[0128] Triple loop in formula Traversing the convolution kernel In each cycle, the convolution kernel Elements in With the input tensor The corresponding element in Multiply.

[0129] Here , and Determine the input tensor The position corresponding to the convolution kernel element in is, correspond , correspond , correspond In this way, the convolution kernel slides over the input tensor and performs a weighted summation of the area around each position (mentioned above The calculation process) gets the output tensor The element value at the corresponding position in .

[0130] In actual calculations, the step size of the convolution needs to be considered and fill If the step size is not 1, the convolution kernel The distance moved each time when sliding is If there is padding , then padding will be done on the boundaries of the input tensor Layer data, usually filled with 0 or other specified values ​​to control the output tensor The size and shape of. At this time, the above formula , and The value range will depend on the step size and fill Make adjustments accordingly.

[0131] For step length , if the step length Greater than 1, the convolution kernel moves in the corresponding dimension each time it slides. units. This will cause the size of the output tensor to become smaller in the corresponding dimension because the convolution kernel skips some elements of the input tensor. For example, in a certain dimension, if the input tensor size is , the convolution kernel size is , the step length is , then the size of the output tensor in this dimension is .

[0132] For the filling , padding is adding extra layers at the boundaries of the input tensor. If the padding value is 0, then when calculating the output tensor, the convolution kernel can perform the convolution operation completely at the boundary of the input tensor, and the size of the input tensor becomes larger. Padding can control the size of the output tensor to keep a certain relationship with the size of the input tensor, or to meet specific computing requirements. For example, in some cases, through appropriate padding, the size of the output tensor can be made the same as the input tensor. If the padding amount in each dimension is , then the effective size of the input tensor in this dimension becomes , this new size needs to be taken into account when calculating the output tensor size.

[0133] Three-dimensional convolution slides the convolution kernel on the input tensor at a certain step size and performs weighted summation calculation. At the same time, it combines the padding operation to adjust the size of the input tensor to obtain the output tensor. It has a wide range of applications in processing three-dimensional data, such as three-dimensional images and video data.

[0134] S4. Establish a hybrid training optimization mechanism: jointly fine-tune multimodal AI models;

[0135] Specifically, determine the fine-tuning goal: clearly determine whether the model's classification performance is to be optimized. For example, in an image-text retrieval task, the goal may be to improve the retrieval accuracy; in an image description generation task, the goal may be to generate a more accurate and detailed text description.

[0136] Set fine-tuning parameters: including learning rate, batch size, and number of training rounds. Usually, the learning rate is much smaller than that of the initial pre-training to avoid excessive damage to the weights of the pre-trained model. The batch size should be set reasonably according to the hardware resources and the amount of data. A larger batch size can improve training efficiency, but may cause insufficient memory. The number of training rounds needs to be determined through experiments to avoid overfitting.

[0137] Choose an optimizer: Commonly used optimizers such as Adam or SGD are used for joint fine-tuning. The Adam optimizer usually performs better when processing multimodal data because it can adaptively adjust the learning rate.

[0138] During fine-tuning, data from different modalities need to be fused, which can be achieved in a variety of ways, such as early fusion, where image and text features are concatenated at the input layer, or late fusion, where features from different modalities are fused at a higher level in the model. In the VisualBERT model, image features and text features are fused through a specific module and then jointly learned in subsequent layers.

[0139] Back propagation and parameter update: In each training step, the gradient is calculated according to the cross entropy loss function, and the model parameters are updated through the back propagation algorithm. The choice of loss function depends on the specific task. For classification tasks, the cross entropy loss function is used. In multimodal tasks, the cross entropy loss function is combined to optimize the model performance.

[0140] S5. Establish a trusted assurance system: Generate a credibility assessment report for large multimodal AI models.

[0141] Intra-modal consistency refers to the accuracy of the content generated by the model in a single modality (such as text, images, audio) and the degree of match with real data, ensuring that the generated content meets expectations in a single modality without obvious errors or deviations.

[0142] 1. Evaluation indicators for text generation accuracy: BLEU (Bilingual Evaluation Understudy) is calculated by comparing the similarity between the generated text and the reference text based on n-gram (n consecutive words) matching. Range: 0~1 (the higher the score, the closer it is to the manual reference text). ROUGE (Recall-Oriented Understudy for Gisting Evaluation) is calculated by focusing on the recall rate (keyword coverage) of the generated text and the reference text.

[0143] 2. Evaluation indicators for image generation accuracy: FID (Frechet Inception Distance); Calculation method: Use the Inception-v3 model to extract the feature vectors of the generated image and the real image. Calculate the Frechet distance (the difference in mean and covariance) between the two feature distributions. Range: 0~∞ (the smaller the better, 0 means complete consistency). Calculation method of IS (Inception Score): Based on the prediction results of the image classification model, evaluate the diversity and clarity of the generated image. That is ,in, is the class distribution of all generated images.

[0144] 3. Evaluation indicators for audio generation accuracy: WER (Word Error Rate) calculation method: After converting the generated speech to text, count the ratio of inserted, deleted, and replaced erroneous words to the total number of words. That is, WER = (insertion number + deletion number + replacement number) / total number of words in the reference text; MOS (Mean Opinion Score) calculation method: manual scoring (1 to 5 points) to evaluate the naturalness and clarity of the speech.

[0145] Cross-modal alignment refers to the logical consistency and semantic matching between the cross-modal content (such as text-image, text-audio) generated by the model, ensuring that information from different modalities supports each other without contradiction or ambiguity.

[0146] 1. Evaluation indicators of text-image alignment: CLIP Score is the similarity between the text description and the generated image calculated using the CLIP model (pre-trained multimodal model). Calculation method: The text and image are respectively passed through the CLIP encoder to obtain feature vectors. The cosine similarity of the two is calculated (range: -1~1, the higher the better the match); Human Alignment Score is the matching degree between the manually annotated generated image and the text description (e.g. 1~5 points).

[0147] Example: 5 points: completely consistent with the description, details are accurate; 1 point: the description has nothing to do with the image.

[0148] 2. Text-audio alignment evaluation indicators: The audio-video synchronization error rate is calculated by detecting the time alignment error (in milliseconds) between the lip shape and action in the generated audio and video, that is, using an audio-video alignment algorithm (such as dynamic time warping, DTW). Semantic consistency scoring method: Manually determine whether the audio content (such as narration) is consistent with the text theme.

[0149] In addition, the specific credibility assessment report of the multimodal AI large model is shown in Table 1:

[0150] Table 1 Credibility evaluation report of multimodal AI large model

[0151]

[0152] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope of the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.

Claims

1. A software system development method based on a multimodal AI large model, characterized in that: The steps include: S1. Build a multimodal AI model: Build a multimodal AI model with a convolutional neural network for image data processing, a recurrent neural network for speech data processing, and a Transformer architecture for text data processing; S2. Construction of unified representation of multimodal data: Build a heterogeneous data fusion engine, perform spatiotemporal alignment on the fused heterogeneous data, and generate a multimodal tensor sequence; S3. Development of software system functions: Based on the multimodal AI big model, develop the analysis function of the software system; among them, develop the data analysis function of the software system, conduct comprehensive analysis of the multimodal AI big model, mine potential data information, and provide support for the decision-making of the software system; S4. Establish a hybrid training optimization mechanism: jointly fine-tune multimodal AI models; S5. Establish a trusted assurance system: Generate a credibility assessment report for large multimodal AI models.

2. According to claim 1, a software system development method based on a multimodal AI large model is characterized in that: In step S1, features are extracted by performing convolution operation on the convolution kernel and the input image. Assume that the input image is , the convolution kernel is , the output feature map is , for a two-dimensional image, the formula for the convolution operation is: ; in, are the coordinates in the output feature map, are the coordinates in the convolution kernel, and are the height and width of the convolution kernel respectively.

3. According to claim 1, a software system development method based on a multimodal AI large model is characterized in that: In step S1, RNN processes continuous speech input, and the loop calculation is: for each time step , perform the following calculations: Calculate the current time step The hidden state of , by inputting and the hidden state at the previous time step Perform linear transformation and pass through activation function get ,in, and is the weight matrix, It is the bias vector and can also calculate the output as needed ; Choose a suitable loss function according to the specific task. Taking classification task as an example, the cross entropy loss function is commonly used: ; in, is the sample size; is the number of categories; It is a sample Belongs to category The true label of is the model prediction sample Belongs to category probability; For back propagation: calculate the gradient through the time back propagation algorithm to update the weight matrix and the bias vector .

4. According to claim 1, a software system development method based on a multimodal AI large model is characterized in that: In step S1, TF calculation: Chinese The number of occurrences is ,document The total number of words is , then the word In the documentation The word frequency in for: ;in, Words In the documentation The number of times it appears in For Documentation Total number of words; IDF Calculation: Document Collection There are documents containing the word The number of documents is , then the word Inverse document frequency for: ; TF-IDF value: word In the documentation The TF-IDF value in is the product of term frequency and inverse document frequency. .

5. The software system development method based on a multimodal AI big model according to claim 1 is characterized in that: In step S2, the heterogeneous data fusion engine uses the CRF modeling sequence labeling method, the goal is to When , predict the optimal label sequence ,in, Indicates Entity labels for locations; Entity fusion is specifically a rule-based weighted fusion: the core of entity fusion is to calculate the and Similarity , determine whether they are the same entity; The similarities of multiple attributes are weighted and combined to obtain the comprehensive similarity: ;in, It is The similarity of the attributes, is the total number of attribute similarities; is the attribute weight, satisfying , determined by manual experience or machine learning.

6. The software system development method based on a multimodal AI big model according to claim 1 is characterized in that: In step S3, the development of the analysis function mainly adopts the tensor product fusion algorithm to analyze the multimodal fusion: the data of different modes are represented as tensors, and the image data is represented as a three-dimensional tensor, where the dimensions correspond to the height, width and number of channels of the image respectively; the text data will be represented as a two-dimensional tensor, where one dimension represents the length of the text and the other dimension represents the dimension of the word vector. There are two modes of data, which are represented as tensors respectively. and , The shape is , The shape is ; Tensor product operation: Perform a tensor product operation on two tensors. The operation rule of the tensor product is that for two tensors and , their tensor product Elements ,in, is the tensor product Each element of the matrix in For tensors Each element of the matrix in For tensors The tensor product of the elements of the matrix The shape will become , is the size of the first dimension, is the size of the second dimension, is the size of the third dimension, is the size of the fourth dimension, is the size of the fifth dimension.

7. The software system development method based on a multimodal AI big model according to claim 6 is characterized in that: For a 4D tensor: is a The matrix of is a The tensor product of is a The matrix, tensor product of Each element in The calculation formula is: ; Among them, i=1, 2,…m; j=1, 2,…n; k=1, 2,…,p; o=1, 2,…,q.

8. The software system development method based on a multimodal AI big model according to claim 6 is characterized in that: For multidimensional tensors, assuming is a The tensor of is a tensor, then their tensor product is a A multidimensional tensor is a first-order matrix.

9. The software system development method based on a multimodal AI big model according to claim 1, characterized in that: In step S4, the multimodal AI large model is jointly fine-tuned, and the steps adopted are: determining the fine-tuning target, setting the fine-tuning parameters, selecting the optimizer, and back-propagation and parameter updating.

10. The software system development method based on a multimodal AI big model according to claim 1, characterized in that: In step S5, the credibility assessment report of the multimodal AI large model generates intra-modal consistency and intra-modal alignment.

Citation Information

Patent Citations

  • Fever disease multi-mode feature fusion cognitive system, equipment, and storage medium

    CN113488183A

  • License classification method and system based on multi-modal feature fusion

    CN115115883A

  • Multi-hardware mixed large model training method, system and related device

    CN119539012A

  • Large model-based power industry data analysis method, apparatus and device, and storage medium

    CN119829988A

  • Real-time contextually aware artificial intelligence (AI) assistant system and a method for providing a contextualized response to a user using ai

    US20240412720A1

Cited By

  • Network security threat intelligent detection method based on big data analysis

    CN120498872A