Text generation method and device based on multi-modal data, equipment and medium
By extracting and fusing the features of multimodal data and time series data, generating target features and performing text mapping, the problem of inefficient processing of long-sequence multimodal time series data in the prior art is solved, and efficient computing resource utilization and high-quality text generation are achieved.
Patent Information
- Application Number
- CN202411915902.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-24
- Publication Date
- 2025-05-06
AI Technical Summary
The prior art is difficult to efficiently process long sequence multimodal time series data, resulting in huge computing resource consumption and low processing efficiency.
By extracting features of multimodal data and time series data and fusing them, target features are generated so as to generate target text information for describing multimodal data. Specific steps include obtaining multimodal data and time series data, extracting visual and time series features, fusing features and mapping them to text space through feature mappers.
It significantly reduces the consumption of computing resources, improves processing efficiency, can effectively deal with high computing complexity, and generates more accurate and detailed text descriptions, improving the quality of text generation.
Smart Images

Figure CN119940305A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a method and device for generating text based on multimodal data, an electronic device, and a storage medium. Background Art
[0002] Currently, Transformer (a deep learning model architecture) is often used as the mainstream architecture of large AI (Artificial Intelligence) models to build models for processing time series data. Time series data usually has complex time patterns, changing distributions, many confounding variables, and long-term dependencies, which may result in very long sequence lengths.
[0003] However, in practical applications, time series data is often combined with multimodal data (such as images, text, etc.). For example, multiple image data may have corresponding time frames to form a time series. This multimodal time series data not only needs to deal with time dependency, but also needs to consider the fusion of image features and other modal features at the same time.
[0004] Due to the quadratic computational complexity of the Transformer's self-attention mechanism, the computational efficiency is low when processing long sequences, especially when processing multimodal time series data. The sequence length may be very long, resulting in huge consumption of computing resources. It will also lead to inefficient processing of these data, making it difficult to efficiently process time series and multimodal data at the same time. Summary of the invention
[0005] The embodiment of the present application provides a text generation method based on multimodal data to solve the problem that time series and multimodal data are difficult to be processed simultaneously.
[0006] Correspondingly, an embodiment of the present application also provides a text generation device based on multimodal data, an electronic device and a storage medium to ensure the implementation and application of the above method.
[0007] In order to solve the above problems, the present application discloses a method for generating text based on multimodal data, the method comprising:
[0008] Acquire multimodal data and time series data corresponding to the multimodal data;
[0009] Extracting a first feature from the multimodal data and a second feature from the time series data;
[0010] fusing the first feature and the second feature to obtain a target feature;
[0011] According to the target features, target text information corresponding to the multimodal data is obtained.
[0012] Optionally, extracting the first feature from the multimodal data includes:
[0013] Build visual encoder and object detection models;
[0014] extracting visual features from the multimodal data using the visual encoder;
[0015] The target detection model is used to extract multi-scale features of the target object in the multimodal data at multiple scales and global features at the overall level;
[0016] The visual feature, the multi-scale feature and the global feature are integrated to obtain the first feature.
[0017] Optionally, the fusing the first feature and the second feature to obtain a target feature includes:
[0018] Build feature mapper;
[0019] Mapping the first feature and the second feature into a text space using the feature mapper to obtain a first text feature corresponding to the first feature and a second text feature corresponding to the second feature;
[0020] The first text feature and the second text feature are fused to obtain the target feature.
[0021] Optionally, the feature mapper has corresponding parameters, and the feature mapper is trained by the following steps:
[0022] Obtain multimodal sample data;
[0023] Extracting a first sample feature from the multimodal sample data;
[0024] Extracting a text sample entity from the first sample feature, and using the text sample entity as a reference point of the text space;
[0025] Mapping the first sample feature into the text space using the feature mapper to obtain a first target feature;
[0026] Calculate the loss value corresponding to the feature mapper according to the first target feature and the text sample entity;
[0027] The parameters corresponding to the feature mapper are adjusted according to the loss value.
[0028] Optionally, extracting a second feature in the time series includes:
[0029] Performing minimum scaling and maximum scaling on the time series data to obtain standard time series data;
[0030] A second feature is extracted from the standard time series data.
[0031] Optionally, obtaining target text information for describing the multimodal data according to the target feature includes:
[0032] Dividing the target feature into a plurality of sub-features;
[0033] Determining text information corresponding to the sub-feature;
[0034] The text information corresponding to the plurality of sub-features is fused to obtain the target text information.
[0035] Optionally, the time series data includes a plurality of variables, and performing minimum scaling and maximum scaling on the time series data to obtain standard time series data includes:
[0036] determining a smallest first variable and a largest second variable among the plurality of variables;
[0037] Calculate the standard variable corresponding to the variable according to the first variable and the second variable;
[0038] The standard variables are used to replace the variables in the time series data to obtain the standard time series data.
[0039] Optionally, extracting the second feature from the standard time series data includes:
[0040] Splitting the standard time series data into a plurality of sub-time series data;
[0041] Extracting a second sub-feature from the sub-time series data;
[0042] The second feature is obtained by fusing multiple second sub-features.
[0043] Optionally, the sub-feature has a corresponding time step, each time step has a corresponding hidden state, and the determining of the text information corresponding to the sub-feature includes:
[0044] Calculate the hidden state corresponding to the current time step based on the hidden state corresponding to the previous time step and the sub-feature corresponding to the current time step;
[0045] According to the hidden state corresponding to the current time step, the text information corresponding to the sub-feature at the current time step is calculated.
[0046] The embodiment of the present application also discloses a text generation device based on multimodal data, the device comprising:
[0047] An acquisition module, used to acquire multimodal data and time series data corresponding to the multimodal data;
[0048] An extraction module, used to extract a first feature from the multimodal data and a second feature from the time series data;
[0049] A fusion module, used for fusing the first feature and the second feature to obtain a target feature;
[0050] The description module is used to obtain target text information corresponding to the multimodal data according to the target features.
[0051] An embodiment of the present application also discloses an electronic device, including: a processor; and a memory, on which executable code is stored. When the executable code is executed, the processor executes the text generation method based on multimodal data as described in any one of the embodiments of the present application.
[0052] The embodiments of the present application also disclose one or more machine-readable media on which executable codes are stored. When the executable codes are executed, the processor executes the text generation method based on multimodal data as described in any one of the embodiments of the present application.
[0053] Compared with the prior art, the embodiments of the present application have the following advantages:
[0054] In an embodiment of the present application, multimodal data and time series data corresponding to the multimodal data are obtained; a first feature in the multimodal data and a second feature in the time series data are extracted; the first feature and the second feature are fused to obtain a target feature; and according to the target feature, a target text information corresponding to the multimodal data is obtained. The embodiment of the present application avoids directly processing multimodal time series data by extracting features of multimodal data and time series data separately and fusing them, significantly reducing the consumption of computing resources and improving processing efficiency, especially when processing long sequence data, and can effectively deal with the problem of high computational complexity. By fusing the features of multimodal data and time series data, the rich information in different modal data can be fully utilized, which helps to capture the correlation between multimodal data and long-term dependencies in time series, generate more accurate and detailed text descriptions, and thus improve the quality of text generation. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 It is a flowchart of the steps of an embodiment of a method for generating text based on multimodal data of the present application;
[0056] Figure 2 It is a schematic diagram of the architecture of a text generation method based on multimodal data of the present application;
[0057] Figure 3 It is a structural block diagram of an embodiment of a text generation device based on multimodal data of the present application;
[0058] Figure 4 It is a schematic diagram of the structure of a device provided in one embodiment of the present application. DETAILED DESCRIPTION
[0059] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below in conjunction with the accompanying drawings and specific implementation methods.
[0060] Reference Figure 1 , is a flowchart of a method for generating text based on multimodal data according to an embodiment of the present application, comprising the following steps:
[0061] Step 101: Acquire multimodal data and time series data corresponding to the multimodal data;
[0062] In the embodiment of the present application, it is necessary to collect and ensure the temporal synchronization of multimodal data (such as images, text, audio, etc.) and corresponding time series data. Specifically, it can be video data (including images and audio) and its timestamp.
[0063] Step 102: extracting a first feature from the multimodal data and a second feature from the time series data;
[0064] In the embodiment of the present application, it is necessary to extract features from data of different modalities for subsequent fusion.
[0065] In one embodiment of the present application, the efficiency of feature extraction is improved through a linear attention mechanism and semi-separable matrix block decomposition.
[0066] Step 103: Fusing the first feature and the second feature to obtain a target feature;
[0067] In the embodiment of the present application, it is necessary to fuse the features of the multimodal data and the time series data to generate the target feature. Specifically, the first feature and the second feature can be fused together by weighted fusion, splicing fusion, and cross-connection fusion.
[0068] Weighted fusion refers to assigning different weights to different features, and then performing weighted summation or other combination operations. These weights are usually learned during the training process to optimize the performance of the fused features. The weights can be automatically adjusted according to the importance of the modality to improve the fusion effect.
[0069] Concatenation and fusion refers to directly concatenating different features into a high-dimensional vector, which retains information of all modalities but may lead to dimensional expansion and increased computational complexity.
[0070] Cross-connection refers to the interaction and fusion of different features through a certain mechanism at different layers or stages of the network. It can better capture the relationship between features and improve the expressiveness of the model.
[0071] In one embodiment of the present application, when fusing the first feature and the second feature, a state space model (SSM) can be introduced to model the dynamic characteristics of time series data. By combining structured mask attention (SMA), the long-term dependency between multimodal data can be effectively captured, thereby improving the quality of the fusion of the first feature and the second feature.
[0072] In another embodiment of the present application, a new combined algorithm based on semi-separable matrix block decomposition is proposed, which utilizes the linear recursion and quadratic dual form of the state space model to achieve the best trade-off in the main efficiency indicators. Compared with the traditional Transformers optimized selective scanning implementation, the new combined algorithm has improved implementation speed while allowing the use of larger loop state sizes with almost no speed loss. In addition, the new combined algorithm is highly competitive with optimized softmax attention implementations (such as FlashAttention-2), and the longer the sequence length, the better the performance.
[0073] In order to further optimize the parallel processing capability of multimodal data, the Transformers block can also be modified to achieve tensor parallelism, mainly by introducing the grouped-value attention (GVA) head structure, thereby improving the parallel efficiency and scalability of the model. Through these algorithm-level optimizations, the embodiments of the present application can achieve higher efficiency and better performance when processing large-scale multimodal and time series data.
[0074] In the step of fusing the first feature and the second feature to obtain the target feature, the new combination algorithm is used to efficiently process the fused features, ensuring that the computational efficiency is greatly improved while maintaining high accuracy.
[0075] Specifically, the semi-separable state space model matrix is first divided into blocks of size Q×Q. Then, the properties of the semi-separated matrix are used to decompose each low-rank non-diagonal block, and a low-rank approximation is performed on each low-rank non-diagonal block. Among them, the diagonal blocks are regarded as smaller semi-separable matrices, which can be calculated using the quadratic form of the above-mentioned new combined algorithm (similar to the attention mechanism), and the non-diagonal blocks are calculated by batch matrix multiplication, which greatly reduces the computational complexity and improves the efficiency of feature fusion.
[0076] Step 104: Obtain target text information corresponding to the multimodal data according to the target features.
[0077] In the embodiment of the present application, text information describing the multimodal data is eventually generated. Specifically, based on the Transformer generative model, a text description is generated according to the fused features.
[0078] In order to further optimize the generation efficiency of target text information, tensor parallelism and sequence parallelism technologies can be introduced.
[0079] Tensor parallelism: By introducing a parallel projection structure, we split the input projection and output projection matrices into multiple fragments, depending on the tensor parallelism. This allows each layer to only require one all-reduce instead of two. At the same time, group normalization is used to ensure that each GPU can complete the normalization operation independently.
[0080] Sequence parallelism: For residual and normalization operations, reduce-scatter, residual + normalization, and then all-gather are used to replace all-reduce in tensor parallelism. For attention or state-space model operations, we adopt context parallelism (CP) by splitting along the sequence dimension through ring attention and transferring states between GPUs.
[0081] In the embodiments of the present application, by optimizing the feature extraction and fusion process, multimodal time series data can be efficiently processed and high-quality text descriptions can be generated, which are suitable for complex scenarios.
[0082] In one embodiment of the present application, in the step of describing the generation of target text information, optimization can also be achieved through the above-mentioned new combination algorithm, which can maintain a fast generation speed and high-precision text description capability when processing long sequence data.
[0083] In one embodiment of the present application, extracting the first feature from the multimodal data includes:
[0084] Build visual encoder and object detection models;
[0085] extracting visual features from the multimodal data using the visual encoder;
[0086] The target detection model is used to extract multi-scale features of the target object in the multimodal data at multiple scales and global features at the overall level;
[0087] The visual feature, the multi-scale feature and the global feature are integrated to obtain the first feature.
[0088] In an embodiment of the present application, the visual encoder refers to a pre-trained convolutional neural network (CNN) model, such as ResNet, which is used to extract visual features of an image.
[0089] The target detection model refers to a model that combines SSD (Single Shot MultiBox Detector) and YOLO (You Only Look Once), which is used to detect target objects in images and extract features of these targets at different scales.
[0090] SSD is an object detection algorithm that performs well in small object detection. SSD can effectively detect small objects in images by making predictions on feature maps of different scales. This is because feature maps of different scales correspond to different receptive fields and can better capture objects of different sizes.
[0091] YOLO is an object detection algorithm known for its speed. The advantage of YOLO is that it can achieve fast detection because it directly predicts on the entire image without generating and classifying candidate regions like other methods. This end-to-end detection method makes YOLO very popular in application scenarios with high real-time requirements.
[0092] In an embodiment of the present application, when the visual encoder extracts the first feature of multimodal data, a linear attention mechanism can be used to optimize the feature extraction process of the visual encoder. Specifically, by introducing the semi-separable matrix block decomposition technology, the computational complexity can be effectively reduced, thereby improving the extraction efficiency of visual features.
[0093] When extracting the first feature of multimodal data for the target detection model, multi-scale features can be obtained through SSD, and global features can be obtained through YOLO. By integrating the detection results of SSD and YOLO, it is possible to maintain the advantage of SSD in detecting small objects while taking advantage of the speed of YOLO. Exemplarily, the detection results of SSD and YOLO are processed by non-maximum suppression (NMS) respectively to remove redundant frames, and the detection results of the two are merged, usually by setting a threshold to retain the detection frames with higher confidence, and finally outputting the merged detection results to improve the overall detection accuracy and speed.
[0094] In one embodiment of the present application, the parameters of SSD and YOLO can be adjusted to optimize the overall performance of the model, thereby improving the model's processing efficiency for multimodal data. For example, when the task mainly relies on small object detection, the parameters of SSD may be adjusted first to optimize its small object detection performance; for another example, when the task has high speed requirements, the parameters of YOLO may be adjusted first to optimize its speed performance; for another example, the parameters of SSD and YOLO are adjusted at the same time: when the task needs to take into account both small object detection and speed, the parameters of both may be adjusted at the same time to achieve the best balance.
[0095] In one embodiment of the present application, the step 103 of fusing the first feature and the second feature to obtain a target feature includes:
[0096] Build feature mapper;
[0097] Mapping the first feature and the second feature into a text space using the feature mapper to obtain a first text feature corresponding to the first feature and a second text feature corresponding to the second feature;
[0098] The first text feature and the second text feature are fused to obtain the target feature.
[0099] In an embodiment of the present application, the feature mapper is a neural network module for converting features (first features and second features) from different modalities into a common text space, so that the features of different modalities are comparable and fusible.
[0100] A feature mapper is used to map the first feature (such as image feature) and the second feature (such as time series feature) into the text space respectively, to obtain the first text feature and the second text feature, so as to ensure that the mapped features are suitable for the text generation task.
[0101] The mapped first text feature and the second text feature are fused, and concatenation, weighted summation or a more complex fusion method (such as attention mechanism) can also be used to generate the final target feature.
[0102] The embodiment of the present application ensures that the features of multimodal and time series data are semantically aligned by mapping to the text space, thereby improving the fusion effect. All features are in the same space and can be directly used in the text generation model, simplifying subsequent processing steps. The fused target features contain rich semantic information, which helps to generate more accurate and detailed text descriptions.
[0103] In one embodiment of the present application, the feature mapper has corresponding parameters, and the feature mapper is trained by the following steps:
[0104] Obtain multimodal sample data;
[0105] Extracting a first sample feature from the multimodal sample data;
[0106] Extracting a text sample entity from the first sample feature, and using the text sample entity as a reference point of the text space;
[0107] Mapping the first sample feature into the text space using the feature mapper to obtain a first target feature;
[0108] Calculate the loss value corresponding to the feature mapper according to the first target feature and the text sample entity;
[0109] The parameters corresponding to the feature mapper are adjusted according to the loss value.
[0110] In the embodiment of the present application, in order to optimize the performance of the feature mapper, it is necessary to train the feature mapper, and the specific training process is as follows:
[0111] Collect sample data containing multiple modalities (such as images, texts, etc.), which will be used to train the feature mapper; extract initial features from the multimodal sample data, which may include image features, text features, etc.; extract key entities from the text data and use them as reference points in the text space. These entities represent important information in the text and are used to guide the learning of the feature mapper; use the feature mapper to map the extracted first sample features to the text space to generate a first target feature; calculate the loss value by comparing the first target feature with the text sample entity, which measures the difference between the mapped feature and the actual text entity; according to the calculated loss value, adjust the parameters of the feature mapper through the back propagation algorithm to optimize its performance.
[0112] In an embodiment of the present application, LLM (Large Language Model) is used to generate the target entity name. For example, the image features (first sample features) output by the visual encoder are input into the LLM, and then the LLM generates a text sample entity.
[0113] The embodiment of the present application uses text sample entities as reference points to ensure that the mapped features are semantically consistent with the actual text information, thereby improving the accuracy of the generated text. This training method improves the generalization ability of the feature mapper, enabling it to handle more diverse multimodal data and adapt to a wider range of scenarios. Through the calculation of loss values and parameter adjustment, the feature mapper can better learn the semantic representation of multimodal data, improving the performance of the overall system.
[0114] In one embodiment of the present application, in order to enhance the generalization ability of the feature mapper, we use a multimodal instance library and context examples. The multimodal instance library provides a rich set of cross-data type examples, while the context examples ensure that the model can be better generalized in different scenarios and data distributions. By combining the multimodal instance library and the context examples, the feature mapper performs better in practical applications, thereby improving the more accurate features obtained after feature mapping.
[0115] Exemplarily, given a multimodal mention context, the target entity name is directly generated using a fixed-parameter LLM, and n retrieved multimodal instances are used as context examples. In this way, the feature mapper is able to more accurately represent image features in the text space, thereby improving the quality of overall text generation.
[0116] In one embodiment of the present application, extracting the second feature in the time series time includes:
[0117] Performing minimum scaling and maximum scaling on the time series data to obtain standard time series data;
[0118] A second feature is extracted from the standard time series data.
[0119] In an embodiment of the present application, it is necessary to standardize the time series data through minimum scaling and maximum scaling, and scale the time series data to a fixed range (usually [0, 1]) to eliminate the scale differences between different variables, thereby speeding up the generation of text information based on multimodal data and avoiding the impact of excessive differences in eigenvalues on generation efficiency.
[0120] Extract feature representation from the standardized time series data, possibly using models such as CNN and RNN. By extracting the second feature, complex patterns in the time series, such as trends and periodicity, can be captured, providing high-quality feature representation for multimodal data fusion.
[0121] The embodiment of the present application combines standardization and feature extraction, providing a better foundation for the fusion of multimodal data and time series data, and improving the quality and accuracy of the final text generation.
[0122] In an embodiment of the present application, when extracting the second feature in the time series data, the scaled input values and time information (standard time series data) can be embedded by using convolutional layers with different expansion rates to ensure that the representation of subsequent layers has a larger receptive field, which helps to capture long-term dependencies in the time series data and complex patterns of multimodal data.
[0123] In one embodiment of the present application, the step 104, obtaining target text information for describing the multimodal data according to the target feature, includes:
[0124] Dividing the target feature into a plurality of sub-features;
[0125] Determining text information corresponding to the sub-feature;
[0126] The text information corresponding to the plurality of sub-features is fused to obtain the target text information.
[0127] In the embodiment of the present application, the processing of the target feature is achieved by block division and state transfer. Specifically, the target feature is first divided into blocks of size Q, and the local output of each block is calculated (assuming that the initial state is 0, what is the output of each block); the final state of each block is calculated (assuming that the initial state is 0, what is the final state of each block); the final state of all blocks is calculated recursively, using any desired algorithm, such as parallel or sequential scanning (taking into account all previous inputs, what is the actual final state of each block); for each block, its actual final state is output.
[0128] It should be noted that in the above process of processing the target features, the calculation of local data, the calculation of the final state and the output of the final state all utilize matrix multiplication, that is, the use of tensor cores, and can be calculated in parallel. Only when calculating the recursion of the final state does scanning need to be done, but since only a very small feature block needs to be scanned, it usually takes very little time.
[0129] For example, suppose there is video data (including images and audio) and corresponding time series data. The target feature is divided into multiple sub-features, and each sub-feature generates a text description, such as "a person is running", "background music is playing", "high intensity of exercise", and finally merged into a text description of "a person is running intensely, and background music is playing."
[0130] In the embodiment of the present application, by dividing the target feature into sub-features, each part can be processed more finely, thereby improving the accuracy of the text description. Block division and state transfer make the processing of complex data more efficient, and each sub-feature can be processed independently, reducing the burden on computing resources. By considering the context and state transfer, the semantic consistency and coherence of the generated text information are ensured.
[0131] In one embodiment of the present application, the time series data includes multiple variables, and performing minimum scaling and maximum scaling on the time series data to obtain standard time series data includes:
[0132] determining a smallest first variable and a largest second variable among the plurality of variables;
[0133] Calculate the standard variable corresponding to the variable according to the first variable and the second variable;
[0134] The standard variables are used to replace the variables in the time series data to obtain the standard time series data.
[0135] In the embodiment of the present application, the time series data includes multiple variables X, and the time series data can be standardized by performing min-max scaling on these variables.
[0136] Among the multiple variables, the minimum value X_min and the maximum value X_max are determined, and the standard variable X_1 = (X-X_min) / (X_max-X_min) of each variable is calculated.
[0137] The standardized time series data is obtained by replacing the variable X in the original time series data with the standardized standard variable X_1.
[0138] The embodiment of the present application ensures that each variable has equal importance in model training by scaling all variables to the same range. For machine learning models (such as neural networks) that are sensitive to the scale of input data, standardization of time series data can improve the performance and accuracy of the model, thereby improving the speed and accuracy of extracting the second feature, and also avoiding variables with larger scales dominating the distance calculation, ensuring that the model will not be biased towards certain features due to differences in variable scales.
[0139] In one embodiment of the present application, extracting the second feature from the standard time series data includes:
[0140] Splitting the standard time series data into a plurality of sub-time series data;
[0141] Extracting a second sub-feature from the sub-time series data;
[0142] The second feature is obtained by fusing multiple second sub-features.
[0143] In an embodiment of the present application, the standardized time series data is divided into multiple sub-time series data so as to be processed in parallel on different computing units. This is suitable for very long time series data and avoids the problem that a single model is difficult to process overly long sequences.
[0144] By processing multiple subsequences in parallel, the overall computing time can be significantly reduced. Multi-core CPU or distributed computing resources can be fully utilized to improve the utilization of computing resources.
[0145] Specific second sub-features, such as statistical features, frequency domain features, etc., are extracted from each sub-time series data to capture different local patterns and data characteristics. Each sub-sequence may contain different local patterns, and extracting sub-features can better capture these local information. In addition, different sub-sequences may extract different features, increasing the diversity and richness of features.
[0146] The second sub-features extracted from each sub-sequence are integrated to obtain a comprehensive feature representation through splicing, weighted averaging or other fusion strategies. Fusion of multiple sub-features can integrate the information of each sub-sequence to obtain a more comprehensive feature representation. Through fusion, the noise impact of a single sub-sequence feature can be reduced and the robustness of the overall feature can be improved.
[0147] The embodiment of the present application uses the idea of sequence parallelism to effectively handle the problem of feature extraction of long sequence data, improve the computational efficiency and the quality of feature representation, and is suitable for the fusion processing of multimodal data and time series data, providing strong support for the analysis of complex data.
[0148] In one embodiment of the present application, the sub-feature has a corresponding time step, each time step has a corresponding hidden state, and the determining of the text information corresponding to the sub-feature includes:
[0149] Calculate the hidden state corresponding to the current time step based on the hidden state corresponding to the previous time step and the sub-feature corresponding to the current time step;
[0150] According to the hidden state corresponding to the current time step, the text information corresponding to the sub-feature at the current time step is calculated.
[0151] In time series data, data is usually arranged in chronological order, and the data at each time point can be regarded as a time step, and each time step has a corresponding hidden state. In the state space model (recurrent neural network RNN), the hidden state is an important concept that contains information from previous time steps and helps the model remember long-term dependencies in the sequence.
[0152] In the embodiment of the present application, the generation of text is achieved by updating the hidden state in the RNN. Specifically, the hidden state is continuously updated with the input of each time step, and then the output, such as text, is generated according to the final hidden state.
[0153] Specifically, each sub-feature is treated as an input for one time step, and then the hidden state is updated in some recursive way.
[0154] In the embodiment of the present application, the hidden state corresponding to the current time step is calculated by the following recursive formula:
[0155] h t =(A t h t-1 +B t x t )×m+n
[0156] Among them, h t is the hidden state corresponding to the current time step, A t is the transition matrix between hidden states, h t-1 is the hidden state corresponding to the previous time step, B t is the transformation matrix representing the input to the hidden state, x t is the input of the current time step, that is, the sub-feature corresponding to the current time step, m is the model coefficient, and n is the accuracy threshold. Specifically, m and n are related parameters used to adjust the hidden state.
[0157] In the above recursive formula of the hidden state, any method for calculating the forward propagation of the state space model can be regarded as a matrix multiplication algorithm on a semi-separable matrix, including a linear time semi-separable matrix multiplication algorithm and a quadratic time naive matrix multiplication. In practical applications, different matrix multiplication algorithms can be selected according to specific computing requirements to optimize computing efficiency. Specifically, a linear time semi-separable matrix multiplication algorithm is used to calculate the multiplication of low-rank non-diagonal blocks, and a quadratic time naive matrix multiplication is used to calculate the multiplication of diagonal blocks.
[0158] In an embodiment of the present application, different matrix multiplications are selected to calculate the hidden state to balance computational efficiency and model performance, thereby ensuring that the hidden state can be updated efficiently when processing time series data while maintaining the accuracy of the model.
[0159] In the embodiments of the present application, not only can the complex patterns and long-term dependencies of time series data be efficiently captured, but also the computational complexity in the data processing process is significantly reduced, and the efficiency of processing long sequence data is improved. By fusing time series data, the model can more accurately map image features to text space and generate high-quality text descriptions. When processing video frames or multimodal sequence data, the model can capture dynamic changes in the time dimension and improve the generalization ability of the model.
[0160] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present application, the following examples are provided for illustrative explanation:
[0161] Scenario: Video frame description generation task
[0162] 1) Input data
[0163] Image data: a frame of an image in a video, such as an image containing "a cat sleeping on the sofa".
[0164] Text data: Text information associated with the image, such as “the cat is sleeping on the sofa” or “the cat’s movements are gradually becoming slower.”
[0165] Time series data: Time series data related to video frames, such as the timestamp of the video frame (such as the 10th second), changes in objects over time (such as changes in the movements of a cat), sensor data (such as temperature changes), etc.
[0166] Among them, image data and text data are part of multimodal data.
[0167] 2) Data processing
[0168] Image data processing: extract image features and capture visual information in the image, such as the shape, color, and position of the cat. Image features are represented as high-dimensional vectors for subsequent feature fusion.
[0169] Text data processing: Perform preprocessing such as word segmentation and embedding on text data to generate text feature representation. Text features are represented as high-dimensional vectors for fusion with image features and time series features.
[0170] Time series data processing: Standardize the time series data (such as minimum-maximum scaling) and extract time features (such as timestamps, time differences, etc.) to extract time series features and capture dynamic changes in the time dimension, such as changes in a cat’s movements.
[0171] 3) Feature Fusion
[0172] Image features, time series features, and text features are fused to form multimodal time series features. For example, image features capture the visual information of a cat, text features provide a description of the cat's behavior, and time series features capture the cat's movement changes. The fused features can simultaneously reflect the characteristics of images, time series, and text.
[0173] 4) Block matrix decomposition and state transfer
[0174] The fused multimodal time series features are subjected to block matrix decomposition to optimize computational efficiency. Specifically, the state space model matrix is divided into multiple small blocks, decomposed using the properties of semi-separable matrices, and the local output and final state of each block are calculated through state transfer to capture the long-term dependency of the time series.
[0175] 5) Output data
[0176] Generated text description, such as "A cat falls asleep on the sofa."
[0177] In the embodiment of the present application, through feature extraction, fusion and mapping, the model can combine these multimodal information to generate high-quality text descriptions. The text description is the output of the model, which is used to explain or describe the content of the input data (multimodal data and time series data corresponding to the multimodal data), such as generating natural language sentences related to images and time series data.
[0178] Reference Figure 2 , is a schematic diagram of the architecture of a text generation method based on multimodal data in this application. Through a series of algorithm optimization and model design, it can efficiently process multimodal and time series data and ultimately generate high-quality text descriptions.
[0179] Efficiency algorithms: Algorithms used to optimize computational efficiency, such as the semi-separable matrix decomposition and attention mechanism mentioned above, aiming to reduce the computational complexity when processing large-scale data.
[0180] Time Series: represents time series data, which is time-dependent and needs to be captured through methods such as state space models.
[0181] Structured matrices: Structured representations of data, such as semi-separable matrices or sparse matrices, to optimize computational efficiency.
[0182] Structured Masked Attention: represents a sparse attention mechanism that focuses on the most important parts of the input data and reduces the consumption of computing resources.
[0183] Input Embedding: It is used to convert multimodal data (such as images, text, etc.) and time series data into numerical representation for subsequent processing.
[0184] Semi-separable matrices: Decomposition methods for representing semi-separable matrices are used to optimize the efficiency of matrix operations, especially when dealing with large-scale time series data.
[0185] State-space duality: It represents the duality of the state-space model, which may be used to process time series data in both directions and capture its dynamic changes.
[0186] Attention: represents the attention mechanism, which is used to focus on the important parts of the input data and improve the accuracy of text generation.
[0187] State-space models: used to model the dynamic characteristics of time series data and capture long-term dependencies.
[0188] LLMs: represents a large set of language models, which are used to generate the final text description.
[0189] like Figure 2 As shown in the figure, the system takes multimodal data (such as images, text, etc.) and time series data as input, which are converted into numerical representations through the Input Embedding module for subsequent calculations and processing. In order to improve computational efficiency, the system adopts structured matrix and semi-separable matrix representation methods, which can effectively optimize the processing of large-scale data. Based on data representation, the system introduces efficiency algorithms such as semi-separable matrix decomposition and attention mechanism to reduce computational complexity and improve processing speed. In particular, the structured masked attention mechanism can focus on the most important part of the input data and reduce unnecessary consumption of computing resources. For time series data, the system captures its dynamic characteristics through a state space model and uses state space duality for bidirectional processing to better capture long-term dependencies.
[0190] During the processing, the attention mechanism is widely used to ensure that the system can focus on the key parts of the input data, thereby improving the accuracy of text generation. Finally, all processed data is input into LLMs (Large Language Model Sets), which are based on advanced language generation technology and can generate high-quality text descriptions. The entire process achieves efficient processing of multimodal and time series data through a series of algorithm optimization and model design, and finally outputs text content that meets the requirements.
[0191] In an embodiment of the present application, multimodal data and time series data corresponding to the multimodal data are obtained; a first feature in the multimodal data and a second feature in the time series data are extracted; the first feature and the second feature are fused to obtain a target feature; and according to the target feature, a target text information corresponding to the multimodal data is obtained. The embodiment of the present application avoids directly processing multimodal time series data by extracting features of multimodal data and time series data separately and fusing them, significantly reducing the consumption of computing resources and improving processing efficiency, especially when processing long sequence data, and can effectively deal with the problem of high computational complexity. By fusing the features of multimodal data and time series data, the rich information in different modal data can be fully utilized, which helps to capture the correlation between multimodal data and long-term dependencies in time series, generate more accurate and detailed text descriptions, and thus improve the quality of text generation.
[0192] It should be noted that, for the method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of the present application are not limited by the described order of actions, because according to the embodiments of the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present application.
[0193] On the basis of the above-mentioned embodiments, this embodiment further provides a text generation device based on multimodal data, which is applied to electronic devices such as terminal devices and servers.
[0194] Reference Figure 3 , shows a structural block diagram of an embodiment of a text generation device based on multimodal data of the present application, which may specifically include the following modules:
[0195] An acquisition module 301 is used to acquire multimodal data and time series data corresponding to the multimodal data;
[0196] An extraction module 302, configured to extract a first feature from the multimodal data and a second feature from the time series data;
[0197] A fusion module 303 is used to fuse the first feature and the second feature to obtain a target feature;
[0198] The description module 304 is used to obtain target text information corresponding to the multimodal data according to the target features.
[0199] The embodiment of the present application also provides a non-volatile readable storage medium, which stores one or more modules (programs). When the one or more modules are applied to a device, the device can execute instructions (instructions) of each method step in the embodiment of the present application.
[0200] The present application embodiment provides one or more machine-readable media on which instructions are stored, and when executed by one or more processors, an electronic device executes one or more of the methods described in the above embodiments. In the present application embodiment, the electronic device includes various types of devices such as terminal devices and servers (clusters).
[0201] The embodiments of the present disclosure may be implemented as a device configured as desired using any appropriate hardware, firmware, software, or any combination thereof, and the device may include electronic devices such as terminal devices and servers (clusters). Figure 4 An exemplary apparatus 400 that may be used to implement various embodiments described in this application is schematically illustrated.
[0202] For one embodiment, Figure 4 An exemplary apparatus 400 is shown having one or more processors 402, a control module (chip set) 404 coupled to at least one of the processor(s) 402, a memory 406 coupled to the control module 404, a non-volatile memory (NVM) / storage device 408 coupled to the control module 404, one or more input / output devices 410 coupled to the control module 404, and a network interface 412 coupled to the control module 404.
[0203] The processor 402 may include one or more single-core or multi-core processors, and the processor 402 may include any combination of general-purpose processors or special-purpose processors (such as graphics processors, application processors, baseband processors, etc.). In some embodiments, the device 400 can be used as a terminal device, server (cluster), etc. described in the embodiments of the present application.
[0204] In some embodiments, the apparatus 400 may include one or more computer-readable media (e.g., memory 406 or NVM / storage device 408) having instructions 414 and one or more processors 402 combined with the one or more computer-readable media and configured to execute the instructions 414 to implement a module to perform the actions described in the present disclosure.
[0205] For one embodiment, control module 404 may include any suitable interface controller to provide any suitable interface to at least one of processor(s) 402 and / or any suitable device or component in communication with control module 404 .
[0206] The control module 404 may include a memory controller module to provide an interface to the memory 406. The memory controller module may be a hardware module, a software module, and / or a firmware module.
[0207] The memory 406 may be used, for example, to load and store data and / or instructions 414 for the device 400. For one embodiment, the memory 406 may include any suitable volatile memory, such as a suitable DRAM. In some embodiments, the memory 406 may include a double data rate type four synchronous dynamic random access memory (DDR4 SDRAM).
[0208] For one embodiment, control module 404 may include one or more input / output controllers to provide interfaces to NVM / storage device 408 and input / output device(s) 410 .
[0209] For example, NVM / storage 408 may be used to store data and / or instructions 414. NVM / storage 408 may include any suitable non-volatile memory (e.g., flash memory) and / or may include any suitable non-volatile storage device(s) (e.g., one or more hard disk drives (HDDs), one or more compact disk (CD) drives, and / or one or more digital versatile disk (DVD) drives).
[0210] NVM / storage device 408 may include storage resources that are physically part of the device on which apparatus 400 is installed, or it may be accessible to the device without being part of the device. For example, NVM / storage device 408 may be accessed via input / output device(s) 410 over a network.
[0211] (One or more) input / output devices 410 may provide an interface for the apparatus 400 to communicate with any other appropriate device, and the input / output device 410 may include a communication component, an audio component, a sensor component, etc. The network interface 412 may provide an interface for the apparatus 400 to communicate through one or more networks, and the apparatus 400 may wirelessly communicate with one or more components of a wireless network according to any of one or more wireless network standards and / or protocols, for example, accessing a wireless network based on a communication standard, such as WiFi, 2G, 3G, 4G, 5G, etc., or a combination thereof for wireless communication.
[0212] For one embodiment, at least one of the processor(s) 402 may be packaged together with the logic of one or more controllers (e.g., a memory controller module) of the control module 404. For one embodiment, at least one of the processor(s) 402 may be packaged together with the logic of one or more controllers of the control module 404 to form a system-in-package (SiP). For one embodiment, at least one of the processor(s) 402 may be integrated on the same die with the logic of one or more controllers of the control module 404. For one embodiment, at least one of the processor(s) 402 may be integrated on the same die with the logic of one or more controllers of the control module 404 to form a system-on-chip (SoC).
[0213] In various embodiments, the device 400 may be, but is not limited to, a terminal device such as a server, a desktop computing device, or a mobile computing device (e.g., a laptop computing device, a handheld computing device, a tablet computer, a netbook, etc.). In various embodiments, the device 400 may have more or fewer components and / or different architectures. For example, in some embodiments, the device 400 includes one or more cameras, a keyboard, a liquid crystal display (LCD) screen (including a touch screen display), a non-volatile memory port, multiple antennas, a graphics chip, an application-specific integrated circuit (ASIC), and a speaker.
[0214] Among them, the main control chip can be used as a processor or control module in the detection device, sensor data, location information, etc. are stored in a memory or NVM / storage device, the sensor group can be used as an input / output device, and the communication interface may include a network interface.
[0215] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0216] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0217] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable text generation terminal device based on multimodal data to produce a machine, so that the instructions executed by the processor of the computer or other programmable text generation terminal device based on multimodal data generate instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0218] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable text generation terminal device based on multimodal data to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0219] These computer program instructions can also be loaded onto a computer or other programmable terminal device for generating text based on multimodal data, so that a series of operation steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable terminal device provide for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0220] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present application.
[0221] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or terminal device including the elements.
[0222] The above is a detailed introduction to a text generation method and device based on multimodal data, an electronic device and a storage medium provided by the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for general technical personnel in this field, according to the idea of the present application, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. A text generation method based on multimodal data, characterized in that: The method comprises: Acquire multimodal data and time series data corresponding to the multimodal data; Extracting a first feature from the multimodal data and a second feature from the time series data; fusing the first feature and the second feature to obtain a target feature; According to the target features, target text information corresponding to the multimodal data is obtained.
2. The method according to claim 1, characterized in that The extracting a first feature from the multimodal data includes: Build visual encoder and object detection models; extracting visual features from the multimodal data using the visual encoder; The target detection model is used to extract multi-scale features of the target object in the multimodal data at multiple scales and global features at the overall level; The visual feature, the multi-scale feature and the global feature are integrated to obtain the first feature.
3. The method according to claim 1, characterized in that The fusing the first feature and the second feature to obtain a target feature includes: Build feature mapper; Mapping the first feature and the second feature into a text space using the feature mapper to obtain a first text feature corresponding to the first feature and a second text feature corresponding to the second feature; The first text feature and the second text feature are fused to obtain the target feature.
4. The method according to claim 3, characterized in that The feature mapper has corresponding parameters, and the feature mapper is trained by the following steps: Obtain multimodal sample data; Extracting a first sample feature from the multimodal sample data; Extracting a text sample entity from the first sample feature, and using the text sample entity as a reference point of the text space; Mapping the first sample feature into the text space using the feature mapper to obtain a first target feature; Calculate the loss value corresponding to the feature mapper according to the first target feature and the text sample entity; The parameters corresponding to the feature mapper are adjusted according to the loss value.
5. The method according to claim 1, characterized in that Extracting a second feature of the time series includes: Performing minimum scaling and maximum scaling on the time series data to obtain standard time series data; A second feature is extracted from the standard time series data.
6. The method according to claim 1, characterized in that The step of obtaining target text information corresponding to the multimodal data according to the target feature includes: Dividing the target feature into a plurality of sub-features; Determining text information corresponding to the sub-feature; The text information corresponding to the plurality of sub-features is fused to obtain the target text information.
7. The method according to claim 5, characterized in that The time series data includes a plurality of variables, and performing minimum scaling and maximum scaling on the time series data to obtain standard time series data includes: determining a smallest first variable and a largest second variable among the plurality of variables; Calculate the standard variable corresponding to the variable according to the first variable and the second variable; The standard variables are used to replace the variables in the time series data to obtain the standard time series data.
8. The method according to claim 5, characterized in that The extracting the second feature from the standard time series data comprises: Splitting the standard time series data into a plurality of sub-time series data; Extracting a second sub-feature from the sub-time series data; The second feature is obtained by fusing multiple second sub-features.
9. The method according to claim 6, characterized in that The sub-feature has a corresponding time step, each time step has a corresponding hidden state, and the determining of the text information corresponding to the sub-feature includes: Calculate the hidden state corresponding to the current time step based on the hidden state corresponding to the previous time step and the sub-feature corresponding to the current time step; According to the hidden state corresponding to the current time step, the text information corresponding to the sub-feature at the current time step is calculated.
10. A text generation device based on multimodal data, characterized in that: The device comprises: An acquisition module, used to acquire multimodal data and time series data corresponding to the multimodal data; An extraction module, used to extract a first feature from the multimodal data and a second feature from the time series data; A fusion module, used for fusing the first feature and the second feature to obtain a target feature; The description module is used to obtain target text information corresponding to the multimodal data according to the target features.
11. An electronic device, characterized in that: include: processor; and A memory having executable codes stored thereon, which, when executed, enables the processor to execute the text generation method based on multimodal data as described in any one of claims 1-9.
12. One or more machine-readable media having executable codes stored thereon, which, when executed, enable a processor to execute the text generation method based on multimodal data as described in any one of claims 1-9.