A multi-modal data generation method, system and electronic device for multi-dimensional data
By integrating and synchronously labeling multi-dimensional data, and using weighted fusion and deep learning models to generate multimodal data, the alignment and coordination problems in single-modal data processing are solved, and efficient fusion and accurate analysis of multimodal data are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING YISHENGZE TECHNOLOGY CO LTD
- Filing Date
- 2026-05-09
- Publication Date
- 2026-07-10
AI Technical Summary
In existing technologies, single-modal data processing suffers from problems such as incompatibility, lack of coordination, and labeling errors, which limit the accuracy and robustness of multimodal data applications and make it difficult to effectively integrate and fuse different types of data.
After acquiring, integrating, and synchronizing multi-dimensional data, corresponding analysis methods are used for annotation to extract semantic information and structured features. Multimodal data is generated using weighted fusion and deep learning models, and iterative optimization is performed based on feedback from historical data.
It achieves alignment, coordination and synchronization of multimodal data, improves the accuracy and robustness of data processing, reduces annotation errors and time costs, and ensures the reliability and accuracy of analysis results.
Smart Images

Figure CN122364772A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the field of data processing, and in particular to a method, system and electronic device for generating multi-dimensional data in a multimodal manner. Background Technology
[0002] The rapid development of automation and robotics has led to the widespread application of data collection in various scenarios. Sensor data, image data, and audio data are all collected and applied in a single manner. However, this single-dimensional data processing has limitations: Most current data generation systems rely on single-modality or single-dimensional data for analysis and generation, which may limit the accuracy and robustness of the system when dealing with complex, cross-domain data; single-dimensional data cannot fully reflect the actual situation and may overlook the multimodal correlations and mutual influences between data; data alignment is difficult: different types of modal data (such as text, images, and audio) differ significantly in semantic level, temporal information, and physical properties, making it difficult for traditional methods to align multimodal data; how to effectively integrate these multi-dimensional data for accurate analysis and generation remains a bottleneck for current technology. Various analytical models (such as Transformer, CNN, etc.) usually work independently and lack a collaborative mechanism, making it difficult to fuse the results of different models. How to achieve effective information sharing and result fusion in multimodal data processing is a challenge in current systems. When labeling different types of data, such as images, audio, and sensor data, there are problems of high error and time costs. In some cases, single-modal data labeling may not be accurate enough, especially in complex scenarios, where the labeling process is time-consuming and prone to errors.
[0003] Therefore, how to process data from different modalities is an urgent problem to be solved. Summary of the Invention
[0004] The purpose of this invention is to provide at least one method, system, and electronic device for generating multimodal data of multidimensional data, which can at least solve the technical problems such as lack of alignment, lack of coordination, and labeling errors that exist when multiple single-modal data are applied, and can at least achieve the technical effect of alignment, coordination, and synchronization when multimodal data is applied.
[0005] To address the aforementioned technical problems, at least one embodiment of this application provides a method for multi-dimensional data... The modal data generation method includes: acquiring raw data in at least two dimensions; integrating and synchronizing the raw data in each dimension; labeling the data characteristics of different dimensions using corresponding analysis methods; extracting semantic information and structured features of each dimension; converting the data in each dimension into structured information to obtain the labeling results of each dimension, forming initial multimodal data; analyzing each labeling result using a corresponding analysis model to obtain analysis results for each modality; assigning appropriate weights to each modal analysis result based on the credibility, historical accuracy, and importance of the raw data; fusing the analysis results using a weighted fusion method to generate multimodal data; evaluating the analysis results; and revising the analysis methods and models based on historical labeling and analysis results to obtain iterative analysis methods, models, and weights for generating final multimodal data.
[0006] At least one embodiment of this application also provides a multi-dimensional data multimodal data generation method system, including: a data acquisition module for acquiring data of various modalities; a data annotation module for annotating data of different modalities; a data analysis module for analyzing the annotated data of different modalities to obtain analysis results; a weight adjustment module for adjusting the weights of the analyzed data of different modalities; a generation module for generating new data based on the weights of the data of different modalities and the analysis results; and an optimization module for optimizing the weights and updating the system based on historical data feedback.
[0007] At least one embodiment of this application also provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the above-described method for generating multimodal data of multidimensional data.
[0008] At least one embodiment of this application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for generating multimodal data of multidimensional data.
[0009] The embodiments of this application provide a method, system, and electronic device for generating multimodal data of multidimensional data. By annotating multidimensional data, the original data of different dimensions is transformed into structured information, providing valuable input for the multimodal data generation system. A deep learning model is used to analyze the data, realizing the fusion of relationships between different modal data. Combined with weights, new data is generated, thus realizing the fusion of multimodal data.
[0010] In some optional embodiments, the at least two dimensions of raw data include: text data, image data, audio data, and environmental data. The synchronization of the raw data of each dimension includes: time synchronization and spatial synchronization. Time synchronization is used to ensure the correspondence of data of each dimension at the same time. Spatial synchronization is used to spatially align image data, environmental data, and video data, aligning images from different perspectives to the same coordinate system.
[0011] By synchronizing different modal data spatially and temporally, the correlation of multimodal data during analysis is ensured.
[0012] In some optional embodiments, the annotation of data for each modality using corresponding analysis methods includes: text annotation of text data using natural language description, image annotation of image data using computer vision models, audio annotation of audio data using audio processing models, environmental annotation of environmental data, analysis of data for each dimension using computational methods, and third-domain annotation of the analysis results; text annotation includes text description annotation and text association annotation; text description annotation is used to convert the collected data into natural language descriptions, and text association annotation is used to generate more related text information for the data based on existing knowledge bases or contextual information; image annotation is used to annotate image data and extract key features and information from the image; audio annotation is used... The method for annotating audio data and extracting audio feature information includes: FFT, neural network model, and Z-transform. FFT can be used to convert time-domain data to frequency domain and extract frequency domain features; neural network model is used to learn the nonlinear mapping relationship between various modal data and encode multi-source inputs into a joint embedding space; Z-transform is used to eliminate scale differences caused by different units and sampling frequencies, so that data from different sources can be compared and fused under a unified statistical distribution; after calculation and analysis, the obtained cross-modal association results are annotated in a third domain. The third domain annotation results are used as a key component of the state space of the deep reinforcement learning policy network, or as a direct basis for generating personalized training plans, realizing a complete closed loop from multimodal data to intelligent decision-making.
[0013] By labeling data of different modalities, effective and structured semantic information representations are provided for different types of raw data. The labeling results are integrated according to different task requirements, ultimately providing valuable input for multimodal data generation systems. Using different analysis methods to label different types of data enables feature extraction based on different data types, ensuring feature accuracy.
[0014] In some optional embodiments, the analysis of each annotation result using a corresponding analysis model includes: using a first model to analyze text description annotations and text association annotations to obtain a preliminary understanding and information extraction of the text content; using a second model to analyze image annotations to obtain feature information in the image; and using a third model to analyze audio annotations to extract features from the audio data; the first model includes: a Transformer model and / or an LSTM model, the second model includes: a convolutional neural network, and the third model includes: a convolutional neural network.
[0015] By employing different analytical models to analyze different types of data, feature extraction can be achieved based on the different types of data, ensuring the accuracy of the features.
[0016] In some optional embodiments, the method of assigning appropriate weights to each modality analysis result based on the credibility, historical accuracy, and importance of the original data includes: assigning higher weights to data with high credibility; assigning higher weights to data sources that have shown high accuracy in tasks or analyses prior to the current time; assigning higher weights to data from specific sources when data from specific sources is more important than other data in a specific task; and assigning random weights to data based on public opinion preferences.
[0017] By analyzing the importance of data from different dimensions and assigning weights based on their importance, a scientific allocation of weights was achieved, improving the accuracy of the fusion results. The weighting of data from different modalities ensured the accuracy and reliability of the analysis results, better reflecting the contribution of each data dimension to the final outcome.
[0018] Using deep learning models for data analysis improves the ability to understand data, establishes relationships and dependencies between different types of data, and achieves the fusion of data from different modalities.
[0019] In some optional embodiments, the weighted fusion method includes: weighted average and weighted summation; the method of using weighted fusion to fuse the results of each modality analysis to generate the analysis result includes using linear average and / or square average and / or cubic average and / or exponential average to weight and accumulate each modality analysis result separately, and calculating the analysis result according to the weighted fusion formula based on the weight of each modality analysis result; the weight allocation is continuously adjusted according to the performance of historical data and task requirements.
[0020] Different fusion methods are used for fusion, and weights are adjusted to improve the accuracy of fusion.
[0021] In some optional embodiments, the correction of annotations and weights includes: during the annotation and analysis process, for data with annotation errors or excessively high time costs, correction is performed using existing data similar to the target data; the correction methods include: proportional assignment change and cross-domain mapping change; wherein, proportional assignment change includes: calculating the difference between the target data and similar data, adjusting it according to a certain proportion to generate more accurate annotations; cross-domain mapping change includes: mapping the target data from one modality to another, and then using data from similar domains for inference and correction; combining the corrected multimodal initial data and analysis results, using a deep learning model to generate the final multimodal data; and using model evaluation metrics to evaluate the quality of the generated final multimodal data results.
[0022] By correcting the data of different modalities, the accuracy of data annotation in each dimension was ensured, the time cost of the annotation process was reduced, and the final multimodal data results were generated through a deep learning model, achieving perfect integration of multimodal data.
[0023] In some optional embodiments, the iterative analysis method, analysis model, and weights include: obtaining historical data feedback results based on historical annotation data and historical analysis results, the historical data feedback results including: the accuracy of the annotation data, the correctness of the analysis results, and the evaluation efficiency; and performing error analysis, improving the model, optimizing the annotation process, and adjusting parameters and weights based on the historical data feedback results to obtain the iterative analysis method, analysis model, and weights.
[0024] Using historical data for iteration improves the accuracy of the iteration, provides a precise analytical model for subsequent analysis, and increases the accuracy of subsequent analysis. Attached Figure Description
[0025] One or more embodiments are illustrated by way of example with reference to the accompanying drawings, and these illustrative descriptions do not constitute a limitation on the embodiments.
[0026] Figure 1 This is a schematic diagram of the system structure provided in one embodiment of this application. Detailed Implementation
[0027] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the various embodiments of this application will be described in detail below with reference to the accompanying drawings. However, those skilled in the art will understand that many technical details have been provided in the various embodiments of this application to help readers better understand this application. However, the technical solutions claimed in this application can be implemented even without these technical details and various changes and modifications based on the following embodiments. The division of the various embodiments below is for the convenience of description and should not constitute any limitation on the specific implementation of this application. The various embodiments can be combined with and referenced by each other without contradiction.
[0028] To address the aforementioned technical problems of misalignment, errors, and difficulty in coordinating multiple single data sets, this invention proposes a multimodal data generation method for multidimensional data. The implementation details of this multimodal data generation method are described below. The following content is provided for ease of understanding and is not essential for implementing this solution.
[0029] Example 1:
[0030] This embodiment of the multi-dimensional data multimodal data generation method can be applied to electronic devices with communication, computing and data storage capabilities.
[0031] This includes: acquiring raw data from at least two dimensions in the same scenario, labeling, integrating and synchronizing the raw data from each dimension, labeling the data from different dimensions using corresponding analysis methods, obtaining the labeling results of the data from each dimension, and forming multimodal initial data.
[0032] Raw data from different dimensions, including: text data, image data, audio data, environmental data, etc.
[0033] The original data of each dimension is synchronized, including time synchronization and spatial synchronization. Time synchronization is used to ensure the correspondence of data of each dimension at the same time. Spatial synchronization is used to spatially align image data, environmental data and video data, aligning images from different perspectives to the same coordinate system.
[0034] The original data of different dimensions are labeled using corresponding analysis methods. Text data is labeled using natural language description, image data is labeled using computer vision models, audio data is labeled using audio processing models, and environmental data is labeled using environmental methods. Computational methods are used to analyze the data of each dimension, and the results are labeled in the third domain. Text labeling includes text description labeling and text association labeling. Text description labeling is used to convert the collected data into natural language descriptions, and text association labeling is used to generate more related text information for the data based on existing knowledge bases or context information. Image labeling is used to label image data and extract key features and information from the images. Audio labeling is used to label audio data and extract audio feature information. Third domain labeling is used to abstractly classify or state the results of the fusion analysis and construct a unified representation space across modalities.
[0035] The computational methods include FFT, neural networks, and Z-computation.
[0036] At the level of raw annotation of multimodal data, it is necessary to use corresponding analysis methods to independently annotate the data characteristics of different modalities in order to extract semantic information and structured features of each dimension.
[0037] Text annotation includes text description annotation and text association annotation.
[0038] For text data, such as athlete training logs, coach tactical notes, and post-match interview records, natural language processing technology is used for text annotation, specifically including entity recognition, sentiment polarity analysis, topic modeling, and keyword extraction, generating structured text feature vectors with timestamps and semantic labels. Entity recognition includes: athlete name, tactical terminology, and injury location; sentiment polarity analysis includes: positive, negative, and neutral; topic modeling includes: fatigue complaints and tactical execution evaluations.
[0039] At the level of text annotation refinement, in addition to conventional entity recognition and sentiment analysis, a fine-tuning strategy for large language models based on cue learning is further introduced. Dedicated annotation templates are constructed for special corpora in the sports domain, such as handwritten tactical board recognition, abbreviations in video subtitles, and non-standard expressions on athletes' social media. Specifically, LoRA or Adapter techniques are used to perform lightweight fine-tuning of LLaMA or ChatGLM-like models, enabling the models to extract quantitative labels of fatigue levels from metaphorical expressions like "my legs felt like lead in the last five minutes." Simultaneously, thought chain cues guide the model to output interpretable annotation basis, such as appending the mapping logic between key phrases in the original text and fatigue levels in natural language to the annotation results. Furthermore, text annotation can incorporate a time-aware mechanism, jointly annotating different events in the same text sequence according to their temporal relationships. For example, annotating the temporal evolution pattern of fatigue accumulation from a progressive description like "felt great in the first half, started cramping after halftime, and was forced to leave the field in the 20th minute of the second half," rather than annotating isolated single sentences.
[0040] Image annotation includes object recognition and scene analysis.
[0041] For image data, such as keyframes of competition videos, pose capture images, and venue environment photos, computer vision models are used for image annotation. Common methods include object detection and classification based on convolutional neural networks (such as ResNet and EfficientNet), pose estimation based on OpenPose or MediaPipe, and action recognition based on temporal convolutional networks, which output the position of people in the image, joint coordinates, action category, and confidence score.
[0042] At the level of image and video annotation refinement, the focus should be on addressing occlusion, rapid motion blur, and multi-view fusion issues in moving scenes. For limb occlusion, a common phenomenon in multi-player adversarial scenarios, a graph convolution-based pose completion network can be used. This network utilizes the human skeleton to perform geometric inference to complete occlusion key points, while outputting completion confidence as a weight reference for subsequent fusion. Kinematic constraints include: constant bone length and limited joint angle range. For image blurring caused by high-speed motion, an asynchronous low-latency data stream based on an event camera can be introduced as a supplementary modality. The event camera outputs a pixel-level brightness change event stream with a temporal resolution down to the microsecond level. Through spatiotemporal alignment and fusion with traditional frame images, sharper motion trajectories can be reconstructed from high-speed sprints, swings, and other actions. In multi-view video analysis, an implicit 3D reconstruction method based on neural radiation field is adopted to project the annotation results of 2D images from multiple fixed perspectives into a 3D spatial field, thereby obtaining the continuous movement trajectory of all athletes without blind spots. On this basis, the relative velocity vectors between players, the dynamic changes in defensive distance, and the spatiotemporal heat zone of local numerical advantage are annotated. The 2D image annotation results include: key points and ball position.
[0043] Audio annotation includes sound classification and speech recognition.
[0044] For audio data, such as athletes' breathing sounds, coach's instruction recordings, and audience noise, audio processing models are used for audio annotation. For example, Mel frequency cepstral coefficients (MFCC) combined with short-time Fourier transform are used to extract acoustic features. Then, pre-trained audio classification models are used to identify labels such as the degree of breathing rapidity, the content of voice instructions, and the intensity of environmental noise. Audio classification models include VGGish and AudioMAE.
[0045] At the level of audio annotation refinement, attention should be paid to the separation of target audio under environmental noise and the refined extraction of physiological acoustic features. Actual competition environments often contain a large amount of background noise (such as spectator shouts, broadcasts, and wind noise), which poses a serious challenge to the extraction of weak signals such as breathing sounds and footsteps. Background noise includes spectator shouts, broadcasts, and wind noise. To address this, a deep learning-based blind source separation model (such as Sudo RM-RF or Conv-TasNet) can be used to decompose the original mixed audio, separating the main audio channels collected by the athlete's near-field microphone. Then, respiratory cycle detection and gait frequency analysis can be performed on the separated signals. Blind source separation models include Sudo RM-RF or Conv-TasNet. Furthermore, vibration signals can be collected using bone conduction microphones, bypassing airborne environmental interference, to directly obtain the athlete's voice and heart rate rhythm. These sensors are worn behind the ear or on the neck and are not sensitive to motion interference. The collected vibration signals, after wavelet denoising, can be used to extract the ultra-low frequency components in heart rate variability, providing more robust physiological indicators for fatigue assessment. In addition, audio annotation should include semantic transcription and intent recognition of coach instructions and on-field communication, chronologically aligning brief instructions such as "marking," "covering," and "passing back" with the on-field situation, thereby providing a semantic reference for tactical execution evaluation.
[0046] For environmental data, such as temperature, humidity, air pressure, wind speed, and site hardness, environmental labeling is performed, including discretization of continuous values, such as high temperature / normal temperature / low temperature, outlier marking, and environmental risk level classification related to athletic performance.
[0047] At the level of environmental data annotation, the concepts of micro-environment perception and dynamic risk prediction are introduced. Traditional environmental annotation often relies on data from fixed weather stations around the venue, but there is a spatiotemporal discrepancy between the sampling points and the athlete's actual location. An improved approach is to integrate a miniature environmental sensor array onto the athlete or mobile device to collect real-time data on temperature, humidity, airflow speed, and UV intensity around the individual, forming a personalized micro-environment trajectory. When annotating this micro-environment data, in addition to conventional classification and grading, the temporal rate of change of environmental parameters should be calculated, including the rate of temperature rise and the magnitude of humidity drop, as drastic environmental changes often have a more significant impact on athletic performance than steady-state extreme values. Combined with the athlete's physiological data, a personalized thermal load response model is established, annotating the individual's "safe exercise time window" and "forced rest trigger threshold" under current environmental conditions. This dynamic environmental annotation result will be directly input into the real-time strategy generation module, along with physiological data, including core body temperature estimates and sweating rates.
[0048] After completing the independent annotation of each dimension, the data of each dimension are analyzed by computational methods and the analysis results are annotated in the third domain to construct a unified representation space across modalities.
[0049] Commonly used computational methods include: Fast Fourier Transform (FFT), neural network models, and signal processing and statistical techniques such as Z-transform.
[0050] FFT is used to convert time-domain motion data such as accelerometer, gyroscope, and electromyography signals into the frequency domain, and extract frequency domain features such as power spectral density, dominant frequency distribution, and harmonic energy. These features can reveal the athlete's stride frequency stability, muscle vibration patterns, and fatigue-related frequency shift phenomena.
[0051] Neural network models, such as variational autoencoders, LSTMs, or graph neural networks, are used to learn the nonlinear mapping relationships between various modal data, encoding multiple source inputs such as text sentiment tags, image pose keypoints, audio fatigue features, and environmental risk levels into a joint embedding space.
[0052] Z-transform (normalization) is used to eliminate scale differences caused by different units and sampling frequencies, enabling data from different sensors to be compared and fused under a unified statistical distribution.
[0053] After the above calculation and analysis, the obtained cross-modal association results are labeled with a third domain, such as "heart rate frequency domain main peak decrease + respiratory audio spectrum entropy increase + negative text sentiment → high-risk fatigue state" and other multimodal joint patterns.
[0054] Third-domain annotation is essentially an abstract category label or state classification of the fusion analysis results. It can include quantifiable meta-labels such as fatigue level (low / medium / high), injury risk probability value, tactical execution effectiveness score, and attention concentration level. Unlike the original unimodal annotation, third-domain annotation has the characteristic of cross-modal comprehensive inference, representing high-level cognitive states or situational categories directly related to athletic performance extracted from multi-source heterogeneous data. The results of third-domain annotation will serve as a key component of the state space of deep reinforcement learning policy networks, or as a direct basis for generating personalized training plans, thereby truly realizing a complete closed loop from multimodal data to intelligent decision-making.
[0055] At the level of refining computational methods and third-domain annotation, special emphasis needs to be placed on adversarial alignment and uncertainty quantification at the feature level. Since data from different modalities differ significantly in acquisition frequency, latency, and noise characteristics, direct concatenation or weighted fusion may lead to a high-noise modality dominating the final annotation result. The solution is to employ domain adversarial neural networks, training a domain discriminator during the feature extraction stage. This forces the feature extractors of each modality to learn shared feature representations that are independent of the modality's origin and only relevant to the task, thereby enhancing the generalization robustness of third-domain annotation. Simultaneously, for traditional signal processing methods such as FFT and Z-transform, they can be parameterized and embedded into an end-to-end differentiable architecture. For example, the window length and overlap rate of FFT can be designed as learnable hyperparameters, automatically optimized by backpropagation of the validation loss during training, rather than being manually fixed. For the uncertainty of the third domain label, the prediction of the distribution form should be output. The prediction of the distribution form includes obtaining the prediction variance through Monte Carlo dropout or deep ensemble model, rather than a single point estimate, and decomposing it into random uncertainty and cognitive uncertainty. Random uncertainty is caused by data noise, while cognitive uncertainty is caused by insufficient model knowledge. The former can be reduced by sensor calibration and data cleaning, while the latter suggests the need to collect more training samples in this scenario or introduce prior knowledge constraints.
[0056] Finally, all the above annotation results and third-domain annotations should be organized into standardized data frames with timestamps, confidence levels, modality sources, and version identifiers, and stored in a traceable time-series database for subsequent policy model queries and training. The annotation process itself is designed as an incremental iterative mode, that is, after each competition or training session, the newly generated multimodal data and corresponding third-domain annotations will be used to fine-tune the annotation models of each modality, enabling the system to gradually adapt to the individual characteristics of the athlete, the special rules of the sport, and the unique attributes of the venue environment, thereby achieving continuous improvement in annotation accuracy and enhanced cross-scene transfer capabilities.
[0057] By labeling data from different dimensions, the raw data is transformed into structured data information, resulting in text-labeled data, image-labeled data, audio-labeled data, environmental-labeled data, and third-domain-labeled data, forming multimodal initial data.
[0058] The corresponding analysis model is applied to the annotation results of each modality, and the annotation data of different modalities are analyzed to obtain the analysis results of each modality, including: The first pair of text-annotated data was analyzed to obtain the text data.
[0059] The first model includes: Transformer model and / or LSTM (Long Short-Term Memory Network).
[0060] The second model is used to analyze the image annotation data to obtain the image data.
[0061] The third model was used to analyze the audio annotation data to obtain the audio data.
[0062] The second and third models respectively include: convolutional neural networks and / or graph neural networks.
[0063] Based on the reliability, historical accuracy, and importance of the original data for each modality, appropriate weights are assigned to the analysis results for each modality, including: Data with high credibility should be assigned higher weights; data sources that demonstrate high accuracy in tasks or analyses prior to the current time should be assigned higher weights; data from specific sources should be assigned higher weights when they are more important than other data in a specific task; and data should be assigned random weights based on public opinion preferences.
[0064] The results of the modal analyses are merged using a weighted fusion method to generate multimodal data. Weighted fusion methods include: weighted average and weighted summation.
[0065] Weighted averages include: linear average, squared average, cubic average, and exponential average.
[0066] During the annotation and analysis process, outlier data with errors or excessively high time costs are corrected by using existing data similar to the target data.
[0067] Correction methods include proportional assignment correction or cross-modal mapping correction.
[0068] Proportional assignment changes include: calculating the differences between the target data and similar data, adjusting them according to a certain proportion, and generating more accurate annotations.
[0069] Cross-domain mapping transformations include mapping target data from one modality to another, and then using data from similar domains for inference and correction.
[0070] Using a deep learning model, based on the corrected initial multimodal data, analysis results, and weights, the final multimodal data is generated and output in different formats. The quality of the generated final multimodal data results is evaluated using model evaluation metrics.
[0071] Model evaluation metrics include: multimodal consistency detection and evaluation of generated results.
[0072] We optimize based on historical data, adjust system parameters based on feedback from real-world applications, and continuously learn and improve by incorporating new data.
[0073] This includes: obtaining historical data feedback results based on historical annotation data and historical analysis results, including: the accuracy of the annotation data, the correctness of the analysis results, and the evaluation efficiency; and conducting error analysis, improving the model, optimizing the annotation process, and adjusting parameters and weights based on the historical data feedback results to obtain the iterative analysis method, analysis model, and weights.
[0074] Example 2: This embodiment is a detailed description of Embodiment 1.
[0075] Multi-dimensional data acquisition is not limited to a single type of data. In practical applications, data can come from multiple fields, including but not limited to various data sources such as vision, hearing, and environmental perception. Each data type has its specific characteristics and application scenarios. Combining these multimodal data can improve the intelligence level of the system and the accuracy of data generation.
[0076] Data Acquisition: Data acquisition methods include using sensors and text to acquire multi-dimensional data.
[0077] Sensor data is one of the fundamental sources of multidimensional data. Different types of sensors can acquire different forms of data information.
[0078] Sensors come in various forms, including visual sensors, sound sensors, and environmental sensors, and acquire different types of data, including visual data, sound data, and environmental data.
[0079] Visual sensors, such as cameras, are used to acquire image or video data. The visual data acquired by cameras is a two-dimensional image or video stream, which can be processed and analyzed using computer vision techniques such as convolutional neural networks (CNNs). Visual data acquired by cameras is widely used in tasks such as object recognition and image generation; the image data is typically in the form of a two-dimensional matrix.
[0080] The expression for image data I is as follows: (1); in, The function representing the data captured by the camera. Represents the spatial coordinates of the image. Represents the time coordinate.
[0081] Sound sensors, such as microphones, are used to acquire audio data and capture sounds in the environment. This audio data can be analyzed using audio processing techniques such as convolutional neural networks (CNNs). The audio data acquired by microphones is used for tasks such as sound classification, speech recognition, and environmental sound analysis; the audio data is typically a time-series signal.
[0082] The expression for audio data A is as follows: (2); Represents a timestamp. This represents a function that allows the microphone to acquire audio signals.
[0083] Environmental sensors, such as temperature and humidity sensors, are used to acquire physical data about the environment, such as temperature, humidity, and air pressure. This sensor data is typically continuous numerical data that describes the environmental state at a specific moment.
[0084] Environmental sensors acquire data such as temperature, humidity, air pressure, and light, which are typically numerical data sampled at regular intervals and are suitable for fields such as environmental monitoring and smart homes.
[0085] The expression for the temperature sensor data T is as follows: (3); This function represents the data collected by the temperature sensor, where T represents the temperature value. Represents a timestamp.
[0086] Natural language data, i.e., text descriptions:
[0087] In some scenarios, it is also necessary to acquire text data describing natural language, such as user-input commands or descriptions, or text output from sensors, such as "temperature too high." This data is usually unstructured text data, suitable for tasks such as text generation and sentiment analysis.
[0088] Multi-dimensional data acquisition equipment includes cameras, microphones, and environmental sensors, such as temperature sensors, humidity sensors, and barometric pressure sensors.
[0089] Cameras include still cameras, moving cameras, thermal imaging cameras, etc., and are used to capture image or video data.
[0090] A microphone is used to record audio signals and can capture sounds from different directions and frequency ranges.
[0091] Environmental sensors include temperature sensors, humidity sensors, and barometric pressure sensors, which can output analog or digital signals.
[0092] Sensor interface technologies, such as I2C, SPI, USB, and Bluetooth, are used for connecting sensors to computing devices and transmitting data.
[0093] Data integration and synchronization:
[0094] How to effectively integrate and synchronize data from different sources is a key problem that needs to be solved in multimodal systems.
[0095] Sensor data is typically collected in different time and spatial dimensions, so effective time synchronization and spatial alignment are required to ensure the correlation between modes during multimodal analysis.
[0096] For example, image and audio data need to be guaranteed to be sequential.
[0097] Synchronization methods include time synchronization and spatial synchronization.
[0098] Time synchronization: Timestamps are used to mark data to ensure the correspondence between data of different dimensions at the same time.
[0099] Spatial alignment: For image and video data, spatial alignment may be required to align image data from different perspectives to the same coordinate system.
[0100] Synchronization processes typically rely on signal processing techniques, including algorithms such as Kalman filtering and timestamp matching.
[0101] Data annotation: Data labeling is one of the core steps in a multimodal data generation system. It involves labeling raw sensor data or other multidimensional data using various methods to facilitate subsequent analysis and model training.
[0102] Data annotation encompasses multiple dimensions, such as text annotation, image annotation, and audio annotation. Text annotation includes descriptive text annotation and associative text annotation. Each annotation method corresponds to a specific task and technique, aiming to provide effective semantic information and structured representations for different types of raw data.
[0103] Text annotation: Transform sensor data or raw data into natural language descriptions and perform text annotation.
[0104] Text description annotation: Generate a corresponding natural language description for each segment of sensor data.
[0105] Text association annotation: Based on existing knowledge bases or contextual information, it generates more related textual information for the data, including related events and scenario descriptions in the context. This process takes into account the data's contextual environment, historical information, background knowledge, etc., thereby providing the data with more semantic information.
[0106] Image annotation: For image data, annotation is performed using computer vision models (such as CNN) to extract key features and information, such as using CNN convolutional neural networks to extract object recognition, scene analysis, etc.
[0107] Audio annotation: For audio data, the sound is annotated using audio processing models (such as CNN) to extract the feature information of the sound. For example, CNN convolutional neural networks can be used to extract sound classification, speech recognition, etc.
[0108] Specifically, Natural Language Generation (NLG) technology is used to extract key features from the raw data and transform them into descriptive text.
[0109] Image annotation: Image annotation involves transforming image data into structured labels or descriptions.
[0110] Image annotation tasks typically employ computer vision techniques to extract key features from images and then label the images based on these features. The goals of annotation can include tasks such as object recognition, scene understanding, and object detection.
[0111] Computer vision techniques include convolutional neural networks (CNNs).
[0112] Audio annotation: Audio annotation aims to extract speech or audio features from audio data and generate annotation information based on these features.
[0113] Common audio annotation tasks include sound classification and speech recognition. Audio data is typically analyzed using audio processing models to extract temporal or frequency domain features.
[0114] Audio processing models include convolutional neural networks (CNNs).
[0115] After annotation, all the raw data is transformed into structured information, resulting in text annotation, image annotation, and audio annotation results. These annotation results are then integrated according to different task requirements, ultimately providing valuable input for the multimodal data generation system.
[0116] The goal of the model analysis phase is to conduct in-depth analysis and information extraction of different types of labeled data using deep learning models.
[0117] To enhance data understanding capabilities, in addition to using traditional deep learning models such as Transformer, LSTM, and CNN, this application also introduces Graph Neural Networks (GNNs) to process and fuse relationships between different modalities. GNNs are capable of handling graph-structured data and are particularly suitable for establishing complex associations and dependencies between multimodal data.
[0118] The goal of text annotation data analysis is to extract information from text description annotations and text association annotations, understand the text content, and provide valuable features for subsequent tasks.
[0119] Text data analysis mainly uses Transformer and LSTM models for model training and information extraction.
[0120] The goal of image data analysis is to extract feature information from images using convolutional neural networks (CNNs).
[0121] In multimodal data analysis, the relationships between image data and other modalities are also important. Graph neural networks (GNNs) can play a role in these tasks, especially in establishing connections between multimodal data, helping to convey information in graph-structured data.
[0122] The goal of audio data analysis is to extract feature information from audio, such as audio type and speech content, using convolutional neural networks (CNNs). Similar to images, audio signals can be converted into spectrograms, and then CNNs can be used for feature extraction.
[0123] To better handle the connections between audio and other modalities (such as text modality and image modality) in multimodal data, GNNs can be used to fuse cross-modal features.
[0124] Weight assignment and result fusion: In a multimodal data generation system, weight assignment and result fusion are key steps to ensure that the analysis results of different modalities can be effectively combined and output as a final decision.
[0125] In this application, appropriate weights are assigned to the analysis results of different modalities (text, image, audio), and a weighted fusion method is used to generate the final analysis results. This process not only considers the credibility of the data in each dimension, but also combines historical accuracy and the importance of the original data, thereby ensuring the accuracy and reliability of the analysis results.
[0126] Weighting: Weights are assigned to each analysis result based on factors such as the reliability of the data dimensions, historical accuracy, and original data values. Appropriate weights are assigned to data analysis results from different modalities to better reflect the contribution of each data dimension to the final result.
[0127] Weighting is typically based on the following factors: Data Dimension Credibility: Data dimensions with higher credibility should be given higher weight in the result fusion. The level of credibility can be determined by assessing the quality of data in each dimension; high-quality sensor data or model analysis results usually have higher credibility.
[0128] Historical accuracy: Historical accuracy refers to the accuracy of a dimension in past tasks or analyses. For example, if a particular sensor has consistently demonstrated high accuracy, it should be given a higher weight in the current analysis.
[0129] The importance of raw data values: The importance of raw data values from different modalities varies across different application scenarios. For example, certain sensor data may be more important than other data in a specific task, and therefore these data should receive higher weight.
[0130] Dimensions with higher credibility and higher historical accuracy will be given higher weight.
[0131] Results fusion: The analysis results obtained from different modalities are weighted and accumulated to generate the final analysis result.
[0132] The weighted fusion method takes into account the importance of data from each dimension, thus yielding a more accurate final analysis.
[0133] Multimodal similarity association protection: In multimodal data generation systems, the core objective of multimodal similarity association protection is to improve system efficiency and accuracy by correcting data annotation errors in a certain dimension or reducing annotation time costs. When there are errors in the annotation of a certain data dimension or the annotation process takes a long time, the system can automatically find existing data similar to the target data and correct the target data based on these similar data.
[0134] Similar data correction: During the annotation and analysis process, if there are errors in the annotation of data in a certain dimension or the time cost is too high, the system will select existing data that is similar to the target data for correction.
[0135] Correction method: Proportional value assignment: Based on the existing labeled data, make corrections proportionally.
[0136] Cross-domain mapping transformation: Perform cross-modal mapping on data, and use data from similar domains to make inferences and corrections.
[0137] Multimodal similarity association protection leverages the similarity relationships between different modalities, primarily including similar data correction and two correction methods: proportional assignment change and cross-domain mapping change. Through this approach, the system can automatically correct annotation errors or fill in missing data, ensuring more accurate final analysis results.
[0138] Similar data correction: During the annotation and analysis of multimodal data, if inaccurate annotation of a certain dimension is found, or if the annotation process consumes too much time, the system will correct it based on existing similar data. Specifically, the system will identify and select the existing data sample most similar to the target data for reference and correction.
[0139] The selection of similar data is based on a similarity measure between data dimensions. Common similarity measurement methods include: Euclidean distance: used to calculate the distance between two data points in space, suitable for numerical data.
[0140] Cosine similarity: used to measure the similarity between two vectors, often used for similarity calculation of text data.
[0141] Accard similarity: Used to calculate the similarity between sets, suitable for discrete or set-type data.
[0142] By calculating the similarity between each labeled data sample and the target data, the system can select the most relevant sample for reference correction.
[0143] Proportional changes: In some cases, existing labeled data may not perfectly match the target data, but still possess a certain degree of similarity. To correct the labeling of the target data, a proportional adjustment method can be used to modify the existing data's labeling values. This method calculates the differences between the target data and similar data and adjusts them proportionally, thereby generating more accurate labels.
[0144] Data generation and output: In a multimodal data generation system, data generation and output are among the system's final tasks. Through the aforementioned analysis and correction processes, the system integrates data and analysis results from multiple modalities to ultimately generate multimodal data output that meets the requirements. This generated data can be new text descriptions, images, audio synthesis, and other multimodal content, providing support for subsequent application systems or users. This stage not only requires integrating the aforementioned annotation, analysis, and correction results but also ensuring that the generated data meets expectations in terms of semantics, format, and quality.
[0145] Based on the aforementioned analysis and correction process, the final multimodal data results are generated and output to the application system for further processing.
[0146] The output data can be new text descriptions, image generation, audio synthesis, and other multimodal content for use by other systems or users.
[0147] The data generation process involves the system's previous analysis and correction processes, and the final data generated depends on the requirements of the target task and data modality.
[0148] Generative tasks are typically achieved using deep learning models such as Generative Adversarial Networks (GANs), Variational Autoencoders (VAEs), Recurrent Neural Networks (RNNs), and Transformers.
[0149] Text generation: Generate new text descriptions by analyzing text annotations and model inference results.
[0150] Image generation: Based on the image analysis results, use a generative model to generate new images or modify existing images.
[0151] Audio synthesis: generating new audio data through audio analysis models, or generating new audio content based on existing audio.
[0152] Text generation: Text generation is typically accomplished using language models (such as Transformer and GPT series models), which generate new text content based on input annotation information or analysis results. The generated text needs to be adaptively corrected based on the previously labeled data.
[0153] Language models include: Transformer and GPT series models.
[0154] Output data format: The generated data results can be output in different formats, depending on the application requirements. The generated multimodal data can be output through API interfaces, file storage systems, or real-time streaming services for further processing by other systems or users.
[0155] Text data: can be output as plain text files, JSON format, or CSV format.
[0156] Image data: can be output as image files, such as PNG, JPEG, SVG, etc., or embedded in web pages or applications.
[0157] Audio data: It can be output as audio files, such as MP3, WAV, etc., or provided to users in the form of streaming media.
[0158] Data quality control: The generated multimodal data output needs to be not only accurate but also of high quality. To ensure data quality, the following methods can be used: Multimodal consistency detection: Ensure consistency among modalities such as text, images, and audio, meaning they should collectively describe the same object or event, avoiding semantic conflicts.
[0159] Evaluation of generated results: The quality of generated text, images and audio is evaluated using model evaluation metrics (such as BLEU score, FID score, etc.) to ensure that they meet the generation standards.
[0160] Model evaluation metrics include: Bilingual Evaluation Alternate (BLEU) score and Fraser Initial Distance (FID) score.
[0161] Bilingual Evaluation Understudy (BLEU) was initially used for automatic evaluation in machine translation tasks and is now widely used in various text generation scenarios, such as tactical description generation, automatic writing of athlete comments, and game commentary text generation. Its core idea is to measure generation quality by comparing the n-gram overlap between the generated text and one or more manually annotated reference texts. Specifically, the generated text and reference texts are first segmented into unigrams (single words), bigrams (two consecutive words), and so on up to n-grams (usually n=1 to 4). Then, the frequency of each n-gram in the generated text appearing in any reference text is counted, and the precision is calculated. To prevent cheating by repeating high-frequency words, such as repeatedly outputting the word "good" to obtain a high score for unigrams, BLEU introduces a shortening penalty factor, which discounts the score when the generated text is shorter than the reference text. The final BLEU score is the geometrically weighted average of the n-gram precision rates of each order multiplied by the shortening penalty factor, with the score ranging from 0 to 1. The closer the value is to 1, the more similar the generated text is to the reference text. Taking sports as an example, to assess the accuracy of an automatically generated "description of athlete fatigue," a professional assessment written by a coach or team doctor can be used as a reference text. A higher BLEU score indicates that the generated text is closer to the expression habits of experts at the vocabulary and phrase level. It is worth noting that the BLEU score focuses on the surface matching of words and phrases and has lower sensitivity to semantic fluency and logical coherence. Therefore, in actual assessments, it is usually necessary to combine it with human evaluation or other semantic similarity indicators.
[0162] The Fréchet Inception Distance (FID) is a core metric for evaluating the quality and diversity of generated images, particularly suitable for image generation tasks such as motion pose generation, tactical heatmap synthesis, and virtual scene rendering. FID calculation is built upon a pre-trained Inception v3 image classification network. This network, trained on large-scale image datasets, effectively captures semantic information from its intermediate layer features. Specifically, the real and generated image sets are input into the Inception v3 network, extracting 2048-dimensional feature vectors before the pooling layer. The real image set includes actual motion pose images extracted from competition videos, while the generated image set includes simulated pose images generated by the model based on tactical descriptions. Then, the mean vector and covariance matrix of the two feature vector sets are calculated separately. Finally, the Fréchet distance between these two multivariate Gaussian distributions is calculated, mathematically expressed as the square of the L2 norm of the difference between the real and generated feature means, plus the sum of the traces of the two covariance matrices, minus the square root trace of twice the product of the covariance matrices. A lower FID score indicates that the generated image is closer to the real image in terms of feature distribution, meaning it possesses both high visual quality and good diversity. Taking a specific sports scene as an example, to evaluate a model that generates a "basketball shooting sequence," we can use the captured postures of athletes in actual games as the real image set, and the postures generated by the model based on the text description "stop-and-shoot" as the generated image set. A lower FID score indicates that the generated postures are highly similar to those of real athletes in terms of joint angles, body coordination, and movement trajectory, and that not all generated results are copies of the same action template. FID is sensitive to common generation defects such as image noise, blur, and artifacts, and does not require individual manual scoring for each image, thus offering advantages in large-scale evaluation.
[0163] Feedback and optimization: This includes historical data feedback and model iteration.
[0164] Historical data feedback refers to the large amount of labeled data and analysis results accumulated during long-term operation. This data and results, as historical data, can serve as the basis for model optimization.
[0165] Optimize based on historical data and feedback, identify weaknesses or potential errors in the current model and annotation process, adjust model weight allocation and data annotation process, improve the accuracy and efficiency of the overall system, and thus provide a basis for subsequent optimization.
[0166] During operation, the system records each data generation and analysis result, forming historical data records. These records include: Accuracy of labeled data: Record the accuracy of each data labeling, such as the matching degree of text description, the accuracy of image recognition, and the correctness of audio recognition.
[0167] Correctness of analysis results: Record the model's prediction results in each modality and compare them with the actual results to evaluate the deviation of the analysis results.
[0168] System performance feedback: Evaluate the system's efficiency in actual tasks, including response time, processing speed, and resource consumption.
[0169] Applications of data feedback: By analyzing feedback from historical data, the system can extract the following types of information: Error analysis: Identify the types of errors that occur during the system annotation process and determine whether the decrease in accuracy is due to unreasonable model weight allocation, missing data, or deviations in the annotation process.
[0170] Model improvement: By evaluating the performance of different modal analysis results, the weight allocation and fusion strategy of different modal models are adjusted to improve the weighted accumulation process.
[0171] Annotation process optimization: Optimize the data annotation process based on historical feedback. For example, if the annotation accuracy of image data is low, the training of the image analysis model can be enhanced, or the image annotation strategy can be adjusted.
[0172] Historical data feedback is analyzed periodically to generate feedback reports, thereby helping to iteratively optimize the model.
[0173] Model iteration: This is the process of updating and optimizing an existing model based on feedback from historical data.
[0174] By analyzing the feedback data, the system automatically adjusts parameters, trains the model, and optimizes the process, thereby improving performance. Iteration typically involves the following steps: parameter adjustment, optimizing the data annotation process, and continuous learning using new data.
[0175] Parameter adjustments include: adjusting the weights of the modal analysis model and adjusting the parameters of the analysis algorithm.
[0176] Adjusting the weights of modal analysis models: If a certain mode performs poorly, its weight can be adjusted to increase its influence on the final result.
[0177] Adjust the parameters of the analysis algorithm: By analyzing the model performance in the feedback, adjust the parameter settings of the analysis algorithm, such as the learning rate, regularization term, batch size, etc.
[0178] Optimizing the data annotation process and establishing a feedback mechanism can help the system identify bottlenecks or inefficiencies in the annotation process, allowing for improvements.
[0179] If the accuracy of the text annotation process is not high, the system can enhance the training of the text annotation model or adjust the text annotation strategy, such as selecting more accurate natural language processing tools or adding contextual information.
[0180] By continuously learning from new data and training it with new data, the model can continuously learn and improve. New data further supplements the deficiencies of historical data and helps the system discover new patterns and rules. The core goal of continuous learning is to enable the system to gradually acquire adaptive capabilities, allowing it to operate effectively in different environments and tasks.
[0181] Based on feedback from practical applications, we adjust model parameters, optimize the data annotation process, and continuously learn and improve by incorporating new data.
[0182] In multimodal data generation systems, feedback and optimization are crucial for ensuring long-term effective operation. Through continuous learning and optimization, the system can adapt to new data and environmental changes, improving overall accuracy and efficiency. Feedback mechanisms not only help the system identify and correct potential errors but also effectively adjust model weight allocation, data labeling processes, and algorithm performance.
[0183] By using feedback from historical data, the system can continuously adjust the model's parameters and weights to produce better prediction results in practical applications.
[0184] Example 3:
[0185] This embodiment is a detailed description of Embodiments 1 and 2.
[0186] Natural language generation (NLG) technology is used to extract key features from raw data and transform them into descriptive text. This process usually requires combining contextual information and generating natural language descriptions that conform to syntax and semantics through a certain model.
[0187] Let the sensor data be , of which each A raw data point, such as temperature or humidity, can be represented using a text generation model, such as Transformer or GPT, as follows:
[0188] Here, Generate represents the process of generating natural text.
[0189] Specifically, assuming the sensor data indicates a temperature of 25°C, the text description could be: "Current temperature is 25 degrees Celsius".
[0190] Input data: , Output text description: .
[0191] Text association annotation takes into account the context, historical information, and background knowledge of the data, thereby providing more semantic information to the data.
[0192] The principle is as follows: Let the text association annotation model be The model uses existing context C and sensor data D to generate associative text L: (5); Where C represents contextual information related to the current sensor data; the text association annotation model is... This includes models based on knowledge graphs or contextual learning.
[0193] Specifically, let's assume the input data is: Context information: C .
[0194] Output suggested text: L = "Summer high temperature, the temperature has reached 35 degrees Celsius, the air conditioner may be turned on."
[0195] Object recognition uses convolutional neural networks (CNNs) to extract features from images and identify the features of objects within those images.
[0196] The principle is as follows: Let the image data be I, the CNN model be C, and the image annotation result be... The annotation process can then be represented as: (6); Where C represents the object detection or scene analysis process performed by the convolutional neural network (CNN).
[0197] Specifically, the input image could be a scene depicting a person and a car.
[0198] Output image annotation: .
[0199] Through audio processing technology, the system can identify these sound categories and generate corresponding tags.
[0200] The principle is as follows: Let the audio signal be A, and the audio processing model be... The audio annotation results are The annotation process can then be represented as: (7); Specifically, for example, input audio: containing the voice message "turn on the air conditioner".
[0201] Output audio annotations: .
[0202] The Transformer model computes relationships between words in a sequence through its self-attention mechanism, enabling it to capture global information in text and excelling particularly at modeling long-distance dependencies and contextual information. When processing text data, the Transformer model effectively understands the relationships between contexts and extracts key information.
[0203] Text data analysis mainly uses Transformer and LSTM models for model training and information extraction.
[0204] The principle of using the Transformer model for text data analysis is as follows: Given a text sequence X The core of the Transformer model is to calculate the attention of each word to all other words. The weights are calculated using the following formula: (8); Where Q, K, and V are the query, key, and value matrices, respectively. This represents the dimension of the key vector.
[0205] The advantage of the Transformer model lies in its ability to compute all word relationships in parallel, making it suitable for handling global dependency information in long texts.
[0206] Specifically, for example, input text: T="Temperature too high, turn on the air conditioner".
[0207] Output analysis: The Transformer model understands the contextual information in the text and extracts the causal analysis between "the temperature is too high" and "the air conditioner is turned on".
[0208] LSTM (Long Short-Term Memory) networks are able to capture temporal dependencies in text, with particularly significant effects in sequence generation and sentiment analysis tasks. LSTMs effectively handle long-term dependencies by introducing gating mechanisms such as input gates, forget gates, and output gates.
[0209] The principle of using the LSTM model for text data analysis is as follows: LSTM computation consists of three main gating parts: the input gate ( Forgotten Gate ), output gate ( These gates are used to control the flow of information: (9); (10); (11); This represents the hidden state at the previous time step. This indicates the current input.
[0210] Specifically, for example, input text: T = "The weather forecast shows there will be thunderstorms tonight"; Output Analysis: LSTM can extract the correlation between "weather forecast" and "thunderstorm" by analyzing the time series relationship of text.
[0211] Convolutional Neural Networks (CNNs) are used for image data analysis to extract spatial features from images. CNNs can identify details such as objects, edges, and textures in images and generate high-dimensional feature representations.
[0212] The principle is as follows: Input image I, after being processed by convolutional layer C, yields feature map F: (12); The convolution operation C includes convolution kernels, pooling, and activation functions.
[0213] CNNs gradually abstract high-level features of images through multiple layers of convolution and pooling operations, and finally map them into classification labels or regression values through fully connected layers.
[0214] Specifically, the input image is a photo of a cat.
[0215] Output analysis: CNN extracts the "cat" object features from the image.
[0216] Graph Neural Networks (GNNs) are effective tools for processing graph-structured data. In multimodal data analysis, GNNs are particularly well-suited for handling relationships between modalities, such as the connections between text and images, or images and audio.
[0217] The principle behind using graph neural networks (GNNs) for image data analysis is as follows: Given a graph G = (V, E), where V are nodes representing different modalities of data, such as images and text, and E are edges representing the relationships between these modalities, the features of each node can be represented as follows: The characteristics of each edge are: .
[0218] Graph neural networks update the features of each node using information from its neighboring nodes. The update process is as follows:
[0219] (13); in, Represents a node Features at the k-th layer It is a node The neighboring nodes, It is a message passing function.
[0220] Specifically, the input consists of text and image modal data. The graph neural network (GNN) associates and fuses the feature information of the image and text to obtain a common feature representation.
[0221] Convolutional Neural Networks (CNNs) are used to extract feature information from audio, such as audio type and speech content. Similar to images, audio signals can also be converted into spectrograms, and then CNNs can be used for feature extraction. To better handle the relationships between audio and other modalities (such as text and images) in multimodal data, Generative Neural Networks (GNNs) can be used for cross-modal feature fusion.
[0222] The principle behind using a convolutional neural network (CNN) to extract feature information from audio is as follows: The audio signal A is subjected to a short-time Fourier transform (STFT) to obtain its spectrum S: S = STFT(A) (14); A convolutional neural network (CNN) is used to process the spectrogram and extract audio features: F=C(S) (15); Where C represents the audio features obtained by CNN processing the spectrogram.
[0223] Specifically, input audio: a voice message containing "Hello, how's the weather today?"
[0224] Output analysis: CNN can identify speech content in audio and perform speech recognition or classification.
[0225] By employing a graph neural network (GNN), the features of the audio signal are combined with features from other modalities, such as image and text modalities, to perform cross-modal information fusion. GNN helps to transmit information between nodes, enabling the model to integrate features from various modalities and generate a unified representation.
[0226] The principle is as follows: Assume the feature representation of the audio modality is as follows The features of the image modality are The feature representation of the text modality is as follows Then GNN fuses these features through a message passing mechanism: (16); in, Includes audio, image, and text nodes. Edge features representing relationships between modalities.
[0227] Specifically, the input includes audio, image, and text modal data; Output: A comprehensive cross-modal feature representation is obtained through the fusion of GNN models.
[0228] The principle of weight allocation is as follows: Suppose there are three different data dimensions , , , representing text analysis results, image analysis results, and audio analysis results, respectively. The weights for each dimension are calculated based on a comprehensive assessment of credibility, historical accuracy, and the importance of the original data. , , The weighted calculation formula is as follows: (17); in, This represents the data credibility of the i-th dimension. Indicates the historical accuracy of the i-th dimension. This indicates the importance of the original data in the i-th dimension. These are weighting coefficients used to balance the importance of various factors. Typically, these coefficients are adjusted according to the needs of the task.
[0229] In one specific embodiment of this application, it is assumed that in a certain scenario, the credibility of text data... =0.85, historical accuracy =0.90, Importance of raw data =0.80, then the weight of the text data The calculation formula is as follows: ; The same method was used to calculate the weights for image data and audio data respectively.
[0230] After obtaining the weights of each data dimension, the analysis results of different modalities are fused according to the weighted cumulative method to generate the final analysis result. The fusion method includes weighted average, weighted summation, or other forms of weighted fusion.
[0231] The principle behind the weighted average fusion of data analysis results is as follows: Assume the analysis results for each data dimension are as follows: , , The corresponding weights are respectively , , So, the final fusion result The calculation formula is as follows: (18); in, It is the final weighted fusion result, representing the decision after comprehensively considering the data analysis results of the three dimensions of text, image, and audio.
[0232] In one specific embodiment of this application, it is assumed that the text analysis result... 0.75, weight Image analysis results 0.65, weight Audio analysis results 0.9, weight According to the weighted average calculation formula:
[0233] Therefore, the final analysis results , representing the decision value after comprehensive system analysis.
[0234] Weighting and optimization: To improve the accuracy of the fusion results, the weight allocation is continuously adjusted based on the performance of historical data and task requirements. For example, the weight coefficients are dynamically adjusted based on the model's validation results on different datasets. When the data of a certain dimension performs well in certain tasks, the weight of that dimension can be increased, and vice versa.
[0235] For a given task T, after multiple training cycles, the system will adjust the weights based on the performance of each dimension on that task.
[0236] The formula for dynamically adjusting the weights is as follows:
[0237]
[0238] in, It is the weight of the i-th dimension in the previous period; It is the amount of weight adjustment based on the analysis results of task T, which is usually related to the change in the accuracy of that dimension; It is the learning rate, which controls the magnitude of weight adjustments.
[0239] In one specific embodiment of this application, if the accuracy of text analysis results improves by 35% in a certain task, while the accuracy of image and audio analysis does not change significantly, the weight of text may increase, while the weight of images and audio remains unchanged or decreases.
[0240] The principle of similar data correction is as follows:
[0241] Let the target data be The labeled dataset is , where each data All have corresponding labels. This represents text or image labels, etc. It involves calculating the similarity between target data and labeled data, and selecting the data sample most similar to the target data. And correct the dip in the target data.
[0242] Similarity calculation methods include Euclidean distance and cosine similarity.
[0243] Similarity is calculated using Euclidean distance, and the principle is as follows:
[0244] (twenty one); in, This represents the value of the target data in the k-th feature dimension. This represents the value of the labeled data in the k-th feature dimension.
[0245] The principle behind using cosine similarity to calculate similarity is as follows:
[0246] (twenty two); in, Represents the vector dot product. It represents the magnitude of the vector.
[0247] By calculating similarity, the data with the highest similarity is selected for correction, resulting in the corrected annotation. .
[0248] The principle behind the calculation of proportional changes is as follows:
[0249] Let the original label of the target data be Similar data are labeled as The difference between the two for: (twenty three); Then it can be corrected proportionally:
[0250] (twenty four); in, It is a scaling factor that controls the degree of correction; typically, A smaller scaling factor means a smaller correction, while a larger scaling factor means a larger correction.
[0251] In one specific embodiment of this application, the text description annotation of the target data Text description annotation for similar data, such as "temperature is 20 degrees". If the temperature is 25 degrees, then the difference between them is... The corrected annotation is: .
[0252] Assumption If the value is 0.8, then the corrected label is 24 degrees.
[0253] Similarity calculation is performed using cross-domain mapping transformation: In some complex scenarios, data annotation may involve cross-modal data types. For example, the inference of image annotation may need to be corrected based on text descriptions or other modalities. In this case, cross-domain mapping changes can be made by mapping the target data from one modality to another, and then using data from similar domains for inference and correction.
[0254] Cross-domain mapping methods include generative models or transfer learning. For example, image data can be transformed into information related to text descriptions through trained generative models or image-to-text mapping models, and then corrected. The core idea of cross-domain mapping is to use information obtained from one modality to help correct the annotations of another modality.
[0255] The principle of cross-domain mapping is as follows:
[0256] Let the data of the target mode be Its target is labeled as If the generative model mapped to another modality is f, then the corrected annotation obtained by cross-domain mapping is... for: (25); In the cross-domain mapping correction process, the generation model f can be a text generation model, an image description generation model, etc., and the mapping result will be used to update the annotation of the target data.
[0257] The text generation process is as follows:
[0258] Input data: The corrected text annotation information is used as input to form an input sequence. ; Model inference: Based on the input sequence, utilize the Transformer model or other generative models. Perform reasoning to generate new text sequences. ; Output data: Final output text sequence This is the generated multimodal data.
[0259] Text generation formula:
[0260] Let the input text be The model generates new text sequences. ; (26); in, The t-th word in the text sequence.
[0261] Image generation: Image generation can be performed using Generative Adversarial Networks (GANs) or Variational Autoencoders (VAEs). These models can generate new images based on input features or modify existing images based on analysis results.
[0262] The image generation process is as follows: Input data: Features are extracted from text description annotations or image analysis results and passed as input to the generative model;
[0263] Generative models: employing Generative Adversarial Networks (GANs) or Variational Autoencoders (VAEs), or similar generative models. Generate new images based on input features ; Output data: The final output image. .
[0264] Image generation formula:
[0265] Let the input features be The input features can be feature vectors described in text or feature vectors of other modalities, and the generated image is... : (27); in, This represents the image generation model.
[0266] Audio generation: Audio generation methods include sequence-to-sequence (Seq2Seq) models or deep learning models such as WaveNet. These models can generate new audio data based on audio analysis results or text descriptions. The audio data can be speech, ambient sounds, etc.
[0267] The audio generation process is as follows: Input data: Pass the audio analysis results or text descriptions as input features to the audio generation model;
[0268] Audio generation model: using models such as WaveNet Generate new audio data based on the input. ; Output data: The final output is the generated audio data. .
[0269] The audio generation formula is as follows:
[0270] Let the input features be The generated audio is : (28); The formula for historical data feedback is as follows:
[0271] Let the historical data collected in the system be... Each of these data points Including annotation information Analysis results and system performance feedback By calculating the error of each data point, the system can optimize the model based on feedback from historical data.
[0272] ; Errors include labeling errors and prediction errors.
[0273] Based on these errors, the system adjusts the model parameters and annotation strategy to make subsequent data processing more accurate and efficient.
[0274] The principle of model iteration is as follows: Each iteration of the model involves updating the parameters.
[0275] Let the initial model parameters be After each iteration, the model parameters are updated to The update process is typically based on gradient descent or other optimization algorithms: (29); in, For learning rate, Indicates the current parameter gradient, This represents the loss function of the model.
[0276] In a multimodal data generation system, the loss function involves not only the prediction error of each modality, but also the weighted fusion error of the results of each modality. By continuously optimizing the loss function, the model can continuously improve its accuracy.
[0277] The specific strategies for feedback optimization include dynamic weight adjustment and optimization of the data annotation process.
[0278] Dynamic weight adjustment: The weights of different modalities are dynamically adjusted based on the analysis results of historical data feedback.
[0279] Let the mode The current weights are respectively The weighting ratio is adjusted based on the feedback error: , (30; (31); in, and These are the updated weight values calculated based on historical data. This is the learning rate.
[0280] Data annotation workflow optimization: Adjustments will be made to annotation workflows with low standardization rates. For example, for image annotation, more image preprocessing steps may be added, or more advanced image segmentation algorithms may be used to improve annotation accuracy.
[0281] Example 4: Another embodiment of this application relates to a multi-dimensional data multimodal data generation method system. The implementation details of this embodiment's multi-dimensional data multimodal data generation method device are described below. The following implementation details are provided for ease of understanding and are not essential for implementing this solution. This embodiment's multi-dimensional data multimodal data generation method system, as... Figure 1 As shown, it includes a data acquisition module, a data annotation module, a data analysis module, a weight adjustment module, a generation module, and an optimization module.
[0282] The data acquisition module is used to collect data from various modalities, including sensors.
[0283] The data annotation module is used to annotate data of different modalities.
[0284] The data analysis module is used to analyze the labeled modal data and obtain the analysis results.
[0285] The weight adjustment module is used to adjust the weights of data from different modal analyses.
[0286] The generation module is used to generate new data based on the weights and analysis results of different modal data.
[0287] The optimization module is used to optimize the weights and update the system based on historical data feedback.
[0288] It is worth mentioning that all modules involved in this embodiment are logical modules. In practical applications, a logical unit can be a physical unit, a part of a physical unit, or a combination of multiple physical units. Furthermore, to highlight the innovative aspects of this application, this embodiment does not introduce units that are not closely related to solving the technical problems proposed in this application; however, this does not mean that other units are absent in this embodiment.
[0289] Example 5: Another embodiment of this application relates to an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform a multimodal data generation method for multidimensional data in the above embodiments.
[0290] The memory and processor are connected via a bus, which can include any number of interconnecting buses and bridges, connecting various circuits of one or more processors and memories. The bus can also connect various other circuits, such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and will not be described further herein. The bus interface provides an interface between the bus and the transceiver. The transceiver can be a single element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium. Data processed by the processor is transmitted over the wireless medium via an antenna, which further receives data and transmits it to the processor.
[0291] The processor manages the bus and general processing, and also provides various functions, including timing, peripheral interfaces, voltage regulation, power management, and other control functions. Memory is used to store data used by the processor during operation.
[0292] Example 6: Another embodiment of this application relates to a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the method embodiments described above.
[0293] That is, those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program instructing related hardware. This program is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0294] Those skilled in the art will understand that the above embodiments are specific embodiments for implementing this application, and in practical applications, various changes can be made to them in form and detail without departing from the spirit and scope of this application.
Claims
1. A method for generating multimodal data from multidimensional data, characterized in that, include: At least two dimensions of raw data are acquired. After integrating and synchronizing the raw data of each dimension, the data characteristics of different dimensions are labeled using corresponding analysis methods. Semantic information and structured features of each dimension are extracted, and the data of each dimension is converted into structured information to obtain the labeling results of each dimension, forming multimodal initial data. The labeling results are analyzed using corresponding analysis models to obtain the analysis results of each modality. Based on the credibility, historical accuracy and importance of the original data of each modality, appropriate weights are assigned to each modality analysis result. The analysis results of each modality are fused using a weighted fusion method to generate multimodal data. Evaluate the analysis results; Based on historical annotation and analysis results, the analysis method and model are revised and iterated to obtain the iterated analysis method, model, and weights, which are used to generate the final multimodal data.
2. The method for generating multimodal data of multidimensional data according to claim 1, characterized in that, The original data of at least two dimensions includes: text data, image data, audio data, and environmental data. The synchronization of the original data of each dimension includes: time synchronization and spatial synchronization. Time synchronization is used to ensure the correspondence of data of each dimension at the same time. Spatial synchronization is used to spatially align image data, environmental data, and video data, aligning images from different perspectives to the same coordinate system.
3. The method for generating multimodal data of multidimensional data according to claim 2, characterized in that, The annotation of data for each modality is performed using corresponding analysis methods, including: text annotation using natural language description, image annotation using computer vision models, audio annotation using audio processing models, environmental annotation of environmental data, analysis of data for each dimension using computational methods, and third-domain annotation of the analysis results; text annotation includes text description annotation and text association annotation; text description annotation is used to convert collected data into natural language descriptions, and text association annotation is used to generate more related text information for the data based on existing knowledge bases or contextual information; image annotation is used to annotate image data and extract key features and information from the image; audio annotation is used to annotate audio data... The process involves annotation and extraction of audio feature information. The calculation method includes: FFT, a neural network model, and Z-transform. FFT can be used to convert time-domain data to the frequency domain and extract frequency domain features. The neural network model is used to learn the nonlinear mapping relationship between various modal data, encoding multi-source inputs into a joint embedding space. The Z-transform is used to eliminate scale differences caused by different units and sampling frequencies, enabling data from different sources to be compared and fused under a unified statistical distribution. After calculation and analysis, the obtained cross-modal association results are annotated in a third domain. The third-domain annotation results are used as a key component of the state space of the deep reinforcement learning policy network, or as a direct basis for generating personalized training plans, achieving a complete closed loop from multimodal data to intelligent decision-making.
4. The method for generating multimodal data of multidimensional data according to claim 1, characterized in that, The analysis of each annotation result using a corresponding analysis model includes: using a first model to analyze text description annotations and text association annotations to obtain a preliminary understanding of the text content and information extraction; using a second model to analyze image annotations to obtain feature information in the image; and using a third model to analyze audio annotations to extract features from the audio data. The first model includes a Transformer model and / or an LSTM model, the second model includes a convolutional neural network, and the third model includes a convolutional neural network.
5. The method for generating multimodal data of multidimensional data according to claim 1, characterized in that, The modality analysis results are assigned appropriate weights based on the reliability, historical accuracy, and importance of the original data for each modality, including: Data with high credibility should be assigned higher weights; data sources that demonstrate high accuracy in tasks or analyses prior to the current time should be assigned higher weights; data from specific sources should be assigned higher weights when they are more important than other data in a specific task; and data should be assigned random weights based on public opinion preferences.
6. The method for generating multimodal data of multidimensional data according to claim 1, characterized in that, The weighted fusion methods include: weighted average and weighted summation; The method of using weighted fusion to fuse the results of each modal analysis to generate the analysis result includes using linear average or / and square average or / and cubic average or / and exponential average to weight and accumulate each modal analysis result separately, and combining the weights of each modal analysis result to calculate the analysis result according to the weighted fusion formula; The weight allocation is continuously adjusted based on historical data performance and task requirements.
7. The method for generating multimodal data of multidimensional data according to claim 1, characterized in that, The correction of labeling and weights includes: during the labeling and analysis process, for data with labeling errors or excessive time costs, using existing data similar to the target data for correction; Correction methods include: proportional assignment changes and cross-domain mapping changes; The proportional assignment change includes: calculating the difference between the target data and similar data, adjusting it according to a certain proportion, and generating more accurate annotations; Cross-domain mapping transformation includes: mapping target data from one modality to another, and then using data from similar domains for inference and correction; Based on the combined and corrected initial multimodal data and analysis results, a deep learning model was used to generate the final multimodal data; the quality of the generated final multimodal data results was evaluated using model evaluation metrics.
8. The method for generating multimodal data of multidimensional data according to claim 1, characterized in that, The iterative analysis method, analysis model, and weights include: obtaining historical data feedback results based on historical annotation data and historical analysis results, including the accuracy of the annotation data, the correctness of the analysis results, and the evaluation efficiency; and conducting error analysis, improving the model, optimizing the annotation process, and adjusting parameters and weights based on the historical data feedback results to obtain the iterative analysis method, analysis model, and weights.
9. A system for generating multimodal data from multidimensional data, characterized in that, include: The data acquisition module is used to collect data from various dimensions; The data annotation module is used to annotate data of different dimensions using different analysis methods, obtain the annotation results of data of each dimension, and form multimodal initial data; The data analysis module is used to analyze the initial multimodal data using the analysis model corresponding to each mode, and obtain the analysis results for each mode. The weight adjustment module is used to adjust the weights of different modal analysis results; The generation module is used to generate multimodal intermediate data and multimodal final data based on the weights of each modality's data and the analysis results. The optimization module is used to optimize the analysis methods, analysis models, and weights based on historical data feedback, and to update the system.
10. An electronic device, characterized in that, include: At least one processor; as well as, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform a multimodal data generation method for multidimensional data as described in any one of claims 1 to 8.