An artificial intelligence driven multi-modal data analysis method and system
By constructing a cross-modal semantic association model and a hierarchical processing module, the problems of spatiotemporal alignment, feature fusion, and semantic analysis in multimodal data analysis are solved, enabling efficient and accurate analysis and real-time response of multimodal data, and adapting to complex and ever-changing application scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- XIAN DASHENG TECH CO LTD
- Filing Date
- 2026-01-09
- Publication Date
- 2026-04-21
AI Technical Summary
Traditional multimodal data analysis systems face challenges in processing multimodal data, including complex spatiotemporal characteristics, difficulties in feature extraction and fusion, insufficient semantic consistency analysis, and inadequate real-time performance. These issues result in low accuracy and reliability of the analysis, making it difficult to adapt to complex and ever-changing application scenarios.
The AI-driven multimodal data analysis system includes a data acquisition module, a multimodal fusion module, a feature extraction module, and a hierarchical processing module. By constructing a cross-modal semantic association model, using alignment fusion technology to process spatiotemporal characteristics, combining feature embedding technology to fuse multimodal features, and performing standardized processing and hierarchical analysis, it achieves multi-dimensional semantic collaborative evaluation and contextual factor analysis.
It achieves efficient and accurate fusion and analysis of multimodal data, can evaluate the semantic consistency of the target scene in real time, and provides timely response and processing measures, thereby improving the system's analytical accuracy and ability to adapt to complex scenarios.
Smart Images

Figure CN121479577B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, specifically to an artificial intelligence-driven multimodal data analysis method and system. Background Technology
[0002] In today's digital age, with the rapid development of information technology, people are facing increasingly diverse data types, no longer limited to single text or image data, but exhibiting a trend of multimodal data fusion. Multimodal data includes information in various forms such as images, text, and voice, which together describe different aspects of a target scene and contain rich semantic information. However, traditional data analysis systems face many challenges when processing multimodal data.
[0003] Multimodal data presents complex spatiotemporal characteristics. Data from different modalities often exhibits differences and inconsistencies in time and space. For example, the acquisition time of image data and the recording time of audio data may differ; different devices may be located at different points in the target scene, resulting in varying spatial coordinates and semantic relationships in the acquired data. Traditional systems struggle to effectively handle these spatiotemporal differences, making it difficult to establish correlations between multimodal data and accurately reflect the true situation of the target scene.
[0004] Feature extraction and fusion of multimodal data present challenges. Each modality of data has its unique feature representation, such as image texture, text semantics, and speech emotion. Traditional feature extraction methods are often designed for a single modality, failing to fully leverage the complementarity between multimodal data. Moreover, when fusing features from different modalities, semantic gaps and information loss can easily occur due to differences in feature dimensions and representations, resulting in fused features that cannot accurately represent the comprehensive semantics of the multimodal data.
[0005] Traditional systems lack comprehensive and in-depth semantic consistency analysis of target scenarios. When analyzing multimodal data, they typically only perform preliminary comparative analysis without considering the impact of contextual factors. When semantic inconsistencies arise in the target scenario, they cannot further analyze the causes and extent of their impact, making it difficult to provide effective handling measures and solutions, resulting in low analytical accuracy and reliability.
[0006] As industries increasingly demand higher levels of data analytics capabilities, data analysis systems are required to process multimodal data accurately and in real time, and to respond and process data promptly based on the analysis results. Traditional systems struggle to meet these demands in terms of real-time performance and processing efficiency, and are ill-suited to complex and ever-changing application scenarios. Summary of the Invention
[0007] The purpose of this invention is to provide an artificial intelligence-driven multimodal data analysis method and system to solve the problems mentioned in the background art.
[0008] To achieve the above objectives, the present invention provides the following technical solution: an artificial intelligence-driven multimodal data analysis system, the system comprising:
[0009] The system includes a data acquisition module, a multimodal fusion module, a feature extraction module, a comprehensive analysis module, and a hierarchical processing module.
[0010] The data acquisition module is used to set up acquisition nodes in the interactive interface between the target scene and the multimodal device, and deploy acquisition terminals to collect multimodal data of the target scene and operating parameters of the multimodal device in real time, and convert the acquired data into a format.
[0011] The multimodal fusion module is used to construct a cross-modal semantic association model, and uses alignment fusion technology to process the spatiotemporal characteristics of multimodality. At the same time, it uses feature embedding technology to fuse the representation features of multimodality, and performs association simulation of target scene and multimodal device based on real-time collected multimodal data and device operating parameters.
[0012] The feature extraction module is used to construct image feature extraction algorithms, text semantic extraction algorithms, and voice emotion extraction algorithms, and transmits the real-time collected multimodal data and device operating parameters to the constructed extraction algorithms to calculate and obtain image features, text semantics, and voice emotion.
[0013] The comprehensive analysis module is used to standardize the acquired image features, text semantics and voice emotion, perform correlation calculation to obtain a first analysis value and set a preset first benchmark threshold, and perform preliminary comparative analysis of the semantic consistency of the target scene.
[0014] The first analysis value is generated by weighted summation of standardized image features, text semantics, and speech sentiment. Its weight coefficients are dynamically adjusted according to the modal importance of multimodal data. This analysis value is used to quantitatively evaluate the degree of synergy of each modal information in the target scene at the semantic level. By comparing it with a preset first benchmark threshold, it can quickly identify the modal combinations and specific locations of semantic consistency deviations.
[0015] The analysis of the target scenario specifically includes: the degree of semantic synergy between multimodal data, i.e., the correlation and complementarity of different modal information in the spatiotemporal dimensions; the impact of device interaction status on scenario semantics, such as the semantic representation differences caused by changes in parameters like sensor acquisition frequency and data transmission latency; and the scenario's dynamic response capability, i.e., the efficiency of maintaining semantic consistency in real-time updates of multimodal data. Through the above multi-dimensional analysis, the stability and reliability of the target scenario in complex interactive environments can be comprehensively evaluated.
[0016] The hierarchical processing module is used to further calculate and obtain a second analysis value by combining contextual factors when the semantic inconsistency of the target scene is detected, and to perform a second comparison analysis with the second analysis value by setting a second benchmark threshold, so as to further analyze the performance of the target scene under different multimodal data and device interaction states.
[0017] When the semantics of the target scene are found to be consistent, the multimodal data will continue to be collected and processed.
[0018] Preferably, the data acquisition module includes a multimodal data acquisition unit, a modal feature acquisition unit, and a data preprocessing unit;
[0019] The multimodal data acquisition unit includes an image acquisition unit, a text acquisition unit, and a voice acquisition unit. It is used to acquire and transmit multimodal data of the target scene in real time by deploying a group of perception sensors on the interactive interface of the target scene, and transmit the data to the data preprocessing unit wirelessly. The group of perception sensors includes an image sensor group, a text input device group, and a voice microphone group. The multimodal data includes image frame data, text string data, and voice waveform data.
[0020] The image acquisition unit is used to acquire image frame data of the target scene in real time based on the image sensor group. The image sensor group includes a camera, an optical lens and an image encoder, which respectively acquire the resolution, color channels and frame rate of the image data.
[0021] The text acquisition unit is used to acquire text string data in real time based on the text input device group. The text input device group includes a keyboard input module, a touch input module, and an OCR recognition module. The text string data includes character sequences, input timestamps, and semantic context.
[0022] The voice acquisition unit is used to acquire voice waveform data in real time based on the voice microphone group, which includes a microphone array, an audio amplifier, and an analog-to-digital converter. The voice waveform data includes sampling rate, bit depth, and audio duration.
[0023] Preferably, the data preprocessing unit is used to filter out noise and outliers from the collected multimodal data and device operating parameters, unify the data formats of different modalities, and perform feature filtering on the collected multimodal data and device operating parameters by timestamp alignment to obtain keyframe images, effective text segments, and effective speech intervals.
[0024] Preferably, the multimodal fusion module includes a spatiotemporal alignment unit, a feature fusion unit, and a visualization integration unit;
[0025] The spatiotemporal alignment unit, feature fusion unit, and visualization integration unit work together to construct a cross-modal semantic association model;
[0026] The spatiotemporal alignment unit includes a coordinate mapping unit and a timing synchronization unit;
[0027] The coordinate mapping unit extracts the spatial coordinates of the target scene and the device location information from the positioning system, and uses an alignment algorithm to establish a multimodal spatiotemporal mapping model to map the spatial coordinates of the scene, the device pose, and the interaction area. At the same time, it adds typical multimodal features to the target scene, including image semantic regions, text key nodes, and speech emotion intervals. After the initial alignment is completed, a calibration tool is used to define the timestamp offset, spatial coordinate deviation, and semantic association weight of the constructed spatiotemporal mapping model. At the same time, the acquisition start point, processing end point, and interaction range of the multimodal device are set, and static alignment, dynamic alignment, and continuous alignment are performed to align the spatiotemporal response of the multimodal data of the target scene.
[0028] The timing synchronization unit is used to input multimodal timing parameters, including image acquisition time, text input time, and voice recording time. After input, timing analysis is performed to synchronize the timing features, semantic features, and emotional features of the target scene at different interaction stages.
[0029] The feature fusion unit is used to establish a multimodal embedding model, including image feature dimension, text word vector dimension and speech Mel coefficient dimension, and then apply a feature fusion network to simulate multimodal semantic association. The feature embedding technology is used to analyze the multimodal fusion results and evaluate the dimensional distribution, semantic consistency and sentiment matching of the fused features.
[0030] The visualization integration unit is used to import the target scene model into the multimodal fusion model for integration to obtain a digital twin model. Then, it transmits the real-time collected multimodal data of the target scene and the operating parameters of the multimodal devices to time series analysis and feature embedding for dynamic fusion. The dynamic fusion result is then imported into the digital twin model to update the status of the target scene and multimodal devices in real time. The fused data is displayed through a visualization interface, providing a user interaction interface.
[0031] Preferably, the feature extraction module includes an image feature extraction unit, a text semantic extraction unit, and a speech emotion extraction unit;
[0032] The image feature extraction unit is used to construct an image feature extraction algorithm, calculate and obtain image features of key frames of the target scene based on preprocessed multimodal data, and extract the texture distribution and target detection information of the image;
[0033] The text semantic extraction unit is used to construct a text semantic extraction algorithm, which calculates and obtains text semantics based on preprocessed multimodal data, and extracts the semantic intent and key information of the interaction process in the target scene.
[0034] The voice emotion extraction unit is used to construct a voice emotion extraction algorithm, which calculates and obtains voice emotions based on preprocessed device operating parameters, and extracts the emotional tendency and intensity of the interaction process in the target scene.
[0035] Preferably, the image feature extraction unit is used to calculate and obtain image features by analyzing the multimodal fused image frame data and combining it with the preprocessed text semantic data, and to extract the correspondence between the image semantic region and the text key node.
[0036] Preferably, the text semantic extraction unit is used to calculate and obtain text semantics by comparing text string data in adjacent interaction stages and combining preprocessed voice emotion data, and to extract the semantic change rate of the target scene interaction process.
[0037] Preferably, the comprehensive analysis module includes a multimodal correlation unit and a preliminary discrimination unit;
[0038] The multimodal association unit is used to standardize the acquired image features, text semantics and voice sentiment, then perform association calculation to obtain the first analysis value, and perform comprehensive analysis of the semantic consistency of the target scene.
[0039] The preliminary discrimination unit is used to set a preset first benchmark threshold based on industry standards and historical data for multimodal data analysis, and to perform a preliminary comparative analysis with the obtained first analysis value to analyze the semantic consistency of the target scene. The specific analysis scheme is as follows: when the first analysis value is greater than the first benchmark threshold, it indicates that the target scene is semantically consistent under the current interaction conditions and real-time monitoring continues; when the first analysis value is less than or equal to the first benchmark threshold, it indicates that the target scene is semantically inconsistent under the current interaction conditions and processing measures and further verification operations are required.
[0040] Preferably, the hierarchical processing module includes a processing index calculation unit and a level determination unit;
[0041] The processing index calculation unit is used to combine the obtained first analysis value to further analyze the semantic consistency of the target scene under different multimodal data and device interaction states, and perform correlation calculation to obtain the second analysis value;
[0042] The level determination unit is used to perform a second comparative analysis between a preset second baseline threshold and the acquired second analysis value, further analyzing the consistency of the target scene's semantics under the influence of multiple contextual factors, and generating a corresponding processing level. The specific analysis scheme is as follows: When the second analysis value is greater than the second baseline threshold, it indicates that the target scene is still semantically consistent under the interaction conditions after considering all contextual factors. At this time, a level 3 processing message is generated to prompt monitoring personnel to continuously observe the target scene. When the second analysis value is equal to the second baseline threshold, it indicates that the target scene is semantically inconsistent under the interaction conditions after considering all contextual factors, and there is a potential misunderstanding. At this time, a level 2 processing message is generated to prompt relevant personnel to immediately conduct a detailed verification of the target scene. When the second analysis value is less than the second baseline threshold, it indicates that the target scene is significantly semantically inconsistent under the interaction conditions after considering all contextual factors. At this time, a level 1 processing message is generated to automatically trigger the processing mechanism and notify relevant modules to start the semantic correction plan.
[0043] Preferably, the present invention also includes an artificial intelligence-driven multimodal data analysis method, applied to the aforementioned artificial intelligence-driven multimodal data analysis system, the method comprising the following steps:
[0044] Step 1: Set up acquisition nodes in the interaction interface between the target scene and the multimodal device, and deploy acquisition terminals to collect multimodal data of the target scene and operating parameters of the multimodal device in real time, and convert the format of the collected data;
[0045] Step 2: Construct a cross-modal semantic association model, use alignment and fusion technology to process the spatiotemporal characteristics of multimodality, use feature embedding technology to fuse the representation features of multimodality, and perform association simulation of the target scene and multimodal devices based on real-time collected multimodal data and device operating parameters;
[0046] Step 3: Construct image feature extraction algorithm, text semantic extraction algorithm, and speech emotion extraction algorithm, and transmit the real-time acquired multimodal data and device operating parameters to the constructed extraction algorithms to calculate and obtain image features, text semantics, and speech emotion;
[0047] Step 4: After standardizing the acquired image features, text semantics, and speech sentiment, perform association calculations to obtain the first analysis value and make a preliminary comparison with the preset first benchmark threshold to analyze the semantic consistency of the target scene;
[0048] Step 5: If the analysis reveals that the semantics of the target scene are inconsistent, then the second analysis value is obtained by further combining the context factors. The second baseline threshold is preset and the second analysis value is compared and analyzed again to further analyze the performance of the target scene under different multimodal data and device interaction states.
[0049] Compared with the prior art, the beneficial effects of the present invention are:
[0050] In the data acquisition phase, the data acquisition module sets up acquisition nodes and deploys acquisition terminals in the target scene and multimodal device interaction interface, enabling real-time acquisition of multimodal data and device operating parameters. It can also perform format conversion on the acquired data. Furthermore, the data acquisition unit is subdivided into image, text, and voice acquisition units, which, together with corresponding sensor groups, can accurately acquire various types of data. The preprocessing unit can filter out noise, unify the format, and align key information through timestamps, ensuring the comprehensiveness, accuracy, and standardization of data acquisition, laying a solid foundation for subsequent analysis.
[0051] The multimodal fusion module constructs a cross-modal semantic association model, uses alignment fusion technology to process spatiotemporal characteristics, and fuses representational features using feature embedding technology. It can also simulate the association between target scenes and devices. Among them, the spatiotemporal alignment unit establishes and calibrates a spatiotemporal mapping model through coordinate mapping and time synchronization to achieve accurate alignment of the spatiotemporal response of multimodal data; the feature fusion unit establishes an embedding model and applies a fusion network to simulate semantic association and evaluate fusion features; the visualization integration unit generates a digital twin model and dynamically updates its status and displays it visually, making multimodal data fusion more efficient, accurate, and easier for users to understand and interact with.
[0052] The feature extraction module constructs various extraction algorithms that can obtain image features, text semantics, and speech emotion from multimodal data. Furthermore, each extraction unit can combine other modal data for more in-depth feature extraction. For example, the image feature extraction unit combines text semantic data to extract correspondences, and the text semantic extraction unit combines speech emotion data to extract semantic change rates, fully exploring the correlation information between multimodal data and improving the depth and accuracy of feature extraction.
[0053] The comprehensive analysis module standardizes and correlates the acquired features to obtain a first analysis value, which is then compared with a preset threshold. This preliminary analysis of semantic consistency allows for a quick determination of the semantic state of the target scene under the current interaction conditions, providing a basis for subsequent processing.
[0054] When semantic inconsistencies occur, the hierarchical processing module calculates a second analysis value based on contextual factors and performs a secondary comparison to generate different processing levels. It provides corresponding processing measures for different degrees of semantic inconsistency, achieving in-depth analysis and precise processing of the target scene's semantics. This improves the system's analysis accuracy and reliability, and enables it to better adapt to complex and ever-changing application scenarios.
[0055] This system and method effectively solves the problems of spatiotemporal alignment difficulties, insufficient feature fusion, and incomplete semantic analysis in traditional multimodal data analysis systems through the collaborative work of various modules. It improves the efficiency and accuracy of data analysis, can analyze the semantic consistency of the target scene in real time and comprehensively, and make timely and effective responses and processing based on the analysis results. It has broad application prospects and high practical value. Attached Figure Description
[0056] Figure 1 This is a schematic diagram illustrating the working principle of the AI-driven multimodal data analysis system described in this invention.
[0057] Figure 2 Design diagram of the data acquisition module;
[0058] Figure 3 A flowchart illustrating the work of the data preprocessing unit;
[0059] Figure 4 Design diagram of the multimodal fusion module;
[0060] Figure 5 This is a design diagram of the feature extraction module. Detailed Implementation
[0061] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0062] Please see Figures 1-5 This invention relates to an artificial intelligence-driven multimodal data analysis system. The system includes a data acquisition module, a multimodal fusion module, a feature extraction module, a comprehensive analysis module, and a hierarchical processing module. Specific implementation details are as follows:
[0063] Data acquisition module: Set up acquisition nodes and deploy acquisition terminals in the interactive interface of the target scene and multimodal devices to collect multimodal data of the target scene and operating parameters of the multimodal devices in real time, and convert the acquired data into a format.
[0064] Multimodal fusion module: Constructs a cross-modal semantic association model, uses alignment fusion technology to process the spatiotemporal characteristics of multimodality, uses feature embedding technology to fuse the representation features of multimodality, and performs association simulation of target scene and multimodal device based on real-time collected multimodal data and device operating parameters.
[0065] Feature extraction module: Constructs image feature extraction algorithm, text semantic extraction algorithm and voice emotion extraction algorithm, and transmits the real-time collected multimodal data and device operating parameters to the constructed extraction algorithm to calculate and obtain image features, text semantics and voice emotion.
[0066] The comprehensive analysis module standardizes the acquired image features, text semantics, and voice sentiment, performs correlation calculations to obtain a first analysis value, and compares it with a preset first benchmark threshold to analyze the semantic consistency of the target scene.
[0067] Hierarchical processing module: If the analysis finds that the semantics of the target scene are inconsistent, it will further combine contextual factors to calculate and obtain a second analysis value, and perform a second comparison analysis with the second analysis value based on a preset second benchmark threshold, so as to further analyze the performance of the target scene under different multimodal data and device interaction states.
[0068] Example 1:
[0069] The system's data acquisition module consists of a multimodal data acquisition unit, a modal feature acquisition unit, and a data preprocessing unit. The multimodal data acquisition unit, as the core of the data acquisition process, comprises three sub-units: an image acquisition unit, a text acquisition unit, and a voice acquisition unit. These sub-units work together to achieve comprehensive acquisition of multimodal data from the target scene.
[0070] In terms of image acquisition, the image acquisition unit utilizes an image sensor array to perform the task of acquiring real-time image frame data of the target scene. The image sensor array is a collection of various devices, specifically including a camera, an optical lens, and an image encoder. Each of these devices plays a crucial role: the camera captures images of the target scene; the optical lens processes light to ensure image clarity and quality; and the image encoder encodes the acquired image information for subsequent transmission and storage. During acquisition, the image sensor array can acquire key parameters such as image resolution, color channels, and frame rate. Resolution determines the image's clarity, color channels affect color reproduction, and frame rate relates to image smoothness. By acquiring these parameters, high-quality image frame data can be obtained.
[0071] The text acquisition unit relies on a text input module, which includes a keyboard input module, a touch input module, and an OCR recognition module. The keyboard input module allows users to input text information via a keyboard, while the touch input module is suitable for touchscreen devices, allowing users to directly input text on the screen. The OCR recognition module recognizes text in images and converts it into editable text strings. These text strings contain character sequences, input timestamps, and semantic context. The character sequences represent the specific content of the text, the input timestamps record the time of text input, and the semantic context provides important background information for subsequent semantic analysis.
[0072] The voice acquisition unit uses a microphone array to acquire voice waveform data in real time. The microphone array consists of a microphone array, an audio amplifier, and an analog-to-digital converter (ADC). The microphone array can acquire sound signals from different directions, improving the accuracy and comprehensiveness of voice acquisition; the audio amplifier amplifies the acquired sound signals to ensure signal strength; and the ADC converts the analog sound signals into digital signals for computer processing. The voice waveform data includes parameters such as sampling rate, bit depth, and audio duration. The sampling rate determines the sampling accuracy of the sound signal, the bit depth affects the sound quality, and the audio duration records the duration of the speech.
[0073] The multimodal data acquisition unit achieves real-time acquisition and transmission of multimodal data from the target scene by deploying a sensor array on the interactive interface of the target scene. The sensor array is a combination of an image sensor array, a text input array, and a voice microphone array, which work together to simultaneously acquire data from multiple modalities, including images, text, and voice. The acquired data is transmitted wirelessly to the data preprocessing unit. This wireless transmission method offers advantages such as high flexibility and convenient deployment, adapting to different target scenes and device layouts. The multimodal feature acquisition unit is used to specifically acquire the operating parameters of the multimodal device. These operating parameters include hardware status parameters (such as sensor operating voltage, current, temperature, processor utilization, memory usage, etc.) and data transmission parameters (such as wireless signal strength, data transmission rate, packet loss rate, etc.). This unit establishes a connection with the control interface of the multimodal device to acquire various dynamic parameters of the device in real time during operation. Hardware status parameters reflect the physical operating condition of the device, ensuring it is within its normal operating range and preventing data distortion due to device malfunction. Data transmission parameters are used to evaluate the transmission quality of data from the acquisition end to the preprocessing unit, providing a basis for subsequent analysis of data integrity and timeliness. The modal feature acquisition unit associates and tags the acquired device operating parameters with the image, text, and voice data acquired by the multimodal data acquisition unit. By adding the same timestamp and device identifier, it achieves a precise correspondence between multimodal data and device operating status, laying the foundation for subsequent analysis of the impact of different device states on data quality.
[0074] The data preprocessing unit is a crucial component of the data acquisition module. It performs a series of processing operations on the acquired multimodal data and equipment operating parameters. The data preprocessing unit filters out noise and outliers from the data. Noise may be caused by sensor malfunctions, environmental interference, etc., while outliers may be caused by errors or abnormal conditions during the data acquisition process. These noises and outliers can affect subsequent data analysis and processing, and therefore need to be filtered out.
[0075] The data preprocessing unit unifies the data formats of different modalities. Since data of different modalities such as images, text, and speech have different formats and characteristics, they need to be converted into a unified format to facilitate subsequent fusion and analysis.
[0076] The data preprocessing unit uses timestamp alignment to perform feature filtering on the collected multimodal data and device operating parameters, obtaining keyframe images, valid text fragments, and valid speech intervals. Timestamp alignment ensures temporal consistency across different modalities, facilitating correlation analysis. Feature filtering extracts key information from large amounts of data, improving the efficiency and accuracy of data processing.
[0077] For keyframe images, the data preprocessing unit first samples the video stream uniformly at a fixed frame rate, for example, extracting one frame per second. Then, by calculating the visual differences between adjacent frames, such as histogram differences and structural similarity index (SSIM), frames with visual differences exceeding a preset threshold are identified as keyframes. These keyframes represent significant changes in the video content in terms of scene, action, or object presentation, covering the core information in the video. For valid text fragments, the data preprocessing unit uses natural language processing technology to analyze the grammatical structure, semantic coherence, and relevance to the overall context of the multimodal data. Text paragraphs with complete semantics and contributing to the expression of core content are defined as valid text fragments, such as removing redundant embellishments, meaningless repetitions, and parts with low relevance to the main content. Regarding the effective speech interval, the data preprocessing unit uses Voice Endpoint Detection (VAD) technology to identify the start and end points of speech from a continuous speech stream based on the time-domain and frequency-domain characteristics of the speech signal, such as energy fluctuations, zero-crossing rate, fundamental frequency, and spectral composition. This determines the effective speech interval and eliminates silence, noise, and irrelevant information from the speech. During the acquisition of keyframe images, effective text segments, and the effective speech interval, the data preprocessing unit utilizes timestamp alignment technology to precisely match multimodal data and device operating parameters in the time dimension. This ensures the correlation and synchronization between different modal data, providing a reliable foundation for subsequent multimodal data fusion and analysis.
[0078] In practical applications, the various units of the data acquisition module collaborate closely to form a complete data acquisition and preprocessing workflow. For example, in an intelligent customer service system, the image acquisition unit can capture customers' facial expressions and movements through a camera, the text acquisition unit can capture customers' text inquiries through a keyboard and touch input module, and the voice acquisition unit can capture customers' voice inquiries through a microphone array. This data is transmitted to the data preprocessing unit for processing, filtering out noise and outliers, standardizing the data format, and aligning keyframe images, valid text segments, and valid voice intervals using timestamps. The processed data is then transmitted to the subsequent multimodal fusion module for further analysis and processing, thereby providing accurate and comprehensive data support for the intelligent customer service system, enabling accurate understanding and response to customer needs and emotions.
[0079] Through the aforementioned workflow of the data acquisition module, efficient acquisition and preprocessing of multimodal data from the target scene can be achieved, laying a solid data foundation for subsequent work of the entire multimodal data analysis system. The module's design fully considers the characteristics and acquisition requirements of different modalities, and through reasonable hardware deployment and data processing algorithms, ensures the accuracy, integrity, and consistency of the data, providing strong support for the overall system performance.
[0080] Example 2:
[0081] The multimodal fusion module comprises a spatiotemporal alignment unit, a feature fusion unit, and a visualization integration unit. Each unit achieves deep fusion of multimodal data through hierarchical processing. The cross-modal semantic association model is the core framework for the multimodal fusion module to achieve cross-modal semantic association. Its specific construction and implementation are accomplished through the spatiotemporal alignment unit, feature fusion unit, and visualization integration unit. Specifically, the spatiotemporal alignment unit establishes the spatiotemporal association foundation of multimodal data through coordinate mapping and temporal synchronization, providing spatiotemporal constraints for the model; the feature fusion unit simulates multimodal semantic association through embedded models and fusion networks, serving as the core link for feature fusion in the model; and the visualization integration unit visualizes the association results through a digital twin model, completing the model's simulation of association with the target scene and device. The spatiotemporal alignment unit, as a fundamental component, consists of a coordinate mapping unit and a temporal synchronization unit, which respectively establish the association framework for multimodal data from the spatial and temporal dimensions.
[0082] The coordinate mapping unit begins by extracting the spatial coordinates of the target scene and the device location information from the positioning system. This positioning system can be an outdoor system such as GPS or BeiDou, or an indoor system such as UWB or RFID. The extracted information includes the scene's three-dimensional coordinates and the device's attitude parameters (such as pitch, yaw, and roll angles). Based on this information, the coordinate mapping unit uses alignment algorithms (such as the Iterative Closest Point (ICP) algorithm and the multimodal transformation matrix algorithm) to establish a spatiotemporal mapping model. This model can accurately map the scene's spatial coordinates, device pose, and interaction area. Taking a smart factory scenario as an example, the model digitally models the spatial position of the production equipment, the range of motion of the robotic arm, and the interaction area of the control panel.
[0083] After the initial modeling is completed, the coordinate mapping unit adds multimodal typical features to the target scene. For example, in the image modality, it marks image semantic regions such as material recognition areas and equipment fault warning areas; in the text modality, it sets key text nodes such as operation instruction keywords and equipment parameter thresholds; and in the speech modality, it divides speech emotion regions such as operator instruction intervals and equipment abnormal alarm tone intervals. The addition of these features provides anchors for subsequent semantic association.
[0084] To ensure the accuracy of spatiotemporal mapping, the system uses calibration tools to fine-tune the model. These calibration tools are specifically a software toolset for fine-tuning the spatiotemporal mapping model, including a parameter configuration interface, an error analysis module, and an alignment verification component. The parameter configuration interface allows users to input timestamp offsets (e.g., the sampling time difference between the camera and microphone ±50ms), spatial coordinate deviation values (e.g., the origin offset of different sensors (x:0.2m, y:0.1m, z:0.05m)), and semantic association weight matrices (e.g., image-to-text association weight 0.6, text-to-speech association weight 0.4). The error analysis module evaluates the adjustment effect by calculating the deviation rate of multimodal data before and after alignment (e.g., spatial coordinate deviation rate ≤3%). The alignment verification component generates an alignment report, including static / dynamic alignment error curves and the cumulative error value of continuous alignment (e.g., 24-hour cumulative error ≤0.5m). During calibration, it is necessary to define the timestamp offset of multimodal data (such as the sampling time difference between the camera and the microphone), spatial coordinate deviation (such as the offset of the coordinate origin of different sensors), and semantic association weight (such as the importance of the association between device status data and image data). At the same time, it is also necessary to set the acquisition start point of multimodal devices (such as the device initialization state when the production line starts), processing end point (such as the data cutoff condition when the task is completed), and interaction range (such as the working radius of the robot arm). Through three methods, namely static alignment (one-time calibration for fixed scenarios), dynamic alignment (real-time calibration to adapt to device movement), and continuous alignment (ensuring cumulative error correction over long-term operation), the spatiotemporal response alignment of multimodal data in the target scenario is achieved.
[0085] Static alignment is performed during system initialization. For fixed-position devices and spatial structures in the target scene, such as fixed production lines in a factory or stationary office equipment in an office, a single calibration operation establishes the spatial coordinate reference and temporal zero point for multimodal data. For example, it maps the field of view center of an image sensor, the physical location of a text input device, and the pickup range of a voice acquisition device to a unified three-dimensional coordinate system, and sets the initial timestamp as the synchronization starting point for all devices, ensuring the basic consistency of multimodal data in spatial layout and temporal starting point in a fixed scene. Dynamic alignment is used to handle situations where devices move or the scene changes, such as mobile robots operating in a workshop or wearable devices changing position during human activity. The system dynamically updates the spatial coordinate mapping relationship of multimodal data by tracking the device's pose change data in real time (such as motion parameters fed back by gyroscopes and accelerometers). Simultaneously, it adjusts the timestamp offset in real time according to the device's moving speed and data acquisition frequency, ensuring the spatiotemporal synchronization of different modal data during movement. For example, when a moving camera changes its shooting angle, dynamic alignment immediately corrects the spatial correspondence between the semantic regions of the image and key nodes in the text. Continuous alignment periodically corrects the accumulated errors generated during long-term system operation. By analyzing the spatiotemporal deviation trends in historical data, such as the accumulation of timestamp deviations caused by device clock drift and spatial positioning errors caused by mechanical wear, the calibration tool is invoked to recalculate the parameters of the spatiotemporal mapping model at preset time intervals (e.g., once per hour) or when the deviation exceeds a threshold. The results of static and dynamic alignment are then fine-tuned. For example, after 8 hours of continuous operation, the coordinate offset of the image sensor is corrected by comparing the position information of reference markers at known locations in each modal data, ensuring that the spatiotemporal consistency of multimodal data remains within an acceptable range during long-term operation.
[0086] The timing synchronization unit primarily addresses the temporal consistency issue of multimodal data. It first inputs multimodal temporal parameters, including image acquisition time (e.g., sampling times at 30 frames per second from a camera), text input time (e.g., timestamp sequences of keystrokes), and voice recording time (e.g., timestamps for audio sampling). These time parameters are then imported into the timing analysis module. This module uses techniques such as Dynamic Time Warping (DTW) and sliding time window matching to synchronize the temporal characteristics (e.g., process time nodes in a production workflow), semantic characteristics (e.g., the temporal sequence logic of operation instructions), and emotional characteristics (e.g., the temporal variation curve of the operator's voice emotion) of the target scene at different interaction stages. For example, in a medical diagnostic scenario, the timing synchronization unit aligns the patient's image acquisition time, medical record text entry time, and consultation voice recording time to ensure that doctors obtain consistent multimodal information in terms of time dimension during analysis.
[0087] The feature fusion unit is the core component for achieving multimodal semantic association. It integrates feature spaces from different modalities by establishing a multimodal embedding model. This model includes image feature dimensions (such as visual feature vectors extracted via ResNet), text word vector dimensions (such as semantic vectors generated via BERT), and speech Mel coefficient dimensions (such as acoustic features extracted via Mel frequency cepstral coefficients). These dimensions are coupled through feature fusion networks (such as attention mechanism networks or cross-modal transformation networks). The network simulates the semantic associations between multimodal data. For example, in a smart home scenario, it associates the user's gesture image features captured by a camera with the Mel coefficient features of voice commands to identify the user's control intent.
[0088] Feature embedding technology plays a crucial role in the fusion process. It can analyze the fusion results of multimodal features and evaluate the dimensionality distribution of the fused features (such as the proportion of each modality feature in the fusion vector), semantic consistency (such as the degree of matching between the image and the text description), and sentiment matching (such as the fit between the voice sentiment and the image scene). Through this evaluation, the system can dynamically adjust the fusion parameters and optimize the integration effect of multimodal features.
[0089] The feature embedding technique is used to map the representational features of multimodal data to a unified semantic space, achieving comparability and correlation between features of different modalities. For image data, after extracting high-dimensional visual features through a convolutional neural network, the embedding layer is used to convert them into fixed-dimensional vectors, preserving key semantic information such as texture and shape of the image. For text data, the text string is first converted into word embedding vectors through a word vector model, and then processed by a recurrent neural network or Transformer model to obtain sentence-level semantic embedding vectors, covering the semantic intent and contextual association of the text. For speech data, after extracting acoustic features such as Mel-frequency cepstral coefficients, deep neural networks are used to map them into emotion embedding vectors, reflecting the emotional tendency and intensity in the speech. During the mapping process, this technique optimizes the embedding vectors through cross-modal loss functions (such as contrastive loss and triplet loss), making semantically related features from different modalities closer together in a unified space and semantically unrelated features farther apart. At the same time, it dynamically adjusts the distribution weights of the embedding vectors by combining real-time collected multimodal data and device operating parameters, ensuring that the fused features can retain the unique information of each modality and reflect cross-modal semantic associations, thus providing an accurate feature foundation for subsequent association simulation and semantic consistency analysis.
[0090] The visualization integration unit transforms abstract fusion data into an intuitive digital twin model. First, it imports the target scene model (such as Building Information Modeling (BIM) or 3D models of industrial equipment) into the multimodal fusion model for integration, forming the basic framework of the digital twin. Then, it transmits real-time collected multimodal data from the target scene (such as image frames, text commands, and voice signals) and multimodal equipment operating parameters (such as real-time sensor data and equipment operating status indicators) to the time-series analysis and feature embedding module for dynamic fusion. The fused result is then imported into the digital twin model, enabling real-time updates of the target scene and the status of the multimodal equipment.
[0091] At the visualization level, the system presents the fused data in an intuitive way through a 3D visualization interface. For example, in smart city management, the digital twin model will display the fusion results of traffic flow image data, road condition text reports and traffic command voice instructions in real time. Managers can interact with the system in real time through interactive interfaces (such as touch screens and voice control) to obtain detailed data analysis reports or issue control instructions.
[0092] The entire multimodal fusion module exhibits a high degree of synergy and dynamic adaptability. Taking a smart retail scenario as an example, the coordinate mapping unit spatially models the store's shelf layout and camera positions; the temporal synchronization unit aligns the time of customer movement trajectory images, product label text scanning time, and consultation voice time; the feature fusion unit uses an embedding model to correlate and analyze customer facial expression features with product review text and consultation voice emotion; and finally, the visualization integration unit displays customer shopping behavior preferences in real time within a digital twin model, providing data support for merchants' marketing strategies. This fusion approach not only achieves spatiotemporal alignment of multimodal data but also uncovers deeper information behind the data through semantic association and feature integration, providing a comprehensive and accurate basis for subsequent analysis and decision-making.
[0093] Example 3:
[0094] The feature extraction module, as a key component of the multimodal data analysis system, consists of an image feature extraction unit, a text semantic extraction unit, and a speech emotion extraction unit. Each unit uses a customized algorithm architecture to achieve feature parsing and semantic association of different modal data.
[0095] The core function of the image feature extraction unit is to construct targeted image feature extraction algorithms. These algorithms calculate image features of keyframes in the target scene based on preprocessed multimodal data (such as image frames and text semantic tags). In its implementation, the algorithm first performs preprocessing operations such as noise reduction and size normalization on the input image frame data. Then, it uses convolutional neural networks (such as ResNet and YOLO series) to extract the texture distribution and object detection information of the image. This preprocessing refers to a series of processing operations performed by the data preprocessing unit on the acquired multimodal data and device operating parameters. Specifically, this includes filtering out noise and outliers, unifying the data formats of different modalities, and performing feature filtering on the acquired multimodal data and device operating parameters through timestamp alignment to obtain keyframe images, effective text segments, and effective speech intervals. The preprocessed multimodal data comes from the multimodal data acquisition unit. This unit deploys a group of perception sensors (including image sensors, text input devices, and voice microphones) on the interactive interface of the target scene to collect image frame data, text string data, and voice waveform data of the target scene in real time. These data are then transmitted wirelessly to the data preprocessing unit, where they are processed to obtain the preprocessed multimodal data. Taking a smart security scenario as an example, the algorithm filters key frames (such as frames where abnormal human behavior occurs) from the surveillance video stream, extracts texture features (such as wall materials and floor paving patterns) through convolutional layers, and uses an object detection module to identify the location and category of objects such as people and vehicles, while also recording the spatial distribution relationships of each object.
[0096] Specifically, the image feature extraction unit also analyzes the multimodal fused image frame data and combines it with preprocessed textual semantic data (such as scene description text and tag keywords) to perform cross-modal feature association. For example, in medical image analysis, the system matches the pixel features of CT images with lesion descriptions in medical record text, focuses on the anatomical regions mentioned in the text through an attention mechanism, extracts the texture features of these regions (such as density distribution and edge contours), and establishes a correspondence between image semantic regions and key nodes in the text, thereby improving the targeting and semantic accuracy of feature extraction.
[0097] The text semantic extraction unit constructs a text semantic extraction algorithm to achieve semantic parsing of preprocessed multimodal data (such as text strings and speech emotion features). The algorithm first performs basic processing on the text string data, such as word segmentation and part-of-speech tagging, and then uses recurrent neural networks (such as LSTM and GRU) or Transformer architectures to extract the semantic intent and key information of the text. Taking a customer service dialogue scenario as an example, the algorithm identifies the problem type (such as product complaint or feature inquiry), extracts key information (such as order number or fault description) from the customer's inquiry text, and analyzes the semantic logic of the statements (such as causal relationships and conditional relationships).
[0098] Furthermore, this unit compares text string data from adjacent interaction stages (such as the content of previous and subsequent dialogues) and combines it with preprocessed voice emotion data (such as intonation and speech rate features) to perform cross-modal semantic analysis. Specifically, the algorithm calculates the semantic similarity of adjacent text segments, identifies key points of semantic change, and uses voice emotion features as weighting factors to adjust the focus of semantic analysis. For example, when a user's voice contains anger, the system will focus on extracting keywords expressing dissatisfaction from the text and analyze the rate of semantic change (such as the speed of semantic transition from normal inquiry to complaint), thereby more accurately grasping the user's true needs and emotional tendencies.
[0099] The foundation of the speech emotion extraction unit is the construction of a speech emotion extraction algorithm. This algorithm calculates speech emotion features based on preprocessed device operating parameters (such as audio sampling data and device status indicators). The algorithm first performs frame segmentation and windowing on the speech waveform data, extracting acoustic features such as Mel-frequency cepstral coefficients (MFCC), fundamental frequency, and energy. Then, it uses classification models such as Support Vector Machines (SVM) and neural networks to identify the emotional tendency (e.g., happiness, anger, neutral) and emotional intensity (e.g., mild dissatisfaction, extreme excitement) of the speech. In a meeting recording scenario, the algorithm analyzes the speaker's speech in real time, labels each statement with an emotion tag, and generates an emotion intensity curve, providing a reference for the emotional dimension of the meeting summary.
[0100] In terms of cross-modal collaboration, the voice emotion extraction unit forms a data closed loop with the image and text feature extraction units. For example, in an in-vehicle interaction system, when the driver's voice shows anxiety, the system automatically retrieves in-vehicle image features (such as the driver's facial expressions) and dashboard text data (such as speed and fuel warning information) at the current moment. The accuracy of emotion recognition is verified through multimodal feature fusion, while providing comprehensive basis for subsequent driving safety warnings.
[0101] The entire feature extraction module's workflow embodies the interactivity and complementarity of multimodal data. Taking an educational scenario as an example, the image feature extraction unit extracts students' facial expressions (such as focus and confusion) and body language (such as raising hands and lowering heads) from classroom monitoring videos. The text semantic extraction unit analyzes the semantic depth of students' questions and assignments, and the speech emotion extraction unit identifies the emotional state of students' voices when answering questions. These three types of features, through cross-modal correlation analysis, can comprehensively assess students' learning status: when the image detects a student frowning (confused expression), the text question involves a knowledge gap, or the tone of voice shows hesitation, the system will comprehensively judge that the student has difficulty understanding the current knowledge point, thereby triggering a personalized tutoring process.
[0102] The feature extraction module's algorithm design fully considers the characteristics of different modalities and the needs of various application scenarios. For image data, it focuses on extracting spatial features and target semantics; for text data, it emphasizes semantic logic and contextual relationships; and for speech data, it focuses on recognizing emotional features and acoustic patterns. Simultaneously, through cross-modal feature fusion, it overcomes the limitations of single-modal analysis, making the extracted features more semantically complete and valuable for application, providing rich feature dimensions for subsequent comprehensive analysis and decision-making.
[0103] Example 4:
[0104] The comprehensive analysis module is the core component of the multimodal data analysis system for achieving semantic consistency assessment. It consists of a multimodal association unit and a preliminary discrimination unit. Through standardization processing, association calculation, and threshold comparison, it achieves preliminary analysis of the semantic consistency of the target scene.
[0105] The multimodal association unit (MAU) begins by standardizing the image features, text semantics, and speech sentiment output from the feature extraction module. This standardization process addresses the differences in feature distribution across different modalities by employing normalization (e.g., Min-Max normalization) or standardization (e.g., Z-Score standardization) methods to map each modal feature to a uniform numerical range, ensuring fairness in subsequent calculations. For example, in a smart conferencing scenario, image features might be 2048-dimensional ResNet feature vectors, text semantics might be 768-dimensional BERT word vectors, and speech sentiment might be 128-dimensional MFCC feature vectors. Standardization converts these features of different dimensions and scales into comparable numerical representations.
[0106] After standardization, the multimodal association unit obtains the first analysis value through association calculation. The association calculation employs a cross-modal attention mechanism or a fusion neural network to mine semantic associations between features of different modalities. Specifically, the algorithm calculates the similarity between image features and text semantics (e.g., cosine similarity), the matching degree between text semantics and speech sentiment (e.g., through sentiment dictionary mapping), and the fit between speech sentiment and image scene (e.g., through scene classification models), and generates a comprehensive first analysis value through weighted summation. Taking an e-commerce customer service scenario as an example, when a user sends a product image (image features) and inputs inquiry text (text semantics), and the customer service representative responds via voice (voice sentiment), the association calculation analyzes the degree of matching between product details in the image and the text inquiry content, as well as the consistency between the customer service representative's voice sentiment and the text semantics, generating a first analysis value reflecting the strength of multimodal semantic association.
[0107] When calculating the similarity between image features and text semantics, the image feature vectors output by the image feature extraction unit (such as vectors containing texture distribution and object detection information) and the text semantic vectors output by the text semantic extraction unit (such as word vectors containing semantic intent and key information) are transformed into the same high-dimensional space. The cosine value of the angle between the two vectors is calculated using the cosine similarity formula. The closer this value is to 1, the stronger the semantic association between the content presented by the image and the description in the text. For example, the cosine similarity between the feature of a "smiling face" detected in the image and the semantics of "happy expression" in the text will be significantly higher than the similarity with the semantics of "sad expression". When calculating the matching degree between text semantics and speech emotion, an association mapping between text semantics and speech emotion is established based on an emotion dictionary. Emotional tendency keywords (such as "happy" and "angry") are extracted from the text semantics. At the same time, the emotion category and intensity identified by the speech emotion extraction unit are determined. By querying the correspondence between words and emotion categories in the emotion dictionary, the number and degree of matching between text emotion keywords and speech emotion categories are calculated. For example, when words such as "pleasant" and "excited" appear in the text, the matching degree with the emotion "joy" identified by the speech emotion will be correspondingly higher. When calculating the fit between voice emotion and image scene, a pre-trained scene classification model is used to analyze image features and determine the scene type to which the image belongs (such as "wedding scene", "argument scene", "library", etc.). Each scene type corresponds to a preset typical emotion tendency (such as "wedding scene" corresponds to "joy", "argument scene" corresponds to "anger"). The emotion tendency obtained by the voice emotion extraction unit is compared with the typical emotion of the scene. The higher the degree of consistency or similarity between the two, the higher the fit. For example, when the voice emotion is "joy", the fit with the image scene classified as "wedding scene" is higher than the fit with the "library" scene.
[0108] The preliminary discrimination unit, based on industry standards and historical data from multimodal data analysis, presets a first benchmark threshold and compares the first analysis value with this threshold to achieve a preliminary analysis of the semantic consistency of the target scenario. Industry standards may come from specific domain standards (such as semantic consistency standards for medical image diagnosis), while historical data is obtained through statistical analysis of a large number of labeled cases (such as in the smart education scenario, probabilistic statistics are performed on the semantic consistency labeling results of 100,000 multimodal teaching data to determine the threshold).
[0109] The specific analysis scheme presents a binary discrimination logic: when the first analysis value is greater than the first baseline threshold, the target scene is determined to be semantically consistent under the current interaction conditions, and the system continues to monitor the scene in real time. For example, in an autonomous driving scenario, if the correlation calculation value of road image features collected by the camera, text data from vehicle sensors (such as speed and fuel level), and voice commands (such as navigation prompts) is higher than the threshold, it indicates that the descriptions of road conditions by each modality are consistent, and the system maintains normal driving mode; when the first analysis value is less than or equal to the first baseline threshold, the target scene is determined to be semantically inconsistent under the current interaction conditions, and the system triggers processing measures and further verification operations. For example, in a smart healthcare scenario, a patient's CT image features show lung nodules, but the medical record text does not mention this abnormality, and the doctor's consultation voice does not involve any related descriptions. In this case, the correlation calculation value may be lower than the threshold, and the system will prompt medical staff to re-examine the image and verify the completeness of the medical record.
[0110] In practical applications, the two units of the comprehensive analysis module form a closely collaborative processing chain. Taking a smart security scenario as an example, when a surveillance camera captures an image of a suspicious person (image features), the access control system records the person's card-swiping text information (text semantics), while the voice recognition system acquires the person's dialogue with security personnel (voice emotion). The multimodal association unit first standardizes the person's features in the image (such as facial expressions and clothing features), the identity information in the text (such as name and department), and the dialogue content in the voice (such as purpose of visit and emotional state). Then, it analyzes the identity matching degree between the image and text (such as the consistency between the face recognition result and the card-swiping information) and the semantic consistency between the text and voice (such as the degree of consistency between the description of the purpose of visit and the voice content) through association calculation, generating a first analysis value. The preliminary judgment unit compares this value with a preset security threshold. If the analysis value is higher than the threshold, it indicates that the person's identity and behavior description are consistent, and the system allows normal passage; if the analysis value is lower than the threshold, such as the image showing that the person's face is obscured and there is tension in the voice, or the text identity information does not match the historical records, then it is determined that the semantics are inconsistent, the system automatically triggers an alarm, and notifies security personnel to conduct on-site verification.
[0111] Another typical application scenario is intelligent financial customer service. When a user conducts business through video customer service, the camera captures the user's facial image (image features), the user inputs a textual request for business processing (textual semantics), and simultaneously communicates with the customer service representative via voice (voice emotion). The multimodal association unit analyzes the matching degree between the user's facial expressions (such as whether they are focused) and the textual request (such as a confused expression corresponding to a complex business request), and the consistency between the textual semantics and the voice content (such as whether verbal confirmation is consistent with the terms entered in the text), generating a first analysis value. The preliminary judgment unit presets a threshold according to the compliance requirements of the financial industry. If the analysis value is higher than the threshold, it indicates that the user's request is clearly expressed and the information across all modalities is consistent, and the business is processed normally. If the analysis value is lower than the threshold, such as the user mentioning "I don't understand the terms" in their voice but confirming agreement in the text, the system will pause the business process and prompt the customer service representative to further explain the terms to the user to ensure the consistency between the user's true intentions and the information across all modalities.
[0112] The comprehensive analysis module is designed to fully consider the complementarity of multimodal data and the complexity of semantic relationships. It eliminates data heterogeneity through standardization, uncovers deep semantic connections through association calculations, and achieves rapid discrimination through threshold comparisons. This design ensures both analytical efficiency and the ability to identify semantic contradictions in multimodal data to a certain extent. This provides necessary preliminary judgments for subsequent hierarchical processing, enabling the system to dynamically adjust processing strategies based on semantic consistency, thereby improving the reliability and practicality of multimodal data analysis.
[0113] Example 5:
[0114] The hierarchical processing module is a core component in the multimodal data analysis system for realizing in-depth semantic analysis and differentiated responses. It consists of a processing index calculation unit and a level determination unit. By combining secondary calculation of contextual factors and comparison of multi-level thresholds, it realizes semantic consistency analysis of the target scene under complex interactive states.
[0115] The processing index calculation unit operates based on the first analysis value output by the comprehensive analysis module. Its core function is to further analyze the semantic consistency of the target scene under different multimodal data and device interaction states, and obtain a second analysis value through correlation calculation. The second analysis value is obtained as follows: based on the comprehensive analysis module's finding of semantic inconsistency in the target scene (i.e., the first analysis value is less than or equal to the first baseline threshold), the processing index calculation unit performs further correlation calculation on the first analysis value in conjunction with contextual factors. Contextual factors include historical interaction data of the target scene, spatial environment parameters (such as illumination and noise at sensor deployment locations), operating status parameters of multimodal devices (such as camera focal length shift and microphone signal-to-noise ratio), and user interaction habits. The processing index calculation unit integrates these contextual factors with real-time collected multimodal data (image features, text semantics, and voice emotion) and device operating parameters, and adjusts the first analysis value through weighted calculation to generate the second analysis value. The weight coefficients of this weighted calculation are dynamically allocated according to the degree of influence of contextual factors on semantic consistency. For example, when a device's operating state is abnormal, its corresponding weight will increase to highlight the impact of that factor on semantic analysis, thereby enabling the second analysis value to more comprehensively reflect the semantic consistency state of the target scene after integrating contextual factors. This unit incorporates a wealth of contextual factors during the calculation process. These factors include, but are not limited to, historical interaction data in the time dimension (such as the semantic consistency change trend of the target scene over a period of time), environmental parameters in the spatial dimension (such as changes in illumination and noise interference at the sensor deployment location), operating parameters in the device state dimension (such as the focal length shift of the camera and the signal-to-noise ratio fluctuation of the microphone), and interaction habits in the user behavior dimension (such as the expression style and operation preferences of a specific user).
[0116] Taking a smart factory equipment fault early warning scenario as an example, when the comprehensive analysis module determines that the first analysis value of the currently collected equipment image features (such as dashboard indicator light status), maintenance log text semantics (such as recent maintenance records), and operating noise voice emotion (such as abnormal vibration sounds) is lower than the benchmark threshold, indicating semantic inconsistency, the processing index calculation unit will retrieve the equipment's past week's fault history data (temporal context), the operating status of adjacent equipment (spatial context), the equipment's current load parameters (equipment status context), and the operator's routine inspection process (user behavior context). Through a multimodal association algorithm, these contextual factors are fused and calculated with the currently collected data to analyze whether the abnormal indicator light is related to excessive load, whether the abnormality not recorded in the maintenance log is due to omissions in the inspection process, and whether the abnormal sound is affected by vibration interference from adjacent equipment, ultimately generating a second analysis value that comprehensively considers the contextual factors.
[0117] The core task of the grading unit is to preset a second baseline threshold and perform a secondary comparative analysis between the second analysis value and this threshold to generate different processing levels. Setting the second baseline threshold is more complex than setting the first baseline threshold, requiring comprehensive consideration of the stringency of industry standards, the security level of the application scenario, and the influence weight of contextual factors. For example, in medical surgical navigation scenarios, the second baseline threshold is set more strictly to avoid surgical risks due to semantic misunderstandings; while in ordinary consumer-grade smart home appliance scenarios, the threshold setting is relatively lenient, allowing for a certain degree of semantic ambiguity.
[0118] The specific analysis scheme presents a three-level judgment logic: when the second analysis value is greater than the second baseline threshold, it indicates that the target scene remains semantically consistent under the interaction conditions after considering comprehensive contextual factors. At this time, a third-level processing information is generated to prompt monitoring personnel to continue observing the target scene. For example, in an intelligent transportation system, monitoring images of a certain road segment show abnormal vehicle trajectories, text traffic flow data records a surge in traffic volume on that road segment, and voice command and dispatch mention temporary construction. The first analysis value of the comprehensive analysis module is lower than the threshold, but the second analysis value calculated by the processing index calculation unit, combining the lane narrowing caused by construction (spatial context) and trajectory fluctuation data during historical construction periods (temporal context), is higher than the threshold. The level judgment unit generates a third-level processing information to prompt monitoring personnel to continue monitoring rather than immediately taking road closure measures.
[0119] When the second analysis value equals the second baseline threshold, it is determined that the target scenario is semantically inconsistent under the interaction conditions after considering comprehensive contextual factors, indicating a potential misunderstanding. At this point, secondary processing information is generated to prompt relevant personnel to immediately conduct a detailed verification of the target scenario. Taking the cargo sorting scenario in smart warehousing as an example, the image recognition system detects that a package label is damaged (image feature), the destination of the package in the text sorting instruction is ambiguous (text semantics), and there is background noise in the voice sorting prompt (voice emotion). The first analysis value is lower than the threshold. The processing index calculation unit calculates the second analysis value, which equals the threshold, by combining the package's logistics history trajectory (temporal context) and the destination distribution of adjacent packages (spatial context). The level determination unit generates secondary processing information, prompting sorting personnel to suspend operations and manually verify the package information.
[0120] When the second analysis value is less than the second benchmark threshold, it is determined that the target scene has significant semantic inconsistency under the interaction conditions after considering comprehensive context factors. At this time, first-level processing information is generated, the processing mechanism is automatically triggered, and the relevant modules are notified to start the semantic correction plan. For example, in an autonomous driving system, the image of the road ahead captured by the camera shows an obstacle (image feature), the text data of the vehicle radar does not identify the obstacle (text semantics), and the voice alarm system does not issue a prompt (voice emotion). The first analysis value is lower than the threshold. The second analysis value calculated by the processing index calculation unit, combined with the radar detection angle deviation (device state context) and the image recognition error caused by weather (environmental context), is still significantly lower than the threshold. The level determination unit generates first-level processing information, the system automatically starts the emergency braking procedure and switches to the backup sensor fusion mode.
[0121] In practical applications, the two units of the hierarchical processing module form a deep decision chain. Taking intelligent medical diagnosis as an example, when the comprehensive analysis module finds that the first analysis value of the patient's medical imaging features (such as nodular shadows on lung CT), electronic medical record text semantics (such as the chief complaint being cough), and consultation voice emotion (such as no obvious abnormal emotions) is lower than the threshold, indicating semantic inconsistency (abnormal imaging but not mentioned in the chief complaint), the processing index calculation unit will retrieve the patient's past medical history (temporal context), family history (spatial context), current medication status (device status context, here referring to the effect of drugs on physiological indicators), and details of the consultation (user behavior context, such as whether symptoms were deliberately concealed). Through correlation calculation, it analyzes whether the nature of the nodules is related to the past medical history and whether the patient did not mention it in the chief complaint due to insufficient understanding of the condition, generating a second analysis value.
[0122] If the second analysis value is greater than the second baseline threshold, it indicates semantic consistency after comprehensive consideration of the context (e.g., the nodule is benign and unrelated to the current cough symptoms), generating tertiary treatment information and recommending regular follow-up. If it is equal to the threshold, it indicates potential diagnostic bias (e.g., the nodule has a malignant tendency but the patient has not provided a complete medical history), generating secondary treatment information and prompting the doctor to conduct in-depth consultation and supplementary examinations. If it is less than the threshold, it indicates significant semantic inconsistency (e.g., imaging features are highly suggestive of malignancy but are not reflected in the medical record and consultation), generating primary treatment information, automatically triggering a multidisciplinary consultation process and initiating further pathological testing plans.
[0123] Another typical scenario is user intent recognition in intelligent customer service systems. When the first analysis value of the user's question image (image features), input text (text semantics), and voice consultation (voice emotion) is lower than a threshold (e.g., a contradiction between the text content and the voice expression), the processing index calculation unit will combine the user's historical consultation records (temporal context), the current conversation's contextual context (spatial context, such as the previously discussed product model), the user's identity tag (device status context, such as VIP user or new user), and dialect features in the voice (user behavior context) to calculate and generate a second analysis value. The level determination unit generates three levels of processing information based on different thresholds: if the semantics are consistent, continue with a regular response; if there is a deviation, prompt customer service personnel to pay attention to the user's potential needs; if there is a significant inconsistency, automatically transfer to senior customer service and retrieve historical interaction records for auxiliary processing.
[0124] The hierarchical processing module effectively solves the semantic misjudgment problem of the comprehensive analysis module when facing complex scenarios by introducing in-depth analysis of contextual factors and a multi-level response mechanism, achieving an upgrade from "preliminary judgment" to "deep decision-making." This design not only improves the system's accuracy in identifying semantic inconsistencies but also allows for differentiated processing measures based on the severity of the inconsistencies. While ensuring system reliability, it optimizes resource allocation and response efficiency, enabling the multimodal data analysis system to better adapt to the complexity and uncertainty of real-world scenarios.
[0125] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0126] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. An artificial intelligence-driven multimodal data analysis system, characterized in that: It includes a data acquisition module, a multimodal fusion module, a feature extraction module, a comprehensive analysis module, and a hierarchical processing module; The data acquisition module is used to set up acquisition nodes in the interactive interface between the target scene and the multimodal device, and deploy acquisition terminals to collect multimodal data of the target scene and operating parameters of the multimodal device in real time, and convert the acquired data into a format. The multimodal fusion module is used to construct a cross-modal semantic association model, and uses alignment fusion technology to process the spatiotemporal characteristics of multimodal data. At the same time, it uses feature embedding technology to fuse the representation features of multimodal data, and performs association simulation of target scene and multimodal device based on real-time collected multimodal data and device operating parameters. The feature extraction module is used to construct image feature extraction algorithms, text semantic extraction algorithms, and speech emotion extraction algorithms, and transmits the real-time collected multimodal data and device operating parameters to the constructed extraction algorithms to calculate and obtain image features, text semantics, and speech emotion. The comprehensive analysis module is used to standardize the acquired image features, text semantics, and voice emotion, perform correlation calculations to obtain a first analysis value, and perform a preliminary comparison with a preset first benchmark threshold to analyze the semantic consistency of the target scene. The hierarchical processing module is used to further calculate and obtain a second analysis value by combining contextual factors when the semantic inconsistency of the target scene is detected, and to perform a second comparison analysis with the second analysis value by setting a second benchmark threshold, so as to further analyze the performance of the target scene under different multimodal data and device interaction states. The multimodal fusion module includes a visualization integration unit. The visualization integration unit is used to import the target scene model into the multimodal fusion model for integration to obtain a digital twin model. Then, it transmits the real-time collected multimodal data of the target scene and the operating parameters of the multimodal devices to time series analysis and feature embedding for dynamic fusion. The dynamic fusion result is then imported into the digital twin model to update the status of the target scene and the multimodal devices in real time, so as to simulate the correlation between the target scene and the multimodal devices. The fused data is displayed through a visualization interface, providing a user interaction interface. The target scene model includes Building Information Modeling (BIM) and 3D models of industrial equipment; The target scene multimodal data includes image frames, text commands, and voice signals; The operating parameters of the multimodal device include real-time sensor data and device operating status indicators.
2. The AI-driven multimodal data analysis system according to claim 1, characterized in that: The data acquisition module includes a multimodal data acquisition unit, a modal feature acquisition unit, and a data preprocessing unit; The multimodal data acquisition unit includes an image acquisition unit, a text acquisition unit, and a voice acquisition unit. It is used to acquire and transmit multimodal data of the target scene in real time by deploying a group of perception sensors on the interactive interface of the target scene, and transmit the data to the data preprocessing unit wirelessly. The group of perception sensors includes an image sensor group, a text input device group, and a voice microphone group. The multimodal data includes image frame data, text string data, and voice waveform data. The image acquisition unit is used to acquire image frame data of the target scene in real time based on the image sensor group. The image sensor group includes a camera, an optical lens and an image encoder, which respectively acquire the resolution, color channels and frame rate of the image data. The text acquisition unit is used to acquire text string data in real time based on the text input device group. The text input device group includes a keyboard input module, a touch input module, and an OCR recognition module. The text string data includes character sequences, input timestamps, and semantic context. The voice acquisition unit is used to acquire voice waveform data in real time based on the voice microphone group, which includes a microphone array, an audio amplifier, and an analog-to-digital converter. The voice waveform data includes sampling rate, bit depth, and audio duration.
3. The AI-driven multimodal data analysis system according to claim 2, characterized in that: The data preprocessing unit is used to filter out noise and outliers from the collected multimodal data and equipment operating parameters, unify the data formats of different modalities, and perform feature filtering on the collected multimodal data and equipment operating parameters by timestamp alignment to obtain keyframe images, effective text segments, and effective speech intervals.
4. The AI-driven multimodal data analysis system according to claim 1, characterized in that: The multimodal fusion module also includes a spatiotemporal alignment unit and a feature fusion unit; The spatiotemporal alignment unit, feature fusion unit, and visualization integration unit work together to construct a cross-modal semantic association model; The spatiotemporal alignment unit includes a coordinate mapping unit and a timing synchronization unit; The coordinate mapping unit extracts the spatial coordinates of the target scene and the device location information from the positioning system, and uses an alignment algorithm to establish a multimodal spatiotemporal mapping model to map the spatial coordinates of the scene, the device pose, and the interaction area. At the same time, it adds typical multimodal features to the target scene, including image semantic regions, text key nodes, and speech emotion intervals. After the initial alignment is completed, a calibration tool is used to define the timestamp offset, spatial coordinate deviation, and semantic association weight of the multimodal data. At the same time, the acquisition start point, processing end point, and interaction range of the multimodal device are set. The constructed spatiotemporal mapping model is statically aligned, dynamically aligned, and continuously aligned to align the spatiotemporal response of the multimodal data of the target scene. The timing synchronization unit is used to input multimodal timing parameters, including image acquisition time, text input time, and voice recording time. After input, timing analysis is performed to synchronize the timing features, semantic features, and emotional features of the target scene at different interaction stages. The feature fusion unit is used to establish a multimodal embedding model, including image feature dimension, text word vector dimension and speech Mel coefficient dimension, and then apply a feature fusion network to simulate multimodal semantic association. The feature embedding technology is used to analyze the multimodal fusion results and evaluate the dimensional distribution, semantic consistency and sentiment matching of the fused features.
5. The AI-driven multimodal data analysis system according to claim 1, characterized in that: The feature extraction module includes an image feature extraction unit, a text semantic extraction unit, and a speech emotion extraction unit; The image feature extraction unit is used to construct an image feature extraction algorithm, calculate and obtain image features of key frames of the target scene based on preprocessed multimodal data, and extract the texture distribution and target detection information of the image; The text semantic extraction unit is used to construct a text semantic extraction algorithm, which calculates and obtains text semantics based on preprocessed multimodal data, and extracts the semantic intent and key information of the interaction process in the target scene. The voice emotion extraction unit is used to construct a voice emotion extraction algorithm, which calculates and obtains voice emotions based on preprocessed device operating parameters, and extracts the emotional tendency and intensity of the interaction process in the target scene.
6. The AI-driven multimodal data analysis system according to claim 5, characterized in that: The image feature extraction unit is used to analyze the multimodal fused image frame data and combine it with the preprocessed text semantic data to calculate and obtain image features, and extract the correspondence between the image semantic region and the text key node.
7. The AI-driven multimodal data analysis system according to claim 5, characterized in that: The text semantic extraction unit is used to calculate and obtain text semantics by comparing text string data in adjacent interaction stages and combining preprocessed voice emotion data, and to extract the semantic change rate of the target scene interaction process.
8. The AI-driven multimodal data analysis system according to claim 1, characterized in that: The comprehensive analysis module includes a multimodal correlation unit and a preliminary discrimination unit; The multimodal association unit is used to standardize the acquired image features, text semantics and voice sentiment, then perform association calculation to obtain the first analysis value, and perform comprehensive analysis of the semantic consistency of the target scene. The preliminary discrimination unit is used to perform a preliminary comparison analysis with the obtained first analysis value based on industry standards and historical data for multimodal data analysis, and to analyze the semantic consistency of the target scene. The specific analysis scheme is as follows: when the first analysis value is greater than the first benchmark threshold, it indicates that the target scene is semantically consistent under the current interaction conditions; when the first analysis value is less than or equal to the first benchmark threshold, it indicates that the target scene is semantically inconsistent under the current interaction conditions.
9. The AI-driven multimodal data analysis system according to claim 1, characterized in that: The hierarchical processing module includes a processing index calculation unit and a level determination unit; The processing index calculation unit is used to combine the obtained first analysis value to further analyze the semantic consistency of the target scene under different multimodal data and device interaction states, and perform correlation calculation to obtain the second analysis value; The level determination unit is used to preset a second benchmark threshold and the obtained second analysis value, perform a second comparative analysis, further analyze the consistency of the target scene semantics after the influence of multiple contextual factors, and generate a corresponding processing level. The specific analysis scheme is as follows: When the second analysis value is greater than the second benchmark threshold, it means that the target scene is still semantically consistent under the interaction conditions after comprehensive contextual factors. At this time, a third-level processing information is generated to prompt the monitoring personnel to continuously observe the target scene. When the second analysis value equals the second baseline threshold, it indicates that the target scenario is semantically inconsistent under the interaction conditions after considering the comprehensive context factors, and there is a potential misunderstanding. At this time, a secondary processing message is generated to prompt relevant personnel to immediately conduct a detailed verification of the target scenario. When the second analysis value is less than the second baseline threshold, it indicates that the target scenario is significantly semantically inconsistent under the interaction conditions after considering the comprehensive context factors. At this time, a primary processing message is generated to automatically trigger the processing mechanism and notify the relevant modules to start the semantic correction plan.
10. An artificial intelligence-driven multimodal data analysis method, applied to an artificial intelligence-driven multimodal data analysis system as described in any one of claims 1 to 9, characterized in that, Includes the following steps: Step 1: Set up the acquisition node in the interaction interface between the target scene and the multimodal device, and deploy the acquisition terminal to collect multimodal data of the target scene and the operating parameters of the multimodal device in real time, and convert the format of the collected data; Step 2: Construct a cross-modal semantic association model, use alignment and fusion technology to process the spatiotemporal characteristics of multimodality, use feature embedding technology to fuse the representation features of multimodality, and perform association simulation of the target scene and multimodal devices based on real-time collected multimodal data and device operating parameters; The visualization integration unit imports the target scene model into the multimodal fusion model to obtain a digital twin model. Then, it transmits the real-time collected multimodal data of the target scene and the operating parameters of the multimodal devices to the time series analysis and feature embedding for dynamic fusion. The dynamic fusion results are then imported into the digital twin model to update the status of the target scene and multimodal devices in real time, so as to simulate the correlation between the target scene and multimodal devices. The fused data is displayed through a visualization interface, providing a user interaction interface. The target scene model includes Building Information Modeling (BIM) and 3D models of industrial equipment; The target scene multimodal data includes image frames, text commands, and voice signals; The operating parameters of the multimodal device include real-time sensor data and device operating status indicators; Step 3: Construct image feature extraction algorithm, text semantic extraction algorithm, and speech emotion extraction algorithm, and transmit the real-time acquired multimodal data and device operating parameters to the constructed extraction algorithms to calculate and obtain image features, text semantics, and speech emotion; Step 4: After standardizing the acquired image features, text semantics, and speech sentiment, perform association calculations to obtain the first analysis value and make a preliminary comparison with the preset first benchmark threshold to analyze the semantic consistency of the target scene; Step 5: If the analysis reveals that the semantics of the target scene are inconsistent, then the second analysis value is obtained by further combining the context factors. The second baseline threshold is preset and the second analysis value is compared and analyzed again to further analyze the performance of the target scene under different multimodal data and device interaction states.
Citation Information
Patent Citations
Multi-modal data intelligent analysis system and method
CN116881335A
Complex scene-oriented end-to-end multi-modal content unified perception method and system
CN121051686A