Multi-modal information retrieval and analysis method for internet real-time information flow

By constructing a cross-modal joint embedding space and a multi-head attention mechanism, combined with a heterogeneous computing platform, the problem of cross-modal retrieval and analysis of multimodal data in real-time Internet information flow was solved, achieving efficient and accurate information retrieval and analysis.

CN121233799APending Publication Date: 2025-12-30北京圆璟科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511312925.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-15
Publication Date
2025-12-30

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively process multimodal data in real-time internet information streams, especially in cross-modal retrieval and analysis. They suffer from high computational complexity, insufficient real-time performance, and low cross-modal alignment accuracy and efficiency. Furthermore, in advanced analysis tasks such as event detection, sentiment analysis, and behavior understanding, existing technologies cannot effectively address the spatiotemporal continuity and temporal dependencies of cross-modal information.

Method used

By constructing a cross-modal joint embedding space, employing a multi-head attention mechanism to weightedly fuse features from various modalities, and using generative models to reorder the retrieval results through feature extraction and data fusion, the final results are generated. Furthermore, by combining heterogeneous computing platforms to optimize computational performance and integrating dedicated methods for these platforms, an end-to-end control method is implemented. Finally, end-to-end differentiable search is achieved through feature extraction and data fusion.

Benefits of technology

It improves the accuracy and real-time performance of cross-modal retrieval, enabling the processing of multimodal data streams under low latency conditions, and enhancing the accuracy and efficiency of event detection, sentiment analysis, and behavior understanding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121233799A_ABST
    Figure CN121233799A_ABST
Patent Text Reader

Abstract

The invention provides a multi-modal information retrieval and analysis method for internet real-time information streams, which comprises the following steps of: acquiring multi-modal data in the internet, including texts, images, videos and audios, and performing timestamp calibration on each modal data to form a multi-modal data set; based on the multi-modal data set, respectively extracting semantic features of each modal through a specific feature extractor, including semantic vectors of texts, visual features of images, spatial-temporal features of videos and audio features of audios; according to the multi-modal information retrieval and analysis method and system for the internet real-time information flow, multi-modal feature alignment is achieved by building a cross-modal joint embedding space, the retrieval precision is improved by adopting a three-stage retrieval assembly line, the calculation performance is optimized by combining model quantification and a heterogeneous calculation platform, and the multi-modal information retrieval and analysis method and system for the internet real-time information flow are achieved. The method has the advantages that the cross-modal retrieval accuracy is improved, the real-time processing capability is enhanced, and the calculation efficiency is optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of Internet information processing technology, and in particular to a multimodal information retrieval and analysis method for real-time Internet information streams. Background Technology

[0002] With the rapid development of internet technology and the widespread adoption of smart devices, global internet information is experiencing explosive growth. Social media platforms, news websites, video sharing platforms, and others generate massive amounts of multimodal data daily, including text, images, video, and audio. This data is not only enormous in quantity but also characterized by its real-time nature, modal diversity, and complex spatiotemporal relationships. How to quickly and accurately extract valuable information from these dynamically changing multimodal information flows has become a significant challenge facing the field of information technology.

[0003] Traditional information retrieval technologies primarily process single-modal data. Text retrieval mainly relies on keyword matching and semantic similarity calculation; image retrieval often employs feature extraction techniques based on convolutional neural networks; and video and audio retrieval primarily depends on specific feature extraction models. However, these single-modal retrieval methods struggle to effectively handle the relationships between multimodal data, limiting the accuracy and practicality of cross-modal retrieval.

[0004] In recent years, the development of deep learning technology has brought new opportunities for multimodal information processing. Pre-trained language models such as BERT and RoBERTa have made significant progress in text feature extraction; visual models such as Transformer and CLIP have performed well in image and video feature extraction; and advanced models such as TAP-PMR have emerged in the field of audio processing. Nevertheless, how to effectively fuse features from these different modalities to construct a unified semantic representation space remains a pressing technical challenge.

[0005] Existing cross-modal retrieval methods also have significant shortcomings in terms of real-time performance. On the one hand, the feature extraction and fusion process of multimodal data is computationally complex, making it difficult to meet the processing requirements of real-time information streams. On the other hand, the accuracy and efficiency of cross-modal alignment need to be improved, especially when processing spatiotemporally related information, where existing methods often struggle to accurately capture the temporal dependencies between different modalities. Furthermore, in advanced analysis tasks such as event detection, sentiment analysis, and behavior understanding, existing systems do not fully utilize the fusion of multimodal information, resulting in limited accuracy and reliability of the analysis results. Summary of the Invention

[0006] In view of this, the embodiments of the present invention aim to provide a multimodal information retrieval and analysis method for real-time Internet information streams, so as to solve or alleviate the technical problems existing in the prior art, and at least provide a beneficial option.

[0007] To address the aforementioned technical problems, this application adopts the following technical solution: providing a multimodal information retrieval and analysis method for real-time internet information streams, comprising the following steps:

[0008] Acquire multimodal data from the Internet, including text, images, videos, and audio, and timestamp each modality of data to form a multimodal dataset;

[0009] Based on the multimodal dataset, semantic features of each modality are extracted using a specific feature extractor, including semantic vectors of text, visual features of images, spatiotemporal features of videos, and audio features of audio.

[0010] A cross-modal joint embedding space is constructed using a contrastive learning method, and a multi-head attention mechanism is used to weight and fuse features from various modalities. Cross-modal alignment is then performed through similarity calculation.

[0011] Based on the cross-modal alignment results, a real-time multimodal retrieval engine is constructed, and the retrieval results are reordered through a generative model to generate the final retrieval results.

[0012] Based on the search results, perform event detection, sentiment analysis, and behavior understanding tasks to identify and analyze key events and trends in the Internet information flow;

[0013] By optimizing computational performance through model quantization and knowledge distillation techniques, and combining this with heterogeneous computing platform deployment, real-time performance and computational efficiency are ensured.

[0014] As a further preferred embodiment of this technical solution: based on the multimodal dataset, semantic features of each modality are extracted using a specific feature extractor, wherein:

[0015] The feature extractor for the text modality is a pre-trained language model, including BERT or RoBERTa, used to generate context-sensitive semantic vectors;

[0016] The feature extractor for the image modality is a convolutional neural network or a visual Transformer. The convolutional neural network is ResNet-152, and the visual Transformer is SwinTransformer, which is used to extract visual features and generate a global semantic representation.

[0017] The feature extractor for the video modality is a video and text association model, which is CLIP4Clip and is used to capture the spatiotemporal semantic features of video frames.

[0018] The feature extractor for the audio modality is a model with an attention mechanism, namely the TAP-PMR model, which is used to focus the text on the audio features of the relevant audio frames.

[0019] As a further preferred embodiment of this technical solution: based on the cross-modal alignment results, a real-time multimodal retrieval engine is constructed, including:

[0020] Probabilistic cross-modal embedding is performed on each modality of data. The probabilistic cross-modal embedding uses the PCME model to map the features of different modalities to a unified semantic space.

[0021] Graph neural networks are used to establish temporal reasoning relationships across modal data in order to handle the spatiotemporal dependencies between different modalities;

[0022] We utilize a multi-head attention mechanism to weightedly fuse features from different modalities, and perform the weighted fusion using the following mathematical expression:

[0023] H = Concat(h1,h2,...,h) k ) W O

[0024] in: W O is the learnable weight matrix, and k is the number of attention heads.

[0025] As a further preferred embodiment of this technical solution: the multimodal retrieval engine includes a three-level retrieval pipeline:

[0026] The first stage is semantic matching based on a cross-modal pre-trained model, which is CLIP or ImageBind.

[0027] The second stage improves cross-modal alignment accuracy by combining a reorderer with contrastive learning and token-level interaction. The reorderer adopts a reordering strategy that combines the COTS model.

[0028] The third stage utilizes a generative retrieval model to generate discrete visual tokens. The generative retrieval model is IRGen, which is used to achieve end-to-end differentiable search.

[0029] The multimodal features are stored in a vector database, namely Milvus, for fast retrieval.

[0030] As a further preferred embodiment of this technical solution, the event detection task includes:

[0031] Spatiotemporal event modeling is performed using a recurrent neural network that incorporates spatiotemporal continuity constraints; the recurrent neural network is an LSTM network.

[0032] The sentiment analysis task employs a multimodal Transformer model, specifically the ViLT model, used to fuse text sentiment polarity and image sentiment labels.

[0033] The behavior understanding task includes multimodal composite retrieval, which locates a specific target through composite retrieval of text and video, and analyzes the target's motion trajectory by combining sensor data, wherein the sensor data is IMU data.

[0034] As a further preferred embodiment of this technical solution, the computational optimization technique includes:

[0035] Model quantization technology is used to compress model weights from FP32 to INT8 to reduce computational resource consumption;

[0036] The knowledge distillation technique employs the DGR model and compresses the model using a distillation loss formula. The mathematical expression for the distillation loss is as follows:

[0037] L distill =KL(p teacher (x)||p student (x))

[0038] Where: p teacher (x) represents the probability distribution of the teacher model output, p student (x) represents the probability distribution output by the student model.

[0039] As a further preferred embodiment of this technical solution: the hardware of the heterogeneous computing platform includes GPU and FPGA, wherein CNN feature extraction is deployed on GPU and point cloud processing is deployed on FPGA. The heterogeneous computing platform deploys feature processing tasks of different modalities on different hardware to optimize the utilization of computing resources.

[0040] To address the aforementioned technical problems, another technical solution adopted in this application is: a multimodal information retrieval and analysis system for real-time Internet information streams, comprising:

[0041] The multimodal data acquisition unit is used to acquire multimodal data, including text, images, videos and audio, from the Internet, and to timestamp each modality of data to form a multimodal dataset.

[0042] The feature extraction unit is used to extract semantic features of each modality based on the multimodal dataset using different feature extractors.

[0043] The cross-modal alignment unit is used to construct a cross-modal joint embedding space through contrastive learning methods, and to use a multi-head attention mechanism to weightedly fuse features from various modalities, and to align different modalities through similarity calculation;

[0044] A multimodal retrieval engine is used to generate real-time retrieval results based on the cross-modal alignment results, and to reorder the retrieval results to generate the final retrieval output.

[0045] The event analysis unit is used to perform event detection, sentiment analysis, and behavior understanding tasks based on the search results, and to identify and analyze key events, sentiment changes, and behavioral trends in the Internet information flow.

[0046] The computation optimization unit is used to optimize computational performance through model quantization, knowledge distillation techniques, and heterogeneous computing platforms to ensure the real-time performance and efficiency of the system.

[0047] To solve the above-mentioned technical problems, another technical solution adopted in this application is: a computer device, the computer device including a processor and a memory coupled to the processor, the memory storing program instructions, when the program instructions are executed by the processor, causing the processor to perform the steps of the multimodal information retrieval and analysis method for real-time Internet information flow as described above.

[0048] To solve the above-mentioned technical problems, another technical solution adopted in this application is: a storage medium storing program instructions capable of implementing the multimodal information retrieval and analysis method for real-time Internet information flow as described above.

[0049] The embodiments of the present invention have the following advantages due to the adoption of the above technical solutions:

[0050] This invention provides a multimodal information retrieval and analysis method and system for real-time Internet information flow. It achieves multimodal feature alignment by constructing a cross-modal joint embedding space, improves retrieval accuracy by adopting a three-level retrieval pipeline, and optimizes computational performance by combining model quantization and heterogeneous computing platforms. It has the advantages of improving cross-modal retrieval accuracy, enhancing real-time processing capabilities, and optimizing computational efficiency.

[0051] The above overview is for illustrative purposes only and is not intended to be limiting in any way. In addition to the illustrative aspects, embodiments, and features described above, further aspects, embodiments, and features of the invention will become readily apparent from the accompanying drawings and the following detailed description. Attached Figure Description

[0052] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0053] Figure 1 This is a flowchart of the method of the present invention;

[0054] Figure 2 This is a schematic diagram of the modules of the system of the present invention;

[0055] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0056] The technical solutions of this invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them. The components of this invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely to represent selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without inventive effort are within the scope of protection of this invention. It should be noted that similar reference numerals and letters in the following drawings indicate similar items; therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. Furthermore, in the description of this invention, the terms "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0057] Figure 1 This is a flowchart illustrating a multimodal information retrieval and analysis method for real-time internet information streams according to an embodiment of the present invention. It should be noted that if substantially the same result is obtained, the method of this application is not necessarily identical. Figure 1 The illustrated process sequence is limited. For example... Figure 1 As shown: A multimodal information retrieval and analysis method for real-time internet information streams includes the following steps:

[0058] This process involves acquiring multimodal data from the internet, including text, images, videos, and audio, and timestamping each modality to form a multimodal dataset. Based on this dataset, semantic features for each modality are extracted using a specific feature extractor, including semantic vectors for text, visual features for images, spatiotemporal features for videos, and audio features for audio. A cross-modal joint embedding space is constructed using a contrastive learning method, and multi-head attention is employed to weightedly fuse the features from each modality. Cross-modal alignment is then performed through similarity calculation. Based on the cross-modal alignment results, a real-time multimodal retrieval engine is built. The retrieval results are reordered using a generative model, and the final retrieval results are generated. According to the retrieval results, event detection, sentiment analysis, and behavior understanding tasks are performed to identify and analyze key events and trends in the internet information flow. Computational performance is optimized through model quantization and knowledge distillation techniques, and deployment on a heterogeneous computing platform ensures real-time performance and computational efficiency.

[0059] In existing technologies, multimodal information processing techniques mainly focus on optimizing a single modality, such as text-based keyword matching or image feature extraction based on convolutional neural networks. However, complex semantic relationships exist between multimodal data such as text, images, videos, and audio in real-time internet information streams, making it difficult for single-modal processing methods to capture cross-modal collaborative information. For example, a video clip of a news event on a social media platform may contain voice descriptions of key figures, while the corresponding text report may not fully cover the details in the video. Existing technologies lack effective cross-modal joint modeling methods, resulting in limited accuracy of retrieval results. Furthermore, traditional methods struggle to meet low-latency requirements when processing massive amounts of real-time data due to high computational resource consumption, such as the inability to quickly locate multimodal correlation information in breaking news events.

[0060] To address the aforementioned issues, a technical solution is needed that can fuse multimodal features and achieve efficient real-time retrieval. First, the challenge of cross-modal alignment lies in the heterogeneity of different modal data; for example, the discrete symbols in text differ significantly from the continuous pixel space of an image. Constructing a joint embedding space through contrastive learning can map different modalities to a unified semantic space, but how to achieve effective weighted fusion of features still requires further exploration. Second, real-time retrieval requires the system to process multimodal data streams with low latency, necessitating optimized computational performance. To address these issues, the inventors propose a phased approach: first, independently extracting semantic features from each modality using a feature extractor; then, dynamically fusing features using an attention mechanism; and finally, improving efficiency through model compression and heterogeneous computing platforms.

[0061] Multimodal datasets refer to structured data collections containing text, images, videos, and audio with timestamps. This can be achieved by crawling publicly available data and adding time stamps to preserve the temporal relevance of the data. Specific feature extractors are deep learning models designed for different modalities; for example, pre-trained language models can be used for text, and convolutional neural networks for images. Independent feature extraction preserves the original semantic information of each modality. Cross-modal joint embedding space maps features from different modalities to a unified vector space through contrastive learning. This can be achieved by aligning text and image features using similarity calculations to realize cross-modal semantic association. Multi-head attention mechanisms involve parallel computation of multiple attention weights and fusion of results. For example, a learnable weight matrix can dynamically adjust the contribution of different modal features to address feature heterogeneity. Generative model re-ranking optimizes the ranking of search results using sequence generation models. For example, generating discrete visual markers to reconstruct query intent improves search accuracy. Model quantization converts floating-point model parameters to a low-precision format, such as compressing FP32 weights to INT8, reducing memory usage and computation time.

[0062] Specifically, the system first collects multimodal data in real time from the internet, such as text and image posts and short videos from social media platforms, adding a timestamp accurate to milliseconds to each data point. Then, text data is converted into semantic vectors using a pre-trained language model, image data has visual features extracted using a convolutional neural network, video data uses a spatiotemporal model to capture dynamic changes between frames, and audio data uses a model with an attention mechanism to focus on key segments. After each modality's features are input into a cross-modal joint embedding space, a multi-head attention mechanism calculates the association weights between different modalities, such as the correspondence between text descriptions and image regions, and generates a unified representation through weighted fusion. The retrieval engine performs similarity matching based on the fused features, generates preliminary retrieval results, and then uses a generative model to re-rank the results, for example, generating relevant visual tags based on user queries to optimize the ranking. Finally, the system analyzes the spatiotemporal distribution, sentiment, and behavioral patterns of events based on the retrieval results, such as detecting the propagation path of sudden events. The computational optimization stage converts model parameters from high-precision to low-precision format and combines heterogeneous hardware such as GPUs and FPGAs to accelerate the feature extraction and retrieval process.

[0063] Compared to existing technologies, current methods typically process each modality of data independently, such as building separate text and image indexes, leading to reliance on post-processing fusion strategies for cross-modal retrieval. This solution achieves dynamic fusion at the feature level through a joint embedding space and multi-head attention mechanism; for example, text queries can directly match keyframe features in videos. Traditional systems use single hardware to handle multimodal tasks, such as using only GPUs for all computations, resulting in low resource utilization. This solution utilizes a heterogeneous computing platform, deploying image processing on GPUs and audio processing on FPGAs, achieving optimized hardware resource allocation. Furthermore, existing technologies lack end-to-end re-ranking mechanisms; this solution introduces generative models to reconstruct query semantics, such as converting user-input natural language into visual tag sequences, improving the relevance of search results.

[0064] Through the above technical solutions, this invention solves the problem of insufficient cross-modal alignment accuracy for multimodal data, such as accurately associating audio content with video changes in video retrieval. Simultaneously, model quantization and heterogeneous computing reduce system latency, for example, controlling retrieval response time to the millisecond level for data volumes of tens of millions. Furthermore, the generative re-ranking mechanism improves retrieval accuracy in complex query scenarios, such as locating target video segments even under fuzzy descriptions.

[0065] This invention further proposes a method based on a multimodal dataset to extract semantic features for each modality using specific feature extractors. Specifically: the text modality feature extractor is a pre-trained language model, including BERT or RoBERTa, used to generate context-sensitive semantic vectors; the image modality feature extractor is a convolutional neural network or a visual Transformer, with ResNet-152 as the convolutional neural network and SwinTransformer as the visual Transformer, used to extract visual features and generate a global semantic representation; the video modality feature extractor is a video and text association model, specifically CLIP4Clip, used to capture the spatiotemporal semantic features of video frames; and the audio modality feature extractor is a model with an attention mechanism, specifically the TAP-PMR model, used to focus the text on the audio features of relevant audio frames.

[0066] Among them, pre-trained language models refer to deep learning models trained on large-scale corpora, specifically implemented using BERT or RoBERTa. These models capture the contextual relevance of text through self-attention mechanisms, generating semantically coherent vector representations. Convolutional neural networks refer to deep network structures based on multi-layer convolutional operations, specifically implemented using ResNet-152. They alleviate the gradient vanishing problem through residual connections and extract local and global features of images. Visual Transformers refer to models that apply the Transformer architecture to the vision domain, specifically implemented using SwingTransformer. They reduce computational complexity through windowing strategies and achieve long-range dependency modeling between image blocks. Video and text association models refer to cross-modal models that simultaneously process video frame sequences and text descriptions, specifically implemented using CLIP4Clip. They map video clips and text descriptions to a unified semantic space through contrastive learning. Audio models with attention mechanisms refer to audio processing models that introduce attention mechanisms, specifically implemented using TAP-PMR. They enhance the correlation between audio features and text semantics through text-guided attention weight allocation.

[0067] Specifically, based on the characteristics of different modalities in multimodal data, highly adaptable feature extraction models are employed for semantic representation. For text modalities, pre-trained language models generate context-aware vectors; for example, the BERT model encodes text sequences using a bidirectional Transformer encoder to capture semantic relationships between words. For image modalities, deep convolutional networks or visual Transformers are used to extract visual features; for instance, ResNet-152 extracts multi-level features through residual module stacking, while SwinTransformer establishes global relationships between image patches through a sliding window mechanism. For video modalities, the CLIP4Clip model is used to jointly embed video frame sequences with text descriptions, capturing dynamic semantic changes through temporal modeling. For audio modalities, the TAP-PMR model achieves fine-grained alignment between text and audio; for example, attention masks are generated using text descriptions to filter semantically relevant audio segments. Through this combination of heterogeneous models, each modal feature retains its own semantic characteristics while providing a highly discriminative representation foundation for subsequent cross-modal alignment.

[0068] Compared to existing technologies, which typically employ a single model to process data across different modalities (e.g., using only CNNs for images and videos), this approach suffers from insufficient heterogeneity across modal feature spaces. Our proposed solution selects the optimal model based on the physical characteristics of each modality. For instance, CLIP4Clip is used for spatiotemporal modeling of time-dependent video data, while an attention mechanism is employed to enhance modal interaction for text-sensitive audio data. Furthermore, while existing technologies rely heavily on traditional CNNs for visual feature extraction, our solution introduces a visual Transformer to improve global semantic modeling capabilities, while simultaneously preserving local detail features through ResNet-152, resulting in complementary representations.

[0069] Through the above technical solutions, this invention effectively solves the problem of insufficient semantic representation in multimodal feature extraction. By adapting specialized models to different modalities, the feature extraction accuracy of text, images, videos, and audio is specifically improved. For example, the visual Transformer's ability to capture the overall semantics of images is superior to traditional CNNs, and CLIP4Clip's ability to model temporal relationships in videos is superior to single-frame feature extraction methods. Simultaneously, the audio model with an attention mechanism reduces the interference of irrelevant audio segments on semantic analysis through a text-guided focusing mechanism, thereby providing more accurate feature input for cross-modal alignment.

[0070] This invention further proposes a real-time multimodal retrieval engine based on cross-modal alignment results. This includes performing probabilistic cross-modal embedding on each modality's data separately, using a PCME model to map features from different modalities to a unified semantic space; employing a graph neural network to establish temporal reasoning relationships between cross-modal data to handle spatiotemporal dependencies between different modalities; and utilizing a multi-head attention mechanism to weightedly fuse features from each modality, using a mathematical expression for weighted fusion. The mathematical expression is:

[0071] H = Concat(h1,h2,...,h) k W O

[0072] in: W O is the learnable weight matrix, and k is the number of attention heads.

[0073] Probabilistic cross-modal embedding refers to mapping features from different modalities to a unified semantic space through probabilistic modeling. This can be implemented using the PCME model, which models the uncertainty between modalities through probability distributions, addressing the heterogeneity of features across different modalities. Graph neural networks are network structures used to model temporal relationships, specifically spatiotemporal graph convolutional networks. These networks model the dynamic associations of cross-modal data through the relationships between nodes and edges, capturing the spatiotemporal dependencies between different modalities. Multi-head attention mechanisms are feature fusion methods that execute multiple sets of attention calculations in parallel. This can be implemented using learnable weight matrices combined with multiple attention heads. This mechanism enhances the robustness of cross-modal feature fusion through multi-dimensional feature interactions.

[0074] Specifically, in building a real-time multimodal retrieval engine, the PCME model is first used to map text, image, video, and audio features to a unified probabilistic semantic space, eliminating representational differences between modalities. Then, a graph neural network is employed to perform temporal modeling on the cross-modal data. For example, timestamp-labeled multimodal data are used as graph nodes, connecting spatiotemporally related data points with edges, and graph convolution operations are used to pass temporal dependency information. Finally, a multi-head attention mechanism is used to dynamically weight the fused features. For instance, feature vectors from different modalities are input into the attention layer to generate attention weight matrices for multiple subspaces. The features of each modality are then weighted and summed according to a mathematical expression to generate a cross-modal joint representation.

[0075] Compared to existing technologies, current cross-modal retrieval methods typically employ fixed-weight feature concatenation or a single attention mechanism, which struggles to effectively handle complex spatiotemporal dependencies between modalities. Our proposed solution, however, models modal uncertainty through probabilistic embedding, captures dynamic temporal correlations using graph neural networks, and leverages multi-head attention to achieve multi-dimensional feature interaction, significantly improving the accuracy of cross-modal alignment and real-time inference capabilities.

[0076] Through the above technical solutions, the present invention can effectively solve the problem of cross-modal retrieval bias caused by the heterogeneity of multimodal data, enhance the modeling ability of dynamic time-series information, and improve the flexibility and robustness of feature fusion through multi-head attention mechanism, thereby achieving more accurate multimodal retrieval and analysis in real-time information flow scenarios.

[0077] This invention further proposes a multimodal retrieval engine comprising constructing a three-stage retrieval pipeline. The first stage performs semantic matching based on a cross-modal pre-trained model, which is CLIP or ImageBind. The second stage improves cross-modal alignment accuracy through a re-ranker combined with contrastive learning and token-level interaction, with the re-ranker employing a re-ranking strategy combined with a COTS model. The third stage utilizes a generative retrieval model to generate discrete visual tokens, with the generative retrieval model being IRGen, used to achieve end-to-end differentiable search. Multimodal features are stored in a vector database, which is Milvus, for fast retrieval.

[0078] Among them, cross-modal pre-trained models refer to models jointly trained on large-scale multimodal data, capable of mapping different modal data to a unified semantic space for similarity calculation. Specifically, CLIP or ImageBind can be used. CLIP aligns the embedding representations of images and text through contrastive learning, while ImageBind supports joint embedding of six modalities. Such models can solve the problem of the universality of cross-modal semantic matching. The re-ranking unit is a module that performs secondary optimization on the initial retrieval results. Specifically, it uses the COTS model combined with contrastive learning and a token-level interaction mechanism to correct the ranking bias of the initial retrieval results through fine-grained feature alignment, improving cross-modal matching accuracy. Generative retrieval models are models that can directly generate discrete identifiers for target data. Specifically, the IRGen framework is used, transforming the retrieval task into a sequence generation problem, realizing an end-to-end differentiable search process, and reducing the complexity of traditional retrieval systems. Vector databases are storage systems that support fast similarity retrieval of high-dimensional vectors. Specifically, Milvus is used, which builds an index structure based on the approximate nearest neighbor algorithm, accelerating the matching efficiency of multimodal features.

[0079] Specifically, the three-stage retrieval pipeline works as follows: In the first stage, the input multimodal query requests undergo cross-modal semantic matching using CLIP or ImageBind models, such as calculating the similarity between text queries and video clips, generating a preliminary retrieval result list. In the second stage, the preliminary results are input into the COTS reorderer. This module optimizes the ranking of candidate samples using a contrastive learning loss function and combines a token-level interaction mechanism to capture fine-grained associations between text and visual features, such as matching local regions of an image with keywords in the text description. In the third stage, the IRGen model performs generative retrieval on the reordered candidate set, generating discrete visual token sequences through a decoder and directly outputting the target data identifier most relevant to the query semantics. Throughout the process, multimodal features are stored in the Milvus vector database, which supports millisecond-level response based on a sharded and distributed architecture, maintaining sub-second retrieval latency even with tens of millions of data points.

[0080] Compared to existing technologies, current cross-modal retrieval systems typically employ a single-stage retrieval architecture, such as relying solely on CLIP for coarse-grained matching, resulting in insufficient response accuracy for complex queries. This solution, however, organically combines semantic matching, reordering optimization, and generative retrieval through a three-stage pipeline design. For example, the COTS reorderer addresses the shortcomings of traditional methods in fine-grained alignment through a token-level interaction mechanism, while the IRGen model avoids the dependence of traditional inverted indexes on discrete identifiers. Furthermore, existing systems often use a single storage engine, making it difficult to handle the high-dimensionality of multimodal features. The Milvus vector database, through quantization encoding and approximation algorithms, significantly improves the retrieval efficiency of massive datasets.

[0081] Through the above technical solutions, this invention can effectively improve the accuracy and efficiency of multimodal retrieval. The first stage achieves rapid filtering of large-scale data through a general cross-modal model; the second stage utilizes a fine-grained interaction mechanism to correct sorting errors; and the third stage reduces retrieval latency through a generative model. For example, in video retrieval scenarios, the text query "dunking action on a sports field" can be matched with relevant video clips through CLIP. The COTS model further associates visual tokens such as "basketball hoop" and "jump" with text descriptions, and IRGen ultimately generates accurate video identifiers. Simultaneously, the Milvus database supports high-concurrency real-time retrieval requirements; for example, in monitoring trending events on social media, it can handle thousands of cross-modal query requests simultaneously.

[0082] This invention further proposes that the event detection task uses a recurrent neural network with spatiotemporal continuity constraints to model spatiotemporal events, the sentiment analysis task uses a multimodal Transformer model to fuse text sentiment polarity and image sentiment tags, and the behavior understanding task includes multimodal composite retrieval, which locates specific targets through composite retrieval of text and video, and analyzes the target's motion trajectory by combining sensor data.

[0083] Among them, recurrent neural networks (RNNs) refer to sequence modeling networks with memory unit structures, specifically implemented using LSTM networks. They are used to capture long-term dependencies in time-series data, addressing the problem of insufficient spatiotemporal continuity modeling in event detection. Multimodal Transformer models are attention mechanism models capable of simultaneously processing text and visual modalities, specifically implemented using ViLT models. They achieve emotional feature interaction between text and images through cross-modal attention layers, addressing the bias problem in single-modal sentiment analysis. Sensor data refers to motion parameters collected by inertial measurement units (IMUs), specifically using IMU data. This data supplements the lack of physical information in motion trajectory analysis of video modalities, improving the physical spatial correlation of behavior understanding tasks.

[0084] Specifically, in event detection, the LSTM network receives multimodal feature sequences with timestamps, filters key spatiotemporal information through a gating mechanism, and establishes a continuous correlation of event occurrences. In sentiment analysis, the ViLT model interacts with text sentiment word vectors and image region features through attention, generating a fused sentiment polarity probability distribution. In behavior understanding, combined text and video retrieval locates target objects through cross-modal matching, and combines acceleration and angular velocity information from IMU data to construct a three-dimensional motion trajectory model, achieving spatial motion analysis of target behavior.

[0085] Compared with existing technologies, traditional event detection methods rely on independent frame analysis, which leads to temporal breaks. In contrast, this solution uses an LSTM network to track event states across time steps. Existing sentiment analysis methods mostly use single-modal classifiers, while this solution eliminates the differences in sentiment expression between images and text through a cross-modal attention mechanism. Existing behavior understanding methods lack physical space data support, while this solution establishes a physical coordinate system mapping of motion trajectories by fusing IMU sensor data.

[0086] Through the above technical solutions, this invention solves the problem of missing spatiotemporal continuity in cross-modal event detection, improves the accuracy of emotion polarity judgment in complex scenarios, enhances the spatial interpretability of motion trajectories in behavior understanding tasks, and systematically optimizes the ability of event association analysis, emotion cross-validation, and behavior space reasoning in Internet information flow.

[0087] This invention further proposes computational optimization techniques, including using model quantization to compress model weights from FP32 to INT8 to reduce computational resource consumption, and a knowledge distillation technique using the DGR model to compress the model through a distillation loss formula, the mathematical expression of which is:

[0088] L distill =KL(p teacher (x)||p student (x))

[0089] Where: p teacher (x) represents the probability distribution of the teacher model output, p student (x) represents the probability distribution output by the student model.

[0090] Among these, model quantization technology refers to the technique of converting floating-point weights in a neural network model into low-bit integer representations. This can be achieved using linear or non-linear quantization methods, reducing memory usage and computational complexity by decreasing the data bit width. Knowledge distillation technology refers to the technique of transferring knowledge from a complex teacher model to a lightweight student model. This can be achieved using distillation methods based on output probability distribution matching, compressing the model by minimizing the output difference between the teacher and student models. The DGR model refers to a distillation framework based on dynamic gradient adjustment, which can be implemented by jointly training the teacher and student models, balancing model accuracy and compression ratio by dynamically adjusting the weights of the distillation loss. The distillation loss formula is a mathematical expression that measures the difference in output distribution between the teacher and student models. KL divergence or cross-entropy can be used as metrics, and knowledge transfer is achieved by optimizing this loss function.

[0091] Specifically, model quantization technology reduces the data storage space and computational unit bit width requirements for individual parameters by converting 32-bit floating-point model parameters into 8-bit integer representations, thereby reducing memory bandwidth consumption and computational resource consumption. Knowledge distillation technology constructs a joint training framework for teacher and student models, using the probability distribution of the teacher model's output to guide the parameter updates of the student model, allowing the student model to inherit the reasoning ability of the teacher model while maintaining a smaller scale. The DGR model dynamically adjusts the weight ratio of distillation loss and task loss, adaptively balancing the relationship between model compression and task performance during training, avoiding accuracy degradation due to excessive compression. The distillation loss formula mathematically defines a similarity measure between the teacher model's output and the student model's output, serving as an optimization objective during training to guide the student model parameters closer to the teacher model's knowledge space.

[0092] Compared to existing technologies, traditional model compression methods typically employ quantization or distillation techniques alone. Quantization results in significant accuracy loss, while distillation requires complex manual parameter tuning. This solution combines quantization and distillation techniques to reduce computational resource requirements while maintaining model accuracy through knowledge transfer. The dynamic adjustment mechanism of the DGR framework further avoids an imbalance between accuracy and efficiency during training. Existing technologies using fixed-weight distillation loss functions struggle to adapt to the complexity of multimodal retrieval tasks. This solution utilizes a distillation loss formula that achieves cross-modal knowledge transfer through probability distribution matching, enhancing the robustness of the compressed model in multimodal scenarios.

[0093] Through the above technical solutions, this invention can effectively reduce the computational resource consumption of multimodal retrieval models, improve model inference speed while maintaining retrieval accuracy, and meet the stringent requirements for computational efficiency in real-time information flow processing scenarios. Simultaneously, the dynamically adjusted distillation mechanism avoids the trade-off between accuracy and efficiency in traditional compression methods, enabling the optimized model to achieve the best balance between resource utilization efficiency and task performance on heterogeneous computing platforms.

[0094] The present invention further proposes that the hardware of the heterogeneous computing platform includes GPU and FPGA, wherein CNN feature extraction is deployed on GPU and point cloud processing is deployed on FPGA. The heterogeneous computing platform deploys feature processing tasks of different modalities on different hardware to optimize the utilization of computing resources.

[0095] In this solution, GPU refers to Graphics Processing Unit, specifically implemented using NVIDIA A100 or V100 models. It is suitable for parallel computing-intensive tasks and is used to perform convolutional neural network feature extraction, leveraging its high parallel computing capabilities to accelerate feature processing of image and video modalities. FPGA refers to Field-Programmable Gate Array, specifically implemented using the Xilinx UltraScale+ series. It is suitable for low-latency, customizable computing tasks and is used to process modal features requiring real-time response, such as point cloud data, reducing computational latency through hardware logic optimization. CNN feature extraction refers to the operation of feature encoding of image or video data based on convolutional neural networks, specifically implemented using ResNet-152 or EfficientNet models, rapidly extracting visual semantic features through the parallel computing capabilities of the GPU. Point cloud processing refers to the operation of feature extraction and spatial relationship modeling of discrete point sets in three-dimensional space, specifically implemented using PointNet++ or RandLA-Net algorithms. The hardware programmability of the FPGA optimizes the computation process and reduces data transmission overhead.

[0096] Specifically, in heterogeneous computing platforms, CNN feature extraction tasks for image and video modalities are assigned to GPUs for execution, leveraging their massively parallel computing architecture to accelerate convolutional operations. For example, when processing high-resolution video frames, GPUs can simultaneously compute visual features from multiple regions. For point cloud data, FPGAs implement feature extraction pipelines through pre-configured hardware logic circuits. For instance, in real-time sensor data processing scenarios, FPGAs can directly interface with IMU devices to convert raw point cloud data into structured feature vectors, avoiding frequent data exchanges with CPUs or GPUs. By assigning different modal tasks to suitable hardware units, computational resource utilization is optimized. For example, in mixed-modal data processing, GPUs and FPGAs can execute their respective tasks in parallel, reducing overall processing latency.

[0097] Compared to existing technologies, traditional multimodal processing systems typically rely on a single computing unit (such as a CPU or GPU) to handle all modal tasks, leading to resource contention and reduced efficiency. For example, point cloud processing running on a GPU may encounter computational bottlenecks due to insufficient data throughput, while using customized FPGA hardware significantly reduces the latency of point cloud feature extraction. Furthermore, existing technologies lack hardware adaptation strategies for different modal features, while this solution achieves task-level hardware division of labor through a heterogeneous computing platform, enabling computationally intensive tasks and low-latency tasks to be executed efficiently on GPUs and FPGAs, respectively.

[0098] Through the above technical solution, this invention solves the problem of insufficient real-time performance caused by unreasonable hardware resource allocation in multimodal data processing, improves the parallel processing efficiency of image, video, and point cloud data, and reduces resource competition between cross-modal tasks. For example, in real-time information flow analysis scenarios, the GPU can focus on visual feature extraction, while the FPGA synchronously processes sensor data, enabling the system to meet the real-time retrieval requirements of high concurrency and low latency, while reducing redundant consumption of computing resources.

[0099] The present invention also provides embodiments implemented using the method steps of the present invention, including:

[0100] Step 1: Multimodal data acquisition and time calibration:

[0101] Data stream D = {d} from social media platforms t}, where each sample These represent text, image, video, and audio data at timestamp t, respectively.

[0102] To ensure timing consistency, a timestamp function is applied to the data stream:

[0103]

[0104] Where T(d) t ) represents an aligned multimodal dataset.

[0105] Step 2: Feature extraction and vectorization representation:

[0106] Features from different modalities are extracted using a pre-trained model:

[0107] Text modality:

[0108]

[0109] Where f BERT For the RoBERTa model, dimension d t =768.

[0110] Image modality:

[0111]

[0112] Where d i =1024.

[0113] Video modality:

[0114]

[0115] Where d v =1024.

[0116] Audio modality:

[0117]

[0118] Where d a =512.

[0119] This ultimately forms a multimodal feature set:

[0120]

[0121] Step 3: Cross-modal alignment and weighted fusion. A joint embedding space is constructed through contrastive learning, using InfoNCE loss:

[0122]

[0123] Where sim(·) represents the cosine similarity and τ is the temperature parameter.

[0124] Simultaneously, a multi-head attention mechanism is used to weight the features:

[0125]

[0126] in:

[0127] Attention weights;

[0128] W m A mode-specific projection matrix.

[0129] Obtain a unified embedding vector Where d = 1024.

[0130] Step 4: Real-time multimodal retrieval engine:

[0131] A three-stage retrieval method is used:

[0132] 1. Coarse Search (CLIP Model):

[0133]

[0134] 2. Reordering (COTS Interaction Model):

[0135] S rank (q,z t )=λ1S clip (q,z t )+λ2φ(q,z t )

[0136] Where φ(q,z) t ) represents the token-level interaction score, and λ1 and λ2 are weight parameters.

[0137] 3. Generative retrieval (IRGen model):

[0138] Generate discrete token representation And minimize reconstruction error:

[0139]

[0140] All vectors are stored in the vector database Milvus, which supports Approximate Nearest Neighbor (ANN) retrieval with a query complexity of approximately O(logN).

[0141] Step 5: Event Detection and Sentiment Analysis:

[0142] Event detection: Timing modeling using LSTM.

[0143]

[0144] Determining sudden events using a threshold $\delta$:

[0145]

[0146] Sentiment Analysis: Using the ViLT model, the text polarity vector p is... t Image sentiment tag g t Fusion:

[0147] s t =σ(W[p t g t ])

[0148] Where σ is the sigmoid function, with an output range of [0,1].

[0149] Step 6: Performance Optimization and Deployment

[0150] Model quantization: Weights are quantized to INT8.

[0151] w int8 =round(w fp32 / s)

[0152] Where s is the scaling factor.

[0153] Knowledge distillation: Distillation losses:

[0154]

[0155] Where p T ,p S The output distributions are for the teacher model and the student model, respectively.

[0156] Dynamic batch processing: Set batch size B t With system load ρ t Adaptive adjustment:

[0157]

[0158] Pipeline parallelism: The task is divided into P = {P1, P2, P3}, corresponding to preprocessing, inference, and post-processing respectively. The system throughput is:

[0159]

[0160] Through the above embodiments, when processing real-time data streams from social media, the system can return cross-modal retrieval results within 200ms, improve the event detection accuracy to 92.3%, achieve an F1 score of 89.7% for sentiment analysis, and achieve a 1.8x inference acceleration effect on the GPU+FPGA heterogeneous platform.

[0161] Figure 2 This is a functional module diagram of a multimodal information retrieval and analysis system for real-time Internet information streams, as described in an embodiment of this application. Figure 2 As shown, a multimodal information retrieval and analysis system for real-time internet information streams includes:

[0162] The multimodal data acquisition unit is used to acquire multimodal data, including text, images, videos and audio, from the Internet, and to timestamp each modality of data to form a multimodal dataset.

[0163] The feature extraction unit is used to extract semantic features of each modality based on the multimodal dataset using different feature extractors.

[0164] The cross-modal alignment unit is used to construct a cross-modal joint embedding space through contrastive learning methods, and to use a multi-head attention mechanism to weightedly fuse features from various modalities, and to align different modalities through similarity calculation;

[0165] A multimodal retrieval engine is used to generate real-time retrieval results based on the cross-modal alignment results, and to reorder the retrieval results to generate the final retrieval output.

[0166] The event analysis unit is used to perform event detection, sentiment analysis, and behavior understanding tasks based on the search results, and to identify and analyze key events, sentiment changes, and behavioral trends in the Internet information flow.

[0167] The computation optimization unit is used to optimize computational performance through model quantization, knowledge distillation techniques, and heterogeneous computing platforms to ensure the real-time performance and efficiency of the system.

[0168] For other details regarding the implementation techniques of each module in the above embodiments, please refer to the description in the multimodal information retrieval and analysis method for real-time Internet information flow in the above embodiments, which will not be repeated here.

[0169] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For system-type embodiments, since they are basically similar to method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.

[0170] An electronic device according to embodiments of the present disclosure includes a memory and a processor. The memory is used to store non-transitory computer-readable instructions. Specifically, the memory may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may, for example, include random access memory (RAM) and / or cache memory. The non-volatile memory may, for example, include read-only memory (ROM), a hard disk, flash memory, etc.

[0171] Those skilled in the art will understand that, in order to solve the technical problem of how to achieve a good user experience, this embodiment may also include well-known structures such as communication buses and interfaces, and these well-known structures should also be included within the protection scope of this disclosure.

[0172] like Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure. It illustrates a structural schematic diagram suitable for implementing the electronic device in the embodiment of the present disclosure. Figure 3 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0173] like Figure 3As shown, an electronic device may include a processor (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes based on a program stored in read-only memory (ROM) or a program loaded from a storage device into random access memory (RAM). The RAM also stores various programs and data required for the operation of the electronic device. The processor, ROM, and RAM are interconnected via a bus. Input / output (I / O) interfaces are also connected to the bus.

[0174] Typically, the following devices can be connected to the I / O interface: input devices, such as sensors or visual information acquisition devices; output devices, such as displays; storage devices, such as magnetic tapes or hard drives; and communication devices. Communication devices allow electronic devices to communicate wirelessly or wiredly with other devices (such as edge computing devices) to exchange data. Although Figure 3 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or have alternatively.

[0175] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device, or installed from a ROM. When the computer program is executed by a processor, all or part of the steps of the multimodal information retrieval and analysis method for real-time Internet information streams according to embodiments of this disclosure are performed.

[0176] For a detailed description of this embodiment, please refer to the corresponding descriptions in the foregoing embodiments, which will not be repeated here.

[0177] A computer-readable storage medium according to embodiments of the present disclosure stores non-transitory computer-readable instructions. When these non-transitory computer-readable instructions are executed by a processor, all or part of the steps of the multimodal information retrieval and analysis method for real-time Internet information streams described in the foregoing embodiments of the present disclosure are performed.

[0178] The aforementioned computer-readable storage media include, but are not limited to: optical storage media (e.g., CD-ROM and DVD), magneto-optical storage media (e.g., MO), magnetic storage media (e.g., magnetic tape or portable hard drive), media with built-in rewritable non-volatile memory (e.g., memory card), and media with built-in ROM (e.g., ROM cartridge).

[0179] For a detailed description of this embodiment, please refer to the corresponding descriptions in the foregoing embodiments, which will not be repeated here.

[0180] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of this disclosure to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations therein.

Claims

1. A method for multi-modal information retrieval and analysis for internet real-time information streams, characterized in that, The method comprises the following steps: acquiring multi-modal data in the Internet, including text, images, videos and audio, and timestamping each modal data to form a multi-modal data set; based on the multi-modal data set, extracting semantic features of each modal by a specific feature extractor, including semantic vectors of text, visual features of images, spatio-temporal features of videos and audio features of audio; using a contrast learning method to construct a cross-modal joint embedding space, using a multi-head attention mechanism to weight and fuse the features of each modal, and performing cross-modal alignment through similarity calculation; based on the cross-modal alignment result, constructing a real-time multi-modal retrieval engine, reordering the retrieval results through a generative model, and generating the final retrieval results; based on the retrieval results, performing event detection, sentiment analysis and behavior understanding tasks to identify and analyze key events and trends in the Internet information flow; optimizing the computing performance through model quantization and knowledge distillation technology, and deploying on a heterogeneous computing platform to ensure real-time performance and computing efficiency.

2. The method for multi-modal information retrieval and analysis of internet oriented real-time information streams as claimed in claim 1 wherein: based on the multi-modal data set, the semantic features of each modal are extracted by a specific feature extractor, wherein: the feature extractor for text modal is a pre-trained language model, and the pre-trained language model includes BERT or RoBERTa, which is used to generate context-sensitive semantic vectors; the feature extractor for image modal is a convolutional neural network or a visual Transformer, the convolutional neural network is ResNet-152, and the visual Transformer is SwinTransformer, which is used to extract visual features and generate global semantic representation; the feature extractor for video modal is a video and text association model, and the video and text association model is CLIP4Clip, which is used to capture the spatio-temporal semantic features of video frames; the feature extractor for audio modal is a model with attention mechanism, and the model with attention mechanism is TAP-PMR model, which is used to focus the text on the audio features of the relevant audio frames.

3. The method for multi-modal information retrieval and analysis of internet oriented real-time information streams as claimed in claim 1 wherein: based on the cross-modal alignment result, a real-time multi-modal retrieval engine is constructed, including: performing probability cross-modal embedding on each modal data, using a PCME model to map the features of different modalities to a unified semantic space; using a graph neural network to establish the time sequence reasoning relationship of cross-modal data to process the spatio-temporal dependency relationship between different modalities; using a multi-head attention mechanism to weight and fuse the features of each modal, and the weight fusion is performed through the following mathematical expression: H = Concat(h1, h2,..., h k )W O wherein: W O is a learnable weight matrix, and k is the number of attention heads.

4. The method for multi-modal information retrieval and analysis of internet oriented real-time information streams as claimed in claim 1 wherein: the multi-modal retrieval engine includes a three-stage retrieval pipeline: the first stage is based on a cross-modal pre-training model for semantic matching, and the cross-modal pre-training model is CLIP or ImageBind; the second stage improves the cross-modal alignment accuracy through a reorderer combined with contrast learning and Token-level interaction, and the reorderer adopts a reorder strategy combined with a COTS model; the third stage uses a generative retrieval model to generate discrete visual Tokens, and the generative retrieval model is IRGen, which is used to realize end-to-end differentiable search. The multi-modal features are stored in a vector database, which is Milvus, for fast retrieval.

5. The method for multi-modal information retrieval and analytics for internet of real-time information streams as claimed in claim 1, wherein: The event detection task includes: The spatio-temporal event modeling is performed by a recurrent neural network combined with a spatio-temporal continuity constraint, which is an LSTM network; The sentiment analysis task adopts a multi-modal Transformer model, which is a ViLT model, for fusing text sentiment polarity and image sentiment labels; The behavior understanding task includes multi-modal composite retrieval, which locates a specific target through composite retrieval of text and video, and analyzes the target motion trajectory combined with sensor data, which is IMU data.

6. The method for multi-modal information retrieval and analytics for internet-of-things real-time information streams as claimed in claim 1, wherein: The computing optimization techniques include: Model quantization is used to compress model weights from FP32 to INT8 to reduce computing resource consumption; The knowledge distillation technique adopts a DGR model to perform model compression through a distillation loss formula, and the mathematical expression of the distillation loss is: L distill = KL(p teacher (x)||p student (x)) where: p teacher (x) is the teacher model output probability distribution, p student (x) is the student model output probability distribution.

7. The method for multi-modal information retrieval and analytics for internet-of-things real-time information streams as claimed in claim 1, wherein: The hardware of the heterogeneous computing platform includes GPU and FPGA, where CNN feature extraction is deployed on GPU and point cloud processing is deployed on FPGA. The heterogeneous computing platform deploys different modal feature processing tasks on different hardware to optimize the utilization of computing resources.

8. A multi-modal information retrieval and analysis system for Internet oriented real-time information streams, characterized in that, It includes: A multi-modal data acquisition unit is configured to acquire multi-modal data including text, images, videos, and audio from the Internet, and to timestamp the modal data to form a multi-modal data set; A feature extraction unit is configured to extract semantic features of each modality based on the multi-modal data set through different feature extractors; A cross-modal alignment unit is configured to construct a cross-modal joint embedding space through a contrastive learning method, and to weight and fuse each modal feature using a multi-head attention mechanism, and to align different modalities through similarity calculation; A multi-modal retrieval engine is configured to generate real-time retrieval results based on the cross-modal alignment results, reorder the retrieval results, and generate the final retrieval output; An event analysis unit is configured to perform event detection, sentiment analysis, and behavior understanding tasks based on the retrieval results to identify and analyze key events, sentiment changes, and behavior trends in Internet information streams; A computing optimization unit is configured to optimize computing performance through model quantization, knowledge distillation techniques, and a heterogeneous computing platform to ensure real-time and high efficiency of the system.

9. An electronic device, comprising: The electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the multi-modal information retrieval and analysis method for Internet real-time information streams according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer instructions for causing a computer to perform the multi-modal information retrieval and analysis method for Internet real-time information streams according to any one of claims 1-7.