Entity mapping-oriented multi-modal data resource management equipment

By parsing and extracting features from multimodal data, generating encoded data packets and performing semantic transformation, the problem of fusion and consistency maintenance in existing multimodal data management systems when entities change is solved, achieving efficient data resource management and retrieval, and improving the value and reliability of data assets.

CN121277976APending Publication Date: 2026-01-06SHANGHAI YUANFANG TIANSHU INTEGRATED EQUIPMENT CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511441185.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-10
Publication Date
2026-01-06

AI Technical Summary

Technical Problem

Existing multimodal data management systems struggle to achieve deep integration and consistent maintenance at the entity level when faced with entity changes, resulting in the loss or disconnection of historical data, the inability to construct a complete entity historical timeline, and the lack of explicit modeling of entity evolution behavior, which affects the value mining and reliability of data assets.

Method used

By performing multimodal parsing and feature extraction on the raw data units uploaded by users, an encoded data package containing data fragment IDs, entity IDs, image vectors, audio vectors, text vectors, and keywords is generated. Semantic conversion is then performed using the image and audio encoded copies to form an indexed document, supporting efficient retrieval and entity-level association.

Benefits of technology

It enables semantic management and efficient retrieval of multimodal data, enhances the asset value and utilization efficiency of data resources, solves the problem of deep integration and consistency maintenance at the entity level, and ensures the integrity and traceability of data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121277976A_ABST
    Figure CN121277976A_ABST
Patent Text Reader

Abstract

The invention discloses entity mapping-oriented multi-modal data resource management equipment, which is characterized in that fine multi-modal analysis and feature extraction are carried out on an original data unit uploaded by a user to generate an encoded data packet containing a data fragment ID (Identity), an entity ID, an image vector, an audio vector, a text vector and a keyword; therefore, a foundation is laid for unified management of different modal data; furthermore, semantic transformation is carried out on the multi-modal vector in the coded data packet by utilizing the image coding book and the audio coding book, an index document is formed, and an index system is updated to support subsequent efficient retrieval, when a user submits a complex query, the system can carry out two-stage mixed query execution based on the updated index system, and the user experience is improved. And the sorted result set can be quickly and accurately obtained. In this way, semantic management, entity-level association and efficient retrieval of the multi-modal data are achieved, and the capitalization value and the utilization efficiency of data resources are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent management, and more specifically, to a multimodal data resource management device oriented towards entity mapping. Background Technology

[0002] With the rapid development of artificial intelligence, big data, and knowledge graph technologies, the demand for cross-modal data fusion and management is growing across various industries. In many fields such as smart cities, financial risk control, cultural tourism, and the industrial internet, data often coexists in multiple modalities, including text, images, audio, and video. How to effectively identify, manage, and assetize the same entity (such as a specific person, vehicle, or product) described in different modalities has become a core challenge in current multimodal data applications. Traditional single-modal or weakly correlated data management systems often struggle to achieve deep entity-level fusion and consistency maintenance when faced with dynamic evolutions in entity lifecycle management, such as changes in customer information, organizational restructuring, and mergers and acquisitions.

[0003] Currently, multimodal data management mainly adopts methods such as modality-independent storage, feature-aligned fusion, and knowledge graph-based association. However, these existing solutions generally suffer from significant technical flaws. Traditional systems tend to assign static IDs to each entity. When an entity undergoes changes such as a company merger, the system faces a dilemma: directly overwriting old records will result in the loss of important historical data and relationships; creating new IDs will disconnect historical data from the new entity, making complete historical tracing difficult. Simultaneously, an entity's multimodal data is often scattered across different systems, leading to asynchronous update operations during entity state changes and a lack of transactional guarantees, easily causing data inconsistencies. Existing systems also lack explicit modeling of the entity's evolutionary behavior itself, only recording the current state, making it impossible to construct a complete entity historical timeline. These problems collectively lead to extreme difficulties in historical data tracing, severely impacting the value mining and reliability of data assets.

[0004] Therefore, there is a need for an optimized multimodal data resource management device oriented towards entity mapping. Summary of the Invention

[0005] To address the aforementioned technical problems, this application is proposed. Embodiments of this application provide a multimodal data resource management device oriented towards entity mapping. This device performs refined multimodal parsing and feature extraction on user-uploaded raw data units, generating encoded data packets containing data fragment IDs, entity IDs, image vectors, audio vectors, text vectors, and keywords, thus laying the foundation for unified management of different modalities. Furthermore, it utilizes image and audio encoding copies to perform semantic transformation on the multimodal vectors in the encoded data packets, forming index documents and updating the index system to support subsequent efficient retrieval. When a user submits a complex query, the system can perform a two-stage hybrid query execution based on the updated index system, quickly and accurately obtaining the sorted result set. In this way, semantic management, entity-level association, and efficient retrieval of multimodal data are achieved, significantly improving the asset value and utilization efficiency of data resources.

[0006] According to one aspect of this application, a multimodal data resource management device for entity mapping is provided, comprising: The raw data unit acquisition module is used to acquire raw data units uploaded by users. The multimodal parsing and multimodal feature extraction module is used to perform multimodal parsing and multimodal feature extraction on the original data units to obtain encoded data packets. The encoded data packets include data fragment IDs, entity IDs, image vectors, audio vectors, text vectors, and keywords. The semantic conversion module is used to perform semantic conversion on the image vectors and audio vectors in the encoded data packets based on the image encoding book and the audio encoding book to obtain the index document, and to update the index system based on the index document to obtain the updated index system; The complex query retrieval module is used to retrieve complex queries submitted by users. The two-stage hybrid query execution module is used to perform two-stage hybrid query execution on complex queries based on an updated index system to obtain a sorted result set.

[0007] Compared with existing technologies, this application provides a multimodal data resource management device oriented towards entity mapping. It performs refined multimodal parsing and feature extraction on user-uploaded raw data units to generate encoded data packets containing data fragment IDs, entity IDs, image vectors, audio vectors, text vectors, and keywords, thus laying the foundation for unified management of different modalities. Furthermore, it uses image and audio encoding books to perform semantic transformation on the multimodal vectors in the encoded data packets, forming index documents and updating the index system to support subsequent efficient retrieval. When a user submits a complex query, the system can perform a two-stage hybrid query execution based on the updated index system, quickly and accurately obtaining the sorted result set. In this way, semantic management, entity-level association, and efficient retrieval of multimodal data are achieved, significantly improving the asset value and utilization efficiency of data resources. Attached Figure Description

[0008] The above and other objects, features, and advantages of this application will become more apparent from the more detailed description of the embodiments of this application in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this application and form part of the specification. They are used together with the embodiments of this application to explain this application and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.

[0009] Figure 1 This is a block diagram of a multimodal data resource management device for entity mapping according to an embodiment of this application; Figure 2 This is a schematic diagram of data flow in a multimodal data resource management device for entity mapping according to an embodiment of this application. Detailed Implementation

[0010] Hereinafter, exemplary embodiments according to this application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this application, and not all embodiments of this application. It should be understood that this application is not limited to the exemplary embodiments described herein.

[0011] As indicated in this application and claims, unless the context clearly indicates otherwise, the words "a," "an," "an," and / or "the" are not specifically singular and may include plural forms. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of explicitly identified steps and elements, which do not constitute an exclusive list, and the method or apparatus may also include other steps or elements.

[0012] While this application makes various references to certain modules of the systems according to embodiments of this application, any number of different modules can be used and run on user terminals and / or servers. The modules described are merely illustrative, and different aspects of the systems and methods may use different modules.

[0013] Flowcharts are used in this application to illustrate the operations performed by the system according to embodiments of this application. It should be understood that the preceding or following operations are not necessarily performed in exact order. Instead, various steps can be processed in reverse order or simultaneously as needed. Furthermore, other operations can be added to these processes, or one or more steps can be removed from them.

[0014] Hereinafter, exemplary embodiments according to this application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this application, and not all embodiments of this application. It should be understood that this application is not limited to the exemplary embodiments described herein.

[0015] The technical solution of this application proposes a multimodal data resource management device oriented towards entity mapping. Figure 1 This is a block diagram of a multimodal data resource management device for entity mapping according to an embodiment of this application. Figure 2 This is a system architecture diagram of a multimodal data resource management device with entity mapping according to an embodiment of this application. Figure 1 and Figure 2 As shown, the entity-mapping-oriented multimodal data resource management device 300 according to an embodiment of this application includes: a raw data unit acquisition module 310, used to acquire raw data units uploaded by users; a multimodal parsing and multimodal feature extraction module 320, used to perform multimodal parsing and multimodal feature extraction on the raw data units to obtain encoded data packets, the encoded data packets including data fragment IDs, entity IDs, image vectors, audio vectors, text vectors, and keywords; a semantic conversion module 330, used to perform semantic conversion on the image vectors and audio vectors in the encoded data packets based on the image encoding and audio encoding to obtain index documents, and to update the index system based on the index documents to obtain an updated index system; a complex query acquisition module 340, used to acquire complex queries submitted by users; and a two-stage hybrid query execution module 350, used to perform two-stage hybrid query execution on the complex queries based on the updated index system to obtain a sorted result set.

[0016] Specifically, the raw data unit acquisition module 310 is used to acquire raw data units uploaded by the user. Raw data units refer to user-uploaded data in its original form, without undergoing multimodal analysis and feature extraction by the system. It encompasses unstructured data across multiple modalities, including but not limited to text, images, videos, audio, and raw data streams from various sensors. Specifically, for video raw data units, it can be understood as a continuous sequence of image frames and synchronized audio signals; for text raw data units, it is the unprocessed raw text content; and for audio data, it is the raw audio waveform. These raw data units are the foundational input for subsequent multimodal analysis, feature extraction, and entity mapping. In practice, the user uploads raw, unstructured data (such as video files, audio files, text documents, etc.) to the device. Upon acquisition, the system immediately performs preliminary preprocessing to form a complete chain for the effective acquisition and preparation of raw data units.

[0017] Specifically, the multimodal parsing and multimodal feature extraction module 320 is used to perform multimodal parsing and multimodal feature extraction on the original data units to obtain encoded data packets. The encoded data packets include data fragment IDs, entity IDs, image vectors, audio vectors, text vectors, and keywords. It should be understood that existing technologies have significant shortcomings in multimodal data management, making it difficult to meet the growing demand for cross-modal data fusion and management. Specifically, current multimodal data management methods are mostly based on independent modal storage, feature alignment-based fusion, or knowledge graph-based association. However, these methods generally suffer from problems such as a lack of unified entity IDs, fragmented management, insufficient ownership confirmation, difficulty in traceability, and poor scalability. For example, most existing solutions only focus on modal features, lacking a unified cross-modal entity identifier, resulting in inaccurate correspondence between objects in different modalities, leading to information redundancy and management dispersion, making efficient access and transactions difficult. Furthermore, existing solutions lack perfect entity-level data ownership confirmation, failing to bind data sources, generation times, and usage permissions, making it difficult to achieve full-chain traceability during data application and transactions. Therefore, the technical solution of this application achieves unified identification, ownership confirmation, storage, retrieval, and circulation of cross-modal data by performing multimodal parsing and multimodal feature extraction on the original data units, thereby enhancing the asset value and circulation efficiency of data resources. Specifically, by transforming the original data into encoded data packets containing data fragment IDs, entity IDs, image vectors, audio vectors, text vectors, and keywords, the foundation is laid for achieving beneficial effects such as unified identification, reliable ownership confirmation, efficient management, convenient circulation, and strong scalability. This eliminates information fragmentation, ensures data immutability and traceability, and supports efficient storage, retrieval, and access.

[0018] In practice, the original data units are first segmented to obtain a set of segments. It should be understood that existing temporal data segmentation methods based on low-order visual features, such as techniques using color histogram differences to determine shot boundaries, rely entirely on the statistical distribution of pixels and therefore cannot perceive the true semantics of the image content. This leads to a large number of meaningless short segments when processing scenes with coherent narratives but frequent visual changes (such as shot-reverse-shot sequences), disrupting the scene's integrity. Simultaneously, this method is extremely insensitive to gradual transitions (such as cross-dissolve), and the slight differences between adjacent frames make it difficult to reach the trigger threshold, resulting in missed detection of key shot boundaries. Furthermore, dynamic noise such as rapid camera movement and sudden changes in lighting can cause drastic fluctuations in color differences between frames, causing the system to misjudge non-scene transition points as boundaries, resulting in numerous erroneous segments. Its fundamental flaw lies in treating video as an isolated stream of visual information, completely ignoring the contextual information contained in other modalities such as audio that provides crucial evidence for scene transitions, making its decision-making process one-sided and easily influenced.

[0019] To overcome the aforementioned technical deficiencies, this application proposes an adaptive scene boundary detection mechanism based on a multimodal semantic manifold in a preferred example. This mechanism segments the original data units, achieving high-precision, semantically aware adaptive segmentation of video temporal data. It accurately identifies various scene boundaries, including hard cuts and gradual transitions, while exhibiting strong robustness to dynamic noise such as camera motion and lighting changes. The generated video segments are semantically coherent and complete, reflecting the narrative structure of the video content. This provides high-quality basic data units for subsequent multimodal data resource management and is a key prerequisite for achieving efficient indexing, accurate retrieval, and deep content understanding.

[0020] In this process, firstly, the raw data units are extracted frame by frame to obtain a frame sequence, and then audio waveforms are extracted from the raw data units. This step forms the basis for data preprocessing, ensuring the multimodal input required for subsequent semantic analysis; Next, multimodal semantic trajectory parameterization is performed on the frame sequence and audio waveform to obtain visual and auditory semantic trajectories. That is, by parameterizing the frame sequence and audio waveform using multimodal semantic trajectories, the raw, unstructured video pixels and audio waveforms are transformed into semantic representations that can be rigorously analyzed mathematically. Specifically, for the input continuous image frame sequence and synchronized audio signal, a pre-trained deep encoder model is used to map the visual and auditory content at each time point into a high-dimensional semantic embedding vector. Thus, the evolution of the video over time is constructed as two parameterized trajectories on their respective semantic manifolds. This process can be expressed by the following formula:

[0021] in, Representing the visual semantic trajectory, it is a vector function that maps time t to the visual semantic space; Representing the auditory semantic trajectory, it is a vector function that maps time t to the auditory semantic space; and These are depth encoders for processing images and audio, respectively. It is the video frame corresponding to time point t; It is a short audio segment centered on t; and These are the semantic manifold spaces of vision and hearing, respectively; thus, the transformation from raw data to structured semantic trajectories is completed. In other words, changes in video content are intuitively expressed as morphological changes in spatial trajectories; smooth content results in a gentle trajectory, while abrupt changes in content result in a turning point in the trajectory. This lays the foundation for subsequent analysis based on differential geometry and fundamentally solves the problem of insensitivity to content semantics. Furthermore, a cross-modal semantic trajectory discontinuity assessment is performed on the visual and auditory semantic trajectories to obtain a discontinuity signal. That is, by assessing the discontinuity of the visual and auditory semantic trajectories across modalities, an index that simultaneously measures hard cuts and gradual transitions and effectively fuses multimodal information is used to quantify the intensity of semantic changes. Specifically, drawing on the concept of the rate of change of a curve in differential geometry, the energy density of the two semantic trajectories at each time point is calculated; this value approximates the square of the magnitude of the trajectory tangent vector. This essentially calculates the first derivative of the semantic vector with time, thereby capturing the instantaneous speed of semantic change. To achieve effective fusion of cross-modal information and suppress single-modal noise, a geometric averaging method is used to combine the visual and auditory energy densities into a unified discontinuity signal. Specifically, the cross-modal semantic trajectory discontinuity assessment of the visual and auditory semantic trajectories is performed using the following formula:

[0022] in, and These are the energy densities of the visual and auditory trajectories at time point t, respectively. This indicates taking the derivative with respect to time; The norm of a vector; This is the final synthesized discontinuous signal; thus, a time-series signal that accurately reflects the degree of joint change across modal semantics is generated. Whether it's a hard cut causing instantaneous trajectory jumps or a gradually changing transition causing continuous trajectory curvature, it can be... The signal is characterized by significant peaks or bulges, and the signal response is strongest only when the audiovisual semantics change significantly at the same time, thus effectively avoiding interference from single-mode noise. Then, adaptive threshold peak detection and boundary determination are performed on the discontinuous signal to obtain a set of boundary timestamps. It should be understood that different videos or different segments of the same video have vastly different content rhythms and switching styles, and a fixed judgment threshold cannot adapt to such variations, leading to unstable detection performance. Therefore, in the preferred example of this application, adaptive threshold peak detection and boundary determination are further performed on the discontinuous signal. That is, the E(t) signal is treated as a random process, and its statistical characteristics are dynamically estimated within a sliding time window. An adaptive dynamic threshold δ(t) is calculated based on the mean and standard deviation of the signal within this window. Subsequently, all points on the E(t) signal that exceed the dynamic threshold and are local maxima are found, and the timestamps of these points are determined as the final scene boundaries; this process can be expressed by the formula:

[0023] in, It is the dynamic threshold at time point t; and These are sliding windows Inside The mean and standard deviation of the signal; It is an adjustable weighting coefficient; It is the final set of boundary timestamps; It is a candidate boundary point; It is a minimal time interval used to determine the local minimum neighborhood; thus, it is possible to accurately and robustly extract the true scene transition points from discontinuous signals. It is worth mentioning that the detection sensitivity can self-adjust according to the real-time dynamics of the video content, maintaining high sensitivity in calm narratives and increasing tolerance in intense scenes, thereby achieving high-accuracy segmentation results in various complex video scenarios. Furthermore, based on the boundary timestamp set, the original data units are segmented to obtain the fragment set. The boundary timestamp set refers to a series of time points determined through adaptive threshold peak detection; these time points precisely identify the location where scene or semantic transformation occurs within the original data units. In other words, in the technical solution of this application, the scene boundaries detected in the preceding steps are applied to the original data, thereby accurately segmenting long video or audio streams into short fragments with independent semantics. Next, keyword extraction and text encoding are performed on the text data of each segment in the fragment set to obtain text vectors and keywords. Keyword extraction extracts concise, high-information-density core concepts from lengthy or complex text, providing direct index terms for traditional retrieval mechanisms such as inverted indexes, thus improving retrieval efficiency. Text encoding maps unstructured text data to a continuous semantic space. In this space, semantically similar text segments will have similar vector representations, which is crucial for supporting complex semantic queries and calculating text similarity. In this way, text data can effectively participate in the entity recognition and mapping process, helping the system identify objects in different modalities and mapping the data of the same object in different modalities to a unique entity ID. This overcomes the problem of lacking a unified cross-modal entity identifier in existing technologies, achieving unified identification and efficient management of data resources.

[0024] In this process, firstly, the most representative and concise words or phrases are automatically identified and extracted from the text data to reflect the core content or theme of the text fragment. This is typically achieved using Natural Language Processing (NLP) techniques. Specifically, various algorithms are employed for keyword extraction. For example, statistical methods such as the TF-IDF model assess word importance by calculating the frequency of a word's occurrence in the current text and its inverse document frequency in the entire corpus. Words with higher importance are more likely to be selected as keywords. Alternatively, graph-based ranking algorithms, such as TextRank, can be used to construct a graph structure from the words in the text, with co-occurrence relationships between words forming edges. Then, the PageRank algorithm is used to rank the words, and words with high scores are selected as keywords. Simultaneously, pre-trained deep learning models can be used. These models learn from the contextual information of the text, enabling them to more accurately understand the semantic contribution of words, thereby generating or extracting high-quality keywords to obtain a keyword list. Keywords are a set of words or phrases obtained after the keyword extraction process, which concisely and clearly reflect the main content or core theme of the text fragment. Secondly, text data is converted into high-dimensional vectors in numerical form, i.e., text vectors. This numerical representation can capture the semantic information and contextual relationships of the text, enabling computers to perform mathematical operations and analyses, thereby achieving functions such as semantic similarity comparison. Specifically, text embedding techniques, such as Word2Vec, GloVe, or FastText, can be used to map individual words to a low-dimensional dense vector space. For a text segment, the overall text vector of the segment can be obtained by averaging, weighted averaging, or using sequence models (such as RNN, LSTM) on the embedding vectors of all words. Alternatively, the entire sentence or document can be directly encoded into a vector, for example, using pre-trained models based on the Transformer architecture, such as BERT and RoBERTa. These models can capture long-distance dependencies and deep semantics between words, generating text vectors containing rich contextual information. Typically, a fixed-length text vector representing the entire text segment can be obtained through a specific output layer of the model (such as the output vector of the [CLS] token) or by pooling all token output vectors. The text vector is a key numerical representation for semantic matching, content classification, and entity recognition and mapping. Furthermore, keyframe images in each segment of the fragment set are visually encoded to obtain image vectors. It should be understood that raw pixel information cannot be directly understood by machines for its semantic content, nor can it be efficiently compared for similarity or associated across modalities. Traditional image processing methods often rely on low-order visual features, which have limited capabilities for object recognition and semantic understanding in complex scenes, thus restricting the depth and breadth of data assetization. In the technical solution of this application, by visually encoding keyframe images, the raw pixel data of the image can be transformed into an abstract, high-dimensional numerical form. This transformation allows the image data to depart from its original form and enter a computable semantic space. The generated image vectors provide a strong visual basis for entity recognition and mapping.

[0025] In this process, firstly, representative keyframe images are identified and extracted from each segment in the fragment set. Then, for each extracted keyframe image, the system uses image recognition technology to perform visual encoding processing. Specifically, firstly, it should be understood that the original keyframe images may have different resolutions, sizes, and color spaces. To adapt to the requirements of the visual encoding model, the images need to undergo standardization preprocessing, such as uniformly adjusting the image size (e.g., scaling to 224x224 pixels), normalizing pixel values ​​(e.g., scaling pixel values ​​to between 0 and 1), and possibly performing color space conversion. Next, the preprocessed keyframe images are input into a pre-trained deep encoder model. The deep encoder model is typically built on advanced architectures such as Convolutional Neural Networks (CNN) or Vision Transformer (ViT). These models have been trained on massive amounts of image data, learning rich visual features from low-level image textures and edges to high-level semantic concepts. During this process, the deep encoder model extracts features from the input image layer by layer through its multi-layer network structure. Shallow networks typically extract low-level features, such as edges and corners; deep networks extract more abstract, semantically meaningful high-level features, such as object shape, texture, and even contextual information of the entire scene. Then, at a specific layer of the deep encoder model (usually before a fully connected layer at the end of the model or after a global average pooling layer), a fixed-length numerical array, i.e., an image vector, is output. Through this visual encoding process, each keyframe image is successfully transformed into a unique image vector, which can be efficiently stored, compared, and analyzed by the computer, thus providing powerful numerical support for subsequent operations such as entity recognition, cross-modal association, and retrieval. Subsequently, auditory encoding is performed on the audio data of each segment in the fragment set to obtain audio vectors. It should be understood that the original audio waveform is a continuous analog signal (or its digital representation), from which machines struggle to directly extract high-level semantic information, nor can they directly perform effective similarity comparisons or fuse with other modal data. Traditional audio analysis methods may remain at the acoustic feature level, lacking the ability to understand the deep semantic content. In the technical solution of this application, auditory encoding is used to convert the original audio signal into an abstract, high-dimensional numerical form, enabling unstructured audio data to enter a computable semantic space, providing important auditory evidence for entity recognition and mapping.

[0026] In this process, firstly, the corresponding audio data is obtained from each segment in the segment set. This audio data may be an audio stream extracted from the original video or an independent audio file. Subsequently, for each obtained audio data, the system applies audio processing techniques to perform auditory encoding. Specifically, firstly, it should be understood that the original audio data may have different sampling rates, bit depths, and number of channels. To adapt to the requirements of the auditory encoding model, the audio needs to be standardized preprocessed. This typically includes: resampling: unifying all audio to the same sampling rate to ensure input consistency; normalization: adjusting the audio volume so that its peak is within a certain range to avoid overload or weak signal; noise reduction: if the original audio contains background noise, noise reduction algorithms can be applied to improve signal quality. Next, the preprocessed audio data is input into a pre-trained deep encoder model. The deep encoder model is usually built based on convolutional neural networks (CNN), recurrent neural networks (RNN, such as LSTM / GRU), or Transformer architectures. These models have been trained on a large amount of audio data and have learned rich auditory features from low-level acoustic properties to high-level semantic concepts. In this process, the deep encoder model extracts features from the input audio layer by layer through its multi-layered network structure. Shallow networks may capture basic acoustic features such as pitch, loudness, and timbre; deep networks extract more abstract and semantically meaningful high-level features, such as the semantic content of the speech, speaker identity, musical genre, or specific environmental sounds (such as vehicle sounds or animal calls). Subsequently, a fixed-length numerical array, i.e., an audio vector, is output from a specific layer of the deep encoder model. This vector is a compact, high-dimensional semantic representation of the original audio segment, capturing both the acoustic features and semantic content of the audio. Through the above auditory encoding process, each audio segment is successfully transformed into a unique audio vector, which can be efficiently stored, compared, and analyzed by the computer, thus providing powerful numerical support for subsequent operations such as entity recognition, cross-modal association, and retrieval.

[0027] Specifically, the semantic conversion module 330 is used to perform semantic conversion on image vectors and audio vectors in the encoded data packet based on the image encoding book and audio encoding book to obtain an index document, and to update the index system based on the index document to obtain an updated index system. It should be understood that although the original image vectors and audio vectors can capture the deep semantic features of multimodal data, they are continuous high-dimensional numerical representations, and directly using them for large-scale indexing and retrieval presents challenges in terms of efficiency and interpretability. For example, in traditional relational databases or simple vector databases, it is difficult to directly associate high-dimensional vectors with semantically understandable concepts. In the technical solution of this application, by introducing image encoding books and audio encoding books and performing semantic conversion on image vectors and audio vectors, the continuous vector space can be quantized into discrete, symbolic visual and audio codes. Discrete codes are easier to construct based on traditional index structures such as inverted indexes, accelerating retrieval speed; in addition, these codes can be regarded as high-level semantic tags or concepts, thereby enhancing the interpretability of the data.

[0028] In practice, firstly, semantic transformation is performed on the image vectors and audio vectors in the encoded data packet based on the image encoding book and audio encoding book to obtain visual and audio codes. During this process, for each image vector in the encoded data packet, the system calculates the Euclidean distance between that image vector and each image center vector in the image encoding book. The image encoding book is a set of representative image feature vectors (i.e., image center vectors) obtained through pre-training or clustering, where each center vector represents a visual concept or pattern. Euclidean distance measures the similarity between two vectors in a multidimensional space; a smaller distance indicates higher similarity. Subsequently, the index of the image center vector corresponding to the minimum Euclidean distance is used as the visual code of that image vector. This means that each image vector is quantized into a discrete identifier of the visual concept most similar to it in the encoding book; similarly, for each audio vector in the encoded data packet, the system calculates the Euclidean distance between the audio vector and each audio center vector in the audio encoding book; the audio encoding book is another set of representative audio feature vectors (i.e., audio center vectors) obtained through pre-training or clustering, each center vector representing an auditory concept or pattern; subsequently, the index of the audio center vector corresponding to the minimum Euclidean distance is used as the visual encoding of the audio vector. This process is also a quantization process, mapping continuous audio vectors to discrete auditory concept identifiers; Next, visual and audio codes are added to the encoded data packet to obtain a symbolic encoded data packet. After the symbolic transformation of image and audio vectors is completed, these newly generated discrete visual and audio codes are integrated back into the original encoded data packet. In this process, the encoded data packet will contain data fragment ID, entity ID, original image vector, original audio vector, text vector, keywords, and the newly added visual and audio codes to obtain a symbolic encoded data packet; Then, based on the symbolized encoded data packets, an index document is constructed. This index document is a data unit tailored for the indexing system, containing all the key information used for fast retrieval and filtering. Subsequently, this index document is submitted to the indexing system for processing. The indexing system parses the information in the index document, such as building inverted indexes or other efficient index structures by associating keywords, visual encodings, and audio encodings with corresponding entity IDs and data fragment IDs. By integrating these new index documents, the indexing system is updated, forming an updated indexing system capable of supporting multimodal and multidimensional retrieval.

[0029] Specifically, the complex query acquisition module 340 is used to acquire complex queries submitted by users. It should be understood that traditional data management systems are mostly based on unimodal or weakly related methods, making it difficult to achieve cross-modal fusion at the entity level. This means that if a user wants to query multimodal information related to a certain entity (such as a person, a vehicle, or a product), traditional systems often cannot provide unified, comprehensive, and accurate results. Problems such as "lack of unified entity IDs," "fragmented management," and "difficulty in traceability" in existing technologies directly lead to users being unable to obtain satisfactory results through simple unimodal queries when faced with complex information needs. To address these shortcomings, the system must be able to process and respond to complex information needs proposed by users. These needs are often not expressible by single text keywords; they may involve image features, audio cues, entity associations, or even combined queries based on specific semantic codes. Therefore, in the technical solution of this application, acquiring complex queries submitted by users allows users to express their query intentions in a more natural and richer way, thereby fully utilizing the multimodal entity mapping and refined indexes already established in the system to achieve more accurate and comprehensive information retrieval.

[0030] In practice, the system first provides a user interface that allows users to submit query elements in multiple modalities. For example: a text input box allows users to enter a list of keywords, such as "red car," "meeting minutes," or "specific person's name"; an image upload / capture function allows users to upload an image (as the source of the "query image vector") or capture a visual area from the current screen, using the image content as a query condition; an audio recording / upload function allows users to record a sound (as the source of the "query audio vector") or upload an audio file, using sound features as a query condition; and an entity ID input function allows users to directly enter a known "entity ID" to search for all multimodal data associated with that entity. Next, after users submit these multimodal inputs on the interface, the system immediately performs preliminary collection and encapsulation of these raw query elements. For example, if a user enters the text keyword "red car" and uploads an image of a red car, the system will collect both as different components of the same complex query. Subsequently, the collected raw complex query (which may contain text strings, image data, audio data, or identified entity IDs, etc.) is transmitted over the network to the system's backend processing module. During transmission, this raw data may be initially encoded into a format suitable for transmission (e.g., images are converted into Base64 encoded strings, or binary data is transmitted directly).

[0031] Specifically, the two-stage hybrid query execution module 350 is used to perform two-stage hybrid query execution on complex queries based on an updated index system to obtain a sorted result set. It should be understood that traditional data management systems are mostly based on single-modal or weakly related methods, making it difficult to achieve cross-modal fusion at the entity level. This results in complex queries submitted by users often failing to receive a unified, comprehensive, and accurate response during the generation, circulation, and application of data products. Specifically, single-modal retrieval methods cannot meet users' needs for cross-modal information. For example, users may want to query text descriptions and audio clips related to an image; furthermore, user-submitted queries may simultaneously contain text keywords, image samples, and audio samples, or directly specify entity IDs, requiring the system to comprehensively consider these multi-source information for retrieval. In the technical solution of this application, by performing two-stage hybrid query execution on complex queries, unified identification, ownership confirmation, storage, retrieval, and circulation of cross-modal data are achieved, improving the asset value and circulation efficiency of data resources. The two-stage hybrid query execution is a query processing strategy that combines coarse-grained filtering and fine-grained scoring. The first stage uses discrete encoding and inverted indexes for rapid filtering, while the second stage uses the original high-dimensional vectors for precise similarity calculation and ranking. That is, on the narrowed candidate set, precise similarity scoring based on the original high-dimensional vectors is performed to ensure the accuracy and semantic relevance of the search results, thus providing users with a ranked result set, achieving efficient indexing, accurate retrieval, and deep content understanding. This hybrid strategy effectively balances retrieval efficiency and accuracy, overcoming the shortcomings of traditional methods and single-vector search methods.

[0032] In practice, the process begins with multimodal query parsing of complex queries to obtain executable queries. These executable queries include entity IDs, keyword lists, query visual encoding, query audio encoding, query image vectors, and query audio vectors. This preprocessing stage transforms the user-submitted raw, unstructured complex queries into a standardized format that the system can understand and execute. During this process, the system first analyzes all modal elements contained in the user-submitted complex query. For example, if the user uploaded images, recorded audio, and entered text, the system identifies these different input modalities. Secondly, for image data in the query, the system performs visual encoding to obtain query image vectors and further generates query visual encoding; for audio data in the query, the system performs auditory encoding to obtain query audio vectors and further generates query audio encoding; for text input, the system extracts a keyword list; if the query explicitly specifies or can be identified through multimodal fusion, it includes entity IDs. Subsequently, all these parsed, extracted, and encoded elements are integrated into a structured executable query object, which contains all the information required for the subsequent two stages of the query. Next, a candidate set is obtained through fast filtering based on the entity ID, keyword list, query visual code, and query audio code from the executable query, using an inverted index. In this process, firstly, the system utilizes a pre-built and continuously updated indexing system containing an inverted index structure for multimodal data. Keywords, visual codes, audio codes, and entity IDs are all used as index terms, pointing to the associated data segment IDs and entity IDs. Then, the system uses the entity ID, keyword list, query visual code, and query audio code from the executable query as query conditions and matches them in the inverted index. For example, if the query visual code is "123" (representing "canines"), the indexing system will quickly return all data segment IDs containing the visual code "123". Similarly, keywords, audio codes, and entity IDs are also used for fast filtering. Subsequently, by logically combining the above matching results (such as AND and OR operations), the system can efficiently filter out a batch of candidate sets that initially match the query conditions. Furthermore, based on the query image vector and query audio vector in the executable query, the candidate set is finely scored based on precise vector similarity to obtain the sorted result set. This process aims to perform high-precision sorting of the results filtered in the first stage to ensure the accuracy and relevance of the final results. In this process, firstly, the system extracts the original image vector and audio vector corresponding to each data segment from the candidate set; then, for each data segment in the candidate set, the system calculates the similarity between its original image vector and the query image vector in the executable query, as well as the similarity between its original audio vector and the query audio vector. Commonly used similarity measurement methods include cosine similarity, the inverse of Euclidean distance, etc.; if the query contains vectors of multiple modalities (such as images and audio), the system will use a fusion strategy to perform weighted summation or more complex fusion of the similarity scores of different modalities to obtain a comprehensive relevance score; subsequently, the system sorts the candidate data segments in descending order according to the comprehensive relevance score, thereby obtaining the final sorted result set, which has been precisely matched and sorted and can most accurately reflect the user's complex query intent.

[0033] As described above, the entity-mapping-oriented multimodal data resource management device 300 according to the embodiments of this application can be implemented in various wireless terminals, such as servers with entity-mapping-oriented multimodal data resource management algorithms. In one possible implementation, the entity-mapping-oriented multimodal data resource management device 300 according to the embodiments of this application can be integrated into the wireless terminal as a software module and / or a hardware module. For example, the entity-mapping-oriented multimodal data resource management device 300 can be a software module in the operating system of the wireless terminal, or it can be an application developed for the wireless terminal; of course, the entity-mapping-oriented multimodal data resource management device 300 can also be one of many hardware modules of the wireless terminal.

[0034] Alternatively, in another example, the entity-mapping multimodal data resource management device 300 and the wireless terminal can also be separate devices, and the entity-mapping multimodal data resource management device 300 can be connected to the wireless terminal via wired and / or wireless networks, and transmit interactive information in accordance with an agreed data format.

[0035] The various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or improvement of the technology in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. An entity mapping oriented multi-modal data resource management device, characterized in that, The method comprises the following steps: An original data unit acquisition module is configured to acquire original data units uploaded by a user; A multi-modal analysis and multi-modal feature extraction module is configured to perform multi-modal analysis and multi-modal feature extraction on the original data units to obtain encoded data packets, the encoded data packets comprising data segment IDs, entity IDs, image vectors, audio vectors, text vectors, and keywords; A semantic conversion module is configured to perform semantic conversion on the image vectors and audio vectors in the encoded data packets based on an image codebook and an audio codebook to obtain index documents, and update an index system based on the index documents to obtain an updated index system; A complex query acquisition module is configured to acquire a complex query submitted by a user; A two-stage hybrid query execution module is configured to perform two-stage hybrid query execution on the complex query based on the updated index system to obtain a sorted result set.

2. The entity-mapping oriented multi-modal data resource management device according to claim 1, characterized in that, The multi-modal analysis and multi-modal feature extraction module comprises: A data segmentation unit is configured to perform data segmentation on the original data units to obtain a segment set; A keyword extraction and text encoding unit is configured to perform keyword extraction and text encoding on text data in each segment in the segment set to obtain text vectors and keywords; A visual encoding unit is configured to perform visual encoding on key frame images in each segment in the segment set to obtain image vectors; An auditory encoding unit is configured to perform auditory encoding on audio data in each segment in the segment set to obtain audio vectors.

3. The entity-mapping oriented multi-modal data resource management device according to claim 2, characterized in that, The data segmentation unit comprises: An audio waveform extraction subunit is configured to perform frame-by-frame extraction on the original data units to obtain a frame sequence, and extract audio waveforms from the original data units; A multi-modal semantic trajectory parameterization subunit is configured to perform multi-modal semantic trajectory parameterization on the frame sequence and the audio waveforms to obtain visual semantic trajectories and auditory semantic trajectories; A cross-modal semantic trajectory discontinuity evaluation subunit is configured to perform cross-modal semantic trajectory discontinuity evaluation on the visual semantic trajectories and the auditory semantic trajectories to obtain discontinuity signals; An adaptive threshold peak value detection and boundary determination subunit is configured to perform adaptive threshold peak value detection and boundary determination on the discontinuity signals to obtain a boundary timestamp set; A data unit segmentation subunit is configured to segment the original data units based on the boundary timestamp set to obtain the segment set.

4. The entity-mapping oriented multi-modal data resource management device according to claim 3, characterized in that, The cross-modal semantic trajectory discontinuity evaluation subunit is configured to perform cross-modal semantic trajectory discontinuity evaluation on the visual semantic trajectories and the auditory semantic trajectories according to the following formula: ; wherein, and are the energy densities of the visual semantic trajectory and the auditory semantic trajectory, respectively, at time point t; denotes the derivation with respect to time; denotes the norm of a vector; is the final synthesized discontinuity signal.

5. The entity-mapping oriented multi-modal data resource management device according to claim 1, characterized in that, The semantic conversion module comprises: A semantic encoding conversion unit is configured to perform semantic conversion on the image vectors and audio vectors in the encoded data packets based on an image codebook and an audio codebook to obtain visual encoding and audio encoding; A symbolization packaging unit is configured to add the visual encoding and the audio encoding to the encoded data packets to obtain symbolized encoded data packets; An index document construction unit is configured to construct index documents based on the symbolized encoded data packets.

6. The entity-mapping oriented multi-modal data resource management device according to claim 5, characterized in that, The semantic encoding conversion unit is configured to: Calculate the Euclidean distance between the image vectors and each image center vector in the image codebook; index of the image center vector corresponding to the minimum Euclidean distance as the visual encoding of the image vector.

7. The entity-mapping oriented multi-modal data resource management device according to claim 5, characterized in that, The semantic encoding conversion unit is further configured to: calculate the Euclidean distance between the audio vector and each audio center vector in the audio codebook; index of the audio center vector corresponding to the minimum Euclidean distance as the visual encoding of the audio vector.

8. The entity-mapping oriented multi-modal data resource management device according to claim 1, characterized in that, The two-stage hybrid query execution module is configured to: perform multi-modal query parsing on the complex query to obtain an executable query, the executable query comprising an entity ID, a keyword list, a query visual encoding, a query audio encoding, a query image vector, and a query audio vector; perform inverted index-based fast filtering based on the entity ID, the keyword list, the query visual encoding, and the query audio encoding in the executable query to obtain a candidate set; perform precise vector similarity-based fine-grained scoring on the candidate set based on the query image vector and the query audio vector in the executable query to obtain the ranked result set.