A knowledge graph-based multimedia data intelligent retrieval method and system
Patent Information
- Application Number
- CN202511424707.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-30
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2045-09-30
AI Technical Summary
[0004]本发明提供了一种基于知识图谱的多媒体数据智能检索方法及系统,以解决现有技术中多媒体数据检索精准度低、动态适应能力不足的问题
1.本发明借助知识图谱构建多媒体数据的结构化语义关联,突破传统关键词匹配的局限,通过实体与关系的深度建模,实现对多媒体内容语义的精准捕捉,为检索提供结构化知识支撑;
Smart Images

Figure CN121350282B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of multimedia data processing and intelligent retrieval technology, and in particular to a multimedia data intelligent retrieval method and system based on knowledge graphs. Background Technology
[0002] The accuracy and real-time performance of multimedia data retrieval directly impact information acquisition efficiency and user interaction experience. Knowledge graphs, as a core tool for structured knowledge representation, are crucial for achieving deep semantic associations of multimodal content and enhancing dynamic retrieval adaptability. They are significant for resolving information location biases in complex scenarios and maximizing the utilization value of multimedia data. Efficient knowledge graph construction and application can provide structured support for entities and relationships in multimedia retrieval, avoiding the semantic fragmentation problems of traditional retrieval methods.
[0003] In a current technology, multimedia data retrieval primarily employs a keyword-matching or simple metadata annotation approach. This involves manually annotating basic audio and video tags to build an index, combined with shallow keyword matching. However, it fails to incorporate knowledge graphs to model entity relationships and cross-modal semantics within multimedia content, storing semantic information only in a single dimension. This results in a lack of structured knowledge support for retrieval. Furthermore, this approach relies on manual annotation and shallow matching, failing to leverage knowledge graphs to uncover deep semantic relationships within multimedia content, and making it difficult to dynamically integrate newly added content based on the graph. When faced with a constant influx of new multimedia data or fuzzy user queries, the system, lacking the semantic association and dynamic update capabilities of a knowledge graph, is prone to insufficient semantic understanding and low accuracy in retrieval results, ultimately failing to meet the demands for efficient and accurate retrieval in complex scenarios. Summary of the Invention
[0004] This invention provides a multimedia data intelligent retrieval method and system based on knowledge graphs to solve the problems of low accuracy and insufficient dynamic adaptability in multimedia data retrieval in the prior art.
[0005] In a first aspect, the present invention provides a multimedia data intelligent retrieval method based on knowledge graphs, comprising: Acquire acoustic signals and video clips from multimedia data sources, extract semantic features to obtain an initial entity set and relational links, and construct a basic knowledge graph; Based on the aforementioned basic knowledge graph, node information is propagated, relation weight calculation is processed, new entity recognition and attribute association expansion are completed, and the coverage of the expanded graph is determined. Acquire new acoustic data of real-time incoming entity behavior. If the similarity with existing nodes in the expanded map coverage exceeds a preset node similarity threshold, fuse new relationship weights and adjust the propagation iteration number to obtain a dynamically adapted map structure. For the dynamically adapted graph structure, focus on key semantic paths and determine the matching degree between the fuzzy query intent and the dynamically adapted graph structure. If the matching degree is less than or equal to a preset intent matching threshold, an optimized query representation vector is obtained. Cross-modal association features are extracted from the optimized query representation vector to determine the semantic mapping fusion relationship of entity acoustic signals. The expanded map coverage is compared with the preset coverage evaluation threshold to obtain a preliminary retrieval result set. If the retrieval accuracy of the preliminary retrieval result set is lower than the preset accuracy threshold, the attention weight is adjusted to obtain the refined semantic mapping fusion relationship. Combined with the expanded map coverage and the preset coverage evaluation threshold, it is determined whether the overall map coverage meets the dynamic adaptation verification. Based on the refined semantic mapping fusion relationship, new content links are integrated and feedback weight adjustments are processed. If the new content links meet the preset link verification threshold, the final retrieval result set is obtained.
[0006] Secondly, the present invention provides a multimedia data intelligent retrieval system based on knowledge graphs, comprising: The graph construction module is used to acquire acoustic signals and video clips from multimedia data sources, extract semantic features to obtain an initial entity set and relation links, and construct a basic knowledge graph. The graph expansion module is used to propagate node information, process relation weight calculations, complete the new entity recognition and attribute association expansion based on the basic knowledge graph, and determine the coverage of the expanded graph. The dynamic update module is used to acquire new acoustic data of entity behavior that is coming in in real time. If the similarity with the existing nodes in the expanded map coverage exceeds the preset node similarity threshold, the new relation weights are fused and the number of iterations is adjusted to obtain a dynamically adapted map structure. The query optimization module is used to focus on key semantic paths and determine the matching degree between the fuzzy query intent and the dynamically adapted graph structure. If the matching degree is less than or equal to a preset intent matching threshold, an optimized query representation vector is obtained. The preliminary retrieval module is used to extract cross-modal association features from the optimized query representation vector, determine the semantic mapping fusion relationship of entity acoustic signals, and compare the expanded map coverage with the preset coverage evaluation threshold to obtain a preliminary retrieval result set. The weight adjustment module is used to adjust the attention weights to obtain a refined semantic mapping fusion relationship if the retrieval accuracy of the preliminary retrieval result set is lower than a preset accuracy threshold. Combined with the expanded map coverage and the preset coverage evaluation threshold, it determines whether the overall map coverage meets the dynamic adaptation verification. The result generation module is used to integrate new content links and process feedback weight adjustments based on the refined semantic mapping fusion relationship. If the new content links meet the preset link verification threshold, the final search result set is obtained.
[0007] Thirdly, the present invention also provides an electronic device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor executes the computer program to implement the knowledge graph-based multimedia data intelligent retrieval method described above.
[0008] Fourthly, the present invention also provides a computer-readable storage medium comprising a stored computer program, wherein, when the computer program is executed, it controls the device where the computer-readable storage medium is located to execute the knowledge graph-based intelligent multimedia data retrieval method described above.
[0009] Compared with the prior art, the present invention has the following beneficial effects: 1. This invention utilizes knowledge graphs to construct structured semantic associations for multimedia data, overcoming the limitations of traditional keyword matching. Through deep modeling of entities and relationships, it achieves accurate capture of the semantics of multimedia content, providing structured knowledge support for retrieval. 2. This invention solves the problem of insufficient dynamic adaptability of traditional retrieval systems by integrating new acoustic data and adjusting the spectrum structure in real time through a spectrum expansion and dynamic update mechanism, ensuring that the spectrum can match the real-time changes of multimedia data; 3. This invention improves the retrieval accuracy in fuzzy query scenarios by combining cross-modal feature extraction and semantic mapping fusion through a closed-loop process of "query optimization - preliminary retrieval - weight refinement". At the same time, it ensures the reliability and comprehensiveness of the retrieval results through multiple rounds of threshold verification. Attached Figure Description
[0010] Figure 1 This is a schematic diagram of a multimedia data intelligent retrieval method based on knowledge graphs provided in the first embodiment of the present invention; Figure 2 This is a schematic diagram of the structure of a multimedia data intelligent retrieval system based on knowledge graphs provided in the second embodiment of the present invention. Detailed Implementation
[0011] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0012] Reference Figure 1 The first embodiment of the present invention provides a multimedia data intelligent retrieval method based on knowledge graphs, comprising the following steps: S1. Acquire acoustic signals and video clips from multimedia data sources, extract semantic features to obtain an initial entity set and relational links, and construct a basic knowledge graph; S2, based on the basic knowledge graph, propagate node information, process relation weight calculation, complete the identification of new entities and attribute association expansion, and determine the coverage of the expanded graph; S3, acquire new acoustic data of real-time incoming entity behavior; if the similarity with existing nodes in the expanded graph coverage exceeds a preset node similarity threshold, fuse new relationship weights and adjust the propagation iteration number to obtain a dynamically adapted graph structure. S4, for the dynamically adapted graph structure, focus on key semantic paths, determine the matching degree between the fuzzy query intent and the dynamically adapted graph structure, and if the matching degree is less than or equal to the preset intent matching threshold, obtain the optimized query representation vector; S5, extract cross-modal association features from the optimized query representation vector, determine the semantic mapping fusion relationship of entity acoustic signals, and compare the expanded map coverage with the preset coverage evaluation threshold to obtain a preliminary retrieval result set; S6. If the retrieval accuracy of the preliminary retrieval result set is lower than the preset accuracy threshold, the attention weight is adjusted to obtain the refined semantic mapping fusion relationship. Combined with the expanded map coverage and the preset coverage evaluation threshold, it is determined whether the overall map coverage meets the dynamic adaptation verification. S7. Based on the refined semantic mapping fusion relationship, integrate new content links and process feedback weight adjustments. If the new content links meet the preset link verification threshold, obtain the final retrieval result set.
[0013] In step S1, acoustic signals and video clips from a multimedia data source are acquired, semantic features are extracted to obtain an initial entity set and relational links, and a basic knowledge graph is constructed, including: S11, Acoustic signals and video clips are collected from a multimedia data source, and the acoustic signals and video clips are filtered according to a preset quality filtering threshold to obtain a filtered multimedia dataset; S12, extract semantic features from the filtered multimedia dataset to obtain an initial entity set; S13, Analyze the correlation of semantic features in the initial entity set. If the similarity exceeds the preset feature correlation threshold, determine the correlation relationship and form a relation link set. S14, import the set of relationship links and the initial set of entities into a pre-constructed graph structure to generate a basic knowledge graph.
[0014] In step S11, acoustic signals and video clips are collected from a multimedia data source, and the acoustic signals and video clips are filtered according to a preset quality filtering threshold to obtain a filtered multimedia dataset.
[0015] It should be noted that the process of collecting acoustic signals and video clips from multimedia data sources and filtering them according to a preset quality screening threshold to obtain a filtered multimedia dataset means that acoustic signals and video clips are first collected from multimedia data sources such as surveillance cameras and recording equipment, and then unqualified data is removed according to the preset quality screening threshold of "acoustic signal-to-noise ratio ≥ 30 dB and video clip clarity ≥ 720p", and the data that meets the requirements is retained to form the filtered multimedia dataset.
[0016] It is worth noting that the preset quality screening threshold is set to "signal-to-noise ratio ≥ 30 dB, resolution ≥ 720p" because data below this standard will impair the accuracy of subsequent semantic feature extraction. Acoustic signals with a signal-to-noise ratio below 30 dB have too high a proportion of noise, making it impossible to clearly extract effective voiceprint features such as "vehicle horns" and "pedestrian footsteps"; video clips with a resolution below 720p lose details, making it difficult to capture the contours of actions such as "vehicles changing lanes" and "pedestrians running" through 3D convolutional networks, both of which will lead to entity recognition errors. Therefore, this threshold is used to ensure the basic quality of the data.
[0017] For example, data from the morning rush hour is collected from the city traffic monitoring system. Construction interference audio with a signal-to-noise ratio of 25 dB and equipment failure video with a resolution of 540p are removed. Finally, 5 qualified data are retained. The filtered dataset contains clearly identifiable traffic sound signals and 720p high-definition road images.
[0018] In step S12, semantic features are extracted from the filtered multimedia dataset to obtain an initial entity set.
[0019] It should be noted that the extraction of semantic features from the selected multimedia dataset to obtain the initial entity set refers to converting the acoustic signal into a Mel spectrogram and then using a convolutional neural network to extract pitch and frequency features. For video clips, a 3D convolutional network is used to extract entity contours and motion features. The two types of features are then mapped into low-dimensional vectors and clustered according to vector similarity (8 clusters) to obtain the initial entity set such as "car", "horn sound", and "pedestrian".
[0020] It is worth noting that the number of clusters is set to 8, which is based on the common semantic categories of multimedia data. It covers 3 core voiceprint categories in the acoustic field, namely "alarm sounds, vehicle sounds, and ambient sounds", and 5 core visual entities in the video field, namely "personal actions, vehicle status, road facilities, event scenes, and object types". A total of 8 categories can cover the basic semantics of most retrieval scenarios. This avoids cluster dispersion due to too many categories and entity confusion due to too few categories, ensuring that the semantic distinction of the initial entity set is clear.
[0021] For example, the filtered traffic data is processed by extracting features such as "car horn sound" and "brake sound" from the audio and "vehicle lane change" and "pedestrian crossing" from the video. After clustering, an initial entity set containing 8 types of entities such as "car", "horn sound" and "pedestrian" is obtained.
[0022] In step S13, the semantic features in the initial entity set are analyzed for correlation. If the similarity exceeds a preset feature correlation threshold, the correlation is determined and a relational link set is formed.
[0023] It should be noted that the analysis of semantic feature correlation in the initial entity set, and the determination of relationship to form a relation link set if the correlation exceeds the threshold, refers to using the cosine similarity algorithm to calculate the similarity of the semantic vectors of each entity. If the similarity is ≥0.8 (preset feature correlation threshold), it is determined that the two entities have a correlation such as "issue" or "accompany", and the relationship link set is formed.
[0024] It is worth noting that the preset feature association threshold is set to 0.8, which is based on practical experience with semantic vector similarity: this value can accurately distinguish between strong and weak associations. For example, the similarity between "car" and "horn sound" is 0.92 (above 0.8), which belongs to a strong association between an entity and a direct soundprint, and should be retained as a valid relationship; the similarity between "car" and "background wind sound" is 0.35 (below 0.8), which belongs to a weak association between unrelated entities, and should be removed to avoid redundancy, thereby ensuring the accuracy of the relationship link set.
[0025] For example, analyzing the initial entity set of the traffic scene, the similarity between "car" and "horn sound" is 0.92, and the similarity between "traffic congestion" and "horn sound" is 0.85, both exceeding 0.8. The relationship between "car horn sound" and "traffic congestion accompanied by horn sound" is determined, forming a link set containing 3 sets of relationships.
[0026] In step S14, the set of relationship links and the initial set of entities are imported into a pre-constructed graph structure to generate a basic knowledge graph.
[0027] It should be noted that the process of importing the set of relationship links and the initial set of entities into a preset graph structure to generate a basic knowledge graph refers to using graph database technology, where nodes represent initial entities and edges represent relationship links, importing into a pre-built "entity-relationship" graph structure, and generating a queryable basic knowledge graph through node association and indexing.
[0028] It is worth noting that graph databases are used to construct the knowledge graph because the graph structure can intuitively present the semantic relationships between entities. Compared with traditional relational databases, it is more suitable for the "node-edge" storage logic of knowledge graphs, which facilitates the propagation of node information and the retrieval of multi-entity association paths in subsequent graph neural networks, laying a structural foundation for the expansion of the knowledge graph.
[0029] For example, a traffic graph can be built using Neo4j, with "car" and "horn sound" as nodes and "sound" as edges. After importing entity and relation data, the generated graph can quickly query the associated entities and relationships of "car" and supports subsequent node information propagation operations.
[0030] In step S2, based on the basic knowledge graph, node information is propagated, relation weight calculations are processed, new entity identification and attribute association expansion are completed, and the coverage of the expanded graph is determined, including: S21, extract semantic features corresponding to entities and relations from the basic knowledge graph, propagate the node information corresponding to the semantic features, and obtain the updated node representation; S22, For the updated node representation, calculate the weighted relation strength to form an expanded relation link set; S23, if the relationship strength exceeds a preset relationship strength threshold, extract new entities from the expanded relationship link set, identify the category of the new entities and obtain associated attributes, and complete the attribute association expansion; S24, integrate the newly added entity, the expanded set of relationship links, and the associated attributes with the basic knowledge graph, determine the coverage of the expanded graph, and obtain the expanded knowledge graph.
[0031] In step S21, semantic features corresponding to entities and relations are extracted from the basic knowledge graph, and node information corresponding to the semantic features is propagated to obtain updated node representations.
[0032] It should be noted that the step of extracting entity and relation semantic features from the basic knowledge graph and propagating node information to obtain the updated node representation refers to extracting the semantic features of entities such as "car" and "horn sound" and relations such as "emitting", applying a graph neural network to propagate node information through weighted aggregation of neighbor node feature vectors, updating the node semantic representation, and obtaining the updated node representation containing the context.
[0033] It is worth noting that graph neural networks are used to propagate node information because the network can achieve cross-neighbor fusion of node features: for example, the "car" node can aggregate scene features from neighboring nodes such as "road" and "traffic light", so that the updated "car" node representation not only includes its own voiceprint / visual features, but also incorporates related information such as "rainy weather environment" and "intersection location", providing a more comprehensive semantic basis for subsequent relation weight calculation and avoiding the limitations of single node features.
[0034] For example, in the traffic map, the "car" node aggregates the features of "road congestion" (weight 0.6) and "rainy weather" (weight 0.4), and the updated "car" node representation can more accurately reflect the entity associations in the actual traffic scenario.
[0035] In step S22, the weighted relation strength is calculated for the updated node representation to form an expanded relation link set.
[0036] It should be noted that the calculation of weighted relation strength for the updated node representation to form an expanded relation link set refers to using the mean aggregation function to process the relation weights between nodes. The relation strength is obtained by multiplying the historical association weights by 0.4 and adding the updated node feature weights by 0.6. The sum of the two is the relation strength. All relation strengths between nodes are sorted to form an expanded relation link set that includes newly added potential associations.
[0037] It is worth noting that the weighting formula sets "historical weight 0.4 and updated feature weight 0.6" because the updated node representation has incorporated the neighbor context information and can better reflect the current entity association status. Higher weights are required to ensure that the relationship strength matches the real-time semantics. At the same time, retaining the historical weight avoids completely severing the original association rules and balances timeliness and stability.
[0038] For example, for the updated "robotic arm" and "part jamming" nodes, the historical association weight is 0.6 and the updated feature weight is 0.8. The relationship strength is calculated as 0.6×0.4+0.8×0.6=0.72. At the same time, the relationship strength between "robotic arm" and "equipment maintenance" is sorted out to form an expanded set of relationship links.
[0039] In step S23, if the relationship strength exceeds a preset relationship strength threshold, new entities are extracted from the expanded relationship link set, the category of the new entities is identified and associated attributes are obtained, and attribute association expansion is completed.
[0040] It should be noted that the phrase "extracting new entities and completing attribute association expansion when the relationship strength exceeds the preset threshold" means that the preset relationship strength threshold is 0.7. If the relationship strength in the expanded relationship link set exceeds 0.7 and the corresponding entity is not in the initial entity set, it is extracted as a new entity. The entity category is identified through semantic feature matching, and attributes such as "occurrence scenario" and "feature parameters" are associated to complete the attribute association expansion.
[0041] It is worth noting that the preset relationship strength threshold is set to 0.7 because relationships below 0.7 are relatively weak. For example, the relationship strength between "car" and "distant thunder" is 0.5, and introducing corresponding entities can easily lead to graph redundancy. Relationships above 0.7, such as the relationship strength between "door and window vibration" and "unidentified person approaching" (0.75), are strong associations. The corresponding entities can fill semantic gaps in the graph, such as "unidentified person approaching" not being in the initial set. At the same time, the associated attributes must match the entity category. For example, "safety hazard entity" is associated with "occurrence time" and "location" to ensure that the attributes and entity semantics are compatible.
[0042] For example, the relationship strength between "door and window vibration" and "unidentified person approaching" in the security map is 0.75 (exceeding 0.7). "Unidentified person approaching" is extracted as a new entity, its category is identified as "security hazard entity", and associated with the attributes of "occurrence time 22:00" and "floor 3".
[0043] In step S24, the newly added entity, the expanded set of relationship links, and the associated attributes are integrated with the basic knowledge graph to determine the coverage of the expanded graph and obtain the expanded knowledge graph.
[0044] It should be noted that the integration of newly added entities, extended relationships, and associated attributes with the basic knowledge graph, determining the coverage of the expanded graph, and obtaining the expanded knowledge graph refers to integrating the newly added entities, extended relationships, and associated attributes with the original graph, updating the node and edge structure, counting the total number of the three core semantic dimensions of the integrated graph ("entity category, attribute dimension, and relationship type"), determining the coverage of the expanded graph, and obtaining the expanded knowledge graph.
[0045] For example, the original knowledge graph contains 8 types of entities, 5 attribute dimensions, and 3 types of relationships. After integrating the newly added entities and relationships, it becomes 9 types of entities, 7 attribute dimensions, and 4 types of relationships. The scope of the expanded knowledge graph is determined to include these dimensions, and an expanded knowledge graph is obtained.
[0046] In step S3, new acoustic data of entity behavior that comes in real time is acquired. If the similarity between this data and existing nodes in the extended knowledge graph exceeds a preset node similarity threshold, new relation weights are fused and the number of propagation iterations is adjusted to obtain a dynamically adapted graph structure, including: S31, acquire new acoustic data of entity behavior in real-time data stream, extract the frequency distribution feature vector of the new acoustic data, and obtain the initial acoustic representation; S32, based on the initial acoustic representation and the existing node feature vectors in the expanded map coverage area, calculate the cosine similarity; if the similarity exceeds a preset node similarity threshold, determine the node set that needs to be updated. S33, Based on the set of nodes to be updated, calculate the new relation weights and merge the new relation weights to obtain the updated weight set; S34, adjust the propagation iteration number according to the updated weight set, recalculate the connection strength between nodes in the node set to be updated based on node similarity, update the extended knowledge graph structure, and obtain a dynamically adapted graph structure.
[0047] In step S31, new acoustic data of entity behavior in the real-time data stream is acquired, and the frequency distribution feature vector of the new acoustic data is extracted to obtain the initial acoustic representation.
[0048] It should be noted that the acquisition of new acoustic data of real-time entity behavior and extraction of frequency distribution feature vectors to obtain the initial acoustic representation refers to acquiring new acoustic data of entity behavior such as "abnormal whistling sound" and "suspicious knocking sound" from the real-time audio stream, converting the time domain signal into the frequency domain through Fourier transform, extracting the frequency distribution feature vectors, and obtaining the initial acoustic representation.
[0049] It is worth noting that Fourier transform is used to extract frequency features because time-domain acoustic signals cannot directly distinguish the essential differences between different voiceprints. For example, both "abnormal honking" and "normal honking" manifest as sound wave vibrations in the time domain, but in the frequency domain, the high-frequency range (1000-1500Hz) of "abnormal honking" has a higher proportion. Fourier transform can accurately capture this difference, providing a feature basis that reflects the essence of voiceprints for subsequent similarity calculations, thus avoiding the limitations of time-domain analysis.
[0050] For example, "continuous rapid honking" is obtained from the real-time audio stream of urban roads, and the frequency distribution feature vector is extracted by Fourier transform. The high-frequency band 1000-1500Hz accounts for 0.7 and the main peak frequency is 1200Hz, thus obtaining the initial acoustic representation.
[0051] In step S32, cosine similarity is calculated based on the initial acoustic representation and the existing node feature vectors in the expanded map coverage area. If the similarity exceeds a preset node similarity threshold, the node set that needs to be updated is determined.
[0052] It should be noted that the calculation of the cosine similarity between the initial acoustic representation and the feature vector of the extended knowledge graph node, and the determination of the node set to be updated if the similarity exceeds the threshold, refers to using the cosine similarity algorithm to calculate the similarity between the initial acoustic representation and the feature vector of existing nodes such as "traffic noise" and "abnormal alarm" in the extended graph. If the similarity is ≥0.8 (preset node similarity threshold), the node and its directly related nodes (such as the "traffic congestion" node associated with "traffic noise") are included in the node set to be updated.
[0053] It is worth noting that the preset node similarity threshold is set to 0.8 because this value can accurately filter nodes that are strongly associated with the new acoustic data: nodes with a similarity below 0.8 have weak correlation and no actual gain in the dynamic adaptation of the graph after updating; nodes with a similarity above 0.8 are highly correlated with the new data and can make the graph fit the real-time semantic changes after updating, while updating only associated nodes can reduce computing power consumption.
[0054] For example, the similarity between the initial acoustic representation of "continuous rapid honking" and the nodes of "traffic noise" and "construction noise" in the extended graph is calculated to be 0.85 and 0.55, respectively. The "traffic noise" node and the associated "traffic congestion" node are then included in the node set that needs to be updated.
[0055] In step S33, based on the set of nodes to be updated, the new relation weights are calculated and fused to obtain the updated weight set.
[0056] It should be noted that the calculation and fusion of new relation weights based on the set of nodes to be updated to obtain the updated weight set means that when calculating the new relation weights of the set of nodes to be updated, the product of the new acoustic data credibility and 70% is taken, and the product of the node historical association weight and 30% is taken, and then the two products are added together.
[0057] It is worth noting that the weighting formula sets "new data credibility 0.7, historical weight 0.3" because new acoustic data directly reflects the current real-time scene and needs to be given higher weights to ensure the timeliness of the weights; at the same time, 30% of the historical weights are retained to avoid sudden changes in weights that could lead to instability in the spectrum structure. For example, the historical correlation pattern of "traffic noise - traffic congestion" needs to be partially inherited to balance real-time performance and structural stability.
[0058] In the example, the "Traffic Noise - Traffic Congestion" node set needs to be updated. The credibility of the new acoustic data is used to quantify the reliability of the relevant new acoustic data. In the example, it is set to 0.9. The historical association weight is the weight of the previous association relationship of the node set, which is the historical basis for calculating the weight of the new relationship. In the example, it is set to 0.6. The weight of the new relationship is calculated according to the formula 0.9×0.7+0.6×0.3=0.81. The updated weight set is obtained by incremental fusion, replacing the original historical weight of 0.6.
[0059] In step S34, the propagation iteration number is adjusted according to the updated weight set, the connection strength between nodes in the node set to be updated is recalculated based on node similarity, the extended knowledge graph structure is updated, and a dynamically adapted graph structure is obtained.
[0060] It should be noted that "adjusting the propagation iteration number based on the updated weight set, recalculating the connection strength and updating the graph to obtain a dynamically adaptive graph structure" means that, based on the updated weight set, if the weight update magnitude is calculated using "new relation weight - historical association weight", in the example 0.81-0.6=0.21>0.2, which exceeds 0.2, the number of graph neural network propagation iterations is adjusted from 1 to 3. Then, based on node similarity, the connection strength between nodes in the node set to be updated is recalculated, replacing the original connection strength, updating the extended knowledge graph structure, and finally obtaining a dynamically adaptive graph structure.
[0061] It is worth noting that the number of iterations was adjusted from 1 to 3 because when the weight update magnitude exceeds 0.2, 1 iteration cannot allow nodes to fully integrate the new weight features. 3 iterations can ensure accurate calculation of connection strength through multiple rounds of neighbor feature aggregation, so that the graph structure can truly adapt to real-time data changes.
[0062] For example, the propagation iteration count is adjusted to 3 times, and the connection strengths between the "traffic noise-traffic congestion" and "traffic noise-vehicle aggregation" nodes are recalculated to 0.81 and 0.78 respectively. The extended knowledge graph structure is then updated to obtain a dynamically adapted graph incorporating real-time horn data. In step S4, for the dynamically adapted graph structure, the key semantic path is focused, and the matching degree between the fuzzy query intent and the dynamically adapted graph structure is determined. If the matching degree is less than or equal to a preset intent matching threshold, an optimized query representation vector is obtained, including: S41, Obtain the fuzzy query data input by the user, calculate the weight distribution of each semantic path, focus on the core path, and obtain the semantic representation vector; S42, Match the semantic representation vector with the nodes in the dynamically adapted graph structure, and use Euclidean distance to measure the matching degree to obtain the matching metric value; S43, if the matching metric is less than or equal to the preset intent matching threshold, integrate the semantic representation vector with the attribute association data in the dynamically adapted graph structure to generate an optimized query representation vector.
[0063] In step S41, the fuzzy query data input by the user is obtained, the weight distribution of each semantic path is calculated, the core path is focused, and the semantic representation vector is obtained.
[0064] It should be noted that the process of obtaining fuzzy query data and calculating the semantic path weight distribution, and focusing on the core path to obtain the semantic representation vector, refers to obtaining user fuzzy query data such as "abnormal sound nearby", using an attention mechanism and a softmax function to calculate the weight distribution of each semantic path in the query, such as "nearby-location-abnormal sound" and "abnormal sound-type-alarm sound", filtering out the core paths with a weight ≥ 0.7, and mapping the core path features into a low-dimensional vector to obtain the semantic representation vector.
[0065] It is worth noting that the core path weight threshold is set at 0.7 because fuzzy queries have semantic ambiguity. Paths with a weight below 0.7 are secondary semantics. Focusing on core paths with a weight ≥ 0.7 can avoid interference from secondary semantics and ensure that the semantic representation vector can accurately reflect the user's core query intent. For example, when the weight of the "abnormal sound" path is 0.75, it is prioritized to construct the vector.
[0066] For example, when a user inputs the fuzzy query "strange noise at the intersection", the weights of each semantic path are calculated: "intersection-location-strange noise" has a weight of 0.78, "strange noise-type-mechanical sound" has a weight of 0.62, the core path with a weight of 0.78 is focused, and its features are mapped into a low-dimensional vector to obtain the semantic representation vector.
[0067] In step S42, the semantic representation vector is matched with the nodes in the dynamically adapted graph structure, and the matching degree is measured by Euclidean distance to obtain the matching metric value.
[0068] It should be noted that the matching of semantic representation vectors with dynamic graph nodes and the use of Euclidean distance to measure the matching degree to obtain a matching metric value refers to matching the semantic representation vectors with the feature vectors of nodes such as "traffic noise" and "abnormal alarm" in the dynamically adapted graph structure, and measuring the difference between the two using the Euclidean distance formula to obtain a matching metric value.
[0069] For example, the semantic representation vector of “strange noise at the intersection” is matched with the nodes of “abnormal horn blasting at the intersection” and “road construction noise” in the dynamic graph. The Euclidean distances are calculated to be 0.12 and 0.35, respectively, and the matching metrics are 0.12 (corresponding to the node of “abnormal horn blasting at the intersection”) and 0.35 (corresponding to the node of “road construction noise”).
[0070] In step S43, if the matching metric is less than or equal to the preset intent matching threshold, the semantic representation vector is integrated with the attribute association data in the dynamically adapted graph structure to generate an optimized query representation vector.
[0071] It should be noted that the phrase "if the matching metric is less than or equal to the preset intent matching threshold, the data will be integrated to generate an optimized query representation vector" means that the preset intent matching threshold is 0.2 (Euclidean distance ≤ 0.2 is considered a high match). If the matching metric is ≤ 0.2, the semantic representation vector and the attribute association data such as "occurrence location" and "sound pressure level" of the nodes in the dynamic graph will be integrated through weighted fusion (vector weight 0.6, attribute data weight 0.4) to generate an optimized query representation vector.
[0072] For example, the matching metric between the semantic vector of “strange noise at the intersection” and the node of “abnormal honking at the intersection” is 0.12 (≤0.2). The node’s attribute data of “location: Chengdong intersection” and “sound pressure level: 78 decibels” are integrated and weighted to generate an optimized query representation vector.
[0073] In step S5, cross-modal association features are extracted from the optimized query representation vector to determine the semantic mapping fusion relationship of the entity acoustic signals. The expanded spectral coverage is then compared with a preset coverage evaluation threshold to obtain a preliminary retrieval result set, including: S51, extract cross-modal association features from the optimized query representation vector, and perform semantic decomposition based on the entity acoustic signals in the cross-modal association features to obtain an acoustic semantic vector; S52, the acoustic semantic vector is converted into a vector form in the graph embedding space, and the cosine similarity is calculated based on the semantic path features in the dynamically adapted graph structure to determine the semantic mapping fusion relationship of the entity acoustic signal. S53, Based on the semantic mapping fusion relationship, and by integrating the acoustic semantic vector with the attribute extension data in the dynamically adapted spectrogram structure, a cross-modal representation vector is generated; S54. Based on the comparison between the number of coverage semantic dimensions of the cross-modal representation vector and the preset coverage evaluation threshold corresponding to the coverage range of the expanded map, if the threshold is exceeded, the associated multimedia data is filtered to obtain a preliminary search result set.
[0074] In step S51, cross-modal association features are extracted from the optimized query representation vector, and semantic decomposition is performed on the entity acoustic signals in the cross-modal association features to obtain an acoustic semantic vector.
[0075] It should be noted that the step of extracting cross-modal association features from the optimized query representation vector and decomposing the entity acoustic signal to obtain the acoustic semantic vector refers to extracting the cross-modal association features of "text position + acoustic signal" from the optimized query representation vector, performing semantic decomposition on the entity acoustic signal through the signal decomposition rule set (time domain analysis + frequency domain transformation), extracting features such as frequency and duration, mapping them into a low-dimensional vector, and obtaining the acoustic semantic vector.
[0076] For example, from the optimized query vector of "abnormal honking at intersection", the cross-modal features of "intersection + abnormal honking" are extracted, and the abnormal honking signal is decomposed: the duration of 2 seconds is obtained in the time domain, the main peak frequency of 1100Hz is obtained in the frequency domain, and it is mapped into a low-dimensional vector to obtain the acoustic semantic vector.
[0077] In step S52, the acoustic semantic vector is converted into a vector form in the graph embedding space, and the cosine similarity is calculated based on the semantic path features in the dynamically adapted graph structure to determine the semantic mapping fusion relationship of the entity acoustic signal.
[0078] It should be noted that the process of converting acoustic semantic vectors to the graph embedding space and calculating cosine similarity to determine semantic mapping fusion relationship refers to converting acoustic semantic vectors into an embedding space form consistent with the dynamically adapted graph structure node vectors through linear transformation, and then calculating the cosine similarity between this vector and semantic path features in the graph such as "abnormal honking → traffic incident" and "mechanical noise → equipment failure". If the similarity is ≥0.75, it is determined to be the semantic mapping fusion relationship of the entity acoustic signal.
[0079] It is worth noting that the cosine similarity threshold is set at 0.75 because this value can effectively distinguish the strong correlation between acoustic signals and semantic paths. Paths with a similarity below 0.75 have weak correlation, such as the path similarity of abnormal horn honking and "equipment failure" being 0.6. Paths with a similarity above 0.75, such as the path similarity of abnormal horn honking and "traffic incident" being 0.82, have strong correlation and can accurately determine the mapping and fusion relationship.
[0080] For example, after converting the acoustic semantic vector of "abnormal honking at intersection" to the graph embedding space, the cosine similarity with the paths "abnormal honking → traffic congestion" and "abnormal honking → pedestrian violation" is calculated to be 0.81 and 0.65, respectively, thus determining "abnormal honking → traffic congestion" as a semantic mapping fusion relationship.
[0081] In step S53, based on the semantic mapping fusion relationship, the acoustic semantic vector is integrated with the attribute extension data in the dynamically adapted spectrogram structure to generate a cross-modal representation vector.
[0082] It should be noted that the aforementioned generation of cross-modal representation vectors based on semantic mapping fusion relationships refers to using semantic mapping fusion relationships such as "abnormal honking → traffic congestion" as a link, integrating acoustic semantic vectors with attribute extension data of corresponding path nodes in the dynamic graph, such as "impact range of 200 meters" and "occurrence time of morning peak" for "traffic congestion", and integrating them through weighted averaging, with an acoustic vector weight of 0.55 and an attribute data weight of 0.45, to generate a cross-modal representation vector containing acoustic, attribute, and semantic path features.
[0083] It is worth noting that the weighting is set to "0.55 for acoustic vector and 0.45 for attribute data" because acoustic vector is the core basis for retrieval and needs to be given higher weight to ensure the accuracy of the retrieval direction. At the same time, attribute data can supplement scene information and avoid the vector semantics being too simplistic. The balance between the two can make the cross-modal representation vector more comprehensively adaptable to retrieval needs.
[0084] For example, based on the mapping relationship of "abnormal honking → traffic congestion", the acoustic vector of "abnormal honking" (weight 0.55) is integrated with the attribute data of "impact range 200 meters" and "morning peak period" of "traffic congestion" (weight 0.45) to generate a cross-modal representation vector.
[0085] In step S54, the number of coverage semantic dimensions of the cross-modal representation vector is compared with the preset coverage evaluation threshold corresponding to the coverage range of the expanded map. If the threshold is exceeded, the associated multimedia data is filtered to obtain a preliminary retrieval result set.
[0086] It should be noted that the comparison of the number of semantic dimensions covered by the cross-modal vectors with a preset threshold, and the filtering of data to obtain a preliminary search result set if the threshold is exceeded, refers to the comparison of the total number of semantic dimensions covered by the cross-modal representation vectors with the preset coverage evaluation threshold (≥8 dimensions) corresponding to the coverage range of the expanded graph. If the number of dimensions is ≥8, audio segments and video segments associated with the vector in the dynamic graph are filtered to form a preliminary search result set.
[0087] It is worth noting that the preset coverage evaluation threshold is set to ≥8 dimensions because the expanded graph covers multiple semantic dimensions such as "acoustics, attributes, and events". Vector semantic coverage is incomplete if it is less than 8 dimensions, which may lead to the omission of key information in the search results. ≥8 dimensions (such as covering 10 dimensions such as frequency, location, and event type) can ensure that the vectors can match the multi-dimensional semantics of the graph, resulting in more comprehensive search results.
[0088] For example, the cross-modal vector of "abnormal honking at intersection" covers 10 semantic dimensions (≥8), such as "frequency 1100Hz, duration 2 seconds, location Chengdong intersection, event type traffic congestion". The "audio of honking at Chengdong intersection during morning rush hour" and "video of traffic congestion" associated in the graph are filtered to obtain a preliminary search result set.
[0089] In step S6, if the retrieval accuracy of the preliminary retrieval result set is lower than a preset accuracy threshold, the attention weights are adjusted to obtain a refined semantic mapping fusion relationship. Combining the expanded map coverage with a preset coverage evaluation threshold, it is determined whether the overall map coverage meets the dynamic adaptation verification, including: S61, calculate the retrieval accuracy based on the preliminary retrieval result set, and extract the entity acoustic signal and node data of the dynamically adapted spectral structure from the preliminary retrieval result set; S62, perform secondary semantic decomposition based on the entity acoustic signal to obtain the acoustic semantic vector; S63, calculate the association strength based on the acoustic semantic vector and the semantic path features in the dynamically adapted graph structure. If the association strength is lower than the preset association strength threshold, adjust the attention weight to obtain the refined semantic mapping fusion relationship. S64, based on the cross-modal representation data corresponding to the refined semantic mapping fusion relationship and the coverage data of the dynamically adapted graph structure, count the total number of semantic dimensions covered by the overall graph. S65, compare the total number of semantic dimensions covered by the overall map with the preset coverage evaluation threshold corresponding to the expanded map coverage range. If the threshold is met, determine that the overall map coverage range meets the dynamic adaptation verification, and output the refined semantic mapping fusion relationship.
[0090] In step S61, the retrieval accuracy is calculated based on the preliminary retrieval result set, and the node data of the entity acoustic signal and the dynamically adapted spectral structure in the preliminary retrieval result set are extracted.
[0091] It should be noted that the calculation of the retrieval accuracy of the preliminary retrieval result set and the extraction of entity acoustic signals and graph node data refers to calculating the retrieval accuracy by "number of results that accurately match the query intent / total number of results × 100%". If the accuracy is < 90% (preset accuracy threshold), entity acoustic signals such as "abnormal horn" and "construction noise" are extracted from the result set, as well as related node data such as "traffic congestion" and "equipment failure" in the dynamically adapted graph structure.
[0092] It is worth noting that the preset accuracy threshold is set to 90% because search results below this value contain a lot of irrelevant data. For example, when the accuracy is 82%, 18% of the results do not match the query intent, which will affect the user's ability to obtain effective information. Using 90% as the standard can ensure the reliability of the initial result set, and at the same time, when triggering subsequent optimization processes, it can specifically solve the problem of insufficient accuracy.
[0093] For example, the initial search result set contains 10 data entries, of which 8 match the query intent of "abnormal honking at intersections". The accuracy is calculated to be 80% (<90%). The acoustic signals of the "abnormal honking" entities in the result set are extracted, as well as the data of the "traffic congestion" and "intersection monitoring" nodes in the map.
[0094] In step S62, a secondary semantic decomposition is performed based on the entity acoustic signal to obtain an acoustic semantic vector.
[0095] It should be noted that the above-mentioned secondary semantic decomposition based on entity acoustic signals to obtain acoustic semantic vectors refers to performing secondary semantic decomposition on the extracted entity acoustic signals using more refined signal decomposition rules (subdividing the frequency band + time domain segmentation)—subdividing the frequency band to the 50Hz interval, segmenting the time domain into 0.5-second segments, extracting the subdivided frequency, amplitude and other features, mapping them into low-dimensional vectors to obtain acoustic semantic vectors.
[0096] For example, the "abnormal honking at intersection" signal is decomposed in two ways: the frequency band is subdivided into the 50Hz range, and the 1100-1150Hz frequency band accounts for 0.65%; the time domain is divided into 0.5-second segments, and the peak amplitude of 85dB in the 0.5-1 second segment is extracted and mapped into a low-dimensional vector to obtain the acoustic semantic vector.
[0097] In step S63, the association strength is calculated based on the acoustic semantic vector and the semantic path features in the dynamically adapted graph structure. If the association strength is lower than the preset association strength threshold, the attention weight is adjusted to obtain the refined semantic mapping fusion relationship.
[0098] It should be noted that the calculation of the association strength between the acoustic semantic vector and the semantic path in the graph involves adjusting the attention weights if the value is below a threshold to obtain a refined mapping relationship. Specifically, Euclidean distance is used to calculate the association strength between the acoustic semantic vector and semantic path features such as "abnormal honking → traffic congestion" in the dynamic graph. The preset association strength threshold is 0.5; a distance greater than 0.5 is considered low association. At this point, a feedback loop is activated, which collects the deviation information between the calculated association strength and the preset association strength threshold and sends this information back to the attention weight adjustment module. Based on the degree of deviation, the module increases the attention weight of high-association paths from 0.4 to 0.6. The feedback loop mechanism positively enhances the attention weight of high-association paths and suppresses the weight of low-association paths according to a preset learning rate based on the association strength deviation, thereby obtaining a refined semantic mapping fusion relationship.
[0099] It is worth noting that the preset correlation strength threshold is set to 0.5 (Euclidean distance ≤ 0.5 is considered high correlation) because when the distance is > 0.5, the acoustic vector and path features differ significantly (e.g., the distance between abnormal horn and "equipment failure" path is 0.62). The weights need to be adjusted to strengthen highly correlated paths (e.g., the distance between "traffic congestion" path is 0.38) to ensure that the refined mapping relationship can accurately reflect the correlation between acoustic signals and spectral semantics.
[0100] For example, the correlation strengths between the acoustic vector of "abnormal honking at intersection" and the paths "abnormal honking → traffic congestion" and "abnormal honking → equipment failure" are calculated to be 0.38 and 0.62, respectively. Since 0.62 > 0.5, the attention weight of the "traffic congestion" path is adjusted from 0.4 to 0.6 to obtain the refined mapping relationship of "abnormal honking → traffic congestion".
[0101] In step S64, based on the cross-modal representation data corresponding to the refined semantic mapping fusion relationship and the coverage data of the dynamically adapted graph structure, the total number of semantic dimensions covered by the overall graph is counted.
[0102] It should be noted that the total number of semantic dimensions of the overall graph coverage based on the refined mapping relationship data and graph coverage data refers to extracting the cross-modal representation data corresponding to the refined mapping relationship, combining it with the coverage data of the dynamically adapted graph structure, such as entity category and relationship type dimensions, and counting the total number of all non-repeating semantic dimensions.
[0103] For example, the refined mapping relationship cross-modal data covers three dimensions: "frequency, location, and event type", while the dynamic graph coverage data covers five dimensions: "entity category, relationship type, and attribute parameters". After deduplication, the total number of semantic dimensions covered by the overall graph is eight.
[0104] In step S65, the total number of semantic dimensions covered by the overall map is compared with the preset coverage evaluation threshold corresponding to the expanded map coverage range. If the threshold is met, the overall map coverage range is determined to meet the dynamic adaptation verification, and the refined semantic mapping fusion relationship is output.
[0105] It should be noted that comparing the total number of semantic dimensions of the overall map coverage with a preset threshold, and determining that the map coverage meets the dynamic adaptation verification and outputting the refined relationship, means comparing the total number of semantic dimensions of the overall map coverage with the preset coverage evaluation threshold (≥8 dimensions) corresponding to the expanded map coverage range. If the number of dimensions is ≥8, it is determined that the overall map coverage range can adapt to dynamic data changes and retrieval needs, meets the dynamic adaptation verification, and outputs the refined semantic mapping fusion relationship.
[0106] It is worth noting that the preset coverage evaluation threshold follows the ≥8 dimensions in step S5 to ensure the consistency of the graph coverage. Step S5 requires the retrieval vector to cover ≥8 dimensions to match the graph, and step S6 verifies that the graph itself covers ≥8 dimensions, which can form a logical closed loop to ensure that the graph can not only adapt to vector retrieval, but also dynamically cover multi-dimensional semantics.
[0107] For example, the total number of semantic dimensions covered by the overall map is 8 (≥8). If the comparison with the preset coverage evaluation threshold meets the requirements, it is determined that the coverage of the overall map meets the dynamic adaptation verification, and the refined semantic mapping fusion relationship of "abnormal honking → traffic congestion" is output.
[0108] In step S7, based on the refined semantic mapping fusion relationship, new content links are integrated and feedback weight adjustments are processed. If the new content links meet the preset link verification threshold, the final retrieval result set is obtained, including: S71, based on the refined semantic mapping fusion relationship, input the data in the dynamically adapted graph structure and the newly added content data in real time into the real-time integration module to generate a set of structured content links; S72, compare the number of links in the structured content link set with a preset link verification threshold. If the threshold is exceeded, it is determined that the verification requirements are met. S73, based on the structured content link set that meets the verification requirements, extract the path information of the associated multimedia data and the dynamically adapted graph structure, sort and integrate them according to the relevance of the query intent, and form a final retrieval result set that includes content link integration and result extraction.
[0109] In step S71, based on the refined semantic mapping fusion relationship, the data in the dynamically adapted graph structure and the newly added content data are input into the real-time integration module to generate a structured content link set.
[0110] It should be noted that the process of integrating data based on the refined mapping relationship to generate a structured content link set refers to using the refined semantic mapping fusion relationship, such as "abnormal horn honking → traffic congestion," as a basis. The data of the "traffic congestion" node in the dynamically adapted graph structure is input into the real-time integration module along with the newly added content data. The module generates a structured content link set, such as "abnormal horn honking audio → traffic congestion node → on-site video," according to the association logic of "entity-content-graph path."
[0111] For example, based on the refined relationship of "abnormal honking → traffic congestion", the data of the "Chengdong intersection morning rush hour congestion" node in the dynamic map is input into the integration module along with the real-time newly added "Chengdong intersection honking audio" and "congestion scene 720p video" to generate a structured content link set.
[0112] In step S72, the number of links in the structured content link set is compared with a preset link verification threshold. If the threshold is exceeded, the verification requirements are deemed met.
[0113] It should be noted that the comparison of the number of structured content links in the set with a preset threshold, and the determination that the verification requirements are met if the threshold is exceeded, means that the preset link verification threshold is 10 links. This threshold is set based on the requirement that most search scenarios need at least 10 valid content links to support the diversity of results. The number of valid links in the structured content link set is counted. If the number is ≥10, it is determined that the content covered by the link set is rich enough and meets the verification requirements.
[0114] It is worth noting that the preset threshold of 10 is because a set of fewer than 10 links may result in insufficient search results, failing to meet the diverse information needs of users; ≥10 links ensure that the results cover related content from different perspectives, improving the search experience.
[0115] For example, the structured content link set contains "3 audio clips of horns, 5 videos of traffic congestion, and 2 graph path associations", totaling 10 valid links (≥10), which meets the verification requirements when compared with the preset threshold.
[0116] In step S73, based on the structured content link set that meets the verification requirements, the path information of the associated multimedia data and the dynamically adapted graph structure is extracted, sorted and integrated according to the relevance of the query intent, and a final retrieval result set including content link integration and result extraction is formed.
[0117] It should be noted that generating the final search result set based on the verified link set refers to extracting the path information of the associated multimedia data and dynamic graph from the link set that meets the verification requirements, sorting them according to the "query intent relevance score", and integrating them to form a final search result set containing "content links + graph paths + relevance ranking".
[0118] For example, the associated audio and video data and the path information of "abnormal honking → traffic congestion" are extracted from the verified link set, sorted by relevance score, with the audio of honking at Chengdong intersection ranking first with a score of 0.92 and the video of congestion ranking second with a score of 0.85, and integrated to form the final search result set.
[0119] In summary, this invention optimizes the entire process of multimedia data retrieval, from collection and screening to graph construction and precise retrieval, through a knowledge graph-driven multimodal data retrieval mechanism. It effectively solves the problems of insufficient semantic understanding and weak dynamic adaptability of traditional retrieval technologies, significantly improves the accuracy and real-time performance of multimedia data retrieval, and provides strong support for the efficient acquisition of multimedia information in complex scenarios such as intelligent security and intelligent urban management.
[0120] Reference Figure 2 The second embodiment of the present invention provides a multimedia data intelligent retrieval system based on knowledge graphs, comprising: The knowledge graph construction module is used to acquire acoustic signals and video clips from multimedia data sources, extract semantic features to obtain an initial entity set and relation links, construct a basic knowledge graph and output it. The graph expansion module is used to propagate node information, process relation weight calculation, complete the new entity recognition and attribute association expansion based on the basic knowledge graph, determine the coverage of the expanded graph and obtain the expanded knowledge graph, and output the expanded knowledge graph and the coverage of the expanded graph. The dynamic update module is used to acquire new acoustic data of entity behavior in real time, compare the new acoustic data with the extended knowledge graph, fuse new relation weights and adjust the number of iterations to obtain a dynamically adapted graph structure and output it. The query optimization module is used to focus on key semantic paths and determine the matching degree based on fuzzy query data for the dynamically adapted graph structure, generate an optimized query representation vector and output it. The preliminary retrieval module is used to extract cross-modal association features based on the optimized query representation vector, and combine the dynamically adapted graph structure with the expanded graph coverage to obtain and output a preliminary retrieval result set; The weight adjustment module is used to verify the accuracy based on the preliminary retrieval result set, adjust the attention weight to obtain the refined semantic mapping fusion relationship, combine the expanded map coverage to determine whether the map coverage meets the dynamic adaptation verification, and output the refined semantic mapping fusion relationship. The result generation module is used to integrate new content links based on the refined semantic mapping fusion relationship and verify them based on a preset link verification threshold, so as to obtain and output the final retrieval result set containing content link integration and result extraction.
[0121] The modules of this system work together in sequence to realize the intelligent processing of the entire process of multimedia data retrieval. Its working principle corresponds one-to-one with the above-mentioned method embodiments, and will not be repeated here.
[0122] This invention also provides an electronic device, including a processor, a memory, and a computer program stored in the memory. When the processor executes the program, it implements all the steps of the aforementioned multimedia data intelligent retrieval method based on knowledge graphs. This electronic device can employ a high-performance computing terminal and acquire data by accessing multimedia data sources such as surveillance cameras and recording devices, providing computing power support for graph construction and retrieval calculations.
[0123] This invention also provides a computer-readable storage medium storing a computer program. When the program runs, it controls the device containing the storage medium to execute the aforementioned knowledge graph-based intelligent multimedia data retrieval method. This medium can be a high-speed solid-state drive, deployed at the edge computing node of the intelligent retrieval system, to achieve localized multimedia data processing and retrieval result generation, reducing data transmission latency.
[0124] It should be noted that the knowledge graph-based intelligent multimedia data retrieval system provided in this embodiment of the invention is used to execute all the process steps of the knowledge graph-based intelligent multimedia data retrieval method in the above embodiments. The working principles and beneficial effects of the two are one-to-one, so they will not be described again.
[0125] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor. The processor is the control center of the electronic device, connecting all parts of the electronic device via various interfaces and lines.
[0126] The memory can be used to store the computer programs and / or modules. The processor implements various functions of the electronic device by running or executing the computer programs and / or modules stored in the memory and by calling data stored in the memory. The memory may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the mobile phone (such as audio data, phonebook, etc.). In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0127] Wherein, if the modules / units integrated in the electronic device are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of the present invention can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium can be appropriately added or removed according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electrical carrier signals and telecommunication signals.
[0128] It should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Furthermore, in the accompanying drawings of the device embodiments provided by this invention, the connection relationship between modules indicates that they have a communication connection, which can be implemented as one or more communication buses or signal lines. Those skilled in the art can understand and implement this without creative effort. The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of this invention. It should be understood that the above descriptions are merely specific embodiments of this invention and are not intended to limit the scope of protection of this invention. In particular, for those skilled in the art, any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A multimedia data intelligent retrieval method based on knowledge graphs, characterized in that, include: Acquire acoustic signals and video clips from multimedia data sources, extract semantic features to obtain an initial entity set and relational links, and construct a basic knowledge graph; Based on the aforementioned basic knowledge graph, node information is propagated, relation weight calculation is processed, new entity recognition and attribute association expansion are completed, and the coverage of the expanded graph is determined. Acquire new acoustic data of real-time incoming entity behavior. If the similarity with existing nodes in the expanded map coverage exceeds a preset node similarity threshold, fuse new relationship weights and adjust the propagation iteration number to obtain a dynamically adapted map structure. For the dynamically adapted graph structure, focus on key semantic paths and determine the matching degree between the fuzzy query intent and the dynamically adapted graph structure. If the matching degree is less than or equal to a preset intent matching threshold, an optimized query representation vector is obtained. Cross-modal association features are extracted from the optimized query representation vector to determine the semantic mapping fusion relationship of entity acoustic signals. The expanded map coverage is compared with the preset coverage evaluation threshold to obtain a preliminary retrieval result set. If the retrieval accuracy of the preliminary retrieval result set is lower than the preset accuracy threshold, the attention weight is adjusted to obtain the refined semantic mapping fusion relationship. Combined with the expanded map coverage and the preset coverage evaluation threshold, it is determined whether the overall map coverage meets the dynamic adaptation verification. Based on the refined semantic mapping fusion relationship, new content links are integrated and feedback weight adjustments are processed. If the new content links meet the preset link verification threshold, the final retrieval result set is obtained. The process involves extracting cross-modal association features from the optimized query representation vector, determining the semantic mapping fusion relationship of entity acoustic signals, and comparing the expanded spectral coverage with a preset coverage evaluation threshold to obtain a preliminary retrieval result set, including: Cross-modal association features are extracted from the optimized query representation vector, and semantic decomposition is performed on the entity acoustic signals in the cross-modal association features to obtain an acoustic semantic vector; The acoustic semantic vector is converted into a vector form in the graph embedding space. Based on the semantic path features in the dynamically adapted graph structure, the cosine similarity is calculated to determine the semantic mapping and fusion relationship of the entity acoustic signal. Based on the semantic mapping fusion relationship, and by integrating the acoustic semantic vector with the attribute extension data in the dynamically adapted spectrogram structure, a cross-modal representation vector is generated. Based on the comparison between the number of coverage semantic dimensions of the cross-modal representation vector and the preset coverage evaluation threshold corresponding to the coverage range of the expanded map, if the threshold is exceeded, the associated multimedia data is filtered to obtain a preliminary search result set.
2. The multimedia data intelligent retrieval method based on knowledge graphs according to claim 1, characterized in that, The process of acquiring acoustic signals and video clips from a multimedia data source, extracting semantic features to obtain an initial entity set and relational links, and constructing a basic knowledge graph includes: Acoustic signals and video clips are collected from a multimedia data source, and the acoustic signals and video clips are filtered according to a preset quality filtering threshold to obtain a filtered multimedia dataset. Semantic features are extracted from the filtered multimedia dataset to obtain an initial entity set; Analyze the semantic features in the initial entity set. If the similarity exceeds a preset feature association threshold, determine the association relationship and form a set of relationship links. The set of relationship links and the initial set of entities are imported into a pre-constructed graph structure to generate a basic knowledge graph.
3. The multimedia data intelligent retrieval method based on knowledge graphs according to claim 1, characterized in that, Based on the aforementioned basic knowledge graph, the process of propagating node information, processing relation weight calculations, completing the identification of new entities and attribute association expansion, and determining the coverage area of the expanded knowledge graph includes: Extract semantic features corresponding to entities and relations from the basic knowledge graph, propagate the node information corresponding to the semantic features, and obtain the updated node representation; For the updated node representation, calculate the weighted relation strength to form an expanded relation link set; If the relationship strength exceeds a preset relationship strength threshold, new entities are extracted from the expanded relationship link set, the category of the new entities is identified and the associated attributes are obtained, and the attribute association expansion is completed. By integrating the newly added entities, the expanded set of relationship links, and the associated attributes with the basic knowledge graph, the coverage of the expanded graph is determined, and the expanded knowledge graph is obtained.
4. The multimedia data intelligent retrieval method based on knowledge graphs according to claim 1, characterized in that, The process of acquiring new acoustic data on real-time influx of entity behavior, and if the similarity between this data and existing nodes within the expanded geographic graph coverage exceeds a preset node similarity threshold, involves fusing new relation weights and adjusting the propagation iteration count to obtain a dynamically adaptive geographic graph structure, including: Acquire new acoustic data of entity behavior from real-time data stream, extract the frequency distribution feature vector of the new acoustic data, and obtain an initial acoustic representation; Based on the initial acoustic representation and the existing node feature vectors in the expanded map coverage area, the cosine similarity is calculated. If the similarity exceeds a preset node similarity threshold, the set of nodes that need to be updated is determined. Based on the set of nodes that need to be updated, calculate the new relation weights and merge the new relation weights to obtain the updated weight set; The propagation iteration number is adjusted according to the updated weight set, the connection strength between nodes in the node set to be updated is recalculated based on node similarity, the knowledge graph structure is updated and extended, and a dynamically adapted graph structure is obtained.
5. The multimedia data intelligent retrieval method based on knowledge graphs according to claim 1, characterized in that, The dynamically adapted graph structure focuses on key semantic paths and determines the matching degree between the fuzzy query intent and the dynamically adapted graph structure. If the matching degree is less than or equal to a preset intent matching threshold, an optimized query representation vector is obtained, including: Obtain fuzzy query data input by the user, calculate the weight distribution of each semantic path, focus on the core path, and obtain the semantic representation vector; The semantic representation vector is matched with the nodes in the dynamically adapted graph structure, and the matching degree is measured by Euclidean distance to obtain the matching metric value. If the matching metric is less than or equal to the preset intent matching threshold, the semantic representation vector is integrated with the attribute association data in the dynamically adapted graph structure to generate an optimized query representation vector.
6. The multimedia data intelligent retrieval method based on knowledge graphs according to claim 1, characterized in that, If the retrieval accuracy of the preliminary retrieval result set is lower than a preset accuracy threshold, the attention weights are adjusted to obtain a refined semantic mapping fusion relationship. Combined with the expanded map coverage and a preset coverage evaluation threshold, it is determined whether the overall map coverage meets the dynamic adaptation verification, including: The retrieval accuracy is calculated based on the preliminary retrieval result set, and the node data of the entity acoustic signal and the dynamically adapted spectral structure in the preliminary retrieval result set are extracted. A secondary semantic decomposition is performed on the entity's acoustic signal to obtain an acoustic semantic vector; The association strength is calculated based on the acoustic semantic vector and the semantic path features in the dynamically adapted graph structure. If the association strength is lower than the preset association strength threshold, the attention weight is adjusted to obtain the refined semantic mapping fusion relationship. Based on the cross-modal representation data corresponding to the refined semantic mapping fusion relationship and the coverage data of the dynamically adapted graph structure, the total number of semantic dimensions covered by the overall graph is counted. Based on the comparison between the total number of semantic dimensions covered by the overall map and the preset coverage evaluation threshold corresponding to the expanded map coverage range, if the threshold is met, it is determined that the overall map coverage range meets the dynamic adaptation verification, and the refined semantic mapping fusion relationship is output.
7. The multimedia data intelligent retrieval method based on knowledge graphs according to claim 1, characterized in that, The process involves integrating new content links based on the refined semantic mapping fusion relationship and processing feedback weight adjustments. If the new content links meet the preset link verification threshold, the final search result set is obtained, including: Based on the refined semantic mapping fusion relationship, the data in the dynamically adapted graph structure and the newly added content data are input into the real-time integration module to generate a set of structured content links; The number of links in the structured content link set is compared with a preset link verification threshold. If the threshold is exceeded, the verification requirement is deemed met. Based on the structured content link set that meets the verification requirements, the path information of the associated multimedia data and the dynamically adapted graph structure is extracted, sorted and integrated according to the relevance of the query intent, and a final retrieval result set including content link integration and result extraction is formed.
8. A multimedia data intelligent retrieval system based on knowledge graphs, characterized in that, include: The graph construction module is used to acquire acoustic signals and video clips from multimedia data sources, extract semantic features to obtain an initial entity set and relation links, and construct a basic knowledge graph. The graph expansion module is used to propagate node information, process relation weight calculations, complete the new entity recognition and attribute association expansion based on the basic knowledge graph, and determine the coverage of the expanded graph. The dynamic update module is used to acquire new acoustic data of entity behavior that is coming in in real time. If the similarity with the existing nodes in the expanded map coverage exceeds the preset node similarity threshold, the new relation weights are fused and the number of iterations is adjusted to obtain a dynamically adapted map structure. The query optimization module is used to focus on key semantic paths and determine the matching degree between the fuzzy query intent and the dynamically adapted graph structure. If the matching degree is less than or equal to a preset intent matching threshold, an optimized query representation vector is obtained. The preliminary retrieval module is used to extract cross-modal association features from the optimized query representation vector, determine the semantic mapping fusion relationship of entity acoustic signals, and compare the expanded map coverage with the preset coverage evaluation threshold to obtain a preliminary retrieval result set. The weight adjustment module is used to adjust the attention weights to obtain a refined semantic mapping fusion relationship if the retrieval accuracy of the preliminary retrieval result set is lower than a preset accuracy threshold. Combined with the expanded map coverage and the preset coverage evaluation threshold, it determines whether the overall map coverage meets the dynamic adaptation verification. The result generation module is used to integrate new content links and process feedback weight adjustments based on the refined semantic mapping fusion relationship. If the new content links meet the preset link verification threshold, the final search result set is obtained. The preliminary retrieval module is specifically used for: Cross-modal association features are extracted from the optimized query representation vector, and semantic decomposition is performed on the entity acoustic signals in the cross-modal association features to obtain an acoustic semantic vector; The acoustic semantic vector is converted into a vector form in the graph embedding space. Based on the semantic path features in the dynamically adapted graph structure, the cosine similarity is calculated to determine the semantic mapping and fusion relationship of the entity acoustic signal. Based on the semantic mapping fusion relationship, and by integrating the acoustic semantic vector with the attribute extension data in the dynamically adapted spectrogram structure, a cross-modal representation vector is generated. Based on the comparison between the number of coverage semantic dimensions of the cross-modal representation vector and the preset coverage evaluation threshold corresponding to the coverage range of the expanded map, if the threshold is exceeded, the associated multimedia data is filtered to obtain a preliminary search result set.
Citation Information
Patent Citations
Knowledge graph-oriented cross-media retrieval system
CN105550190A
Knowledge graph construction method for multi-modal data
CN120296652A