Methods and Systems for Intelligent Tag Generation and Knowledge Graph Construction of Converged Media Content
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-23
- Publication Date
- 2026-08-14
AI Technical Summary
[0002]在当前的融媒体内容管理实践中,针对多源异构数据的处理已具备一定的自动化基础,能够提取人物、物品、语音及字幕等基础信息并生成标签,然而,在面对复杂多变的媒体场景时,现有技术在构建数据间的深层逻辑关联以及动态优化方面,似乎还有进一步探索的空间
Smart Images

Figure CN122220953B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a method and system for generating intelligent tags and constructing knowledge graphs for converged media content. Background Technology
[0002] In current practices of converged media content management, the processing of multi-source heterogeneous data has reached a certain level of automation, capable of extracting basic information such as people, objects, voice and subtitles and generating tags. However, when faced with complex and ever-changing media scenarios, existing technologies seem to have room for further exploration in building deep logical connections between data and in dynamic optimization.
[0003] Taking a typical breaking news video as an example, the footage may simultaneously contain multiple moving figures, rapidly changing street scenes in the background, and scattered specific objects, accompanied by rapid on-site dialogue and scrolling subtitles. In existing conventional processing workflows, systems mostly tend to identify these elements as independent metadata entries. While they can record the appearance of a person, the presence of an object, or a segment of audio content, the existing association mechanisms may not be flexible enough when attempting to reconstruct the spatiotemporal context of a person's interaction with an object at a specific point in time, or to align subtle mappings between changes in tone of voice and semantic shifts in subtitles. Furthermore, as media data accumulates in the resource library, the distribution of data features becomes increasingly diversified. Existing label weighting methods are mostly based on static rules. When faced with segments that have low feature density but high actual semantic value, the system may struggle to automatically perceive and adjust their priority, potentially limiting the efficiency of intelligent utilization of media resources in complex application scenarios during deep retrieval. Summary of the Invention
[0004] This invention provides a method and system for generating intelligent tags and constructing knowledge graphs for converged media content. Through deep fusion of multi-source heterogeneous data and spatiotemporal semantic correlation calibration, it realizes the intelligent transformation and organization of media resources from fragmented features to structured knowledge.
[0005] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows:
[0006] Firstly, a method for generating intelligent tags and constructing knowledge graphs for converged media content, the method comprising:
[0007] Step 1: Map the features of people, items, scene descriptions, voice text, subtitles and sensitive information in the initial analysis result set into structured data to form structured label items, and summarize the structured label items into a structured label set.
[0008] Step 2: Convert the structured tag items in the structured tag set into graph entity nodes. With the media file as the root node, establish the spatiotemporal relationship between the corresponding entity nodes of people, items and scenes as graph edges; form semantic relationship edges by semantic mapping between the corresponding entity nodes of voice text and subtitle text, and mark the propagation path of the entity nodes corresponding to sensitive information features to construct a preliminary knowledge graph topology.
[0009] Step 3: Map the multi-source heterogeneous data of entity nodes and graph edges in the preliminary knowledge graph topology into high-dimensional feature vectors to construct a multi-dimensional association model. Divide the feature intervals according to the spatial distribution density of the feature vectors. Define media data as observation sequences and semantic data as state sequences. Use the preset node weights to predict the state sequence at the current time to obtain the predicted value. Calculate the gain coefficient in combination with the current feature interval density. Use the gain coefficient to weight and fuse the deviation between the observation sequence and the predicted value to update the node weights, thus obtaining the optimized knowledge graph.
[0010] Step 4: Store the optimized knowledge graph in a distributed database, construct a hierarchical retrieval system based on role permissions based on the stored graph data, and execute intelligent recommendation logic to realize the management and application of media resources.
[0011] Secondly, the system for intelligent tag generation and knowledge graph construction of converged media content includes:
[0012] The module is used to map the features of people, items, scene descriptions, voice text, subtitle text and sensitive information in the initial analysis result set into structured data, form structured label items, and summarize the structured label items into a structured label set.
[0013] The topology module is used to convert structured tag items in the structured tag set into graph entity nodes. With the media file as the root node, it establishes the spatiotemporal relationship between the corresponding entity nodes of people, items and scenes as graph edges; it forms semantic relationship edges by semantic mapping between the corresponding entity nodes of voice text and subtitle text, and marks the propagation path of the entity nodes corresponding to sensitive information features, thus constructing a preliminary knowledge graph topology structure.
[0014] The calibration module is used to map multi-source heterogeneous data of entity nodes and graph edges in the initial knowledge graph topology into high-dimensional feature vectors to construct a multi-dimensional association model. It divides feature intervals according to the spatial distribution density of feature vectors, defines media data as observation sequences, and defines semantic data as state sequences. It uses preset node weights to predict the state sequence at the current time to obtain the predicted value, and calculates the gain coefficient in combination with the current feature interval density. It uses the gain coefficient to weight and fuse the deviation between the observation sequence and the predicted value to update the node weights, thus obtaining the optimized knowledge graph.
[0015] The application module is used to store the optimized knowledge graph into a distributed database, build a hierarchical retrieval system based on role permissions based on the stored graph data, and execute intelligent recommendation logic to realize the management and application of media resources.
[0016] The above-described solution of the present invention has at least the following beneficial effects:
[0017] By extracting multi-source data from converged media, mapping structured tags, modeling spatiotemporal and semantic associations, calibrating multi-dimensional feature fusion, and optimizing knowledge graphs, the transformation of media resources from unstructured to structured and from scattered data to related knowledge has been achieved, improving the accuracy and semantic coherence of the knowledge graph. Based on this, a hierarchical retrieval and intelligent recommendation system has been built, realizing the transformation of media resources from discrete storage to structured association, improving retrieval efficiency and resource application value in complex scenarios. Attached Figure Description
[0018] Figure 1 This is a flowchart illustrating the method for generating intelligent tags and constructing knowledge graphs for converged media content provided in an embodiment of the present invention.
[0019] Figure 2 This is a schematic diagram of the intelligent tag generation and knowledge graph construction system for converged media content provided in an embodiment of the present invention. Detailed Implementation
[0020] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.
[0021] like Figure 1 As shown, embodiments of the present invention propose a method for generating intelligent tags and constructing knowledge graphs for converged media content, the method comprising the following steps:
[0022] Step 1: Map the features of people, items, scene descriptions, voice text, subtitles and sensitive information in the initial analysis result set into structured data to form structured label items, and summarize the structured label items into a structured label set.
[0023] Step 2: Convert the structured tag items in the structured tag set into graph entity nodes. With the media file as the root node, establish the spatiotemporal relationship between the corresponding entity nodes of people, items and scenes as graph edges; form semantic relationship edges by semantic mapping between the corresponding entity nodes of voice text and subtitle text, and mark the propagation path of the entity nodes corresponding to sensitive information features to construct a preliminary knowledge graph topology.
[0024] Step 3: Map the multi-source heterogeneous data of entity nodes and graph edges in the preliminary knowledge graph topology into high-dimensional feature vectors to construct a multi-dimensional association model. Divide the feature intervals according to the spatial distribution density of the feature vectors. Define media data as observation sequences and semantic data as state sequences. Use the preset node weights to predict the state sequence at the current time to obtain the predicted value. Calculate the gain coefficient in combination with the current feature interval density. Use the gain coefficient to weight and fuse the deviation between the observation sequence and the predicted value to update the node weights, thus obtaining the optimized knowledge graph.
[0025] Step 4: The optimized knowledge graph is stored in a distributed database. Based on the stored graph data, a hierarchical retrieval system with role-based permissions is constructed, and intelligent recommendation logic is executed to realize the management and application of media resources. In this embodiment of the invention, through the extraction of multi-source data from converged media, structured tag mapping, spatiotemporal and semantic association modeling, multi-dimensional feature fusion calibration, and knowledge graph optimization, the transformation of media resources from unstructured to structured, and from scattered data to associated knowledge, is realized. This improves the accuracy and semantic coherence of the knowledge graph. The hierarchical retrieval and intelligent recommendation system built on this basis realizes the transformation of media resources from discrete storage to structured association, improving retrieval efficiency and resource application value in complex scenarios.
[0026] In a preferred embodiment of the present invention, the person data, audio text data, item data, and subtitle data in the converged media file are extracted and analyzed to obtain an initial analysis result set, which may include:
[0027] In this embodiment of the invention, standardized preprocessing of converged media files is carried out as a preliminary operation for data extraction. This step is compatible with all formats of converged media files such as MP4, MOV, AVI, MP3, and WAV. Non-standard files are converted into high-definition parsable formats through audio and video codec libraries, and three independent data channels are separated: video track, audio track, and subtitle track. Addressing the characteristics of rapid scene switching, moving figures, and noisy environments in dynamic scenes such as breaking news, video stabilization, image enhancement, audio noise reduction, and keyframe fixed sampling are performed. Simultaneously, a global timestamp axis accurate to milliseconds and a unified image coordinate system are generated for the entire media file. Targeted extraction and professional analysis of four core data types are performed, utilizing AI vision, speech recognition, and OCR recognition to extract data on people, objects, audio text, and subtitles, and completing feature refinement and structured recording. For the people data, human / face targets in video frames are identified. The system distinguishes between main characters and background figures, extracting core features such as identity, gender, clothing, and actions, and simultaneously marking the coordinates of figures in the frame and the time intervals of their appearance and disappearance. For object data, it traverses keyframes of the video to identify physical objects within the frame, eliminates invalid background interference, and extracts the object's name, category, state, and relative position with surrounding elements, marking its spatiotemporal attributes. For audio and text data, it first performs noise reduction and sound source separation on the audio track, distinguishes different speakers and speech types, and then completes high-precision text conversion using speech-to-text technology, while extracting additional features such as speech rate, tone, and emotion and binding timestamps. For subtitle data, it is compatible with both soft and hard subtitle formats, directly parsing the subtitle track or recognizing on-screen subtitles through OCR, completing text cleaning, normalization, and semantic segmentation, and recording the subtitle content, style, and corresponding spatiotemporal information. All dimensions of data are marked with attributes according to a unified global spatiotemporal benchmark to avoid data isolation.
[0028] After data extraction, multi-source data fusion and basic analysis are conducted to address the pain point of isolated and unrelated data in traditional technologies, while ensuring data validity and standardization. Using a global timestamp as a benchmark, millisecond-level spatiotemporal alignment is achieved for four major data categories: people, objects, audio text, and subtitles, enabling one-to-one correspondence of multiple elements at specific time points. Simultaneously, data deduplication and error correction are performed, eliminating duplicately identified people and objects, as well as incorrectly transcribed text, and correcting semantic conflicts between audio and subtitles. In accordance with the requirements for compliance management of converged media, sensitive content in text and images is preliminarily scanned, and sensitive information features are marked. Finally, basic feature labels are generated for each data category, completing the structured output of the initial analysis result set. All data after preprocessing, extraction, and fusion analysis are integrated into a standardized and structured initial analysis result set. This result set includes four core datasets: person features, object features, audio text, and subtitle text, along with additional fields such as unique media file identifiers, a global timeline, and preliminary sensitive information screening tags.
[0029] In a preferred embodiment of the present invention, step 1 above maps the features of people, items, scene descriptions, voice text, subtitles, and sensitive information in the initial analysis result set into structured data to form structured tag items, and summarizes the structured tag items into a structured tag set, which may include:
[0030] In this embodiment of the invention, step 110 involves extracting features of people, items, scenes, voice text, subtitles, and sensitive information from the initial analysis result set to obtain a multi-dimensional feature dataset. Specifically, this includes: hierarchically decomposing the initial analysis result set by feature type; classifying and filtering all data in the result set according to the classification criteria for features of people, items, scenes, voice text, subtitles, and sensitive information to ensure that each type of feature is completely extracted without overlap or confusion; when extracting features of people, simultaneously extracting all associated attributes corresponding to that feature, including person identity information, gender, age range, actions, appearance timestamp, disappearance timestamp, and screen coordinates; when extracting item information, simultaneously extracting all associated attributes corresponding to that information, including item name, item category, item feature description, appearance timestamp, disappearance timestamp, screen coordinates, and item status attributes; when extracting scene descriptions, from... In the initial analysis result set, all associated attributes such as scene type, environmental features, spatial range, and corresponding spatiotemporal intervals are extracted from the image environment features and spatiotemporal correlation information. When extracting speech text, all associated attributes corresponding to the text are extracted simultaneously, including speech-to-text content, intonation features, speaker identifier, speech start timestamp, and speech end timestamp. When extracting subtitle text, all associated attributes corresponding to the text are extracted simultaneously, including subtitle text content, subtitle start timestamp, subtitle end timestamp, and corresponding image coordinates. When extracting sensitive information features, all associated attributes corresponding to the feature are extracted simultaneously, including sensitive information type, specific content of sensitive information, location of sensitive information, corresponding spatiotemporal stamp, and other associated feature information. After completing the full extraction of the six types of features, each type of feature and all its associated attributes are organized into independent feature subsets. Then, the six independent feature subsets are integrated to form a multi-dimensional feature dataset.
[0031] Step 111: Based on the preset label mapping rule library, map each type of feature in the feature dataset to its corresponding label item. Map person features to person entity labels, item information to item entity labels, scene descriptions to scene entity labels, speech text to speech text labels, subtitle text to subtitle text labels, and sensitive information features to sensitive labels, generating a preliminary label set including various preliminary labels. Specifically, this includes: retrieving the preset label mapping rule library in the system. This rule library is a pre-built and verified fixed correspondence library between features and labels. The library contains a one-to-one correspondence between feature type, specific feature content, basic label encoding, and standard label name. This correspondence covers all common feature content of the six types of features, including person features and item information, and can directly match various features in the multi-dimensional feature dataset. Process each of the six feature subsets in the multi-dimensional feature dataset one by one, matching the corresponding entries in the label mapping rule library according to the feature type, and processing each feature data in each feature subset. Each feature data point in the character feature subset is mapped to a corresponding tag item. This mapping assigns a character entity tag to each feature data point in the item information subset, an item entity tag to each feature data point in the scene description subset, a scene entity tag to each feature data point in the speech text subset, a speech text tag to each feature data point in the subtitle text subset, and a sensitive information feature subset to each sensitive tag. During the mapping process, each generated tag item retains the unique identifier of the media file corresponding to the original feature data, spatiotemporal association information, and core feature content, ensuring that the tag item is traceable to the original feature data. After completing the tag mapping operation for all feature data in the multi-dimensional feature dataset, all generated character entity tags, item entity tags, scene entity tags, speech text tags, subtitle text tags, and sensitive tags are comprehensively summarized to form a preliminary tag set containing various preliminary tags. The number of tag items in this tag set is consistent with the total number of feature data points involved in the mapping.
[0032] Step 112 involves normalizing the initial tag set, identifying and merging synonymous or near-synonymous tags, eliminating conflicts between tags, unifying the tag expression format and type encoding, and generating a standardized tag item set. Specifically, this includes: identifying and merging synonymous or near-synonymous tags; first, grouping the initial tag set according to tag type, dividing six categories of tags, such as person entity tags and item entity tags, into independent tag groups; identifying synonymous or near-synonymous tags only within the same type of tag group; calculating the semantic similarity between any two tags within the same type of tag group; determining whether the two tags are synonymous or near-synonymous by semantically matching their core feature content and expression connotation; simultaneously setting a semantic similarity threshold; if the semantic matching degree of two tags reaches this threshold, they are determined to be synonymous or near-synonymous tags; merging the tags determined to be synonymous or near-synonymous, retaining the tag name that better fits the original core feature content during merging; and integrating all related information of the two tags (media file identifier, spatiotemporal information, etc.); the number of merged tags is the original number of synonymous or near-synonymous tags minus the actual number of merged tag pairs.
[0033] To eliminate conflicts between tags, the tags after merging synonyms and near-synonyms are first grouped according to the unique identifier of the media file. Then, within the group with the same media file identifier, a second grouping is performed according to the spatiotemporal stamp, so that all tags in the same media file and the same spatiotemporal interval are grouped into the same group. The content of the tags in each group is checked one by one to check for any contradictory or conflicting tags. If a conflict is found, the multi-dimensional feature dataset extracted in step 110 and the original initial analysis result set are used as the sole basis to correct the conflicting tags, either by deleting the erroneous tags or correcting the core content of the tags, to ensure that the content of tags in the same media file and the same spatiotemporal interval is consistent and conflict-free.
[0034] The standardization process involves unifying the representation format and type coding of tags. First, the representation format of all tags is standardized, establishing a fixed structure: core feature content + spatiotemporal identifier + media file identifier. All expressions use formal, written language, eliminating colloquial, fragmented, and abbreviated descriptions to ensure consistency across all tags. Second, fixed numerical prefixes are assigned to six tag categories: 1 for person entities, 2 for item entities, 3 for scene entities, 4 for voice / text tags, 5 for subtitle / text tags, and 6 for sensitive tags. The code consists of a fixed number of consecutive digits, ordered sequentially according to the tag generation order, ensuring a unique type code for each tag and preventing duplication. After completing all the above normalization operations, all processed tag items are comprehensively organized and integrated to generate a standardized tag set. Each tag item in this set possesses a unique type code, a standardized representation format, no content conflicts, no synonyms or near-synonyms, and retains complete related source information.
[0035] Step 113 involves integrating all tags belonging to the same media file from the standardized tag set to form a structured tag set. Specifically, this includes: first, reading the unique media file identifier carried by each tag in the standardized tag set; using this identifier as the classification basis, performing a full-domain classification and aggregation of the standardized tag set; integrating all tags with the same unique media file identifier into an independent tag group, with each tag group uniquely corresponding to one converged media file; and ensuring the number of tag groups after classification and aggregation matches the number of converged media files involved in the integration. Each independent tag group is then internally organized to create a standardized tag index table. This index table includes five core elements: tag number, tag type, tag core content, spatiotemporal association information, and original feature source. Information is generated by sequentially sorting tag item numbers within a tag group according to tag type to ensure quick retrieval of tags within the group. After establishing the internal index table of the tag group, a unique identifier for the media file corresponding to the tag group is created and bound to the tag item index table and all standardized tags within the group, forming a structured tag subset corresponding to a single media file. The tags within this subset are precisely associated with the corresponding media file, and the index is clear and the information is complete. If there are multiple structured tag subsets corresponding to multiple converged media files, all structured tag subsets are comprehensively summarized to form a structured tag set for the converged media content. Each structured tag subset in this set corresponds one-to-one with a single converged media file, and all tags are standardized, the index is clear, and the associated information is complete.
[0036] In this embodiment, six types of features are extracted from the initial analysis result set, and a rule-based mapping from features to tags is achieved based on a preset rule base. Then, multi-dimensional normalization processing is used to solve the problems of tag duplication, conflict, and format disorder. Finally, the structured integration of tag items is completed according to media files, avoiding the deviation in the construction of graph nodes caused by non-standard tags, and improving the accuracy and efficiency of knowledge graph topology construction.
[0037] In a preferred embodiment of the present invention, step 2 above, which involves converting the structured tag items in the structured tag set into graph entity nodes, establishing spatiotemporal relationships between the entity nodes corresponding to characters, items, and scenes as graph edges, with the media file as the root node; forming semantic association edges by semantic mapping between the entity nodes corresponding to the audio text and the subtitle text, and marking the propagation paths of the entity nodes corresponding to sensitive information features, and constructing a preliminary knowledge graph topology, may include:
[0038] In this embodiment of the invention, step 220 involves creating independent graph entity nodes in the knowledge graph for each unique person entity tag, item entity tag, scene entity tag, voice / text tag, subtitle / text tag, and sensitive tag included in the structured tag set, thereby obtaining an entity node set. Specifically, this includes: performing a full-domain, category-by-category traversal of the structured tag set; filtering out all unique tag items under each category according to the classification of person entity tags, item entity tags, scene entity tags, voice / text tags, subtitle / text tags, and sensitive tags; the filtering process uses the unique type code of the tag as the criterion, with each type code corresponding to one unique tag item, ensuring that each filtered tag item is a unique, standardized, and non-repeating structured tag item; and creating a corresponding independent graph entity in the knowledge graph for each filtered unique tag item. When a node is created, all associated attribute information of the original tag item is completely bound to the corresponding entity node. The bound attribute information includes the unique tag type code, core feature content, spatiotemporal association information, unique media file identifier, and original feature source. At the same time, a unique node identifier is assigned to each newly created entity node. This identifier is a globally unique sequence code to ensure that all entity nodes in the knowledge graph are unique and can be accurately traced back to the original tag item and original feature. After the entity nodes of all unique structured tag items are created, the generated person entity nodes, item entity nodes, scene entity nodes, voice text entity nodes, subtitle text entity nodes, and sensitive tag entity nodes are fully integrated to form a graph entity node set containing six independent node types. Each entity node in this set has complete attribute information, a unique identifier, and retains all association relationships with the structured tag items.
[0039] Step 221: Based on the entity node set and using the media file as the root node, extract the spatiotemporal information recorded in the structured tag set. Establish spatiotemporal association edges between character entity nodes, item entity nodes, and scene entity nodes that share a common spatiotemporal co-occurrence relationship, forming a preliminary entity association network centered on the root node. Specifically, this includes: creating a dedicated root node in the knowledge graph for each corresponding multimedia file in the structured tag set. Each root node is bound to the basic attribute information of the corresponding media file, including the media file's unique identifier, file format, total file duration, file storage path, and file acquisition time. The root node serves as the core association hub for all entity nodes in the corresponding media file and is the central node for all subsequent associations. Extract all character entity nodes, item entity nodes, and scene entity nodes from the entity node set, and read all the spatiotemporal association information bound to each of the three types of nodes, specifically including the appearance timestamp, disappearance timestamp, and screen coordinate range of the node's corresponding features in the media file, ensuring that the spatiotemporal information of each node is complete and verifiable. Based on the extracted spatiotemporal information, determine whether there is a common spatiotemporal co-occurrence relationship between the three types of nodes. The criteria for determining whether a node pair has a co-occurrence relationship are: the timestamp intervals of the two nodes to be determined must have a valid overlap, and the screen coordinate ranges of the two nodes must have a valid overlapping area. Only when both temporal and spatial overlap conditions are met simultaneously can the node pair be determined to have a co-occurrence relationship in spatiotemporal space. For all node pairs determined to have a co-occurrence relationship in spatiotemporal space, a spatiotemporal association edge of the knowledge graph is established between them. Each spatiotemporal association edge is labeled with complete association attribute information, including the temporal overlap duration, spatial overlap degree, and co-occurrence intensity of the node pair. The co-occurrence intensity is calculated by multiplying the temporal overlap duration by the spatial overlap degree. The spatiotemporal correlation edges are directed, and the directionality of the edges is marked according to the chronological order of the appearance of the corresponding features in the media file. After the construction of all spatiotemporal correlation edges is completed, basic correlation edges are established between the root node of each converged media file and all the person entity nodes, item entity nodes and scene entity nodes corresponding to the file. Then, the root nodes of all converged media files, the corresponding person, item and scene entity nodes, and all spatiotemporal correlation edges and basic correlation edges between the nodes are fully integrated to form a preliminary entity correlation network centered on the root node of the media file, which only contains the spatiotemporal correlation dimension.
[0040] Step 222: Based on the preliminary entity association network, perform semantic similarity calculation on the entity nodes corresponding to the speech text tags and the entity nodes corresponding to the subtitle text tags to obtain the calculation results; establish semantic association edges between entity node pairs that meet the preset semantic mapping conditions according to the calculation results, and generate an intermediate graph topology including semantic dimensions; specifically, this includes: extracting all speech text entity nodes and subtitle text entity nodes from the preliminary entity association network, grouping the two types of nodes according to the unique identifier of the media file root node, ensuring that only speech text entity nodes and subtitle text entity nodes under the same converged media file root node are used for pairwise semantic similarity calculation, and no cross-file calculation is performed between nodes of different media files; perform semantic similarity calculation on any pair of speech text entity nodes and subtitle text entity nodes in each group. The calculation process is as follows: first, extract full-domain semantic features from the core text content of the two nodes respectively, transform the unstructured text content into standardized semantic feature vectors of the same dimension, then calculate the inner product of the two semantic feature vectors, and divide the calculated inner product result by the two semantic features. The product of vector magnitudes yields the semantic similarity value for the node pair, ranging from 0 to 1. A higher value indicates a stronger semantic association between the two nodes. A fixed semantic similarity threshold is preset, and the semantic similarity calculation result for each node pair is compared with this threshold. If the result is greater than or equal to the threshold, the node pair is deemed to meet the preset semantic mapping conditions. For all speech text entity nodes and subtitle text entity node pairs that meet the semantic mapping conditions, semantic association edges of the knowledge graph are established between them. Each semantic association edge is labeled with complete association attribute information, including the semantic similarity value of the node pair and the spatiotemporal overlap interval of the node's corresponding features in the media file. After completing the construction of all semantic association edges, each semantic association edge and its corresponding speech text and subtitle text entity nodes are integrated into the preliminary entity association network. At the same time, basic association edges are established between the root node of each converged media file and all speech text entity nodes and subtitle text entity nodes corresponding to that file. After integration, an intermediate graph topology containing both spatiotemporal association dimensions and semantic association dimensions is formed.
[0041] Step 223: Based on the intermediate graph topology, identify and determine the propagation path of sensitive information according to the content of sensitive tags and the context information in the media file. Mark the attribute information of the propagation path on the corresponding sensitive tag entity node, and finally form a complete preliminary knowledge graph topology structure. Specifically, this includes: extracting all sensitive tag entity nodes from the intermediate graph topology, grouping the sensitive tag entity nodes according to the unique identifier of the root node of the media file to ensure that only sensitive tag entity nodes under the same converged media file are analyzed for propagation paths separately, reading all attribute information bound to each sensitive tag entity node one by one, including the sensitive information type, the specific content of the sensitive information, the spatiotemporal interval of the sensitive information, and the unique identifier of the corresponding media file. At the same time, retrieve the context information of the sensitive tag entity node in the media file. This context information includes all attribute information of person entity nodes, item entity nodes, scene entity nodes, voice text entity nodes, and subtitle text entity nodes that exist in the same spatiotemporal interval as the sensitive tag entity node, as well as the spatiotemporal and semantic relationships between these nodes and the sensitive tag entity node. Based on the retrieved attribute information and context information of the sensitive tag entity node, comprehensively analyze and determine the complete propagation path of sensitive information in the knowledge graph. The core elements of the propagation path are determined, including the propagation start node, propagation path nodes, propagation end node, and the corresponding spatiotemporal interval. The propagation start node is the entity node that appears for the first time and corresponds to the sensitive information; the propagation path nodes are all other entity nodes that have a spatiotemporal or semantic relationship with the sensitive information; and the propagation end node is the entity node that appears last time and corresponds to the sensitive information. After determining the propagation path, the complete attribute information of the propagation path is uniformly marked on the corresponding sensitive label entity nodes. The marked attribute information includes a unique identifier for the propagation path, a unique identifier for the propagation start node, and a unique identifier for the propagation path nodes. The system includes an identifier set, a unique identifier for the propagation endpoint node, the complete spatiotemporal interval corresponding to the propagation, and the propagation association type. The propagation association type is divided into two categories: spatiotemporal association propagation and semantic association propagation. After marking the propagation path attribute information of all sensitive label entity nodes in the intermediate graph topology, the system integrates all media file root nodes, six types of entity nodes, all spatiotemporal association edges, all semantic association edges, all basic association edges, and all propagation path attribute information marked on sensitive label entity nodes in the graph. This forms a complete preliminary knowledge graph topology structure that simultaneously includes both spatiotemporal and semantic association dimensions and completes the marking of sensitive information propagation paths.
[0042] In this embodiment, standardized structured tags are transformed into knowledge graph entity nodes with complete attributes. A spatiotemporal association network of people, objects, and scenes is built with media files as the root node. Then, semantic similarity calculation is used to construct semantic association edges between audio text and subtitle text. Finally, the propagation path of sensitive information is sorted out and marked by combining contextual information, thereby improving the deep association analysis and mining capabilities of converged media data.
[0043] In a preferred embodiment of the present invention, step 3 above involves mapping the multi-source heterogeneous data of entity nodes and graph edges in the preliminary knowledge graph topology to high-dimensional feature vectors to construct a multi-dimensional association model. Feature intervals are divided according to the spatial distribution density of the feature vectors. Media data is defined as an observation sequence, and semantic data is defined as a state sequence. The current state sequence is predicted using preset node weights to obtain a predicted value. A gain coefficient is calculated based on the current feature interval density. The gain coefficient is then used to weight and fuse the deviation between the observation sequence and the predicted value to update the node weights, resulting in an optimized knowledge graph. This optimized knowledge graph may include:
[0044] In this embodiment of the invention, step 330 involves extracting attribute information of each entity node and association weights of graph edges from the preliminary knowledge graph topology, encoding and fusing the attribute information and association weights, and mapping them to an initial high-dimensional feature vector of a unified dimension. Specifically, this includes: performing a full-domain, node-by-node, and edge-by-edge traversal of the preliminary knowledge graph topology, extracting complete attribute information for each entity node, including unique node identifiers, core tag features, spatiotemporal association information, node type encoding, sensitive attribute markers (exclusive to sensitive tag entity nodes), and original feature sources; simultaneously extracting complete association weights for all graph edges, including the temporal overlap duration, spatial overlap, and co-occurrence strength of spatiotemporal association edges, the semantic similarity value and spatiotemporal overlap interval of semantic association edges, and the default association coefficient of basic association edges; and performing standardized numerical encoding processing on the extracted entity node attribute information, classifying attributes... One-hot encoding is used to convert the data into fixed-length binary numerical vectors. Continuous spatiotemporal stamps, image coordinates, and other attributes are converted into standardized values in the 0-1 range using min-max normalization. Node unique identifiers and type codes are converted into continuous ordered values using sequential encoding. The core feature content of the tags is converted into numerical vectors after semantic feature extraction. Sensitive attribute markers are encoded as 1 / 0 depending on their presence or absence. The extracted graph edge association weights are directly converted into standardized values in the 0-1 range using min-max normalization to eliminate the interference of different units and numerical ranges on subsequent calculations. The encoded entity node attribute information numerical vectors are then concatenated and fused with the normalized association weights of all graph edges corresponding to that node, sequentially according to their dimensions. All the concatenated and fused vectors are uniformly mapped to numerical vectors of the same dimension, which is the initial high-dimensional feature vector.
[0045] Step 331: Based on the initial high-dimensional feature vectors, calculate the Euclidean distance and cosine similarity of the vectors in the feature space to construct a multi-dimensional association model reflecting the spatiotemporal association strength and semantic mapping depth between nodes. Specifically, this includes: for any two initial high-dimensional feature vectors in the initial high-dimensional feature vector set, calculate the Euclidean distance and cosine similarity between them. The Euclidean distance calculation process involves first performing difference operations on the feature values of the corresponding dimensions of the two vectors sequentially, then squaring the difference results of all dimensions, summing all the squared results, and finally taking the square root of the summation result. The numerical value represents the Euclidean distance between two vectors. A larger value indicates a greater difference in spatial features between the two vectors in the feature space. The calculation process for cosine similarity is as follows: First, multiply the feature values of the corresponding dimensions of the two vectors in turn and sum all the product results to obtain the inner product of the two vectors. Then, sum the squares of all feature values of the two vectors in each dimension and take the square root to obtain the magnitude of the two vectors. Finally, divide the inner product result by the product of the magnitudes of the two vectors. The resulting value is the cosine similarity, which ranges from 0 to 1. A larger value indicates a higher degree of semantic association and matching between the two vectors.
[0046] The multidimensional association model is a supervised feature association analysis model. The model adopts a fully connected network structure and consists of three layers: an input layer, a hidden layer, and an output layer. The input data of the input layer is the initial high-dimensional feature vector and the Euclidean distance and cosine similarity value calculated between this vector and other vectors. The hidden layer consists of three fully connected layers. The number of neurons in the first fully connected layer is half the number of the dimension of the input layer, the second layer is half the number of the first layer, and the third layer is half the number of the second layer. Each hidden layer uses a linear rectified activation function to achieve non-linear mapping of features. The output layer has two output neurons, which correspond to the spatiotemporal association strength and semantic mapping depth between entity nodes, respectively. The output values are mapped to the 0-1 interval, representing the tightness of the spatiotemporal association between nodes and the depth of the semantic mapping, respectively.
[0047] We selected manually annotated multimedia knowledge graph data to construct the model training set and test set, which were randomly divided in a 7:3 ratio. The input data for both the training and test sets consisted of the initial high-dimensional feature vectors, Euclidean distance, and cosine similarity corresponding to the annotated data. The output data consisted of the true values of the spatiotemporal association strength and semantic mapping depth between manually annotated nodes. The mean squared error was used as the loss function of the model. This function was calculated by squaring the difference between the model's predicted value and the manually annotated true value, and then averaging the squared results over all samples. The smaller the loss function value, the higher the model's predictive accuracy. The network parameters of the model are iteratively optimized using stochastic gradient descent. The model learning rate is set to 0.01 and the batch size is 32. In each iteration, the training set data is input into the model in batches, the loss function value is calculated, and the network parameters are updated along the gradient descent direction. The training continues iteratively until the loss function value converges to a preset threshold (≤0.001). After training is completed, the test set data is input into the model to verify its generalization ability. When the prediction accuracy of the model on the test set is ≥95%, the model training is considered complete. The trained multidimensional association model can reflect the spatiotemporal association strength and semantic mapping depth between entity nodes.
[0048] Step 332: For each target high-dimensional feature vector in the multidimensional association model, a local neighborhood range centered on the target high-dimensional feature vector is defined. The number of non-self feature vectors falling within this local neighborhood range is counted, and the ratio of this number to the volume of the local neighborhood range is used as the spatial distribution density value of the target high-dimensional feature vector position, thus obtaining the density distribution dataset corresponding to all high-dimensional feature vectors. Specifically, this includes: for each target high-dimensional feature vector in the multidimensional association model, a hypersphere local neighborhood range with the vector as its geometric center is defined, and the radius of the hypersphere is a preset fixed value. Ensure that the local neighborhood size of all target vectors is consistent; the formula for calculating the volume of the local neighborhood of the hypersphere is: ,in for The volume of a hypersphere Let be the uniform dimension number of the initial high-dimensional feature vectors. For gamma function, A predetermined radius is set for the hypersphere; the number of non-self initial high-dimensional feature vectors falling within this local neighborhood is counted, denoted as . The statistically obtained quantity Volume of the local neighborhood Perform a ratio calculation, i.e., spatial distribution density value = quantity. ÷ Neighborhood volume This ratio is the spatial distribution density value of the target high-dimensional feature vector at the corresponding position in the feature space. The spatial distribution density values of all initial high-dimensional feature vectors in the multidimensional association model are calculated one by one according to the above method. All density values are associated and integrated according to the unique vector identifier to form a density distribution dataset containing all vector density features.
[0049] Step 333 involves analyzing the numerical trend of the density distribution dataset, identifying peak points where density values jump from low to high and trough points where they fall from high to low, and dividing the continuous feature space into several feature intervals with different density thresholds based on these peak and trough points. Specifically, this includes: sorting all spatial distribution density values in the density distribution dataset in ascending order; plotting a density value curve with the position of the feature vector in the high-dimensional feature space as the horizontal axis and the spatial distribution density value as the vertical axis; visually presenting the distribution pattern of density values in the feature space; performing a global and segment-by-segment trend analysis on this curve; and identifying peak and trough points in the curve. Peak points are those where density values jump from low to high in the curve. The peak and valley points are the local maximum points reached after the rise, and the valley points are the local minimum points where the density value in the curve falls from high to low. The identification of peak and valley points is based on the change of the first derivative of the curve. The point where the first derivative changes from positive to negative is the peak point, and the point where the first derivative changes from negative to positive is the valley point. This ensures the accuracy of peak and valley point identification. All the identified valley points are used as natural dividing points of the feature space. The continuous high-dimensional feature space corresponding to the preliminary knowledge graph is divided into several independent and non-overlapping feature intervals. Each feature interval corresponds to a unique density threshold range. The spatial distribution density values of all initial high-dimensional feature vectors within the interval have significant similarities, and there are obvious step differences in density values between different intervals. This realizes the division of the feature space according to density features.
[0050] Step 334: For each segmented feature interval, extract the original acquired data stream of the media file corresponding to the feature vector within that feature interval as an observation sequence, and simultaneously extract the semantic logic evolution data stream between entity nodes within that feature interval as a state sequence. Specifically, this includes: performing individual, full-scale processing on each feature interval; tracing back the original acquired data stream of the corresponding converged media file based on the unique identifier of all initial high-dimensional feature vectors within the interval; this original acquired data stream includes original video frame acquired data, original audio sampling data, original subtitle acquired data, and original character / object / scene recognition data, and all data carries the global timestamp of the media file; sorting the original acquired data stream of the converged media file corresponding to the feature interval in order from earliest to latest global timestamp; the sorted complete data stream is the feature interval. The media data observation sequence corresponding to the interval; at the same time, based on the unique identifier of the initial high-dimensional feature vector within the interval, the entity nodes in the corresponding preliminary knowledge graph topology are traced backward, and the semantic logic evolution data stream between these entity nodes as the global timestamp of the media file changes is extracted. This data stream contains the establishment, enhancement, weakening, and disappearance of semantic relationships between nodes, as well as the core features such as semantic logic progression, turning point, and connection driven by spatiotemporal correlation. All semantic logic evolution data carries the corresponding global timestamp. The semantic logic evolution data stream is ordered in order from early to late according to the global timestamp. The ordered complete data stream is the semantic data state sequence corresponding to the feature interval, ensuring that the media data observation sequence and the semantic data state sequence within the same feature interval correspond one-to-one on the global timestamp and are completely synchronized in the time dimension.
[0051] Step 335: Based on the feature interval and the corresponding state sequence within that feature interval, the preset node weights are used to perform forward propagation calculations on the current state sequence to obtain state prediction values representing the semantic logic evolution trend. Specifically, this includes: reading the complete semantic data state sequence corresponding to the feature interval to be processed, confirming the timestamp range and data dimension of the sequence, and simultaneously retrieving the system's preset initial weights for entity nodes. These preset node weights are pre-set according to the feature importance of the converged media content. Key entity nodes such as people, scenes, and core items have higher initial weights than other auxiliary entity nodes, and the preset weight values of all entity nodes are within the range of 0-1. A higher weight value indicates greater importance in the semantic logic evolution trend. The greater the importance of change, the more the semantic data state sequence is calculated by progressively forward propagating the calculation using the global timestamp of the media file as the minimum calculation step. The value of the semantic data state sequence at the previous moment is weighted by the preset weight of the corresponding entity node at the current moment, and then the weighted results of all entity nodes are summed. That is, the current state prediction value = the state sequence value at the previous moment × the sum of the preset weights of the corresponding entity nodes. Following this calculation method, the calculation is progressively advanced from the global timestamp from early to late to complete the forward propagation calculation of the entire time axis within the feature interval, and finally the state prediction value corresponding to the full timestamp within the feature interval is obtained. This prediction value can characterize the natural evolution trend of the semantic logic between entity nodes.
[0052] Step 336: Compare the predicted state value with the media data observation sequence within the same feature interval to obtain the original deviation vector. Based on the original deviation vector, obtain the entity nodes to be updated and their current connection weights in the preliminary knowledge graph topology. Specifically, this includes: comparing the predicted state value with the media data observation sequence within the same feature interval point by point and time by time according to the global timestamp; performing a difference operation between the media data observation sequence value and the predicted state value corresponding to each timestamp, i.e., single timestamp deviation value = media data observation sequence value - predicted state value. A positive difference indicates that the observed value is higher than the predicted value, and a negative difference indicates that the observed value is lower than the predicted value; and calculating the single timestamp deviation values corresponding to all timestamps within the feature interval according to the global timestamp from earliest to latest. The data are integrated sequentially from evening to night to form a numerical vector whose dimensions are completely consistent with the semantic data state sequence. This vector is the original deviation vector. The original deviation vector intuitively reflects the degree and change pattern of the deviation between the state prediction value and the actual observation value. Based on the characteristics of each dimension of the original deviation vector, the corresponding entity nodes in the preliminary knowledge graph topology are traced back. All entity nodes that are directly related to the dimensions of the original deviation vector are the entity nodes in the preliminary knowledge graph that need to be updated in terms of weight. At the same time, from the preliminary knowledge graph topology, all the current graph edge connection weights between all entity nodes to be updated are extracted one by one, including the co-occurrence strength of spatiotemporal related edges, the semantic similarity value of semantic related edges, and the default association coefficient of basic related edges.
[0053] Step 337: Based on the numerical distribution characteristics of the original deviation vector and combined with the spatial distribution density value corresponding to the current feature interval, analyze the dispersion of the spatial distribution density value relative to the global density distribution. Transform the dispersion into an adjustment amplitude parameter and apply this parameter to the original deviation vector to quantify the reliable fluctuation range of the original deviation vector under the current data density environment, thus obtaining the dynamic gain coefficient. Specifically, this includes: analyzing the overall numerical distribution characteristics of the original deviation vector, calculating the mean and standard deviation of the original deviation vector, where the mean reflects the overall offset trend of the deviation, and the standard deviation reflects the overall fluctuation amplitude of the deviation. The mean and standard deviation together characterize the overall characteristics of the original deviation vector. Calculate the dispersion of the spatial distribution density value of the current feature interval relative to the global density distribution, and use the coefficient of variation to characterize this dispersion. The coefficient of variation is calculated by dividing the standard deviation of the global density distribution by the mean of the global density distribution. The larger the coefficient of variation, the greater the difference between the density characteristics of the current feature interval and the global density distribution. The calculated coefficient of variation is directly converted into an adjustment amplitude parameter. The adjustment amplitude parameter is positively correlated with the coefficient of variation; the larger the coefficient of variation, the larger the adjustment amplitude parameter. This adjustment amplitude parameter is applied to the original deviation vector, and the standard deviation of the original deviation vector is normalized to convert it into a variance normalized value in the 0-1 interval, eliminating the difference in deviation fluctuation amplitude between different feature intervals. The adjustment amplitude parameter and the variance normalized value are multiplied, i.e., dynamic gain coefficient = adjustment amplitude parameter × variance normalized value. This dynamic gain coefficient can accurately quantify the reliable fluctuation range of the original deviation vector under the current feature interval density environment.
[0054] Step 338 involves scaling the original deviation vector using a dynamic gain coefficient to obtain a weighted fusion deviation that includes density adaptive characteristics. This weighted fusion deviation is then fed back as a correction signal to the preliminary knowledge graph topology to update the current weights of each entity node and the association strength of graph edges, resulting in an optimized knowledge graph. Specifically, this includes: scaling the original deviation vector element-wise using a dynamic gain coefficient, where each single-timestamp deviation value in the original deviation vector is multiplied by the dynamic gain coefficient. The resulting vector after scaling all elements is the weighted fusion deviation that includes density adaptive characteristics. This deviation eliminates the interference of different density feature intervals on the deviation, better reflecting the actual semantic logic evolution of converged media data; and feeding the weighted fusion deviation back to the preliminary knowledge graph topology as a correction signal to update the current weights of all entity nodes to be updated. The weights are adjusted by setting the new node weight to the average of the original node weight and the weighted fusion deviation, ensuring that the entity node weights can be dynamically adjusted according to the actual characteristics of the converged media data, and that the weight values are more in line with the actual value of the data. At the same time, the correlation strength of all graph edges between the entity nodes to be updated is updated collaboratively, with the new correlation strength being equal to the original connection weight multiplied by (1 + the average of the weighted fusion deviation). This achieves synchronous optimization of graph edge correlation strength and entity node weights, ensuring that the relationship between nodes and edges is more accurate. After updating the weights of all entity nodes to be updated and the correlation strength of corresponding graph edges in the initial knowledge graph topology, all media file root nodes, six types of entity nodes, all spatiotemporal correlation edges, semantic correlation edges, basic correlation edges, and the propagation path attribute information marked on sensitive tag entity nodes in the graph are fully integrated to obtain the optimized knowledge graph.
[0055] In this embodiment, by transforming the multi-source heterogeneous data of the preliminary knowledge graph into high-dimensional feature vectors of a unified dimension, a dedicated multi-dimensional association model is constructed and trained, realizing the quantification of the spatiotemporal association strength and semantic mapping depth between entity nodes; by dividing feature intervals according to the spatial distribution density of feature vectors and defining observation sequences and state sequences, the knowledge graph can adapt to the diverse feature distribution characteristics of converged media data; at the same time, relying on the dynamic gain coefficient, the density adaptive update of entity node weights and graph edge association strength is realized, improving the intelligent utilization efficiency of converged media resources in complex application scenarios.
[0056] In a preferred embodiment of the present invention, step 4 above, which involves storing the optimized knowledge graph in a distributed database, constructing a hierarchical retrieval system for role permissions based on the stored graph data, and executing intelligent recommendation logic to realize the management and application of media resources, may include:
[0057] In this embodiment of the invention, step 440 involves serializing and encapsulating the updated entity node weights and graph edge association strengths in the optimized knowledge graph, storing them in a distributed database to form a graph data storage set for concurrent access. Specifically, this includes: performing a detailed traversal of the optimized knowledge graph across the entire domain, node by node, and edge by edge, extracting the core data of all entity nodes that have completed weight updates and the core data of all graph edges that have completed association strength updates. The core data of entity nodes includes a unique node identifier, label type, core feature content, spatiotemporal association information, updated node weights, sensitive attribute markers, and propagation path attribute information specific to sensitive label nodes. The core data of graph edges includes a unique edge identifier, identifiers of associated entity node pairs, updated spatiotemporal association strength, updated semantic mapping depth, basic association coefficient, and the spatiotemporal interval corresponding to the association. During the extraction process, all redundant and invalid data are removed to ensure that only optimized and valid core data is retained.
[0058] The extracted core data of entity nodes and graph edge core data are subjected to structured serialization and encapsulation. All core data of a single entity node are integrated and encapsulated into an independent node data unit, and all core data of a single graph edge are integrated and encapsulated into an independent edge data unit. At the same time, a unique association mapping identifier is established between each node data unit and its corresponding associated edge data unit. This identifier is consistent with the unique identifier of the entity node pair, ensuring that the association relationship between nodes and edges is not lost after serialization and encapsulation. All encapsulated node data units and edge data units are uniformly converted into a binary serialization format that can be directly parsed, read, and stored by the distributed database, ensuring the efficiency and integrity of data during transmission and storage.
[0059] The binary serialized node and edge data units are stored according to the sharding rules of the distributed database. The unique identifier of the media file is selected as the sharding key. The data is evenly sharded according to the type of the converged media file and the acquisition time to achieve load balancing of the database storage and avoid the impact of single shard data overload on access efficiency. During the storage process, a global retrieval identifier is established for each node and edge data unit. This identifier corresponds one-to-one with the unique identifier of the node and edge, supporting the rapid location and retrieval of data in subsequent retrieval processes. After all the core data units of the optimized knowledge graph are distributed and stored, all node and edge data units in the distributed database, as well as all node-edge association mapping relationships, are integrated to form a graph data storage set that supports simultaneous access by multiple users and data conflict-free access.
[0060] Step 441: Based on the graph data storage set, extract sensitive tag attributes and propagation path information from entity nodes. Combined with preset user role definitions, construct a hierarchical retrieval index system including different access levels and data filtering rules. Specifically, this includes: traversing and filtering all entity node data units in the graph data storage set, locating all sensitive tag entity nodes, and extracting complete sensitive tag attribute information and propagation path information for these nodes. The sensitive tag attribute information includes the sensitive information type, the specific content of the sensitive information, the spatiotemporal range in which the sensitive information appears in the media file, and the identifiers of associated non-sensitive entity nodes. The propagation path information includes the identifiers of the starting node, the nodes along the way, the node at the end, the type of propagation association (spatiotemporal association propagation or semantic association propagation), and the complete spatiotemporal range covered by the propagation. At the same time, according to the compliance requirements of converged media content management, the extracted sensitive information is divided into three levels according to the degree of sensitivity: Level 1 is the highest sensitivity level, Level 2 is the medium sensitivity level, and Level 3 is the lowest sensitivity level.
[0061] The system retrieves user role definition rules pre-defined according to the needs of converged media business management. Based on the work permissions and content access requirements of different positions, system users are divided into three categories: public access roles, restricted access roles, and core control roles. Corresponding to these three user roles, the hierarchical retrieval system is divided into three levels of access permissions: public access layer, restricted access layer, and core control layer. The public access layer only allows access to entity nodes and related graph edges without sensitive attributes; the restricted access layer allows access to level 3 and level 2 sensitive information and all graph data without sensitive information; the core control layer allows access to all levels of sensitive information and all graph data in the graph data storage set.
[0062] Based on the classification of sensitive information levels and the access permission range, four types of graph data filtering rules are formulated to ensure that users of different roles can only access graph data within their authorized access range. These rules are: 1) Node blocking rules based on sensitive attributes, where sensitive tag entity nodes exceeding the user's permission level are directly blocked when a lower-level user initiates a search request; 2) Truncating rules based on propagation paths, where the propagation path of sensitive information is selectively truncated, retaining only propagation path nodes and related edges that match the user's permission level, and eliminating paths exceeding the user's permission level; 3) Threshold filtering rules based on association strength, where a minimum judgment standard for spatiotemporal association strength and semantic mapping depth is set, filtering out weakly associated entity nodes and graph edges below this standard, retaining only strongly associated valid data; and 4) Filtering rules based on spatiotemporal sequences, where only graph data matching the spatiotemporal filtering conditions set in the user's search request is retained, and nodes and edges that do not match the spatiotemporal range are eliminated.
[0063] Independent search indexes are built for the three access levels: public access, restricted access, and core control. Each search index uses unique entity node identifiers, tag types, spatiotemporal association information, association strength, and sensitivity level as core search fields. Four types of data filtering rules are embedded into the corresponding access level's search index to achieve automatic data filtering during the search process. An automatic matching mapping relationship between user roles and access levels is established. When a user initiates a search request, the system can automatically match the corresponding access level and search index based on the user's role identifier. By integrating the three independent search indexes, the matching mapping relationship between user roles and access levels, and the full data filtering rules, a complete and logically clear hierarchical search index system is formed, ensuring that user role permissions and graph data can be matched during the search process.
[0064] Step 442: Utilize the hierarchical retrieval index system to respond to user retrieval requests. Based on the user's role and permissions, dynamically trim and reorganize entity nodes and associated edges in the graph data storage set to obtain candidate resource subgraphs that conform to the current permission scope. Specifically, this includes: after receiving a user's retrieval request, the system performs a comprehensive and detailed analysis of the request, extracting all retrieval elements such as core keywords, spatiotemporal filtering conditions, content type filtering conditions, and resource format requirements. Simultaneously, it reads the user's role identifier and, through the pre-defined matching mapping relationship between user roles and access levels in the hierarchical retrieval index system, matches the user with the corresponding access level. Simultaneously, it retrieves the retrieval index library and data filtering rules corresponding to that access level to clarify the user's permission scope and the boundaries of accessible graph data.
[0065] Image and video entity nodes that initially match the search elements are extracted from the graph data storage set. These nodes correspond to image frames and video keyframes in the original content of the converged media. Feature matching is performed on these nodes using a shape rotation symmetry detection algorithm to filter out effective nodes and associated edges that highly match the search elements. The shape contour features of the target in the original content of the converged media corresponding to these entity nodes are extracted and transformed into a set of two-dimensional coordinate points. Then, the geometric center of the target shape contour is calculated and used as the shape rotation symmetry center. Subsequently, the rotation angle detection range is set to 0 to 360 degrees, and the two-dimensional coordinate point set is rotated around the rotation symmetry center with a fixed step size to obtain a new two-dimensional coordinate point set corresponding to each rotation angle. The shape similarity of the two-dimensional coordinate point set before and after rotation is calculated, and a judgment threshold for shape similarity is set. Entity nodes with similarity reaching the threshold are judged as valid entity nodes that satisfy rotation symmetry, i.e., highly match the search elements. Finally, all the selected valid entity nodes are marked, and all graph edges that have an association with the valid entity nodes are also marked to form a valid node-edge set.
[0066] Strictly adhering to the four types of data filtering rules corresponding to the user's access level, a comprehensive dynamic pruning operation is performed on the selected valid node-edge set. First, a node blocking rule based on sensitive attributes is executed to block all sensitive tag entity nodes and associated edges that exceed the user's permission level. Next, a truncation rule based on propagation path is executed to truncate the propagation path of sensitive information, retaining only propagation path nodes and edges within the permission range. Then, a threshold filtering rule based on association strength is executed to remove all weakly associated nodes and edges whose spatiotemporal association strength and semantic mapping depth are lower than the judgment criteria. Finally, a filtering rule based on spatiotemporal sequence is executed to delete all nodes and edges that do not match the spatiotemporal filtering conditions of the user's search request. During the pruning process, each node and edge is verified one by one to ensure that the remaining nodes and edges all meet the user's role permission range and search element requirements.
[0067] After dynamic pruning, all valid entity nodes and graph edges are completely reorganized according to the spatiotemporal relationships and semantic mapping relationships in the original optimized knowledge graph. This restores the original association logic between nodes and edges, ensuring that the association relationships of the reorganized subgraph are clear and complete. Each reorganized subgraph is assigned a unique identifier. At the same time, the updated weights of all entity nodes, the updated association strengths of all graph edges, and all core data such as spatiotemporal information, core feature content, and tag types of nodes and edges are fully preserved. Finally, a candidate resource subgraph that conforms to the user's current permission scope is formed. This subgraph only contains the association information of converged media resources that highly match the user's search elements and that the user has the right to access.
[0068] Step 443: Based on the spatiotemporal correlation strength and semantic mapping depth between entity nodes in the candidate resource subgraph, calculate the matching degree between the user's past behavior trajectory and the features of the candidate resource subgraph, forming an intelligent recommendation list sorted by matching degree to realize the management and application of media resources; specifically, this includes: performing a comprehensive and detailed analysis of each candidate resource subgraph, extracting the core features in the subgraph, including the updated weights of all entity nodes, the spatiotemporal correlation strength between nodes, and the semantic mapping depth between nodes, and standardizing and quantifying all extracted core features, uniformly converting feature values of different dimensions and ranges into comparable standardized values, ensuring that each core feature can participate in the subsequent matching degree calculation in the same dimension, while retaining the original attributes and relative differences of the features.
[0069] The system retrieves all past behavioral data of the user, including search records, browsing records, collection records, download records, duration of operation on various integrated media resources, and resource forwarding records. Layered feature extraction is performed on this behavioral data. First, the frequency of user operations on integrated media resources of different tag types, content types, and formats is statistically analyzed. Second, the proportion of user's time spent on various integrated media resources relative to the total time spent on operations is calculated. Then, corresponding weight coefficients are assigned to different behavior types according to the needs of integrated media business management. The weight coefficients for deep operations such as collection, download, and forwarding are higher than those for shallow operations such as search and browsing. Combining operation frequency, duration of operation, and corresponding weight coefficients, the user's past behavioral characteristics are standardized and quantified, forming a complete set of user behavioral features.
[0070] Based on the extracted and quantified user behavior feature set and the core features of candidate resource subgraphs, the overall fit between the user's past behavior trajectory and each candidate resource subgraph is comprehensively calculated from three dimensions: spatiotemporal correlation fit, semantic mapping fit, and entity node importance fit. In the calculation process, the matching situation of the three dimensions is comprehensively considered, and the consideration weight of each dimension is set according to the application needs of converged media resources. Finally, the overall matching degree between each candidate resource subgraph and the user's past behavior trajectory is obtained. The matching degree value intuitively reflects the degree of fit between the candidate resource subgraph and the user's needs.
[0071] All candidate resource sub-graphs are sorted from highest to lowest based on their overall matching degree with users' past behavior patterns. Then, the integrated media resources corresponding to the sorted candidate resource sub-graphs are integrated with the core information of the resources, including resource name, content type, core features, storage path, and related resource information. This generates an intelligent recommendation list sorted by matching degree. At the same time, it provides multi-dimensional and refined management functions for the integrated media resource management terminal, including resource access statistics, resource association analysis, dynamic adjustment of user permissions, resource library update and maintenance, and monitoring of the effectiveness of sensitive information control. This enables intelligent management of integrated media resources and their practical application in real business scenarios.
[0072] In this embodiment, by serializing and distributing the optimized knowledge graph, load balancing and high-concurrency access by multiple users are achieved for the graph data; by combining sensitive information and user roles to construct a hierarchical retrieval index system, access control of converged media resources is realized; and by integrating a shape rotation symmetry detection algorithm in the retrieval request response stage, feature matching of entity nodes of image and video converged media resources is realized, thereby improving the matching degree between candidate resource subgraphs and user retrieval needs.
[0073] like Figure 2 As shown, embodiments of the present invention also provide a system for intelligent tag generation and knowledge graph construction of converged media content, including:
[0074] The module is used to map the features of people, items, scene descriptions, voice text, subtitle text and sensitive information in the initial analysis result set into structured data, form structured label items, and summarize the structured label items into a structured label set.
[0075] The topology module is used to convert structured tag items in the structured tag set into graph entity nodes. With the media file as the root node, it establishes the spatiotemporal relationship between the corresponding entity nodes of people, items and scenes as graph edges; it forms semantic relationship edges by semantic mapping between the corresponding entity nodes of voice text and subtitle text, and marks the propagation path of the entity nodes corresponding to sensitive information features, thus constructing a preliminary knowledge graph topology structure.
[0076] The calibration module is used to map multi-source heterogeneous data of entity nodes and graph edges in the initial knowledge graph topology into high-dimensional feature vectors to construct a multi-dimensional association model. It divides feature intervals according to the spatial distribution density of feature vectors, defines media data as observation sequences, and defines semantic data as state sequences. It uses preset node weights to predict the state sequence at the current time to obtain the predicted value, and calculates the gain coefficient in combination with the current feature interval density. It uses the gain coefficient to weight and fuse the deviation between the observation sequence and the predicted value to update the node weights, thus obtaining the optimized knowledge graph.
[0077] The application module is used to store the optimized knowledge graph into a distributed database, build a hierarchical retrieval system based on role permissions based on the stored graph data, and execute intelligent recommendation logic to realize the management and application of media resources.
[0078] It should be noted that this system is a system corresponding to the above method. All implementation methods in the above method embodiments are applicable to this embodiment and can achieve the same technical effect.
[0079] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A method for intelligent tag generation and knowledge graph construction of converged media content, characterized in that, The method includes: Step 1: Map the features of people, items, scene descriptions, voice text, subtitles and sensitive information in the initial analysis result set into structured data to form structured label items, and summarize the structured label items into a structured label set. Step 2: Convert the structured tag items in the structured tag set into graph entity nodes. With the media file as the root node, establish the spatiotemporal relationship between the corresponding entity nodes of people, items and scenes as graph edges; form semantic relationship edges by semantic mapping between the corresponding entity nodes of voice text and subtitle text, and mark the propagation path of the entity nodes corresponding to sensitive information features to construct a preliminary knowledge graph topology. Step 3: Map the multi-source heterogeneous data of entity nodes and graph edges in the preliminary knowledge graph topology to high-dimensional feature vectors to construct a multi-dimensional association model. Divide the feature intervals according to the spatial distribution density of the feature vectors. Define media data as observation sequences and semantic data as state sequences. Predict the state sequence at the current moment using preset node weights to obtain the predicted value. Calculate the gain coefficient based on the density of the current feature interval. Use the gain coefficient to weight and fuse the deviation between the observation sequence and the predicted value to update the node weights, resulting in an optimized knowledge graph. This includes the state sequence at the current moment based on the feature interval and the corresponding state sequence within that feature interval. Call the preset node weights to perform forward propagation calculation on the state sequence at the current moment to obtain the state prediction value representing the semantic logic evolution trend. Compare the state prediction value with the media data observation sequence within the same feature interval to obtain the original... The deviation vector is used to obtain the entity nodes to be updated and their current connection weights in the preliminary knowledge graph topology. Based on the numerical distribution characteristics of the original deviation vector and the spatial distribution density value corresponding to the current feature interval, the dispersion of the spatial distribution density value relative to the global density distribution is analyzed. The dispersion is transformed into an adjustment amplitude parameter, which is then applied to the original deviation vector to quantify the reliable fluctuation range of the original deviation vector under the current data density environment, thus obtaining a dynamic gain coefficient. The original deviation vector is scaled using the dynamic gain coefficient to obtain a weighted fusion deviation amount that includes density adaptive characteristics. This weighted fusion deviation amount is used as a correction signal to feed back to the preliminary knowledge graph topology to update the current weights of each entity node and the association strength of the graph edges, resulting in an optimized knowledge graph. Step 4: Store the optimized knowledge graph in a distributed database, construct a hierarchical retrieval system based on role permissions based on the stored graph data, and execute intelligent recommendation logic to realize the management and application of media resources.
2. The method for generating intelligent tags and constructing knowledge graphs for converged media content according to claim 1, characterized in that, Before step 1, collect and analyze the data of people, audio text, items and subtitles in the converged media files to obtain an initial analysis result set.
3. The method for generating intelligent tags and constructing knowledge graphs for converged media content according to claim 2, characterized in that, The initial analysis results set includes features of people, items, scenes, audio text, subtitles, and sensitive information. These features are mapped to structured data, forming structured label items. These structured label items are then aggregated into a structured label set, including: From the initial analysis results set, features of people, items, scene descriptions, voice text, subtitles, and sensitive information were extracted to obtain a multi-dimensional feature dataset; Based on the preset label mapping rule library, each type of feature in the feature dataset is mapped to the corresponding label item, the person feature is mapped to the person entity label, the item information is mapped to the item entity label, the scene description is mapped to the scene entity label, the voice text is mapped to the voice text label, the subtitle text is mapped to the subtitle text label, and the sensitive information feature is mapped to the sensitive label, generating a preliminary label set including various preliminary labels. The initial tag set is normalized, synonymous or near-synonymous tags are identified and merged, conflicts between tags are eliminated, the tag expression format and type encoding are unified, and a standardized tag item set is generated. All tags belonging to the same media file in the standardized tag set are integrated to form a structured tag set.
4. The method for generating intelligent tags and constructing knowledge graphs for converged media content according to claim 3, characterized in that, The structured tags in the structured tag set are converted into graph entity nodes. Using the media file as the root node, spatiotemporal relationships between corresponding entity nodes representing people, items, and scenes are established as graph edges. Semantic mappings between the entity nodes representing audio text and subtitle text are formed into semantic relationship edges. The propagation paths of entity nodes corresponding to sensitive information features are marked, constructing a preliminary knowledge graph topology, including: Each unique person entity tag, item entity tag, scene entity tag, voice text tag, subtitle text tag, and sensitive tag included in the structured tag set is created as an independent graph entity node in the knowledge graph, resulting in an entity node set; Based on the set of entity nodes and with the media file as the root node, the spatiotemporal information recorded in the set of structured tags is extracted. Spatiotemporal association edges are established between character entity nodes, item entity nodes and scene entity nodes that have a common spatiotemporal co-occurrence relationship, forming a preliminary entity association network centered on the root node. Based on the preliminary entity association network, semantic similarity is calculated between entity nodes corresponding to speech text tags and entity nodes corresponding to subtitle text tags to obtain the calculation results; based on the calculation results, semantic association edges are established between entity node pairs that meet the preset semantic mapping conditions to generate an intermediate graph topology including semantic dimensions. Based on the intermediate graph topology, the propagation path of sensitive information is identified and determined according to the content of sensitive tags and the context information in the media file. The attribute information of the propagation path is marked on the corresponding sensitive tag entity node, and finally a complete preliminary knowledge graph topology structure is formed.
5. The method for generating intelligent tags and constructing knowledge graphs for converged media content according to claim 4, characterized in that, Multi-source heterogeneous data of entity nodes and graph edges in the preliminary knowledge graph topology are mapped to high-dimensional feature vectors to construct a multi-dimensional association model. Feature intervals are divided according to the spatial distribution density of feature vectors. Media data is defined as observation sequences, and semantic data is defined as state sequences, including: The attribute information of each entity node and the association weight of the graph edge are extracted from the preliminary knowledge graph topology. The attribute information and association weight are encoded and fused to map into an initial high-dimensional feature vector of a unified dimension. Based on the initial high-dimensional feature vector, the Euclidean distance and cosine similarity of the vector in the feature space are calculated to construct a multi-dimensional association model that reflects the spatiotemporal association strength and semantic mapping depth between nodes. For each target high-dimensional feature vector in the multidimensional association model, a local neighborhood range centered on the target high-dimensional feature vector is defined. The number of non-self feature vectors falling within the local neighborhood range is counted, and the ratio of this number to the volume of the local neighborhood range is used as the spatial distribution density value of the target high-dimensional feature vector position, thus obtaining the density distribution dataset corresponding to all high-dimensional feature vectors. The numerical trend of the density distribution dataset is analyzed to identify the peak points where the density value jumps from low to high and the valley points where the density value falls from high to low. Based on the peak points and valley points, the continuous feature space is divided into several feature intervals with different density thresholds. For each feature interval after division, the original acquired data stream of the media file corresponding to the feature vector in the feature interval is extracted as the observation sequence, and the semantic logic evolution data stream between entity nodes in the feature interval is extracted as the state sequence.
6. The method for generating intelligent tags and constructing knowledge graphs for converged media content according to claim 5, characterized in that, The optimized knowledge graph is stored in a distributed database. Based on the stored graph data, a hierarchical retrieval system for role-based permissions is constructed, and intelligent recommendation logic is executed to realize the management and application of media resources, including: The updated entity node weights and graph edge association strengths in the optimized knowledge graph are serialized and encapsulated, and stored in a distributed database to form a graph data storage set that can be accessed concurrently. Based on the graph data storage set, sensitive tag attributes and propagation path information are extracted from entity nodes. Combined with the preset user role definition, a hierarchical retrieval index system including different access levels and data filtering rules is constructed. The hierarchical retrieval index system is used to respond to user retrieval requests. Based on the user's role and permissions, the entity nodes and associated edges in the graph data storage set are dynamically trimmed and reorganized to obtain candidate resource subgraphs that conform to the current permission scope. Based on the spatiotemporal correlation strength and semantic mapping depth between entity nodes in the candidate resource subgraph, the matching degree between the user's past behavior trajectory and the features of the candidate resource subgraph is calculated, forming an intelligent recommendation list sorted by matching degree, thereby realizing the management and application of media resources.
7. The method for generating intelligent tags and constructing knowledge graphs for converged media content according to claim 6, characterized in that, The different access levels include the public access layer, the restricted access layer, and the core control layer.
8. The method for generating intelligent tags and constructing knowledge graphs for converged media content according to claim 7, characterized in that, The data filtering rules include node blocking rules based on sensitive attributes, truncation rules based on propagation paths, threshold filtering rules based on spatiotemporal and semantic association strength, and sequence data access rules based on feature intervals.
9. A system for intelligent tag generation and knowledge graph construction of converged media content, wherein the system implements the method as described in any one of claims 1 to 8, characterized in that, include: The module is used to map the features of people, items, scene descriptions, voice text, subtitle text and sensitive information in the initial analysis result set into structured data, form structured label items, and summarize the structured label items into a structured label set. The topology module is used to convert structured tag items in the structured tag set into graph entity nodes. With the media file as the root node, it establishes the spatiotemporal relationship between the corresponding entity nodes of people, items and scenes as graph edges; it forms semantic relationship edges by semantic mapping between the corresponding entity nodes of voice text and subtitle text, and marks the propagation path of the entity nodes corresponding to sensitive information features, thus constructing a preliminary knowledge graph topology structure. The calibration module maps multi-source heterogeneous data of entity nodes and graph edges in the initial knowledge graph topology into high-dimensional feature vectors to construct a multi-dimensional association model. It divides feature intervals according to the spatial distribution density of the feature vectors, defining media data as observation sequences and semantic data as state sequences. Using preset node weights, it predicts the state sequence at the current moment to obtain predicted values, and calculates gain coefficients based on the current feature interval density. The gain coefficients are then used to weight and fuse the deviations between the observation sequences and predicted values to update the node weights, resulting in an optimized knowledge graph. This includes a state sequence based on feature intervals and their corresponding state sequences. Preset node weights are used to perform forward propagation calculations on the current state sequence to obtain predicted state values representing the semantic logic evolution trend. The predicted state values are then compared with the media data observation sequences within the same feature interval to obtain the original... An initial deviation vector is generated, and based on this vector, the entity nodes to be updated and their current connection weights in the preliminary knowledge graph topology are obtained. According to the numerical distribution characteristics of the original deviation vector, combined with the spatial distribution density value corresponding to the current feature interval, the dispersion of the spatial distribution density value relative to the global density distribution is analyzed. This dispersion is transformed into an adjustment amplitude parameter, which is then applied to the original deviation vector to quantify the reliable fluctuation range of the original deviation vector under the current data density environment, resulting in a dynamic gain coefficient. The original deviation vector is then scaled using the dynamic gain coefficient to obtain a weighted fusion deviation quantity that includes density adaptive characteristics. This weighted fusion deviation quantity is used as a correction signal and fed back to the preliminary knowledge graph topology to update the current weights of each entity node and the association strength of the graph edges, resulting in an optimized knowledge graph. The application module stores the optimized knowledge graph in a distributed database, constructs a hierarchical retrieval system based on role permissions based on the stored graph data, and executes intelligent recommendation logic to realize the management and application of media resources.
Citation Information
Patent Citations
Medical knowledge graph construction method for multi-source heterogeneous data fusion and incremental updating
CN121506516A
Rapid predictive analysis of very large data sets using the distributed computational graph
US20170124464A1