A knowledge graph-based cultural digital video content display method and system
By using a knowledge graph-based approach, multimodal data is collected to generate triplet knowledge, analyze the relationships between entities, and construct the video content structure. This solves the problems of fragmented content and insufficient user experience in the digital display of culture, and achieves efficient cultural dissemination and immersive experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-20
- Publication Date
- 2026-03-31
AI Technical Summary
Existing methods of digital cultural display struggle to strike a balance between content depth and user experience. They lack exploration of the intrinsic connections between cultural elements, resulting in fragmented content presentation that makes it difficult for users to form a holistic understanding. Furthermore, different forms of data fail to be effectively combined, resulting in a lack of hierarchy and immersion.
Using a knowledge graph-based approach, we collect multimodal data, generate triplet knowledge, analyze relationships between entities, integrate temporal organization and customary order, construct a video content organization structure, and support user-driven video sequences and interactive display responses through knowledge-driven methods, enabling users to explore autonomously and update dynamically.
It significantly enhances the semantic depth and interactive coherence of cultural data, enabling personalized content recommendations and dynamic dissemination optimization, and promoting an immersive experience and efficient dissemination of cultural heritage.
Smart Images

Figure CN121542464B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of cultural digitization technology, and in particular discloses a method and system for displaying cultural digitization video content based on knowledge graphs. Background Technology
[0002] Cultural digitization, as an important means of inheriting and disseminating traditional culture, has irreplaceable value in today's information age. It not only allows history and culture to reach the public in a more vivid way but also provides strong support for cultural preservation and education. However, despite the significant attention this field receives and its undeniable importance, existing methods still face many challenges in practical application and urgently require breakthroughs.
[0003] Current methods of digital cultural presentation often struggle to strike a balance between content depth and user experience. While many solutions can present rich cultural information, they lack an exploration of the intrinsic connections between cultural elements, resulting in fragmented content that makes it difficult for users to form a holistic understanding. Furthermore, existing technologies often neglect the integration of different data formats when processing cultural content; for example, the ineffective combination of images, text, and sound information leads to a lack of depth and immersion in the storytelling of cultural narratives.
[0004] A deeper issue lies in how to construct logical connections between cultural content through technological means and transform them into an intuitive and interactive presentation method, which has become a key technical challenge in this field. Cultural content often involves complex backgrounds, historical contexts, and multifaceted relationships, such as the production techniques, transmission paths, and connections with other artifacts of the same era of a particular cultural relic. If these relationships cannot be sorted out and presented, users can only see isolated points of information and cannot understand the complete story behind the culture. This technological deficiency directly leads to insufficient appeal of cultural content in dissemination. Specifically, in actual business scenarios, such as when displaying an ancient ceramic piece, the system may only be able to provide basic information such as its name and era, but cannot further reveal its connection with specific historical events, craft transmission, or other related cultural relics. When watching related videos, if users want to understand the production techniques or cultural background behind this ceramic piece, they often need to search for other information on their own, resulting in a fragmented and inefficient experience.
[0005] Therefore, how to effectively construct and present the deep connections between cultural elements in digital cultural videos, so that the content is both logically coherent and supports users' independent exploration and interaction, has become a key issue that this research urgently needs to address. Summary of the Invention
[0006] This invention provides a method and system for displaying cultural digital video content based on knowledge graphs, aiming to solve at least one of the defects existing in the above-mentioned prior art.
[0007] One aspect of the present invention relates to a method for displaying cultural digital video content based on knowledge graphs, comprising the following steps:
[0008] S100. Collect cultural multimodal data, process cultural entities in the cultural multimodal data using named entity recognition method, and generate triple knowledge by combining visual feature encoding and audio semantic alignment to obtain a cultural knowledge graph. The cultural multimodal data includes video documents and oral history records.
[0009] S200. Based on the cultural knowledge graph, the relationship extraction method is used to analyze the relationship between entities, integrate the temporal organization of custom order, drive the logical coherence of video generation, and determine the video content organization structure.
[0010] S300: Obtain the video content organization structure, automatically link relevant video segments from entity associations in the cultural knowledge graph, enhance the semantic association depth, and obtain knowledge-driven video sequences;
[0011] S400: Construct a dual-view interface of main video and knowledge graph panel through knowledge-driven video sequences, synchronously display the attributes when clicking on an entity, and judge the interactive display response.
[0012] S500 If the interactive display response conforms to the inheritance relationship, then jump to the relevant video along the inheritance relationship, incorporate audio alignment elements to enhance dynamic interaction, and obtain the user navigation path;
[0013] S600: Based on the user's navigation path, incrementally update the personalized content recommendation module in the cultural knowledge graph, integrate visual encoding feedback to adjust the priority of related entities, and obtain an optimized propagation sequence.
[0014] Further, step S100 includes:
[0015] S110. Obtain cultural multimodal data from image documents and oral history records. Use image segmentation tools to separate visual segments related to cultural entities from the image portion of the cultural multimodal data. At the same time, extract text content from the audio data in the oral history records using speech-to-text tools to obtain a preliminary set of cultural entity information.
[0016] S120. Based on the preliminary set of cultural entity information, the text content is processed using a named entity recognition tool to extract the names of cultural entities in the text content, and visual feature codes are generated by combining visual fragments with feature extraction tools to obtain multimodal feature descriptions of cultural entities.
[0017] S130. For the multimodal feature description of cultural entities, if the semantic alignment matching degree between the visual feature encoding and the audio-to-text content is higher than the preset threshold, the corresponding triplet knowledge is generated through the semantic alignment tool; otherwise, the mismatched part is corrected by the content completion tool to determine the final triplet knowledge set.
[0018] S140. Integrate the triplet knowledge set through knowledge graph construction tools to obtain the complete mapping of node and edge relationships and generate a cultural knowledge graph.
[0019] Further, step S200 includes:
[0020] S210. Based on the cultural knowledge graph, use a relation extraction tool to analyze the relationships between cultural entities, and obtain at least one key connection path to obtain a preliminary cultural relationship network.
[0021] S220. For the preliminary cultural relationship network, use time sequence analysis tools to integrate the temporal customs in the cultural relationship network, sort out the temporal logic of cultural events, and determine the temporal arrangement framework of cultural content.
[0022] S230. If the semantic consistency between the time arrangement framework and the cultural background is higher than the preset threshold, then the content arrangement tool is used in combination with the visual content coherence to generate coherent video content segments; otherwise, the inconsistent parts are adjusted by the semantic correction tool to obtain the adjusted content segments.
[0023] S240. By combining the adjusted content segments with the organizational structure using structured arrangement tools and integrating them into a dynamic presentation mode, the final video content organizational structure is determined.
[0024] Further, step S300 includes:
[0025] S310. Based on the cultural knowledge graph, an entity linking tool is used to automatically match the relationships between entities, and at least one related video segment is obtained from the video content organization structure to obtain a preliminary semantically related segment group.
[0026] S320. For the initial semantic association fragment group, the semantic depth enhancement tool is used to sort and filter the association strength between the fragments in the initial semantic association fragment group. If the association strength between the fragments in the semantic association fragment group is lower than the preset threshold, irrelevant fragments are removed to obtain a selected semantic association fragment combination.
[0027] S330. Using a temporal integration tool, the selected semantically related segments are adjusted in time order. Combined with the temporal logic in the cultural knowledge graph, the playback order of the segments is determined to obtain an ordered sequence of video segments.
[0028] S340. Using content fusion tools, the ordered video segment sequence is semantically aligned with the background information in the cultural knowledge graph to determine whether it conforms to the overall knowledge-driven logical consistency, thus obtaining the final knowledge-driven video sequence.
[0029] Further, step S400 includes:
[0030] S410. Based on the knowledge-driven video sequence, use an interface building tool to divide the main video display area and the knowledge graph panel area, obtain the interface layout structure after division, and determine the initial display framework.
[0031] S420. For the initial display frame, the video sequence content and the graph entity association data are loaded into the corresponding areas by the content loading tool to obtain the loaded content distribution view.
[0032] S430. Based on the content distribution view, use the interactive response tool to detect entity click response actions. If a click action is detected, trigger the attribute synchronization display function to obtain the relevant attribute data of the clicked entity and determine its display position in the knowledge graph panel.
[0033] S440. Use dynamic adjustment tools to match and optimize the display position with the user interaction logic, obtain the adjusted interface display effect, and determine the final interactive display response.
[0034] Further, step S500 includes:
[0035] S510. If the interactive display response conforms to the inheritance relationship, then according to the jump requirements of the inheritance relationship, the content mapping tool is used to extract the target video segment related to the current content from the pre-established video library and obtain the corresponding video identifier.
[0036] S520: For video identifiers, use audio processing tools to extract sound features from the target video segment, integrate preset audio alignment elements, and obtain adjusted audio and video synchronized content.
[0037] S530. If the adjusted audio and video synchronization content meets the preset playback standard, then the dynamic interaction tool is used to respond to the user's click action in real time, load the adjusted audio and video synchronization content into the playback area, and determine the triggering time of the interaction response.
[0038] S540. Based on the timing of user operation feedback and interaction response, construct the jump path from the current content to the target video segment through the navigation path generation tool to obtain the complete user navigation path.
[0039] Further, step S600 includes:
[0040] S610. Based on the user navigation path, retrieve the corresponding user behavior data from the pre-established database, use a classification tool to divide the user behavior data into behavior patterns, and determine the preliminary correlation with the content of the cultural knowledge graph.
[0041] S620. If the initial correlation is higher than the preset threshold, the behavioral pattern is matched with the entity in the cultural knowledge graph through the content mapping tool to obtain the list of matched related entities and determine the initial priority in the list of related entities.
[0042] S630. For the initial priority, a visual coding feedback tool is used to extract user interaction response data, and a sorting tool is used to dynamically adjust the list of related entities to obtain the adjusted priority sequence.
[0043] S640. Based on the adjusted priority sequence, refresh the personalized recommendation content in the cultural knowledge graph using the incremental update tool of the content recommendation module, obtain the optimized recommendation combination, and determine the final propagation sequence.
[0044] Another aspect of the present invention relates to a knowledge graph-based cultural digital video content display system, used to perform the above-described knowledge graph-based cultural digital video content display method, comprising:
[0045] The cultural knowledge graph acquisition module is used to collect multimodal cultural data. It uses named entity recognition to process cultural entities in the multimodal cultural data, and combines visual feature encoding and audio semantic alignment to generate triple knowledge, thus obtaining a cultural knowledge graph. The multimodal cultural data includes video documents and oral history records.
[0046] The video content organization structure determination module is used to analyze the relationships between entities based on the cultural knowledge graph, use relationship extraction methods, integrate the temporal organization of customs, drive the logical coherence of video generation, and determine the video content organization structure.
[0047] The knowledge-driven video sequence acquisition module is used to obtain the organizational structure of video content, automatically link relevant video segments from entity associations in the cultural knowledge graph, enhance the depth of semantic association, and obtain knowledge-driven video sequences.
[0048] The interactive display response judgment module is used to construct a dual-view interface of main video and knowledge graph panel through knowledge-driven video sequence, synchronously display the attributes when clicking on an entity, and judge the interactive display response.
[0049] The user navigation path acquisition module is used to jump to the relevant video along the inheritance relationship if the interactive display response conforms to the inheritance relationship, and incorporate audio alignment elements to enhance dynamic interaction and obtain the user navigation path;
[0050] The propagation sequence acquisition module is used to incrementally update the personalized content recommendation module in the cultural knowledge graph based on the user's navigation path, and adjust the priority of related entities by integrating visual encoding feedback to obtain an optimized propagation sequence.
[0051] The beneficial effects achieved by this invention are as follows:
[0052] This invention discloses a method and system for displaying digital cultural video content based on a knowledge graph. It addresses the unique business scenarios in cultural transmission, including fragmented processing of multimodal data (such as video documents and oral history records), inconsistent entity relationship analysis, lack of logical organization in video content, and the inability to achieve personalized navigation and dynamic updates in user interaction. This problem is condensed into the logical challenge of how to efficiently integrate multimodal data to construct a knowledge graph to drive video generation and optimize cultural dissemination paths. This invention processes cultural entities using named entity recognition, combines visual feature encoding with audio semantic alignment to generate triplet knowledge, forming a cultural knowledge graph. It then uses relation extraction to analyze relationships between entities, integrates temporal organization of custom sequences, and determines the video content structure. Video segments are automatically linked from the graph to construct a dual-view interface of main video and knowledge graph panel, supporting synchronous display of entity attributes and jumps to transmission relationships, and incorporating audio alignment to enhance interaction. The graph is incrementally updated based on the user's navigation path, and visual encoding feedback is integrated to adjust entity priorities, obtaining an optimized dissemination sequence. This invention significantly enhances the semantic association depth and interactive coherence of cultural data, enabling personalized content recommendation and dynamic dissemination optimization, ultimately promoting an immersive experience and efficient dissemination of cultural heritage. Attached Figure Description
[0053] Figure 1 This is a flowchart illustrating an embodiment of the knowledge graph-based method for displaying digital cultural video content according to the present invention.
[0054] Figure 2 This is a functional block diagram of an embodiment of the knowledge graph-based cultural digital video content display system of the present invention.
[0055] Explanation of icon numbers:
[0056] 10. Cultural Knowledge Graph Acquisition Module; 20. Video Content Organization Structure Determination Module; 30. Knowledge-Driven Video Sequence Acquisition Module; 40. Interactive Display Response Judgment Module; 50. User Navigation Path Acquisition Module; 60. Propagation Sequence Acquisition Module. Detailed Implementation
[0057] To better understand the above technical solutions, the following will provide a detailed explanation of the technical solutions in conjunction with the accompanying drawings and specific implementation methods.
[0058] like Figure 1As shown, the first embodiment of the present invention proposes a method for displaying cultural digital video content based on knowledge graphs, including the following steps:
[0059] Step S100: Collect cultural multimodal data, process cultural entities in the cultural multimodal data using named entity recognition method, and generate triple knowledge by combining visual feature encoding and audio semantic alignment to obtain a cultural knowledge graph. The cultural multimodal data includes video documents and oral history records.
[0060] The system first collects two types of multimodal cultural data: video documents (such as live videos of intangible cultural heritage skills, video archives of ancient buildings, and documentaries of folk activities, etc., which are dynamic / static visual materials) and oral history records (such as audio files and corresponding transcripts of inheritors recounting the origins of their skills and elderly people recalling folk traditions). It then employs named entity recognition methods adapted to multimodal scenarios (such as the BERT-NER (Bidirectional Encoder Representations from Transformers for Named Entity Recognition) model that integrates visual features and a cross-modal entity extraction framework) to accurately extract cultural entities from video images (such as identifying visual entities like "embroidery stitches" and "shadow puppet props") and oral history content (such as extracting semantic entities like "Dragon Boat Festival customs" and "Luban lock making techniques"), uniformly encoding them into standardized entity labels. Finally, it performs visual feature encoding on the video documents (such as using CNN (Convolutional Neural Networks)). A convolutional neural network (CNN) extracts color, texture, and scene structure features from the image to generate high-dimensional visual vectors. Audio semantic analysis is performed on oral history records (e.g., transcribing text using ASR (Automatic Speech Recognition) technology and extracting semantic information by combining voice emotion features). Cross-modal alignment algorithms (e.g., attention mechanisms) are used to establish the correspondence between visual features and audio semantics (e.g., semantic alignment between "paper-cutting action" in the video and "folding and cutting" in the oral history). Based on the alignment results, triple knowledge in the format of "entity-relationship-entity" is generated (e.g., "shadow puppetry - core props - shadow puppets", "Suzhou embroidery - representative stitches - random stitch embroidery"). All triples are integrated according to graph specifications (e.g., RDF format) to construct a structured cultural knowledge graph containing cultural entities, relationships, and multimodal feature mappings. Entity recognition accuracy is required to be ≥93% (for common cultural entities), and the logical consistency of triple generation is required to be ≥95%, providing structured knowledge support for the semantic organization of subsequent video content.
[0061] Step S200: Based on the cultural knowledge graph, analyze the relationships between entities using the relationship extraction method, integrate the temporal organization of custom order, drive the logical coherence of video generation, and determine the video content organization structure.
[0062] Based on the cultural knowledge graph constructed in step S100, deep learning-based relation extraction methods (such as cross-sentence attention mechanism models and Few-Shot relation extraction frameworks) are employed to mine deep connections between cultural entities from the graph's triple knowledge (including semantic relations such as "subordination," "derivation," and "causation," functional relations such as "tool-use" and "skill-process," and spatial relations such as "region-custom"). Simultaneously, combining the unique temporal logic of cultural content (such as the timeline of skill transmission and the order of folk activities) and the order of organizational customs (such as the procedures of festival celebrations and the normative steps of rituals), a two-dimensional ranking based on "relationship priority and temporal weight" is established. Rules (such as core inheritance relationships taking precedence over general connections, and key custom steps having higher chronological weight than secondary steps); this two-dimensional sorting rule drives the logical coherence of video generation, selects core entities and key relationships as the main storyline of the video, and outlines the narrative path of "introducing the core entity at the beginning → developing related content according to relationships / chronological order → summarizing cultural value at the end," ultimately determining the hierarchical video content organization structure (such as the main module being "Dragon Boat Festival Customs," the sub-modules being "picking mugwort → wrapping zongzi → dragon boat racing → hanging sachets," and the sub-modules being further subdivided into specific display units), requiring structural logical consistency ≥96% and chronological / customary order conformity ≥98%, providing a clear framework guide for the connection of subsequent video segments.
[0063] Step S300: Obtain the video content organization structure, automatically link relevant video segments from the entity association in the cultural knowledge graph, enhance the semantic association depth, and obtain a knowledge-driven video sequence.
[0064] First, obtain the hierarchical video content organization structure determined in step S200 (e.g., the main-sub-module framework of "Dragon Boat Festival Customs → Picking Mugwort → Wrapping Zongzi → Dragon Boat Racing"). Using the core entities (e.g., "picking mugwort" and "wrapping zongzi") and their relationships (e.g., "progression of custom process") of this structure as the retrieval basis, match the corresponding entity nodes from the cultural knowledge graph. Through an entity association mapping mechanism (e.g., association retrieval based on the graph adjacency matrix), automatically link video segments directly / indirectly related to the entities (e.g., the entity "wrapping zongzi" is associated with sub-segments such as "zongzi leaf processing," "glutinous rice preparation," and "wrapping techniques"; the entity "dragon boat racing" is associated with "dragon boat making" and "team formation"). The "race process" segment is used to enhance the semantic connection between segments through semantic enhancement algorithms (such as attention weight redistribution). For example, when jumping from the "making zongzi" segment to the "zongzi leaf varieties" segment, the semantic connection between "ingredients and making" is preserved. At the same time, according to the hierarchy and order of the video content organization structure, the linked segments are sorted in time and smoothed through transitions (such as adding transition effects and semantic transition narration). The result is a knowledge-driven video sequence with "structure that fits the framework, semantic coherence of segments, and close knowledge connection". The semantic matching degree of segment association is required to be ≥95%, and the structural framework fit degree is required to be ≥97%, providing core video materials for subsequent interactive presentations.
[0065] Step S400: Construct a dual-view interface of main video and knowledge graph panel through knowledge-driven video sequence, synchronously display the attributes when clicking on an entity, and judge the interactive display response.
[0066] Using the knowledge-driven video sequence generated in step S300 as the core content carrier, a dual-view visualization interface of "main video playback area + knowledge graph interactive panel" is constructed. The main video playback area plays cultural content videos (such as intangible cultural heritage skill demonstrations and folk activity records) in sequence, while the knowledge graph panel is displayed simultaneously on the side or in a floating layer. The panel dynamically presents the cultural entities involved in the current video frame and the relationships between entities (e.g., when the video "Suzhou embroidery making" is played, the panel displays "Suzhou embroidery", "random stitch embroidery", and "silk fabric"). (Entities and related links); When a user clicks on a target entity in the knowledge graph panel (such as "random needle embroidery"), the interface displays the entity's detailed attributes in real time (including structured information such as definition, historical origin, technical characteristics, and related inheritors, along with corresponding images / audio clips); At the same time, the interaction detection module (such as click event listening and operation behavior recognition algorithm) captures the user's interactive display response (such as clicking on entities, dragging related links, querying attribute details, etc.), determines the user's exploration intention (such as gaining a deeper understanding of a certain skill, querying entity relationships), and provides accurate interaction basis for subsequent dynamic navigation jumps. It requires that the entity attribute display delay be ≤300ms and the interaction response accuracy rate be ≥99%, achieving a seamless connection between "video viewing" and "knowledge exploration".
[0067] Step S500: If the interactive display response conforms to the inheritance relationship, then jump to the relevant video along the inheritance relationship, incorporate audio alignment elements to enhance dynamic interaction, and obtain the user navigation path.
[0068] For the interactive display responses captured in step S400 (such as a user clicking the "Inheritance" tag in the entity association link or querying the entity's inheritance lineage), the intent recognition module determines whether it points to an inheritance relationship between cultural entities (such as preset relationship types like skill inheritance, folk custom continuation, and school derivation). If it is determined to be an inheritance relationship, the inheritance link in the cultural knowledge graph is used as the navigation basis (such as "Ancient Paper-cutting Techniques → Modern Paper-cutting Schools → Modern Campus Paper-cutting Teaching"), and the user is automatically redirected to the corresponding video segment in the link (such as jumping from the "Ancient Paper-cutting Techniques" display video to the "Modern Campus Paper-cutting Teaching" recording). Audio alignment elements are integrated during the redirection process to enhance the connection. The dynamic interactive experience aligns audio clips related to heritage from preceding videos (such as the oral narrative of a craftsman passing down his skills from generation to generation) with the opening audio of the target video, or adds transitional sound effects and narration on the theme of heritage (such as "This craft has been passed down for three hundred years and is now entering the campus"). At the same time, it records the user's navigation path in real time (including the starting entity, the jump node, and the target video), forming a structured user navigation path that includes timestamps, entity associations, and operation paths. It requires an accuracy rate of ≥98% in determining the heritage relationship, a video jump delay of ≤500ms, and an audio alignment synchronization error of ≤100ms, allowing users to intuitively experience the historical context of cultural heritage.
[0069] Step S600: Based on the user's navigation path, incrementally update the personalized content recommendation module in the cultural knowledge graph, integrate visual encoding feedback to adjust the priority of related entities, and obtain the optimized propagation sequence.
[0070] Using the user navigation path generated in step S500 (including behavioral data such as user-explored entity nodes, inheritance relationship jump trajectories, and dwell time) as the core input, the personalized content recommendation module in the cultural knowledge graph is incrementally updated. High-frequency access entities (such as users repeatedly jumping to content related to "modern paper-cutting innovation") and preference relationship types (such as emphasizing "craft inheritance" and "regional derivative" relationships) in the navigation path are extracted as user interest feature tags. Simultaneously, visual encoding feedback during user interaction is collected (such as visual feature data corresponding to behaviors like watching a certain type of video clip completely, replaying it, pausing and taking screenshots). Interest feature tags and visual data are then fused using a weight adjustment algorithm (such as a priority calculation model based on collaborative filtering). The coding feedback dynamically adjusts the recommendation priority of related entities in the knowledge graph (e.g., increasing the priority of user-preferred entities like "intangible cultural heritage inheritance in schools" and decreasing the weight of less popular entities that are not interacted with). Based on the adjusted entity priorities, the dissemination logic of cultural content is reorganized, high-priority entities and related video clips are selected, and they are reorganized according to the path of "user interests → association expansion → cultural value extension" to finally obtain an optimized dissemination sequence (e.g., for users who prefer "modern paper-cutting," a personalized dissemination sequence of "modern campus paper-cutting teaching → innovative works by young designers → development of paper-cutting cultural and creative products" is generated). The interest matching degree is required to be ≥94%, and the user click-through rate of the dissemination sequence is increased by ≥30%, so as to achieve precise and personalized dissemination of cultural content.
[0071] Furthermore, the knowledge graph-based method for displaying cultural digital video content provided in this embodiment includes step S100, which includes:
[0072] Step S110: Obtain cultural multimodal data from image documents and oral history records. Use an image segmentation tool to separate visual segments related to cultural entities from the image portion of the cultural multimodal data. At the same time, extract text content from the audio data in the oral history records using a speech-to-text tool to obtain a preliminary set of cultural entity information.
[0073] The isolated visual fragments related to cultural entities are derived using the following formula:
[0074] (1)
[0075] In formula (1), This represents a set of visual fragments of cultural entities obtained by segmenting from a portion of an image. This indicates the total number of input images. Indicates the first One image, This indicates the parameter configuration for the image segmentation tool. The function represents an image segmentation operation. Indicates cultural entity identification, The function represents a mask extraction operation based on cultural entities.
[0076] The text content extracted from oral history audio is derived using the following formula:
[0077] (2)
[0078] In formula (2), This refers to the text content extracted from oral history audio. This refers to speech-to-text utility functions. This represents a collection of oral history audio data. Indicates the parameters of the speech recognition model. Indicates the total number of audio segments. Indicates the first An audio clip, The function represents an audio-to-text conversion operation. Indicates that for the first Recognition parameters for each audio segment.
[0079] The initial set of cultural entity information is derived using the following formula:
[0080] (3)
[0081] In formula (3), This represents a preliminary collection of information about cultural entities. The function represents a multimodal data fusion operation. The weight coefficients representing visual segments, Weighting coefficients representing the text content. Weighting coefficients representing cross-modal overlap information Indicates the number of visual segments, Indicates the number of text segments. Functions representing vision and text information Calculation of the correlation between them.
[0082] When acquiring multimodal cultural data from visual documents and oral histories, consider a specific cultural heritage project, such as video footage and oral histories from elders about traditional Chinese festivals like the Spring Festival. First, collecting visual documents might include old photographs or video clips capturing scenes of dragon lantern dances, while oral histories consist of audio interviews with elders describing the origins and evolution of Spring Festival customs. This approach emphasizes the diversity of multimodal data, ensuring comprehensive coverage of images, audio, and text, thus providing rich foundational information for subsequent analysis. This method not only preserves the visual representation of culture but also captures the narrative details of oral transmission, contributing to the construction of a more comprehensive cultural database.
[0083] To separate visual fragments related to cultural entities from the image portion of multimodal cultural data using image segmentation tools, it's essential to first understand the principles of these tools. They are typically based on deep learning models such as Mask R-CNN (Mask Region-based Convolutional Neural Network), which divides the image into different regions through pixel-level classification. For example, in Spring Festival footage, image segmentation tools can identify and separate visual fragments of dragon lantern props. The process involves inputting the image, generating bounding boxes to detect the object, applying masks to predict the precise contours, and finally outputting independent dragon lantern image fragments. This separation helps focus on cultural entities, avoids irrelevant background interference, and improves the targeting of data processing.
[0084] When extracting textual content from audio data in oral history records using speech-to-text tools to obtain a preliminary set of cultural entity information, such as Transformer-based models, the speech-to-text tools convert the audio waveform into a spectrogram, and then decode it into a text sequence through an attention mechanism. For example, an audio recording of an elderly person recounting "dragon lantern dances for blessings during the Spring Festival" will be converted into the text "dragon lantern dances for blessings during the Spring Festival," thus forming a preliminary set including entities such as "Spring Festival," "dragon lantern," and "blessings." This extraction process ensures that audio information is transformed into a processable textual form, laying the foundation for subsequent entity recognition.
[0085] Step S120: Based on the preliminary set of cultural entity information, the text content is processed using a named entity recognition tool to extract the names of cultural entities in the text content, and visual feature codes are generated by combining visual fragments with feature extraction tools to obtain multimodal feature descriptions of cultural entities.
[0086] The following formula describes the process of extracting cultural entities from text using named entity recognition tools:
[0087] (4)
[0088] In formula (4), This represents the set of extracted cultural entity names. This represents the input text content. This represents a preliminary collection of information about cultural entities. This represents the processing function of the named entity recognition tool.
[0089] The following formula describes the process of converting visual segments into visual feature codes using feature extraction tools:
[0090] (5)
[0091] In formula (5), This represents the generated visual feature encoding. This represents the input visual segment. This represents the encoding function of the feature extraction tool.
[0092] The process of combining text entity features with visual features to generate multimodal feature descriptions is described by the following formula:
[0093] (6)
[0094] In formula (6), Multimodal feature descriptions representing cultural entities This represents the cultural entity features extracted from the text. Represents visual feature encoding. This represents the multimodal feature fusion function.
[0095] Based on a preliminary set of cultural entity information, named entity recognition (NER) tools are used to process the text content, extracting the names of culturally relevant entities. NER tools, such as BERT-based NER, tokenize the text and then label entity types through a classification layer. Specifically, in the text "Dragon lantern dance for blessings during the Spring Festival," the BERT-based NER model identifies "Spring Festival" as a festival entity and "dragon lantern" as an object entity, extracting them into a list. This processing accurately captures cultural keywords and improves the structuring level of the data.
[0096] When combining visual fragments with feature extraction tools to generate visual feature codes and obtain multimodal feature descriptions of cultural entities, feature extraction tools such as ResNet (Residual Network) models perform convolution operations on dragon lantern image fragments to extract feature vectors such as color and shape, for example, generating a 512-dimensional encoding vector. This vector is then combined with textual entities to form a multimodal description, such as "Dragon Lantern: Red, Long, Symbol of Blessing." This generation process achieves multidimensional representation of culture by fusing visual and textual features.
[0097] Step S130: For the multimodal feature description of cultural entities, if the semantic alignment matching degree between the visual feature encoding and the audio-to-text content is higher than the preset threshold, the corresponding triplet knowledge is generated through the semantic alignment tool; otherwise, the mismatched part is corrected by the content completion tool to determine the final triplet knowledge set.
[0098] The semantic alignment and matching degree between visual feature encoding and audio-to-text content is derived by the following formula:
[0099] (7)
[0100] In formula (7), This represents the semantic alignment matching score. A semantic vector representing the audio-to-text content. Represents visual feature encoding transpose, This represents the preset matching threshold. When the calculated cosine similarity is greater than or equal to the threshold, it is considered a high matching degree.
[0101] The following formula is used to define the conditions for generating triplet knowledge:
[0102] (8)
[0103] In formula (8), This represents the generated triplet knowledge. This indicates the semantic alignment tool's generation function. This indicates the correction function for the content completion tool. The identifier vector represents the mismatched part, and different knowledge generation strategies are selected based on the matching degree.
[0104] The final set of triplet knowledge is derived using the following formula:
[0105] (9)
[0106] In formula (9), This represents the final set of triplet knowledge. Indicates the first The generated triplet knowledge This represents the total number of cultural entities processed. This represents a knowledge verification function that ensures the integrity and accuracy of the knowledge set through union operations and verification filtering.
[0107] For multimodal feature descriptions of cultural entities, if the semantic alignment between the visual feature encoding and the audio-to-text content exceeds a preset threshold, a corresponding triplet knowledge set is generated using a semantic alignment tool. Otherwise, a content completion tool corrects the mismatch, determining the final triplet knowledge set. Semantic alignment tools, such as the CLIP (Contrastive Language-Image Pre-training) model, calculate the cosine similarity between the visual encoding and the text embedding. If it exceeds a threshold of 0.8, a triplet is generated, such as (dragon lantern, belongs to, Spring Festival custom). If there is a mismatch, such as a visual display of a red dragon lantern but a text description of a green one, the content completion tool uses a generative model such as GPT (Generative Pre-trained Transformer) to fill in the gaps and output "Red dragon lanterns are used for prayer," ultimately forming the triplet knowledge set. This mechanism ensures the accuracy and completeness of knowledge, effectively reducing misleading information due to bias in cultural preservation.
[0108] Step S140: Integrate the triplet knowledge set through a knowledge graph construction tool to obtain the complete mapping of node and edge relationships and generate a cultural knowledge graph;
[0109] The process of constructing a complete knowledge graph by integrating all triple knowledge is described by the following formula:
[0110] (10)
[0111] In formula (10), This represents the constructed cultural knowledge graph. This indicates the total number of triples. Represents the head entity node. Indicates relation edges, Represents the tail entity node. Represents a set of entity nodes. Represents a set of relations.
[0112] The following formula defines the mapping mechanism from each node in a knowledge graph to the complete information of that node:
[0113] (11)
[0114] In formula (11), This represents the complete mapping relationship of the nodes. Indicates the first in the atlas 1 node Represents the set of all nodes. A unique identifier representing a node. Represents the attribute information of a node. The type label of the node.
[0115] The following formula describes the mapping structure of edge relationships to detailed information about those edge relationships in a knowledge graph:
[0116] (12)
[0117] In formula (12), Represents a complete mapping of edge relationships. Indicates the first in the atlas Edge, Represents the set of all edges. Indicates the source node of the edge. Indicates the relationship type of the edge. This represents the target node of the edge. This represents the weight value of the edge.
[0118] By integrating triplet knowledge sets through knowledge graph construction tools to obtain a complete mapping of node and edge relationships, a cultural knowledge graph can be generated. When such tools are used, tools like Neo4j import the triples into a graph database, with nodes representing entities such as "Spring Festival" and "dragon lantern," and edges representing relationships such as "containment." This mapping process generates a queryable graph. This integration process not only visualizes cultural connections but also facilitates searching, such as retrieving the connection path between dragon lanterns and blessings, thus providing efficient knowledge retrieval support in cultural studies.
[0119] Furthermore, the knowledge graph-based method for displaying cultural digital video content provided in this embodiment includes step S200 as follows:
[0120] Step S210: Based on the cultural knowledge graph, use a relation extraction tool to analyze the relationships between cultural entities, and obtain at least one key connection path to obtain a preliminary cultural relationship network.
[0121] The following formula is used to obtain key connection paths from cultural entity relationships:
[0122] (13)
[0123] In formula (13), The score represents the importance of a key link in a cultural relationship network. This represents the total number of edges in the path. Indicates the first Cultural relevance of the strip Indicates the first The type weight of the edge Indicates the first The distance metric between the entities at both ends of the stripe is increased by 1 to avoid the denominator being zero.
[0124] A preliminary cultural relationship network is derived using the following formula:
[0125] (14)
[0126] In formula (14), It represents a global connectivity index for the entire cultural relationship network. It represents the set of all cultural entities in a cultural relationship network. This represents a cultural entity node in a cultural relationship network. Representation and entity The set of all adjacent entities, Representing entities Importance weight, Representing entities and The relationship type coefficient This represents a measure of the feature similarity between two entities.
[0127] When extracting relationships from a cultural knowledge graph, consider a specific cultural theme, such as the traditional customs of the Dragon Boat Festival. Relation extraction tools, such as those based on BERT (Bidirectional Encoder Representations from Transformers), first tokenize the nodes and edges in the graph, and then identify the types of relationships between entities through a classification layer. This relation extraction tool takes triplet data as input, such as (Dragon Boat Festival, containment, dragon boat racing), calculates attention weights, and outputs relationship labels such as "custom association," thereby obtaining key paths, such as the link from "Dragon Boat Festival" to "dragon boat racing" and then to "commemorating Qu Yuan." This analysis process iteratively traverses the graph nodes, ensuring that at least one complete path is extracted, forming a preliminary cultural relationship network. This preliminary cultural relationship network uses nodes to represent entities and edges to represent relationships, helping to reveal the inherent connections between cultural elements.
[0128] Step S220: For the preliminary cultural relationship network, integrate the temporal customs sequence in the cultural relationship network through time sequence analysis tools, sort out the temporal logic of cultural events, and determine the temporal arrangement framework of cultural content.
[0129] Step S220: For the preliminary cultural relationship network, integrate the temporal customs sequence in the cultural relationship network through time sequence analysis tools, sort out the temporal logic of cultural events, and determine the temporal arrangement framework of cultural content.
[0130] The following formula is used to define the optimization objective function for the time-series arrangement framework of cultural content:
[0131] (15)
[0132] In formula (15), Indicates inclusion The objective function for optimizing the temporal arrangement framework of cultural content over a given time period. This represents the set of all possible time arrangements. This represents one specific permutation scheme. Indicates the first The importance weight of each time period Indicates time period Quality assessment function for Chinese-language content The penalty coefficient representing temporal continuity. Indicates adjacent time periods and The cost function for the discontinuity between them.
[0133] When integrating chronological customs, time series analysis tools such as LSTM (Long Short-Term Memory) sequence models input events into the network as time series. First, they encode the temporal label of each event, such as the dragon boat race occurring on the morning of the Dragon Boat Festival. Then, they capture dependencies through recurrent layers, clarifying the logic, such as making zongzi (sticky rice dumplings) before the dragon boat race, and finally determining the chronological arrangement framework. This framework arranges events in chronological order, ensuring the accuracy of the cultural content's temporal sequence and providing a foundation for subsequent arrangement.
[0134] Step S230: If the semantic consistency between the time arrangement framework and the cultural background is higher than the preset threshold, then the content arrangement tool is used in combination with the visual content coherence to generate coherent video content segments; otherwise, the inconsistent parts are adjusted by the semantic correction tool to obtain the adjusted content segments.
[0135] The semantic consistency between the chronological arrangement framework and the cultural context is derived using the following formula:
[0136] (16)
[0137] In formula (16), The semantic consistency score between the chronological arrangement framework and the cultural context is represented by the score. This represents the total number of semantic matching points to be evaluated. Indicates the first The weight coefficients of each matching point Indicates the first Semantic vectors of a time-ordered framework This represents a semantic vector representing the corresponding cultural background. The function calculates the cosine similarity between two vectors.
[0138] The coherence index for generating coherent video content segments is derived using the following formula:
[0139] (17)
[0140] In formula (17), This is an indicator of the coherence of generated video content segments. This indicates the total number of video clips. and These represent the weight parameters for visual coherence and content coherence, respectively. Function to calculate adjacent segments and Visual coherence between them Function evaluates a single fragment Content quality.
[0141] The adjusted content segment is derived using the following formula:
[0142] (18)
[0143] In formula (18), This indicates the adjusted content segment after semantic correction. Represents the semantic correction transformation matrix. This represents the original inconsistency content vector. This indicates the adjustment parameter for the correction intensity. The gradient of the semantic error function is represented. This represents the set of data containing the detected inconsistencies.
[0144] When determining semantic consistency, if the similarity between the dragon boat race and the Dragon Boat Festival background in the framework exceeds a threshold of 0.7 based on the calculated cosine value, then content arrangement tools, such as generators based on GANs (Generative Adversarial Networks), will incorporate visual coherence, for example, smoothly transitioning the dragon boat race video clip with the zongzi-making clip to generate a coherent video. Otherwise, the semantic correction tool uses a Transformer decoder to adjust inconsistencies, such as correcting time misalignments, and outputs the adjusted clips. This mechanism enhances the authenticity of the content through matching verification.
[0145] Step S240: Combine the adjusted content segments with the organizational structure using structured editing tools, integrate them into the dynamic presentation mode, and determine the final video content organizational structure.
[0146] The responsiveness of the dynamic rendering mode is calculated using the following formula:
[0147] (19)
[0148] In formula (19), This indicates the responsiveness of the dynamic rendering mode. Indicates the number of types of dynamic elements. Indicates the first Frequency parameters of dynamic elements Indicates the first The range of change of dynamic elements Indicates the first The baseline value of a dynamic element.
[0149] When constructing the organizational structure, using structured orchestration tools such as a graph-based organizer combines segments with a tree structure and incorporates dynamic modes such as interactive transitions to determine the completeness of the final structure, such as evaluating path coverage. This enables efficient video presentation in cultural dissemination and supports dynamic narratives for users to explore Dragon Boat Festival customs.
[0150] Furthermore, the knowledge graph-based method for displaying cultural digital video content provided in this embodiment includes step S300 as follows:
[0151] Step S310: Based on the cultural knowledge graph, use an entity linking tool to automatically match the relationships between entities, and obtain at least one related video segment from the video content organization structure to obtain a preliminary semantically related segment group.
[0152] The initial semantically related fragment group is obtained using the following formula:
[0153] (20)
[0154] In formula (20), This represents the clustering results of semantically related fragment groups. Indicates the segment grouping scheme, This indicates the total number of groups. Indicates the first A collection of fragments within a group. Indicates the first The semantic representation vector of each segment. Indicates the first The center vector of each group, This represents the distance metric function between a segment and the group center.
[0155] The entity linking tool first identifies and matches entities in the cultural knowledge graph. For example, for the Mid-Autumn Festival theme, the graph contains entities such as "Mid-Autumn Festival," "moon gazing," "eating mooncakes," and "Chang'e flying to the moon." The tool then calculates similarity using embedded vectors and automatically associates these entities with segments within the video content's organizational structure.
[0156] After inputting graph nodes, the entity linking tool uses a model such as Word2Vec to generate vector representations. For example, it calculates the cosine similarity between "moon gazing" and the moon scene in a video clip. If the similarity is higher than 0.8, a match is made, thus extracting relevant segments from the structure, such as a video of a dinner party under the moon, forming a preliminary group of semantically related segments. This matching process ensures that the association between entities is based on semantic distance, avoiding random links.
[0157] Step S320: For the preliminary semantic association fragment group, the semantic depth enhancement tool is used to sort and filter the association strength between the fragments in the preliminary semantic association fragment group. If the association strength between the fragments in the semantic association fragment group is lower than the preset threshold, irrelevant fragments are removed to obtain a selected semantic association fragment combination.
[0158] The entity linking tool first identifies and matches entities in the cultural knowledge graph. For example, for the Mid-Autumn Festival theme, the graph contains entities such as "Mid-Autumn Festival," "moon gazing," "eating mooncakes," and "Chang'e flying to the moon." The tool then calculates similarity using embedded vectors and automatically associates these entities with segments within the video content's organizational structure.
[0159] After inputting graph nodes, the entity linking tool uses a model such as Word2Vec to generate vector representations. For example, it calculates the cosine similarity between "moon gazing" and the moon scene in a video clip. If the similarity is higher than 0.8, a match is made, thus extracting relevant segments from the structure, such as a video of a dinner party under the moon, forming a preliminary group of semantically related segments. This matching process ensures that the association between entities is based on semantic distance, avoiding random links.
[0160] Step S320: For the preliminary semantic association fragment group, the semantic depth enhancement tool is used to sort and filter the association strength between the fragments in the preliminary semantic association fragment group. If the association strength between the fragments in the semantic association fragment group is lower than the preset threshold, irrelevant fragments are removed to obtain a selected semantic association fragment combination.
[0161] The following formula is used to calculate the correlation strength between segments using cosine similarity and Gaussian distance decay function:
[0162] (twenty one)
[0163] In formula (21), Representing fragments and fragments The strength of semantic association between them and Each represents a segment and fragments semantic vector representation, This represents the distance between two segments in the semantic space. This represents the distance attenuation parameter.
[0164] The following formula is used to define the retention criteria for semantically related fragment groups:
[0165] (twenty two)
[0166] In formula (22), Indicates the first The average association strength of a group of semantically related fragments Indicates the first The set of all segments contained in a segment group This represents the total number of fragment pairs in a fragment group. Represents any two segments within a group and The strength of the association, This indicates a preset threshold for association strength. When the average association strength is greater than or equal to the threshold, the semantic association fragment group is retained.
[0167] While processing the initial semantically related fragment groups, the semantic depth enhancement tool assesses the strength of the associations between fragments. For example, in the Mid-Autumn Festival group, it calculates association scores for the fragments about "eating mooncakes" and "Chang'e flying to the moon," encoding the textual descriptions of the fragments and outputting strength values such as 0.6 to 0.9 through an attention mechanism similar to BERT. After sorting by the semantic depth enhancement tool, fragments with strengths below a preset threshold of 0.7 are discarded, such as irrelevant modern party fragments, retaining core elements such as combinations of traditional moon-viewing and mythological narratives, resulting in a carefully selected set of semantically related fragments. This selection process, through iterative comparison, enhances the cohesion within each group.
[0168] Step S330: Use a temporal integration tool to adjust the temporal order of the selected semantically related segments, and combine the temporal logic in the cultural knowledge graph to determine the playback order of the segments, thus obtaining an ordered sequence of video segments.
[0169] An ordered sequence of video clips is derived using the following formula:
[0170] (twenty three)
[0171] In formula (23), This represents the optimal sequence of video clips. This represents the sequence to be optimized. Indicates the total number of video clips. Indicates the first position in the sequence A video clip, This represents the temporal distance cost between adjacent segments. Indicates the first The penalty value for each segment that violates cultural logic. This represents the weighting coefficient of the logical penalty.
[0172] The time-series integration tool then adjusts the temporal order of the selected combinations. For example, combining the Mid-Autumn Festival's temporal logic within the graph, such as prioritizing "reunion dinner" before "moon viewing," the tool uses sequence models like RNNs to process segment labels and capture dependencies. Specifically, after inputting segment timestamps, the tool uses sequence models to predict the order, such as placing "eating mooncakes" after dinner, ensuring the sequence progresses from daytime preparation to nighttime climax, forming an ordered sequence of video segments. This adjustment, based on graph logic, maintains the natural flow of cultural events.
[0173] Step S340: Using a content fusion tool, semantically align the ordered video segment sequence with the background information in the cultural knowledge graph, determine whether it conforms to the overall knowledge-driven logical consistency, and obtain the final knowledge-driven video sequence.
[0174] The logical consistency evaluation value of the video clip sequence is obtained using the following formula:
[0175] (twenty four)
[0176] In formula (24), This represents the logical consistency evaluation value of a video sequence. Indicates the total length of the video segment sequence. Indicates the first Knowledge-driven feature representation of a video segment Indicates the first Each video clip features a feature representation predicted based on a knowledge graph. A sensitivity parameter representing logical consistency. This represents an exponential function.
[0177] Knowledge-driven video sequences are derived using the following formula:
[0178] (25)
[0179] In formula (25), This represents the final knowledge-driven video sequence. The embedding vector representing a video segment. A vector representing the contextual background information of a cultural knowledge graph. Feature vectors representing the relationship between videos and knowledge graphs. Indicates the fusion weight of video content. The fusion weights represent the relational features.
[0180] Finally, semantic alignment is performed using content fusion tools. For example, ordered sequences are matched with background information such as "the origin of the Mid-Autumn Festival." The Transformer model is used to calculate the overall consistency score. If the moon-viewing segment in the ordered sequence has a high alignment with the Chang'e background, logical consistency is confirmed, and a knowledge-driven video sequence is output. The content fusion tool's fusion process includes embedding alignment and validation loops to ensure that the sequence reflects a complete cultural narrative, such as from the beginning of Mid-Autumn Festival customs to the end of the myth. This mechanism enhances the educational value of the video through background injection.
[0181] Furthermore, the knowledge graph-based method for displaying cultural digital video content provided in this embodiment includes step S400 as follows:
[0182] Step S410: Based on the knowledge-driven video sequence, use an interface building tool to divide the main video display area and the knowledge graph panel area, obtain the divided interface layout structure, and determine the initial display framework.
[0183] The optimal division ratio between the main video display area and the knowledge graph panel area is determined using the following formula:
[0184] (26)
[0185] In formula (26), This indicates the area ratio between the video region and the knowledge graph region. Indicates the width of the video display area. Indicates the height of the video display area. Indicates the width of the knowledge graph panel. Indicates the height of the knowledge graph panel.
[0186] The following formula is used to evaluate and determine the initial display frame quality of a knowledge-driven video sequence interface:
[0187] (27)
[0188] In formula (27), This represents the overall evaluation value of the initial display frame. Indicates the total number of UI components. Indicates the first The importance weight of each component Indicates the first The compatibility index of each component, Indicates the first Performance metrics for each component.
[0189] The interface building tool first divides the main video display area and the knowledge graph panel area based on the characteristics of the knowledge-driven video sequence. For example, for a video sequence about Spring Festival culture, the tool allocates the left side of the screen as the main video area to play core segments such as "pasting Spring Festival couplets" and "staying up late on New Year's Eve," while the right side is set as the knowledge graph panel, displaying the association graphs of related entities such as "Spring Festival," "red envelopes," and "dragon and lion dances." This division process involves the application of layout algorithms. Specifically, the tool calculates the screen resolution and applies a grid system, such as dividing the total width into a 70% video area and a 30% knowledge graph area, ensuring the balance of the initial display framework and thus obtaining the divided interface layout structure, providing a foundation for subsequent content loading.
[0190] Step S420: For the initial display frame, the video sequence content and the graph entity association data are loaded into the corresponding areas by the content loading tool to obtain the loaded content distribution view.
[0191] The content distribution view after loading is derived using the following formula:
[0192] (28)
[0193] In formula (28), This represents the content distribution view after loading. This indicates the total amount of content in the video sequence. Indicates the first The weighting coefficient of each video content. Indicates the first Content of a video sequence, Indicates the total number of entities in the graph. Indicates the first The correlation strength of each graph entity. Indicates the first Data related to each entity in the graph.
[0194] The content loading tool loads the video sequence content into the main video area on the initial display framework. For example, after inputting a Spring Festival video sequence, the content loading tool extracts segment data through the API (Application Programming Interface). For instance, it first loads the "New Year's Eve Dinner" video segment and simultaneously loads entity association data in the knowledge graph panel. For example, the "New Year's Eve" entity is linked to its historical origin and custom descriptions. This loading mechanism relies on a data mapping process. The content loading tool parses the metadata tags of the sequence and matches them with nodes in the knowledge graph. For example, it uses key-value pairs to map video timestamps to entity IDs (Identities) to obtain a loaded content distribution view, ensuring visual synchronization between the video and the knowledge graph.
[0195] Step S430: Based on the content distribution view, use the interactive response tool to detect entity click response actions. If a click action is detected, trigger the attribute synchronization display function to obtain the relevant attribute data of the clicked entity and determine its display position in the knowledge graph panel.
[0196] The following formula is used to define the click response detection conditions:
[0197] (29)
[0198] In formula (29), Indicates the clicked entity Click response detection results This indicates the distance between the clicked location and the entity's bounding box. This indicates the threshold range for click detection. It returns 1 when the distance is less than or equal to the threshold to indicate that a click was detected, and 0 otherwise to indicate that no click was detected.
[0199] The relevant attribute data of the clicked entity is obtained using the following formula:
[0200] (30)
[0201] In formula (30), Indicates the clicked entity The attribute collection is displayed synchronously. Indicates the first Individual attribute data, Indicates from the clicked entity Through relationships The attribute extraction function is obtained. Indicates the clicked entity Total number of associated attributes.
[0202] The display priority score for a position in the knowledge graph panel is calculated using the following formula:
[0203] (31)
[0204] In formula (31), Represents the coordinates in the knowledge graph panel. Location display priority score, This indicates the relevance weight to the clicked entity. Indicates the panel layout constraint weights. Indicates spatial availability weight. , , These are the corresponding weighting coefficients.
[0205] The interactive response tool detects entity click responses. For example, when a user clicks the "red envelope" entity in the knowledge graph panel, the tool captures the action through an event listener and triggers the attribute synchronization display function. This function retrieves relevant attribute data of the clicked entity from the graph database, such as the origin, symbolic meaning, and modern usage of "red envelope," and then determines its display position in the knowledge graph panel. For example, it uses coordinate calculations to determine the position of the pop-up window to avoid obstructing the main video area.
[0206] Step S440: Use a dynamic adjustment tool to match and optimize the display position with the user interaction logic, obtain the adjusted interface display effect, and determine the final interactive display response.
[0207] The following formula is used to optimize the matching between display location and user interaction logic:
[0208] (32)
[0209] In formula (32), This indicates the optimized best display location configuration. This indicates the display location needs optimization. Indicates the total number of display positions. Indicates the first The weight coefficient of each position, Indicates the first The coordinate parameters of each display location, Indicates the user's position in the first month. Interactive behavior characteristics at each location, The dynamic adjustment tool indicates the first Adjustment parameters for each position, This represents a function that evaluates the matching degree between location and user interaction logic.
[0210] The final interactive display response is derived using the following formula:
[0211] (33)
[0212] In formula (33), This indicates the final, determined interactive display response value. This represents the total number of dimensions in the interactive response. Indicates the first Influence factors in each dimension Indicates the first User operation commands in each dimension Indicates the first System response actions in each dimension Indicates the first Data processing results from each dimension A comprehensive calculation function representing the interactive response.
[0213] The dynamic adjustment tool matches and optimizes the display position with user interaction logic. For example, after detecting a click, the tool assesses the user's device type; if it's a mobile device, it reduces the size of the pop-up window and optimizes the interaction logic, such as adding support for swipe gestures, to achieve the adjusted interface display effect. This optimization process involves an iterative feedback loop. The tool simulates the user path, calculates response latency, and adjusts layout parameters to ensure a smooth and user-friendly final interactive display.
[0214] Furthermore, the knowledge graph-based method for displaying cultural digital video content provided in this embodiment includes step S500:
[0215] Step S510: If the interactive display response conforms to the inheritance relationship, then according to the jump requirement of the inheritance relationship, the content mapping tool is used to extract the target video segment related to the current content from the pre-established video library and obtain the corresponding video identifier.
[0216] The following formula is used to quantify whether an interactive display response conforms to a predefined inheritance relationship structure:
[0217] (34)
[0218] In formula (34), The score indicates the degree of conformity in the inheritance relationship. Indicates the total number of inheritance relationships. Indicates the first The weighting coefficient of each inheritance relationship. Indicates the current content With parent content The control logic of formula (34) is based on the weight priority of inheritance relationship and the strength of relationship between content. Through the process of "weight assignment → correlation strength quantification → weighted summation scoring", it quantifies the degree of fit between interactive display response and predefined inheritance relationship structure, and provides a basis for determining whether to trigger video jump in the future. The core is to highlight the influence of core inheritance relationship through weighted summation and accurately measure the inheritance relationship conformity of response content.
[0219] The target video clip extracted from the video library is obtained using the following formula:
[0220] (35)
[0221] In formula (35), This indicates the target video segment extracted from the video library. This represents a pre-established collection of video libraries. Indicates the current query requirement. Indicates the current content context. Indicates query and video similarity, Indicates content and video The correlation, Indicates video Quality rating , , These are the corresponding weight parameters. The control logic of formula (35) is based on the video library, query requirements, and content context. Through the process of "weighted fusion of multi-dimensional indicators → calculation of comprehensive video score → selection of the best-scoring video", it extracts the target video segments that match the query, fit the context, and meet the quality standards. The core is to use weighted summation to integrate the values of the three dimensions of "query matching degree, contextual relevance degree, and video quality" and take the maximum value to ensure that the extracted video is the best overall.
[0222] The following formula is used to generate a unique video identifier by integrating redirection requirements, content features, and mapping relationships:
[0223] (36)
[0224] In formula (36), This indicates the obtained video identifier. Represents a hash mapping function. The feature vector representing the jump requirement. The feature encoding representing the current content, ⊕ represents the mapping matrix of the content mapping tool, and ⊕ represents the feature fusion operation. The control logic of formula (36) takes the jump requirements, content features, and mapping relationship as the core information source. Through the process of "feature element preparation → multi-dimensional feature fusion → hash mapping to generate unique identifier", a unique video identifier strongly bound to the current requirements is generated. The core is to use a hash function to ensure the uniqueness of the identifier after fusing key features, so as to achieve accurate correspondence between video clips and requirements and content.
[0225] When interactive responses involve cultural heritage relationships, such as in a video system showcasing knowledge of traditional Chinese festivals, if a user clicks on an entity related to "Mid-Autumn Festival," the response is judged to conform to a heritage relationship because Mid-Autumn Festival customs, such as moon gazing, have a direct connection to ancient poetry. In this case, based on the redirection requirement, the content mapping tool extracts relevant segments from a pre-built video library. This tool first parses the metadata of the current content; for example, if the current video is playing a "mooncake making" scene, its metadata contains the keyword "Mid-Autumn Festival heritage." Then, it searches the video library for related segments using a matching algorithm, such as extracting a video segment about the myth of "Chang'e flying to the moon." This extraction process involves database queries; the tool uses SQL-like statements to filter videos in the library tagged with "mythological heritage" to obtain their unique video identifiers, such as "VID-2023-MID-AUTUMN-001," thereby ensuring the cultural continuity between the redirected content and the current theme.
[0226] Step S520: For the video identifier, use an audio processing tool to extract the sound features of the target video segment, and integrate preset audio alignment elements to obtain the adjusted audio and video synchronization content.
[0227] The following formula is used to extract sound features from a target video segment:
[0228] (37)
[0229] In formula (37), This represents the extracted audio feature vector. Indicates the duration of the target video segment. Represents the amplitude of the time-domain audio signal. Represents the frequency domain weighting function. The time-domain window function is represented. The control logic of formula (37) is based on the time-domain audio signal of the target video. Through the process of "time-domain signal preprocessing → frequency domain weighting enhancement → time-domain integration summarization → averaging feature extraction", the feature vector that can represent the core characteristics of video audio is extracted. The core is to combine the window function and frequency domain weighting to accurately capture the key features of the audio and provide a foundation for subsequent audio and video synchronization adjustment.
[0230] The following formula is used to achieve the fusion of the original audio with the preset aligned elements:
[0231] (38)
[0232] In formula (38), This represents the audio alignment parameters after merging. Represents the original audio weighting coefficients. Represents the original audio parameter vector. This represents the preset audio alignment element parameters. The control logic of formula (38) is based on the original audio characteristics and the preset alignment standard. Through the process of "weight allocation → dual-source parameter weighting → fusion parameter output", it balances the original style of the original audio with the standard requirements of the preset alignment elements and generates audio alignment parameters that are adapted to the display requirements. The core is to use a weighted average method to both preserve the characteristics of the original audio of the video and make the audio conform to the audio-video synchronization display standard.
[0233] Based on the acquired video identifiers, the audio processing tool begins extracting sound features from the target video clips. For example, for the "Chang'e Flying to the Moon" clip, the tool analyzes the audio track, extracting features such as pitch variations and rhythm points. Specifically, this involves using Fourier transform to decompose the audio signal into frequency components, identifying the high-frequency components representing the narrative voice of the mythology and the low-frequency components representing ancient musical elements in the background music. Then, preset audio alignment elements are integrated, such as adding synchronized sound effects or adjusting volume balance, to ensure the new clip's audio style matches the original video, resulting in adjusted audio-visual synchronized content. This integration mechanism relies on timeline alignment; the tool calculates the start timestamp of the clip and inserts alignment elements to ensure a natural sound transition and avoid abruptness.
[0234] Step S530: If the adjusted audio and video synchronization content meets the preset playback standard, then a dynamic interactive tool is used to respond to the user's click action in real time, load the adjusted audio and video synchronization content into the playback area, and determine the triggering time of the interactive response.
[0235] The following formula is used to define the conditions under which synchronized audio and video content conforms to playback standards:
[0236] (39)
[0237] In formula (39), This represents the audio-visual synchronization quality assessment index. Represents the audio timestamp. Indicates the video timestamp. This represents the preset synchronization threshold standard, which is used to evaluate the audio and video synchronization quality. When less than or equal to 1, it indicates that the adjusted audio and video synchronization content meets the preset playback standard. The control logic of formula (39) is based on the audio and video timestamp deviation and the preset threshold. Through the process of "calculating timestamp deviation → quantifying the ratio of deviation to threshold → standard compliance judgment", it judges whether the audio and video synchronization content meets the playback requirements. The core is to limit the synchronization error range with the ratio of "deviation / threshold" to ensure the timing consistency of audio and video playback.
[0238] The following formula can be used to determine when to trigger an interactive response for the best user experience:
[0239] (40)
[0240] In formula (40), Indicates the optimal timing for triggering an interactive response. Indicates the parameters for predicting user behavior. Indicates the buffer state parameters, This parameter represents the content loading completion rate. , , These represent the weight coefficients of each parameter. The control logic of formula (40) takes user behavior, playback status, and content loading as the core dimensions. Through the process of "multi-dimensional parameter weighted fusion → trigger timing quantitative calculation → optimal experience timing determination", it comprehensively balances playback smoothness and user operation experience, and determines the best trigger node for interactive response. The core is to use weighted and integrated three types of key state parameters to make the trigger timing both adapt to user behavior habits and ensure playback stability.
[0241] If the adjusted audio and video synchronization content meets preset playback standards, such as a resolution of at least 1080p and an audio bitrate exceeding 128kbps, the dynamic interaction tool will respond to user clicks in real time. When a user clicks on the "Chang'e" entity, the dynamic interaction tool captures the action through a JavaScript event listener, immediately loads the synchronized content into the playback area, and simultaneously determines the trigger timing, such as triggering at the end of the current video frame, to prevent interruption of the user experience. This determination process includes monitoring user operation logs and assessing whether the response latency is less than 500ms, thereby optimizing the smoothness of the interaction. From a business perspective, this real-time response can bring about a technical effect of improved user engagement because it allows for seamless exploration of cultural details and promotes a deeper understanding of knowledge.
[0242] Step S540: Based on the triggering timing of user operation feedback and interaction response, construct the jump path from the current content to the target video segment using the navigation path generation tool to obtain the complete user navigation path.
[0243] The complete user navigation path is derived using the following formula:
[0244] (41)
[0245] In formula (41), This represents the optimal complete user navigation path. Indicates the path candidate index. This represents the total number of dimensions for path evaluation. Indicates the first Weighting factors for each dimension Indicates the first The path in the first Scoring function on each dimension Weight parameters representing path history information Indicates the first The historical navigation effect evaluation value of each path. The control logic of formula (41) is based on the multi-dimensional evaluation and historical effect of candidate navigation paths. Through the process of "candidate path traversal → multi-dimensional score weighting + historical effect fusion → optimal path selection", navigation paths that are both suitable for the current content jump needs and have a good historical experience are selected. The core is to integrate the comprehensive score of "current path quality" and "historical navigation effect" to ensure the rationality and experience of the navigation path.
[0246] Based on the timing of user action feedback and interaction responses, the navigation path generation tool constructs jump paths. For example, if feedback indicates that the user has paused multiple times to view entities, the tool generates a path from "Mooncake Making" to "Chang'e Flying to the Moon." Specifically, it uses graph algorithms such as variants of Dijkstra's Algorithm to calculate the shortest navigation sequence, considering path nodes such as intermediate prompt pages, and finally obtains the complete user navigation path, ensuring logical consistency and supporting backtracking functionality.
[0247] Furthermore, the knowledge graph-based method for displaying cultural digital video content provided in this embodiment includes step S600 as follows:
[0248] Step S610: Based on the user navigation path, obtain the corresponding user behavior data from the pre-established database, use a classification tool to divide the user behavior data into behavior patterns, and determine the preliminary correlation with the content of the cultural knowledge graph.
[0249] User behavior data can be segmented into behavioral patterns using the following formula:
[0250] (42)
[0251] In formula (42), Representing user behavior data The behavioral pattern category that was classified into This represents the total number of predefined behavior patterns. The number of dimensions representing user behavior characteristics. Indicates the first User behavior data in the first Values on the dimensional features, Indicates the first Activation function for dimensional features, Indicates the first Class behavior patterns in the first Classification parameters on dimensional features.
[0252] The preliminary correlation between user behavior data and the content of the cultural knowledge graph is obtained using the following formula:
[0253] (43)
[0254] In formula (43), Representing user behavior data With cultural knowledge graph nodes The initial correlation between them This indicates the total number of user activity records. Indicates the first The importance score of each user behavior Indicates the first Semantic tags for user behavior, The function is used to determine the degree of matching between user behavior and knowledge graph nodes. Represents knowledge graph nodes The number of related concepts Indicates the first The weight values of each related concept.
[0255] Based on the user's navigation path, corresponding user behavior data is retrieved from a pre-established database. For example, in a cultural knowledge dissemination system, the path record formed after a user clicks on a video related to the "Dragon Boat Festival" will be stored in the database. This database typically uses a relational structure such as MySQL and contains fields such as user ID, timestamp, and operation type. The retrieval process involves querying path data. For instance, for the record with user ID "USER-001", the system extracts the navigation sequence from "Dragon Boat Race" to "Legend of Qu Yuan". This data is analyzed to capture user interests, thus providing a basis for subsequent behavior segmentation.
[0256] For user behavior data, a classification tool is used to segment behavioral patterns. This tool can be a machine learning-based algorithm such as K-means clustering, which groups behavioral data into patterns, such as "exploratory" or "deep learning." For path data obtained from a database, the classification tool first preprocesses the data to remove noise, then calculates feature vectors such as click frequency and dwell time. Through the clustering process, user behavior is categorized into "festival custom exploration pattern," which helps identify user preferences for cultural elements. In one embodiment, a preliminary correlation with the content of a cultural knowledge graph is determined. Here, the cultural knowledge graph is a graph database such as Neo4j, storing entities such as "Dragon Boat Festival" and relationships such as "originating from Qu Yuan." The correlation calculation involves the cosine similarity method, comparing the behavioral pattern vector with the graph entity vector. For example, if the user pattern emphasizes "historical legends," the correlation with the entity "Qu Yuan" might be 0.75. If it is higher than the threshold of 0.6, the next matching step is performed.
[0257] Step S620: If the initial correlation is higher than the preset threshold, the behavior pattern is matched with the entities in the cultural knowledge graph through the content mapping tool to obtain the list of matched related entities and determine the initial priority of the related entities in the list.
[0258] The matching score between behavioral patterns and entities in the cultural knowledge graph is obtained using the following formula:
[0259] (44)
[0260] In formula (44), The score represents the matching score between behavioral patterns and entities in the cultural knowledge graph. Indicates the number of feature dimensions involved in the matching. Indicates the first Weights of each feature dimension, The behavioral pattern is indicated in the first Feature values in each dimension Representing the knowledge graph entity in the th... Feature values in each dimension The function represents the function for calculating the similarity between two feature values.
[0261] The initial priority in the list of associated entities is derived using the following formula:
[0262] (45)
[0263] In formula (45), This indicates the initial priority value of an entity in the associated entity list. Represents the frequency weighting coefficient. This indicates the frequency of an entity's appearance in the cultural knowledge graph. Indicates the importance weight coefficient. Indicators representing the importance of an entity This represents the distance weighting coefficient. This represents the semantic distance between an entity and a behavioral pattern.
[0264] If the initial relevance exceeds a preset threshold, a content mapping tool is used to match behavioral patterns with entities in the cultural knowledge graph, generating a list of matched related entities. This content mapping tool employs graph traversal algorithms such as Breadth-First Search (BFS) to search for relevant entities starting from the behavioral pattern keywords. In the Dragon Boat Festival scenario, the pattern "exploring festival customs" matches entities such as "zongzi making" and "mugwort ornaments," generating a list containing these entities and their matching scores to ensure the list covers the user's potential interests. The initial priority of entities in the related entity list is determined based on their centrality in the graph, such as their PageRank value. For example, the entity "Qu Yuan" has a higher initial priority of 0.8 because it connects to multiple festival nodes, while "Dragon Boat Race" might have a priority of 0.6. This determination lays the foundation for subsequent adjustments.
[0265] Step S630: For the initial priority, use a visual coding feedback tool to extract user interaction response data, and use a sorting tool to dynamically adjust the list of related entities to obtain the adjusted priority sequence.
[0266] The adjusted priority of related entities is derived using the following formula:
[0267] (46)
[0268] In formula (46), Indicates the first The adjusted priority of each related entity Indicates the first The initial priority of each associated entity, This represents the weighting coefficients of the visual encoding feedback. Indicates the first Visual encoding feedback score for each entity, This indicates the total number of user interaction responses. Indicates the first The weight of each interactive response, Indicates the first The entity in the first The rating value in each interactive response.
[0269] Based on the initial priority, a visual encoding feedback tool is used to extract user interaction response data. This tool analyzes user interface interactions such as eye-tracking data, encoding visual attention into feedback vectors. During video playback, if a user's gaze lingers on the "Qu Yuan portrait" for more than 5 seconds, the visual encoding feedback tool extracts response data such as attention duration and location coordinates to quantify the interaction intensity.
[0270] The sorting tool dynamically adjusts the list of related entities to obtain an adjusted priority sequence. This sorting tool applies sorting algorithms such as quicksort and adjusts priorities based on feedback data. For example, it can move the highly popular entity "Qu Yuan" to the top of the list, forming a sequence such as "Qu Yuan-Zongzi-Dragon Boat".
[0271] Step S640: Based on the adjusted priority sequence, refresh the personalized recommendation content in the cultural knowledge graph using the incremental update tool of the content recommendation module, obtain the optimized recommendation combination, and determine the final propagation sequence;
[0272] The incrementally updated cultural knowledge graph is derived using the following formula:
[0273] (47)
[0274] In formula (47), Indicates at time The state of the knowledge graph after incremental updates. This represents the initial state of the basic knowledge graph. Indicates at time Incremental update volume This indicates the number of feature factors involved in the update. Indicates the first The influence coefficient of each update factor Indicates the first Each characteristic factor at time... The change value.
[0275] The propagation sequence is derived using the following formula:
[0276] (48)
[0277] In formula (48), This represents a definite final propagation sequence. Represents the space of all possible combinations of propagation sequences. This represents the total number of content nodes in the sequence. Indicates the first The propagation weight of each node, Indicates in sequence The Middle Quality assessment value of each node Indicates the first The timeliness factor of each node.
[0278] Based on the adjusted priority sequence, the personalized recommendations in the cultural knowledge graph are refreshed using the incremental update tool of the content recommendation module. This content recommendation module is a variant of a recommendation system framework such as collaborative filtering. The incremental update tool modifies only the affected nodes through a delta update mechanism. For the first node in the sequence, "Qu Yuan," the incremental update tool adds new relationships such as "linked to modern poetry," refreshes the graph, and generates personalized content.
[0279] The optimized recommended combination is obtained, and the final propagation sequence is determined. This optimized recommended combination is a set of entities extracted from the updated graph, such as "Qu Yuan legend + Zongzi video". The propagation sequence is defined as a sequential push path, such as pushing historical videos first and then customs demonstrations, to ensure the coherent dissemination of cultural knowledge.
[0280] Please see Figure 2 This embodiment provides a knowledge graph-based cultural digital video content display system for executing the aforementioned knowledge graph-based cultural digital video content display method. It includes a cultural knowledge graph acquisition module 10, a video content organization structure determination module 20, a knowledge-driven video sequence acquisition module 30, an interactive display response judgment module 40, a user navigation path acquisition module 50, and a propagation sequence acquisition module 60. The cultural knowledge graph acquisition module 10 collects multimodal cultural data, processes cultural entities in the multimodal data using named entity recognition, and generates triplet knowledge by combining visual feature encoding and audio semantic alignment to obtain a cultural knowledge graph. The multimodal cultural data includes video documents and oral history records. The video content organization structure determination module 20 analyzes the relationships between entities based on the cultural knowledge graph using relation extraction methods and integrates temporal organization of custom sequences. The system drives the logical coherence of video generation and determines the organizational structure of video content. A knowledge-driven video sequence acquisition module 30 acquires the video content organizational structure, automatically links relevant video segments from entity associations in the cultural knowledge graph, enhances semantic association depth, and obtains a knowledge-driven video sequence. An interactive display response judgment module 40 constructs a dual-view interface of the main video and knowledge graph panel based on the knowledge-driven video sequence, synchronously displaying the attributes of clicked entities and judging the interactive display response. A user navigation path acquisition module 50, if the interactive display response conforms to a transmission relationship, jumps to the relevant video along the transmission relationship, incorporates audio alignment elements to enhance dynamic interaction, and obtains the user navigation path. A propagation sequence acquisition module 60 incrementally updates the personalized content recommendation module in the cultural knowledge graph based on the user navigation path, integrates visual encoding feedback to adjust the priority of associated entities, and obtains an optimized propagation sequence.
[0281] This embodiment provides a knowledge graph-based method and system for displaying cultural digital video content. Compared with existing technologies, it processes cultural entities using named entity recognition, combines visual feature encoding with audio semantic alignment to generate triplet knowledge, and forms a cultural knowledge graph. It then employs relation extraction to analyze relationships between entities, integrates temporal organization of custom sequences, and determines the video content structure. Video segments are automatically linked from the graph to construct a dual-view interface consisting of a main video and a knowledge graph panel, supporting simultaneous display of entity attributes and navigation through inheritance relationships, while incorporating audio alignment to enhance interactivity. The graph is incrementally updated based on the user's navigation path, and visual encoding feedback is used to adjust entity priorities, resulting in an optimized dissemination sequence. This embodiment significantly improves the semantic depth and interactive coherence of cultural data, enabling personalized content recommendations and dynamic dissemination optimization, ultimately promoting an immersive experience and efficient dissemination of cultural heritage.
[0282] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention. Clearly, those skilled in the art can make various alterations and modifications to the invention without departing from its spirit and scope. Thus, if these modifications and modifications of the invention fall within the scope of the claims and their equivalents, the invention is also intended to include these modifications and modifications.
Claims
1. A knowledge graph-based cultural digital video content display method, characterized in that, The method comprises the following steps: S100, collecting cultural multi-modal data, processing cultural entities in the cultural multi-modal data by using a named entity recognition method, generating triples knowledge by combining visual feature coding and audio semantic alignment, and obtaining a cultural knowledge graph, wherein the cultural multi-modal data comprises image documents and oral history records; S200, according to the cultural knowledge graph, analyzing the relationship between entities by using a relation extraction method, fusing the time sequence to organize the order of customs, driving the logical coherence of the video generation, and determining the video content organization structure; S300, obtaining the video content organization structure, automatically linking related video segments from the entity association in the cultural knowledge graph, enhancing the semantic association depth, and obtaining a knowledge-driven video sequence; S400, constructing a main video plus knowledge graph panel double-view interface through the knowledge-driven video sequence, synchronously displaying the attributes when a clicked entity is clicked, and judging an interactive display response; S500, if the interactive display response conforms to the inheritance relationship, jumping to related videos along the inheritance relationship, integrating audio alignment elements to enhance dynamic interaction, and obtaining a user navigation path; S600, according to the user navigation path, incrementally updating a personalized content recommendation module in the cultural knowledge graph, fusing visual coding feedback to adjust the priority of associated entities, and obtaining an optimized propagation sequence; Step S500 comprises: S510, if the interactive display response conforms to the inheritance relationship, according to the jump requirement of the inheritance relationship, using a content mapping tool to extract a target video segment related to the current content from a pre-established video library, and obtaining a corresponding video identifier; Whether the interactive display response conforms to the predefined inheritance relationship structure is quantified by the following formula: ; wherein, represents a heritage relationship conformity score, represents a total number of heritage relationships, represents a weight coefficient of the th heritage relationship, represents an association strength between the current content and the parent content, and the parent content. The target video segment extracted from the video library is obtained by the following formula: ; wherein, represents a target video clip extracted from a video library, represents a pre-established video library set, represents a current query requirement, represents a current content context, represents a similarity between the query and the video , represents a relevance between the content and the video , represents a quality score of the video , , , are corresponding weight parameters, respectively; The following formula is used to generate a unique video identifier by fusing the jump requirement, the content feature and the mapping relationship: ; wherein, denotes the acquired video identifier, denotes a hash mapping function, denotes a feature vector of the jump demand, denotes a feature encoding of the current content, denotes a mapping matrix of the content mapping tool, and denotes a feature fusion operation; S520, for the video identifier, sound feature extraction is performed on the target video segment by using an audio processing tool, preset audio alignment elements are integrated, and adjusted audio-video synchronous content is obtained; The sound feature is extracted from the target video segment by the following formula: ; wherein, represents an extracted audio feature vector, represents a time length of a target video segment, represents a time domain audio signal amplitude, represents a frequency domain weighting function, represents a time domain window function; The fusion processing of the original audio and the preset alignment element is realized by the following formula: ; wherein, represents the fused audio alignment parameter, represents the original audio weight coefficient, represents the original audio parameter vector, represents the preset audio alignment element parameter; S530, if the adjusted audio-video synchronous content conforms to the preset playing standard, a dynamic interaction tool is used to respond to the user's clicking action in real time, the adjusted audio-video synchronous content is loaded into a playing area, and the triggering time of the interactive response is judged; The following formula is used to define the condition that the audio-video synchronous content conforms to the playing standard: ; wherein, represents an audio-video synchronization quality evaluation index, represents an audio timestamp, represents a video timestamp, represents a preset synchronization threshold standard, when the audio-video synchronization quality evaluation index is less than or equal to 1, indicating that the adjusted audio-video synchronization content meets the preset playing standard. The following formula is used to judge the time when the interactive response is triggered to obtain the best user experience: ; wherein, represents an optimal trigger timing of an interaction response, represents a user behavior prediction parameter, represents a buffer status parameter, represents a content loading completion degree parameter, , , respectively represent weight coefficients of each parameter; S540, according to the user operation feedback and the triggering time of the interactive response, a navigation path generation tool is used to construct a jump path from the current content to the target video segment, and a complete user navigation path is obtained; The complete user navigation path is obtained by the following formula: ; wherein, represents the optimal complete user navigation path, represents the path candidate index, represents the total number of path evaluation dimensions, represents the weight factor of the th dimension, represents the score function of the th path on the th dimension, represents the weight parameter of the path history information, represents the history navigation effect evaluation value of the th path.
2. The knowledge graph based cultural digitization video content presentation method as claimed in claim 1, wherein, Step S100 comprises: S110, obtain cultural multi-modal data from image documents and oral history records, separate visual fragments related to cultural entities from image parts of the cultural multi-modal data using an image segmentation tool, and extract text content from audio data in the oral history records by a speech-to-text tool to obtain a preliminary cultural entity information set; S120, according to the preliminary cultural entity information set, process the text content using a named entity recognition tool to extract cultural related entity names in the text content, and generate visual feature encodings by a feature extraction tool in combination with the visual fragments to obtain multi-modal feature descriptions of cultural entities; S130, for the multi-modal feature descriptions of cultural entities, if the semantic alignment matching degree of the visual feature encodings and the audio-to-text content is higher than a preset threshold, generate corresponding triple knowledge through a semantic alignment tool, otherwise, correct the unmatched part using a content completion tool to determine a final triple knowledge set; S140, integrate the triple knowledge set by a knowledge graph construction tool to obtain complete mapping of nodes and edge relationships, and generate a cultural knowledge graph. 3.The knowledge graph based cultural digitalized video content presentation method of claim 1, wherein, Step S200 includes: S210, according to the cultural knowledge graph, analyze cultural entity relationships using a relationship extraction tool to obtain at least one key contact path to obtain a preliminary cultural relationship network; S220, for the preliminary cultural relationship network, integrate the time sequence custom order in the cultural relationship network by a time sequence analysis tool to sort out the time logic of cultural events and determine the time arrangement framework of cultural content; S230, if the semantic consistency of the time arrangement framework and the cultural background matching is higher than a preset threshold, generate a coherent video content fragment using a content arrangement tool in combination with visual content coherence, otherwise, adjust the inconsistent part by a semantic correction tool to obtain an adjusted content fragment; S240, combine the adjusted content fragment with the organization structure construction by a structured arrangement tool, integrate a dynamic presentation mode, and determine the final video content organization structure. 4.The knowledge graph-based cultural digital video content presentation method of claim 1, wherein, Step S300 includes: S310, according to the cultural knowledge graph, automatically match the relevance between entities using an entity linking tool to obtain at least one associated video fragment from the video content organization structure to obtain a preliminary semantic association fragment group; S320, for the preliminary semantic association fragment group, sort and filter the association strength between the preliminary semantic association fragment group fragments by a semantic depth enhancement tool, if the association strength between the semantic association fragment group fragments is lower than a preset threshold, eliminate irrelevant fragments to obtain a selected semantic association fragment combination; S330, adjust the time sequence of the selected semantic association fragment combination using a time sequence integration tool, determine the order of fragment playback in combination with the time sequence logic in the cultural knowledge graph, and obtain an ordered video fragment sequence; S340, judging whether the logic consistency of the whole knowledge driving is met by performing semantic alignment on the ordered video segment sequence and the background information in the cultural knowledge graph through a content fusion tool, to obtain a final knowledge driving video sequence. 5.The knowledge graph based cultural digitalized video content presentation method of claim 1, wherein, Step S400 includes: S410, dividing a main video display area and a knowledge graph panel area according to the knowledge driving video sequence, obtaining a divided interface layout structure, and determining an initial display framework by using an interface construction tool; S420, loading video sequence content and graph entity associated data into corresponding areas respectively through a content loading tool for the initial display framework, to obtain a loaded content distribution view; S430, detecting entity click response actions on the basis of the content distribution view by using an interactive response tool, triggering attribute synchronous display functions if the click actions are detected, obtaining related attribute data of the clicked entity, and judging a display position in the knowledge graph panel; S440, matching and optimizing the display position and user interaction logic through a dynamic adjustment tool, obtaining an adjusted interface display effect, and determining a final interactive display response. 6.The knowledge graph based cultural digitalized video content presentation method of claim 1, wherein, Step S600 includes: S610, obtaining corresponding user behavior data from a pre-established database according to a user navigation path, dividing behavior patterns for the user behavior data by using a classification tool, and determining a preliminary correlation degree with cultural knowledge graph content; S620, if the preliminary correlation degree is higher than a preset threshold, matching the behavior patterns and entities in the cultural knowledge graph through a content mapping tool, obtaining a matched associated entity list, and judging an initial priority in the associated entity list; S630, extracting user interaction response data for the initial priority by using a visual coding feedback tool, dynamically adjusting the associated entity list through a sorting tool, and obtaining an adjusted priority sequence; S640, refreshing personalized recommendation content in the cultural knowledge graph through an incremental update tool of a content recommendation module according to the adjusted priority sequence, obtaining an optimized recommendation combination, and determining a final propagation sequence. 7.A knowledge graph-based cultural digital video content presentation system configured to perform the knowledge graph-based cultural digital video content presentation method according to any one of claims 1 to 6. It includes: A cultural knowledge graph acquisition module (10) is configured to collect cultural multi-modal data, process cultural entities in the cultural multi-modal data by using a named entity recognition method, generate triples knowledge by combining visual feature coding and audio semantic alignment, and obtain a cultural knowledge graph. The cultural multi-modal data includes image documents and oral history records. A video content organization structure determination module (20) is configured to analyze the relationship between entities by using a relationship extraction method according to the cultural knowledge graph, fuse the order of customs according to time sequence, drive video generation logic coherence, and determine a video content organization structure. A knowledge driving video sequence acquisition module (30) is configured to obtain the video content organization structure, automatically link related video segments from the cultural knowledge graph according to entity association, enhance semantic association depth, and obtain a knowledge driving video sequence. The interactive display response judgment module (40) is configured to construct a main video plus knowledge graph panel double-view interface through the knowledge-driven video sequence, display attributes when a clicked entity is displayed synchronously, and judge an interactive display response; The user navigation path acquisition module (50) is configured to jump to a related video along a heritage relationship if the interactive display response meets the heritage relationship, integrate an audio alignment element to enhance dynamic interaction, and obtain a user navigation path. The propagation sequence acquisition module (60) is configured to incrementally update a personalized content recommendation module in the cultural knowledge graph according to the user navigation path, integrate visual coding feedback to adjust a priority of an associated entity, and obtain an optimized propagation sequence.
Citation Information
Patent Citations
Lingnan culture visual display method and system based on mapping knowledge domain
CN117634605A
Integrated multi-mode culture resource intelligent data governance and management system
CN120973989A