Semantic and situational knowledge collaborative modeling declarative knowledge construction method and device, computer equipment and readable storage medium
By constructing a dual-graph system that integrates a semantic network and a contextual graph, the shortcomings of traditional knowledge graphs in processing high-order declarative knowledge are addressed. This enables deep integration of multimodal knowledge and semantic contextual collaboration, thereby improving the accuracy and practicality of knowledge representation.
Patent Information
- Application Number
- CN202511000750.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-21
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-07-21
AI Technical Summary
Traditional knowledge graph construction methods suffer from shallow semantic parsing, lack of cross-modal associations, separation of knowledge nodes from context, and insufficient multi-dimensional information collaborative modeling when dealing with high-order declarative knowledge. This results in insufficient three-dimensional expression of knowledge elements and poor domain adaptability.
We adopt a method of co-modeling semantic and contextual knowledge. By parsing multimodal documents and extracting chapter summaries, we construct a dual graph that combines semantic network and contextual graph. We perform graph community network clustering and associate communities, entities, and events with chapter summaries and multimodal knowledge points to form associated knowledge. Finally, we perform vectorization and establish a vector knowledge index.
It achieves deep fusion of multimodal knowledge and collaborative modeling of semantic context, improves the degree of knowledge structuring and application reliability, and enhances the accuracy and practicality of knowledge representation.
Smart Images

Figure CN120975199A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of data processing, in particular to a declarative knowledge construction method and device for semantic and situational knowledge collaborative modeling, computer equipment and readable storage medium. BACKGROUND
[0002] The traditional knowledge graph construction method has systematic bottlenecks in processing high-order declarative knowledge: semantic analysis is shallow and cross-modal correlation is missing, resulting in insufficient three-dimensional expression of knowledge elements; knowledge nodes and context scenes are separated, dynamic scene elements are not effectively coupled with semantic units, causing cross-scene semantic adaptation barriers; multi-dimensional information collaborative modeling is insufficient, and the contradiction between static representation and dynamic learning needs is prominent, with poor field adaptability. Therefore, it is urgent to propose a declarative knowledge construction method that integrates semantic and situational features to improve the accuracy and practicality of knowledge representation. SUMMARY
[0003] The present application relates to the field of data processing, in particular to a declarative knowledge construction method and device for semantic and situational knowledge collaborative modeling, computer equipment and readable storage medium.
[0004] In a first aspect, the present application provides a declarative knowledge construction method for semantic and situational knowledge collaborative modeling, comprising:
[0005] The multi-modal document is parsed and chapter abstracts are extracted to obtain structured content and chapter abstracts;
[0006] Based on the structured content, entities, events and multi-modal knowledge points are extracted and cross-modal fusion is performed;
[0007] Based on the fused multi-modal knowledge points, entity and event resolution is performed to construct a dual graph that cooperates semantic network and situational graph, the dual graph includes knowledge and affair graph and four-dimensional relationship triplets;
[0008] The dual graph is clustered into a graph community network, and communities related to topics or fields are divided;
[0009] The communities, triplets, entities and events are respectively associated with the chapter abstracts and multi-modal knowledge points to form associated knowledge;
[0010] The associated knowledge is vectorized and a vector knowledge index is established and stored in a knowledge vector library;
[0011] The declarative knowledge in the vector knowledge index is evaluated for compression rate and accuracy by a declarative knowledge evaluation system, and the process of multi-modal knowledge point fusion, dual graph construction, graph community network clustering and association mapping is optimized according to the evaluation results.
[0012] In a second aspect, the embodiment of the present application provides a declarative knowledge construction device for semantic and situational knowledge collaborative modeling, comprising:
[0013] The acquisition module is configured to parse a multi-modal document and extract a chapter summary to obtain structured content and a chapter summary; extract entities, events and multi-modal knowledge points from the structured content and perform cross-modal fusion;
[0014] The construction module is configured to perform reference resolution of the entities and events based on the fused multi-modal knowledge points, construct a dual graph atlas of a semantic network and a situational graph atlas, and the dual graph atlas comprises a knowledge graph atlas, a matter graph atlas and four-dimensional relationship triplets; perform graph community network clustering on the dual graph atlas to divide a community related to a theme or a field; associate and map the community, the triplets, the entities and the events with the chapter summary and the multi-modal knowledge points respectively to form associated knowledge; perform vectorization processing on the associated knowledge and establish a vector knowledge index to store in a knowledge vector library; and perform compression rate and accuracy rate evaluation on declarative knowledge in the vector knowledge index through a declarative knowledge evaluation system, and feed back the evaluation results to optimize the processes of multi-modal knowledge point fusion, dual graph atlas construction, graph community network clustering and association mapping.
[0015] In a third aspect, the embodiment of the present application provides a computer device, comprising a processor and a non-volatile memory storing computer instructions, and when the computer instructions are executed by the processor, the computer device executes the method of the first aspect.
[0016] In a fourth aspect, the embodiment of the present application provides a readable storage medium comprising a computer program, and when the computer program runs, controls a computer device where the readable storage medium is located to execute the method of the first aspect.
[0017] Compared with the prior art, the present application has the following beneficial effects: by using the declarative knowledge construction method, device, computer device and readable storage medium for semantic and situational knowledge collaborative modeling disclosed by the present application, the multi-modal document is parsed and the chapter summary is extracted to obtain the structured content; the entities, events and multi-modal knowledge points are extracted from the structured content and cross-modal fusion is performed; the reference resolution is performed based on the fused knowledge points to construct a dual graph atlas comprising a knowledge graph atlas, a matter graph atlas and four-dimensional relationship triplets; the dual graph atlas is clustered to obtain a theme community, and the community, triplets, entities and events are associated and mapped with the chapter summary and multi-modal knowledge points to form associated knowledge; the associated knowledge is vectorized and a vector knowledge index is established; the knowledge is evaluated for compression rate and accuracy rate by an evaluation system, and the whole knowledge construction process is fed back and optimized. The present application realizes multi-modal knowledge deep fusion and semantic and situational collaborative modeling, and improves the structured degree and application reliability of the knowledge. Attached Figure Description
[0018] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of the present invention and should not be considered as limiting the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 A flowchart illustrating the steps of the declarative knowledge construction method for co-modeling semantic and contextual knowledge provided in this embodiment of the invention;
[0020] Figure 2 A schematic diagram of the construction framework of the declarative knowledge construction system provided in the embodiments of the present invention;
[0021] Figure 3 This is a schematic flowchart of the phased collaborative construction process of dual-maps provided in an embodiment of the present invention;
[0022] Figure 4 A schematic block diagram of the declarative knowledge construction device for co-modeling semantic and contextual knowledge provided in an embodiment of the present invention;
[0023] Figure 5 A schematic block diagram of the structure of a computer device provided in an embodiment of the present invention. Detailed Implementation
[0024] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0025] The specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings.
[0026] In order to solve the technical problems mentioned in the background art Figure 1 This is a flowchart illustrating the declarative knowledge construction method for co-modeling semantic and contextual knowledge provided in this embodiment of the disclosure. The following is a detailed description of this declarative knowledge construction method for co-modeling semantic and contextual knowledge.
[0027] Step S201: Parse the multimodal document and extract chapter summaries to obtain structured content and chapter summaries;
[0028] Step S202: Extract entities, events, and multimodal knowledge points based on the structured content and perform cross-modal fusion;
[0029] Step S203: Based on the fused multimodal knowledge points, entity and event referencing is resolved, and a dual graph of semantic network and context graph is constructed. The dual graph includes knowledge and event graph and four-dimensional relation triplet.
[0030] Step S204: Perform graph community network clustering on the dual graphs to divide them into communities related to themes or domains;
[0031] Step S205: Associate and map the community, the triples, entities, and events with the chapter summary and the multimodal knowledge points respectively to form associated knowledge;
[0032] Step S206: Vectorize the associated knowledge and establish a vector knowledge index, then store it in the knowledge vector library;
[0033] Step S207 involves evaluating the compression rate and accuracy of the declarative knowledge in the vector knowledge index using a declarative knowledge evaluation system, and then optimizing the process of multimodal knowledge point fusion, dual-graph construction, graph community network clustering, and association mapping based on the evaluation results.
[0034] In this embodiment of the invention, for example, a server is used as the execution entity to apply this method to construct a declarative knowledge system in the field of traffic accident liability determination. The specific implementation process is as follows:
[0035] The server first initiates a multimodal data processing flow, receiving heterogeneous documents from multiple sources related to traffic accident scenarios. These documents include accident scene surveillance videos (MP4 format), handwritten investigation reports from traffic police (scanned PDFs), vehicle speed detection photos (JPG format), voice transcripts from involved parties (WAV format), and relevant legal and regulatory provisions (Word format). During the data collection and preprocessing phase, the server cleans the raw documents, removing redundant frames from the videos, scanning noise from the PDFs, silent segments from the audio, and irrelevant formatting characters (such as extra spaces and garbled text) from the text, ensuring data quality. Subsequently, the server performs modality recognition and separation operations: using a pre-trained modality classification model (based on a fusion architecture of ResNet-50 and BERT) to identify document types and separate modal content such as video, images, audio, and text—for example, automatically distinguishing between surveillance video streams, PDF image layers containing handwritten text, JPG images recording vehicle speed data, and plain text legal and regulatory provisions from mixed-format accident files.
[0036] For text modalities, the server uses a natural language processing toolchain for parsing: it performs word segmentation (Jieba word segmentation tool), part-of-speech tagging (LTP model), and named entity recognition (BERT-based Chinese NER model) on legal and regulatory Word documents to extract entities such as "motor vehicle", "non-motor vehicle", and "speed limit sign", while uniformly encoding the text into UTF-8 format; for the PDF image layer of the traffic police's handwritten report, it converts it into text through an OCR engine (Tesseract-OCR optimized version), and then corrects the recognition error through context semantic correction (based on BiLSTM-CRF model) to obtain structured information such as "Accident time: 9:20 am on October 15, 2023" and "Parties involved: A (license plate number A12345), B (license plate number B67890)". For non-text modalities, the server calls a dedicated parsing module: For surveillance video, it extracts 10 keyframes per second using a 3D convolutional network (C3D model), and then uses an object detection algorithm (YOLOv8) to identify visual entities such as "Vehicle A", "Vehicle B", "road markings", and "traffic lights" in the frames, simultaneously recording the entity's location coordinates and timestamps; For vehicle speed photos, it extracts the digital information "130km / h" from the photos using image text detection (EAST algorithm), and combines it with image metadata to confirm that the shooting time was 3 seconds before the accident; For the parties' voice transcripts, it uses a speech recognition model (WeNet) to convert WAV format audio into text, and then uses semantic role labeling (SRL) to extract key statements such as "A said: I did not notice the speed limit sign" and "B said: The other party suddenly changed lanes".
[0037] After completing multimodal document parsing, the server enters the chapter summary extraction stage. For legal and regulatory Word documents, the server automatically divides them into chapters based on heading levels (e.g., "Chapter 1 General Provisions," "Chapter 2 Regulations on Motor Vehicle Traffic"). For traffic police report texts without a clear structure, the server calculates paragraph similarity using the TextRank algorithm (with a cosine similarity threshold of 0.65), clustering semantically coherent paragraphs into chapters such as "Basic Accident Information," "Party Statements," and "On-site Investigation Records." Subsequently, the server extracts key features from each chapter: for the "Regulations on Motor Vehicle Traffic" chapter, the server uses the TF-IDF algorithm to filter keywords such as "speed limit," "overtaking," and "yielding," and combines this with the BERT model to generate sentence vectors, selecting the top 30% of sentences with the highest vector cosine similarity as candidate key sentences; for the surveillance video chapter, the server extracts the frequency of entity occurrences in keyframes (e.g., "Vehicle A" appears 85 times) and changes in motion trajectory (e.g., "Vehicle B's acceleration was -5 m / s² 2 seconds before the collision"). 2Visual features such as “” are used to generate chapter summaries. Based on these features, the server calls a pre-trained large language model (Qwen3-7B) to generate chapter summaries: Inputting the text and keywords of the “Motor Vehicle Traffic Regulations” chapter, the model outputs “This chapter stipulates that motor vehicles must obey speed limit signs when driving on roads, confirm a safe distance when overtaking, and slow down and yield at intersections”; inputting keyframe features and timestamps from surveillance video, the model outputs “Surveillance shows: 2023-10-15 09:19:58, vehicle B entered a 120km / h speed limit section at 130km / h; 09:20:00, vehicle A suddenly changed lanes to the left, vehicle B's brakes failed, and the two vehicles collided at 09:20:01.” Finally, the server optimizes and integrates the generated summaries, adjusting sentence fluency using a grammar correction tool (LangCorrect) and checking the logical consistency of each chapter summary (e.g., “accident time” is “09:20:01” in both the report summary and the video summary), forming a structured storage association between the content and the chapter summaries.
[0038] Based on the structured content obtained in step one, the server initiates a multimodal knowledge point extraction process. For text-based legal provisions, the server designs a dedicated Prompt template ("Extract the core knowledge points corresponding to the legal provisions from the following text, in the format 'Knowledge Point Name: Specific Content'") and calls the DeepSeek-7B large model for processing. For example, inputting the corresponding legal provision XX, "Motor vehicles shall not exceed the maximum speed limit indicated by the speed limit sign when driving on the road," the model outputs the knowledge point "Speeding: The speed of a motor vehicle exceeds the maximum speed limit indicated by the speed limit sign." For image-based vehicle speed detection photos, the server extracts the knowledge point "Vehicle B's speed: 130km / h" through a combination of object detection and OCR. For video-based keyframe sequences, the server extracts the knowledge point "Collision event: Vehicle A and Vehicle B had a head-on collision at 09:20:01 on 2023-10-15" through temporal relation reasoning (an event detection model based on LSTM). For speech-based transcripts, the server extracts the knowledge point "Party A's statement: I did not observe Vehicle B speeding" through a combination of sentiment analysis (TextCNN model) and keyword extraction.
[0039] After extracting multimodal knowledge points, the server performs cross-modal fusion. For knowledge points related to "speeding," the server constructs a fusion rule base: it associates the legal definition of "speeding" in the text (knowledge point A), "vehicle B's speed is 130 km / h" in the image (knowledge point B), and "the speed limit sign on this road section shows 120 km / h" in the video (knowledge point C). Through semantic similarity calculation (BERT sentence vector cosine similarity 0.92), it confirms that the three point to the same core concept, and fuses them to generate a comprehensive knowledge point: "Vehicle B is speeding: its speed is 130 km / h, exceeding the maximum speed limit of 120 km / h on this road section, violating relevant regulations XX." For knowledge points related to "collision cause," the server fuses "vehicle A suddenly changes lanes" in the video (visual knowledge point) and "vehicle A claims not to have observed" in the audio (textual knowledge point), and generates a fused knowledge point: "vehicle A's failure to observe caused a sudden lane change, which, together with the speeding vehicle B, caused a collision due to insufficient avoidance." During the fusion process, the server records the confidence scores of each modality's knowledge points (0.95 for text knowledge points, 0.90 for image knowledge points, and 0.88 for video knowledge points). A weighted average method (with weights allocated according to modal reliability: 0.4 for text, 0.3 for image, and 0.3 for video) is used to calculate the overall confidence score of the fused knowledge points (e.g., the confidence score of the fused knowledge point "speeding" is 0.93) to ensure the reliability of the fusion results.
[0040] Based on the fused multimodal knowledge points, the server initiates the entity and event referencing resolution process. First, it constructs an entity and event set: extracting entities (such as "vehicle A", "vehicle B", "Zhang San (driver A)", "Li Si (driver B)", and "speed limit sign") and events (such as "speeding", "sudden lane change", "brake failure", and "collision") from the fused knowledge points, forming an initial entity type set T. e ^(0) = {vehicles, drivers, road facilities} and event type set T v^(0) = {speeding, lane changing, brake failure, collision}. Subsequently, the server handles cross-document referential ambiguity: for example, in a traffic police report, "driver of vehicle A" is described as "Zhang San," while in the party's statement it is described as "Jia"; "speeding" is described as "excessive speed" in video analysis, but as "exceeding the speed limit" in legal texts. The server calls a Transformer-based semantic encoder (Sentence-BERT) to generate embedding vectors for these referentials and calculates cosine similarity: sim("Zhang San", "Jia") = 0.91 (> threshold 0.85), sim("speeding", "excessive speed") = 0.89 (> threshold 0.85). For candidate matches, the server performs metadata verification: the ID card number and driver's license number of "Zhang San" and "Jia" are completely consistent in the system database; the timestamps (09:19:58-09:20:01) and speed values (130km / h>120km / h) corresponding to "speeding" and "excessive speed" are both matched. For key nodes (such as the timestamp of the "collision occurred" event), the server triggers a manual verification process, where traffic experts confirm that "2023-10-15 09:20:01" is the accurate collision time, and finally outputs the optimized entity type set T. e ^ = {vehicles (including subtypes: cars, trucks), drivers (including attributes: name, driver's license number), road facilities (including subtypes: speed limit signs, traffic lights)} and event type set T v ^={Speeding (including attributes: speed, speed limit), sudden lane change (including attributes: lane change direction, duration), brake failure (including attributes: braking start time, deceleration), collision (including attributes: time, location, collision type)}.
[0041] Based on the optimized set of entity and event types, the server constructs a dual-graph system that integrates a semantic network (KG) and a contextual graph (EG). By designing a Promptv3.0 template ("Based on the following entity / event type sets, extract four-dimensional relation triples from text: entity-entity (KG-KG), entity-event (KG-EG), event-entity (EG-KG), event-event (EG-EG), in the format (head entity / event, relation, tail entity / event)"), a large language model is invoked to process the entire corpus. For example, extracting from surveillance videos and report texts:
[0042] Entity-Entity Relationship (KG-KG): ("Vehicle A (small car)", "Own Driver", "Zhang San (Driver)"), ("Speed Limit Sign", "Location", "Accident Section K12+300");
[0043] Entity-Event Relationship (KG-EG): ("Zhang San (Driver)", "Implementation", "Sudden Lane Change (Event)"), ("Vehicle B (Sedan)", "Experience", "Brake Failure (Event)");
[0044] Event-Entity Relationship (EG-KG): ("Collision Occurred (Event)", "Vehicles Involved", "Vehicle A (Sedan)"), ("Speeding (Event)", "Violation of Regulations", "Relevant Regulations XX");
[0045] Event-Event Relationships (EG-EG): ("Speeding (Event)", "Causes", "Brake Failure (Event)"), ("Sudden Lane Change (Event)", "Triggers", "Collision Occurs (Event)").
[0046] The server stores these triples as graph-structured data, where entity nodes contain type and attributes (e.g., attributes of "Vehicle B": license plate number B67890, speed 130km / h), event nodes contain type and attributes (e.g., attributes of "speeding": start time 09:19:58, duration 3 seconds), and relation edges contain relation type and confidence (calculated based on the comprehensive confidence of fused knowledge points, e.g., the relation confidence of "speeding → causing → brake failure" is 0.87).
[0047] The server invokes the Leiden community discovery algorithm to perform cluster analysis on the constructed bi-graph. First, the bi-graph is represented as a weighted undirected graph G = (V, E), where nodes V contain entity nodes (e.g., "Vehicle A", "Zhang San") and event nodes (e.g., "speeding", "collision occurred"), and the weights of edges E are based on the confidence of the relationships between nodes (e.g., the edge weight between "Vehicle B" and "speeding" is 0.92, and the edge weight between "speeding" and "brake failure" is 0.87). The server initializes each node as an independent community and iteratively optimizes the modularity function (Q = Σ[(e...]). i ia i 2 )], where e i i represents the percentage of edge weights within community i, and a i Adjust node affiliation (based on the percentage of edge weights for node i in community): When a node moves from community A to community B, resulting in a modularity increment ΔQ > 0, a move operation is performed. After 50 iterations, the modularity converges to 0.58 (> the threshold of 0.5), forming 3 highly cohesive communities:
[0048] Community 1: Accident Liability Elements Community: Includes event nodes "speeding", "sudden lane change", and "brake failure" and entity nodes "relevant regulations" and "speed limit sign". The average edge weight between nodes is 0.85, and the core event is "speeding" (occurring 42% of the frequency).
[0049] Community 2: Vehicle Dynamics Analysis Community: Contains entity nodes "Vehicle A", "Vehicle B", "Zhang San", "Li Si" and event nodes "Lane Change Start" and "Braking Start". The average edge weight between nodes is 0.81. The core entity is "Vehicle B" (with an average connection weight of 0.83 with other nodes).
[0050] Community 3: Accident Consequence Assessment Community: Includes event nodes "Collision Occurred", "Vehicle Damage", and "Personnel Injury" as well as entity nodes "Accident Section K12+300" and "Insurance Company". The average edge weight between nodes is 0.78, and the core event is "Collision Occurred" (12 associated nodes).
[0051] The server further analyzes cross-community relationships: through the shared entity "vehicle A", community 2 (vehicle dynamic analysis) and community 3 (accident consequence assessment) are connected, revealing the relationship between "vehicle A's lane change trajectory deviation" and "collision damage level"; through the shared event "collision occurred", community 3 and community 1 (accident liability elements) are connected, supporting the logical link between "liability determination" and "consequence assessment".
[0052] The server establishes association mappings between clustered communities, triples, and entities / events and chapter summaries and multimodal knowledge points, respectively. For the association between communities and chapter summaries, the server uses an attention mechanism to calculate the similarity between the semantic tags of the community and the keywords of the chapter summary: the semantic tag of community 1 is "basis for liability determination," and its overlap with the keywords "speed limit," "yield," and "responsibility" in the legal regulations chapter summary "motor vehicle traffic regulations" reaches 0.79 (>threshold 0.75), thus establishing a mapping relationship; the semantic tag of community 2 is "vehicle behavior analysis," and its overlap with the keywords "lane change," "speed," and "braking" in the surveillance video chapter summary "vehicle dynamic trajectory" reaches 0.82, also establishing a mapping relationship.
[0053] For the association between entities / events and multimodal knowledge points, the server uses knowledge embedding technology (TransE model) to convert entity / event nodes into low-dimensional vectors and calculates the L2 distance with the knowledge point vectors: the vector of the entity node "speed limit sign" has a vector distance of 0.28 (< threshold 0.3) with the knowledge point "the maximum speed limit on this road section is 120km / h", thus establishing an association; the vector of the event node "speeding" has a vector distance of 0.25 with the knowledge point "vehicle B is traveling at a speed of 130km / h", thus establishing an association. For the association between triples and knowledge points, the server achieves this through rule matching: the semantic matching degree between the triple ("speeding", "caused", "brake failure") and the fused knowledge point "vehicle B's brakes failed due to speeding" reaches 0.91, thus establishing an association. Finally, the server generates an associated knowledge network, where each community, triple, and entity / event is linked to the corresponding chapter summary and multimodal knowledge point, forming a multi-level association structure of "community-summary-knowledge point-entity / event-triple".
[0054] The server vectorizes the associated knowledge using a GraphSAGE graph neural network to generate high-dimensional vectors for entities / events. Taking two graphs as input, it samples neighbor nodes (each node samples second-order neighbors, and each layer samples 10 nodes) and aggregates neighbor features (using the mean aggregation function) to output a 128-dimensional vector. For example, the cosine similarity between the "speeding" event vector and the "excessive speed" speech knowledge point vector reaches 0.90, and the cosine similarity between the "vehicle B" entity vector and the "130km / h" image knowledge point vector reaches 0.88.
[0055] Based on the vectorized results, the server builds a hierarchical intelligent index:
[0056] Global Index: Stores vectors of all entities / events, supporting global retrieval across communities (e.g., querying "all events involving speeding");
[0057] Community Index: Vector subsets are divided by community. For example, the vector subset of community 1 only contains nodes related to "responsibility elements", which supports fast retrieval within the domain (such as querying "legal basis for accident liability determination").
[0058] Node Index: An inverted index is built for each entity / event vector, recording its associated chapter summary, knowledge points and triple ID, supporting fine-grained queries (such as querying "the specific time when vehicle B was speeding").
[0059] The server stores hierarchical indexes in a knowledge vector library (using the Milvus vector database) and configures a dynamic scaling strategy (automatic sharding when the number of vectors exceeds 1 million) to ensure retrieval latency is less than 100ms.
[0060] The server initiates a declarative knowledge assessment system to evaluate the compression rate and accuracy of the knowledge in the vector knowledge index. The compression rate assessment is achieved through information entropy calculation: the information entropy of the original multimodal data (videos, documents, images, etc.) is H1 = 5.2 bits, and the information entropy of the associated knowledge after knowledge point fusion and graph construction is H2 = 1.04 bits. The calculated compression rate CR = 1 - H2 / H1 = 80% (the original data compression rate is only 35%), indicating a significant reduction in information redundancy and an improvement in fidelity.
[0061] Accuracy evaluation was conducted using a scoring function based on a knowledge graph embedding model (TransE), combined with a manually labeled test set (containing 1000 traffic accident triples): for each triple (h, r, t), the model outputs a score Score(h, r, t), and a score > 0.7 is considered correct. Test results showed that the accuracy of the manually labeled data reached 91.5%, with scores of 0.82 and 0.79 respectively for causal chains such as "speeding → leading to → brake failure" and "sudden lane change → triggering → collision" (both > threshold 0.7), indicating good semantic consistency.
[0062] The server feeds the evaluation results back to the knowledge construction process: For the "Vehicle Damage Level" community (compression rate 65% < 80%) where the compression rate is below standard, the multimodal knowledge point fusion strategy is optimized, and cross-modal fusion of image damage assessment reports and repair records is added; for the triples with low accuracy ("Li Si → Implementation → Braking Failure", score 0.68 < 0.7), the referential resolution process is re-executed to correct the association error between "Li Si" and "Vehicle B Driver"; for the "Personnel Injury" sub-community where the modularity in community clustering is below standard, the resolution parameter of the Leiden algorithm is adjusted (increased from 1.0 to 1.2) to optimize the community partitioning results. Through closed-loop feedback, the server continuously iterates and optimizes each stage of knowledge construction, ultimately forming a high-quality declarative knowledge system for determining traffic accident liability.
[0063] In this embodiment of the invention, the step of parsing the multimodal document and extracting chapter summaries to obtain structured content and chapter summaries can be implemented through the following example.
[0064] Data collection and preprocessing of raw multimodal documents, cleaning up useless information and standardizing character encoding formats;
[0065] Modality recognition and separation are performed on the preprocessed document to separate text, table, image, audio and video modal content, and non-text modalities are converted into text descriptions through image recognition, speech recognition or video analysis technology;
[0066] The document is divided into chapters based on its structure or semantic coherence. Keywords and key sentences are extracted using a pre-trained language model to generate chapter summaries, and the text descriptions and chapter summaries are integrated into the structured content.
[0067] In this embodiment of the invention, executively speaking, with the server as the executing entity, in the multimodal document processing flow in the field of traffic accident liability determination, the server first performs data collection and preprocessing operations on the original multimodal documents. The server receives the traffic accident file package through the government data interface, which includes accident scene monitoring video (25 minutes in MP4 format, 30fps), a traffic police handwritten investigation report (15-page scanned PDF, including signatures and alterations), vehicle damage assessment report (Excel spreadsheet), voice transcripts of both parties (10 minutes each in WAV format, including ambient noise), and relevant legal provisions (2021 revised edition Word document). The server initiated a data cleaning module, which used inter-frame differencing (threshold set to 0.15) to remove consecutive similar redundant frames from video files, retaining the critical video footage within 5 minutes before and after the accident (a total of 3000 frames); for scanned PDF documents, Gaussian filtering (kernelsize = 3×3) was used to remove scanning noise, and morphological processing (dilation operation iterations = 2) was used to enhance the edges of handwritten text; for audio files, short-time energy thresholding (energy threshold set to 0.02) was used to remove silent segments, retaining a total of 18 minutes of effective audio; and for text documents, a unified character encoding conversion was performed, converting the original GBK-encoded Word documents to UTF-8 format, and using regular expressions to remove special symbols and redundant spaces from the text (continuous spaces were compressed into single spaces).
[0068] After the preprocessing is completed, the server performs modality recognition and separation based on an improved CLIP model (Contrastive Language-Image Pretrained model). The model takes the document byte stream features as input and outputs modality classification probabilities: for a 2.3GB file in the archive, it is recognized as the video modality (probability 0.98); for a 1.2MB scanned document, it is recognized as the image modality (probability 0.96); for an 800KB tabular file, it is recognized as the structured data modality (probability 0.99); for a 45MB audio file, it is recognized as the audio modality (probability 0.97); for a 300KB legal provision, it is recognized as the plain text modality (probability 1.0). For non-text modalities, the server invokes a dedicated conversion toolchain: for the handwritten investigation report in the image modality, the Tesseract-OCR engine (loaded with a Chinese handwritten font training set) is used for text recognition, and the context semantic correction model (BiLSTM-CRF architecture, with 500,000 traffic police handwritten samples in the training data) is used to correct the recognition errors, correcting the misrecognition of "Driver A did not yield" as "Driver A did not drive and yield" to the correct text; for the party's transcript in the audio modality, it is converted to text through the WeNet speech recognition model (pre-trained on 100,000 hours of Chinese dialogue corpus), and the speaker segmentation result is output synchronously (distinguishing the voices of Party A, Party B, and the recorder); for the video modality, the FFmpeg tool is used to extract 2 key frames per second, and visual entities such as "motor vehicle", "non-motor vehicle", "traffic signal" in the frame are recognized through the YOLOv8 object detection model (confidence threshold 0.85), and the optical flow method is combined to calculate the entity movement trajectory, generating a text description of "During the period from 09:19:55 to 09:20:05, Vehicle A drove straight along the main road, and Vehicle B merged into the main road from the side road"; for the loss assessment form in the tabular modality, the PyPDF2 tool is used to extract the table structure, and structured data such as "The front bumper of Vehicle A is damaged" and "Repair cost: 5000 yuan" are converted into key-value pair text descriptions.
[0069] Based on the separated modal content, the server performs chapter division and summary generation. For plain text legal document Word documents, the server parses the Document Object Model (DOM) structure and automatically divides them into 8 chapters such as "General Provisions," "Motor Vehicle Traffic Regulations," and "Traffic Accident Handling" according to the heading style ("Level 1 Heading," "Level 2 Heading"). For traffic police report texts without a clear structure, the server uses the TextRank algorithm to calculate the paragraph similarity matrix (with a window size of 5 sentences), and clusters paragraphs with a semantic similarity higher than 0.65 into 4 chapters: "Basic Accident Information," "On-site Investigation Record," "Statement of Parties Involved," and "Preliminary Liability Determination." After the chapter division is completed, the server calls the Qwen3-7B pre-trained language model to generate chapter summaries: For the chapter "Motor Vehicle Traffic Regulations", the input text content and the Top 10 keywords such as "speed limit", "yield", and "overtaking" extracted by TF-IDF are used, and the model outputs "This chapter clarifies the speed limit standards for motor vehicles on different road types, stipulates that a safe distance must be confirmed when overtaking, and vehicles on auxiliary roads should give way to vehicles on main roads"; For the text description of video modality conversion, the input keyframe entity motion trajectory data and timestamps are used, and the model outputs "The monitoring video shows: 09:19:58, vehicle B merges from the auxiliary road into the main road without using the turn signal; 09:20:00, vehicle A brakes suddenly; 09:20:01, the two vehicles collide in the middle lane of the main road". Finally, the server associates and integrates the text descriptions of each modality conversion (such as handwritten report OCR text, speech-to-text text, and video trajectory description) with the corresponding chapter summaries, establishing a mapping relationship through chapter IDs. This forms structured content with a three-level structure of "chapter title - summary text - multimodal original description," which is stored in the "traffic_accident_chapter" table of a relational database (MySQL), with the field "chapter." i d (primary key), chapter_title, summary_text, modal_type (text / image / video / audio), original_content (original converted text).
[0070] In this embodiment of the invention, the extraction of entities, events, and multimodal knowledge points based on the structured content and the cross-modal fusion can be implemented through the following examples.
[0071] Based on a pre-trained large language model, deep semantic understanding and analysis are performed on the text modal content in the structured content, extracting entities, events and implicit knowledge points from the text modal. The entities include entity names and attribute features, the events include event types and feature descriptions, and the implicit knowledge points include association rules or background knowledge that are not explicitly stated in the text but are obtained through contextual semantic reasoning.
[0072] The entity attribute features, event feature descriptions, and implicit knowledge points extracted from the text modality are respectively integrated with the visual features of the image modality, the event voiceprint features of the audio modality, and the action sequence features of the video modality in the structured content. By associating the object features described in the text with the visual features of image recognition, matching the feature descriptions of events in the text with the event features in the audio or video, and using the implicit knowledge points to supplement the semantic associations between cross-modal features, the complete multimodal knowledge points are formed.
[0073] In this embodiment of the invention, for example, a server acts as the execution entity. In a traffic accident liability determination scenario, the server initiates a multimodal knowledge point extraction and fusion process based on structured content. For text modalities within the structured content (such as traffic police report OCR text, legal provisions, and speech-to-text text), the server calls the DeepSeek-7B pre-trained large language model to perform deep semantic parsing through the Prompt project ("Extract entities (including names and attributes), events (including types and features), and implicit knowledge points (including association rules and background knowledge) from the following text"). For example, given the input traffic police report text "Vehicle A (license plate number A12345, small sedan) collided with vehicle B at 09:20:01 at accident section K12+300, driver Zhang San failed to use turn signals as required," the model outputs:
[0074] Entities: Vehicle A (Attribute: License Plate Number = A12345, Type = Small Car), Vehicle B (Attribute: Type = Small Car), Zhang San (Attribute: Identity = Driver), Accident Section K12+300 (Attribute: Location = Middle Lane of Main Road);
[0075] Event: Collision incident (Type = Traffic accident, Characteristics: Time = 09:20:01, Location = Accident section K12+300, Parties involved = Vehicle A, Vehicle B), Failure to use turn signal (Type = Violation, Characteristics: Subject of behavior = Zhang San, Time = 5 seconds before collision);
[0076] Implicit knowledge points: According to relevant regulations, motor vehicles should turn on their turn signals in advance when changing lanes (related rules); the accident section at K12+300 is the intersection of main and auxiliary roads, which is a high-risk area for accidents (background knowledge).
[0077] After completing text modality extraction, the server performs cross-modal association integration. For the image modality, the server extracts visual features from the vehicle damage assessment photos using the ResNet-50 model: the damaged area of the front bumper of vehicle A is 30% (visual feature vector: [0.82, 0.15, ..., 0.07]), and performs cosine similarity calculation with the entity attribute features of "damaged front bumper of vehicle A" extracted from the text (similarity 0.91 > threshold 0.85) to establish an association mapping. For the audio modality, the server uses Mel-frequency cepstral coefficients (MFCC) to extract the voiceprint features of the party's speech transcript: when driver Zhang San stated "did not notice the car behind," the speech rate was 180 words / minute, and the speech energy fluctuation amplitude was 0.35 (higher than 0.21 for calm statements), which matches the feature description "distracted attention" of the text event "did not observe road conditions" (mutual information value of acoustic and semantic features 0.78 > threshold 0.7). For the video modality, the server extracts the motion sequence features of vehicle B using a 3D convolutional network (C3D): within 2 seconds before the collision, vehicle B's lateral displacement reaches 1.2 meters, with an acceleration of -4.8 m / s². 2 (Braking features) are time-aligned with the feature description "braking reaction time 0.5 seconds" of the text event "emergency braking" (timestamp error < 0.3 seconds) to confirm action matching.
[0078] The server further utilizes implicit knowledge points to supplement cross-modal semantic association: based on the background knowledge that "the accident section is the intersection of the main and auxiliary roads," it associates the action sequence features of vehicle B in the video of "merging from the auxiliary road into the main road" with the violation event of "failing to yield to vehicles on the main road" in the text; through the association rule of "turn signals must be turned on when changing lanes," it fuses the visual feature of "vehicle A not turning on its turn signal" in the image (the average pixel value of the turn signal area is <50, indicating it is in an off state) with the event feature of "failure to use turn signals as required" in the text to generate a complete multimodal semantic association. The knowledge point is: "At 09:19:56, when driver Zhang San of vehicle A changed lanes at the accident site K12+300 (intersection of main and auxiliary roads), he did not turn on his turn signal (image verification) and did not observe oncoming vehicles (audio emotion feature verification), which caused vehicle B to brake suddenly 0.5 seconds before the collision (video action sequence verification), ultimately leading to a collision between the two vehicles (comprehensive event)." This knowledge point is stored in the multimodal knowledge base, and the associated ID field includes the text summary ID, image frame ID, audio segment ID, and video keyframe timestamp.
[0079] In this embodiment of the invention, the process of resolving entity and event referencing based on the fused multimodal knowledge points and constructing a dual graph that combines semantic network and context graph can be implemented through the following example.
[0080] Based on the initial corpus, an initial entity and event extraction prompt template is designed. The initial entity set and initial event set are jointly extracted from the multimodal knowledge points to construct an initial type dictionary containing basic entity categories and basic event types.
[0081] Based on the category coverage of the initial entity set, initial event set, and initial type dictionary, an extended corpus is introduced and the initial entity and event extraction Prompt template is iteratively optimized through an incremental learning strategy to expand the category set of entities or events and generate an updated entity set and event set.
[0082] The updated entity set and event set are subjected to referential resolution. Embedding vectors of entities and events are generated based on the Transformer semantic encoder. Cross-document semantic disambiguation is achieved through cosine similarity calculation and a three-level disambiguation strategy to obtain the disambiguated standardized entity set and event set. The three-level disambiguation strategy includes a progressive disambiguation process of sequential automatic resolution, metadata verification based on document timestamp and source, and expert verification.
[0083] Based on the optimized entity and event extraction Prompt template, and after dereference resolution, a standardized entity set and event set are obtained. Combining the four-dimensional relationship definition of entity-entity, entity-event, event-entity, and event-event, a triplet is designed to generate the Prompt template.
[0084] The three-tuples are used to generate a Prompt template to extract the four-dimensional relationship from the multimodal knowledge points, generate the four-dimensional relationship three-tuples, and construct the knowledge and reasoning graph of the semantic network and the context graph.
[0085] In an embodiment of the present invention, for example, a server is used as the execution subject. In a traffic accident liability determination scenario, the server initiates entity and event processing flow based on the fused multimodal knowledge points. The server first designs an initial entity and event extraction prompt template ("Extract entities (name + attributes) and events (type + features) from traffic accident text. Entity categories include [vehicle, driver, road facilities], and event types include [speeding, lane changing behavior, collision event]"). The input is the multimodal knowledge point text "Vehicle A (license plate number A12345) driver Zhang San failed to observe road conditions and collided with vehicle B at K12+300". The initial entity set E0 = {vehicle A (attribute: license plate number = A12345), Zhang San (attribute: identity = driver), K12+300 (attribute: type = road location)} and the initial event set V0 = {failure to observe road conditions (type: violation), collision event (type: traffic accident)} are extracted using the GPT-4 model. The initial type dictionary T0 = {entity category: 3 types, event type: 2 types} is constructed.
[0086] To address the issue of insufficient coverage of the "traffic facilities" category in the initial type dictionary, the server introduced an expanded corpus (containing 500 accident cases including traffic lights and speed limit signs). An incremental learning strategy was used to optimize the Prompt template: the "traffic facilities" entity category and the "signal violation" event type were added to the original template, and the expanded corpus was re-inputted for iterative extraction. After three rounds of iteration, the entity categories expanded to 5 (adding "traffic signs, traffic signals"), and the event types expanded to 4 (adding "signal violation, failure to maintain a safe distance"), generating an updated entity set E1 (containing the "speed limit sign 120km / h" entity) and an event set V1 (containing the "violation of traffic light instructions" event).
[0087] The server performs referential resolution on E1 and V1: It generates embedding vectors (768 dimensions) for entities "Vehicle A" and "Hit-and-Run Car" using a Sentence-BERT encoder, calculates a cosine similarity of 0.92 (>threshold 0.85), triggering automatic resolution; for the referential relationship between "Zhang San" and "Driver A," it verifies the document metadata and finds that both have the same ID number: "110 Semantic and Contextual Knowledge Collaborative Modeling Declarative Knowledge Construction Device 110XXXX1234," resolving the issue through metadata verification; for the difference in timestamps between "collision event" and "accident occurrence" (09:20:01 vs 09:20:03), it initiates an expert verification process. Traffic engineers, based on video frame analysis, confirm that "09:20:01 is the contact moment, and 09:20:03 is the complete stop moment," retaining the main timestamp of the "collision event" and adding secondary attributes. The final standardized entity set E = {vehicle A, vehicle B, Zhang San, Li Si, speed limit sign 120km / h, K12+300} and event set V = {speeding, failure to observe road conditions, lane changing behavior, collision event} are obtained.
[0088] Based on the optimized entity and event set, the server designs a four-dimensional relation triplet to generate a Prompt template ("Extract (head entity / event, relation, tail entity / event) from the text, relation types include: entity-entity [belonging, location], entity-event [implementation, experience], event-entity [involved, violated], event-event [caused, triggered]"). The input is the integrated knowledge point text "Vehicle B (driver Li Si) was speeding at 130km / h, and a collision occurred due to vehicle A's sudden lane change causing brake failure." The model outputs a four-dimensional relation triplet:
[0089] Entity-Entity: (Vehicle B, Owner: Li Si), (Speed Limit Sign 120km / h, Location: K12+300);
[0090] Entity-Event: (Li Si, committing speeding), (Vehicle A, experiencing lane changing behavior);
[0091] Event-Entity: (Speeding, violation of speed limit sign 120km / h), (Collision incident, involving vehicle A);
[0092] Event-Event: (Speeding, resulting in brake failure), (Lane change, causing a collision).
[0093] The server stores the above triples as graph structure data. Entity nodes contain a unique ID (e.g., VEH001 = vehicle A), type label, and attribute key-value pairs. Event nodes contain an event ID (e.g., EVE001 = speeding), timestamp, and confidence score (0.92). Relationship edges contain relationship type and weight value (calculated based on semantic similarity). Finally, a dual graph is constructed that combines a semantic network (knowledge graph) and a context graph (event graph), and stored in the Neo4j graph database, supporting node attribute query and path relationship reasoning.
[0094] In this embodiment of the invention, the step of performing graph community network clustering on the dual graphs to divide them into topic- or domain-related communities includes:
[0095] Based on the strength of the association between entities and events in the dual graph and the relevance of the topic, the Leiden algorithm is used to perform graph community network clustering on the knowledge and reasoning graph. By optimizing the modularity function to maximize the association between nodes within the community, multiple densely connected subgraphs with consistent internal entity and event topics are obtained.
[0096] For each densely connected subgraph, core entities and events are extracted. Based on the semantic tags of the extracted core entities and events, the topic or domain corresponding to the subgraph is determined, and each subgraph is marked as a community related to the topic or domain.
[0097] In this embodiment of the invention, for example, a server acts as the execution entity. In a traffic accident liability determination scenario, the server initiates a community clustering process based on a constructed semantic network and a contextual graph. The server first quantifies the strength of the relationships between nodes in the two graphs: the weight of the connection edge between entity node "Vehicle B" and event node "speeding" is set to 0.92 (based on the comprehensive confidence of fused knowledge points); the weight of the causal relationship edge between event node "speeding" and "brake failure" is set to 0.87 (based on the causal reasoning score of the contextual graph); and the weight of the "violation" relationship edge between entity node "speed limit sign 120km / h" and event node "speeding" is set to 0.95 (based on legal provision matching degree). Simultaneously, topic relevance is calculated: through the BERT model to generate node semantic vectors, the mean cosine similarity of the vectors for "Vehicle B," "speeding," and "speed limit sign" reaches 0.83, classifying them as highly topic-relevant nodes.
[0098] The server invokes the Leiden community discovery algorithm to cluster the bi-graph (containing 28 entity nodes, 15 event nodes, and 89 relation edges). The algorithm represents the bi-graph as a weighted undirected graph G = (V, E), where the node set V contains entities and events, and the weight of the edge set E represents the strength of the association. The server initializes the modularity function Q (initial value 0.32), sets the resolution parameter γ = 1.0, and sets the number of iterations to 50. The modularity is optimized through local moves and community merging: in the first iteration, the "speeding" event node is merged from its independent community into the "vehicle B" entity community, increasing the modularity to 0.41; in the third iteration, the "brake failure" event node is merged into the "speeding-vehicle B" community, achieving a modularity of 0.56; when the final iteration converges, the modularity stabilizes at 0.59 (> the threshold of 0.5), forming three densely connected subgraphs.
[0099] For the first densely connected subgraph (containing 12 nodes and 35 edges), the server extracts core nodes: node importance scores are calculated using the PageRank algorithm; the "speeding" event node scores 0.18 (highest), the "speed limit sign 120km / h" entity node scores 0.15, and the "violation of regulations" relation edge scores 0.17. Based on the semantic labels of the core nodes "speeding," "speed limit sign," and "regulatory clauses," the server uses the TextRank algorithm to generate subgraph topic vectors, which are matched with a pre-defined domain dictionary (containing keywords such as "liability determination," "violation," and "legal basis") to determine the "accident liability elements" community corresponding to this subgraph.
[0100] For the second densely connected subgraph (containing 14 nodes and 31 edges), the core nodes are "Vehicle A" (PageRank score 0.16), "Lane Change Behavior" (score 0.15), and "Unobserved Road Conditions" (score 0.14). The semantic tags focus on vehicle operation and dynamic processes, matching the topic "Vehicle Behavior Analysis," and are labeled as the "Vehicle Dynamic Behavior" community. For the third densely connected subgraph (containing 17 nodes and 23 edges), the core nodes are "Collision Occurrence" (score 0.19), "Vehicle Damage" (score 0.16), and "Insurance Claim" (score 0.13). The topic vector matches "Accident Consequences and Handling," and is labeled as the corresponding community. The server stores the results for each community in a graph database. Each community record includes a community ID, topic tags, a list of core nodes, and the average weight of internal edges, supporting subsequent association mapping and knowledge retrieval.
[0101] In this embodiment of the invention, the step of associating and mapping the community, the triples, entities, and events with the chapter summary and the multimodal knowledge points respectively to form associated knowledge can be implemented through the following example.
[0102] The semantic similarity between the community semantic tags after the dual-graph clustering and the keywords in the chapter summary is calculated using the attention mechanism, and the association between the community and the chapter summary between the document parsing layer and the dual-graph layer is established.
[0103] The vector distance between the entity and event nodes and the multimodal knowledge points is calculated using knowledge embedding technology, and the association links between entities and events and multimodal knowledge points are established in the knowledge point fusion layer and the dual graph layer.
[0104] The associations between the community and chapter summaries, the associations between entities and events and multimodal knowledge points, and the four-dimensional relationship triples are integrated across layers to form the associated knowledge that includes document parsing information, multimodal knowledge points, and dual-graph relationships.
[0105] In this embodiment of the invention, for example, the server acts as the execution entity. In a traffic accident liability determination scenario, the server initiates an association mapping process based on graph community clustering results and structured content. The server first uses an attention mechanism to calculate the semantic similarity between community semantic tags and chapter summary keywords: For the "Accident Liability Elements" community (semantic tag vector: [0.82, 0.15, 0.91, 0.07], dimension 128), it extracts keyword vectors from the chapter summary library for "Laws and Regulations Chapter Summary" (including "speed limit," "liability determination," "legal clauses," etc., vector: [0.79, 0.18, 0.88, 0.05]), and calculates a weighted similarity score of 0.89 (> threshold 0.8) using a multi-head attention mechanism (8 heads). Simultaneously, it calculates the similarity score between this community and the "Surveillance Video Chapter Summary" as 0.42 (< threshold 0.8). Finally, it establishes an association mapping between the "Accident Liability Elements" community and the "Laws and Regulations Chapter Summary," storing the association table (community). i d = COM001, chapter i d=CH003, similarity_score=0.89).
[0106] To address the association between entities and events and multimodal knowledge points, the server uses the TransE knowledge embedding model to generate entity / event node vectors and knowledge point vectors: the embedding vector of the entity node "Speed Limit Sign 120km / h" (entity ID = ENT005) is [0.32, -0.15, 0.67, ..., 0.21], and the knowledge point vector of "Image Modality - Speed Limit Sign Photo" in the multimodal knowledge point database is [0.30, -0.17, 0.65, ..., 0.23]. The L2 vector distance is calculated to be 0.28 (< threshold 0.3), establishing an association link; the embedding vector of the event node "Speeding" (event ID = EVE001) has a distance of 0.25 with the knowledge point vector of "Video Modality - Vehicle B Speed Analysis" and a distance of 0.22 with the knowledge point vector of "Text Modality - Legal Provision Violation Description", both satisfying the threshold condition, forming a multi-source knowledge point association. The server generates a list of associated knowledge points for each entity / event node, including the knowledge point ID, modality type (image / video / text), and vector distance value. For example, the associated list for ENT005 is (KN008, image, 0.28; KN012, text, 0.31).
[0107] The server integrates community-chapter associations, entity / event-knowledge point associations, and four-dimensional relationship triples across layers: the "Accident Liability Elements" community (COM001) is linked to the "Laws and Regulations Chapter Summary" text content via chapter ID=CH003; the entity "Speed Limit Sign 120km / h" (ENT005) within the community is linked to the speed limit sign photo in the image modality and the "120km / h" value extracted by OCR via knowledge point ID=KN008; the event "Speeding" (EVE001) is associated with the entity via the triple (EVE001, violation, ENT005), and the video speed analysis data (130km / h) is integrated with the legal provision description via knowledge point links. The resulting associated knowledge has a three-layer structure: a document parsing layer (chapter summary text), a knowledge point fusion layer (original data of multimodal knowledge points), and a dual graph layer (entity / event and relationship triples), with each layer connected by a unique identifier (such as community). i d, entity i d. knowledge i d) Forming a closed-loop reference. The server stores related knowledge in a hybrid database architecture, where a relational database stores relational tables, a graph database stores a dual-graph structure, and a vector database stores knowledge point embedding vectors, supporting multi-level knowledge tracing and retrieval from "community → chapter → knowledge point → entity / event → triple".
[0108] In this embodiment of the invention, the process of evaluating the compression rate and accuracy of the declarative knowledge in the vector knowledge index through the declarative knowledge evaluation system, and optimizing the multimodal knowledge point fusion, dual graph construction, graph community network clustering and association mapping based on the evaluation results, can be implemented through the following example.
[0109] The declarative knowledge in the vector knowledge index is processed and compressed to evaluate the redundancy of the knowledge vector library by calculating information entropy, and the information fidelity during the knowledge compression process is measured.
[0110] The accuracy of processing the declarative knowledge in the vector knowledge index is evaluated, the semantic rationality score of the triple is calculated based on the scoring function of the knowledge graph embedding model, and the semantic consistency of the entity-event relationship is verified by combining the manually labeled test set.
[0111] Based on the combined results of the processing compression rate assessment and the processing accuracy assessment, the key links leading to quality problems in the knowledge construction process are identified. The assessment results are then fed back to the steps of multimodal knowledge point fusion, dual graph construction, graph community network clustering, and association mapping to iteratively optimize the knowledge construction process.
[0112] In this embodiment of the invention, for example, a server acts as the execution entity. In a traffic accident liability determination scenario, the server initiates a declarative knowledge assessment system to evaluate the quality of declarative knowledge in the vector knowledge index. During the compression rate assessment stage, the server quantifies the redundancy of the knowledge vector library using information entropy calculation: first, it performs information entropy analysis on the original multimodal data (including surveillance videos, handwritten reports, voice transcripts, etc.) to measure the degree of disorder in the data; then, it calculates the information entropy of the associated knowledge in the vector knowledge index (including communities, triples, and multimodal knowledge points). By comparing the two, it evaluates the information integration effect during the knowledge compression process and determines whether core knowledge is retained while removing redundant information.
[0113] In the accuracy evaluation stage, the server uses a scoring function based on the knowledge graph embedding model to judge the semantic rationality of four-dimensional relation triples: the embedded vectors of entities and events are input into the model, and the semantic matching scores of the head node, relation, and tail node in the triple are calculated. Low-quality triples with scores below a set threshold are filtered out. Simultaneously, the server retrieves a manually annotated domain knowledge test set and compares the entity and event relations in the vector knowledge index with the standard relations in the test set to verify semantic consistency and identify issues such as incorrect entity referencing and inverted event relations.
[0114] Based on the comprehensive evaluation results of compression rate and accuracy, the server identifies key optimization points in the knowledge construction process: if the similarity between a community and a chapter summary is below a threshold, it indicates a bias in the topic segmentation during the graph community network clustering stage; if the vector distances between entities and multimodal knowledge points are generally large, it suggests insufficient cross-modal feature association during the multimodal knowledge point fusion stage; if the semantic rationality scores of a large number of "event-event" relationship triples are low, it indicates a defect in event relationship extraction during the dual-graph construction stage. The server feeds these evaluation conclusions back to the corresponding stages: to address the insufficient cross-modal fusion, it adjusts the association weights of text and non-text modal features; to address the community clustering bias, it optimizes the modularity calculation parameters of the Leiden algorithm; and to address the relationship extraction defects, iteratively updates the relationship type definitions in the triple generation Prompt template. By continuously feeding the evaluation results back to the entire knowledge construction process, the server gradually optimizes the relevance of multimodal knowledge point fusion, the accuracy of dual-graph relationships, the topic consistency of community clustering, and the matching accuracy of association mapping, achieving a closed-loop improvement in the quality of declarative knowledge.
[0115] To more clearly describe the solutions provided in the embodiments of the present invention, a more complete implementation method is provided below. Please refer to the following reference. Figure 2 and 3 , Figure 2 This is a schematic diagram of the construction framework of the declarative knowledge construction system provided in the embodiments of the present invention. Figure 3 This is a schematic diagram of the phased collaborative construction process of dual-maps provided in an embodiment of the present invention.
[0116] Step 1: Multimodal document parsing and chapter summary extraction:
[0117] Multimodal document parsing:
[0118] Data Collection and Preprocessing: Collect raw documents in various formats, such as PDF, Word, HTML, and tables (text documents), as well as non-text documents such as images, audio, and video. Perform preliminary cleaning on these documents to remove useless information, such as redundant whitespace characters and noisy data.
[0119] Modality recognition and separation: Identifying different modalities in a document, such as text, tables, images, audio, and video. For documents with mixed modalities, appropriate techniques are used to separate the content of different modalities for separate processing. For example, for a document containing text and charts, image recognition technology is used to distinguish the chart portion from the text portion.
[0120] Text parsing: For text content, natural language processing techniques, such as word segmentation, part-of-speech tagging, and named entity recognition, are used to convert the document text into structured text data, making it easier for computers to understand and process. Simultaneously, the text is encoded to unify the character encoding format, ensuring correct reading and parsing of the text.
[0121] Non-text parsing: For non-text content such as tables, images, audio, and video, appropriate parsing techniques are applied. For images, image recognition algorithms are used to extract key information such as objects, scenes, and text. For audio and video, speech recognition and video content analysis are performed to convert them into corresponding text descriptions or keyframe information for subsequent fusion processing with other modalities.
[0122] Chapter summary extraction:
[0123] Chapter division: Based on the document's structural information, such as headings, paragraph formatting, and page numbers, the document is divided into chapters. For documents without explicit structural information, text segmentation techniques from natural language processing are used to automatically divide the document into different chapters based on semantic coherence and thematic consistency.
[0124] Feature extraction: Key features, including keywords, key sentences, and topic terms, are extracted from the text content of each chapter. Important information in the text is identified using text ranking algorithms and deep learning models as candidate features for chapter summaries.
[0125] Summary Generation: Based on the extracted features, large-scale modeling techniques, such as pre-trained language models based on the Transformer architecture, are used to understand and summarize the chapter content, generating concise and accurate chapter summaries. Large-scale models can generate highly generalized summaries that conform to human language habits based on contextual semantic information, ensuring that the chapter summaries effectively convey the core points of the chapter.
[0126] Summary Optimization and Integration: The generated chapter summaries are optimized by adjusting sentence structure and vocabulary selection to make them smoother and more coherent. Simultaneously, the logical relationships and consistency between the chapter summaries are checked to ensure that the entire document's chapter summaries form a cohesive whole, providing a clear and accurate structured content foundation for subsequent knowledge extraction and integration.
[0127] Step Two: Extraction and Integration of Knowledge Points from Multi-Modal Documents
[0128] Knowledge extraction tools like Deepseek or Qwen3 perform deep semantic understanding and analysis of document content. By inputting document text into a large model, the semantic information output by the model is obtained, allowing for the mining and extraction of potential knowledge points. Large models possess powerful language understanding and generation capabilities, enabling them to capture complex semantic relationships and implicit information in text, thus more accurately identifying knowledge points in documents, including some deep-level knowledge that is difficult to extract using traditional methods.
[0129] Knowledge point integration:
[0130] Cross-modal fusion: This involves fusing knowledge points extracted from different modalities (such as text, images, audio, and video). For example, it involves associating and integrating the features of an object described in text with the visual features of that object identified in an image to form a complete knowledge point about that object. For content that includes both audio explanations and text descriptions, it involves fusing key information from the audio with the textual knowledge points to enrich the connotation and expression of the knowledge points.
[0131] Step 3: Collaborative Modeling of Two Graphs
[0132] This invention employs a collaborative modeling framework based on knowledge graphs and event graphs, using a "hint engineering + dynamic iteration" approach to construct a paradigm and achieve multi-dimensional modeling of structured knowledge. Collaborative modeling includes the following core stages:
[0133] Bimodal knowledge modeling:
[0134] Bimodal knowledge is a knowledge network constructed by jointly modeling static entity knowledge (KG) and dynamic event logic (EG), which includes the relationships between entities, between entities and events, between events and entities, and between events and entities: entity-entity relationship (KG-KG); entity-event relationship (KG-EG); event-entity relationship (EG-KG); event-event relationship (EG-EG).
[0135] Phased build process:
[0136] Phase 1: Basic Model Guidance
[0137] First, from the initial declarative text corpus (D1), we construct entity and event extraction cue words v1.0. Using a large model + cue word approach, we extract entities and events from the text, forming entity category sets and event category sets. The specific operations are as follows:
[0138] Design the Promptv1.0 template to integrate instructions for the joint extraction of entities and events;
[0139] Joint extraction based on semantic parsing of large models:
[0140] Generate entity set E = {e i}|e i ∈ Entity Type Space T e ;
[0141] Generate event set E = {e i}|e i ∈ Event type space T e ;
[0142] Construct an initial type dictionary: T e ^(0), T v ^(0).
[0143] Phase Two: Iterative Evolution of Patterns
[0144] Building upon prompt word v1.0, prompt word v2.0 is constructed by adding entity category sets and event category sets to extract entities and events. This is implemented on the extended corpus (D2). The same large model + prompt word approach is used: entities and events are first selected from the sets; if not found, they are added, and the entity category set and event category set are updated. This process iterates through prompt word v2.x until all D2 corpora are extracted. The specific operations are as follows:
[0145] Upgrade Prompt to v2.0, integrating entity / event category sets and joint entity / event extraction commands:
[0146] Entity extraction: P(e|T) e ^(k));
[0147] Event extraction: P(v|T) v ^(k));
[0148] Implement an incremental learning strategy:
[0149] ife∈T e ^(k):
[0150] T e ^(k+1)=T e ^(k)∪{e};
[0151] if v∈T v ^(k):
[0152] T v ^(k+1)=T v ^(k)∪{v};
[0153] Dynamically update the Prompt to version v2.x.
[0154] Phase 3: Cross-document semantic disambiguation:
[0155] Because the same entity or event may have slightly different names in different documents, such as "speeding" and "driving at excessive speeds," this difference increases the complexity of data processing and necessitates name resolution. This is achieved by generating candidate matches using text embedding. Embedding reflects semantic similarity. Once candidate objects are identified, further evaluation is performed, comparing metadata or, in critical cases, combining manual review, to determine whether these names refer to the same entity or event, thereby optimizing the entity category set and event category set. The specific steps are as follows:
[0156] Constructing a reference resolution model: Generate embedding vectors φ(·) based on a semantic encoder using Transformer; Similarity calculation: sim(x, y) = cos(φ(x), φ(y));
[0157] Implement a three-level disambiguation strategy: automatic resolution: candidate set C = top_k(sim(x, y) > θ); metadata verification: aligning time / space / attribute features.
[0158] Manual verification: Expert review of key nodes: Output optimization type dictionary: T e T v .
[0159] Phase Four: Structured Knowledge Generation
[0160] Based on the optimized category set, the full corpus (D) is processed, selecting entities and events from the set during extraction and generating multi-relation triples. The Promptv3.0 template is designed, constraining the output format: Entity Relationship Group: Relationships between entities: (head:T) e ,relation,tail:T e );
[0161] Relationship Group: Relationship between Entities and Events: (head:T) e ,triggers,event:T v The relationship between events and entities: (event:T) v ,affects,tail:T e Relationship between events: (event:T) v ,leads_to,event:T v ).
[0162] Step 4: Community Detection Algorithm
[0163] Community detection algorithms are used to perform clustering analysis on the constructed dual graph. This algorithm can intelligently identify closely related sets of nodes in the graph, i.e., communities. For example, using the graph-based clustering algorithm Leiden, community partitioning is achieved by continuously optimizing the modularity function. This algorithm continuously adjusts the community affiliation of nodes according to specific rules, causing the modularity of the graph to gradually increase until multiple relatively independent but internally connected communities are formed. Each community corresponds to a unique theme or specific domain in the dual graph.
[0164] Step 5: Knowledge Association Mapping:
[0165] The clustered graph communities, triples, individual entities, or events are mapped to chapter summaries and document knowledge points. For example, an attention mechanism is used to match semantic tags of communities with keywords in chapter summaries, and knowledge embedding techniques are used to associate entity nodes with corresponding knowledge points.
[0166] Step Six: Knowledge Vectorization and Intelligent Indexing
[0167] The rules of events and entity relationships in the dual graph are transformed into high-dimensional knowledge vectors, and a multi-granularity vector library is constructed using hierarchical indexing technology.
[0168] Step Seven: Declarative Knowledge Assessment System
[0169] The declarative knowledge assessment system mainly evaluates the processing compression rate and processing accuracy of declarative knowledge.
[0170] The assessment of declarative knowledge processing compression rate mainly uses the information entropy calculation method to quantify the redundancy of the knowledge base and measure the information fidelity of the knowledge construction system that co-models semantic and contextual knowledge.
[0171] The accuracy evaluation of declarative knowledge processing is based on a scoring function of a knowledge graph embedding model, combined with a manually annotated test set to verify the semantic consistency of knowledge triples.
[0172] The declarative knowledge assessment system further feeds back the assessment results to optimize the entire knowledge building process.
[0173] Implementation Case:
[0174] The following is an example of determining liability in a traffic accident. The process is as follows:
[0175] Step 1: Multimodal document parsing and chapter summary extraction:
[0176] The system receives accident scene surveillance video (MP4 format), traffic police handwritten reports (scanned PDF), vehicle speed data (JPG), and voice transcripts of the parties involved (WAV), as well as legal regulations (Word). The video stream is processed by a 3D convolutional network to extract keyframes and identify the relative positions of the two vehicles at the moment of collision. The handwritten report is converted using OCR to obtain basic accident information, basic information of the parties involved, and the course of the event. The photos are used for object detection to identify that vehicle B was speeding. The voice transcript is converted to text through speech recognition to extract key time points of the narrator. The system parses the text content of legal regulations using Word. The system receives the text content of legal regulations, such as the most common regulations and local traffic laws and regulations, and extracts summaries from each chapter or article, such as regulations for motor vehicle traffic, regulations for non-motor vehicle traffic, and regulations for pedestrians and passengers.
[0177] Step Two: Extraction and Integration of Knowledge Points from Multi-Modal Documents
[0178] For example, based on the relevant regulations, by writing prompt words, XX items are taken as input, and the Deepseek large model is used to extract knowledge points. The extracted knowledge point is speeding. The camera detected that the vehicle speed was 130km / h, while the maximum speed limit of the road section is 120km / h. These three information are then fused together.
[0179] Step 3: Collaborative Modeling of Two Graphs
[0180] Entities (such as "Vehicle A", "Vehicle B", "Speed Limit Sign") and events (such as "Collision Occurred", "Speeding") are extracted from keyframes of surveillance video. These are then combined with text entities (such as "Driver A", "Road Sign") and events (such as "Emergency Braking Failure") from the traffic police's handwritten report. After joint extraction using the Promptv1.0 template, an initial set of entity categories (such as T) is formed. e ^(0) = {vehicles, drivers, road facilities}) and event category set (such as T) v ^(0) = {collision, speeding, brake failure}
[0181] In expanded corpora (such as more accident case texts), Promptv2.0 prioritizes matching existing entity / event types. For example, if a new document mentions "unnoticed observation," and T... v If an event is not found in ^(k), it is added to the event set, and the type dictionary is dynamically updated. The incremental learning strategy ensures that entity and event types adaptively evolve as data expands.
[0182] For instances where "speeding" is expressed as "overspeeding driving" in text and "excessive vehicle speed" in voice recordings, an embedding vector φ is generated using a Transformer encoder. The similarity sim(x, y) = cos(φ(x), φ(y)) is calculated to automatically identify candidate matches. For example, if sim("speeding", "excessive vehicle speed") > θ (θ is set to 0.85), metadata verification (such as timestamp consistency and speed value matching) is performed to ultimately confirm that both refer to the same event.
[0183] Based on the optimized type dictionary, four-dimensional relation triples are generated from the full corpus:
[0184] Entity-Entity Relationship (KG-KG): ('[Event Background](Entity Type)', 'Time', '[June 24, 2020, 16:33](Entity Type)');
[0185] Entity-Event Relationship (KG-EG): ('[Zhang San](Entity Type: Driver)', 'Implementation', '[Speeding](Event Type)');
[0186] Event-Entity Relationship (EG-KG): ('[Liability Determination](Event Type)', 'Parties Involved', '[Zhang San](Entity Type)');
[0187] Event-Event Relationship (EG-EG): ('[Speeding](Event Type)', '[Cause](Cause)', '[Failure to Yield in Time](Event Type)').
[0188] Step 4: Community Detection Algorithm
[0189] In traffic accident cases, the Leiden algorithm is applied to perform community partitioning based on a dual graph (containing entity relation groups and event relation groups):
[0190] Community clustering: Multiple highly cohesive subgraphs were identified, such as the "Accident Liability Elements" community (containing event nodes such as "speeding" and "brake failure") and the "Vehicle Dynamics Analysis" community (containing entity nodes such as "Vehicle A" and "speed limit sign").
[0191] Feature analysis: Statistics on core entities and events in each community. For example, in the "Accident Liability Elements" community, the high-frequency event is "speeding" (occurring in 40% of the events). In the "Vehicle Dynamics Analysis" community, the core entity is "Vehicle B" (with an average connection weight of 0.82 with other nodes).
[0192] Cross-community linkage: By sharing the entity "Vehicle A", the "Vehicle Dynamics Analysis" community and the "Accident Consequence Assessment" community are linked to reveal the causal relationship between "vehicle damage level" and "collision occurrence" event.
[0193] Step 5: Knowledge Association Mapping:
[0194] In the field of traffic accidents, mapping is achieved through the following methods:
[0195] Community and Abstract Association: Using an attention mechanism, we matched the semantic tags (such as "speeding") of the "Accident Liability Elements" community with the keyword overlap (similarity threshold > 0.75) of the chapter abstract "Motor Vehicle Traffic Regulations".
[0196] Community and knowledge point association: The entity "speed limit sign" in the "vehicle dynamics analysis" community is embedded with the extracted knowledge point "the maximum speed limit for this road section is 120km / h" using the TransE model, and the vector distance is calculated (a link is established when the L2 norm is <0.3).
[0197] Step Six: Knowledge Vectorization and Intelligent Indexing
[0198] Implementation of dual-map system for traffic accidents:
[0199] Hierarchical vectorization: The GraphSAGE algorithm is used to generate 128-dimensional vectors of entities / events, among which the cross-modal cosine similarity between the "speeding" event vector and the "excessive speed" speech description reaches 0.89.
[0200] Multi-granularity index: Construct a three-level vector library (global-community-node) to support fine-grained question answering such as "relative position of two vehicles" and complex reasoning tasks such as "basis for determining accident liability".
[0201] Step Seven: Declarative Knowledge Assessment System
[0202] In the verification of the traffic accident knowledge base:
[0203] Compression rate assessment: Through information entropy calculation, the compression rate after knowledge fusion increased from 35% of the original data to 80%, indicating a significant improvement in information fidelity.
[0204] Accuracy assessment: Based on the TransE scoring function, the accuracy of manual annotation on 1,000 triplet test sets reached 91.5%, among which the confidence score of the causal chain "speeding → failure to avoid in time → collision occurred" was higher than the threshold (0.78 / 1.0).
[0205] Please refer to the following: Figure 4 , Figure 4 A declarative knowledge construction device 110 for co-modeling semantic and contextual knowledge, provided in an embodiment of the present invention, includes:
[0206] The acquisition module 1101 is used to parse multimodal documents and extract chapter summaries to obtain structured content and chapter summaries; based on the structured content, entities, events and multimodal knowledge points are extracted and cross-modal fusion is performed;
[0207] The construction module 1102 is used to resolve the referentiality of entities and events based on the fused multimodal knowledge points, and construct a dual graph that combines a semantic network and a contextual graph. The dual graph includes a knowledge and event graph and four-dimensional relation triples. Graph community network clustering is performed on the dual graph to divide it into communities related to topics or domains. The communities, triples, entities, and events are associated with the chapter summary and the multimodal knowledge points, respectively, to form associated knowledge. The associated knowledge is vectorized and a vector knowledge index is established and stored in a knowledge vector library. The compression rate and accuracy of the declarative knowledge in the vector knowledge index are evaluated by a declarative knowledge evaluation system, and the evaluation results are used to optimize the process of multimodal knowledge point fusion, dual graph construction, graph community network clustering, and association mapping.
[0208] It should be noted that the implementation principle of the aforementioned declarative knowledge construction device 110 for co-modeling semantic and contextual knowledge can refer to the implementation principle of the aforementioned declarative knowledge construction method for co-modeling semantic and contextual knowledge, and will not be repeated here. It should be understood that the division of the various modules in the above device is merely a logical functional division; in actual implementation, they can be fully or partially integrated into a single physical entity, or physically separated. Furthermore, these modules can all be implemented in software through processing element calls; they can all be implemented in hardware; or some modules can be implemented through processing element calls to software, and some modules can be implemented in hardware. For example, the declarative knowledge construction device 110 for co-modeling semantic and contextual knowledge can be a separately established processing element, or it can be integrated into a chip within the aforementioned device. Alternatively, it can be stored as program code in the memory of the aforementioned device, and its functions can be called and executed by a processing element of the aforementioned device. The implementation of other modules is similar. Furthermore, these modules can be fully or partially integrated together, or implemented independently. The processing element described here can be an integrated circuit with signal processing capabilities. In the implementation process, each step or module of the above method can be completed by the integrated logic circuit in the hardware of the processor element or by instructions in the form of software.
[0209] For example, these modules can be one or more integrated circuits configured to implement the above methods, such as one or more application-specific integrated circuits (ASICs), one or more digital signal processors (DSPs), or one or more field-programmable gate arrays (FPGAs). As another example, when a module is implemented using processing element scheduler code, the processing element can be a general-purpose processor, such as a central processing unit (CPU) or other processor capable of calling program code. Furthermore, these modules can be integrated together to implement a system-on-a-chip (SOC).
[0210] This invention provides a computer device 100, which includes a processor and a non-volatile memory storing computer instructions. When the computer instructions are executed by the processor, the computer device 100 executes the aforementioned declarative knowledge construction device 110 for co-modeling semantic and contextual knowledge. Figure 5 As shown, Figure 5 This is a structural block diagram of a computer device 100 provided in an embodiment of the present invention. The computer device 100 includes a declarative knowledge construction device 110 for co-modeling semantic and contextual knowledge, a memory 111, a processor 112, and a communication unit 113.
[0211] To enable data transmission or interaction, the memory 111, processor 112, and communication unit 113 are electrically connected to each other directly or indirectly. For example, these components can be electrically connected through one or more communication buses or signal lines. The declarative knowledge construction device 110 for semantic and contextual knowledge co-modeling includes at least one software functional module that can be stored in the memory 111 or embedded in the operating system (OS) of the computer device 100 in the form of software or firmware. The processor 112 is used to execute the declarative knowledge construction device 110 for semantic and contextual knowledge co-modeling stored in the memory 111, such as the software functional modules and computer programs included in the declarative knowledge construction device 110 for semantic and contextual knowledge co-modeling.
[0212] This invention provides a readable storage medium, which includes a computer program. When the computer program runs, it controls the computer device where the readable storage medium is located to execute the aforementioned declarative knowledge construction device 110 for co-modeling semantic and contextual knowledge.
[0213] For illustrative purposes, the foregoing description has been made with reference to specific embodiments. However, the foregoing illustrative discussions are not intended to be exhaustive or to limit the present disclosure to the precise forms disclosed. Numerous modifications and variations are possible in accordance with the foregoing teachings. These embodiments were chosen and described in order to best illustrate the principles of the present disclosure and its practical application, thereby enabling those skilled in the art to best utilize the disclosure and to employ various embodiments with different modifications to suit a particular intended application.
Claims
1. A declarative knowledge construction method that uses co-modeling of semantic and contextual knowledge, characterized in that, include: The document is parsed and chapter summaries are extracted to obtain structured content and chapter summaries. Based on the structured content, entities, events, and multimodal knowledge points are extracted and cross-modal fusion is performed; Based on the fused multimodal knowledge points, entity and event referencing is resolved, and a dual graph is constructed that combines semantic network and context graph. The dual graph includes knowledge and event graph and four-dimensional relation triplet. The dual graphs are subjected to graph community network clustering to divide them into communities related to specific topics or domains. The community, the triples, entities, and events are respectively associated and mapped with the chapter summary and the multimodal knowledge points to form associated knowledge; The associated knowledge is vectorized and a vector knowledge index is established and stored in a knowledge vector library; The declarative knowledge in the vector knowledge index is evaluated for compression rate and accuracy using a declarative knowledge evaluation system. Based on the evaluation results, the process of multimodal knowledge point fusion, dual-graph construction, graph community network clustering, and association mapping is optimized.
2. The method according to claim 1, characterized in that, The process of parsing and extracting chapter summaries from multimodal documents to obtain structured content and chapter summaries includes: Data collection and preprocessing of raw multimodal documents, cleaning up useless information and standardizing character encoding formats; Modality recognition and separation are performed on the preprocessed document to separate text, table, image, audio and video modal content, and non-text modalities are converted into text descriptions through image recognition, speech recognition or video analysis technology; The document is divided into chapters based on its structure or semantic coherence. Keywords and key sentences are extracted using a pre-trained language model to generate chapter summaries, and the text descriptions and chapter summaries are integrated into the structured content.
3. The method according to claim 2, characterized in that, The process of extracting entities, events, and multimodal knowledge points based on the structured content and performing cross-modal fusion includes: Based on a pre-trained large language model, deep semantic understanding and analysis are performed on the text modal content in the structured content, extracting entities, events and implicit knowledge points from the text modal. The entities include entity names and attribute features, the events include event types and feature descriptions, and the implicit knowledge points include association rules or background knowledge that are not explicitly stated in the text but are obtained through contextual semantic reasoning. The entity attribute features, event feature descriptions, and implicit knowledge points extracted from the text modality are respectively integrated with the visual features of the image modality, the event voiceprint features of the audio modality, and the action sequence features of the video modality in the structured content. By associating the object features described in the text with the visual features of image recognition, matching the feature descriptions of events in the text with the event features in the audio or video, and using the implicit knowledge points to supplement the semantic associations between cross-modal features, the complete multimodal knowledge points are formed.
4. The method according to claim 3, characterized in that, The process of resolving entity and event referencing based on the fused multimodal knowledge points, and constructing a dual-graph system that integrates semantic network and context graph, includes: Based on the initial corpus, an initial entity and event extraction prompt template is designed. The initial entity set and initial event set are jointly extracted from the multimodal knowledge points to construct an initial type dictionary containing basic entity categories and basic event types. Based on the category coverage of the initial entity set, initial event set, and initial type dictionary, an extended corpus is introduced and the initial entity and event extraction Prompt template is iteratively optimized through an incremental learning strategy to expand the category set of entities or events and generate an updated entity set and event set. The updated entity set and event set are subjected to referential resolution. Embedding vectors of entities and events are generated based on the Transformer semantic encoder. Cross-document semantic disambiguation is achieved through cosine similarity calculation and a three-level disambiguation strategy to obtain the disambiguated standardized entity set and event set. The three-level disambiguation strategy includes a progressive disambiguation process of sequential automatic resolution, metadata verification based on document timestamp and source, and expert verification. Based on the optimized entity and event extraction Prompt template, and after dereference resolution, a standardized entity set and event set are obtained. Combining the four-dimensional relationship definition of entity-entity, entity-event, event-entity, and event-event, a triplet is designed to generate the Prompt template. The three-tuples are used to generate a Prompt template to extract the four-dimensional relationship from the multimodal knowledge points, generate the four-dimensional relationship three-tuples, and construct the knowledge and reasoning graph of the semantic network and the context graph.
5. The method according to claim 4, characterized in that, The step of performing graph community network clustering on the dual graphs to divide them into topic- or domain-related communities includes: Based on the strength of the association between entities and events in the dual graph and the relevance of the topic, the Leiden algorithm is used to perform graph community network clustering on the knowledge and reasoning graph. By optimizing the modularity function to maximize the association between nodes within the community, multiple densely connected subgraphs with consistent internal entity and event topics are obtained. For each densely connected subgraph, core entities and events are extracted. Based on the semantic tags of the extracted core entities and events, the topic or domain corresponding to the subgraph is determined, and each subgraph is marked as a community related to the topic or domain.
6. The method according to claim 3, characterized in that, The process of associating and mapping the community, the triples, entities, and events with the chapter summary and the multimodal knowledge points to form associated knowledge includes: The semantic similarity between the community semantic tags after the dual-graph clustering and the keywords in the chapter summary is calculated using the attention mechanism, and the association between the community and the chapter summary between the document parsing layer and the dual-graph layer is established. The vector distance between the entity and event nodes and the multimodal knowledge points is calculated using knowledge embedding technology, and the association links between entities and events and multimodal knowledge points are established in the knowledge point fusion layer and the dual graph layer. The associations between the community and chapter summaries, the associations between entities and events and multimodal knowledge points, and the four-dimensional relationship triples are integrated across layers to form the associated knowledge that includes document parsing information, multimodal knowledge points, and dual-graph relationships.
7. The method according to claim 1, characterized in that, The process of evaluating the compression rate and accuracy of the declarative knowledge in the vector knowledge index using a declarative knowledge evaluation system, and optimizing the multimodal knowledge point fusion, dual-graph construction, graph community network clustering, and association mapping based on the evaluation results, includes: The declarative knowledge in the vector knowledge index is processed and compressed to evaluate the redundancy of the knowledge vector library by calculating information entropy, and the information fidelity during the knowledge compression process is measured. The processing accuracy of the declarative knowledge in the vector knowledge index is evaluated, the semantic rationality score of the triple is calculated based on the scoring function of the knowledge graph embedding model, and the semantic consistency of the entity-event relationship is verified by combining the manually labeled test set. Based on the combined evaluation results of processing compression rate and processing accuracy, the key links leading to quality problems in the knowledge construction process are identified. The evaluation results are then fed back to the steps of multimodal knowledge point fusion, dual graph construction, graph community network clustering and association mapping to iteratively optimize the knowledge construction process.
8. A declarative knowledge construction device for co-modeling semantic and contextual knowledge, characterized in that, include: The acquisition module is used to parse multimodal documents and extract chapter summaries to obtain structured content and chapter summaries; Based on the structured content, entities, events, and multimodal knowledge points are extracted and cross-modal fusion is performed; The module is used to resolve the referentiality of entities and events based on the fused multimodal knowledge points, and to construct a dual graph that combines a semantic network and a context graph. The dual graph includes a knowledge graph, a context graph, and a four-dimensional relation triplet. The dual graphs are clustered using graph community networks to identify communities related to specific topics or domains. These communities, triples, entities, and events are then mapped to the chapter summaries and multimodal knowledge points to form associated knowledge. The associated knowledge is vectorized and a vector knowledge index is established and stored in a knowledge vector library. The compression rate and accuracy of the declarative knowledge in the vector knowledge index are evaluated by a declarative knowledge evaluation system, and the process of multimodal knowledge point fusion, dual graph construction, graph community network clustering and association mapping is optimized based on the evaluation results.
9. A computer device, characterized in that, The computer device includes a processor and a non-volatile memory storing computer instructions. When the computer instructions are executed by the processor, the computer device performs the method according to any one of claims 1-7.
10. A readable storage medium, characterized in that, The readable storage medium includes a computer program, which, when executed, controls the computer device on which the readable storage medium is located to perform the method described in any one of claims 1-7.
Citation Information
Patent Citations
Multi-modal knowledge graph construction method
CN112200317A
Knowledge graph construction method, knowledge graph construction system and computing equipment
CN114817553A
Data integration method based on knowledge graph
CN118333059A
Method for constructing vertical domain knowledge graph and related device
CN118504679A
Knowledge graph construction method based on fine-tuning large language model
CN119808917A
Cited By
Text processing method and related equipment
CN121279296A
Data processing method and device for VR interactive training of cross-country container forklift
CN121640580A
Educational resource big data visual analysis method and system based on knowledge graph
CN121658669A
Knowledge graph-based education resource big data visual analysis method and system
CN121658669B
Knowledge base updating method and system
CN121835853A