Hierarchical data classification processing and deep correlation analysis system and method based on knowledge graph

By using a hierarchical data classification and deep correlation analysis system based on knowledge graphs, the problems of unified preprocessing, classification storage, and deep correlation of multimodal forensic data have been solved, enabling efficient and dynamic data analysis and decision support.

CN120974343APending Publication Date: 2025-11-18BEIJING XINSI NETWORK TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202511117882.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-11
Publication Date
2025-11-18

Smart Images

  • Figure CN120974343A_ABST
    Figure CN120974343A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, and discloses a hierarchical data classification processing and deep correlation analysis system and method based on a knowledge graph, and the system comprises an acquisition module which is configured to obtain multi-modal evidence obtaining data, and carries out the preprocessing of the multi-modal evidence obtaining data; the data processing module is configured to respectively construct data sets according to the preprocessed multi-modal evidence obtaining data; the graph construction module is configured to extract main body data of the data sets, divide the data sets with the same main body data into the same category, compare all data in the data sets with the same category, and construct a knowledge graph according to a comparison result; and the association analysis module is configured to update the knowledge graph according to the multi-level association mining result. According to the method, the data processing efficiency is improved, and the mining capability of deep association of the data is enhanced by constructing and continuously optimizing the knowledge graph.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data processing, in particular to a hierarchical data classification processing and deep correlation analysis system and method based on a knowledge graph. BACKGROUND

[0002] Multimodal forensic data refers to information from different data sources or data types, such as text, images, audio, and video, etc. These data often contain rich information, but at the same time bring complexity in processing.

[0003] In monitoring, forensics and other work, we often need to face massive multimodal forensic data, which covers text, images, video, audio and other types. However, the existing technology has many limitations in processing such multimodal data. Traditional data processing methods often lack unified preprocessing standards. For different modal data, independent processing methods are usually used, resulting in chaotic data formats and uneven data quality, which is difficult to meet the requirements of subsequent analysis on data quality. At the same time, in terms of data management, the existing technology does not specially classify and store multimodal data by modality, resulting in chaotic data management, making it difficult to efficiently carry out special analysis of different modal data, greatly affecting the efficiency and accuracy of data processing. In terms of data correlation analysis, the existing technology cannot effectively extract the subject information in the data and establish correlations. Due to the complexity of the source of multimodal data and the diversity of the format, the internal relationship between the data is hidden, and the traditional method cannot construct a knowledge graph that can clearly reflect the deep correlation of the data, making it difficult to fully mine the valuable information hidden behind the data, and difficult to meet the needs of government projects for data deep analysis to support decision-making. In addition, the existing system has deficiencies in the continuity and dynamics of correlation mining, and cannot update the data correlation relationship based on new correlation mining results, making the analysis results lag behind, unable to adapt to the increasing requirements of government projects for data real-time and accuracy, and difficult to achieve dynamic and accurate analysis and decision support for data.

[0004] Therefore, it is necessary to provide a hierarchical data classification processing and deep correlation analysis system and method based on a knowledge graph to solve the limitations of the existing technology in processing and managing massive multimodal forensic data, lack of unified preprocessing standards, chaotic data classification storage, difficulty in establishing deep correlation between data, and inability to achieve dynamic and accurate analysis and decision support for data. SUMMARY

[0005] In view of this, the present application proposes a hierarchical data classification processing and deep correlation analysis system and method based on a knowledge graph, aiming to solve the problems that the prior art has limitations in processing and managing massive multi-modal forensic data, lacks unified preprocessing standards, data classification storage is chaotic, it is difficult to establish deep correlation between data, and dynamic and accurate analysis and decision support of data cannot be realized.

[0006] In one aspect, the present application proposes a hierarchical data classification processing and deep correlation analysis system based on a knowledge graph, comprising:

[0007] The acquisition module is configured to obtain multi-modal forensic data and preprocess the multi-modal forensic data;

[0008] The data processing module is configured to construct data sets according to the preprocessed multi-modal forensic data; wherein one data set stores forensic data of one modality;

[0009] The graph construction module is configured to extract subject data of the data sets, divide data sets of the same subject data into the same category, compare all data in the data sets of the same category, and construct a knowledge graph according to the comparison result;

[0010] The correlation analysis module is configured to perform multi-level correlation mining based on the constructed knowledge graph and update the knowledge graph according to the multi-level correlation mining result.

[0011] Further, the acquisition module is configured to obtain multi-modal forensic data and preprocess the multi-modal forensic data, comprising:

[0012] The multi-modal forensic data includes text data, image data, video data and audio data;

[0013] The text information of the text data is extracted through OCR technology, and the unstructured text is converted into structured entries through sentence division and stop word removal processing combined with NLP tools;

[0014] The key frames of the video data are extracted using OpenCV, the entities of the key frames and the image data are identified through a target detection algorithm, and the entity information is respectively bound with a timestamp and a geographic location;

[0015] The audio data is converted into text information through a speech-to-text engine;

[0016] The converted or recognized data is converted into a unified format.

[0017] Further, the data processing module is configured to construct data sets according to the preprocessed multi-modal forensic data, comprising:

[0018] Reserve the attributes of the multi-modal forensic data in a unified format;

[0019] Divide the multi-modal forensic data in a unified format into different data sets according to the modalities;

[0020] Remove the repeated data with the same attributes in the same data set.

[0021] Further, the graph construction module is configured to extract the subject data of the data set, and when the data sets with the same subject data are divided into the same category, it includes:

[0022] Extract the subject data and the associated data in the data set, and label them respectively; wherein the subject data includes entities, events, characters and places; the associated data includes the relationship and attributes between the subject data;

[0023] According to the labeling result, the data sets with the same subject data are divided into the same category;

[0024] If there are multiple data sets in the same category, data comparison between the data sets is performed, otherwise, no data comparison between the data sets is performed.

[0025] Further, when all the data in the data sets of the same category are compared, and a knowledge graph is constructed according to the comparison result, it includes:

[0026] Compare the data sets of the same category two by two;

[0027] Among them, the data with the same attributes in the data set are compared, and it is determined whether to retain the data according to the comparison result.

[0028] Further, when the data with the same attributes in the data set are compared, and it is determined whether to retain the data according to the comparison result, it includes:

[0029] If the comparison result shows that the data with the same attributes are the same, remove the repeated data and retain any one of the data;

[0030] If the comparison result shows that the data with the same attributes are different, compare the collection time of the data, if the collection time is the same, mark the data as abnormal, if the collection time is different, retain the latest data.

[0031] Further, when all the data in the data sets of the same category are compared, and a knowledge graph is constructed according to the comparison result, it further includes:

[0032] The retained data is constructed into a knowledge graph according to the subject data and the associated data, and a sub-knowledge graph is constructed for each category of data set;

[0033] If there is same subject data between the sub-knowledge graphs, the sub-knowledge graphs are combined, and the combined knowledge graph is taken as the final knowledge graph.

[0034] The construction of the knowledge graph includes the establishment of nodes and edges, the nodes represent subject data, and the edges represent associated data.

[0035] Further, the correlation analysis module is configured to perform multi-level correlation mining based on the constructed knowledge graph, and when updating the knowledge graph according to the multi-level correlation mining result, the method comprises:

[0036] The multi-level correlation mining of the knowledge graph is performed, the semantic similarity between entities is calculated through a graph neural network, the cross-modal entity correlation is established, the causal logic relationship between events is mined by using a time sequence reasoning algorithm, and an event evolution chain is constructed.

[0037] According to the result of the multi-level correlation mining, the nodes and edges in the knowledge graph are automatically updated to reflect new data and associated relationships.

[0038] Further, the multi-level correlation mining based on the constructed knowledge graph is performed, and after the knowledge graph is updated according to the multi-level correlation mining result, the method comprises:

[0039] The knowledge graph and the update result thereof are displayed through a visualization tool.

[0040] The multi-level correlation relationship between entities is presented through a hierarchical structure view, different colors and line styles are used to distinguish entity types, relationship strengths and time sequence characteristics, and an interactive query function is provided.

[0041] According to subject data, associated data or a time range, relevant nodes and edges are searched, the display content is updated in real time according to the operation, and an associated analysis report is automatically generated.

[0042] Compared with the prior art, the present application has the beneficial effects that the acquisition module of the present application is responsible for acquiring multi-modal forensic data and performing necessary preprocessing to ensure the quality and availability of the data. This stage is the foundation of the entire system, which guarantees the accuracy and reliability of subsequent analysis. The data processing module constructs data sets according to the preprocessed data, and each data set stores forensic data of one modality, which not only helps data management and maintenance, but also provides convenience for subsequent analysis. By classifying data by modality, special processing and analysis can be performed according to the characteristics of different modalities, thereby improving the accuracy of analysis. The graph construction module is one of the cores of the system, extracts the main data in the data set, and divides the data sets of the same main data into the same category. By comparing all data in the data sets of the same category, the system can construct a knowledge graph. This knowledge graph not only contains the association between data, but also reveals the deep connection behind the data, providing strong support for decision-making. The correlation analysis module performs multi-level correlation mining based on the constructed knowledge graph, and through in-depth analysis of the correlation between data, it can discover the potential connection and pattern between data. This analysis result can be used to update the knowledge graph, forming a dynamic and self-improving analysis system. With the continuous updating and optimization of the knowledge graph, the analysis capability of the system will gradually be enhanced, thereby realizing more accurate prediction and decision-making. In summary, the present application realizes efficient processing and deep analysis of multi-modal forensic data through modular design, not only improves the efficiency of data processing, but also enhances the mining ability of deep-level correlation of data by constructing and continuously optimizing the knowledge graph.

[0043] On the other hand, the present application also provides a hierarchical data classification processing and deep correlation analysis method based on a knowledge graph, comprising:

[0044] acquiring multi-modal forensic data and preprocessing the multi-modal forensic data;

[0045] constructing data sets according to the preprocessed multi-modal forensic data; wherein one data set stores forensic data of one modality;

[0046] extracting the main data of the data set, dividing the data sets of the same main data into the same category, comparing all data in the data sets of the same category, and constructing a knowledge graph according to the comparison result;

[0047] performing multi-level correlation mining based on the constructed knowledge graph, and updating the knowledge graph according to the multi-level correlation mining result.

[0048] It can be understood that the hierarchical data classification processing and deep correlation analysis system and method based on the knowledge graph provided by the present application have the same beneficial effects, which will not be described here. BRIEF DESCRIPTION OF DRAWINGS

[0049] Various other advantages and benefits will become apparent to those of ordinary skill in the art upon reading the following detailed description of the preferred embodiments with reference made to the accompanying drawings. The drawings are for purposes of illustration only and are not intended to limit the present application thereto. Like reference numerals have been used throughout the drawings to denote like parts. In the drawings:

[0050] Figure 1 A functional block diagram of a hierarchical data classification processing and deep correlation analysis system based on a knowledge graph is provided for an embodiment of the present application;

[0051] Figure 2 A flowchart of a hierarchical data classification processing and deep correlation analysis method based on a knowledge graph is provided for an embodiment of the present application. DETAILED DESCRIPTION

[0052] Exemplary embodiments of the present disclosure will be described below in greater detail with reference to the accompanying drawings. Although exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be accurately conveyed to those skilled in the art. It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict. The present application will be described in detail below with reference to the accompanying drawings and in conjunction with the embodiments.

[0053] In some embodiments of the present application, referring to Figure 1 The present embodiment provides a hierarchical data classification processing and deep correlation analysis system based on a knowledge graph, which comprises:

[0054] The acquisition module is configured to acquire multi-modal forensic data and pre-process the multi-modal forensic data;

[0055] The data processing module is configured to construct data sets according to the pre-processed multi-modal forensic data; wherein one data set stores forensic data of one modality;

[0056] The graph construction module is configured to extract subject data of the data sets, divide data sets of the same subject data into the same category, compare all data in the data sets of the same category, and construct a knowledge graph according to the comparison results;

[0057] The correlation analysis module is configured to perform multi-level correlation mining based on the constructed knowledge graph, and update the knowledge graph according to the multi-level correlation mining results.

[0058] It can be understood that first, the acquisition module is responsible for obtaining multi-modal forensic data and performing necessary preprocessing to ensure the quality and availability of the data. This stage is the foundation of the entire system, which guarantees the accuracy and reliability of subsequent analysis. The data processing module constructs data sets based on the preprocessed data, each of which stores forensic data of one modality, which not only helps data management and maintenance, but also provides convenience for subsequent analysis work. By classifying data by modality, special processing and analysis can be performed according to the characteristics of different modalities, thereby improving the accuracy of analysis. The graph construction module is one of the cores of the system, which extracts the main data in the data set and divides the data sets of the same main data into the same category. By comparing all data in the data sets of the same category, the system can construct a knowledge graph. This knowledge graph not only contains the association between data, but also reveals the deep connection behind the data, providing strong support for decision-making. The correlation analysis module performs multi-level correlation mining based on the constructed knowledge graph, and through in-depth analysis of the correlation between data, it can discover the potential connection and pattern between data. This analysis result can be used to update the knowledge graph, forming a dynamic and self-improving analysis system. With the continuous updating and optimization of the knowledge graph, the analysis capability of the system will gradually be enhanced, thereby realizing more accurate prediction and decision-making. In summary, the present application realizes efficient processing and deep analysis of multi-modal forensic data through modular design, not only improves the efficiency of data processing, but also enhances the mining ability of deep-level correlation of data by constructing and continuously optimizing the knowledge graph.

[0059] In some embodiments of the present application, the acquisition module is configured to obtain multi-modal forensic data, and when preprocessing the multi-modal forensic data, it includes:

[0060] The multi-modal forensic data includes text data, image data, video data and audio data;

[0061] The text information of the text data is extracted through OCR technology, and the unstructured text is converted into structured entries through sentence division and stop word removal processing combined with NLP tools;

[0062] The key frames of the video data are extracted using OpenCV, the entities of the key frames and image data are identified through a target detection algorithm, and the entity information is bound with timestamps and geographical locations respectively;

[0063] The audio data is converted into text information through a speech-to-text engine;

[0064] The converted or recognized data is converted into a unified format.

[0065] It can be understood that the acquisition module is configured to obtain multi-modal forensic data, which includes text, image, video and audio data. The text information in the text data is extracted by OCR technology, and the unstructured text information is converted into structured terms by combining with natural language processing (NLP) tools for sentence segmentation and stop word removal. The key frames of video data are extracted by using OpenCV, and the entities in the key frames and image data are identified by target detection algorithm, and these entity information is bound with time stamp and geographic location information. The audio data is converted into text information by speech-to-text engine. Finally, these converted or identified data are uniformly formatted for further analysis and processing. The benefit of this multi-modal data acquisition and preprocessing method is that it can comprehensively capture and analyze various aspects of the event, improve the accuracy and efficiency of forensics, and provide rich and structured information for subsequent data analysis.

[0066] In some embodiments of the present application, when the data processing module is configured to construct data sets according to the preprocessed multi-modal forensic data respectively, it includes:

[0067] Preserving the attributes of the multi-modal forensic data in uniform format;

[0068] Dividing the multi-modal forensic data in uniform format into different data sets according to modalities;

[0069] Removing duplicate data with the same attributes in the same data set.

[0070] It can be understood that the data processing module constructs data sets by preserving the attributes of the multi-modal forensic data in uniform format, dividing the multi-modal forensic data in uniform format into different data sets according to modalities, and removing duplicate data with the same attributes in the same data set. First, maintaining the consistency of data attributes helps the accuracy and consistency of subsequent analysis; second, dividing data sets by modalities can be more fine-grained processing and analysis according to the characteristics of different modalities; finally, removing duplicate data can reduce data redundancy and improve data processing efficiency and analysis accuracy. These steps work together to make data processing more efficient and accurate, laying a solid foundation for subsequent data analysis and forensic work.

[0071] In some embodiments of the present application, when the graph construction module is configured to extract the subject data of the data sets and divide the data sets with the same subject data into the same category, it includes:

[0072] Extracting subject data and associated data in the data sets and labeling them respectively; wherein the subject data includes entities, events, characters and places; the associated data includes relationships and attributes between subject data;

[0073] According to the annotation result, data sets with the same subject data are divided into the same category;

[0074] If there are multiple data sets in the same category, data comparison between the data sets is performed, otherwise, data comparison between the data sets is not performed.

[0075] In some embodiments of the present application, when comparing all data in the data sets of the same category and constructing a knowledge graph according to the comparison result, it includes:

[0076] The data sets of the same category are compared in pairs;

[0077] Among them, the data of the same attribute in the data set are compared, and whether to retain the data is determined according to the comparison result.

[0078] In some embodiments of the present application, when comparing the data of the same attribute in the data set and determining whether to retain the data according to the comparison result, it includes:

[0079] If the comparison result shows that the data of the same attribute is the same, the duplicate data is removed and any one of the data is retained;

[0080] If the comparison result shows that the data of the same attribute is different, the collection time of the data is compared, if the collection time is the same, the data is marked as abnormal, if the collection time is different, the latest data is retained.

[0081] It can be understood that the graph construction module effectively divides the data sets into categories with the same subject data by extracting the subject data and associated data in the data sets and annotating. The subject data includes entities, events, characters and places, and the associated data relates to the relationships and attributes between these subject data. Through this classification, data sets with the same subject data can be classified into the same category, thereby simplifying subsequent data processing and analysis work. When there are multiple data sets in the same category, the system will compare the data sets to ensure the accuracy and consistency of the data. In the comparison process, the data of the same attribute is compared one by one to determine whether to retain the data. If the comparison result shows that the data is the same, the duplicate is removed and one is retained; if the data is different, the latest data is retained or marked as abnormal according to the collection time. This processing method not only improves the efficiency of data processing, but also ensures the accuracy and reliability of the knowledge graph.

[0082] In some embodiments of the present application, when comparing all data in the data sets of the same category and constructing a knowledge graph according to the comparison result, it further includes:

[0083] The retained data is used to construct a knowledge graph according to the subject data and the associated data, and a sub-knowledge graph is constructed for each category of data set.

[0084] If there are same subject data between the sub-knowledge graphs, the sub-knowledge graphs are combined, and the combined knowledge graph is taken as the final knowledge graph.

[0085] The construction of the knowledge graph includes the establishment of nodes and edges, the nodes represent subject data, and the edges represent associated data.

[0086] It can be understood that by comparing all data in the same category dataset and constructing a knowledge graph, in-depth analysis and effective integration of data can be achieved. First, the retained data is used to construct a knowledge graph according to subject data and associated data, which enables each category dataset to construct a sub-knowledge graph, thereby realizing the classification management of data. Second, if there are same subject data between the sub-knowledge graphs, these sub-knowledge graphs are combined to form a more comprehensive and integrated knowledge graph, which helps to reveal the internal relationship and interaction between different datasets. The construction of the knowledge graph includes the establishment of nodes and edges, wherein the nodes represent subject data and the edges represent associated data. This structured approach helps to clearly show the relationship between data, facilitating further data mining and analysis. In summary, this method of constructing a knowledge graph can improve the efficiency and quality of data processing.

[0087] In some embodiments of the present application, the correlation analysis module is configured to perform multi-level correlation mining based on the constructed knowledge graph, and when updating the knowledge graph according to the multi-level correlation mining result, it includes:

[0088] Performing multi-level correlation mining on the knowledge graph, calculating the semantic similarity between entities through a graph neural network, establishing cross-modal entity correlation, and using a time sequence reasoning algorithm to mine the causal logic relationship between events to construct an event evolution chain.

[0089] According to the result of multi-level correlation mining, automatically update the nodes and edges in the knowledge graph to reflect new data and associated relationships.

[0090] In some embodiments of the present application, after performing multi-level correlation mining based on the constructed knowledge graph and updating the knowledge graph according to the multi-level correlation mining result, it includes:

[0091] Displaying the knowledge graph and its update result through a visualization tool;

[0092] Presenting the multi-level correlation relationship between entities through a hierarchical structure view, distinguishing entity types, relationship strengths, and time sequence characteristics through different colors and line styles, and providing an interactive query function;

[0093] The related nodes and edges can be retrieved according to subject data, associated data or a time range, and the display content can be updated in real time according to the operation to automatically generate an associated analysis report.

[0094] It can be understood that the associated analysis module performs multi-level association mining through the constructed knowledge graph, calculates the semantic similarity between entities by using a graph neural network, establishes cross-modal entity association, and mines the causal logic relationship between events by using a time sequence reasoning algorithm to construct an event evolution chain. The mining result is used to automatically update the nodes and edges in the knowledge graph to reflect new data and associated relationships. In addition, the knowledge graph and its update result are displayed by using a visualization tool, the multi-level association relationship between entities is presented in a hierarchical structure view, and different colors and line styles are used to distinguish entity types, relationship strengths and time sequence characteristics. An interactive query function is also provided. Users can retrieve related nodes and edges according to subject data, associated data or a time range, update the display content in real time, and automatically generate an associated analysis report. The system not only enhances the depth and breadth of data processing, but also improves the intuitiveness and convenience of user interaction.

[0095] On the other hand, referring to Figure 2 It can be understood that the associated analysis module performs multi-level association mining through the constructed knowledge graph, calculates the semantic similarity between entities by using a graph neural network, establishes cross-modal entity association, and mines the causal logic relationship between events by using a time sequence reasoning algorithm to construct an event evolution chain. The mining result is used to automatically update the nodes and edges in the knowledge graph to reflect new data and associated relationships. In addition, the knowledge graph and its update result are displayed by using a visualization tool, the multi-level association relationship between entities is presented in a hierarchical structure view, and different colors and line styles are used to distinguish entity types, relationship strengths and time sequence characteristics. An interactive query function is also provided. Users can retrieve related nodes and edges according to subject data, associated data or a time range, update the display content in real time, and automatically generate an associated analysis report. The system not only enhances the depth and breadth of data processing, but also improves the intuitiveness and convenience of user interaction.

[0096] S100, acquiring multi-modal forensic data and preprocessing the multi-modal forensic data;

[0097] S200, constructing data sets according to the preprocessed multi-modal forensic data; wherein the forensic data of one modality is stored in one data set;

[0098] S300, extracting subject data of the data set, dividing data sets with the same subject data into the same category, comparing all data in the data set of the same category, and constructing a knowledge graph according to the comparison result;

[0099] S400, performing multi-level association mining based on the constructed knowledge graph, and updating the knowledge graph according to the multi-level association mining result.

[0100] It can be understood that first, by acquiring multi-modal forensic data and preprocessing, the method can integrate different types of data sources, thereby providing a more comprehensive information perspective. Multi-modal data processing not only increases the richness of data, but also improves the accuracy and reliability of analysis, because data of different modalities often complement each other to provide a more complete scene description. Second, the method realizes the classification and preliminary correlation analysis of data by constructing a data set and extracting subject data. Dividing the data set of the same subject into the same category helps to organize complex data sets into a more manageable and analyzable form. This classification process not only improves the efficiency of data processing, but also lays a solid foundation for subsequent in-depth analysis. Further, as a powerful data organization and correlation tool, the knowledge graph can reveal the complex relationships and patterns between data. By performing multi-level correlation mining on the knowledge graph, deep connections between data can be discovered, which is crucial for understanding complex systems and pattern recognition. In addition, updating the knowledge graph according to the mining results ensures the dynamic and real-time nature of the knowledge graph, allowing it to reflect the latest data relationships and patterns. In summary, the method of the present application not only improves the efficiency and accuracy of data processing, but also realizes in-depth mining and understanding of complex data relationships through the construction and dynamic updating of the knowledge graph. This has broad application prospects in fields such as forensic analysis, intelligence analysis, intelligent recommendation, and can provide more scientific and accurate basis for decision-making.

[0101] Those skilled in the art will appreciate that embodiments of the present application can be provided as methods, systems or computer program products. Accordingly, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment combining software and hardware aspects. Furthermore, the present application can take the form of a computer program product on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage devices, etc.) embodying computer-readable program code.

[0102] The present application is described with reference to the flowcharts and / or block diagrams according to the methods, devices (systems) and computer program products of the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing apparatus to produce a machine, so that the instructions executed by the computer or other programmable data processing apparatus produce a device that implements the functions specified in the flowcharts and / or block diagrams. Figure 1 The device that implements the functions specified in one flow or multiple flows and / or blocks. Figure 1 The device that implements the functions specified in one flow or multiple flows and / or blocks.

[0103] These computer program instructions can also be stored in a computer readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer readable memory produce an article of manufacture including instructions which implement the flow Figure 1 of the flow or flows and / or blocks Figure 1 of the block or blocks specified in the flow.

[0104] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the flow Figure 1 of the flow or flows and / or blocks Figure 1 of the block or blocks specified in the flow.

[0105] Finally, it should be noted that the above-mentioned embodiments are merely used to illustrate the technical solutions of the present application, rather than limit the present application. Although the present application has been described in detail with reference to the above-mentioned embodiments, those skilled in the art should understand that the technical solutions of the present application can be modified or equivalent replaced without departing from the spirit and scope of the present application, and any modification or equivalent replacement should be covered in the protection scope of the claims of the present application.

Claims

1. A hierarchical data classification and deep association analysis system based on knowledge graphs, characterized in that, include: The acquisition module is configured to acquire multimodal forensic data and preprocess the multimodal forensic data. The data processing module is configured to construct datasets based on the preprocessed multimodal forensics data; each dataset stores forensics data for one modality. The knowledge graph construction module is configured to extract the main data of the dataset, divide datasets with the same main data into the same category, compare all data within the same category, and construct a knowledge graph based on the comparison results. The association analysis module is configured to perform multi-level association mining based on the constructed knowledge graph, and update the knowledge graph based on the results of multi-level association mining.

2. The hierarchical data classification processing and deep association analysis system based on knowledge graphs according to claim 1, characterized in that, The acquisition module is configured to acquire multimodal forensic data, and the preprocessing of the multimodal forensic data includes: The multimodal forensic data includes text data, image data, video data, and audio data; Text information is extracted from text data using OCR technology, and then combined with NLP tools for sentence segmentation and stop word removal to transform unstructured text into structured terms. Keyframes from video data are extracted using OpenCV. Entities in keyframes and image data are identified using object detection algorithms, and entity information is bound to timestamps and geographic locations, respectively. Audio data is converted into text information using a speech-to-text engine; Transform or identify the data into a unified format.

3. The hierarchical data classification processing and deep association analysis system based on knowledge graphs according to claim 2, characterized in that, When the data processing module is configured to construct datasets based on the preprocessed multimodal forensics data, it includes: Preserve the attributes of multimodal forensic data in a unified format; Multimodal forensic data in a unified format is divided into different datasets based on modality; Remove duplicate data with the same attributes from the same dataset.

4. The hierarchical data classification processing and deep association analysis system based on knowledge graphs according to claim 3, characterized in that, The graph construction module is configured to extract the main data of the dataset, and when dividing datasets with the same main data into the same category, it includes: Extract the main data and related data from the dataset and label them respectively; wherein, the main data includes entities, events, people and places; the related data includes the relationships and attributes between the main data; Data sets with the same main data are grouped into the same category based on the annotation results; If multiple datasets exist within the same category, then data comparison between datasets is performed; otherwise, data comparison between datasets is not performed.

5. The hierarchical data classification processing and deep association analysis system based on knowledge graphs according to claim 4, characterized in that, When comparing all data within the same category of datasets and constructing a knowledge graph based on the comparison results, the following is included: Compare datasets of the same category pairwise; This involves comparing data with the same attributes in the dataset and determining whether to retain the data based on the comparison results.

6. The hierarchical data classification processing and deep association analysis system based on knowledge graphs according to claim 5, characterized in that, The step of comparing data with the same attribute in the dataset and determining whether to retain the data based on the comparison results includes: If the comparison results show that the data with the same attribute are the same, then remove the duplicate data and keep any one of the data; If the comparison results show that the data for the same attribute are different, then the data collection time is compared. If the collection times are the same, the data is marked as abnormal; if the collection times are different, the latest data is retained.

7. The hierarchical data classification processing and deep association analysis system based on knowledge graphs according to claim 6, characterized in that, When comparing all data within the same category of datasets and constructing a knowledge graph based on the comparison results, the process also includes: The retained data is used to construct a knowledge graph based on the main data and related data, and a sub-knowledge graph is constructed for each category of dataset; If there is the same main data among the sub-knowledge graphs, the sub-knowledge graphs are combined and the combined knowledge graph is used as the final knowledge graph. The construction of a knowledge graph includes the establishment of nodes and edges, where nodes represent the main data and edges represent related data.

8. The hierarchical data classification processing and deep association analysis system based on knowledge graphs according to claim 7, characterized in that, The association analysis module is configured to perform multi-level association mining based on the constructed knowledge graph. When updating the knowledge graph based on the multi-level association mining results, it includes: Multi-level association mining is performed on the knowledge graph, semantic similarity between entities is calculated through graph neural networks, cross-modal entity associations are established, and temporal reasoning algorithms are used to mine causal logical relationships between events and construct event evolution chains. Based on the results of multi-level association mining, the nodes and edges in the knowledge graph are automatically updated to reflect new data and relationships.

9. The hierarchical data classification processing and deep association analysis system based on knowledge graphs according to claim 8, characterized in that, The process of performing multi-level association mining based on the constructed knowledge graph, and updating the knowledge graph according to the results of the multi-level association mining, includes: Visualization tools are used to display the knowledge graph and its update results. The hierarchical structure view presents the multi-level relationships between entities, distinguishes entity types, relationship strengths, and time-series features with different colors and line styles, and provides interactive query functions. It supports retrieving relevant nodes and edges based on main data, related data, or time range, and updates the displayed content in real time according to the operation, automatically generating a correlation analysis report.

10. A hierarchical data classification and deep association analysis method based on knowledge graphs, applied to the hierarchical data classification and deep association analysis system based on knowledge graphs as described in any one of claims 1-9, characterized in that, include: Acquire multimodal forensic data and preprocess the multimodal forensic data; Data sets are constructed based on the preprocessed multimodal forensics data; each dataset stores forensics data for one modality. Extract the main data from the dataset, divide datasets with the same main data into the same category, compare all data within the same category, and construct a knowledge graph based on the comparison results; Multi-level association mining is performed based on the constructed knowledge graph, and the knowledge graph is updated according to the results of multi-level association mining.

Citation Information

Patent Citations

  • A knowledge graph system construction method

    CN109697233A

  • Dynamic ontology-based knowledge graph analysis and application method, platform and device

    CN113326345A

  • Metadata-based knowledge graph construction method and device, equipment and storage medium

    CN114840686A

  • Intelligent visualization and text association method for multi-modal knowledge graph

    CN119441281A

  • An information processing method and system based on multiple information sources

    CN119760572A