Normalization method and system for multi-source, multi-mode and multi-format heterogeneous data

By normalizing the heterogeneous data of multi-source, multi-modal, and multi-formats, standardizing it into a soc security event model, the problem of inconsistent log formats of security equipment within the enterprise is solved, and data parsing efficiency and unified management capabilities are improved.

CN120492419APending Publication Date: 2025-08-15YUNNAN POWER GRID CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510316120.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

In the prior art, due to historical systems and information barriers between departments, there are many types of security equipment used within enterprises and inconsistent log formats, resulting in uneven data governance quality, difficulty in integration, and inefficient utilization.

Method used

Using multi-source, multi-modal, multi-format heterogeneous data normalization methods, through data acquisition, preprocessing, analytical strategies and visualization operations, log data from different sources and formats is standardized into a soc security event model, and stored on a big data platform for unified management.

Benefits of technology

It improves data analysis efficiency and accuracy, simplifies analysis strategy management, enhances the flexibility and adaptability of the method, and realizes unified collection and aggregation of data from different sources, types and formats, providing a foundation for subsequent processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120492419A_ABST
    Figure CN120492419A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-source, multi-mode and multi-format heterogeneous data normalization method and system, and relates to the technical field of data processing, and the method comprises the steps: collecting data aggregation of different sources, different types and different formats; the method comprises the following steps: preprocessing data, distinguishing log sources, extracting effective logs, introducing an analysis strategy to log data, standardizing the logs into a soc security event model, normalizing the log data, and carrying out visual operation on the log data through log access configuration; and loading the normalized data into a big data platform layer for storage and calculation. According to the method, the analysis is preferentially performed through the specific analysis strategy, so that unnecessary ordinary analysis strategy attempts are avoided, computing resources are saved, and the analysis efficiency and accuracy are improved. When a specific analysis strategy fails, a standby scheme is provided through a common analysis strategy, multiple analysis strategies are supported, the flexibility and adaptability of the method are enhanced, and the overall analysis success rate is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a method and system for normalizing multi-source, multi-modal, and multi-format heterogeneous data. Background Art

[0002] In security situation awareness systems, data plays an increasingly central role, becoming a critical element in corporate strategic decision-making. However, due to legacy systems and interdepartmental information barriers, internal security devices are used in a wide variety of formats and from different vendors. Logs for each device type vary, and even log data for different versions of the same device type from the same vendor can differ. This results in inconsistent data quality, integration difficulties, and inefficient utilization for network provincial companies. In this context, data standardization is a key component in addressing these challenges. A unified data model aims to improve data quality and reduce data development and maintenance costs, thereby better supporting responsible data analysis and computational logic. Summary of the Invention

[0003] In view of the above-mentioned problems, the present invention is proposed.

[0004] Therefore, the problem this invention aims to solve is that, due to legacy systems and information barriers between departments, the current situation is that internal security devices used are diverse and come from different manufacturers. Each type of device has unique logs, and even different versions of the same device from the same manufacturer have different log data. This leads to inconsistent data quality, integration difficulties, and low utilization efficiency in network provincial companies.

[0005] To solve the above technical problems, the present invention provides the following technical solutions: a method for normalizing heterogeneous data with multiple sources, multiple modalities and multiple formats, which includes collecting data aggregation from different sources, different types and different formats; preprocessing the data, distinguishing log sources and extracting valid logs, introducing parsing strategies for log data, standardizing logs into soc security event models, normalizing log data, and visualizing log data through log access configuration; loading the normalized data into the big data platform layer for storage and calculation.

[0006] As a preferred solution of the method for normalizing multi-source, multi-modal, and multi-format heterogeneous data described in the present invention, the different sources include but are not limited to traffic log data obtained by traffic probes through DPI, synchronization logs or event information logs from third-party network devices and security devices; after collecting data, a log source identifier is added to the data based on the collection source of the data aggregation.

[0007] As a preferred solution of the method for normalizing heterogeneous data with multiple sources, multiple modalities and multiple formats described in the present invention, the data preprocessing includes distinguishing log sources and extracting valid logs. Through the high-level extraction configuration of log access, the log sources are distinguished for separating and labeling different device logs when different security device log sources are mixed together, and Latin1 encoding is used to solve the problem of inconsistent encoding of different devices. When a forwarding header is added to the security device log through forwarding, the forwarding header is removed by extracting the valid log, so that the normalization parsing strategy focuses on the parsing of the security device log, thereby realizing the reusability of the parsing strategy.

[0008] As a preferred solution of the method for normalizing heterogeneous data with multiple sources, multiple modalities and multiple formats described in the present invention, wherein: the parsing strategy corresponds to different types of logs of the same security device, supports import and export, and the strategies under the same parsing strategy group support priority adjustment and quick matching; the parsing strategy group contains several parsing strategies; the parsing strategy supports visual configuration, regenerates groovy scripts, or directly writes groovy scripts and then imports them; the parsing methods include delimiters, regular expressions, JSON, CEF and LEEF, and supports nested parsing of multiple parsing methods; the field mapping in the parsing rules supports direct mapping, time conversion, IP conversion, Base64 conversion and association mapping; the regular expression parsing method supports word extraction, and selects the corresponding field regular expression from the regular expression library.

[0009] As a preferred solution of the method for normalizing heterogeneous data with multiple sources, multiple modalities and multiple formats described in the present invention, the method includes: introducing a parsing strategy into the log data and standardizing the log into a soc security event model, locating the parsing strategy group for the log data to be normalized, obtaining all quick matching strategies, selecting strategies from high to low priority for quick matching, observing whether the quick match is hit, and if not, selecting the next level parsing strategy according to priority until the quick match is hit or the remaining parsing strategies in the parsing strategy group are zero; if the quick match is hit, parsing the log data to determine whether the parsing is successful, and if the parsing is successful, terminating the normalization process; if the parsing is unsuccessful or the remaining parsing strategies in the parsing strategy group are zero, obtaining all ordinary parsing strategies, selecting ordinary strategies from high to low priority for parsing, and terminating the normalization process after the parsing is successful or all ordinary parsing strategies have participated in the parsing.

[0010] As a preferred solution of the method for normalizing heterogeneous data with multiple sources, multiple modalities and multiple formats described in the present invention, the visualization operation of log data through log access configuration includes: security operations and customers completing the entire log access process through a visualization page, configuring multiple log sources through advanced extraction for the same access method, supporting priority adjustment, monitoring and viewing the most recently accessed logs, and verifying advanced extraction rules based on the accessed logs; decoupling the parsing policy group configuration and the log access configuration, and selecting to configure an existing parsing policy group during the log access process; or not setting the parsing policy group, and setting the parsing policy group for the log source after completing the matching of the parsing policy group.

[0011] As a preferred solution of the method for normalizing multi-source, multi-modal, and multi-format heterogeneous data described in the present invention, the normalized data is loaded into the big data platform layer for storage and calculation, including classifying and storing different types of collected data to meet the requirements of data analysis; data calculation provides general real-time analysis and offline calculation capabilities to meet the requirements of security analysis; based on the different values of different data, different storage strategies and analysis strategies are adopted for different values.

[0012] Another object of the present invention is to provide a normalization system for heterogeneous data with multiple sources, multiple modalities, and multiple formats, which can normalize and store heterogeneous data with different sources, different modalities, and different formats.

[0013] To solve the above technical problems, the present invention provides the following technical solutions: a system for a normalization method of heterogeneous data with multiple sources, multiple modalities and multiple formats, comprising: a data acquisition module, a data processing module and a data storage module; the data acquisition module aggregates data from different sources, different types and different formats; the data processing module pre-processes the data, distinguishes the log sources and extracts valid logs, introduces parsing strategies for the log data, standardizes the logs into a soc security event model, normalizes the log data, and visualizes the log data through log access configuration; the data storage module loads the normalized data into the big data platform layer for storage and calculation.

[0014] A computer device includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, it implements the steps of the above-mentioned method for normalizing heterogeneous data with multiple sources, multiple modalities, and multiple formats.

[0015] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the above-mentioned method for normalizing heterogeneous data with multiple sources, multiple modalities, and multiple formats.

[0016] The present invention has the following beneficial effects: It prioritizes parsing using specific parsing strategies, avoiding unnecessary attempts at common parsing strategies, saving computing resources, and improving parsing efficiency and accuracy. It also provides a backup plan using common parsing strategies when specific parsing strategies fail, supporting multiple parsing strategies, enhancing the method's flexibility and adaptability, and improving the overall parsing success rate.

[0017] By configuring parsing policy groups, this invention simplifies the management and maintenance of parsing policies, enabling rapid and accurate parsing of specific data types. Through a visual interface, it supports real-time monitoring and viewing of logs, streamlining the log access process and improving operational convenience. Ultimately, it achieves the unified collection and aggregation of data from diverse sources, types, and formats, providing a foundation for subsequent processing. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0019] Figure 1 This is a flowchart of a method for normalizing multi-source, multi-modal, and multi-format heterogeneous data in Example 1.

[0020] Figure 2 This is a log access function diagram for a normalization method for multi-source, multi-modal, and multi-format heterogeneous data in Example 1.

[0021] Figure 3 This is a data processing module framework diagram for a method for normalizing heterogeneous data from multiple sources, multiple modalities, and multiple formats in Example 1.

[0022] Figure 4 This is a flow chart of the analytical strategy for a normalization method for multi-source, multi-modal, and multi-format heterogeneous data in Example 1.

[0023] Figure 5 This is a visualization log configuration diagram for a normalization method for multi-source, multi-modal, and multi-format heterogeneous data in Example 1. DETAILED DESCRIPTION

[0024] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.

[0025] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0026] Example 1, reference Figure 1-Figure 5 This is the first embodiment of the present invention, which provides a method for normalizing heterogeneous data from multiple sources, modalities, and formats. The prerequisite for security incident analysis and security response is the ability to uniformly and effectively monitor and manage the multi-source, heterogeneous, complex, and large-scale data in an enterprise IT environment. To this end, the platform has built a log normalization engine that enables real-time normalization and contextualization and intelligence enrichment of various third-party security data. The standardized output is SecurityLogs / Events (security logs), a standardized analysis object that supports multi-dimensional risk monitoring.

[0027] First, based on the daily work experience of security experts, we have summarized over 290 commonly used security log fields, which can be used for subsequent threat classification, detection, analysis, statistics, and report output. The basic security log field types include the following dimensions: log source, data type, basic attributes, account authentication, URL domain name, source and destination IP information, host, file, container, registry, event, traffic log, threat intelligence, correlation event, database event, etc. The fields and some field attribute information under each type dimension are shown in Tables 1 and 2:

[0028] Table 1: Log data field table

[0029]

[0030] Table 2: Field attribute information table

[0031]

[0032]

[0033] Based on the above log field model design, the platform completes the standardized processing of data through three steps: data collection, data processing, and data storage, and supports the use of various back-end scenarios. Figure 1 As shown:

[0034] Step 1: Aggregate data from different sources, types, and formats. Specifically, input sources include, but are not limited to, traffic log data obtained by traffic probes through DPI (deep packet analysis), synchronized logs from third-party network devices and security devices, or event information logs.

[0035] Data access supports multiple methods such as syslog / tcp / udp / kafka / beat, and can be connected to the existing unified log collection platform within the enterprise, greatly reducing the cost of equipment re-purchase and configuration. Data access includes modules such as log collection, preprocessing, normalization, and API services. The log collection module collects logs from various security devices and adds log source identifiers; preprocessing is used to distinguish log sources and extract valid logs; normalization standardizes data, asset discovery, asset enrichment, and IP geographic location enrichment; API services are used for visual configuration. The log access function design is as follows: Figure 2 shown.

[0036] Step 2: Preprocess the data, distinguish the log sources and extract valid logs, introduce parsing strategies for log data, standardize the logs into SOC security event models, normalize the log data, and visualize the log data through log access configuration. Specifically, it includes:

[0037] Data processing is the process of standardizing and normalizing collected data to facilitate storage and analysis on the big data platform. Key operations include cleaning / filtering, standardization, association completion, tagging, and loading standardized data into data storage.

[0038] The product has a large number of out-of-the-box parsing strategies built in, covering most device types and models on the market. For specific device logs to be identified, the platform supports a variety of fast parsing solutions, powerful and easy-to-use parsing configuration capabilities, and supports multiple parsing methods, such as json, cef, delimiter-based, regular expression-based, etc. At the same time, the platform has a built-in script parsing engine to facilitate the expansion of customized parsing needs. The framework diagram of the data processing module is as follows Figure 3 shown.

[0039] Preprocessing has two main functions: distinguishing log sources and extracting valid logs. This is configured through advanced extraction in log access. Distinguishing log sources can be used when logs from different security devices are mixed together. This allows you to separate and label logs from different devices, using Latin1 encoding to resolve encoding inconsistencies across devices. When security device logs have been forwarded with a forwarding header, extracting valid logs removes the forwarding header, allowing the normalized parsing strategy to focus on parsing security device logs and making parsing strategies reusable.

[0040] Data normalization involves the application of parsing strategies and the operation of the normalization engine. The parsing strategies specifically include:

[0041] A parsing policy group contains multiple parsing policies corresponding to different types of logs on the same security device. The group supports import and export. Policies within the same parsing policy group support priority adjustment and quick matching.

[0042] The parsing strategy supports visual configuration and generates Groovy scripts. You can also directly write Groovy scripts and then import them.

[0043] Parsing methods include delimiter, regular expression, JSON, CEF, and LEEF, and support nested parsing of various parsing methods.

[0044] Field mapping in parsing rules supports direct mapping, time conversion, IP conversion, Base64 conversion, and association mapping.

[0045] The regular expression parsing method supports word extraction, and the corresponding field regular expression can be selected from the regular expression library to improve the efficiency of parsing configuration.

[0046] When the normalization engine is running, the normalization loads the parsing strategy corresponding to the groovy script, standardizes the log into the SOC (Security Operations Center) security event model, and includes asset discovery, asset enrichment, IP geolocation enrichment and other functions. Specifically including Figure 4 As shown:

[0047] Locate the parsing strategy group for the log data to be normalized, obtain all quick matching strategies, select strategies from high to low priority for quick matching, observe whether the quick match is hit, and if not, select the next level of parsing strategy according to priority until a quick match is hit or the remaining parsing strategies in the parsing strategy group are zero.

[0048] If a quick match is found, the log data is parsed to determine whether the parsing is successful. If the parsing is successful, the normalization process ends.

[0049] If the parsing is unsuccessful or the remaining parsing strategies in the parsing strategy group are zero, all common parsing strategies are obtained and selected from high to low priority for parsing until the parsing is successful or all common parsing strategies have participated in the parsing. The normalization process ends.

[0050] Log access configuration: Figure 5 As shown, security operations and customers complete the entire log access process through a visual page. The same access method can configure multiple log sources through advanced extraction and support priority adjustment. The most recently accessed logs can be viewed through immediate monitoring, and advanced extraction rules can be verified based on the accessed logs.

[0051] Parsing policy group configuration is decoupled from log access configuration. During log access, you can choose to configure an existing parsing policy group. You can also choose not to set one and then set a parsing policy group for the log source after completing the parsing policy group configuration.

[0052] Step 3: Load the normalized data into the big data platform layer for storage and calculation.

[0053] The big data platform layer includes data storage and data computing. Data storage is used to categorize and store different types of collected data to meet data analysis requirements. Data computing provides general real-time analysis and offline computing capabilities to meet security analysis requirements.

[0054] Based on the varying value of different data, and the need for different storage and analysis strategies, the platform leverages the capabilities of each component to build a storage system based on MySQL, Elasticsearch, ClickHouse, and HDFS. Elasticsearch stores tens of billions of data points, primarily for near-real-time search and analysis, investigations, and forensics, while ClickHouse stores hundreds of billions of data points for compliance, cold data retrieval, and long-term analysis and investigations.

[0055] Example 2 is the second embodiment of the present invention, which is different from the first embodiment in that: a system for a normalization method of multi-source, multi-modal, and multi-format heterogeneous data, including a data acquisition module, a data processing module, and a data storage module; the data acquisition module aggregates data from different sources, different types, and different formats; the data processing module pre-processes the data, distinguishes the log sources and extracts valid logs, introduces parsing strategies for the log data, standardizes the logs into a soc security event model, normalizes the log data, and visualizes the log data through log access configuration; the data storage module loads the normalized data into the big data platform layer for storage and calculation.

[0056] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0057] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device.

[0058] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic devices), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disc read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering, or processing in another suitable manner as necessary, and then stored in a computer memory.

[0059] It should be understood that various parts of the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0060] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. A method for normalizing heterogeneous data from multiple sources, multiple modalities, and multiple formats, characterized by: include, Collect and aggregate data from different sources, types, and formats; Preprocess the data, distinguish log sources and extract valid logs, introduce parsing strategies for log data, standardize logs into SOC security event models, normalize log data, and visualize log data through log access configuration; The normalized data is loaded into the big data platform layer for storage and calculation.

2. The method for normalizing multi-source, multi-modal, and multi-format heterogeneous data according to claim 1, wherein: The different sources include but are not limited to traffic log data obtained by traffic probes through DPI, synchronization logs or event information logs from third-party network devices and security devices; After collecting the data, a log source identifier is added to the data based on the collection source of the data aggregation.

3. The method for normalizing multi-source, multi-modal, and multi-format heterogeneous data according to claim 2, characterized in that: The data preprocessing includes distinguishing log sources and extracting valid logs. Through the advanced extraction configuration of log access, the log sources are distinguished to separate and label the logs of different devices when the log sources of different security devices are mixed together. Latin1 encoding is used to solve the problem of inconsistent encoding of different devices. When the security device log adds a forwarding header through forwarding, the forwarding header is removed by extracting the valid log, so that the normalized parsing strategy focuses on the parsing of the security device log, and the parsing strategy is reusable.

4. The method for normalizing multi-source, multi-modal, and multi-format heterogeneous data according to claim 3, wherein: The parsing policy corresponds to different types of logs of the same security device, supports import and export, and the policies under the same parsing policy group support priority adjustment and fast matching; The parsing strategy group includes several parsing strategies; The parsing strategy supports visual configuration and regeneration of Groovy scripts, or direct writing of Groovy scripts and then importing; Parsing methods include delimiter, regular expression, JSON, CEF and LEEF, and support nested parsing of multiple parsing methods; Field mapping in parsing rules supports direct mapping, time conversion, IP conversion, Base64 conversion and association mapping; The regular expression parsing method supports word extraction and selects the corresponding field regular expression from the regular expression library.

5. The method for normalizing multi-source, multi-modal, and multi-format heterogeneous data according to claim 4, characterized in that: The method of introducing a parsing strategy into the log data and normalizing the log into a SOC security event model includes locating a parsing strategy group for the log data to be normalized, obtaining all quick matching strategies, selecting strategies from high to low priority for quick matching, observing whether a quick match is hit, and if not, selecting a next-level parsing strategy according to priority, until a quick match is hit or the number of remaining parsing strategies in the parsing strategy group is zero; If a quick match is found, the log data is parsed to determine whether the parsing is successful. If the parsing is successful, the normalization process ends. If the parsing is unsuccessful or the remaining parsing strategies in the parsing strategy group are zero, all common parsing strategies are obtained and selected from high to low priority for parsing until the parsing is successful or all common parsing strategies have participated in the parsing. The normalization process ends.

6. The method for normalizing multi-source, multi-modal, and multi-format heterogeneous data according to claim 5, characterized in that: The visualization of log data through log access configuration includes security operations and customers completing the entire log access process through a visualization page, configuring multiple log sources through advanced extraction for the same access method, and supporting priority adjustment, monitoring and viewing of recently accessed logs, and verifying advanced extraction rules based on accessed logs. The parsing policy group configuration is decoupled from the log access configuration. During the log access process, you can choose to configure an existing parsing policy group. Alternatively, you can choose not to set a parsing policy group and then set a parsing policy group for the log source after the parsing policy group is matched.

7. The method for normalizing multi-source, multi-modal, and multi-format heterogeneous data according to claim 6, characterized in that: The normalized data is loaded into the big data platform layer for storage and calculation, including classifying and storing different types of collected data to meet the requirements of data analysis; Data computing provides general real-time analysis and offline computing capabilities to meet security analysis requirements; Based on the different values of different data, different storage and analysis strategies are adopted for different values.

8. A system using the method for normalizing multi-source, multi-modal, and multi-format heterogeneous data according to any one of claims 1 to 7, characterized in that: It includes data acquisition module, data processing module and data storage module; The data collection module collects data aggregation from different sources, different types and different formats; The data processing module pre-processes the data, distinguishes the log sources and extracts valid logs, introduces parsing strategies to the log data, standardizes the logs into the SOC security event model, normalizes the log data, and visualizes the log data through log access configuration; The data storage module loads the normalized data into the big data platform layer for storage and calculation.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the processor implements the steps of a normalization method for multi-source, multi-modal, and multi-format heterogeneous data according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of a normalization method for multi-source, multi-modal, and multi-format heterogeneous data according to any one of claims 1 to 7 are implemented.