A multi-source heterogeneous data processing method and system for business security monitoring

CN122432150BActive Publication Date: 2026-09-22CHINA INT DATA SYST CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202610557880.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-04-24
Publication Date
2026-09-22
Estimated Expiration
2046-04-24

AI Technical Summary

Technical Problem

[0005]鉴于此,本发明提出了一种用于业务安全监测的多源异构数据处理方法及系统,旨在解决在业务安全监测场景下,多源异构业务数据在业务系统升级、接口调整或日志输出方式变化后,容易出现业务语义识别不稳定、业务对象关联连续性不足以及监测链路还原准确性下降的问题

Benefits of technology

[0015]与现有技术相比,本发明的有益效果在于:通过对原始业务数据提取当前结构特征并与历史结构特征进行比对,能够在业务系统升级、接口调整或日志输出方式变化时及时识别字段拆分漂移,从而避免继续沿用失配的旧映射规则导致范式字段识别失准;在判定发生字段拆分漂移后,通过字段切分、业务语义关联和语义恢复,将因位置变化、层级变化或拆分表达而分散的业务信息重新组织为规范化语义值,使同一业务语义在结构变化前后仍可被稳定识别;基于规范化语义值生成稳定替代标识,不仅有利于在不同业务系统和不同数据版本之间保持同一业务对象的持续关联,还能兼顾敏感信息处理需求,降低直接使用原始敏感字段带来的风险;通过将新映射结果与历史映射结果进行并行验证,并在验证满足预设切换条件后再输出目标范式数据,能够减少因直接切换新处理结果而引起的误识别和误关联问题,提高目标范式数据投入业务安全监测时的可靠性和稳定性。由此,本方案能够提升多源异构业务数据在结构变化场景下的适应能力、范式字段识别稳定性、业务链路还原准确性以及异常行为持续关联分析能力,从而为业务安全监测提供连续、准确和可信的数据基础。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122432150B_ABST
    Figure CN122432150B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of data processing, and discloses a multi-source heterogeneous data processing method and system for business security monitoring, which comprises the following steps: collecting original business data to perform feature extraction and obtain current structural features; comparing the current structural features with historical structural features, and performing drift determination in combination with the hit change of a preset paradigm field; when the field splitting drifts, performing field reorganization processing on the original business data to obtain normalized semantic values; generating stable alternative identifiers according to the normalized semantic values, and forming new mapping results and historical mapping results; performing parallel verification on the new mapping results and the historical mapping results to obtain verification results; when the verification results meet preset switching conditions, outputting target paradigm data, and using the target paradigm data for business security monitoring. The application improves the normalization processing stability of multi-source heterogeneous business data, the business object continuous association capability and the business security monitoring accuracy in a structural change scenario.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and more specifically, to a method and system for processing multi-source heterogeneous data for business security monitoring. Background Technology

[0002] With the continuous increase in business systems such as office management, customer management, resource planning, work order processing, and API calls, business security monitoring is no longer limited to logs from a single security device. Instead, it increasingly relies on the unified access, understanding, and analysis of multi-source heterogeneous business data. Only by processing business data from different sources, with different structures and semantic definitions into comparable, correlateable, and sustainably monitorable target paradigm data can subsequent business object identification, behavior chain reconstruction, risk event correlation, and compliance auditing have a reliable data foundation. Therefore, multi-source heterogeneous business data processing has become a key foundational step in the field of business security monitoring.

[0003] In existing technologies, for example, Chinese invention patent CN113157994A discloses a method for processing multi-source heterogeneous platform data. This method, through data collection, log standardization processing, and data storage, performs real-time normalization, classification, filtering, and merging of logs and events from multiple sources to achieve unified management of multi-source heterogeneous data. However, in business security monitoring scenarios, the monitoring object is not merely the general log format itself, but relies more on the continuous identification and correlation of business entities, business objects, operational behaviors, and their preceding and following relationships. When business systems are upgraded, interfaces are adjusted, or log output methods change, the same business semantics may exhibit changes in field position, hierarchy, split expression, or inconsistent expression across systems. While existing technologies focusing on standardization and normalization can still achieve unified access and format organization, they are prone to problems such as decreased stability in normalized field identification, insufficient continuity of cross-system correlation for the same business object, and decreased accuracy in reconstructing monitoring links. This affects the ability of business security monitoring to continuously identify and analyze abnormal behavior.

[0004] Therefore, it is necessary to design a multi-source heterogeneous data processing method and system for business security monitoring to solve the problems existing in the current technology. Summary of the Invention

[0005] In view of this, the present invention proposes a multi-source heterogeneous data processing method and system for business security monitoring, aiming to solve the problems that multi-source heterogeneous business data is prone to cause after business system upgrades, interface adjustments or changes in log output methods in business security monitoring scenarios, such as unstable business semantic recognition, insufficient continuity of business object associations and decreased accuracy of monitoring link reconstruction.

[0006] This invention proposes a method for processing multi-source heterogeneous data for business security monitoring, comprising: Collect raw business data output from multiple business systems, extract structural features from the raw business data, and obtain the current structural features; The current structural features are compared with historical structural features, and drift determination is performed by combining the hit changes of preset paradigm fields to obtain the drift determination result. When the drift determination result is field split drift, the original business data is reorganized to obtain a normalized semantic value. The field reorganization process includes: splitting the original business data into fields, associating the splitting results with business semantics to obtain a target field set, and restoring the semantics of the target field set to obtain the normalized semantic value. A stable substitution identifier is generated based on the normalized semantic value. A new mapping result is formed based on the normalized semantic value and the stable substitution identifier. The original business data is then subjected to historical mapping processing to obtain a historical mapping result. The new mapping result and the historical mapping result are then verified in parallel to obtain a verification result. When the verification result meets the preset switching conditions, target paradigm data is output and used for business security monitoring.

[0007] Furthermore, the current structural features include: field path relationships, field hierarchy relationships, field order relationships, and field value type distribution.

[0008] Furthermore, when performing drift determination, the following steps are included: Calculate the path consistency between the current structural features and historical structural features, and statistically analyze the continuous decrease in the hit rate of the preset normalization field; When the path consistency is lower than a preset threshold and the hit rate decreases continuously to a preset number of times, it is determined to be a field splitting drift. When the path consistency is not lower than a preset threshold or the hit rate decreases continuously without reaching a preset number of times, it is determined that the field splitting has not drifted.

[0009] Furthermore, the drift determination results include: When the drift determination result indicates that the field has split and drifted, the field recombination process is performed. When the drift determination result is that the field split has not drifted, historical mapping processing is performed on the original business data to obtain the historical mapping result, and the target paradigm data is generated based on the historical mapping result.

[0010] Furthermore, when performing field segmentation on the original business data, the process includes: segmenting the original business data into multiple field fragments, and combining them into a candidate field set according to the same parent path relationship, adjacent position relationship, and complementary field value relationship.

[0011] Furthermore, when performing business semantic association on the segmentation results to obtain the target field set, the process includes: matching the candidate field set based on business identifiers, operation actions, time correspondences, and co-occurrence relationships of accompanying fields to determine the target field set.

[0012] Furthermore, the semantic recovery includes: merging and correcting the target field set according to a preset field order and normalization rules to obtain the normalized semantic value.

[0013] Furthermore, when performing parallel verification of the new mapping result and the historical mapping result, the following steps are included: Based on the original business data of the same batch, corresponding paradigm data are generated according to the new mapping result and the historical mapping result respectively; the matching completeness of business identifier, operation action, time correspondence and co-occurrence relationship of accompanying field in the paradigm data are counted respectively to obtain the new mapping link closure result and the historical mapping link closure result. The identification of preset sensitive fields in the new mapping result is statistically analyzed to obtain the omission results; The verification results are obtained based on the link closure results and the missed identification results.

[0014] Furthermore, when obtaining the verification result based on the link closure result and the missed detection result, it includes: When the new mapping link closure result is not lower than the historical mapping link closure result, and the new mapping link closure result meets the preset closure condition, and the missed identification result meets the preset missed identification condition, the verification result is determined to meet the preset switching condition; otherwise, the verification result is determined not to meet the preset switching condition.

[0015] Compared with existing technologies, the advantages of this invention are as follows: By extracting current structural features from the original business data and comparing them with historical structural features, it can promptly identify field splitting drift when business systems are upgraded, interfaces are adjusted, or log output methods change, thereby avoiding inaccurate identification of paradigm fields due to the continued use of mismatched old mapping rules; After determining that field splitting drift has occurred, through field segmentation, business semantic association, and semantic recovery, the business information scattered due to changes in position, hierarchy, or split expression is reorganized into standardized semantic values, so that the same business semantics can still be stably identified before and after structural changes; Based on the standardized semantic values, a stable replacement identifier is generated, which not only helps to maintain the continuous association of the same business object between different business systems and different data versions, but also takes into account the needs of sensitive information processing and reduces the risks brought about by directly using the original sensitive fields; By verifying the new mapping results with the historical mapping results in parallel, and outputting the target paradigm data only after the verification meets the preset switching conditions, it can reduce the problems of misidentification and misassociation caused by directly switching the new processing results, and improve the reliability and stability of the target paradigm data when it is put into business security monitoring. Therefore, this solution can improve the adaptability of multi-source heterogeneous business data in scenarios of structural changes, the stability of paradigm field identification, the accuracy of business link reconstruction, and the ability to continuously correlate and analyze abnormal behaviors, thereby providing a continuous, accurate, and reliable data foundation for business security monitoring.

[0016] On the other hand, this application also provides a multi-source heterogeneous data processing system for business security monitoring, used to apply the above-mentioned multi-source heterogeneous data processing method for business security monitoring, including: The acquisition unit is configured to acquire raw business data output by multiple business systems, extract structural features from the raw business data, and obtain the current structural features. The judgment unit is configured to compare the current structural features with the historical structural features, and combine the hit changes of the preset paradigm field to make a drift judgment and obtain a drift judgment result. The processing unit is configured to perform field reorganization processing on the original business data to obtain normalized semantic values ​​when the drift determination result is field split drift. The field reorganization processing includes: performing field segmentation on the original business data, performing business semantic association on the segmentation results to obtain a target field set, performing semantic recovery on the target field set, and obtaining the normalized semantic values. The verification unit is configured to generate a stable substitution identifier based on the normalized semantic value, form a new mapping result based on the normalized semantic value and the stable substitution identifier, perform historical mapping processing on the original business data to obtain a historical mapping result, and perform parallel verification on the new mapping result and the historical mapping result to obtain a verification result. The output unit is configured to output target paradigm data when the verification result meets the preset switching conditions, and to use the target paradigm data for business security monitoring.

[0017] It is understandable that the above-mentioned methods and systems for processing multi-source heterogeneous data for business security monitoring have the same beneficial effects, and will not be elaborated further here. Attached Figure Description

[0018] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings: Figure 1 A flowchart of a multi-source heterogeneous data processing method for business security monitoring provided in an embodiment of the present invention; Figure 2 This is a functional block diagram of a multi-source heterogeneous data processing system for business security monitoring provided in an embodiment of the present invention. Detailed Implementation

[0019] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the disclosure to those skilled in the art. It should be noted that, unless otherwise specified, embodiments and features in the embodiments of the present invention can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.

[0020] In some embodiments of this application, see Figure 1 As shown, this application proposes a method for processing multi-source heterogeneous data for business security monitoring, including: S100: Collect raw business data output from multiple business systems, extract structural features from the raw business data, and obtain the current structural features.

[0021] S200: Compare the current structural features with historical structural features, and combine the hit changes of the preset paradigm field to determine the drift and obtain the drift determination result.

[0022] S300: When the drift determination result is field split drift, the original business data is reorganized to obtain normalized semantic values. The field reorganization process includes: splitting the original business data into fields, associating the splitting results with business semantics to obtain the target field set, and restoring the semantics of the target field set to obtain normalized semantic values.

[0023] S400: Generate stable substitution identifiers based on normalized semantic values. Form a new mapping result based on the normalized semantic values ​​and stable substitution identifiers. Perform historical mapping processing on the original business data to obtain historical mapping results. Perform parallel verification of the new mapping results and historical mapping results to obtain verification results.

[0024] S500: When the verification results meet the preset switching conditions, output the target paradigm data and use the target paradigm data for business security monitoring.

[0025] In some embodiments, when the drift determination result is that the field splitting did not drift, field reorganization processing and parallel verification of the old and new mappings can be skipped. Instead, historical mapping rules can be directly invoked to perform historical mapping processing on the original business data to obtain the historical mapping result, and target paradigm data can be generated based on the historical mapping result. By processing the non-drifted branch and the drifted branch separately, the accuracy of processing in structural change scenarios can be guaranteed while taking into account the data processing efficiency in structurally stable scenarios.

[0026] Specifically, raw business data refers to unprocessed initial data directly output by multiple business systems, which may have different formats, structures, and semantic definitions.

[0027] Structural feature extraction refers to the process of analyzing raw business data to identify its inherent structural attributes, such as field paths, levels, order, and the distribution of field value types.

[0028] Drift determination refers to the process of judging whether the business data structure or semantics has changed by comparing the current data structure features with the historical data structure features and combining the changes in the hit status of preset normalized fields.

[0029] Preset paradigm fields refer to target fields pre-defined for business security monitoring. They are used to represent at least one of the following in the target paradigm data: business entity, business object, operation action, time information, or sensitive attribute. Examples include user identifier, order identifier, operation type, timestamp, contact information, or account identifier. Preset paradigm fields are preferably determined jointly based on business security monitoring rules, historical stable mapping results, and manual annotation results.

[0030] Field splitting drift refers to a situation in data drift where a field that is semantically complete in business context is split into multiple independent fields in the original business data.

[0031] Field reorganization refers to a series of operations that recombine the split fields in the original business data to restore their original business semantics in response to the phenomenon of field splitting and drifting.

[0032] Normalized semantic values ​​refer to data values ​​with unified semantics and normalized format obtained after field reorganization, which can accurately express business meaning.

[0033] A stable alternative identifier is an alternative identifier generated based on normalized semantic values. This alternative identifier is used to maintain a consistent association for the same business object within a preset monitoring domain and reduce the risk of directly exposing the original sensitive fields.

[0034] In some embodiments, the preset monitoring domain may be determined by at least one of the tenant identifier, business system identifier, and business scenario identifier, which is used to limit the association range of stable alternative identifiers, so that the same business objects maintain stable association within the same monitoring domain and are isolated from other monitoring domains.

[0035] Parallel verification refers to running two different data processing logics simultaneously and comparing and evaluating the paradigm data generated by each to verify the effectiveness and reliability of the new processing logic.

[0036] Target paradigm data refers to the data format that, after being processed by this method, conforms to preset specifications, has a unified structure and semantics, and can be directly used for business security monitoring.

[0037] This embodiment provides a method for processing multi-source heterogeneous data for business security monitoring. The method first collects raw business data output from multiple business systems and extracts field path relationships, field hierarchical relationships, field order relationships, and field value type distributions according to preset parsing rules to form the current structural features. Historical structural features are preferably generated by statistically analyzing multiple batches of historical data from the same business system during its stable operation phase, and serve as the baseline structural features before the current version switch of the business system. Subsequently, the current structural features are compared with the historical structural features, and drift determination is performed based on changes in the hit rate of preset normalized fields. Drift determination is not based on a single field name change, but rather examines both changes in structural position and changes in the ability to identify key normalized fields to identify field splitting drift.

[0038] When the drift determination result is field split drift, field reorganization processing is performed on the original business data. Field reorganization processing includes: first, dividing the original business data into field fragments; then, forming a target field set based on the structural and business relationships between the field fragments; subsequently, merging and correcting the target field set according to a preset field order and normalization rules to obtain normalized semantic values. Stable replacement identifiers are generated based on the monitoring domain identifier, field category identifier, and normalized semantic values, ensuring that the normalized semantic values ​​of the same business object within the same monitoring domain and field category generate the same stable replacement identifier. Subsequently, a new mapping result is formed based on the normalized semantic values ​​and stable replacement identifiers, and historical mapping processing is performed on the same batch of original business data to obtain historical mapping results. When verifying the new mapping results and historical mapping results in parallel, corresponding normalization data are generated based on the same batch of original business data, and the new mapping link closure result, historical mapping link closure result, and missed detection result are calculated. When the closure result of the new mapping link is not lower than the closure result of the historical mapping link, and the preset closure condition is met, and the omission result meets the preset omission condition, the verification result is determined to meet the preset switching condition, and then the target paradigm data is output for business security monitoring.

[0039] This method effectively addresses field splitting and drift issues caused by business system upgrades or interface adjustments by extracting structural features, determining drift, reorganizing fields, and verifying the mapping results between old and new data from multiple heterogeneous sources. As a result, this method ensures the stability of paradigm field identification, the continuity of cross-system associations of the same business object, and the accuracy of monitoring link reconstruction, thereby improving the ability of business security monitoring to continuously identify and analyze abnormal behavior.

[0040] This application further proposes that the current structural features include: field path relationships, field hierarchy relationships, field order relationships, and field value type distribution.

[0041] In this context, field path relationships refer to the unique identifying path from the root node to a specific field in hierarchical data (such as hierarchical data formats or nested key-value pairs). For example, in a hierarchical data structure representing user information, where the user identifier field is located under the user information field, its path relationship can be represented as "user identifier field under user information". Extracting field path relationships helps identify changes in the position of a field within the data structure, such as when a field is moved to a different parent node or its nesting depth changes. This is typically achieved by recursively traversing the data structure and recording the complete access path for each field during the traversal.

[0042] Field hierarchy refers to the nesting depth and parent-child relationships between fields in a data structure. For example, in a hierarchical structure of user information data, the user information field could be at the first level, and the user identifier field could be at the second level below the user information field. Extracting field hierarchy can reflect changes in the complexity of the data structure, such as a field changing from a flat structure to a nested structure, or its nesting level increasing or decreasing. This can be achieved by maintaining hierarchy count information during data structure parsing to record the hierarchy depth and parent-child relationships of each field.

[0043] Field order relationships refer to the arrangement of fields at the same level in a data structure. In certain data formats (such as CSV, fixed-length records) or specific business scenarios, the order of fields may have important business implications. Extracting field order relationships helps identify situations where the original order has been disrupted due to field swapping, adding, or deleting fields. This is typically achieved by parsing the data structure, obtaining an ordered list of sibling fields, and recording their index positions within the list.

[0044] Field value type distribution refers to the data type (e.g., string, integer, floating-point, boolean, date, etc.) of a specific field across a large dataset, and the frequency or proportion of these values. For example, a field initially expected to be an integer might suddenly show a large number of string values ​​in a new batch of data. Extracting field value type distribution can reflect semantic changes in the field's data content itself; even if the field's path and hierarchy remain unchanged, its data type may still shift. This can be achieved by sampling and inferring the type of field values, and then statistically analyzing the frequency of each data type.

[0045] By refining the current structural features into field path relationships, field hierarchy relationships, field order relationships, and field value type distribution, this application can comprehensively and meticulously characterize the structural features of original business data from multiple dimensions. This multi-dimensional feature extraction method enhances the ability to perceive subtle changes in data structure. For example, even if the field name remains unchanged, changes in its nesting position in the data, its relative order with other fields, or the data type it carries can accurately capture these shifts. This makes the shift determination process more accurate, avoiding missed or false judgments due to insufficient feature dimensions, thereby improving the adaptability and robustness of the business security monitoring system to the evolution of multi-source heterogeneous data structures and ensuring the effectiveness of subsequent field reorganization and mapping processing.

[0046] This application further proposes to more accurately identify field splitting drift by calculating the path consistency between the current structural features and historical structural features and statistically analyzing the continuous decline in the hit rate of the preset normalization field when making drift determination.

[0047] In some embodiments, historical structural features are statistically obtained from the same business system over the most recent ten to thirty consecutive stable data batches. Path consistency can be obtained by comparing the valid field paths in the current structural features with the corresponding field paths in the historical structural features item by item. Continuous decreases in hit rate can be calculated within three to eight consecutive statistical batches. The preset threshold and preset number of times condition are preferably determined based on the historical stable samples of the business system, replay data before and after version switching, and manual annotation results. Preferably, the path consistency threshold can be set between 60% and 80%, the continuous decrease in hit rate condition can be set to three to five consecutive statistical batches, and the hit rate decrease threshold can be set to a decrease of more than 30% relative to the historical stable average.

[0048] Specifically, the path consistency level can be determined by the proportion of field paths in the current batch that maintain consistency with paths corresponding to historical structural features to the total number of paths in the comparison. Preferably, the valid field paths in the comparison are those that can be parsed in both the current batch and the historical benchmark, and whose meanings are aligned. A continuous decline in the hit rate can be determined by the change in the recognition hit rate of the preset paradigm field in consecutive statistical batches, preferably represented by the decline relative to the historical stable average and the number of consecutive declining batches. When the path consistency level is below a preset threshold, and the hit rate of the preset paradigm field declines by more than 30% relative to the historical stable average within three to five consecutive statistical batches, it can be determined as field splitting drift. When neither of the above conditions is met simultaneously, it is determined as field splitting not drifting. By combining the judgment of changes in structural path and changes in the recognition capability of the paradigm field, the risk of misjudgment caused by fluctuations in a single indicator can be reduced.

[0049] This application further proposes how to efficiently and accurately select subsequent data processing paths based on different judgment results after drift determination, in order to ensure the continuity of business security monitoring and data quality. Unclear processing logic or untimely response may lead to data processing errors or inefficiency.

[0050] Specifically, when the drift determination result is that the field splitting has drifted, field reorganization is performed. When the drift determination result is that the field splitting has not drifted, historical mapping is performed. The original business data is then mapped historically to obtain the historical mapping result, and the target normalized data is generated based on the historical mapping result.

[0051] When the drift determination result is field split drift, a dedicated field reorganization process will be immediately initiated. Field split drift typically refers to a field in the original data being split into multiple subfields, or its internal structure changing, causing the original mapping rules to become invalid. Field reorganization processing involves a series of operations, including field segmentation, business semantic association, and semantic recovery, to reintegrate or normalize the split fields, making them conform to the expected semantic structure, thereby obtaining normalized semantic values. For example, if an "address" field in the original data is split into multiple fields such as "province," "city," and "street," the field reorganization process will identify these split fields and recombine them into a complete and normalized address information according to business semantics.

[0052] On the other hand, when the drift determination result indicates that the field splitting has not drifted, the existing and stable historical mapping processing flow will continue to be used. Historical mapping processing refers to converting the original business data into target normalized data using pre-established and verified mapping rules. Historical mapping processing includes field mapping, data type conversion, and data cleaning. In this way, unnecessary complex field reorganization can be avoided when the data structure has not changed, thereby improving the efficiency of data processing and resource utilization. For example, if the structure of the original business data is highly consistent with the characteristics of the historical structure, the existing mapping rules can be directly applied to map the "user ID" in the original data to the "user identifier" in the target normalized data.

[0053] Through the above technical solution, after obtaining the drift determination result, this application can intelligently select different data processing paths according to the type of determination result. When field split drift is detected, field reorganization processing can be initiated in a timely manner to perform fine-grained segmentation, association, and semantic recovery of the original business data, effectively addressing the challenges brought about by changes in data structure and ensuring that the data can be correctly standardized, thereby providing high-quality input for subsequent business security monitoring. Conversely, when no field split drift is detected, historical mapping processing can be efficiently used, avoiding unnecessary complex calculations, thus improving data processing efficiency while ensuring data processing accuracy. This adaptive processing mechanism enables multi-source heterogeneous data processing methods to respond more flexibly and robustly to changes in the data structure of business systems, improving the accuracy and real-time performance of business security monitoring.

[0054] This application further proposes a method for field segmentation of raw business data, which includes segmenting the raw business data into multiple field fragments and combining them into a candidate field set according to the same parent path relationship, adjacent position relationship and complementary field value relationship.

[0055] Specifically, when segmenting raw business data into multiple field fragments, for raw business data with a hierarchical structure, it is preferable to use leaf node fields and their corresponding field paths as field fragments. For raw business data in key-value pairs, table records, or free text format, it is preferable to segment based on delimiters, fixed field boundaries, field labels, and value patterns. Field fragments preferably contain at least two of the following: field name, field path, field value, field position, and parent node information.

[0056] Based on this, the segmented field fragments are combined according to the following relationships: shared parent path, adjacent position, and complementary field values, to form a candidate field set. The shared parent path relationship reflects that the field fragments originate from the same parent structure; the adjacent position relationship reflects the adjacent arrangement of the field fragments in the original data; and the complementary field value relationship reflects that the combination of multiple field fragments can jointly express a complete business semantic. Preferably, multiple field fragments are combined into the same candidate field set only when they simultaneously satisfy at least two of these relationship conditions.

[0057] In some embodiments, a stable alternative identifier can be determined jointly by a monitoring domain identifier, a field category identifier, and a normalized semantic value. Before generating a stable alternative identifier, it is preferable to first perform a standardization process on the normalized semantic value. This standardization process includes removing invalid spaces, standardizing character case, standardizing time and number formats, and removing separators that do not affect business semantics. Then, the monitoring domain identifier, field category identifier, and the normalized semantic value are combined in a preset order, and an irreversible transformation is performed based on fixed generation parameters within the same monitoring domain to obtain a stable alternative identifier. For normalized semantic values ​​within the same monitoring domain and corresponding to the same business object, the same stable alternative identifier is generated. For normalized semantic values ​​in different monitoring domains, different stable alternative identifiers are generated by using different monitoring domain identifiers or different generation parameters within the domain. Through this method, the continuous association capability of the same business object within the target monitoring domain can be maintained without directly exposing the original sensitive field values.

[0058] The above technical solution not only decomposes the original business data into multiple field fragments during field segmentation, but also intelligently combines these fragments into a candidate field set by considering relationships such as shared parent paths, adjacent positions, and complementary field values. This preliminary combination based on multi-dimensional relationships improves the accuracy and efficiency of field reorganization processing. It ensures that even in the event of field splitting drift, fragmented information of the original business data can be identified and pre-aggregated, providing high-quality input for subsequent business semantic association and semantic recovery. This solves the difficulty of data reorganization caused by field splitting drift, guarantees the accurate generation of standardized semantic values, and thus improves the reliability of business security monitoring.

[0059] This application further proposes to perform business semantic association on the segmentation results to obtain a target field set. Specifically, this process includes matching the candidate field set based on business identifiers, operation actions, time correspondences, and co-occurrence relationships of accompanying fields, thereby determining the target field set.

[0060] The business semantic association of the segmentation results refers to further utilizing business domain knowledge and rules to identify and establish actual business meaning relationships between field fragments after the initial segmentation and combination to form a candidate field set. This step aims to elevate field fragments that are only structurally related to information units with completeness and consistency in business logic. Through business semantic association, it can be ensured that the field set processed subsequently is not only complete in form but also accurately reflects the true situation of business events in content.

[0061] Obtaining the target field set refers to the set of fields with clear business meaning and completeness that are selected, combined, and finally confirmed from the candidate field set after business semantic association processing. This set represents the core information in the business data that truly needs to be extracted, standardized, and used for security monitoring, laying a solid foundation for subsequent semantic recovery and the generation of standardized semantic values.

[0062] Matching candidate fields based on business identifiers refers to using fields or combinations of fields in business data that can uniquely identify a business event, business object, or business entity as the basis. Examples include order numbers, user IDs, and transaction serial numbers. When performing business semantic association, identifying and matching these business identifiers allows disparate field fragments to be grouped under the same business event or object, ensuring data integrity and consistency. This can be achieved through methods such as matching based on preset format rules, comparing identifier dictionaries, and performing association verification based on business primary keys.

[0063] Matching candidate fields based on action actions refers to identifying the specific behavior or operation that occurs within a business event, such as "login," "purchase," "transfer," or "delete." By identifying these actions, the nature and purpose of the business event can be understood. In semantic association, action fields are often closely related to fields such as business identifiers and time, collectively forming a complete description of the business event. This can be achieved through methods such as action keyword matching, comparison with predefined operation type tables, and validation of operation field context rules.

[0064] Matching candidate field sets based on time correspondence refers to analyzing the chronological order, timestamp associations, or co-occurrence relationships within time windows of different field fragments in business data. For example, a business event typically includes occurrence time, creation time, and modification time. By analyzing this time information, it's possible to determine whether different field fragments belong to the same business event or their logical order within the business process. This can be achieved through timestamp parsing and comparison, preset time window validation, and chronological order rule validation within the same business process.

[0065] Matching candidate field sets based on co-occurrence relationships refers to leveraging the phenomenon that certain field fragments tend to appear simultaneously in large amounts of business data, and that they collectively describe a specific business attribute or entity. For example, in user registration information, "username" usually appears alongside "email address" and "phone number." This co-occurrence relationship reflects the inherent semantic connection between fields. Implementation methods can include comparing co-occurrence frequencies based on historical samples, validating field association rule tables, and judging using a pre-defined business rule base.

[0066] Matching the candidate field set involves applying the semantic rules, such as business identifiers, operational actions, time correspondences, and co-occurrence relationships of accompanying fields, to each field fragment in the candidate field set to identify and confirm their business relevance. This process may involve multiple iterations and complex logical judgments, aiming to accurately identify truly business-related field combinations from structurally related fragments.

[0067] Determining the target field set refers to selecting the combination of fields that meet the preset matching conditions from the candidate field set and using them as input for subsequent semantic recovery and normalization processing.

[0068] In some embodiments, when performing business semantic association on a candidate field set, the matching results for business identifiers, operation actions, time correspondences, and co-occurrences of accompanying fields can be calculated separately. When a candidate field set satisfies both the business identifier and time correspondence matching results, or any three of the business identifier, operation action, and co-occurrences of accompanying fields matching results, the candidate field set is determined to be the target field set. In cases where multiple candidate field sets can satisfy the matching conditions, the candidate field set that is associated with the corresponding paradigm field most frequently in historically stable structural features is preferably selected.

[0069] By introducing business semantic association rules such as business identifiers, operational actions, time correspondences, and co-occurrence relationships of accompanying fields, this application can accurately identify and determine the target field set with actual business significance from the initially segmented and combined candidate field set. This overcomes the limitations of relying solely on structural features for field combination, improving the accuracy of field reorganization processing and the effectiveness of semantic recovery. Ultimately, the obtained normalized semantic values ​​can more realistically reflect the business meaning of the original business data, thereby providing more accurate and reliable target paradigm data for business security monitoring, improving the accuracy and efficiency of security monitoring, and reducing the risk of false positives and false negatives.

[0070] This application further proposes a semantic recovery step, which includes: merging and correcting the target field set according to a preset field order and normalization rules to obtain normalized semantic values.

[0071] The preset field order refers to a standardized arrangement of fields defined for the final normalized semantic value during semantic recovery. For example, for a log event, the preset field order can be "event time, user ID, operation type, source IP address". This order can be pre-configured in the system's data dictionary, metadata management module, or defined through the business rule engine to ensure structural consistency of data processed from different sources or at different times. Normalization rules are a set of rules used to unify data formats, data types, data value domains, or perform data transformations. For example, all timestamp formats can be unified to a unified time format standard, the case of user IDs can be standardized, or different strings representing the same business meaning (such as "login", "user login", "user login") can be mapped to a unified "login" value. These rules can be implemented based on regular expressions, lookup tables, data type conversion functions, or custom scripts and stored in a rule base for the semantic recovery module to call. Merging refers to combining multiple field fragments that may exist in the target field set but logically belong to the same semantic unit into a complete field. For example, if the date and time in the original data are split into two independent fields, the merge operation will combine them into a complete date and time field. Or, if a complex field is split into multiple subfields, the merge operation will recombine them according to a preset structure definition. The merge operation usually refers to the preset field order to determine the combination method and the final structure. Correction refers to performing data cleaning, format correction, or value range validation on the merged fields to ensure that the data conforms to normalization rules and business requirements. For example, the format of the merged date and time field is validated to ensure that it conforms to the preset time format. Range checks are performed on numeric fields to correct outliers. Or, text fields are processed to remove extra spaces and unify capitalization. Correction aims to eliminate noise and inconsistencies in the data and improve data quality. Finally, a normalized semantic value is obtained, which is a data representation with a clear structure, uniform format, and accurate semantics after the preset field order is arranged, normalization rules are validated, fields are merged, and data is corrected.

[0072] The above technical solution enables the merging and correction of target field sets after obtaining them, based on preset field order and normalization rules. This resolves potential issues such as inconsistent structure, non-standard format, or incomplete semantics within the target field set, ensuring that key business information extracted from multi-source heterogeneous data can be represented in a standardized and consistent manner. Through the merging operation, previously scattered semantic fragments are integrated into complete business fields. The correction operation improves data quality, eliminating format differences and potential errors. The resulting normalized semantic values ​​possess high accuracy and consistency, significantly enhancing the reliability of subsequent stable replacement identifier generation and providing high-quality, reliable input data for business security monitoring, thereby improving the accuracy and efficiency of monitoring.

[0073] This application further proposes a specific method for parallel verification of new and historical mapping results, including: generating corresponding paradigm data based on the same batch of original business data, respectively, according to the new and historical mapping results. The matching completeness of business identifiers, operational actions, time correspondences, and co-occurrence relationships of accompanying fields in the paradigm data is statistically analyzed to obtain the new mapping link closure result and the historical mapping link closure result. In some embodiments, preset sensitive fields may include mobile phone numbers, ID card numbers, email addresses, account identifiers, device identifiers, shipping addresses, and other fields that can directly or indirectly identify business entities, business objects, or key business attributes. Fields already marked as preset sensitive fields are preferably determined based on manually labeled samples, sensitive field rule bases, or field labels in historical stable mapping results. The identification loss of preset sensitive fields in the new mapping results is statistically analyzed to obtain the omission results. Verification results are obtained based on the link closure results and the omission results.

[0074] Specifically, during parallel verification, it is preferable to select three to five consecutive verification batches of original business data under the current version of the same business system, and generate corresponding paradigm data based on the new mapping results and historical mapping results respectively. For each verification batch, the proportion of the number of records that simultaneously meet the conditions of business identifier location, operation action identification, time correspondence verification, and co-occurrence relationship of accompanying fields meeting preset conditions is counted, and this proportion is used as the link closure result of the corresponding batch.

[0075] Simultaneously, the identification failure rate of preset sensitive fields in the new mapping results is statistically analyzed. The omission result is preferably represented by the proportion of the number of fields marked as preset sensitive fields in the original business data but not identified in the new mapping results to the total number of preset sensitive fields in the verification batch. By statistically analyzing the link closure results and omission results for multiple consecutive verification batches, the reliability of the new mapping results in actual operation can be more stably evaluated, providing a basis for subsequent switchover decisions.

[0076] Through the above technical solutions, this application can comprehensively and objectively evaluate the quality and reliability of the new mapping results. By comparing the link closure results of the old and new mappings, it can be ensured that the new mapping is at least as good as, or even better than, the historical mapping in terms of extracting key business semantic information. Simultaneously, by specifically statistically analyzing the omission of sensitive fields, the omission of important security information due to data drift processing is avoided, thereby ensuring the accuracy and completeness of business security monitoring. This multi-dimensional, parallel verification method provides solid data support for the subsequent safe and stable switch to the new mapping results, reducing the business security risks introduced by improper mapping switching.

[0077] This application further proposes that when obtaining verification results based on link closure results and omission results, if the new mapped link closure result is not lower than the historical mapped link closure result, and the new mapped link closure result meets the preset closure condition, and the omission result meets the preset omission condition, then the verification result is determined to meet the preset switching condition. Otherwise, the verification result is determined not to meet the preset switching condition.

[0078] When obtaining verification results based on link closure results and missed detection results, the following rules are preferably followed: If the new mapping link closure results in three consecutive verification batches are all no lower than the historical mapping link closure results of the corresponding batch, and the average new mapping link closure result is no less than 95%, and the average missed detection result is no higher than 1%, then the verification result is determined to meet the preset switching conditions. Otherwise, the verification result is determined not to meet the preset switching conditions. In some embodiments, when the verification result does not meet the preset switching conditions, the historical mapping results are used to generate the target paradigm data for the current output, and the new mapping results are retained as mapping results to be reviewed. Parallel verification is then performed again after subsequent verification batches have accumulated to a preset number. By setting a rollback process after verification failure, the continuity of business security monitoring can be avoided by directly switching before the new mapping results have reached a stable requirement. The preset closure conditions and preset missed detection conditions are preferably determined based on historical stable operation data, manually annotated field samples, and trial operation playback results.

[0079] It should be noted that the values ​​for path consistency, hit rate decrease, number of consecutive statistical batches, link closure results, and missed identification results mentioned above are only example values. In actual applications, they can be adjusted according to the type of business system, the size of historical samples, and the results of trial operation.

[0080] Through the above technical solution, this application provides a multi-dimensional and highly reliable switching decision mechanism. This mechanism comprehensively considers the relative performance of the new mapping scheme compared to historical mapping schemes, the absolute quality of the new mapping scheme itself, and its performance on key security indicators. This comprehensive evaluation method avoids the risk of misjudgment caused by a single indicator, ensuring that switching only occurs when the new mapping scheme meets stringent requirements in terms of overall data quality and identification of security-sensitive information. This improves the stability and accuracy of business security monitoring data processing, reduces business risks caused by improper data drift handling, and thus guarantees the continuous effectiveness of business security monitoring.

[0081] The following example will provide a more detailed explanation of the above technical solution: A large enterprise has deployed an order management system, a customer relationship management system, and a work order workflow system. The raw business data output by each system is connected to a business security monitoring platform to identify risk events such as abnormal access, abnormal operations, and abnormal associations of sensitive business objects. Before the system upgrade, the order management system output customer phone information as a contact phone number field under customer information. The business security monitoring platform uniformly maps this field to the target normal form field "customer phone number" using existing historical mapping rules.

[0082] With the upgrade of the order management system, the original contact phone number field under customer information is no longer directly output. Instead, it has been split into country code, area code, phone number, and extension number fields under customer contact information. Because the single business semantics have been split into multiple field fragments, the original historical mapping rules can no longer reliably extract the target paradigm field "customer phone number" directly from the upgraded original business data.

[0083] This method first collects the upgraded original business data and extracts the current structural features, including the field path relationships, field hierarchy relationships, field order relationships, and field value type distributions corresponding to the aforementioned multiple field fragments. Then, it compares the current structural features with the historical structural features before the upgrade and analyzes the changes in the hit rate of the preset normalized field "customer phone number". Since the original contact phone number field path disappeared, and the hit rate of the preset normalized field corresponding to the customer phone number continuously decreased relative to the historical stable average over three to five consecutive statistical batches, reaching the preset decrease condition, it was determined that field splitting drift had occurred.

[0084] After determining that field splitting drift has occurred, the original business data undergoes field reorganization. First, the original business data is divided into multiple field fragments, obtaining country code, area code, phone number, and extension number fragments, while retaining the field path, field value, field position, and parent node information for each fragment. Then, based on the same parent path relationship, adjacent position relationship, and complementary field value relationship, these field fragments are combined into a candidate field set.

[0085] Next, the candidate field set is semantically associated with business information. Specifically, based on the business identifier of the order number, the action of placing an order or modifying contact information, the time correspondence in the business records, and the co-occurrence relationship between the customer name and contact information fields, the candidate field set is matched to determine that the country code field fragment, area code field fragment, number field fragment, and extension field fragment belong to the same business object's semantic set of contact information, and this set is determined as the target field set.

[0086] Subsequently, semantic recovery is performed on the target field set. Following a preset field order and normalization rules, country codes, area codes, phone numbers, and extension numbers are merged and corrected, invalid separators are removed, number formats are completed, and a standardized contact information representation is formed, thus obtaining the normalized semantic value corresponding to "customer phone number". Next, a stable replacement identifier is generated based on the monitoring domain identifier, field category identifier, and normalized semantic value, and a new mapping result is formed based on this stable replacement identifier. Simultaneously, historical mapping processing is performed on the original business data of the same batch to obtain historical mapping results. Then, corresponding normalized data are generated based on the new mapping result and the historical mapping result, respectively. The link closure result between the two is statistically analyzed, and the identification of the preset sensitive field "customer phone number" by the new mapping result is statistically analyzed to obtain the omission results.

[0087] When the closure results of the new mapping links in three consecutive verification batches are all no lower than the closure results of the historical mapping links in the corresponding batches, and the average closure rate of the new mapping links is no less than 95%, and the average missed identification rate is no higher than 1%, the verification results are deemed to meet the preset switching conditions. Target paradigm data is then output and used for business security monitoring. This example demonstrates that even after a single business semantic in the original business data is split into multiple fields, this method can still achieve continuous identification and stable association of key business objects through field segmentation, business semantic association, semantic recovery, stable replacement identifier generation, and parallel verification of new and old mappings. If the drift determination result is that the field splitting has not drifted, historical mapping processing is directly performed on the original business data to obtain historical mapping results. Target paradigm data is then generated based on these historical mapping results, without the need for field reorganization and new mapping verification. This ensures that existing mapping rules can be efficiently used in the absence of drift.

[0088] Based on another preferred embodiment described above, see [link to preferred embodiment]. Figure 2 As shown, this embodiment provides a multi-source heterogeneous data processing system for business security monitoring, which applies the above-described multi-source heterogeneous data processing method for business security monitoring, including: The acquisition unit is configured to acquire raw business data output from multiple business systems, extract structural features from the raw business data, and obtain the current structural features.

[0089] The judgment unit is configured to compare the current structural features with the historical structural features and combine the hit changes of the preset paradigm field to make a drift judgment and obtain the drift judgment result.

[0090] The processing unit is configured to perform field reorganization processing on the original business data to obtain normalized semantic values ​​when the drift determination result is field split drift. The field reorganization processing includes: splitting the original business data into fields, associating the splitting results with business semantics to obtain a target field set, and restoring the semantics of the target field set to obtain normalized semantic values.

[0091] The verification unit is configured to generate a stable substitution identifier based on the normalized semantic value, form a new mapping result based on the normalized semantic value and the stable substitution identifier, perform historical mapping processing on the original business data to obtain the historical mapping result, and perform parallel verification of the new mapping result and the historical mapping result to obtain the verification result.

[0092] The output unit is configured to output target paradigm data when the verification result meets the preset switching conditions, and use the target paradigm data for business security monitoring.

[0093] It is understandable that the above-mentioned methods and systems for processing multi-source heterogeneous data for business security monitoring have the same beneficial effects, and will not be elaborated further here.

[0094] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the protection scope of the present invention.

Claims

1. A method for processing multi-source heterogeneous data for business security monitoring, characterized in that, include: Collect raw business data output from multiple business systems, extract structural features from the raw business data, and obtain the current structural features; The current structural features are compared with historical structural features, and drift determination is performed by combining the hit changes of preset paradigm fields to obtain the drift determination result. When the drift determination result is field split drift, the original business data is reorganized to obtain a normalized semantic value. The field reorganization process includes: splitting the original business data into fields, associating the splitting results with business semantics to obtain a target field set, and restoring the semantics of the target field set to obtain the normalized semantic value. A stable substitution identifier is generated based on the normalized semantic value. A new mapping result is formed based on the normalized semantic value and the stable substitution identifier. The original business data is then subjected to historical mapping processing to obtain a historical mapping result. The new mapping result and the historical mapping result are then verified in parallel to obtain a verification result. When the verification result meets the preset switching conditions, target paradigm data is output and used for business security monitoring.

2. The multi-source heterogeneous data processing method for business security monitoring according to claim 1, characterized in that, The current structural features include: field path relationships, field hierarchy relationships, field order relationships, and field value type distribution.

3. The multi-source heterogeneous data processing method for business security monitoring according to claim 1, characterized in that, When performing drift detection, the following are included: Calculate the path consistency between the current structural features and historical structural features, and statistically analyze the continuous decrease in the hit rate of the preset normalization field; When the path consistency is lower than a preset threshold and the hit rate decreases continuously to a preset number of times, it is determined to be a field splitting drift. When the path consistency is not lower than a preset threshold or the hit rate decreases continuously without reaching a preset number of times, it is determined that the field splitting has not drifted.

4. The multi-source heterogeneous data processing method for business security monitoring according to claim 3, characterized in that, After the drift determination result, it includes: When the drift determination result indicates that the field has split and drifted, the field recombination process is performed. When the drift determination result is that the field split has not drifted, historical mapping processing is performed on the original business data to obtain the historical mapping result, and the target paradigm data is generated based on the historical mapping result.

5. The multi-source heterogeneous data processing method for business security monitoring according to claim 4, characterized in that, When performing field segmentation on the original business data, the process includes: segmenting the original business data into multiple field fragments, and combining them into a candidate field set according to the same parent path relationship, adjacent position relationship, and complementary field value relationship.

6. The multi-source heterogeneous data processing method for business security monitoring according to claim 5, characterized in that, When performing business semantic association on the segmentation results to obtain the target field set, the process includes: matching the candidate field set based on business identifiers, operation actions, time correspondences, and co-occurrence relationships of accompanying fields to determine the target field set.

7. The multi-source heterogeneous data processing method for business security monitoring according to claim 1, characterized in that, The semantic recovery includes: merging and correcting the target field set according to a preset field order and normalization rules to obtain the normalized semantic value.

8. The multi-source heterogeneous data processing method for business security monitoring according to claim 1, characterized in that, When performing parallel verification of the new mapping result and the historical mapping result, the following is included: Based on the original business data of the same batch, corresponding paradigm data are generated according to the new mapping result and the historical mapping result respectively; the matching completeness of business identifier, operation action, time correspondence and co-occurrence relationship of accompanying field in the paradigm data are counted respectively to obtain the new mapping link closure result and the historical mapping link closure result. The identification of preset sensitive fields in the new mapping result is statistically analyzed to obtain the omission results; The verification results are obtained based on the link closure results and the missed identification results.

9. The multi-source heterogeneous data processing method for business security monitoring according to claim 8, characterized in that, When obtaining the verification result based on the link closure result and the missed detection result, it includes: When the new mapping link closure result is not lower than the historical mapping link closure result, and the new mapping link closure result meets the preset closure condition, and the missed identification result meets the preset missed identification condition, the verification result is determined to meet the preset switching condition; otherwise, the verification result is determined not to meet the preset switching condition.

10. A multi-source heterogeneous data processing system for business security monitoring, used to apply the multi-source heterogeneous data processing method for business security monitoring as described in any one of claims 1-9, characterized in that, include: The acquisition unit is configured to acquire raw business data output by multiple business systems, extract structural features from the raw business data, and obtain the current structural features. The judgment unit is configured to compare the current structural features with the historical structural features, and combine the hit changes of the preset paradigm field to make a drift judgment and obtain a drift judgment result. The processing unit is configured to perform field reorganization processing on the original business data to obtain normalized semantic values ​​when the drift determination result is field split drift. The field reorganization processing includes: performing field segmentation on the original business data, performing business semantic association on the segmentation results to obtain a target field set, performing semantic recovery on the target field set, and obtaining the normalized semantic values. The verification unit is configured to generate a stable substitution identifier based on the normalized semantic value, form a new mapping result based on the normalized semantic value and the stable substitution identifier, perform historical mapping processing on the original business data to obtain a historical mapping result, and perform parallel verification on the new mapping result and the historical mapping result to obtain a verification result. The output unit is configured to output target paradigm data when the verification result meets the preset switching conditions, and to use the target paradigm data for business security monitoring.

Citation Information

Patent Citations

  • Multi-source heterogeneous platform data processing method

    CN113157994A

  • Management method and system based on multi-source heterogeneous data

    CN121030053A

  • Intelligent and automatic test case generation method based on large language model

    CN121255655A