A data governance method and system for communication big data

By aggregating data from all data sources and integrating data from multiple network domains, and by building a two-dimensional thematic library, the problem of data silos in mobile communication big data has been solved, achieving unified data aggregation and multi-service applications, and improving data utilization efficiency and the richness of application scenarios.

CN115221114BActive Publication Date: 2026-04-24HENAN XINDA WANGYU TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HENAN XINDA WANGYU TECH CO LTD
Filing Date
2022-03-31
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

In the current governance of mobile communication big data, data is scattered across multiple systems, lacking a unified data governance method. This results in a severe data silo effect, poor data application effectiveness, inconsistent data standards, high costs, narrow application areas, and the failure to fully unleash the potential of data.

Method used

By adopting the methods of full-domain data source aggregation, multi-domain data fusion, two-dimensional theme library construction, and business application library construction, data is uniformly aggregated into a distributed file storage system, fused according to data type, and theme libraries are built based on resource libraries and multiple application directions, driving the construction of behavioral theme libraries from the bottom up and top down.

Benefits of technology

It enables unified data aggregation and multi-business applications, reduces redundant development, lowers development costs, improves data utilization efficiency, enriches application scenarios, and enhances data reusability and portability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115221114B_ABST
    Figure CN115221114B_ABST
Patent Text Reader

Abstract

The application provides a data management method and system for communication big data, which comprises the following steps: step 1, mobile communication data of different data sources is stored in a unified distributed file storage system according to a standardized format to form an original library; step 2, the mobile communication data of the original library is extracted, fused, transformed and stored based on resource types to form a resource library; step 3, event behavior subject libraries of different themes are constructed from bottom to top based on the resource library data, whether the established event behavior subject libraries meet the needs of business scenes is analyzed, and when the needs of business scenes cannot be met, abstract behavior subject libraries are constructed from top to bottom based on business scene requirements; and step 4, corresponding event behavior subject libraries and / or abstract behavior subject libraries are selected for different business scenes, and business application libraries are constructed and generated based on the selected subject libraries.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a data governance method, specifically, to a data governance method and system for big data in communications. Background Technology

[0002] Current mobile communication big data governance is still primarily based on actual systems, with data scattered across multiple vendors' systems. Data silos remain significant across different sectors, and most telecom operators employ a "siloed" system construction model, such as... Figure 1 As shown, a reasonable data governance and classification method has not yet been formed, and effective correlation analysis cannot be carried out between various data. This results in data being scattered across various data systems, with a serious data silo effect, poor data application effectiveness, and failure to leverage the value and advantages of mobile communication big data. Consequently, unscientific and non-standard practices are found in data collection, storage, integration, and analysis, leading to frequent duplication of development and other problems, which ultimately affect the overall governance and development of mobile communication big data.

[0003] Secondly, mobile communication data originates from equipment and systems from numerous manufacturers, resulting in inconsistent data standards, inconsistencies in data formats and types, and failure to classify and grade data from multiple sources according to the characteristics of big data communication itself, leading to high costs for data understanding and use.

[0004] Finally, the application of mobile communication data is not yet in-depth, the application field is relatively narrow, and the integration of data with scenarios is insufficient, making it difficult for the "sand" of data to be gathered into a "tower", the massive data resources cannot be activated, and the data potential cannot be fully released.

[0005] In order to solve the above problems, people have been seeking an ideal technological solution. Summary of the Invention

[0006] The purpose of this invention is to address the shortcomings of existing technologies by providing a data governance method and system for big data communication.

[0007] To achieve the above objectives, the technical solution adopted by this invention is: a data governance method for big data in communications, comprising the following steps:

[0008] Step 1: Collect and store mobile communication data from different data sources in a standardized format into a unified distributed file storage system to form the original library;

[0009] Step 2: Extract, merge, transform, and store the mobile communication data from the original library based on the resource type to form a resource library;

[0010] Step 3: Based on the resource library data, build event behavior theme libraries with different themes from the bottom up, analyze whether the established event behavior theme libraries meet the needs of business scenarios, and if they cannot meet the needs of business scenarios, build abstract behavior theme libraries from the top down based on the needs of business scenarios.

[0011] Step 4: Select the corresponding event behavior theme library and / or abstract behavior theme library for different business scenarios, and build and generate the business application library based on the selected theme library.

[0012] Furthermore, when building event behavior theme libraries on different topics from the bottom up based on resource library data, the following steps are executed:

[0013] Based on the preset classification rules, the dimensional data fields in each behavioral event data are classified, and the classification results are natural dimension fields and growth dimension fields;

[0014] All natural dimension fields are integrated and compressed to obtain the integrated and compressed value;

[0015] The integrated compressed value is concatenated with all growth dimension fields to form an abstract key-value pair for each behavioral event;

[0016] Synchronize the abstract key-value pairs of each behavioral event data to all topic libraries.

[0017] Furthermore, when building an abstract behavior theme library from top to bottom based on business scenario requirements, the following steps are executed:

[0018] Determine the topic table for the abstract behavior topic library, treat the dimension fields to be added to the topic table as random events, and calculate the probability distribution P of the random events.

[0019]

[0020] Among them, w x The dimension fields to be added to the topic library table are w1, w2, ..., w1, w2, ... are all known dimension fields in the topic library table, s is the topic already determined in the topic library table, Z is the normalization factor, and λ1 and λ2 are the set weights.

[0021] Compare the probability distribution P with the preset entropy value. When the probability distribution P is greater than the preset entropy value, add the dimension field to the topic library table to form the abstract behavior topic library.

[0022] This invention also provides a data governance system for big data in communications, comprising:

[0023] The global data source aggregation module is used to aggregate mobile communication data from different data sources into a unified distributed file storage system in a standardized format to form the original library;

[0024] The multi-domain data fusion module is used to extract, fuse, transform, and store mobile communication data from the original database based on resource type, forming a resource database.

[0025] The two-dimensional theme library construction module is used to build event behavior theme libraries of different themes from the bottom up based on resource library data, analyze whether the established event behavior theme libraries meet the needs of business scenarios, and build abstract behavior theme libraries from the top down based on business scenario requirements when they cannot meet the needs of business scenarios.

[0026] The business application library construction module is used to select the corresponding event behavior theme library and / or abstract behavior theme library for different business scenarios, and build and generate the business application library based on the selected theme library.

[0027] This invention possesses significant substantive features and substantial advancements compared to existing technologies. Specifically, addressing the complex data sources and diverse data structures of mobile communication big data, this invention utilizes techniques such as full-domain data source aggregation, multi-network domain data fusion, two-dimensional subject library construction, and business application library building to construct a bottom-up, full-lifecycle data governance model for mobile communication big data. This model unifies data aggregation into a distributed file storage system, integrates detailed data according to data type, and builds a unique subject library for the mobile communication field based on a resource library and multiple application directions, providing support for richer application scenarios. This data governance method has several advantages.

[0028] In this invention, multi-domain data fusion is categorized according to data type, integrating multiple data points from a single event and data from multiple interfaces to form various types of resource libraries. The construction of the two-dimensional theme library proceeds from the bottom up, generating specific behaviors based on data, while simultaneously abstracting behavioral events from the top down based on business needs. Under both data-driven and business-driven approaches, multiple theme library outcomes are formed, including app behaviors, call behaviors, regional behaviors, and number behaviors. This provides reusable data support for multi-business applications, reducing the current situation of "siloed" repetitive development across various business lines, and offering strong reusability and portability.

[0029] This patent addresses the complex data sources and diverse data structures of mobile communication big data by dividing the data into four levels: raw data library, resource data library, subject data library, and application data library. The data hierarchy is reasonable, the model development cost is low, the actual data governance effect is obvious, and the solution has strong portability, which can achieve the goal of reducing costs and increasing efficiency. Attached Figure Description

[0030] Figure 1 It is a traditional mobile communication data governance model.

[0031] Figure 2 This is a flowchart illustrating the present invention.

[0032] Figure 3 This is a schematic diagram of two behavioral events in the embodiment.

[0033] Figure 4 This is a schematic diagram of the structure of Embodiment 2 of the present invention. Detailed Implementation

[0034] Concept Explanation

[0035] The original database refers to the data set generated by extracting original data from the entire domain according to the data source, ensuring data integrity.

[0036] Resource repository: refers to the data set generated by extracting and transforming original database data according to resource type, ensuring data consistency, integrity, security, and usability;

[0037] Thematic repositories refer to data collections generated by aggregating, reducing dimensionality, and performing statistics on data from resource repositories, ensuring data reusability and integrity.

[0038] An application library refers to a collection of data generated by analyzing data from a theme library or resource library for different business scenarios, which can support business needs.

[0039] The technical solution of the present invention will be further described in detail below through specific embodiments.

[0040] Example 1

[0041] like Figure 2 As shown, this embodiment provides a data governance method for big data in communications, including the following steps:

[0042] Step 1: Collect and store mobile communication data from different data sources in a standardized format into a unified distributed file storage system to form the original library;

[0043] Step 2: Extract, merge, transform, and store the mobile communication data from the original library based on the resource type to form a resource library;

[0044] Step 3: Based on the resource library data, build event behavior theme libraries with different themes from the bottom up, analyze whether the established event behavior theme libraries meet the needs of business scenarios, and if they cannot meet the needs of business scenarios, build abstract behavior theme libraries from the top down based on the needs of business scenarios.

[0045] Step 4: Select the corresponding event behavior theme library and / or abstract behavior theme library for different business scenarios, and build and generate the business application library based on the selected theme library.

[0046] In practical implementation, based on different business scenarios, we analyze whether the existing topic database can meet the needs of the business scenario through light calculations. Here, light calculations refer to the fact that the fields in the topic database can be combined using simple addition, subtraction, multiplication, division, and other methods to obtain the data required by the business. For example, a topic table partitioned by day may contain fields such as mobile phone number, total call duration, and total number of callers. For business scenarios that require "average call duration of all mobile phone numbers in the last x days", the topic table can be reused.

[0047] In practice, the following principles should be followed during step 1:

[0048] Standardization: Data aggregation is carried out in strict accordance with the standard data access process.

[0049] Uniformity: Aggregate data from across the entire domain, categorize it according to its source, and use it as a unified entry point for data access;

[0050] Periodicity: Access and update data periodically based on actual data conditions.

[0051] Traceability refers to marking the source and dissemination of data so that the source of data quality problems can be quickly located when problems occur.

[0052] Based on the above principles, in this embodiment, the specific steps of step 1 are as follows:

[0053] Step 1.1: Organize the mobile communication data from all data sources according to network type, identify the interfaces and fields corresponding to each network type, confirm the data storage method and storage information, and form a resource table; specifically, network types include telecommunications networks, mobile internet, and other data service networks, etc. By parsing the data from telecommunications networks and mobile internet through various parsing devices, 2G, 3G, 4G, and 5G network data can be further obtained.

[0054] Step 1.2: Based on the resource table, determine the data aggregation method and manage data tasks uniformly in the task scheduling center. Offline data adopts the offline data access method, and real-time data adopts the real-time data access method.

[0055] Specifically, during implementation, the task naming convention is: task_{ods / dwd / dws / ads}_task description_scheduling period;

[0056] Task Description: The English description of this task.

[0057] Scheduling cycle: day, month, week, hour, year, minute, second

[0058] For example, the task 'task_dws_number_portrait_day' is a task that extracts number profiles from the subject database, and its scheduling cycle is once a day.

[0059] Step 1.3: Using a batch processing engine or a streaming computing engine, mobile communication data from different data sources are aggregated and stored in a unified distributed file storage system through corresponding storage methods and storage information to form the original library.

[0060] Specifically, the naming conventions for the original database tables are as follows:

[0061] Table naming convention: ods_{data source}_{original table name in English};

[0062] Field specifications: Fields use the source system field names by default.

[0063] Data source: Abbreviation for data source system

[0064] Belongs to: ODS layer

[0065] For example: ods_cs_cdr cs: Comleader Shield Lingdun system

[0066] illustrate:

[0067] []: This indicates an optional field, which can be left blank if not specified. If specified, it is required.

[0068] {}: indicates a required field.

[0069] When storing data, the source information of the data is preserved through methods such as table naming, so that the data of each domain interface is stored in a unified distributed file storage system, and the data is unified at the storage and usage levels.

[0070] It is understandable that the purpose of step 1 is to inventory and statistically analyze existing data, and to plan for available or necessary data resources. After aggregating data sources across the entire domain, it is necessary to build a mobile communication resource library covering multiple data types to further explore the value of data resources and provide data support for the construction of thematic libraries.

[0071] In this embodiment, step 2 is as follows:

[0072] Step 2.1: Extract mobile communication data from the original database based on resource type;

[0073] Step 2.2: Perform multi-domain interface fusion and multi-domain record fusion on the extracted mobile communication data. The multi-domain interface fusion includes fusing mobile communication data from multiple domain interfaces that belong to the same behavioral event according to communication principles. The multi-domain record fusion includes merging mobile communication data based on multiple device IDs and event occurrence times in the mobile communication data using an approximation analysis method.

[0074] Step 2.3: Store the merged mobile communication data in the corresponding resource library according to the preset standard format.

[0075] It's understandable that the evolution of communication technology from 2G to 5G doesn't mean the disappearance of 2G, 3G, and 4G technologies. Rather, the coexistence of 2G, 3G, 4G, and 5G technologies leads to a significant problem. For example, in call scenarios, there's call data from 2G, 3G, and 4G technologies, as well as call data from 5G technologies, and their data formats, fields, and naming conventions are not consistent. Therefore, simply recording and accumulating all call data is insufficient; fusion is necessary.

[0076] Since the data for the same behavioral event is scattered across multiple network interfaces, i.e. multiple tables in the original database, for example, the location update data of 2G, 3G and 4G mobile devices is in the A / IuCS interface, while the location update data of 5G fallback mobile devices is in the S1-MME interface, it is necessary to merge the data from different network interfaces according to the communication principle.

[0077] In practice, based on communication principles, mobile communication data from multiple network domain interfaces belonging to the same behavioral event are merged. This includes: statistically analyzing the fields under each network domain, horizontally concatenating the fields under all network domains to obtain a superset of the call service process fields, establishing mapping relationships between fields based on the superset, deleting and merging fields, and finally obtaining the merged data.

[0078] In practice, the same behavioral event may generate multiple data entries. For example, in a call event, the data for a single call is scattered across multiple records based on the calling and called parties. In such cases, a multi-domain record fusion method must be employed. Based on multiple device IDs and event times within the data, an approximation analysis method is used to merge the data, achieving a one-to-one mapping between user behavior and data records.

[0079] The reason for using an approximation analysis method here is that the communication data contains several unique device IDs such as phone number, IMCI, and IMSI, and each record also has an event occurrence time (accurate to milliseconds). Ideally, the data would be perfectly matched by associating the device ID and the occurrence time. In reality, due to various reasons, the event occurrence time may not match perfectly, so it is necessary to push the occurrence time forward or backward by a certain interval to perform the association.

[0080] Understandably, since data quality is the foundation for ensuring data application, it is necessary to guarantee data quality when building a resource repository. Its evaluation criteria mainly include four aspects: completeness, consistency, accuracy, and timeliness. Among them, completeness means that the structure of various data resources after integration should be complete and all elements should be present; consistency means that the data in the resource repository should be as consistent as possible with the data in the original repository; usability means that the naming of each resource category should be clear enough to facilitate understanding and use; and security means that the data in the resource repository should have corresponding levels of data anonymization and desensitization strategies based on the security level of different data.

[0081] Based on the above principles, before executing step 2, the original database data is assessed according to the data quality rules to determine whether there is any mobile communication data that needs to be governed, and the mobile communication data that needs to be governed is governed in accordance with the data quality rules; only the governed data can be stored and merged.

[0082] Furthermore, after completing the construction of the resource repository, it is necessary to aggregate, reduce the dimensionality, and perform statistics on the resource repository data to generate thematic libraries, ensuring the reusability and integrity of the data. The construction of thematic libraries is also a core step, which is directly related to whether the data can fully support the needs of upper-level business applications, meet data governance requirements, reduce costs and increase efficiency, and maximize data effectiveness.

[0083] This invention employs a bottom-up approach to building event behavior theme libraries based on resource library data, and a top-down approach to building abstract behavior theme libraries based on business scenario requirements. In other words, it performs behavior definition collisions from two-dimensional perspectives: business event-driven and data attribute-driven. When the two are consistent, they are aggregated into the same theme domain, thus opening up the transformation channel between data and business.

[0084] This aggregation method presents two prominent problems:

[0085] 1. When data attributes are driven from the bottom up, there may be situations where the same data dimension satisfies multiple behavioral events (subject domains). How should this data dimension be classified into subject domains? If related subject domains also contain this dimension, it will cause problems such as redundant construction, high storage costs, and inconsistent calculation methods, which violates the basic requirements of the data governance system.

[0086] For example, such as Figure 3 As shown, when building an app behavior theme library, the question of "who accessed me" needs to be addressed. Here, "who" refers to the phone number or device ID, and "me" refers to the app. When building a terminal behavior theme library, the question of "who did I access?" needs to be addressed. Here, "who" refers to the app or website address, and "who" refers to the device ID. Both of these behavior events originate from the same dimension in the internet access logs of the resource library, resulting in duplicate calculations and inconsistent definitions. When the upper-layer business application library requires support from both theme libraries simultaneously, conflicts and contradictions will arise.

[0087] 2. When driven from top to bottom by business needs, the number of data dimensions required for the same behavioral event (subject domain) will increase. Business needs are greedy, but data dimensions are limited. How to set boundaries to meet the business needs of behavioral events with the minimum set of data dimensions, and prevent the creation of wide-dimensional tables similar to resource repositories during subject database construction, so that each subject database becomes increasingly similar, which will also cause redundant construction, high storage costs, and unclear classification standards, which violates the basic requirements of the data governance system.

[0088] Regarding the first question, we return to the data itself from the two natural language descriptions "Whom did I visit" and "Who visited me". The subject, verb, and object all come from the same data dimension fields defined in the resource library. However, the relevant dimensions are integrated and calculated using the action dimension of "visit" as the fulcrum to obtain the high-dimensional information required in the subject library.

[0089] Therefore, when constructing event behavior theme libraries of different themes from the bottom up based on resource library data, this invention adopts a solution of dimensionality reduction and convergence:

[0090] S1, according to the preset division rules, classify the dimensional data fields in each behavioral event data, and the classification results are natural dimension fields and growth dimension fields;

[0091] For example, the values ​​of fields such as phone number - 1859585xxxx and app name - WeChat are "inherited" and can be inherited. We define them as natural dimension fields. Under normal circumstances, they will not change in the short term. There can be tens of thousands of lines of internet access log records. The phone number is the same and the app accessed can be the same, but the access time is constantly changing. We define them as growth dimension fields.

[0092] S2, integrate and compress all natural dimension fields to obtain integrated compressed values; concatenate the integrated compressed values ​​with all growth dimension fields to form the abstract key value for each behavioral event;

[0093] Normally, a single piece of information (i.e., a single row of data) contains far more natural dimension fields than growth dimension fields. Furthermore, natural dimensions typically serve as information identifiers, while growth dimensions typically provide business support. The natural dimensions are integrated and compressed (using the function Cps) to retain their unique identifier characteristics, while the growth dimensions remain unchanged. The specific formula is as follows:

[0094] Abstract key value = (growth dimension 1 & growth dimension 2 & ... ):::Cps(natural dimension 1, natural dimension 2, natural dimension 3, ...);

[0095] S3 synchronizes the abstract key-value pairs of each behavioral event data to all subject libraries.

[0096] The applicant replaced the high-dimensional information obtained from the original aggregation calculation with artificially defined abstract key values, forcibly shielding the calculation process of natural logic, and forcibly reducing the original "who visited me" and "who I visited" to "me" and "who". At the same time, because the original growth dimension is preserved, information on the "visit" dimension can be queried at any time, reducing the cost of storage and use, simplifying the logic of business analysis, and achieving the unification of calculation methods.

[0097] Regarding the second question, since the uncertainty of a piece of information is directly related to the amount of information it can express, each topic table hopes to express more information and show the greatest uncertainty, that is, to have the largest information entropy. Therefore, when designing a topic table, it is inevitable that the more dimension fields in the table, the better, so as to better support business needs and reduce the correlation and coupling between different topic tables. As a result, the topic table will become larger and larger, and the number of dimensions involved will increase, which will definitely violate our design principles.

[0098] Maximum entropy theory tells us that, without external forces, things always tend towards the most chaotic direction. Things are a unity of constraint and freedom. Things always strive for maximum freedom under constraints, which is actually a fundamental principle of nature. The contradiction between maximizing information entropy and the finite requirement of dimensional fields in the subject database table is like a company wanting employees to work more overtime to create more value and increase revenue, while also wanting to pay less overtime to reduce expenses. As employees work more overtime, they become increasingly tired, and their earnings decrease. Eventually, a balance point will emerge where the benefit ratio is highest. Similarly, we can find a way to determine the optimal point of benefit that maximizes information entropy when adding fields to the subject database table.

[0099] Hungarian mathematician and Shannon Prize laureate, Szysá, proved that for any set of non-contradictory information, maximum entropy not only exists but is also unique, and all have the same simple form—an exponential function. The maximum entropy principle states that when predicting the probability distribution of a random event, our prediction should satisfy all known conditions, and the probability of the unknown parts should be uniform. This minimizes the risk of prediction because the information entropy is maximized; hence, this model is called the maximum entropy model.

[0100] Therefore, this invention introduces the maximum entropy model into the construction steps of the abstract behavior theme library. Specifically, when constructing the abstract behavior theme library from top to bottom based on business scenario requirements, the following steps are performed:

[0101] Determine the topic table for the abstract behavior topic library, treat the dimension fields to be added to the topic table as random events, and calculate the probability distribution P of the random events.

[0102]

[0103] Among them, w x The dimension fields to be added to the topic library table are w1, w2, ..., w1, w2, ... are all known dimension fields in the topic library table, s is the topic already determined in the topic library table, Z is the normalization factor, and λ1 and λ2 are the set weights.

[0104] Compare the probability distribution P with the preset entropy value. When the probability distribution P is greater than the preset entropy value, add the dimension field to the topic library table to form the abstract behavior topic library.

[0105] Example 2

[0106] This embodiment provides a data governance system for big data in communications, such as... Figure 4 As shown, it includes:

[0107] The global data source aggregation module is used to aggregate mobile communication data from different data sources into a unified distributed file storage system in a standardized format to form the original library;

[0108] The multi-domain data fusion module is used to extract, fuse, transform, and store mobile communication data from the original database based on resource type, forming a resource database.

[0109] The two-dimensional theme library construction module is used to build event behavior theme libraries of different themes from the bottom up based on resource library data, analyze whether the established event behavior theme libraries meet the needs of business scenarios, and build abstract behavior theme libraries from the top down based on business scenario requirements when they cannot meet the needs of business scenarios.

[0110] The business application library construction module is used to select the corresponding event behavior theme library and / or abstract behavior theme library for different business scenarios, and build and generate the business application library based on the selected theme library.

[0111] Specifically, when the two-dimensional theme library construction module builds event behavior theme libraries for different themes from bottom to top based on resource library data, it executes:

[0112] Based on the preset classification rules, the dimensional data fields in each behavioral event data are classified, and the classification results are natural dimension fields and growth dimension fields;

[0113] All natural dimension fields are integrated and compressed to obtain the integrated and compressed value;

[0114] The integrated compressed value is concatenated with all growth dimension fields to form an abstract key-value pair for each behavioral event;

[0115] Synchronize the abstract key-value pairs of each behavioral event data to all topic libraries.

[0116] Specifically, when the two-dimensional theme library construction module builds an abstract behavior theme library from top to bottom based on business scenario requirements, it executes:

[0117] Determine the topic table for the abstract behavior topic library, treat the dimension fields to be added to the topic table as random events, and calculate the probability distribution P of the random events.

[0118]

[0119] Among them, w x The dimension fields to be added to the topic library table are w1, w2, ..., w1, w2, ... are all known dimension fields in the topic library table, s is the topic already determined in the topic library table, Z is the normalization factor, and λ1 and λ2 are the set weights.

[0120] Compare the probability distribution P with the preset entropy value. When the probability distribution P is greater than the preset entropy value, add the dimension field to the topic library table to form the abstract behavior topic library.

[0121] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them; although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications can still be made to the specific implementation of the present invention or equivalent substitutions can be made to some technical features without departing from the spirit of the technical solutions of the present invention, and all such modifications and substitutions should be covered within the scope of the technical solutions claimed in the present invention.

Claims

1. A data governance method for big data in communications, characterized in that, Includes the following steps: Step 1: Collect and store mobile communication data from different data sources in a standardized format into a unified distributed file storage system to form the original library; Step 2: Extract, merge, transform, and store the mobile communication data from the original library based on the resource type to form a resource library; The specific steps are as follows: Mobile communication data is extracted from the original library based on resource type; The extracted mobile communication data is subjected to multi-domain interface fusion and multi-domain record fusion. The multi-domain interface fusion includes fusing mobile communication data from multiple domain interfaces that belong to the same behavioral event according to communication principles. The multi-domain record fusion includes merging mobile communication data based on multiple device IDs and event occurrence times in the mobile communication data using an approximation analysis method; The merged mobile communication data is stored in the corresponding resource library according to a preset standard format; The process involves fusing mobile communication data from multiple network domain interfaces that belong to the same behavioral event, based on communication principles. This includes: statistically analyzing the fields under each network domain, horizontally concatenating the fields under all network domains to obtain a superset of the call service process fields, establishing mapping relationships between the fields based on the superset, gradually deleting and merging fields, and finally obtaining the fused data. Step 3: Based on the resource library data, build event behavior theme libraries with different themes from the bottom up, analyze whether the established event behavior theme libraries meet the needs of business scenarios, and if they cannot meet the needs of business scenarios, build abstract behavior theme libraries from the top down based on the needs of business scenarios. We adopt a bottom-up approach to building event behavior theme libraries with different themes based on resource library data, and a top-down approach to building abstract behavior theme libraries based on business scenario requirements. That is, we perform behavior definition collision from two-dimensional perspectives: business event-driven and data attribute-driven. When the two are consistent, we aggregate the two into the same theme domain, thus opening up the transformation channel between data and business. When building event behavior theme libraries on different topics from the bottom up based on resource library data, the following steps are executed: Based on the preset classification rules, the dimensional data fields in each behavioral event data are classified, and the classification results are natural dimension fields and growth dimension fields; All natural dimension fields are integrated and compressed to obtain the integrated and compressed value; The integrated compressed value is concatenated with all growth dimension fields to form an abstract key-value pair for each behavioral event; Synchronize the abstract key-value pairs of each behavioral event data to all topic libraries; When building an abstract behavior theme library from top to bottom based on business scenario requirements, the following steps are executed: Determine the topic table for the abstract behavior topic library, treat the dimension fields to be added to the topic table as random events, and calculate the probability distribution P of the random events. , Among them, w x λ1, w2, ... are the dimension fields to be added to the topic library table, w1, w2, ... are all known dimension fields in the topic library table, s is the topic already determined in the topic library table, Z is the normalization factor, and λ1 and λ2 are the set weights. Compare the probability distribution P with the preset entropy value. When the probability distribution P is greater than the preset entropy value, add the dimension field to the topic library table to finally form the abstract behavior topic library. Step 4: Select the corresponding event behavior theme library and / or abstract behavior theme library for different business scenarios, and build and generate the business application library based on the selected theme library.

2. The data governance method for big data in communications according to claim 1, characterized in that, The specific steps of step 1 are as follows: Based on network type, mobile communication data from all data sources are sorted out, the interfaces and fields corresponding to each network type are identified, the data storage method and storage information are confirmed, and a resource table is formed. Based on the resource table, determine the data aggregation method and manage data tasks uniformly in the task scheduling center. Offline data adopts the offline data access method, and real-time data adopts the real-time data access method. Using a batch processing engine or a streaming computing engine, mobile communication data from different data sources is aggregated, and the processed mobile communication data is aggregated and stored in a unified distributed file storage system through corresponding storage methods and storage information to form the original library.

3. The data governance method for big data in communications according to claim 1, characterized in that, Before performing step 2, the original database data is assessed for quality according to data quality rules to determine whether there is any mobile communication data that needs to be managed, and the mobile communication data that needs to be managed is managed in accordance with data quality rules.

4. A data governance system for communication big data, characterized in that, include: The global data source aggregation module is used to aggregate mobile communication data from different data sources into a unified distributed file storage system in a standardized format to form the original library; The multi-domain data fusion module is used to extract, fuse, transform, and store mobile communication data from the original database based on resource type, forming a resource database. The specific steps are as follows: Mobile communication data is extracted from the original library based on resource type; The extracted mobile communication data is subjected to multi-domain interface fusion and multi-domain record fusion. The multi-domain interface fusion includes fusing mobile communication data from multiple domain interfaces that belong to the same behavioral event according to communication principles. The multi-domain record fusion includes merging mobile communication data based on multiple device IDs and event occurrence times in the mobile communication data using an approximation analysis method; The merged mobile communication data is stored in the corresponding resource library according to a preset standard format; The process involves fusing mobile communication data from multiple network domain interfaces that belong to the same behavioral event, based on communication principles. This includes: statistically analyzing the fields under each network domain, horizontally concatenating the fields under all network domains to obtain a superset of the call service process fields, establishing mapping relationships between the fields based on the superset, gradually deleting and merging fields, and finally obtaining the fused data. The two-dimensional theme library construction module is used to build event behavior theme libraries of different themes from the bottom up based on resource library data, analyze whether the established event behavior theme libraries meet the needs of business scenarios, and build abstract behavior theme libraries from the top down based on business scenario requirements when they cannot meet the needs of business scenarios. We adopt a bottom-up approach to building event behavior theme libraries with different themes based on resource library data, and a top-down approach to building abstract behavior theme libraries based on business scenario requirements. That is, we perform behavior definition collision from two-dimensional perspectives: business event-driven and data attribute-driven. When the two are consistent, we aggregate the two into the same theme domain, thus opening up the transformation channel between data and business. The business application library building module is used to select the corresponding event behavior theme library and / or abstract behavior theme library for different business scenarios, and build and generate the business application library based on the selected theme library. When the two-dimensional theme library construction module builds event behavior theme libraries of different themes from bottom to top based on resource library data, it executes: Based on the preset classification rules, the dimensional data fields in each behavioral event data are classified, and the classification results are natural dimension fields and growth dimension fields; All natural dimension fields are integrated and compressed to obtain the integrated and compressed value; The integrated compressed value is concatenated with all growth dimension fields to form an abstract key-value pair for each behavioral event; Synchronize the abstract key-value pairs of each behavioral event data to all topic libraries; When the 2D topic library construction module builds an abstract behavior topic library from top to bottom based on business scenario requirements, it executes the following: Determine the topic table for the abstract behavior topic library, treat the dimension fields to be added to the topic table as random events, and calculate the probability distribution P of the random events. , Among them, w x λ1, w2, ... are the dimension fields to be added to the topic library table, w1, w2, ... are all known dimension fields in the topic library table, s is the topic already determined in the topic library table, Z is the normalization factor, and λ1 and λ2 are the set weights. Compare the probability distribution P with the preset entropy value. When the probability distribution P is greater than the preset entropy value, add the dimension field to the topic library table to form the abstract behavior topic library.

Citation Information

Patent Citations

  • Multi-source data fusion system and method

    CN105893526A

  • Method and system for processing user browsing behavior data based on data warehouse

    CN108268565A