Data asset operation method, device, electronic device and storage medium
By screening and optimizing the data asset management model, the problem of underutilizing the mobile location data value is solved, and the effect of quickly finding the right data according to business needs and improving the value of data assets is achieved.
Patent Information
- Application Number
- CN202110823573.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-07-21
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2041-07-21
AI Technical Summary
The value of mobile location data is not fully explored, mainly due to the huge amount of data and the difficulty in effectively using it.
By obtaining business needs, filtering target data groups, determining the matching degree and frequency, optimizing the data asset management model, and enhancing the value of data assets.
It realizes the rapid finding suitable data based on business needs, dynamically optimizes the data asset management model, and enhances the value of data assets.
Smart Images

Figure CN115687438B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of big data technology, and is related to, but not limited to, a data asset operation method, device, electronic device, and storage medium. Background Art
[0002] The widespread adoption of mobile communications has led to an exponential growth in the number of public mobile communication base stations. Human communication activities using these base stations have generated a vast amount of terminal signaling data. This data, linked to the location information of both people and base stations, naturally fulfills the three key elements of "person, place, and time," making it a crucial data asset for mobile operators, often referred to as mobile location data. Compared to spatiotemporal location data from other sources, this data asset inherently offers advantages in long-range coverage and around-the-clock collection. However, due to its massive volume, the value of mobile location data has not been fully realized. Summary of the Invention
[0003] In view of this, embodiments of the present application provide a data asset operation method, device, electronic device, and storage medium.
[0004] In a first aspect, an embodiment of the present application provides a method for operating data assets, the method comprising: obtaining multiple business requirements; for each business requirement, screening out a target data group corresponding to the business requirement from a data asset management model; the target data group including the first table-level metadata and the first field-level metadata in the data asset management model; determining the degree of matching between each business requirement and the corresponding target data group; determining, among each business requirement, a business requirement whose frequency of occurrence and corresponding degree of matching among multiple business requirements meet specific conditions as an asset optimization requirement; and optimizing the data asset management model based on the asset optimization requirement.
[0005] In the second aspect, an embodiment of the present application provides an operation device for data assets, including: an acquisition module for acquiring multiple business requirements; a screening module for screening out a target data group corresponding to each business requirement from a data asset management model; the target data group includes the first table-level metadata and the first field-level metadata in the data asset management model; a first determination module for determining the degree of matching between each business requirement and the corresponding target data group; a second determination module for determining, among each business requirement, the business requirement whose frequency of occurrence and corresponding matching degree in multiple business requirements meet specific conditions as an asset optimization requirement; an optimization module for optimizing the data asset management model based on the asset optimization requirement.
[0006] In a third aspect, an embodiment of the present application provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the program, the steps in the data asset operation method described in any of the embodiments of the present application are implemented.
[0007] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps in the data asset operation method described in any of the embodiments of the present application are implemented.
[0008] In an embodiment of the present application, target data groups can be filtered out from the data asset management model according to business needs, which can help data demanders quickly find suitable data for use; according to business needs, the data asset management model can be dynamically optimized to enhance the value of data assets. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] Figure 1 A flowchart of a data asset operation method according to an embodiment of the present application;
[0010] Figure 2 A flowchart of a method for establishing a data asset management model according to an embodiment of the present application;
[0011] Figure 3 A schematic diagram of a data table modeling of a data asset management model according to an embodiment of the present application;
[0012] Figure 4 This is a flow chart of a method for cyclic operation and management of data assets according to an embodiment of the present application;
[0013] Figure 5 for Figure 4 The detailed execution flow diagram of the cyclic operation management method shown;
[0014] Figure 6 This is a schematic diagram of the structure of a data asset operation device according to an embodiment of the present application;
[0015] Figure 7 This is a schematic diagram of the structure of a device for establishing a data asset management model according to an embodiment of the present application;
[0016] Figure 8 A schematic diagram of a hardware entity of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0017] The technical solution of the present application is further described in detail below with reference to the accompanying drawings and embodiments.
[0018] Figure 1This is a flow chart of a data asset operation method according to an embodiment of the present application, as shown in FIG. Figure 1 As shown, the method includes:
[0019] Step 102: Obtain multiple business requirements;
[0020] Among them, business needs can also be called data application needs, data needs, or simply needs. Business needs can be the data demander's needs for reading and writing data. For example, a business need can be the monthly peak daily traffic volume of major scenic spots in Xi'an in 2020. Business needs can also be the preferred shopping hours of male customers in a shopping mall in Yanta District, Xi'an.
[0021] Step 104: For each business requirement, select a target data group corresponding to the business requirement from the data asset management model; the target data group includes the first table-level metadata and the first field-level metadata in the data asset management model;
[0022] Among them, the data asset management model can be a data model used to manage and maintain data assets, and the data asset management model is the basis of the database; the data assets refer to data resources owned or controlled by individuals or enterprises, which can bring future economic benefits to the enterprise and are recorded in physical or electronic form; the data resources can be metadata of terminal signaling data; the data asset management model is used to manage the metadata of terminal signaling data through data tables; the data tables can be logical tables.
[0023] The data asset management model may include a data warehouse layer, a data mart layer and a data application layer; the data warehouse layer may store pre-processed and organized mobile location detail data; the data mart layer may store aggregated label data processed from the data warehouse layer and available for direct application call; the data application layer may store data processed from the data warehouse layer and / or the data mart layer; the data application layer may store intermediate calculation process tables for commonly used call requirements.
[0024] Terminal signaling data can be communication data between a terminal user and a transmitting base station or a micro station; the terminal signaling data can also be called mobile location data; the metadata of the terminal signaling data can be data used to describe the terminal signaling data, mainly used to describe information on the data attributes of the terminal signaling data, to support functions such as indicating storage location, historical data, resource search, file records, etc.; when doing data asset management, there are many types of metadata of terminal signaling data that can be involved, such as technical metadata and business metadata of terminal signaling data. The data asset management model can manage the business metadata of terminal signaling data through data tables.
[0025] The data warehouse layer, the data mart layer and the data application layer include a plurality of data tables; each of the data tables includes table-level metadata and field-level metadata. The table-level metadata may include the layer to which it belongs, business scenario, data timeliness and standard delay (also referred to as data delay); the field-level metadata may include category, range and granularity; wherein the category may be a three-element keyword (for example, a person, place or time), label attribute, others, etc.; the range may be a regional range, a time span (also referred to as a time range), a person range (divided by the number's place of origin, a certain type of attribute), a label value (for example: occupational label, including label values such as student, civil servant, teacher, etc.); the granularity may be location granularity, time granularity or person granularity.
[0026] Step 106: Determine the degree of match between each business requirement and the corresponding target data set;
[0027] Here, referring to step 104, a target data group can be screened out from the data asset management model according to each of the business requirements. Therefore, the matching degree between each business requirement and the corresponding target data group can be determined, and the occurrence frequency of each of the business requirements can be counted.
[0028] Step 108: Determine, among the various business requirements, the business requirements whose occurrence frequency and corresponding matching degree among the multiple business requirements meet specific conditions as asset optimization requirements;
[0029] It should be noted that, in one embodiment, at least one business requirement whose occurrence frequency and corresponding matching degree meet specific conditions may be determined as an asset optimization requirement.
[0030] In another embodiment, the various business requirements may be classified according to the similarity between the business requirements, and the frequency of occurrence of each type of business requirement (i.e., the frequency of occurrence of similar business requirements) may be determined. Based on the frequency of occurrence of similar business requirements and the degree of matching between the similar business requirements and the corresponding target data group, at least one type of business requirement may be determined from the multiple business requirements as at least one asset optimization requirement. Each type of business requirement may include at least one business requirement. Assuming that a business requirement item is D i , the target data group corresponding to this business requirement item is G a , D i With G a The matching degree between them is M i , D i The frequency of occurrence is m, D i The same demand is F i , F i The frequency of occurrence of G is n, then G aAs F i The corresponding target data group, if F i With G a The matching degree between them is N i , we can determine that the frequency of occurrence of this type of business demand is m+n, and the matching degree between this type of business demand and the corresponding target data group is M i +N i , then at least one asset optimization requirement can be determined based on the frequency of occurrence of this type of business requirement and the corresponding matching degree.
[0031] Step 110: Optimize the data asset management model based on the asset optimization requirements.
[0032] In an embodiment of the present application, target data groups can be screened out from the data asset management model based on business needs, which can help data demanders quickly find suitable data for use; asset optimization needs can be determined based on business needs, the data asset management model can be dynamically optimized, and the value of data assets can be enhanced.
[0033] The present invention also provides a data asset operation method, which includes the following steps:
[0034] Step S202: Acquire multiple business requirements;
[0035] To facilitate understanding, three examples of requirements can be cited for illustration: Business Requirement 1: Find the monthly peak daily visitor flow at major scenic spots in Huqiu District, Suzhou City, in 2019; Business Requirement 2: Find the escape route of a domestic criminal; Business Requirement 3: Find the preferred shopping hours of female customers at a shopping mall in Chaoyang District, Beijing, in July. The business requirements include target table-level metadata and target field-level metadata.
[0036] Step S204: Filtering a first candidate data table from the data asset management model according to the target table-level metadata of each business requirement;
[0037] The target table-level metadata includes the atomic scenario (i.e., target business scenario) and the timeliness requirement to which the business requirement belongs; the atomic scenario can also be called a business scenario; the timeliness requirement includes a timeliness category parameter (i.e., target data timeliness) and a target data delay, and whether the data timeliness is real-time or offline. Based on the atomic scenario in the business requirement and the timeliness category parameter in the data timeliness, various table-level metadata in the data asset management model can be queried to screen out all eligible data tables as the first candidate data tables.
[0038] In one embodiment, the first candidate data table may be screened from the data asset management model according to the target level, target business scenario, and target data timeliness of each business requirement.
[0039] You can first select the required atomic scenarios for each requirement based on business needs. Each business requirement can be composed of one or more atomic scenarios; the atomic scenarios may include: Scenario 1: Find people (groups) given a known area and time range; Scenario 2: Find places / areas given a known person (group) and time range; Scenario 3: Find time (time period special report) given a known person (group) and area range; Scenario 4: Find label attributes about people (gender, age, occupation, permanent residence); Scenario 5: Find label attributes about areas (places); Scenario 6: Find label attributes about the "people-land relationship" of time periods. For the above three requirements, the atomic scenarios selected are: Business requirement 1 corresponds to Scenario 1; Business requirement 2 corresponds to Scenario 2; Business requirement 3 corresponds to the combination of Scenario 1 and Scenario 4.
[0040] Next, based on the business needs, select the data timeliness requirements for each requirement. These requirements include two aspects: first, whether the data is processed in real time with immediate results or processed offline with batch results, i.e., whether the timeliness category is real-time or offline; and second, the maximum allowable data latency. In one embodiment, Business Requirement 2 requires real-time data to meet ongoing real-time queries, and the desired data latency is as low as possible, with a maximum allowable time of 10 minutes. Business Requirement 1 and Business Requirement 3 primarily focus on historical data statistics and only require offline data for calculations. Business Requirement 3 focuses on only one month's data, and the data request is a one-time one. The time it takes for the data to reach the application caller can be controlled through frequent communication between the two parties during the data request and authorization process. For Business Requirement 1, a monthly report on the previous month's daily peak traffic data is clearly required. At the end of each month, the day of the next month when this report is available reflects the data latency requirement. Assuming that, from a management perspective, completion is desired by mid-month, the maximum allowable data latency may be 15 days.
[0041] Step S206: based on the target field-level metadata of each business requirement, filter out a second candidate data table from the first candidate data table, and filter out candidate metadata from the second candidate data table;
[0042] The target field-level metadata of the business requirement is the requirement parameter of the business requirement, and the requirement parameter may include category (also referred to as target category), range (also referred to as target range) and granularity (also referred to as target granularity); the candidate metadata is metadata related to the target field-level metadata in the second candidate data table; based on the various requirement parameters in the business requirement, the metadata attributes of each field can be matched from the first candidate data table to filter out tables that meet the requirement parameters, and the table that meets the requirement parameters after selecting relevant data elements and deleting irrelevant data elements is used as the second candidate data table. When performing data element matching, the range-type metadata must meet the requirement range less than or equal to the data range; the requirement granularity must be greater than or equal to the data granularity.
[0043] In one embodiment, based on the target category, target range, and target granularity of each business requirement, a second candidate data table can be screened out from the first candidate data table, and candidate metadata can be screened out from the second candidate data table; the candidate metadata is metadata in the second candidate data table that is related to the target category, target range, and target granularity.
[0044] When the category is location, the scope may be a regional scope, and the granularity may be a regional granularity; the smallest regional unit required may be selected as the regional granularity, such as: base station / cell, custom area, province, city, district / county; the maximum regional scope covered by the requirement may be selected according to the regional granularity, such as: nationwide, province, city, district / county, etc.
[0045] When the category is time, the range can be a time range, and the granularity can be a time granularity; the minimum time unit of the requirement can be selected as the time granularity; the time granularity is provided to the user for selection based on the actual table information of the data warehouse, and the optional time granularities are, for example: year, month, day, hour, 15 minutes; the maximum time period length (i.e., time range) covered by the requirement can be selected according to the time granularity, for example: January 2020 to September 2020.
[0046] When the category is a person, the scope may be a task scope, and the granularity may be a person granularity; the smallest person unit required may be selected as the person granularity, for example: a person, a team, an organization; the range of people covered by the requirement may be selected according to the person granularity, for example: all mobile users in Jiangsu Province.
[0047] In addition, for certain tag requirements (tag-based business requirements), you can select a tag category that matches the requirement in the tag library in the data mart layer, or fill in a tag category that is not in the tag library. The range corresponding to the tag requirement is the tag range. For each tag category, the corresponding required tag value is listed. For example, for age group tags, the tag range involved in the requirement is limited to people aged 60-80. Table 1 is a requirement parameter table for business requirements 1 to 3. The configurable requirement parameters are shown in Table 1:
[0048] Table 1
[0049] Business requirements Business Requirement 1 Business Requirement 2 Business Requirement 3 Regional granularity District / County Base station cell Base station cell Regional scope Huqiu District Nationwide A shopping mall in Chaoyang District Time granularity day Hour Hour Time Range Full year 2019 7*24 hours 24 hours a day in July Character granularity people people people Character range Unlimited There may be a list of key personnel Unlimited Tag Category none none Gender, group Tag range none none Female, shopping mall customer
[0050] To sum up, based on the target field-level metadata (i.e., target category, target range, and target granularity) required by business needs, a second candidate data table can be screened out from the first candidate data table, and candidate metadata can be screened out from the second candidate data table, and irrelevant metadata can be deleted from the second candidate data table.
[0051] Step S208: performing permutations and combinations on the second candidate data table to obtain multiple candidate data groups;
[0052] Among them, the second candidate data table after deleting irrelevant data elements can be freely arranged and combined to obtain multiple candidate data groups, and all table-level metadata and field-level metadata in each candidate data group can meet business requirements; a second candidate data table can appear repeatedly in multiple candidate data groups.
[0053] Step S210: determining a target data group from the plurality of candidate data groups according to the data delay and data level of each second candidate data table of each candidate data group;
[0054] The data layer may be one of a data warehouse layer, a data mart layer and a data application layer.
[0055] Step S212: Determine the matching degree between each business requirement and the corresponding target data set;
[0056] The matching degree may be the number of field-level metadata that is consistent between the target field-level metadata in the business requirement and the field-level metadata in each target data table of the target data group.
[0057] Step S214: determining, among the various business requirements, the business requirements whose occurrence frequency and corresponding matching degree among the multiple business requirements meet specific conditions as asset optimization requirements;
[0058] Step S216: Optimize the data asset management model based on the asset optimization requirements.
[0059] In this embodiment of the application, because the terminal signaling data originates from the field of telecommunications operations, its field format is unique and professional, and the frequency intervals at which the data is generated are also highly relevant. By filtering the data table based on the target table-level metadata in the business requirements and then matching the data elements based on the target field-level metadata, the data in the filtered target data group can be made more accurate and meet the user's data requirements.
[0060] The present invention also provides a data asset operation method, which includes the following steps:
[0061] Step S302: Acquire multiple business requirements;
[0062] Step S304: Filtering a first candidate data table from the data asset management model according to the target table-level metadata of each business requirement;
[0063] Step S306: based on the target field-level metadata of each business requirement, filter out a second candidate data table from the first candidate data table, and filter out candidate metadata from the second candidate data table;
[0064] Step S308: performing permutations and combinations on the second candidate data table to obtain multiple candidate data groups;
[0065] It is assumed that the plurality of candidate data groups include candidate data group G a , G a The second candidate data tables Table1 to Table5 are included in G a ∈{Table1,Table2,Table3,Table4,Table5}.
[0066] Step S310: determining the data level corresponding to each second candidate data table in each candidate data group;
[0067] Step S312: determining a first recommendation index according to the number of second candidate data tables in each data level and the level coefficient of the corresponding data level;
[0068] Among them, it is assumed that I a1 Indicates the first recommendation index, Indicates that the level coefficient in the candidate data set is L x table; Indicates that the level coefficient in the candidate data set is L x The first recommendation index can be expressed by formula (1):
[0069]
[0070] Step S314: determining a second recommendation index according to the data delay of each second candidate data table in each candidate data group;
[0071] Among them, it is assumed that I a2 represents the second recommendation index, Delay represents the data delay of the second candidate data table, and the second recommendation index can be expressed by formula (2):
[0072]
[0073] Step S316: determining a target data group from the plurality of candidate data groups according to each of the first recommendation indexes and the corresponding second recommendation indexes;
[0074] Among them, it is assumed that I a represents the sum of the first recommendation index and the corresponding second recommendation index, then I a It can be expressed by formula (3):
[0075]
[0076] Step S318: Determine the matching degree between each business requirement and the corresponding target data set;
[0077] Step S320: Determine, among the various business requirements, the business requirements whose occurrence frequency and corresponding matching degree among the multiple business requirements meet specific conditions as asset optimization requirements;
[0078] Step S322: Optimize the data asset management model based on the asset optimization requirements.
[0079] In the embodiment of the present application, the target data group is determined from the multiple candidate data groups based on each first recommendation index and the corresponding second recommendation index, which can make the determination of the target data group more accurate.
[0080] The present invention also provides a data asset operation method, which includes the following steps:
[0081] Step S402: Acquire multiple business requirements;
[0082] Step S404: Filtering a first candidate data table from the data asset management model according to the target table-level metadata of each business requirement;
[0083] Step S406: based on the target field-level metadata of each business requirement, filter out a second candidate data table from the first candidate data table, and filter out candidate metadata from the second candidate data table;
[0084] Step S408: performing permutations and combinations on the second candidate data table to obtain multiple candidate data groups;
[0085] Step S410: determining the data level corresponding to each second candidate data table in each candidate data group;
[0086] Step S412: determining a first recommendation index according to the number of second candidate data tables in each data level and the level coefficient of the corresponding data level;
[0087] Step S414: determining a second recommendation index according to the data delay of each second candidate data table in each candidate data group;
[0088] Step S416: determining the candidate data group with the smallest sum of the first recommendation index and the second recommendation index as the target data group;
[0089] Among them, when selecting the target data group from the candidate data groups, two factors can be considered: data delay and data level. The shorter the data delay and the closer the data level is to the upper layer application, the higher the priority recommendation. Therefore, referring to formula (3), I a The smallest candidate data group is determined as the target data group.
[0090] Assume that the candidate data set G a ∈{Table1,Table2,Table3,Table4,Table5}, and Table1 and Table2 are tables in the data mart layer, Table3 is a table in the data warehouse layer, Table4 and Table5 are tables in the data application layer, and the data delays from Table1 to Table5 are Delay1 to Delay5, that is, the number of tables in the data warehouse layer is 1, and the number of tables in the data mart layer and the data application layer is 2, then the candidate data group G a Recommendation index I a It can be expressed by formula (4):
[0091] I a =L1*1+L2*2+L3*2+Delay1+Delay2+Delay3+Delay4+Delay5 (4);
[0092] Then you can a In the smallest case, G a Select as the target recommendation data group.
[0093] Step S418: Determine the matching degree between each business requirement and the corresponding target data set;
[0094] Step S420: Determine, among the various business requirements, the business requirement with the highest occurrence frequency and the lowest matching degree among the multiple business requirements as the asset optimization requirement;
[0095] Among them, the higher the frequency of occurrence of business needs or similar business needs, the more urgent it is to build an intermediate process table to meet such business needs; the greater the matching degree between each business need and the corresponding target data group, the greater the matching degree gap, indicating that the degree to which the tables in the existing MID (data mart layer) or APP layer (data application layer) meet the business needs is relatively low. Therefore, the business needs with the highest frequency and the lowest matching degree can be identified as asset optimization needs.
[0096] In addition, since the occurrence frequency of multiple business requirements may not be high, or the matching degree is not low, even if the business requirement with the highest occurrence frequency and the lowest matching degree is selected from multiple business requirements, the need to build an intermediate process table for the business requirement is not urgent enough, and the degree to which the tables in the existing MID (data mart layer) or APP layer (data application layer) meet business requirements is relatively high. Therefore, the selected asset optimization requirements do not contribute much to the optimization of the data asset management model. Therefore, it is possible to first determine whether there is at least one business requirement among the multiple business requirements whose occurrence frequency is greater than the frequency threshold or whose matching degree is less than the matching threshold. If so, the at least one business requirement is determined as the target business requirement, and the business requirement with the highest occurrence frequency and the lowest matching degree among the at least one target business requirement is determined as the asset optimization requirement.
[0097] Step S422: Based on the asset optimization requirement, a corresponding first data table is established;
[0098] The first data table may be established according to the target table-level metadata and target field-level metadata in the asset optimization requirement.
[0099] Step S424: adding the first data table to the data asset management model.
[0100] In an embodiment of the present application, by considering the two factors of data latency and data hierarchy, candidate data groups with short latency and data hierarchy close to the upper-level application are preferentially recommended. On the one hand, users can be recommended to use non-detailed data (i.e., data at the data mart layer or data application layer) as much as possible, which can avoid users from accessing sensitive detailed data at the data warehouse layer and ensure data security; on the other hand, the data reading efficiency can be improved; in addition, by determining the business needs with the highest frequency and the lowest matching degree as asset optimization needs, and establishing a first data table based on the asset optimization needs and adding it to the data asset management model, the matching degree between the business needs with urgent needs and the data in the data asset management model can be improved.
[0101] The present invention also provides a data asset operation method, which includes the following steps:
[0102] Step S502: Acquire multiple business requirements;
[0103] Step S504: Filtering a first candidate data table from the data asset management model according to the target table-level metadata of each business requirement;
[0104] Step S506: based on the target field-level metadata of each business requirement, filter out a second candidate data table from the first candidate data table, and filter out candidate metadata from the second candidate data table;
[0105] Step S508: performing permutations and combinations on the second candidate data table to obtain multiple candidate data groups;
[0106] Step S510: determining the data level corresponding to each second candidate data table in each candidate data group;
[0107] Step S512: determining a first recommendation index according to the number of second candidate data tables in each data level and the level coefficient of the corresponding data level;
[0108] Step S514: determining a second recommendation index according to the data delay of each second candidate data table in each candidate data group;
[0109] Step S516: sorting the plurality of candidate data groups according to each of the first recommendation indexes and the corresponding second recommendation indexes;
[0110] The plurality of candidate data groups may be sorted according to the sum of the first recommendation index and the second recommendation index, for example, according to I a Sort from smallest to largest.
[0111] Step S518: determining a plurality of first data groups from the plurality of candidate data groups according to the sorting result;
[0112] Among them, I can be selected from multiple candidate data groups a The top-ranked candidate data groups are used as the first data group.
[0113] Step S520: In response to the received selection instruction, determining a target data group from the plurality of first data groups;
[0114] The first data group can be used for user reference and selection; the user can perform a target data group selection operation on the first data group, and the processor generates a selection instruction after receiving the selection operation to determine the target data group from the first data group.
[0115] Step S522: Determine the matching degree between each business requirement and the corresponding target data set;
[0116] Step S524: Determine the priority of each business requirement based on the occurrence frequency and corresponding matching degree of each business requirement among the multiple business requirements;
[0117] Among them, the higher the frequency of occurrence and the lower the corresponding matching degree, the higher the priority of the corresponding business demand can be.
[0118] Step S526: determining the business requirements whose priorities meet specific conditions as asset optimization requirements;
[0119] Among them, the business needs can be sorted in order of priority from high to low, and several business needs with higher priority can be determined as asset optimization needs; in addition, the business needs can be sorted in order of priority from high to low, and several business needs with higher priority can be determined as target business needs. When the occurrence frequency of the target business need is greater than the frequency threshold, or the matching degree is less than the matching degree threshold, the target business need is determined as the asset optimization need.
[0120] Step S528: Based on the target table-level metadata and target field-level metadata of each of the at least one asset optimization requirement, update the first table-level metadata and first field-level metadata of the corresponding target data table in the data asset management model.
[0121] In an embodiment of the present application, the candidate data groups can also be sorted according to the first recommendation index and the second recommendation index, and the candidate data groups with higher sorting can be provided for user selection, thereby improving the autonomy and flexibility of the user in selecting the target data group; by determining the priority of the business needs according to the frequency of occurrence and matching degree of the business needs, and determining multiple asset optimization needs in order of priority, and updating the data asset management model based on the metadata in the multiple asset optimization needs, the diversity and flexibility of determining the asset optimization needs are improved, and the flexibility of the data asset management model update method is also improved.
[0122] Figure 2 This is a flow chart of a method for establishing a data asset management model according to an embodiment of the present application. Figure 2 As shown, the method includes:
[0123] Step 202: Based on the three elements of the terminal signaling data, data tables are established and stored in the data warehouse layer, the data mart layer, and the data application layer respectively; the three elements include user identification information, user location information, and time point;
[0124] Among them, the data warehouse layer is used to store the detailed data in the metadata of the terminal signaling data that has been preprocessed; the detailed data matches the user location information with the user identification information and the time point, or matches the user identification information with the user location information and the time point; the data mart layer is used to store the label data in the metadata of the terminal signaling data processed by the data warehouse layer; the label data is divided into user identification information, user location information and time point according to the subject; the data application layer is used to store the commonly used intermediate calculation process table in the metadata of the terminal signaling data processed by the data warehouse layer or the data mart layer; the data table can be used to represent the logical relationship between data, and the data table can be a logical table.
[0125] The user identification information can also be referred to as a person or a group of people, and can be various types of user IDs (Identity documents). The user location information can also be referred to as a place, and can be various types of area codes, such as province, city, and community codes. The time point can also be referred to as time, and can be either a moment or a time period, such as a holiday code (for example, the code corresponding to the Spring Festival), a report type code, etc.
[0126] Step 204: creating table-level metadata and field-level metadata corresponding to each data table according to the three elements of the terminal signaling data;
[0127] Step 206: Establish the data asset management model including the data warehouse layer, the data mart layer, and the data application layer.
[0128] In the embodiment of the present application, by establishing a data asset management model by layering and dividing the data tables, the metadata of the terminal signaling data can be managed and maintained more efficiently. Since a logical table is established, which reflects the logical relationship between the data, it can also help data demanders quickly understand the data.
[0129] The present invention also provides a method for establishing a data asset management model, the method comprising the following steps:
[0130] Step S602: In the data warehouse layer, multiple basic real-time data tables are established based on the user identification information and user location information of the terminal signaling data; an offline fact data table is established based on the user identification information, user location information, and time point of the terminal signaling data; and other data tables related to the basic real-time data table and the offline fact data table are established based on the basic real-time data table and the offline fact data table.
[0131] Step S604: In the data mart layer, data in the multiple basic real-time data tables, the offline fact data tables, and the other data tables in the data warehouse layer are aggregated to create an aggregated label data table including multiple label data;
[0132] Step S606: In the data application layer, the data in the multiple basic real-time data tables, the offline fact data table and the other data tables in the data warehouse layer, and / or the summary tag data table in the data mart layer are processed to create a summary granularity data table containing data of multiple granularities;
[0133] Step S608: creating table-level metadata and field-level metadata corresponding to each data table according to the three elements of the terminal signaling data;
[0134] The table-level metadata includes the level, business scenario, data timeliness and data latency; the field-level metadata includes category, scope and granularity.
[0135] Step S610: establishing the data asset management model including the data warehouse layer, the data mart layer and the data application layer.
[0136] In the embodiment of the present application, by placing the summary granularity data table into the data application layer, the summary tag data table into the data mart layer, and the basic real-time data table and the offline fact data table into the data warehouse layer, it is possible to help data demanders more conveniently retrieve the middle layer data from the data application layer, reduce repeated development costs, and shorten data processing time.
[0137] The present invention also provides a method for establishing a data asset management model, the method comprising the following steps:
[0138] Step S702: selecting at least one service scenario from pre-established service scenarios as a service scenario for each data table according to a search relationship between user identification information, user location information, and time points in the terminal signaling data;
[0139] Among them, the retrieval relationship can be to retrieve user location information based on user identification information and time point, or to retrieve user identification information based on user location information and time point, etc. The data table modeling can be performed respectively at the data warehouse layer, the data mart layer and the data application layer. Before performing data table modeling, business scenario modeling can also be performed. Business scenario modeling can enumerate six possible data application atomic scenarios (also called business scenarios) based on the three elements of "people-place-time", namely scenarios 1 to 6 in step S204.
[0140] Step S704: Based on the three elements of the terminal signaling data, data tables are respectively established and stored in the data warehouse layer, the data mart layer, and the data application layer;
[0141] Figure 3 This is a schematic diagram of a data table modeling of a data asset management model according to an embodiment of the present application; see Figure 3 In data table modeling, the data warehouse layer 301 considers the fast retrieval requirements of the common scenarios of [person+time:place] and [place+time:person], and generates two basic real-time tables based on user and location retrieval respectively: user trajectory sequence table 3011 and base station (cell) slice table 3012; the data mart layer 302 mainly stores various light summary tag tables with person or place as key value; the data application layer 303 stores summary tables of various granularities that may be called by most application needs to facilitate direct call by applications.
[0142] Step S706: Establishing a hierarchy of each data table according to the hierarchy of the user identification information, the hierarchy of the user location information, and the hierarchy of the time point of the terminal signaling data; the hierarchy being one of the data warehouse layer, the data mart layer, and the data application layer;
[0143] Step S708: selecting at least one business scenario from the business scenarios as a business scenario for each of the data tables according to a retrieval relationship among the user identification information, the user location information, and the time point of the terminal signaling data;
[0144] The business scenario may be one scenario from scenario 1 to scenario 6 or a combination of multiple scenarios.
[0145] Step S710: establishing a data aging of each data table according to the category to which the time point of the terminal signaling data belongs; the data aging is an offline table or a real-time table;
[0146] Among them, the data age can be an offline table or a real-time table; the categories of the time points include real-time and offline, which are used to indicate whether the data is processed in real time with immediate results or processed offline with batch results; the data delay can be the average processing time (also called average processing time) recorded by the Delay function from the original data to the current table.
[0147] Step S712: Establish the data delay of each data table according to the maximum value of the terminal signaling data at a time point; the data delay is the average processing time from the original data to the corresponding data table; the maximum value of the time point can represent the maximum value allowed for the data delay.
[0148] Step S714: establishing the person category, location category, time category and tag attribute of each data table according to the category of the user identification information, the category of the user location information and the category of the time point of the terminal signaling data;
[0149] The person category of the data table can be determined according to the category of the user identification information, the place category of the data table can be determined according to the category of the user location information, and the time category of the data table can be determined according to the category of the time point.
[0150] Step S716: establishing a person range, location range, time range and tag value for each data table according to the range of the user identification information, the range of the user location information and the range of the time point of the terminal signaling data;
[0151] Among them, the scope of the user identification information can be unlimited or a list of key personnel, the scope of the user location information can be the whole country, XX district, a shopping mall in YY district, and the scope of the time point can be the whole year of 2019, 7 days, etc.
[0152] Step S718: establishing the person granularity, location granularity and time granularity of each data table according to the minimum unit of user identification information, the minimum unit of user location information and the minimum unit of time point of the terminal signaling data.
[0153] The person granularity may be a person, the location granularity may be a base station cell, a district / county, etc., and the time granularity may be a day, an hour, etc.
[0154] It should be noted that each field in each table corresponds to a data element, and the granularity of data openness permissions can be controlled to the field level. Therefore, the metadata information of each field is very important for data users. In view of the business characteristics of mobile location data assets, the business metadata of the standard field (also called field-level metadata) may include: category, range and granularity; wherein the category can be a three-element keyword (for example, it can be a person, place or time), tag attribute, others, etc.; the range can be an area range, a time span (also called a time range), a person range (divided by the number's location, a certain type of attribute), a tag value (for example: occupational label, including label values such as students, civil servants, and teachers); the granularity can be location granularity, time granularity or person granularity. Table 2 is a zipper table of location signaling for all male users in the Jiangsu Province area in June 2020. The configurable field-level business metadata is shown in Table 2:
[0155] Table 2
[0156]
[0157]
[0158] As shown in Table 2, IMSI (International Mobile Subscriber Identity) is the SIM (Subscriber Identity Module) card number; MSISDN is the number a calling user needs to dial to call a mobile user in the GSM (Global System for Mobile Communications) PLMN (Public Land Mobile Network). It has the same function as a PSTN (Public Switched Telephone Network) number and is the number that uniquely identifies a mobile user in the public switched telephone network numbering plan, i.e., the called user's mobile phone number; From_country is the country code; From_Source is the place where the mobile phone number is registered, which is the first 7 digits of the valid mobile phone number; Longitude represents longitude; Latitude represents latitude; TAC (Tracking Area Code) is the TAC of the current cell; Cell ID (Cell Identity), also known as the cell identity, is the ECI (E-UTRAN) of the current cell. CellIdentifier (E-UTRAN cell unique identifier) determines the user location by identifying which cell in the network transmits the user call and translating this information into latitude and longitude. Geohash coding uses the geohash algorithm to encode longitude and latitude, converting two dimensions into one, and partitioning the address location. ProvinceCode represents the province code; City represents the city where the base station is located; Timestamp represents the timestamp of signaling generation; Day represents the day; PROCEDURE TYPE / Event ID (EventIdentity) represents the process type or event code; and Gen represents signaling source tracing.
[0159] Step S720: Establish the data asset management model including the data warehouse layer, the data mart layer and the data application layer.
[0160] In the embodiment of the present application, by establishing table-level metadata and field-level metadata corresponding to each data table according to the three elements of the terminal signaling data, the metadata of the terminal signaling data can be managed and maintained more efficiently.
[0161] In the related art, although some methods for managing mobile location data have been proposed, they have the following technical shortcomings:
[0162] First, the above method is limited to functional optimization at the database level, rather than data management at the data asset level.
[0163] Among them, the above methods are designed for time series and spatiotemporal databases, and belong to database-level functional optimization, focusing on solving one or a category of application requirements, rather than solving the problem of maximizing the value of data assets used by data asset managers in multiple scenarios.
[0164] Secondly, due to the uniqueness of mobile location data, it cannot solve the difficulties faced by data users in understanding the data.
[0165] Because mobile location data originates from the telecommunications industry, its field formats are unique and specialized, and the frequency intervals at which the data is generated are also highly relevant. The aforementioned methods focus on optimizing the physical storage of general spatiotemporal data, but lack the logical understanding of the data's inherent meaning. This approach fails to effectively help data users understand the data and rapidly develop and utilize it.
[0166] Furthermore, the above method cannot solve the problem of designing and using a large amount of intermediate-layer data in data asset management.
[0167] The above method cannot solve the design difficulties faced by data asset managers in data hierarchical management; it cannot help data asset managers decide which intermediate process data exists in the form of logical tables and which needs to be stored on disk; it cannot help data users choose which layer and which data can effectively and compliantly meet business needs and minimize data latency.
[0168] Finally, the above method cannot dynamically adjust the data asset management strategy based on data usage demand.
[0169] The above methods cannot help data asset managers dynamically and scientifically adjust data management strategies based on the data usage needs of upper-level applications and continuously improve data value.
[0170] In response to the above problems, the embodiment of the present application proposes a mobile location data asset operation method based on a dynamic feedback mechanism. This method establishes a set of data asset management models and data application demand collection, evaluation and optimization feedback models (also known as demand evaluation and data recommendation models) featuring mobile location data, focusing on management for maximizing the application value of data assets. The mobile location data asset operation method fits the three-element characteristics of mobile location signaling data: "people-place-time". In the construction of the data asset management model, it focuses on the virtuous cycle of "construction-service-feedback". The initial construction relies on the basic characteristics of location signaling. In the asset operation process, it emphasizes business demand-oriented services. By continuously evaluating the matching degree between new business needs and asset data, and continuously feeding back and optimizing the data asset management model, the goal of maximizing asset value is gradually achieved. At the same time, it serves more data users and achieves a win-win situation for social interests and asset managers.
[0171] The embodiment of the present application provides a method for operating mobile location data assets based on a dynamic feedback mechanism. This method for operating a dynamic feedback mechanism includes a dynamic mechanism for the mutual influence between the data asset management model and data application requirements (also known as business requirements). Under this mechanism, the data asset management model can be continuously optimized and constructed to meet the ever-increasing data application needs, reduce management costs, and improve the value of data utilization; data demanders can also quickly understand the data from the characteristics of mobile location data, and experience data recommendation services by themselves based on the guided demand collection process.
[0172] In the embodiments of the present application, for mobile location data scenarios, by building an effective data asset management model, data needs are automatically evaluated and the best data sharing solution is recommended; by collecting numerous data needs, the asset management model is continuously optimized to form a virtuous cycle operation mechanism of "construction-service-feedback".
[0173] Figure 4 This is a flow chart of a method for circular operation and management of data assets according to an embodiment of the present application; see Figure 4 The cyclic operation mechanism includes the following steps 401 to 403, which are closely linked through the "people-place-time" business metadata that conforms to the characteristics of mobile location data. Figure 5 for Figure 4 Detailed process diagram of the circular operations management method shown.
[0174] Step 401: constructing a data asset management model for mobile location data;
[0175] Among them, see Figure 5 , the step 401 may include steps 5011 to 5013:
[0176] Step 5011: Data assets are constructed in layers.
[0177] Based on the three key characteristics of mobile location data (person, place, and time), a three-tiered data asset management model is constructed: a data warehouse layer, a data mart layer, and a data application layer. The data warehouse layer is designed to store preprocessed and organized mobile location data, with two basic data structures, [person + time: place] and [place + time: person], serving as its primary components. [person + time: place] represents the use of person and time as primary keys to obtain a place, i.e., matching a place with a person and time. [place + time: person] represents the use of place and time as primary keys to obtain a person, i.e., matching a person with a place and time. The data mart layer is designed to process and generate data from the data warehouse layer. This data is aggregated and tagged data that can be directly accessed by applications, categorized into three themes: person, place, and time. The data application layer is designed to process and generate data from the data warehouse layer and / or the data mart layer. The data application layer stores intermediate calculation process tables for commonly used requirements.
[0178] At the data warehouse level, you can use enumeration to select any two of the three elements of mobile location data as primary keys to obtain the other element. However, because [Person + Location: Time] matches time with both person and location, this storage method is far inferior to the other two models in terms of retrieval and data compression. It also has limited application scenarios and is not suitable for basic data. Therefore, we do not recommend placing it at the data warehouse level. In summary, all basic data related to [Person + Time: Location] and [Location + Time: Person] can be placed at the data warehouse level, regardless of the database used. This database can be Hive, HBase, or Redis (Remote Dictionary Server).
[0179] For the data mart layer, mart data with people as the theme can be [person number: person attribute], where "person number" is the key field, such as various user IDs; person attributes can be multiple label fields, such as gender, age, occupation, permanent residence, etc.; mart data with location / region as the theme can be [location / region number: location / region attribute], where "location / region number" is the key field, such as various area codes; location / region attributes can be multiple label fields, such as comfort index, elderly activity area, road condition index, etc.; mart data with time as the theme can be [time period: time period attribute], which is mostly used in various report scenarios, where "time period" is the key field, such as holiday code (for example, the code corresponding to the Spring Festival), report type code, etc., and time period attributes can be multiple label fields, such as population migration index and population regional statistics.
[0180] The data application layer is a process of continuous development, evolving from scratch based on the richness of data assets exposed. Initially, the data application layer consisted of temporary tables that were retained to meet the needs of certain data applications. As the number of users of data assets gradually increased, it became apparent that the logic behind many of the requests was becoming consistent. Some requests even required the use of existing intermediate tables without requiring access to underlying data, significantly reducing development and usage costs. Therefore, consideration was given to making some of these temporary tables available in the data catalog for application use and assigning them to the data application layer.
[0181] Step 5012: Building a data model framework.
[0182] Among them, the construction of the data model framework can include business scenario modeling and data table modeling. Business scenario modeling can enumerate six possible data application atomic scenarios (also called business scenarios) based on the three elements of "people-place-time". In data table modeling, considering the fast retrieval requirements of the common scenarios of [people + time: place] and [place + time: people], the data warehouse layer generates two basic real-time tables based on users and locations for retrieval: user trajectory sequence table and base station (cell) slice table; the data mart layer mainly stores various light summary label tables with people or places as key values; the data application layer stores summary tables of various granularities that may be called by most application needs to facilitate direct calls by applications.
[0183] For business scenario modeling, the business scenarios may include: Scenario 1: Given a known area and time range, search for people (groups); Scenario 2: Given a known person (group) and time range, query for location / area; Scenario 3: Given a known person (group) and area range, query for time (time period special report); Scenario 4: Search for label attributes about people (gender, age, occupation, permanent residence); Scenario 5: Search for label attributes about areas (places); Scenario 6: Search for label attributes about the "people-land relationship" of the time period.
[0184] Any requirement can be met by combining one or more atomic scenarios. For example, if you want to know the hourly foot traffic in a certain business district, Scenario 1 can satisfy it. If you want to know the hourly foot traffic of young people in a certain business district, you need to use a combination of Scenario 1 and Scenario 4.
[0185] For data table modeling, according to the data layering design, the data warehouse layer includes real-time fact tables, offline fact tables, and related dimension tables. Among them, in order to meet the real-time demand for efficient retrieval of three elements, two basic real-time tables are designed based on user and location retrieval respectively: user trajectory sequence table and base station (cell) slice table, which meet the fast retrieval requirements of [person + time: location] and [location + time: person] scenarios. The offline fact table sets the retrieval for [person + time + location] and meets the requirements of scenarios 1, 2, and 3 at the same time. Due to the different time window sizes for pre-processing such as deduplication of raw data and base station ping-pong switching, offline fact tables have smaller storage space than real-time fact tables and can store data for a longer period of time. See Figure 3 , the offline fact table can be a location zipper table 3013; the related dimension table can be a base station working parameter dimension table 3014 and a custom area base station dimension table 315 related to the real-time fact table and the offline fact table.
[0186] The data mart layer primarily stores various lightly aggregated tag tables, keyed by person or location. A well-organized data mart layer is generated from the data warehouse layer's basic data through model calculations or by linking with other data sources. Each lightly aggregated tag table may contain one or more attribute tags. Therefore, the data mart layer serves as a hub for various types of tagged data.
[0187] The data application layer stores summary tables of various granularities that may be called by most application needs to facilitate direct calls by applications, such as summarizing at the regional level or summarizing at the time granularity.
[0188] Step 5013: Metadata specification construction.
[0189] The metadata specification includes table-level metadata models and field-level data meta-models. For the data tables in the data model framework, the metadata items of each data table are standardized and configured, while the business metadata of the data meta-fields is also standardized.
[0190] It should be noted that standard table-level metadata may include: the level to which it belongs, the business scenario, data timeliness, and standard latency.
[0191] Step 402: Data requirements collection and evaluation;
[0192] Among them, in order to meet the vision of maximizing data services, data managers need to be able to accurately collect various data requirements; based on the established data asset management model, first collect data requirements, evaluate the requirements, and then give recommended data.
[0193] Wherein, the step 402 may include the following steps 5021 to 5024:
[0194] Step 5021: Select an atomic scene;
[0195] Among them, the atomic scenarios to which the business requirements belong can be collected;
[0196] Step 5022: Determine data timeliness;
[0197] Among them, the timeliness requirements of the business needs can be collected;
[0198] Step 5023: Configure required parameters;
[0199] Among them, the demand parameters of business requirements can be collected;
[0200] Step 5024: Demand assessment and data recommendation;
[0201] Among them, an evaluation recommendation model can be adopted, which takes data requirements as input, automatically generates candidate data groups through data table screening and data element matching, and finally considers two factors, data delay and data hierarchy, to give combined ranked recommended candidate open data (also called target data group).
[0202] The step 5024 may include steps 50241 to 50244:
[0203] Step 50241: Data table screening;
[0204] Step 50242: data element matching;
[0205] Step 50243: Generate candidate data groups;
[0206] The tables selected in step 50243 can be freely arranged and combined to generate candidate data groups. All tables and fields in each candidate data group can meet the target requirements. A table can appear repeatedly in multiple candidate data groups. All possible candidate data groups are enumerated and recorded as G. n , where n is the number of candidate data groups. For a∈[1,n], each candidate data group can be represented as G a ∈{Table1,Table2,Table3,Tabl4e,Tabl5e}. At this point, any candidate data set provided to the data user can meet the target requirements, but which one to recommend still needs to be considered.
[0207] Step 50244: Select a target recommendation data set;
[0208] Among them, when selecting the target recommended data group from the candidate data groups, two factors can be considered: data latency and data level. The shorter the data latency and the closer the data level is to the upper application, the higher the priority recommendation. Therefore, the level coefficients of the three data levels of PDW (data warehouse layer), MID (data mart layer), and APP (data application layer) are configured as L1, L2, and L3 respectively. For G a , its recommendation index I a It is the sum of the hierarchical coefficients and data delays of all tables in the candidate data group. Recommendation index I a It can be expressed by formula (3):
[0209]
[0210] in, Indicates that the level coefficient in the candidate data set is L x table; Indicates that the level coefficient in the candidate data set is L x The number of tables, represents the cumulative sum of the products of the number of tables in each of the three data levels in the candidate data group multiplied by the level coefficient of the corresponding data level; Represents the cumulative sum of data delays of all tables in the candidate data group.
[0211] Assume that the candidate data set G a ∈{Table1,Table2,Table3,Table4,Table5}, and Table1 and Table2 are tables in the data mart layer, Table3 is a table in the data warehouse layer, Table4 and Table5 are tables in the data application layer, and the data delays from Table1 to Table5 are Delay1 to Delay5, that is, the number of tables in the data warehouse layer is 1, and the number of tables in the data mart layer and the data application layer is 2, then the candidate data group G a Recommendation index I a It can be expressed by formula (4):
[0212] I a =L1*1+L2*2+L3*2+Delay1+Delay2+Delay3+Delay4+Delay5 (4);
[0213] Among them, you can a In the smallest case, G a Select as the target recommended data group; you can also set I aSort from small to large, select the first few items as recommended data groups for user reference and selection, and the user selects the target recommended data group from the recommended data group; in order to recommend users to use non-detailed data as much as possible, the L1 coefficient value can generally be adjusted to a larger value. If you still need to apply for L1 layer data, you can start the sensitive data security approval mechanism.
[0214] Step 403: Feedback and data asset management model optimization;
[0215] The data manager can manage the data from the perspective of data assets, considering that the data will be used for multiple business needs rather than just one. Step 403 may include the following steps 5031 to 5033:
[0216] Step 5031: Statistics on matching degree of similar requirements;
[0217] The frequency of occurrence of each requirement can be collected to calculate the matching degree between the requirements and assets. Business requirements can be categorized based on the similarity of their requirement parameters, and the frequency of occurrence of each type of business requirement (i.e., frequency of occurrence) can be determined. The higher the frequency of occurrence of similar business requirements, the more urgent it is to build an intermediate process table that meets such business requirements. The matching degree between each business requirement and the corresponding target recommended data set can also be determined. The greater the matching degree gap, the lower the degree to which the tables in the existing MID (data mart layer) or APP layer (data application layer) meet the business requirements.
[0218] Assume that a business requirement item is D i , the target recommendation data group selected is G a , then you can traverse G a The granularity and range attributes of the three key data elements in each table are compared with the business requirements. If there is one item that is consistent, it is counted as 1, and the total count is finally M i As D i With G a When a new business requirement item matches the business requirement item D i If they are the same or similar, the two business requirements are considered to be of the same type, and the new business requirement is also considered to be the same as G a The cumulative matching count is in M i middle.
[0219] Step 5032: Sort asset optimization items based on business requirement matching and frequency of occurrence;
[0220] Among them, when statistics show that a certain type of business demand is not well matched with the tables in the existing MID (data mart layer) or APP layer (data application layer) and the business demand occurs frequently, it is necessary to consider optimizing the existing data tables in the MID or APP layer, or to supplement the construction of tables that directly meet the business demand. i The MID and APP layer data tables are constructed in descending order of priority. The construction content includes optimizing the existing tables in the MID and APP layers, updating data element parameters (such as table-level metadata or field-level metadata), and observing whether the overall demand matching situation is improved; or optimizing the data table model and introducing new table construction until the data assets reach the optimal state to meet all data demand retrieval.
[0221] Step 5033: Comprehensively consider the asset adjustment plan and impact.
[0222] Among them, before optimizing the data asset management model, it is necessary not only to evaluate the impact of the optimization and adjustment plan on the entire data asset management model, but also to evaluate the impact of each optimization adjustment on the required machine resources and related supporting facilities.
[0223] Every time a data table design is added or changed, the scope of impact can be deduced through table or record-level lineage analysis. Considering that the data on each production line may affect online business, metadata lineage analysis must be performed before any changes are made to assess the scope of data changes. Changes that may affect online business should be handled with caution, or adjustments should be made at the appropriate time.
[0224] Evaluating the impact of data asset design adjustments on machine resources and related supporting facilities generally requires leveraging an intelligent operations and maintenance monitoring system to statistically analyze data such as daily resource consumption. This allows for a rough estimate of the computing and storage space required for standard quantitative data. Based on this, the potential resource expansion or contraction resulting from data management adjustments can be calculated. Projects involving additional resources require reasonable budget requests. Projects with smaller expansions but significant post-adjustment benefits are more likely to receive financial support from the company.
[0225] In the embodiment of the present application, by sorting asset optimization items according to business demand matching and frequency of occurrence, it is possible to achieve the maximum satisfaction of the demand range, the most convenient data positioning and use, the most appropriate data element matching, the minimum data delay and the highest data quality, etc.
[0226] In the process of continuous demand collection and feedback, the construction of data asset management models will gradually maximize its value and serve asset managers and more data users.
[0227] The embodiment of the present application provides a data asset management model that is composed of a data layer design, a data model framework, and metadata specifications that conform to the characteristics of mobile location data.
[0228] The key points of data layering design include: using two basic data structures [people + time: location] and [location + time: people] as the main components of the data warehouse layer; using the three-element characteristics of mobile location data (people-location-time) to construct label data around the three themes of people, place, and time as the data mart layer; and the data storage application layer is designed to store intermediate calculation process tables retained for frequently used call requirements.
[0229] The key elements of the data model framework include: six atomic data application scenarios designed based on the three elements of mobile location data (person, place, and time) as business scenario models. The data table model is centered around a user trajectory sequence table and a base station (cell) slice table; surrounding it are various lightly aggregated tag tables keyed by person or place; and peripherally storing summary tables of various granularities that most applications may require, facilitating direct application access.
[0230] Key metadata specifications include: table-level metadata configuration items, including the level, business scenario, data timeliness, and standard latency. Field-level metadata configuration items include category (keywords for person, location, and time, tag attributes, and other), scope (regional scope, time span, and person scope), and granularity (location granularity, time granularity, and person granularity).
[0231] The present application provides a demand collection and evaluation model that complies with a dynamic feedback mechanism. The model uses the atomic scenarios of business requirements and the timeliness category parameters in the data timeliness to quickly narrow down the scope of data recommendations. Data element matching is used to screen table combinations that meet the requirements to generate candidate data groups. Finally, the candidate data groups are ranked based on the required data timeliness and data hierarchy.
[0232] The present embodiment provides a method for optimizing an asset management model for mobile location data. The method matches each demand parameter with the granularity and range attributes of the three key data elements, establishes a table comparing the degree of matching between data usage requirements (i.e., business requirements) and the current state of asset management, and prioritizes optimization of the data asset management model for certain types of requirements with low matching and high frequency, after comprehensively considering adjustment options and impacts.
[0233] It should be noted that the embodiments of this application focus on data management at the data asset level, rather than functional optimization at the database level. These embodiments can help data asset managers maximize the value of data assets across multiple application scenarios, rather than focusing on solving a single application requirement or category.
[0234] The embodiments of the present application can help data demanders efficiently understand and use data based on the uniqueness of mobile location data. Because mobile location data originates from the field of telecommunications operations, its field format has certain professional uniqueness, and the frequency intervals of data generation are also rich in practical significance. The design of metadata through the data asset management model helps data demanders quickly understand the data; it provides data demand collection and evaluation, and automatically recommends candidate data based on demand, helping data demanders quickly find suitable data for use.
[0235] In the embodiments of this application, the value of the retained data in the middle layer can be displayed, and the data value conversion rate can be improved. The embodiments of this application solve the design difficulties of data asset managers in data layer management; it can help data asset managers decide which intermediate process data should be stored in the form of logical tables and which should be stored on disk; it can help those who need similar data to have direct access to public middle layer data, reducing repeated development costs and shortening data processing time.
[0236] This embodiment of the application can dynamically adjust data asset management strategies based on data usage needs. Through the virtuous cycle operation mechanism of "build-service-feedback", this embodiment of the application helps data asset managers dynamically and scientifically adjust data management strategies based on the data usage needs of upper-level applications, thereby continuously improving data value.
[0237] Based on the foregoing embodiments, an embodiment of the present application provides an operating device for data assets, which includes the various modules included and can be implemented by a processor in an electronic device; of course, it can also be implemented by a specific logic circuit; in the implementation process, the processor can be a central processing unit (CPU), a microprocessor (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), etc.
[0238] Figure 6 This is a schematic diagram of the structure of the data asset operation device of the embodiment of the present application, such as Figure 6 As shown, the apparatus 600 includes an acquisition module 601, a screening module 602, a first determination module 603, a second determination module 604, and an optimization module 605, wherein:
[0239] Acquisition module 601, used to acquire multiple business requirements;
[0240] A screening module 602 is configured to screen a target data group corresponding to each business requirement from the data asset management model; the target data group includes the first table-level metadata and the first field-level metadata in the data asset management model;
[0241] A first determination module 603 is used to determine the matching degree between each business requirement and the corresponding target data group;
[0242] The second determining module 604 is configured to determine, among the various business requirements, the business requirements whose occurrence frequency and corresponding matching degree among the multiple business requirements meet specific conditions as asset optimization requirements;
[0243] The optimization module 605 is used to optimize the data asset management model based on the asset optimization requirements.
[0244] In one embodiment, the data asset management model includes a data warehouse layer, a data mart layer and a data application layer; the data warehouse layer, the data mart layer and the data application layer include multiple data tables; each data table includes table-level metadata and field-level metadata.
[0245] In one embodiment, the business requirements include target table-level metadata and target field-level metadata; the screening module 602 includes: a first screening submodule, used to screen out a first candidate data table from the data asset management model based on the target table-level metadata of each business requirement; a second screening submodule, used to screen out a second candidate data table from the first candidate data table and screen out candidate metadata from the second candidate data table based on the target field-level metadata of each business requirement; a combination submodule, used to arrange and combine the second candidate data tables to obtain multiple candidate data groups; and a third screening submodule, used to determine a target data group from the multiple candidate data groups based on the data latency and data level of each second candidate data table of each of the multiple candidate data groups.
[0246] In one embodiment, the third screening submodule includes: a first determination unit, used to determine the data level corresponding to each second candidate data table in each candidate data group; a second determination unit, used to determine the first recommendation index based on the number of second candidate data tables in each data level and the level coefficient of the corresponding data level; a third determination unit, used to determine the second recommendation index based on the data delay of each second candidate data table in each candidate data group; a screening unit, used to determine the target data group from the multiple candidate data groups based on each first recommendation index and the corresponding second recommendation index.
[0247] In one embodiment, the screening unit comprises:
[0248] The first determining subunit is configured to determine the candidate data group with the smallest sum of the first recommendation index and the second recommendation index as the target data group.
[0249] In one embodiment, the screening unit includes: a sorting subunit, used to sort the multiple candidate data groups according to each first recommendation index and the corresponding second recommendation index; a first screening subunit, used to determine multiple first data groups from the multiple candidate data groups based on the sorting results; and a second screening subunit, used to determine a target data group from the multiple first data groups in response to a received selection instruction.
[0250] In one embodiment, the second determining module 604 includes: a first determining submodule, configured to determine the business requirement with the highest occurrence frequency and the lowest matching degree among the multiple business requirements as the asset optimization requirement.
[0251] In one embodiment, the second determination module 604 includes: a second determination submodule, used to determine the priority of the corresponding business requirement based on the frequency of occurrence and the corresponding matching degree of each business requirement among the multiple business requirements; and a third determination submodule, used to determine the business requirement whose priority meets specific conditions as an asset optimization requirement.
[0252] In one embodiment, the business requirement includes target table-level metadata and target field-level metadata; the optimization module 605 includes: an establishment submodule for establishing a corresponding first data table based on each of the at least one asset optimization requirement; and an addition submodule for adding the at least one first data table to the data asset management model;
[0253] In one embodiment, the optimization module 605 includes: an update submodule, which is used to update the table-level metadata and field-level metadata of the corresponding target data table in the data asset management model based on the target table-level metadata and target field-level metadata of each asset optimization requirement in the at least one asset optimization requirement.
[0254] In one embodiment, the device also includes: a first establishment module, used to establish data tables stored in the data warehouse layer, data mart layer and data application layer respectively; a second establishment module, used to establish table-level metadata and field-level metadata corresponding to each of the data tables; and a third establishment module, used to establish the data asset management model including the data warehouse layer, the data mart layer and the data application layer.
[0255] Figure 7 This is a schematic diagram of the structure of the device for establishing the data asset management model according to the embodiment of the present application. Figure 7 As shown, the apparatus 700 includes a first establishing module 701, a second establishing module 702, and a third establishing module 703, wherein:
[0256] A first establishing module 701 is configured to generate data tables stored in a data warehouse layer, a data mart layer, and a data application layer, respectively, based on three elements of terminal signaling data; the three elements including user identification information, user location information, and time point;
[0257] The second establishment module 702 is used to establish table-level metadata and field-level metadata corresponding to each data table based on the three elements of the terminal signaling data, wherein the data warehouse layer is used to store detailed data in the metadata of the pre-processed terminal signaling data; the detailed data matches user location information by user identification information and time point, or matches user identification information by user location information and time point; the data mart layer is used to store label data in the metadata of the terminal signaling data processed by the data warehouse layer; the label data is divided into user identification information, user location information and time point by subject; the data application layer is used to store commonly used intermediate calculation process tables in the metadata of the terminal signaling data processed by the data warehouse layer or the data mart layer;
[0258] The third establishing module 703 is used to establish the data asset management model including the data warehouse layer, the data mart layer and the data application layer.
[0259] In one embodiment, the first establishment module 701 includes: a first establishment submodule, used to establish multiple basic real-time data tables in the data warehouse layer based on user identification information and user location information of terminal signaling data; establish an offline fact data table based on the user identification information, user location information and time point of the terminal signaling data; establish other data tables related to the basic real-time data table and the offline fact data table based on the basic real-time data table and the offline fact data table; a second establishment submodule, used to summarize and process the data in the multiple basic real-time data tables, the offline fact data tables and the other data tables in the data warehouse layer in the data mart layer in the data mart layer, and establish a summarized label data table including multiple label data; a third establishment submodule, used to process the data in the multiple basic real-time data tables, the offline fact data tables and the other data tables in the data warehouse layer, and / or the summarized label data table in the data mart layer in the data application layer, and establish a summarized granularity data table including multiple granularity data.
[0260] In one embodiment, the table-level metadata includes the level, business scenario, data timeliness and data delay; the field-level metadata includes category, range and granularity; the second establishment module 702 includes: a fourth establishment submodule, which is used to establish the level of each data table according to the level of the user identification information of the terminal signaling data, the level of the user location information and the level of the time point; the level is one of the data warehouse layer, the data mart layer and the data application layer; a fifth establishment submodule, which is used to select at least one business scenario from the pre-established business scenarios as the business scenario of each data table according to the retrieval relationship between the user identification information, the user location information and the time point of the terminal signaling data; a sixth establishment submodule, which is used to establish the data timeliness of each data table according to the category of the time point of the terminal signaling data; the data timeliness is an offline table or a real-time table ; The seventh establishment submodule is used to establish the data delay of each of the data tables according to the maximum value of the time point of the terminal signaling data; the data delay is the average processing time from the original data to the corresponding data table; the eighth establishment submodule is used to establish the person category, place category, time category and label attribute of each of the data tables according to the category of the user identification information of the terminal signaling data, the category of the user location information and the category of the time point; the ninth establishment submodule is used to establish the person range, place range, time range and label value of each of the data tables according to the range of the user identification information of the terminal signaling data, the range of the user location information and the range of the time point; the tenth establishment submodule is used to establish the person granularity, place granularity and time granularity of each of the data tables according to the minimum unit of the user identification information of the terminal signaling data, the minimum unit of the user location information and the minimum unit of the time point.
[0261] It should be noted that, in the embodiment of the present application, if the above-mentioned data asset operation method is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the relevant technology can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable an electronic device (which can be a mobile phone, tablet computer, desktop computer, personal digital assistant, navigator, digital phone, video phone, television, sensor device, etc.) to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a U disk, a mobile hard disk, a read-only memory (ROM), a magnetic disk or an optical disk. In this way, the embodiment of the present application is not limited to any specific combination of hardware and software.
[0262] The description of the above device embodiment is similar to the description of the above method embodiment and has similar beneficial effects as the method embodiment. For technical details not disclosed in the device embodiment of this application, please refer to the description of the method embodiment of this application for understanding.
[0263] Correspondingly, an embodiment of the present application provides an electronic device, Figure 8 This is a hardware entity diagram of an electronic device according to an embodiment of the present application, such as Figure 8 As shown, the hardware entity of the electronic device 800 includes: a memory 801 and a processor 802, wherein the memory 801 stores a computer program that can be run on the processor 802, and when the processor 802 executes the program, it implements the steps in the data asset operation method or the data asset management model establishment method of the above-mentioned embodiment.
[0264] The memory 801 is configured to store instructions and applications executable by the processor 802, and can also cache data to be processed or processed by the processor 802 and various modules in the electronic device 800 (for example, image data, audio data, voice communication data and video communication data), which can be implemented through flash memory (FLASH) or random access memory (Random Access Memory, RAM).
[0265] Correspondingly, an embodiment of the present application provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, it implements the steps in the data asset operation method or the data asset management model establishment method provided in the above embodiments.
[0266] It should be noted that the description of the above storage medium and device embodiments is similar to the description of the above method embodiments, and has similar beneficial effects as the device embodiments. For technical details not disclosed in the storage medium and method embodiments of this application, please refer to the description of the device embodiments of this application for understanding.
[0267] It should be understood that "one embodiment" or "an embodiment" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. The above-mentioned serial numbers of the embodiments of the present application are for description only and do not represent the advantages and disadvantages of the embodiments.
[0268] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0269] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.
[0270] The units described above as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units; they may be located in one place or distributed across multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. In addition, the functional units in the various embodiments of the present application may all be integrated into one processing unit, or each unit may be separately used as a unit, or two or more units may be integrated into one unit; the above-mentioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.
[0271] Those skilled in the art will appreciate that all or part of the steps in implementing the above-mentioned method embodiments can be accomplished by hardware associated with program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, the steps of the above-mentioned method embodiments are executed; and the aforementioned storage medium includes various media that can store program codes, such as a mobile storage device, a read-only memory (ROM), a magnetic disk, or an optical disk. Alternatively, if the above-mentioned integrated unit of the present application is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application, or the part that contributes to the relevant technology, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for causing a computer device (which can be a mobile phone, tablet computer, desktop computer, personal digital assistant, navigator, digital phone, video phone, television, sensor device, etc.) to execute all or part of the methods described in each embodiment of the present application. And the aforementioned storage medium includes various media that can store program codes, such as a mobile storage device, a ROM, a magnetic disk, or an optical disk.
[0272] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined, if they do not conflict, to obtain new method embodiments. The features disclosed in the several product embodiments provided in this application can be arbitrarily combined, if they do not conflict, to obtain new product embodiments. The features disclosed in the several method or device embodiments provided in this application can be arbitrarily combined, if they do not conflict, to obtain new method embodiments or device embodiments.
[0273] The above is merely an embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A data asset operation method, characterized in that: The method comprises: Acquire multiple business requirements; the business requirements include target table-level metadata and target field-level metadata; Filtering a first candidate data table from the data asset management model according to the target table-level metadata of each business requirement; According to the target field-level metadata of each business requirement, a second candidate data table is screened out from the first candidate data table, and candidate metadata is screened out from the second candidate data table; Performing permutations and combinations on the second candidate data table to obtain a plurality of candidate data groups; determining a target data group from the plurality of candidate data groups based on the data latency and data hierarchy of each second candidate data table of each candidate data group; the target data group including the first table-level metadata and the first field-level metadata in the data asset management model; Determining the degree of match between each of the business requirements and the corresponding target data set; Determine the business demand with the highest occurrence frequency and the lowest matching degree among the multiple business demands as the asset optimization demand; or, determine the priority of the corresponding business demand based on the occurrence frequency and corresponding matching degree of each business demand among the multiple business demands; determine several business demands with higher priorities as asset optimization demands; or, determine several business demands with higher priorities as target business demands, and if the occurrence frequency of the target business demand is greater than a frequency threshold, or the matching degree is less than a matching degree threshold, determine the target business demand as the asset optimization demand; Based on the asset optimization requirements, the data asset management model is optimized.
2. The method according to claim 1, characterized in that The target table-level metadata includes the target level, target business scenario, and target data timeliness of the business requirement; the target field-level metadata includes the target category, target range, and target granularity; The selecting a first candidate data table from the data asset management model based on the target table-level metadata of each business requirement includes: selecting a first candidate data table from the data asset management model based on the target level, target business scenario, and target data timeliness of each business requirement; The filtering out a second candidate data table from the first candidate data table and filtering out candidate metadata from the second candidate data table based on the target field-level metadata of each business requirement includes: filtering out a second candidate data table from the first candidate data table and filtering out candidate metadata from the second candidate data table based on the target category, target range, and target granularity of each business requirement; the candidate metadata is metadata in the second candidate data table that is related to the target category, target range, and target granularity.
3. The method according to claim 1, characterized in that Determining the target data group from the plurality of candidate data groups according to the data delay and data level of each second candidate data table of each candidate data group includes: Determining the data level corresponding to each second candidate data table in each candidate data group; determining a first recommendation index according to the number of second candidate data tables in each of the data levels and the level coefficient of the corresponding data level; determining a second recommendation index according to the data delay of each second candidate data table in each candidate data group; A target data group is determined from the plurality of candidate data groups according to each of the first recommendation indexes and the corresponding second recommendation indexes.
4. The method according to claim 3, characterized in that Determining a target data group from the plurality of candidate data groups according to each of the first recommendation indexes and the corresponding second recommendation indexes includes: Determine the candidate data group with the smallest sum of the first recommendation index and the second recommendation index as the target data group; or, The plurality of candidate data groups are sorted according to each of the first recommendation indexes and the corresponding second recommendation indexes; a plurality of first data groups are determined from the plurality of candidate data groups according to the sorting results; and a target data group is determined from the plurality of first data groups in response to the received selection instruction.
5. The method according to claim 1, characterized in that Optimizing the data asset management model based on the asset optimization requirement includes: Based on the asset optimization requirements, a first data table is established; and the first data table is added to the data asset management model; or, Based on the target table-level metadata and target field-level metadata required by the asset optimization, first table-level metadata and first field-level metadata of the target data table in the data asset management model are updated.
6. The method according to any one of claims 1 to 4, characterized in that The method further comprises: Generate data tables stored in the data warehouse layer, data mart layer, and data application layer based on the three elements of terminal signaling data; the three elements include user identification information, user location information, and time point; Based on the three elements of the terminal signaling data, table-level metadata and field-level metadata corresponding to each data table are established, wherein the data warehouse layer is used to store detailed data in the metadata of the pre-processed terminal signaling data; the detailed data matches user location information with user identification information and time point, or matches user identification information with user location information and time point; the data mart layer is used to store label data in the metadata of the terminal signaling data processed by the data warehouse layer; the label data is divided into user identification information, user location information and time point by subject; the data application layer is used to store commonly used intermediate calculation process tables in the metadata of the terminal signaling data processed by the data warehouse layer or the data mart layer; The data asset management model including the data warehouse layer, the data mart layer and the data application layer is established.
7. The method according to claim 6, characterized in that The data tables stored in the data warehouse layer, data mart layer and data application layer are established based on the three elements of the terminal signaling data, including: In the data warehouse layer, multiple basic real-time data tables are established based on user identification information and user location information of the terminal signaling data; an offline fact data table is established based on the user identification information, user location information and time point of the terminal signaling data; and other data tables related to the basic real-time data table and the offline fact data table are established based on the basic real-time data table and the offline fact data table; In the data mart layer, data in the multiple basic real-time data tables, the offline fact data tables, and the other data tables in the data warehouse layer are aggregated to create an aggregated label data table including multiple label data; In the data application layer, the data in the multiple basic real-time data tables in the data warehouse layer, the offline fact data table and the other data tables, and / or the summary label data table in the data mart layer are processed to establish a summary granularity data table containing data of multiple granularities.
8. The method according to claim 6, characterized in that The table-level metadata includes the level, business scenario, data timeliness, and data latency; the field-level metadata includes category, scope, and granularity; The step of establishing table-level metadata and field-level metadata corresponding to each data table according to the three elements of the terminal signaling data includes: Establishing a hierarchy of each data table according to the hierarchy of the user identification information, the hierarchy of the user location information, and the hierarchy of the time point of the terminal signaling data; the hierarchy being one of the data warehouse layer, the data mart layer, and the data application layer; selecting at least one service scenario from pre-established service scenarios as the service scenario of each data table according to a retrieval relationship among user identification information, user location information, and time points of the terminal signaling data; Establishing a data aging of each data table according to the category to which the time point of the terminal signaling data belongs; the data aging is an offline table or a real-time table; Establishing a data delay for each data table based on the maximum value of the time point of the terminal signaling data; the data delay is the average processing time from the original data to the corresponding data table; Establishing the person category, location category, time category and tag attributes of each data table according to the category of the user identification information, the category of the user location information and the category of the time point of the terminal signaling data; Establishing a person range, location range, time range and tag value for each data table according to the range of user identification information, user location information and time point of the terminal signaling data; The person granularity, location granularity and time granularity of each data table are established according to the minimum unit of user identification information, the minimum unit of user location information and the minimum unit of time point of the terminal signaling data.
9. A data asset operation device, characterized in that: The device comprises: An acquisition module is used to acquire multiple business requirements; the business requirements include target table-level metadata and target field-level metadata; a screening module configured to screen out a first candidate data table from the data asset management model based on the target table-level metadata of each business requirement; screen out a second candidate data table from the first candidate data table and screen out candidate metadata from the second candidate data table based on the target field-level metadata of each business requirement; permutate and combine the second candidate data tables to obtain a plurality of candidate data groups; determine a target data group from the plurality of candidate data groups based on the data latency and data hierarchy of each second candidate data table in each of the plurality of candidate data groups; the target data group includes the first table-level metadata and the first field-level metadata in the data asset management model; A first determination module is used to determine the matching degree between each of the business requirements and the corresponding target data group; The second determination module is configured to determine the business requirement with the highest occurrence frequency and the lowest matching degree among the multiple business requirements as the asset optimization requirement; or, based on the occurrence frequency and corresponding matching degree of each business requirement among the multiple business requirements, determine the corresponding business requirements with higher priorities as the asset optimization requirements; or, determine the business requirements with higher priorities as target business requirements, and if the occurrence frequency of the target business requirements is greater than a frequency threshold, or the matching degree is less than a matching degree threshold, determine the target business requirements as the asset optimization requirements; The optimization module is used to optimize the data asset management model based on the asset optimization requirements.
10. An electronic device comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, wherein: When the processor executes the program, the steps in the data asset operation method described in any one of claims 1 to 8 are implemented.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps in the data asset operation method described in any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Data asset management method and device
CN112395340A
Data management method and device, electronic equipment and readable storage medium
CN113127455A