Data recommendation method, device, equipment and computer storage medium
By constructing data profiles and recommendation models, and based on the full lifecycle characteristics of metadata, the problem of low efficiency and accuracy in wireless network data utilization has been solved, achieving efficient data support for business scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHINA MOBILE GROUP DESIGN INST
- Filing Date
- 2022-06-29
- Publication Date
- 2026-04-24
AI Technical Summary
Existing technologies have low efficiency and accuracy in utilizing wireless network data, and relying on expert experience makes it difficult to meet the increasingly detailed business analysis needs.
By identifying business scenarios, a recommendation model is used to construct data profiles based on the full lifecycle features of metadata, according to the correlation between multiple data profiles and optional scenarios. Cluster analysis and lineage analysis are then performed to determine the target data profile and recommend business support data.
It improved the efficiency and accuracy of data utilization, and enabled the massive wireless network data items to effectively support business scenarios.
Smart Images

Figure CN117349536B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the field of data processing technology, specifically to a data recommendation method, apparatus, device, and computer storage medium. Background Technology
[0002] The operation of wireless networks has accumulated a wealth of data. Operators manage this data in three domains based on its source and business focus: the operations domain, the business domain, and the management domain. These three domains are isolated from each other. Currently, when data application personnel conduct business analysis on wireless networks, they need to extract data from each domain based on expert experience to create datasets that support business scenarios.
[0003] In the process of implementing the embodiments of the present invention, the inventors of the present invention discovered that existing data processing methods have the problem of low efficiency and accuracy in data utilization. Summary of the Invention
[0004] In view of the above problems, embodiments of the present invention provide a data recommendation method to solve the problem of low data utilization efficiency and accuracy in business development in the prior art.
[0005] According to one aspect of the present invention, a data recommendation method is provided, the method comprising:
[0006] Define the business scenario;
[0007] The target data profile corresponding to the business scenario is determined based on the recommendation model; the recommendation model is determined based on the association between multiple data profiles and optional scenarios; wherein, the data profile is used to describe the characteristics of a data item throughout its entire lifecycle; the data profile and the association are determined based on the metadata corresponding to the data item.
[0008] Based on the target data profile, determine the business support data corresponding to the business scenario.
[0009] In one alternative approach, the metadata corresponds to a plurality of the data items; the method further includes:
[0010] Feature extraction is performed on the metadata to obtain the data attribute features and operation timing features corresponding to each data item;
[0011] The data profile is obtained by arranging and combining the data attribute features according to the operation timing features.
[0012] In one alternative approach, the metadata includes business domain metadata, technical domain metadata, and operational domain metadata; the data attribute features include business attribute features and technical attribute features; the method further includes:
[0013] The business features are obtained by extracting features from the business domain metadata.
[0014] The technical features are obtained by extracting features from the metadata of the technical domain.
[0015] The operation domain metadata is subjected to feature extraction to obtain the operation timing features.
[0016] In an alternative approach, the method further includes:
[0017] Cluster analysis is performed on the data profiles to obtain multiple profile classes and at least one class feature corresponding to each profile class;
[0018] The association between each class feature and the optional scenario is determined based on the lineage of the metadata;
[0019] The association between the data profile and the optional scenarios is determined based on the association between each of the aforementioned class features and the optional scenarios.
[0020] In an alternative approach, the method further includes:
[0021] The co-occurrence scenario information for each of the aforementioned class features is determined based on the blood relationship; the co-occurrence scenario is at least one of the optional scenarios;
[0022] The association between each class feature and the optional scenario is determined based on the co-occurrence scenario information.
[0023] In an alternative approach, the method further includes:
[0024] The similarity information between each of the class features is determined based on the co-occurrence scene information;
[0025] The association between each of the class features and the optional scene is determined based on the similarity information.
[0026] In an alternative approach, the method further includes:
[0027] Determine the initial matching weights between the class features and the optional scenarios;
[0028] The distance between each of the class features is calculated based on the similarity information;
[0029] The initial matching weights are adjusted based on the distance to obtain the association relationship.
[0030] According to another aspect of the present invention, a data recommendation apparatus is provided, comprising:
[0031] The first determination module is used to determine the business scenario;
[0032] The recommendation module is used to determine the target data profile corresponding to the business scenario based on the recommendation model; the recommendation model is determined based on the association between multiple data profiles and optional scenarios; wherein, the data profile is used to describe the characteristics of a data item throughout its entire lifecycle; the data profile and the association are determined based on the metadata corresponding to the data item.
[0033] The second determining module is used to determine the business support data corresponding to the business scenario based on the target data profile.
[0034] According to another aspect of the present invention, a data recommendation device is provided, comprising:
[0035] The processor, memory, communication interface, and communication bus are provided, wherein the processor, memory, and communication interface communicate with each other via the communication bus.
[0036] The memory is used to store at least one executable instruction that causes the processor to perform the operation of the data recommendation method as described in any one of the above.
[0037] According to another aspect of the present invention, a computer-readable storage medium is provided, the storage medium storing at least one executable instruction that causes a data recommendation device to perform operations of the data recommendation method as described in any one of the present inventions.
[0038] This invention, through the following embodiments, determines a business scenario; determines a target data profile corresponding to the business scenario based on a recommendation model; the recommendation model is determined based on the association between multiple data profiles and optional scenarios; wherein, the data profile is used to describe the characteristics of a data item throughout its entire lifecycle; the data profile and the association are determined based on the metadata corresponding to the data item; and the business support data corresponding to the business scenario is determined based on the target data profile. This differs from existing technologies that rely on expert experience to determine which data items to use and what operations to perform to support the development and analysis of business scenarios, resulting in low efficiency and accuracy in data utilization. This invention describes the characteristics of the entire lifecycle of data based on metadata, and further trains a recommendation model based on the association between the data profile and optional business development scenarios. This allows the target data profile corresponding to the current business scenario to be determined based on the recommendation model, and finally, the corresponding business support data is determined based on the target data profile. This achieves data support for the business scenario, improves the efficiency and accuracy of data utilization, and enables the effective support of massive wireless network data items for business scenario applications.
[0039] The above description is merely an overview of the technical solutions of the embodiments of the present invention. In order to better understand the technical means of the embodiments of the present invention and to implement them in accordance with the contents of the specification, and to make the above and other objects, features and advantages of the embodiments of the present invention more apparent and understandable, specific embodiments of the present invention are described below. Attached Figure Description
[0040] The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:
[0041] Figure 1 A flowchart illustrating the data recommendation method provided in an embodiment of the present invention is shown;
[0042] Figure 2 A flowchart illustrating a data recommendation method provided in another embodiment of the present invention is shown;
[0043] Figure 3 A schematic diagram illustrating the data profile construction process in a data recommendation method provided by another embodiment of the present invention is shown.
[0044] Figure 4 A schematic diagram of the data recommendation device provided in an embodiment of the present invention is shown;
[0045] Figure 5 A schematic diagram of the data recommendation device provided in an embodiment of the present invention is shown. Detailed Implementation
[0046] Exemplary embodiments of the invention will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the invention are shown in the drawings, it should be understood that the invention can be implemented in various forms and should not be limited to the embodiments set forth herein.
[0047] Before describing the data recommendation method of this invention, the prior art and its problems will be explained.
[0048] The operation of wireless networks has accumulated a wealth of data, which, while facilitating business analysis, also presents significant challenges. Each data domain contains numerous data items, each with multiple dimensions. The relationships between data within and across data domains are difficult to clarify using traditional methods, especially as 5G networks bring finer data granularity and increased data dimensions, leading to more application scenarios. Existing methods, relying on expert experience to determine which data items and operations to use to support business scenario analysis, and requiring collaboration across multiple departments for data metrics, are ill-suited to increasingly refined analytical needs. Furthermore, traditional data management separates storage, processing, and application, with different departments responsible for different data processes. Data users in business scenarios lack a holistic understanding of the data, and their knowledge of data sources, processing methods, and application algorithms relies heavily on expert experience. Therefore, current methods result in low efficiency and accuracy when applying data to business scenarios.
[0049] Figure 1 A flowchart illustrating a data recommendation method provided in an embodiment of the present invention is shown. This method is executed by a computer processing device. The computer processing device may include a mobile phone, a laptop computer, etc. Figure 1 As shown, the method includes the following steps:
[0050] Step 10: Determine the business scenario.
[0051] In one embodiment of the present invention, a business scenario corresponds to a scenario of providing services or development tasks, such as an indoor coverage traffic prediction scenario or a user recommendation scenario.
[0052] Specifically, a preset data search service interface can be displayed, and natural language processing and semantic analysis can be performed on the information entered in the service interface to obtain target keywords, and business scenarios can be determined based on the target keywords.
[0053] Step 20: Determine the target data profile corresponding to the business scenario based on the recommendation model; the recommendation model is determined based on the association between multiple data profiles and optional scenarios; wherein, the data profile is used to describe the characteristics of the data item throughout its entire lifecycle; the data profile and the association are determined based on the metadata corresponding to the data item.
[0054] In one embodiment of the present invention, the recommendation model is used to recommend data profiles corresponding to business scenarios, so that data supporting the business scenario can be found based on the data profiles, thereby enabling the development, provision or optimization of the business scenario.
[0055] Considering that highly relevant data is needed to support business development, and metadata contains a large amount of data, metadata can be processed to obtain supporting data corresponding to various business scenarios, thereby improving the efficiency and accuracy of business development.
[0056] Furthermore, considering that business development is generally sequential and iterative—that is, the entire process of a data item in the metadata from its creation to its use and eventual destruction—can include one or more stages such as integration, cleaning, transformation, processing, analysis, use, archiving, and destruction. Here, a data item refers to a dimension of the data, such as a field like user identifier or access permissions. Therefore, the system development process can be characterized based on the full lifecycle characteristics of the data. To prevent missing scenario-related data and improve the accuracy of the recommendation model, the full lifecycle profile of the data and its relationship with the scenario can be determined based on the most comprehensive metadata.
[0057] Specifically, data features can be extracted from metadata first. Based on these extracted features, data profiles can be constructed for each data item. Then, the data profiles can be categorized according to the historical support scenarios included within them, thus revealing the relationships between multiple data profiles and selectable scenarios. Specifically, data extraction from metadata can be performed from three dimensions: business attributes, technical attributes, and historical operations. This allows data profiles to reflect what the data is (technical attributes), what it is used for (business attributes), and how it has been used historically (historical operations), enabling data recommendations for business support scenarios based on these data profiles.
[0058] Therefore, in another embodiment of the present invention, the metadata corresponds to multiple data items. Before step 20, the method further includes: step 201: extracting features from the metadata to obtain data attribute features and operation timing features corresponding to each data item.
[0059] In one embodiment of the present invention, data attribute features are used to characterize the physical attributes and business attributes of data items. Physical attributes characterize the relevant definitions of data at the physical storage and retrieval level, such as field definitions, field attributes, access permissions, and ETL (Extract-Transform-Load, describing the process of extracting, transforming, and loading data from the source to the destination) job details in the physical database. Business attributes include business rules for the data item, stakeholder information, security levels, and data usage instructions. Business attributes are set according to the needs of the business system. Stakeholders include data owners, data managers, and data users, and stakeholder information may include contact information.
[0060] In one embodiment of the present invention, the operation timing feature is used to characterize the characteristics of the historical operation time sequence of the data item during use, such as the operation time, operation type, operation parameters and operator information corresponding to each historical operation.
[0061] In the field of mobile communications, when conducting data governance, metadata can be divided into business domain metadata, technical domain metadata, and operational domain metadata. Therefore, data features under preset dimensions can be extracted from the metadata of the corresponding domain to obtain data attribute features and operational timing features.
[0062] Therefore, in another embodiment of the present invention, the metadata includes business domain metadata, technical domain metadata, and operational domain metadata; the data attribute features include business attribute features and technical attribute features; step 201 further includes:
[0063] Step 2011: Extract features from the business domain metadata to obtain the business features.
[0064] In one embodiment of the present invention, business features are obtained by extracting features from business domain metadata based on preset business-related dimensions. These business-related dimensions include business rules, data stakeholder information, security levels, and usage instructions.
[0065] Step 2012: Extract features from the technical domain metadata to obtain the technical features;
[0066] In one embodiment of the present invention, technical features are obtained by extracting features from the technical domain metadata based on preset technical-related dimensions. These technical-related dimensions include data storage field definitions, field attributes, access permissions, and data warehouse operation information.
[0067] Step 2013: Extract features from the operation domain metadata to obtain the operation timing features.
[0068] In one embodiment of the present invention, feature extraction of operation-related dimensions is performed on the operation domain metadata to obtain operation time sequence features. The operation-related dimensions include operation parameter information under at least one optional operation type corresponding to each optional stage in the data usage process. The optional stages include at least one of data integration, cleaning, transformation, processing, analysis, use, archiving, and destruction. The optional operation types can be types such as replacement, merging, and calling.
[0069] Step 202: Arrange and combine the data attribute features according to the operation timing features to obtain the data profile.
[0070] In one embodiment of the present invention, the time sequence of each historical operation performed on the data attribute features and the corresponding operation parameters are determined based on the operation time sequence characteristics. The operation information corresponding to each data attribute feature is then arranged sequentially according to this time sequence to obtain a data profile. Since the data profile includes information corresponding to each operation of the data attribute features during the actual historical data usage process, and the data attribute features characterize the business functions and technical composition of the data, the data profile can truly reflect the characteristics of the entire lifecycle of data usage.
[0071] In another embodiment of the present invention, to improve data processing efficiency, the data profile can be sliced, that is, the time series of operation information corresponding to the data attribute features can be sampled according to a preset sampling rate to obtain the data profile. After obtaining the data profile, the association relationship between multiple data profiles and optional scenarios is further determined based on the historical associated scenario information of each data profile included in the metadata; therefore, before step 10, the following is also included:
[0072] Step 210: Perform cluster analysis on the data profile to obtain multiple profile classes and at least one class feature corresponding to each profile class.
[0073] In one embodiment of the present invention, in order to reduce data dimensionality and improve the efficiency of data recommendation, feature extraction can be performed on the data profile, and clustering can be performed on the extracted data features. One of the cluster features can be associated with at least one optional scenario. Clustering analysis can employ methods such as k-means.
[0074] Step 211: Determine the association between each class feature and the optional scenario based on the lineage of the metadata.
[0075] In one embodiment of the present invention, the association between class features and optional scenarios includes information on the frequency of use of the class feature in the optional scenario. The lineage of metadata characterizes the association between metadata and other data, such as whether they originated from the same data source, whether they have been used collaboratively in the same scenario, and whether there are any related influences. Therefore, usage scenario information can be extracted for all class features based on lineage to obtain the scenarios corresponding to each class feature. Collaborative filtering can then be performed based on the scenario information corresponding to all class features to obtain the association.
[0076] Therefore, in another embodiment of the present invention, step 211 further includes:
[0077] Step 2111: Determine the co-occurrence scenario information of each of the aforementioned class features based on the blood relationship; the co-occurrence scenario is at least one of the optional scenarios.
[0078] In one embodiment of the present invention, a co-occurrence scenario refers to a business scenario in which class features coexist, where "occurrence" can mean being used and / or updated in that scenario. The co-occurrence scenario information includes the number of possible scenarios in which each pair of class features coexist.
[0079] Step 2112: Determine the association between each class feature and the optional scene based on the co-occurrence scene information.
[0080] In one embodiment of the present invention, the number of times each class feature appears in the same scene is obtained according to the co-occurrence scene, and collaborative filtering is performed based on the number of times it appears in the same scene. The correlation between the class feature and the optional scene is determined; specifically, the more times it appears, the greater the correlation.
[0081] Therefore, in another embodiment of the present invention, step 2112 further includes:
[0082] Step 220: Determine the similarity information between each of the class features based on the co-occurrence scene information.
[0083] In one embodiment of the present invention, the similarity between each class of features is determined based on the ratio of the number of co-occurring optional scenarios to the total number of all optional scenarios in which that class of features appears.
[0084] Specifically, the similarity W between class feature i and class feature j can be calculated using the following formula. ij :
[0085]
[0086] Where Num(i)∈N * This represents the total number of possible scenarios corresponding to class feature i in the business system. From the above formula, we know that 0 ≤ Wij ≤1, W ij The larger the value, the higher the similarity between class feature i and class feature j.
[0087] Step 221: Determine the association between each of the class features and the optional scene based on the similarity information.
[0088] In one embodiment of the present invention, considering that the similarity of the optional scenes corresponding to the class features with high similarity is also high, the class features and optional scenes can be collaboratively filtered according to the similarity information to obtain the association relationship.
[0089] Therefore, in another embodiment of the present invention, step 221 further includes:
[0090] Step 2210: Determine the initial matching weights between the class features and the optional scenarios.
[0091] In one embodiment of the present invention, the initial matching weights can be preset, such as all of them being set to 1 or randomly set.
[0092] Step 2211: Calculate the distance between each class feature based on the similarity information.
[0093] In one embodiment of the present invention, similarity and distance are inversely proportional, that is, the greater the similarity, the closer the distance. Therefore, the similarity between all class features can be normalized to obtain the distance between each class feature.
[0094] Step 2212: Adjust the initial matching weights according to the distance to obtain the association relationship.
[0095] In one embodiment of the present invention, based on the shortest distance method, for each class feature, the initial matching weight corresponding to the class features whose distance is within a certain range is adjusted by using the optional scene information corresponding to other class features with similar lineage to adjust the association between the current class feature and each optional scene, making it more accurate. In another embodiment of the present invention, gradient descent can also be used iteratively to obtain the matching weight value of each class feature for a specific optional scene.
[0096] Step 212: Determine the association between the data profile and the optional scenarios based on the association between each of the class features and the optional scenarios.
[0097] In one embodiment of the present invention, the distance between the data profile and each class feature is calculated respectively, and the association relationship corresponding to the class feature with the shortest distance is taken as the association relationship between the data screen and the optional scene.
[0098] Step 30: Determine the business support data corresponding to the business scenario based on the target data profile.
[0099] In one embodiment of the present invention, the target data profile includes supporting business scenarios. To further improve the efficiency of data utilization, information related to business support dimensions can be extracted and combined from the target data to obtain a multi-dimensional data cube as business support data. The data cube includes data sources adapted to the application scenario, data processing methods, data analysis methods, dimensional indicators, etc.
[0100] In one embodiment of the present invention, the process of performing data recommendation can also refer to Figure 2 ,like Figure 2 As shown:
[0101] First, step A: raw data acquisition and preprocessing to obtain metadata.
[0102] Step A1: Extract raw data from each data domain and store it in the data warehouse.
[0103] Step A2 involves performing duplicate value removal, missing value handling, and outlier handling operations on the target data.
[0104] Step A3: Perform consistency, granularity verification, and transformation on the data according to unified rules.
[0105] Step A4 records the data hierarchy of the data source in the data domain, the correspondence with the data in the data warehouse, the transformation rules, etc., and stores them in the metadata knowledge base.
[0106] The establishment of metadata is divided into business metadata, technical metadata, and operational metadata. Business metadata focuses on the attributes and categories of data sources, technical metadata focuses on establishing the mapping from the data model to the business model, and operational metadata focuses on the process of data processing to form derived indicators. Business personnel can use metadata to query data lineage, data extraction and transformation relationships, etc.
[0107] Step B: Build a data profile based on metadata.
[0108] Step B can be referred to Figure 3 ,like Figure 3 As shown, first, step B1: Use metadata to describe the entire lifecycle of each data item.
[0109] Metadata includes business metadata, technical metadata, and operational metadata. Business metadata can be used to describe the business attributes of data, technical metadata can be used to describe the technical attributes of data, and operational metadata can be used to describe the time-series operations of data.
[0110] For example, in order to evaluate the value of a site, business personnel can collect technical parameter data, RRC data, ARPU data, etc. from the operations domain (O domain), business domain (B domain), and management domain (M domain) based on expert experience. Each data item contains a large number of dimensions.
[0111] Through step A above, data aggregation across various data domains is completed. For individual data items, such as engineering parameter data, business metadata primarily records the data definition and description, business rules, contact information of stakeholders (such as data owners, data managers, and data users), data security levels, and usage instructions. Technical metadata primarily records physical database field definitions, field attributes, access permissions, and ETL job details. Operational metadata records the execution logs of each stage of data usage, including the results of auditing, balancing, and control metrics. Operational metadata records the entire lifecycle of data integration, cleaning, transformation, processing, analysis, use, archiving, and destruction in a time-series format, with each operation being timestamped.
[0112] Business metadata and technical metadata define and design the data, while operational metadata records the process of data being manipulated. Data profiling uses business metadata and technical metadata as basic category attributes. The values of basic category attributes are relatively stable and only require infrequent, periodic updates; while operational metadata is continuously added, working together with basic category attributes to complete a comprehensive description of the data.
[0113] Step B2: Based on the time-series operations described in the operation metadata, sample the data profile at a fixed sampling rate to generate a data profile slice group.
[0114] The data profiles generated in the preceding steps describe all information throughout the data's lifecycle. To facilitate feature extraction and identification of correlations with usage scenarios, this step samples the data profiles over continuous time at a fixed sampling rate, forming a set of data profile slices.
[0115] Taking site value evaluation as an example, we describe data profiling slicing. First, in the data stitching stage: business traffic, RRC connection count, ARPU, etc., are correlated using sector numbers to form full-dimensional sector data. At this point, the data profiling slice contains both data granularity and data correlation information.
[0116] Secondly, in the data analysis and calculation stage:
[0117] (1) Value grading criteria are obtained through parameter design. At this point, the data profile slice includes data items involved in value grading, the relationship between parameter design and value grading, and the preferences of users in different provinces for value grading, etc.
[0118] (2) Weighted scoring across multiple dimensions. At this point, the data profile slice includes the dimensions involved in the weighted scoring, the weight of each dimension on the score, the impact of different objectives on the dimension weights, and the bias of different provinces on the dimension weights, etc.
[0119] (3) Solution Reporting. At this stage, the data profile slice includes which solutions will be selected for reporting, which solutions will be discarded, and the provincial company's preferences for solutions, etc.
[0120] (4) Solution review. At this stage, the data profile includes which solutions were approved, which solutions were rejected, and the group's preferences for the solutions;
[0121] (5) Modification, submission, and review of the plan. At this point, the data profile slice includes the adjustment process of the rejected plan and the relationship between the group's opinion and the adjusted plan.
[0122] Step B3: Cluster the data profile slices. Extract features from the metadata of the data profile slices, such as traffic-related features in business metadata, or features used for regression calculations on the data, and use clustering algorithms to cluster these features.
[0123] The data profile slices generated by the aforementioned steps contain operation metadata, and operation logs are used to record the data operation process. The operation log records include a global transaction ID, a transaction tracing ID, an operation data object ID, an auxiliary data object ID, scenario information, operator, operation content, and operation time. The global transaction ID is used to ensure transaction uniqueness in a distributed scenario; the tracing ID is used for full-link tracing of each stage and step of the transaction, and each operation records its parent operation tracing ID; the operation data object ID is used to identify the data item being operated on; the auxiliary data object ID is used to identify related data items; the scenario information describes the scenario type; and the operator details are recorded.
[0124] Each transaction is recorded independently, with its own tracking information. Data relationships are generated for each operation by associating it with the operation object, auxiliary operation objects, scenario information, operator, and operation time. Periodic analysis is then performed using business metadata, technical metadata, and the data relationship network to determine the relationships between data and scenarios, and between different data sets.
[0125] Based on the above relationships, automated clustering analysis can be used to obtain cluster features. Then, manual intervention is used to review the clusters and add labels to the features. Operational metadata records all operations performed on the data by application scenarios. Therefore, as application scenarios accumulate data operations and the number of application scenario categories increases, the relationships between data and scenarios, as well as the relationships between data points, will change, and the clustering results will also change accordingly. Scheduled tasks need to be set to update the data profile slices.
[0126] Step B4: Create a sparse index for the clustering. Due to the large number of data items and the even larger number of data slices, a global sparse index is created for the clustering to improve the efficiency of the data retrieval table. Figure 3 Each index tag shown contains 8192 data records, which is adjusted according to the actual number of clusters to balance query efficiency and storage space.
[0127] C: Build a data recommendation model based on data profiles
[0128] Step C1: Obtain the correspondence between cluster features and data application scenarios through metadata lineage, and generate training samples for the data recommendation model. In the training samples, each cluster feature corresponds to multiple application scenarios. For example, traffic features and regression calculation features can support multiple business scenarios respectively, forming a correspondence matrix between cluster features and application scenarios.
[0129] Step C2: Train the samples using a recommendation algorithm to obtain a data recommendation model that matches the clustering features with the data application scenario. The specific steps are as follows:
[0130] Step C21: Compile the co-occurrence matrix of cluster features. Co-occurrence means that different cluster features often correspond to the same application scenario. The number of times two data indicators appear together is statistically analyzed and a matrix is formed. An example is shown below:
[0131]
[0132] Clearly, the co-occurrence matrix is a symmetric matrix, with its diagonal elements set to zero. In the example matrix, n... ij ∈N indicates the number of times feature i and feature j correspond to the same application scenario.
[0133] Step C22: Calculate the similarity matrix of the grouping features.
[0134] Based on the co-occurrence matrix obtained from the above steps, the similarity between feature i and feature j is calculated using the following formula:
[0135]
[0136] Where Num(i)∈N * This represents the total number of application scenarios corresponding to feature i in the business system. By definition, 0 ≤ W ij ≤1, W ij The larger the value, the higher the similarity between feature i and feature j. The similarity matrix is constructed using the similarity calculation results, as shown below:
[0137]
[0138] The above steps use a collaborative filtering algorithm to calculate the similarity of the cluster features and construct an initial similarity matrix.
[0139] Step C23: Train the samples to obtain the data recommendation model.
[0140] The goal of data recommendation models is for data users to use fuzzy search based on application scenarios to derive matching cluster features, thereby obtaining data sources and data operations, and clearly understanding how each step of the data supports the application scenario.
[0141] After clustering and grouping the data profile slices, a large number of features are extracted. The purpose of these features, i.e., their association with application scenarios, is not readily apparent, and only a limited number of samples are available for training under the current business model. Feature-application scenario matching weights are assigned using manual or random assignment methods, constructing a feature-application scenario matching weight matrix. Using existing business scenarios as samples, gradient descent is used to iteratively calculate the matching weight value of each feature for a specific data application scenario. The feature-application scenario matching weight matrix represents the association between features and application scenarios, while the feature similarity matrix constructed in step 3022 represents the association between features. The shortest distance method can then be used to obtain the model of application scenario search and recommendation clustered features.
[0142] Step D: Recommend scenario-related data using a data recommendation model.
[0143] Step D4: Use application scenario keywords to query and use the sparse index established in step 204 to find data profiles. The query results are data cubes composed of data sources, data processing methods, data analysis methods, and dimensional indicators that are suitable for the scenario.
[0144] Search engines provide query entry points for profiling massive amounts of data. Data application personnel perform fuzzy searches for data application scenarios, and the search engine uses Natural Language Processing (NLP) to obtain scenario keywords. The data recommendation model established in step C23 then yields clustering features. Using sparse indexes can significantly improve retrieval efficiency. Through these clustering features, a data cube is obtained, consisting of data sources, data processing methods, data analysis methods, and dimensional indicators suitable for the application scenario.
[0145] Step D2: Based on the search results feedback, update the co-occurrence matrix in step C21 above, and re-enter the data recommendation model construction step.
[0146] Based on the data cube obtained from the preceding steps, data application personnel can intuitively see which data sources employ which data processing and analysis methods to support the task completion of the application scenario. This enables the data platform to support the business capabilities of data application personnel. Simultaneously, based on their business experience, data application personnel determine whether the search results meet the scenario requirements. The search engine then returns the results to the training samples in step 302 for further optimization of the search model.
[0147] The above steps comprehensively utilize metadata to profile the data, then associate it with data application scenarios to form a data recommendation model. For example, if business personnel want to query what data is needed and what operations are required for traffic prediction in an indoor distributed antenna system (DAS) scenario, they can enter "indoor coverage traffic prediction" into a search engine. This will be resolved into keywords such as "indoor DAS," "traffic," and "prediction," and the search will recommend the data items supporting that scenario and the methods for performing those operations. Because metadata records a complete picture of the data, business personnel can find the source of the data and the department that previously performed the data operations through the search results, and seek technical support.
[0148] In summary, this method effectively enhances data value. Because metadata records the frequency and effectiveness of data usage, in the examples above, application users can clearly see the top-ranked data algorithms and data manipulation schemes, facilitating optimal selection. Furthermore, based on the correlations between data profile features, search results can also reveal previously unused data usage schemes, inspiring business applications.
[0149] The data recommendation method provided in this invention identifies a business scenario; determines a target data profile corresponding to the business scenario based on a recommendation model; the recommendation model is determined based on the association between multiple data profiles and optional scenarios; wherein, the data profile is used to describe the characteristics of a data item throughout its entire lifecycle; the data profile and the association are determined based on the metadata corresponding to the data item; and the business support data corresponding to the business scenario is determined based on the target data profile. This differs from existing technologies that rely on expert experience to determine which data items to use and what operations to perform to support the development and analysis of business scenarios, resulting in low efficiency and accuracy in data utilization. This invention describes the characteristics of the entire lifecycle of data based on metadata, and further trains a recommendation model based on the association between the data profile and optional business development scenarios. This allows the target data profile corresponding to the current business scenario to be determined based on the recommendation model, and finally, the corresponding business support data is determined based on the target data profile. This achieves data support for the business scenario, improves the efficiency and accuracy of data utilization, and enables the effective support of massive wireless network data items for business scenario applications.
[0150] Figure 4A schematic diagram of the data recommendation device provided in an embodiment of the present invention is shown. Figure 4 As shown, the device 40 includes: a first determining module 401, a recommending module 402, and a second determining module 403.
[0151] The first determining module 401 is used to determine the business scenario;
[0152] The recommendation module 402 is used to determine the target data profile corresponding to the business scenario based on the recommendation model; the recommendation model is determined based on the association between multiple data profiles and optional scenarios; wherein, the data profile is used to describe the characteristics of a data item throughout its entire lifecycle; the data profile and the association are determined based on the metadata corresponding to the data item.
[0153] The second determining module 403 is used to determine the business support data corresponding to the business scenario based on the target data profile.
[0154] The operation process of the data recommendation device provided in this embodiment of the invention is largely the same as that of the aforementioned method embodiment, and will not be described again.
[0155] The data recommendation device provided in this invention determines a business scenario; determines a target data profile corresponding to the business scenario based on a recommendation model; the recommendation model is determined based on the association between multiple data profiles and optional scenarios; wherein, the data profile is used to describe the characteristics of a data item throughout its entire lifecycle; the data profile and the association are determined based on the metadata corresponding to the data item; and the business support data corresponding to the business scenario is determined based on the target data profile. This differs from existing technologies that rely on expert experience to determine which data items to use and what operations to perform to support the development and analysis of business scenarios, resulting in low efficiency and accuracy in data utilization. This invention describes the characteristics of the entire lifecycle of data based on metadata, and further trains a recommendation model based on the association between the data profile and optional business development scenarios. This allows the target data profile corresponding to the current business scenario to be determined based on the recommendation model, and finally, the corresponding business support data is determined based on the target data profile. This achieves data support for business scenarios, improves the efficiency and accuracy of data utilization, and enables the effective support of massive wireless network data items for business scenario applications.
[0156] Figure 5 The diagram shows a structural schematic of a data recommendation device provided in an embodiment of the present invention. The specific implementation of the data recommendation device is not limited by the specific embodiments of the present invention.
[0157] like Figure 5As shown, the data recommendation device may include: processor 502, communications interface 504, memory 506, and communications bus 508.
[0158] The processor 502, communication interface 504, and memory 506 communicate with each other via communication bus 508. Communication interface 504 is used to communicate with other network elements, such as clients or other servers. The processor 502 executes program 510, specifically performing the relevant steps described above in the data recommendation method embodiment.
[0159] Specifically, program 510 may include program code, which includes computer-executable instructions.
[0160] Processor 502 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention. The data recommendation device may include one or more processors of the same type, such as one or more CPUs; or it may include processors of different types, such as one or more CPUs and one or more ASICs.
[0161] Memory 506 is used to store program 510. Memory 506 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.
[0162] Specifically, program 510 can be called by processor 502 to cause the data recommendation device to perform the following operations:
[0163] Define the business scenario;
[0164] The target data profile corresponding to the business scenario is determined based on the recommendation model; the recommendation model is determined based on the association between multiple data profiles and optional scenarios; wherein, the data profile is used to describe the characteristics of a data item throughout its entire lifecycle; the data profile and the association are determined based on the metadata corresponding to the data item.
[0165] Based on the target data profile, determine the business support data corresponding to the business scenario.
[0166] The operation process of the data recommendation device provided in this embodiment of the invention is largely the same as that of the aforementioned method embodiment, and will not be repeated here.
[0167] The data recommendation device provided in this invention determines a business scenario; determines a target data profile corresponding to the business scenario based on a recommendation model; the recommendation model is determined based on the association between multiple data profiles and optional scenarios; wherein, the data profile is used to describe the characteristics of a data item throughout its entire lifecycle; the data profile and the association are determined based on the metadata corresponding to the data item; and the business support data corresponding to the business scenario is determined based on the target data profile. This differs from existing technologies that rely on expert experience to determine which data items to use and what operations to perform to support the development and analysis of business scenarios, resulting in low efficiency and accuracy in data utilization. This invention describes the characteristics of the entire lifecycle of data based on metadata, and further trains a recommendation model based on the association between the data profile and optional business development scenarios. This allows the target data profile corresponding to the current business scenario to be determined based on the recommendation model, and finally, the corresponding business support data is determined based on the target data profile. This achieves data support for business scenarios, improves the efficiency and accuracy of data utilization, and enables the effective support of massive wireless network data items for business scenario applications.
[0168] This invention provides a computer-readable storage medium storing at least one executable instruction that, when executed on a data recommendation device, causes the data recommendation device to perform the data recommendation method described in any of the above method embodiments.
[0169] Specifically, the executable instructions can be used to cause the data recommendation device to perform the following operations:
[0170] Define the business scenario;
[0171] The target data profile corresponding to the business scenario is determined based on the recommendation model; the recommendation model is determined based on the association between multiple data profiles and optional scenarios; wherein, the data profile is used to describe the characteristics of a data item throughout its entire lifecycle; the data profile and the association are determined based on the metadata corresponding to the data item.
[0172] Based on the target data profile, determine the business support data corresponding to the business scenario.
[0173] The operation process of the executable instructions stored in the computer storage medium provided in this embodiment of the invention is largely the same as that in the aforementioned method embodiments, and will not be described again.
[0174] The executable instructions stored in the computer storage medium provided in this embodiment of the invention determine a business scenario; determine the target data profile corresponding to the business scenario according to a recommendation model; the recommendation model is determined based on the association between multiple data profiles and optional scenarios; wherein, the data profile is used to describe the characteristics of a data item throughout its entire lifecycle; the data profile and the association are determined based on the metadata corresponding to the data item; and the business support data corresponding to the business scenario is determined based on the target data profile. This differs from existing technologies that rely on expert experience to determine which data items to use and what operations to perform to support the development and analysis of business scenarios, resulting in low efficiency and accuracy of data utilization. This embodiment of the invention can describe the characteristics of the entire lifecycle of data based on metadata, and further train a recommendation model based on the association between the data profile and optional business development scenarios. This allows the target data profile corresponding to the current business scenario to be determined based on the recommendation model, and finally, the corresponding business support data is determined based on the target data profile. This achieves data support for business scenarios, improves the efficiency and accuracy of data utilization, and enables the effective support of massive wireless network data items for business scenario applications.
[0175] This invention provides a data recommendation apparatus for executing the above-described data recommendation method.
[0176] This invention provides a computer program that can be invoked by a processor to cause a data recommendation device to execute the data recommendation method in any of the above method embodiments.
[0177] This invention provides a computer program product, which includes a computer program stored on a computer-readable storage medium. The computer program includes program instructions that, when executed on a computer, cause the computer to perform the data recommendation method described in any of the above method embodiments.
[0178] The algorithms or displays provided herein are not inherently related to any particular computer, virtual system, or other device. Various general-purpose systems can also be used in conjunction with the teachings herein. The required structure for constructing such systems is apparent from the above description. Furthermore, the embodiments of the present invention are not directed to any particular programming language. It should be understood that the content of the invention described herein can be implemented using various programming languages, and the above description of specific languages is for the purpose of disclosing the best mode of implementation of the invention.
[0179] Numerous specific details are set forth in the specification provided herein. However, it will be understood that embodiments of the invention may be practiced without these specific details. In some instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this specification.
[0180] Similarly, it should be understood that, in order to streamline the invention and aid in understanding one or more of the various aspects of the invention, features of the embodiments of the invention are sometimes grouped together in a single embodiment, figure, or description thereof in the above description of exemplary embodiments of the invention. However, this disclosure should not be construed as reflecting an intention that the claimed invention requires more features than are expressly recited in each claim.
[0181] Those skilled in the art will understand that modules in the device of the embodiments can be adaptively changed and placed in one or more devices different from that embodiment. Modules, units, or components in the embodiments can be combined into a single module, unit, or component, and can be divided into multiple sub-modules, sub-units, or sub-components. Except where at least some of such features and / or processes or units are mutually exclusive, any combination can be used to combine all features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all processes or units of any method or device so disclosed. Unless expressly stated otherwise, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) may be replaced by an alternative feature that serves the same, equivalent, or similar purpose.
[0182] It should be noted that the above embodiments are illustrative of the invention and not restrictive, and that those skilled in the art can devise alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses should not be construed as limiting the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The invention can be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In the unit claims enumerating several means, several of these means may be embodied by the same item of hardware. The use of the words first, second, and third, etc., does not indicate any order. These words can be interpreted as names. The steps in the above embodiments, unless otherwise specified, should not be construed as limiting the order of execution.
Claims
1. A data recommendation method, characterized in that, The method includes: Define the business scenario; The target data profile corresponding to the business scenario is determined based on a recommendation model; the recommendation model is determined based on the association between multiple data profiles and optional scenarios; wherein, the data profile is used to describe the characteristics of a data item throughout its entire lifecycle; the data profile and the association are determined based on the metadata corresponding to the data item; wherein, before determining the target data profile corresponding to the business scenario based on the recommendation model, the method further includes: performing cluster analysis on the data profile to obtain multiple profile classes and at least one class feature corresponding to each profile class; determining the association between each class feature and optional scenarios based on the lineage of the metadata; and determining the association between the data profile and optional scenarios based on the association between each class feature and the optional scenarios. Based on the target data profile, determine the business support data corresponding to the business scenario.
2. The method according to claim 1, characterized in that, The metadata corresponds to multiple of the data items; Before determining the target data profile corresponding to the business scenario based on the recommendation model, the method further includes: Feature extraction is performed on the metadata to obtain the data attribute features and operation timing features corresponding to each data item; The data profile is obtained by arranging and combining the data attribute features according to the operation timing features.
3. The method according to claim 2, characterized in that, The metadata includes business domain metadata, technical domain metadata, and operational domain metadata; the data attribute features include business attribute features and technical attribute features; the feature extraction of the metadata to obtain the data attribute features and operational timing features corresponding to each data item includes: The business features are obtained by extracting features from the business domain metadata. The technical features are obtained by extracting features from the metadata of the technical domain. The operation domain metadata is subjected to feature extraction to obtain the operation timing features.
4. The method according to claim 1, characterized in that, Determining the association between each class feature and the optional scenario based on the lineage of the metadata includes: The co-occurrence scenario information for each of the aforementioned class features is determined based on the blood relationship; the co-occurrence scenario is at least one of the optional scenarios; The association between each class feature and the optional scenario is determined based on the co-occurrence scenario information.
5. The method according to claim 4, characterized in that, Determining the association between each class feature and the optional scene based on the co-occurrence scene information includes: The similarity information between each of the class features is determined based on the co-occurrence scene information; The association between each of the class features and the optional scene is determined based on the similarity information.
6. The method according to claim 5, characterized in that, Determining the association between each of the class features and the optional scene based on the similarity information includes: Determine the initial matching weights between the class features and the optional scenarios; The distance between each of the class features is calculated based on the similarity information; The initial matching weights are adjusted based on the distance to obtain the association relationship.
7. A data recommendation device, characterized in that, The device includes: The first determination module is used to determine the business scenario; A recommendation module is used to determine the target data profile corresponding to the business scenario based on a recommendation model. The recommendation model is determined based on the association between multiple data profiles and optional scenarios. The data profile describes the characteristics of a data item throughout its entire lifecycle. The data profile and its association are determined based on the metadata corresponding to the data item. Before determining the target data profile corresponding to the business scenario based on the recommendation model, the module further includes: performing cluster analysis on the data profiles to obtain multiple profile classes and at least one class feature corresponding to each profile class; determining the association between each class feature and optional scenarios based on the lineage of the metadata; and determining the association between the data profile and optional scenarios based on the association between each class feature and the optional scenarios. The second determining module is used to determine the business support data corresponding to the business scenario based on the target data profile.
8. A data recommendation device, characterized in that, include: The processor, memory, communication interface, and communication bus are provided, wherein the processor, memory, and communication interface communicate with each other via the communication bus. The memory is used to store at least one executable instruction that causes the processor to perform the operation of the data recommendation method as described in any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The storage medium stores at least one executable instruction, which, when executed on the data recommendation device, causes the data recommendation device to perform the operation of the data recommendation method as described in any one of claims 1-6.
Citation Information
Patent Citations
Data model construction method, device and apparatus and computer readable storage medium
CN110457288A
Information processing method and device
CN113298397A