Discoverability enhancement and recommendation method for data assets

By constructing a unified semantic model across departments, the problem of data demand integration in cross-departmental collaboration was solved, enabling dynamic integration and recommendation of data assets, improving collaboration efficiency and analysis accuracy, and possessing self-optimization capabilities.

CN121743374APending Publication Date: 2026-03-27SHENZHEN YUNCHUANG YOUYI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-20
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing data asset management systems are unable to dynamically integrate semantically heterogeneous data requirements in cross-departmental collaborative tasks, resulting in fragmented recommendation results that cannot directly support integrated analysis, thus affecting collaboration efficiency and analytical accuracy.

Method used

Construct a cross-departmental unified semantic model that supports dynamic semantic mapping. Define the hierarchical relationship and semantic association rules between data tags through semantic ontology to achieve context awareness and semantic decoupling, generate composite semantic query conditions, and recommend integrated data asset packages based on comprehensive matching degree.

Benefits of technology

It enables the dynamic fusion and recommendation of data assets in cross-departmental collaborative tasks, improving the discoverability, collaboration efficiency, and analytical accuracy of data assets, and possessing the ability to continuously adapt to business changes and self-optimize.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121743374A_ABST
    Figure CN121743374A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data asset management, in particular to a data asset discoverability enhancement and recommendation method, which comprises the following steps: constructing a unified semantic model; on the basis of the unified semantic model, context sensing and semantic decoupling are carried out on a cross-department cooperation task to extract differentiated semantic requirements of roles of all parties, and the requirements are fused into a composite semantic query condition; and according to the composite semantic query condition, calculating a comprehensive matching degree through a pre-constructed semantic association matrix, and according to the comprehensive matching degree, dynamically assembling the multi-source data assets meeting a threshold condition into an integrated data asset package based on the semantic association rule and the association field to perform priority recommendation. According to the method, a unified model and a context sensing mechanism which support dynamic semantic mapping can be constructed, cross-department heterogeneous data requirements are fused into a unified composite query condition, and an available integrated data asset package is directly recommended.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data asset management, and in particular to a method for enhancing and recommending the discoverability of data assets. Background Technology

[0002] As enterprises increasingly embrace data-driven operations, cross-departmental and cross-domain data collaboration has become a key scenario for unlocking the value of data assets. In companies like new energy vehicle manufacturers, it's increasingly common for sales and R&D departments to collaborate on tasks such as "paying user charging function satisfaction analysis." In such scenarios, the starting point and core obstacle to collaboration lies in the semantic barriers between departments. For example, the sales department's definition of "customer" specifically refers to "users who have completed vehicle purchases," and their data includes fields such as user ID and purchase time; while the R&D department's definition of "customer" broadly refers to "all users of the in-vehicle app," and their data includes user ID, charging frequency, and other behavioral logs.

[0003] While existing mainstream data asset management and discovery systems possess departmental metadata management and keyword-based retrieval capabilities, and even introduce static ontology for conceptual association, they still have fundamental limitations when handling dynamic and specific cross-departmental collaborative tasks.

[0004] The core flaw in existing technologies lies in their inability to dynamically integrate and transform semantically heterogeneous raw data requirements from different departments into a single, directly executable semantic query condition that simultaneously constrains multiple related data sources and their specific attributes within a specific cross-departmental collaborative task context. This results in fragmented recommended data assets that cannot directly support cross-departmental integrated analysis. For example, when faced with the task of "analyzing the charging function satisfaction of paid users," the system might recommend the "list of paid users" from the sales department and the "product user charging data" from the R&D department. However, because it cannot automatically recognize the business semantics that "paid users" are a subset of "product users," and cannot translate this understanding into a composite query logic of "associating by user ID and filtering only the charging data of paid users," the final recommendation results will inevitably be misaligned: the data provided by the R&D department contains non-paid users such as test drive users, while the data provided by the sales department lacks crucial charging behavior fields. This forces business personnel to perform tedious manual associativity, filtering, and cleaning, severely impacting collaboration efficiency and analytical accuracy. This systemic failure from semantic understanding to requirement integration constitutes the main bottleneck in improving the discoverability and collaborative value of data assets. Summary of the Invention

[0005] To construct a unified model and context-aware mechanism that supports dynamic semantic mapping, integrate heterogeneous data requirements across departments into unified composite query conditions, and directly recommend available integrated data asset packages, this application provides a method for enhancing the discoverability and recommendation of data assets.

[0006] This application provides a method for enhancing and recommending the discoverability of data assets, employing the following technical solution: A method for enhancing and recommending the discoverability of data assets, comprising: Construct a cross-departmental unified semantic model that supports dynamic semantic mapping. The unified semantic model defines the hierarchical relationship and semantic association rules between data tags through a semantic ontology. Based on the unified semantic model, context awareness and semantic decoupling are performed on cross-departmental collaborative tasks to extract the differentiated semantic requirements of each party's role, and these requirements are integrated into a composite semantic query condition that can simultaneously constrain multiple related data sources and their specific attributes, and contains cross-departmental related fields. Based on the composite semantic query conditions, the comprehensive matching degree between data assets from different departments and the current task context is calculated through a pre-constructed semantic association matrix. Based on the comprehensive matching degree, multi-source data assets that meet the threshold conditions are dynamically assembled into an integrated data asset package that can be directly used for cross-departmental analysis and prioritized for recommendation.

[0007] Optionally, constructing the unified semantic model specifically involves: defining a general core class and specific subclasses corresponding to departments based on a semantic ontology; establishing semantic hierarchy and semantic association relationships between the core class and subclasses, as well as between different subclasses; and labeling each subclass with its corresponding specific data attribute fields to form a structured and scalable semantic association network.

[0008] Optionally, when a new business tag appears, the unified semantic model performs dynamic semantic mapping; the dynamic semantic mapping includes: the unified semantic model automatically extracts the semantics of the new business tag, associates it with the corresponding specific subclass already existing in the unified semantic model, and synchronously updates the semantic association relationship and the corresponding data attribute fields.

[0009] Optionally, the context awareness and semantic decoupling specifically involves: parsing the cross-departmental collaborative task based on the hierarchical relationship and semantic association rules defined by the unified semantic model, in order to identify the roles of participating departments and their collaborative goals, and to separate the differentiated data subclasses, attribute requirements and associated field constraints corresponding to each role, thereby forming a structured set of multi-role semantic constraints.

[0010] Optionally, the requirement can be integrated into a composite semantic query condition by: based on the semantic association rules defined by the unified semantic model, performing semantic alignment and logical association on each constraint in the multi-role semantic constraint set to generate a composite semantic query condition.

[0011] Optionally, the calculation of the comprehensive matching degree specifically involves: based on the composite semantic query conditions, obtaining the semantic relevance, demand attribute matching degree, and related field association capability of the corresponding data assets from the pre-constructed semantic association matrix, and performing a fusion calculation in conjunction with the semantic constraint information represented by the multi-role semantic constraint set to obtain the comprehensive matching degree.

[0012] Optionally, the fusion calculation of the semantic constraint information represented by the multi-role semantic constraint set specifically involves: assigning semantic weights to the semantic constraints of different roles based on the multi-role semantic constraint set, and performing weighted calculations on the semantic relevance, demand attribute matching degree, and relevance field association ability obtained from the semantic association matrix according to the semantic weights, so as to complete the fusion of the comprehensive matching degree.

[0013] Optionally, the step of generating a recommendation list may also be included: Based on the comprehensive matching degree calculated by the weighted average, multi-source data assets that meet the preset matching degree threshold are selected from data assets from different departments; Based on the semantic association rules and the association fields, the selected multi-source data assets are dynamically assembled into an integrated data asset package; Based on the total matching degree of all data assets contained in each integrated data asset package, priority is ranked, and a recommendation list is generated and output.

[0014] Optional, also includes: Based on the recommendation list, user feedback on the semantic accuracy, attribute completeness, and business value dimensions of the recommendation data is collected in a structured manner. Based on the evaluation feedback, the priority of data filtering rules, attribute matching weights, and confidence levels of object attributes in the unified semantic model are dynamically adjusted to generate subsequent recommendation lists.

[0015] Optionally, generating and outputting the recommendation list further includes: Each integrated data asset package will be output as an independent item; Each independent item includes: a semantic purpose description of the integrated data asset package, source information of the data assets that make up the package, a list of related fields used for cross-departmental association, and the semantic association logic on which the related fields are based.

[0016] In summary, this application includes the following beneficial technical effects: This application constructs a unified model and context-aware mechanism that supports dynamic semantic mapping. In specific cross-departmental collaborative tasks, it can dynamically integrate and transform semantically heterogeneous raw data requirements from different departments into a directly executable composite semantic query condition that can simultaneously constrain multiple related data sources and their specific attributes. Based on this condition, it directly recommends dynamically assembled integrated data asset packages that can be directly used for cross-departmental analysis. This completely solves the core technical problem of existing technologies, which suffer from fragmented recommendation results and inability to directly support integrated analysis due to semantic barriers and failure of requirement integration. It significantly improves the discoverability, collaboration efficiency, and analytical accuracy of data assets.

[0017] The unified semantic model constructed in this application has dynamic semantic mapping and expansion capabilities. When new data tags or concepts appear in the business, it can automatically extract their semantics and associate them with the existing semantic structure of the model, and update the semantic association network synchronously. This enables the model to continuously adapt to business changes without manual reconstruction, fundamentally overcoming the shortcomings of static ontology models that are rigid and difficult to adapt to dynamic business scenarios, and ensuring the timeliness of semantic benchmarks and the long-term effectiveness of recommendation systems.

[0018] This application establishes a closed-loop optimization mechanism based on user feedback, which can structurally collect user evaluations of recommended data in terms of semantic accuracy, attribute completeness, and business value. Based on this feedback, the system dynamically adjusts the data filtering rules, attribute matching weights, and semantic association confidence in subsequent recommendation processes, enabling the recommendation system to have the ability to continuously learn and self-optimize, constantly getting closer to real business needs and user preferences, thereby continuously improving recommendation accuracy and business practical value in long-term operation. Attached Figure Description

[0019] Figure 1 This is a flowchart of the enhancement and recommendation methods; Figure 2 This is a flowchart of the process for generating compound semantic query conditions; Figure 3 It is a line graph showing the precision and recall rates at different matching thresholds. Detailed Implementation

[0020] The following is in conjunction with the appendix Figure 1-3 This application will be described in further detail.

[0021] This application discloses a method for enhancing the discoverability and recommendation of data assets. For example... Figure 1As shown, a method for enhancing the discoverability and recommendation of data assets is presented. This method addresses issues such as cross-departmental semantic barriers, the inability to dynamically integrate requirements, and fragmented recommendation data in existing technologies by constructing a unified cross-departmental semantic model, integrating heterogeneous cross-departmental needs, calculating data asset matching degrees, generating integrated data asset packages, and iteratively optimizing them. It is suitable for scenarios requiring cross-departmental data collaboration, such as new energy vehicle companies. The following steps are described in detail according to logical progression: S1 constructs a unified semantic model like Figure 1 As shown, this step relies on semantic ontology theory and the enterprise's existing data governance foundation to establish a structured semantic association network. This network can unify the semantic cognition of cross-departmental data tags, provide a benchmark for subsequent parsing and collaboration task requirements, effectively alleviate the problem of semantic isolation of data tags between departments in existing technologies, and also have the ability to adapt to business changes.

[0022] S11 Semantic Ontology Modeling When defining the general core class, the system combines the semantic ontology specifications of the data asset management field with the common business characteristics of new energy vehicle companies, identifying the general core class as "customer". The semantic definition of "customer" assigned by the system is various types of users who interact with the enterprise. This definition covers the basic understanding of user data by multiple departments such as sales, R&D, and customer service, providing a unified benchmark for subsequent subclass expansion.

[0023] The system is based on a core customer class, defining specific subclasses for key departments involved in cross-departmental collaboration, and labeling corresponding data attribute fields. For example, the sales department's subclass is for paying users, with configured attribute fields including user ID, purchase time, vehicle model, and transaction amount; the R&D department's subclass is for product users, with configured attribute fields including user ID, number of charges, function click records, and app version number; and the customer service department's subclass is for inquiry users, with configured attribute fields including user ID, inquiry type, conversation duration, and satisfaction rating.

[0024] Based on the business process logic of each department and the company's cross-departmental collaboration history data over the past three years, the system establishes semantic relationships between core categories and subcategories, as well as between different subcategories. Core category customers and specific subcategories of each department form an inclusion relationship, meaning that all user groups within specific subcategories belong to the customer category. The relationships between different subcategories are determined through historical data statistics: paid users and product users have an inclusion relationship; statistics on user behavior data from new energy vehicle companies over the past year show that over 99% of users who have completed vehicle purchases register and use the in-vehicle app, therefore paid users can be largely included in the product user category; paid users and consulting users have a partial inclusion relationship; statistics show that 60% of users in this group have consulted customer service about vehicle usage issues, a percentage derived through statistical analysis of the sales department's paid user list and customer service work order system data; product users and consulting users have an overlapping relationship; some test drive users who only use the app may initiate inquiries, while some consulting users are paid users who have already purchased vehicles.

[0025] The system structurally integrates the aforementioned core classes, specific subclasses, attribute fields, and relationships to form a scalable semantic association network. The specific integration method is as follows: First, a semantic ontology model is established based on the above definitions, defining core classes and specific subclasses as semantic nodes in the model, and attribute fields as attribute sets of nodes; second, semantic relationships are transformed into connection relationships between nodes, with inclusion relationships defined as parent-child node relationships, and partial inclusion and cross relationships defined as association relationships between sibling nodes; finally, by establishing node indexes and association relationship mapping tables, an association network structure supporting semantic queries is formed. This network helps the system quickly identify the semantic relationships of data tags from different departments when subsequently parsing cross-departmental requirements, avoiding data requirement misalignment due to semantic understanding biases. Compared to existing static ontology models, its association relationships are built based on actual enterprise business data, making it more closely aligned with real business scenarios.

[0026] S12 Dynamic Semantic Mapping Processing When a company adds new business tags as its business expands, the unified semantic model automatically executes a dynamic semantic mapping process. This integrates the new tags into the existing system without requiring manual model reconstruction, significantly improving the model's adaptability to business changes. The following section details the specific process using the scenario of adding new rental user business tags as an example.

[0027] The unified semantic model first extracts the semantics of the new tag, then uses natural language processing to parse the business description document of the rental user, determining that the core semantics of the tag is a user who obtains the right to use a vehicle through rental. Simultaneously, the model identifies the core data attribute fields of the tag from the rental business management system, including user ID, rental start time, rental period, and vehicle model, ensuring the semantics and data dimensions of the new tag are complete.

[0028] Subsequently, the unified semantic model is associated with existing specific subclasses. Based on semantic association rules, rental users are associated with the core customer category, as a specific subclass parallel to paying users. Considering the actual scenario of the rental business, rental users typically use the app to perform functions such as charging and navigation while using the vehicle. The model refers to the association logic between paying users and product users to determine that rental users and product users form an inclusion relationship.

[0029] Finally, the unified semantic model is updated with related information, and the semantic association network is improved simultaneously, supplementing the network structure with the hierarchical relationships between rental users, customers, and product users. The specific update operations are as follows: First, a new semantic node for rental users is added to the semantic ontology model, and its parent node is set to the customer node; second, an inclusion relationship is established between the rental user node and the product user node; third, the attribute fields of rental users are added to the node attribute set; finally, the node index and association mapping table are updated to ensure that new nodes can be quickly retrieved and associated. Simultaneously, the model incorporates the user ID field of rental users into the cross-departmental association field system, ensuring that data for this type of user can be associated through a unified field during subsequent cross-departmental collaboration, maintaining the integrity and consistency of the semantic association network.

[0030] Context awareness and compound semantic query condition generation for S2 cross-departmental collaboration tasks like Figure 1 and Figure 2 As shown, this step uses the unified semantic model built by S1 as the core support. By parsing the context information of collaborative tasks, the heterogeneous semantic requirements of different departments are transformed into unified query conditions, solving the problem of inefficient data matching caused by the misunderstanding of requirements in the existing technology.

[0031] S21 Context Awareness and Semantic Decoupling S211 receives and parses collaborative tasks. The system receives complete information about cross-departmental collaborative tasks, including the task name, participating department identifiers, and a textual description of the collaboration objective. Taking a collaborative task involving a new energy vehicle company as an example, the core information retrieved by the system is "Paid User Charging Function Satisfaction Analysis," with participating department identifiers pointing to the sales and R&D departments, and the collaboration objective described as "Clarifying paid users' feedback and satisfaction levels regarding the on-board charging function." This basic information provides clear business scenario anchors for subsequent semantic parsing, preventing the parsing process from deviating from actual collaboration needs.

[0032] S212 Identify departmental roles and collaborative goals The system invokes the unified semantic model built by S1, and based on the association rules of specific subclasses under the core "Customer" class in the model, parses the roles and collaborative goals of participating departments. The model defines "Paying Users" as belonging to a specific subclass of the sales department and "Product Users" as belonging to a specific subclass of the R&D department, with an inclusion relationship between the two, providing direct evidence for role identification. Based on this, the system determines that the core role of the sales department is to provide accurate "Paying User" identity data, and the core role of the R&D department is to provide "Charging Function Usage" behavior data for this type of user. The shared collaborative goal of both departments is to integrate identity and behavioral data around the "Paying User" group to support satisfaction analysis. This identification process relies on the semantic benchmark of S1, effectively avoiding target deviations between departments due to differences in "User" definitions.

[0033] S213 Separate Differentiated Semantic Constraints Based on the attribute field definitions of specific subcategories for each department in S1, the system separates the differentiated semantic constraints of each party's role from the collaboration objectives. The semantic constraints for the sales department explicitly target the "Paying User" subcategory, requiring data to include three core attributes: User ID, purchase time, and vehicle model. The associated field is limited to User ID, and these attributes completely correspond to the "Paying User" attribute fields configured for the sales department in S1. The semantic constraints for the R&D department target the "Product User" subcategory, requiring data to include three attributes: User ID, number of charges, and function click records. The associated field is also User ID, and the attribute range is consistent with the "Product User" configuration for the R&D department in S1. The system organizes these constraints into a structured set of multi-role semantic constraints, with each constraint item labeled with the corresponding department, data subcategory, and attribute requirements, providing a clear basis for subsequent accurate data matching.

[0034] S22 merges to generate composite semantic query conditions S221 Performs semantic alignment Based on the semantic association rules defined in S1, the system performs semantic alignment on the constraint items in the multi-role semantic constraint set. Addressing the constraint differences between the sales department's "paying users" and the R&D department's "product users," the system invokes the association conclusion in the model that "paying users are a subset of product users," setting the sales department's "paying users" constraint as the core constraint and the R&D department's "product users" constraint as an extended constraint. This alignment method, based on the priority order of business semantics, ensures that query conditions not only meet the identity definition needs of the sales department but also cover the functional data scope of the R&D department, avoiding the problem of incomplete query results caused by semantic conflicts in existing technologies.

[0035] S222 Establish logical connections The system revolves around the collaborative goal of "analyzing charging satisfaction among paying users," establishing clear logical relationships for the aligned constraints. The system defines an "AND" logical relationship between each constraint, meaning the query result must simultaneously satisfy: the data belongs to the "Paying User" subclass, contains the identity attributes required by the sales department, belongs to the "Product User" subclass, contains the behavioral attributes required by the R&D department, and all data must be linked through the user ID field. This logical relationship is entirely based on the semantic network of S1. The rationale for using the user ID as a cross-departmental linking field has been verified using historical data during the S1 construction process, ensuring logical consistency and reliability.

[0036] S223 Generates compound semantic query conditions The system integrates the results of semantic alignment and logical association to generate the final composite semantic query condition. This condition explicitly stipulates that the data must simultaneously contain five attributes: user ID, purchase time, vehicle model, number of charging attempts, and function click records; the user ID field serves as the cross-departmental association benchmark and must be present across all data; and the user group corresponding to the data must simultaneously meet the subclass definitions of "paying user" and "product user." This query condition can simultaneously constrain the data sources of both the sales and R&D departments, avoiding the data fragmentation problem caused by separate queries in existing technologies.

[0037] S3 Overall Matching Degree Calculation like Figure 1 As shown, this step takes the composite semantic query conditions generated by S2 as the core basis, extracts data asset features through a pre-constructed semantic association matrix, and completes the comprehensive matching degree calculation by combining the weights of multi-role semantic constraints, providing a quantitative standard for subsequent data screening.

[0038] S31 Pre-built semantic association matrix Before performing cross-departmental collaborative task matching, the system pre-constructs a two-dimensional semantic association matrix of "data assets - matching indicators". The matrix construction references the metadata information of the company's existing data assets and the matching records of cross-departmental collaborative data over the past three years to ensure that the matrix features closely match the actual business. The matrix's row dimension covers the core data assets of each department of the company, while the column dimension sets three indicators: semantic association degree, requirement attribute matching degree, and association capability of related fields. All indicators range from 0 to 1 and are dimensionless values. The larger the value, the stronger the matching degree or association capability between the data asset and the task requirements.

[0039] The metrics for each data asset were determined through historical data statistics and semantic analysis. The sales department's "Paid User List" semantically aligns highly with the "Paid User" subclass constrained by the sales role in S2. Based on nearly a hundred historical matching records, its semantic relevance is determined to be 0.9. This list contains all attribute fields required by S2, with a required attribute matching degree set to 1.0. It also includes the core related field, User ID, with a related field association capability of 1.0. The R&D department's "Product User Charging Data" closely matches the "Product User" subclass constrained by the R&D role. Based on the correlation analysis of APP logs and task requirements, its semantic relevance is determined to be 0.8. It contains charging-related attributes required by S2, with a required attribute matching degree of 1.0, and the related field association capability is also 1.0. The customer service department's "Consultation User Satisfaction Data" has a weaker correlation with the core task requirements. Based on the crossover ratio of consultation users and paid users, its semantic relevance is determined to be 0.7. This data only contains some necessary attributes, with a required attribute matching degree of 0.5, but it also includes the User ID field, with a related field association capability of 1.0. Non-core data assets such as the "non-paying user list" in the sales department have all metrics below 0.1 because they are not directly related to S2 requirements.

[0040] This matrix provides a quantitative basis for subsequent accurate calculation of matching degree. Compared with the existing method that only relies on keyword matching, its indicator dimensions are more comprehensive and the numerical basis is more sufficient, which can more objectively reflect the suitability of data assets and tasks.

[0041] S32 Assign semantic weights The system assigns semantic weights to the semantic constraints of different departments based on the multi-role semantic constraint set formed by S21. The weight allocation comprehensively considers three factors: in the importance ranking of cross-departmental collaboration within the industry, user identity definition is a prerequisite for analysis and usually has a higher weight than behavioral data; in the setting of business priorities within the enterprise, the accuracy of users provided by the sales department directly affects the validity of the analysis conclusions.

[0042] Based on these factors, the system determined the semantic constraint weight for the sales department to be 0.6, and the semantic constraint weight for the R&D department to be 0.4. The customer service department, as a supplementary role, has data that has some auxiliary value for satisfaction analysis. Referring to business experts' value assessment of auxiliary data, its semantic constraint weight is set to 0.5. This way, its role is reflected while avoiding secondary data interfering with the core matching results.

[0043] S33 Weighted Calculation of Overall Matching Degree The system first extracts the three indicator values ​​corresponding to each data asset from the semantic association matrix constructed in S31 based on the composite semantic query conditions in S2. Then, based on the semantic weights assigned in S32, it calculates the overall matching degree using a weighted summation method. The calculation of the overall matching degree is essentially a weighted sum of each indicator value and its corresponding single semantic weight for its department role; that is, for a data asset belonging to a specific department role, its overall matching degree = semantic weight of that department role × (semantic association degree + requirement attribute matching degree + association ability of associated fields). Since all indicators and weights are dimensionless values, the calculation result is also a dimensionless scaled value, and its magnitude reflects the relative degree of matching between the data asset and the task requirements.

[0044] The specific calculation results for each data asset are as follows: The overall matching degree for the sales department's "Paid User List" is 0.9 × 0.6 + 1.0 × 0.6 + 1.0 × 0.6 = 1.74; the overall matching degree for the R&D department's "Product User Charging Data" is 0.8 × 0.4 + 1.0 × 0.4 + 1.0 × 0.4 = 1.12; and the overall matching degree for the customer service department's "Consultation User Satisfaction Data" is 0.7 × 0.5 + 0.5 × 0.5 + 1.0 × 0.5 = 1.1. Due to the characteristics of semantic weight allocation and indicator values, the overall matching degree may be greater than 1. This is a normal result after weighted calculation, and its value effectively reflects the relative matching degree between the data asset and the task requirements. These calculation results provide an intuitive quantitative basis for data filtering.

[0045] S34 sets the matching threshold. like Figure 3 As shown, the system determined the preset matching threshold through simulation tests of 100 different types of cross-departmental collaborative tasks. The tests covered various scenarios such as sales and R&D collaboration, and R&D and customer service collaboration, focusing on the accuracy and recall of recommended data under different thresholds: when the threshold was set to 0.8, the recall rate reached 90% but the accuracy rate was only 75%, resulting in a large amount of redundant data; when the threshold was set to 1.2, the accuracy rate increased to 95% but the recall rate dropped to 70%, making it easy to miss valid data; when the threshold was set to 1.0, the accuracy of recommended data reached 92% and the recall rate reached 85%, which not only ensured the accuracy of recommended data but also covered most of the valid data assets related to the task, which better met the accuracy requirements of data asset recommendation in the industry. Therefore, the system determined the preset matching threshold to be 1.0.

[0046] S4 Integrated Data Asset Package Construction and Recommendation List Generation like Figure 1 As shown, this step uses the comprehensive matching degree and preset threshold obtained from S3 as the core basis to complete the screening, integration and priority ranking of data assets, and finally generate a recommendation list that can directly support cross-departmental analysis.

[0047] S41 Screening Multi-Source Data Assets The system compares the overall matching degree of each department's data assets with the 1.0 threshold set by S34, filtering out multi-source data assets that meet the criteria. Among them, the overall matching degree of the sales department's "Paid User List" is 1.74, the R&D department's "Product User Charging Data" is 1.12, and the customer service department's "Consultation User Satisfaction Data" is 1.1, all three exceeding the threshold. However, non-core data assets such as the sales department's "Non-Paid User List" have an overall matching degree of only 0.06, far below the threshold, and are directly filtered out by the system. This quantitative indicator-based screening method ensures that the selected data assets are highly relevant to the "Paid User Charging Function Satisfaction Analysis" task, laying a high-quality foundation for subsequent integration work.

[0048] S42 Dynamic Assembly and Integration of Data Asset Packages The system invokes the semantic association rules built by S1, using user ID as the core cross-departmental association field to dynamically assemble the multi-source data assets filtered by S41. First, the system extracts user IDs from the sales department's "Paid User List" as filtering identifiers, accurately matching the charging behavior data of corresponding users from the R&D department's "Product User Charging Data." Then, using the user ID, the two sets of data are aligned and merged to form a "Paid User Charging Function Basic Analysis Package," which includes five core fields: user ID, purchase time, vehicle model, number of charging attempts, and function click records. Next, using the same user ID as the association basis, the system extracts the satisfaction rating information of paid users from the customer service department's "Consultation User Satisfaction Data," forming a separate "Paid User Charging Satisfaction Supplementary Package," containing both user ID and satisfaction rating fields. The integrated asset package has completed the cross-departmental data association and cleaning, allowing business personnel to directly use it for analysis without additional manual processing, significantly improving collaboration efficiency.

[0049] S43 Priority sorting generates a recommendation list The system integrates data asset packages as units, calculating the total comprehensive matching degree of the data assets contained in each package, and using this as the core basis for priority ranking. The total comprehensive matching degree of the "Paid User Charging Function Basic Analysis Package" is 1.74 + 1.12 = 2.86, and the total of the "Paid User Charging Satisfaction Supplement Package" is 1.1. The system sorts the packages from highest to lowest total, determining the "Basic Analysis Package" as the first-level recommendation and the "Supplement Package" as the second-level recommendation. If the difference in the totals of multiple integrated packages is less than 0.2, the system will refer to the latest update time of the data assets, prioritizing the asset packages with more recent updates. This ranking logic ensures both data and task compatibility and data timeliness.

[0050] S44 Recommended List Output The system outputs two integrated data asset packages as independent items to the recommendation list. Each independent item contains four core pieces of information to facilitate quick user understanding and use. The "Basic Analysis Package" is defined as "core analysis supporting the usage of the charging function by paid users." The data comes from the sales department's list of paid users and the R&D department's product user charging data. The core related field is the user ID, and the semantic association logic is "based on the S1-defined rule that 'paid users are a subset of product users,' ensuring that data only covers the paid user group through the user ID." The "Supplement Package" is defined as "supplementing paid user satisfaction feedback data on the charging function." The data comes from customer service department's inquiry user satisfaction data, and the related fields and semantic association logic are consistent with the basic package. This clear information presentation allows users to quickly grasp the data value and the basis for its association, further improving data utilization efficiency.

[0051] S5 Feedback Iterative Optimization like Figure 1 As shown, this step uses the recommendation list output by S4 as the core interaction carrier. By collecting user feedback on the recommendation data, the recommendation-related parameters are dynamically adjusted to form a closed-loop mechanism of "recommendation-feedback-optimization".

[0052] S51 collects evaluation feedback The system integrates a structured feedback portal into the recommendation list interface, allowing users to evaluate each integrated data asset package from three core dimensions: semantic accuracy, attribute completeness, and business value. The evaluation uses a 1-5 point rating system, with 1 representing the worst and 5 representing the best. Users can also provide supplementary text explanations or suggestions regarding specific issues.

[0053] S511 clarifies evaluation dimensions and feedback scenarios. The semantic accuracy dimension focuses on the semantic matching degree between data and task requirements. Users can mark issues such as "the basic analysis package contains charging data from non-paying users." The attribute completeness dimension addresses whether the data contains all fields required by the S2 query conditions. Users can report issues such as "the basic analysis package is missing charging data for 2024 models." The business value dimension focuses on the supporting role of data in the analysis conclusions. Users can provide feedback such as "the satisfaction score of the supplementary package directly improves the credibility of the analysis conclusions."

[0054] S512 structured storage feedback data After users submit feedback, the system automatically structures the ratings and text descriptions. The system stores feedback categorized by the unique identifier of the integrated data asset package, while also associating it with the corresponding task scenario, data source, and user role information. For example, the system will bind a "basic analysis package attribute completeness score of 3 points" to the "paid user charging function satisfaction analysis" task and the "sales + R&D" data source, forming a standardized evaluation feedback dataset, providing a precise basis for subsequent parameter adjustments.

[0055] S52 dynamically adjusts relevant parameters The system uses a structured feedback dataset generated by S51 to determine the types and magnitudes of parameters requiring adjustment through statistical analysis. The adjustment logic is based on feedback frequency and impact; when a certain type of feedback occurs in more than 30% of similar tasks, the parameter adjustment process is triggered. This threshold was validated through 50 sets of feedback optimization experiments, demonstrating its ability to strike a balance between responding to changing demands and maintaining system stability.

[0056] S521 Adjusting the priority of data filtering rules If users repeatedly report semantic accuracy issues, such as "recommended data contains information about non-paying users," the system will increase the priority of the "semantic relevance" metric in the S41 data filtering process. The system will increase the filtering weight of semantic relevance by 20% to ensure that data assets with low semantic matching are filtered first in subsequent filtering processes.

[0057] S522 Adjusts Attribute Matching Weights If users collectively report missing attributes, such as "missing vehicle model field affects user segmentation analysis," the system will adjust the attribute matching weights in the S33 comprehensive matching score calculation. Taking the "vehicle model" field as an example, the system will increase the weight of its corresponding attribute matching score from 1.0 to 1.2, allowing data assets containing this field to obtain a higher score in the comprehensive matching score calculation and be more likely to be prioritized in subsequent recommendations.

[0058] S523 Adjusting Semantic Association Confidence If user feedback indicates that the correlation value between two types of data is higher than expected—for example, "the correlation data between paying users and consulting users is crucial for satisfaction attribution"—the system will adjust the confidence level of the corresponding relationships in the S1 semantic association network. The system increases the confidence level of the correlation between "paying users" and "consulting users" from 0.6 to 0.7, enhancing the influence of this correlation in subsequent semantic alignment and data assembly stages, making the correlated data easier to integrate and recommend. After feedback optimization, user satisfaction with the recommended data significantly improves, and the actual utilization rate of data assets also increases, further highlighting the practical value of the closed-loop optimization mechanism.

[0059] The implementation principle of a data asset discoverability enhancement and recommendation method according to an embodiment of this application is as follows: First, this application constructs a cross-departmental unified semantic model that supports dynamic semantic mapping, establishing a unified semantic benchmark and association rules for data tags from different departments, thereby directly breaking down semantic barriers between departments; then, based on this model, by performing context awareness and semantic decoupling on collaborative tasks, it accurately extracts and integrates the differentiated needs of each party's roles, generating composite semantic query conditions that can simultaneously constrain multiple data sources, effectively solving the problem that heterogeneous needs cannot be dynamically aligned and uniformly expressed; subsequently, based on these composite conditions, it calculates the comprehensive matching degree of data assets through a pre-constructed semantic association matrix, and uses this to dynamically assemble the selected multi-source data assets into an integrated data asset package that can be directly used for cross-departmental analysis for priority recommendation, thereby fundamentally changing the situation where traditional recommendation results are fragmented and cannot directly support integrated analysis; in addition, by collecting user feedback and dynamically optimizing the model and recommendation parameters, a closed-loop mechanism for continuously improving recommendation accuracy and business value is formed.

[0060] The above are all preferred embodiments of this application, and are not intended to limit the scope of protection of this application. Therefore, all equivalent changes made in accordance with the structure, shape and principle of this application should be covered within the scope of protection of this application.

Claims

1. A method for enhancing the discoverability and recommendation of data assets, characterized in that, include: Construct a cross-departmental unified semantic model that supports dynamic semantic mapping. The unified semantic model defines the hierarchical relationship and semantic association rules between data tags through a semantic ontology. Based on the unified semantic model, context awareness and semantic decoupling are performed on cross-departmental collaborative tasks to extract the differentiated semantic requirements of each party's role, and these requirements are integrated into a composite semantic query condition that can simultaneously constrain multiple related data sources and their specific attributes, and contains cross-departmental related fields. Based on the composite semantic query conditions, the comprehensive matching degree between data assets from different departments and the current task context is calculated through a pre-constructed semantic association matrix. Based on the comprehensive matching degree, multi-source data assets that meet the threshold conditions are dynamically assembled into an integrated data asset package that can be directly used for cross-departmental analysis and prioritized for recommendation.

2. The method according to claim 1, characterized in that, The construction of the unified semantic model is specifically as follows: based on the semantic ontology, a general core class and specific subclasses corresponding to departments are defined; a semantic hierarchy and semantic association relationship are established between the core class and subclasses, as well as between different subclasses; and each subclass is labeled with its corresponding specific data attribute fields to form a structured and scalable semantic association network.

3. The method according to claim 2, characterized in that, When a new business tag is added, the unified semantic model performs dynamic semantic mapping. The dynamic semantic mapping includes: the unified semantic model automatically extracts the semantics of the new business tag, associates it with the corresponding specific subclass that already exists in the unified semantic model, and synchronously updates the semantic association relationship and the corresponding data attribute fields.

4. The method according to claim 1, characterized in that, The context awareness and semantic decoupling are specifically as follows: based on the hierarchical relationship and semantic association rules defined by the unified semantic model, the cross-departmental collaborative task is parsed to identify the roles of participating departments and their collaborative goals, and to separate the differentiated data subclasses, attribute requirements and associated field constraints corresponding to each role, thereby forming a structured set of multi-role semantic constraints.

5. The method according to claim 4, characterized in that, The specific method of integrating requirements into composite semantic query conditions is as follows: based on the semantic association rules defined by the unified semantic model, semantic alignment and logical association are performed on each constraint in the multi-role semantic constraint set to generate composite semantic query conditions.

6. The method according to claim 5, characterized in that, The calculation of the comprehensive matching degree is specifically as follows: based on the composite semantic query conditions, the semantic relevance, demand attribute matching degree and related field association ability of the corresponding data assets are obtained from the pre-constructed semantic association matrix, and the semantic constraint information represented by the multi-role semantic constraint set is fused and calculated to obtain the comprehensive matching degree.

7. The method according to claim 6, characterized in that, The specific steps for performing the fusion calculation by combining the semantic constraint information represented by the multi-role semantic constraint set are as follows: based on the multi-role semantic constraint set, semantic weights are assigned to the semantic constraints of different roles, and the semantic relevance, demand attribute matching degree and relevance ability of the associated fields obtained from the semantic association matrix are weighted according to the semantic weights to complete the fusion of the comprehensive matching degree.

8. The method according to claim 7, characterized in that, It also includes the step of generating the recommendation list: Based on the comprehensive matching degree calculated by the weighted average, multi-source data assets that meet the preset matching degree threshold are selected from data assets from different departments; Based on the semantic association rules and the association fields, the selected multi-source data assets are dynamically assembled into an integrated data asset package; Based on the total matching degree of all data assets contained in each integrated data asset package, priority is ranked, and a recommendation list is generated and output.

9. The method according to claim 8, characterized in that, Also includes: Based on the recommendation list, user feedback on the semantic accuracy, attribute completeness, and business value dimensions of the recommendation data is collected in a structured manner. Based on the evaluation feedback, the priority of data filtering rules, attribute matching weights, and confidence levels of object attributes in the unified semantic model are dynamically adjusted to generate subsequent recommendation lists.

10. The method according to claim 8, characterized in that, The generation and output of the recommendation list also includes: Each integrated data asset package will be output as an independent item; Each independent item includes: a semantic purpose description of the integrated data asset package, source information of the data assets that make up the package, a list of related fields used for cross-departmental association, and the semantic association logic on which the related fields are based.