A data asset intelligent recommendation method combined with user behavior

CN122777779APending Publication Date: 2026-09-18NANJING LINGYU INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610720588.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-25
Publication Date
2026-09-18

AI Technical Summary

Technical Problem

[0008]本发明的一个目的在于提出一种结合用户行为的数据资产智能推荐方法,针对现有技术中权限控制与多租户隔离未融入推荐链路、导致推荐结果“看得见用不了”及存在越权暴露风险的问题,提出了将访问控制属性与向量召回、候选集构造与排序输出一体化融合的技术方案

Benefits of technology

[0055] 1. By associating tenant identifiers, access control rules, and permission tokens with asset vectors, and enabling near nearest neighbor retrieval to skip or prune based on filtering conditions during the retrieval process, multi-tenant isolation and access control are achieved during the recommendation recall stage, reducing invalid recommendations that are "visible but unusable" and lowering the risk of unauthorized exposure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122777779A_ABST
    Figure CN122777779A_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of data asset management and intelligent recommendation, and discloses a data asset intelligent recommendation method combined with user behaviors, to solve the problem of visible but unusable or over-power exposure caused by the fact that permission control and multi-tenant isolation are not integrated into the recommendation link, the application constructs a support filtering approximate nearest neighbor index by associating a data asset vector with a tenant identifier, an access control rule and a permission token, generates a user vector and a filtering condition based on user behaviors and access control information, and performs filtering and constraint learning in the retrieval and sorting stage, so that the technical effect of outputting only a recommendation list meeting the requirements of tenant isolation and access control, improving recommendation availability and reducing the risk of over-power exposure is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data asset management, and in particular to a method for intelligent recommendation of data assets that combines user behavior. Background Technology

[0002] With the development of enterprise data platforms, data catalogs, and data asset management platforms, organizations typically implement unified metadata management for data assets such as tables, views, metrics, reports, and APIs. They also combine search, tagging, and recommendation technologies to improve the discovery efficiency and reusability of data assets. Existing data asset recommendation solutions often utilize user behavior logs (search, clicks, previews, downloads, calls, or queries) to build user profiles, and achieve recall based on vectorized representations and Approximate Nearest Neighbor (ANN) retrieval. Finally, a ranking model is used to output a recommendation list. Meanwhile, the trends of cloudification and SaaS have led to the widespread adoption of multi-tenant architectures in data platforms. Access control rules are set for different tenants, roles, and resources through permission management systems to meet isolation and compliance requirements.

[0003] However, existing technologies still have the following shortcomings:

[0004] 1. Permissions and multi-tenant isolation are often not deeply integrated with the vector recall process. The common practice is to filter permissions after recall or before display, which can easily lead to invalid recommendations that are "visible but unusable" and may even cause the risk of unauthorized exposure in the middle of the process.

[0005] 2. When the recommended index or model is not synchronized with the permission policy, the problem of recommending inaccessible assets may occur even after the permission changes. This requires additional secondary verification, which increases latency and system complexity.

[0006] 3. Ranking model training usually focuses on behavioral relevance and lacks explicit learning and penalty mechanisms for access constraints, resulting in insufficient usability and stability of ranking results under permission constraints.

[0007] Therefore, there is a need for an intelligent recommendation method for data assets that can address the shortcomings of existing technologies. Summary of the Invention

[0008] One objective of this invention is to propose an intelligent data asset recommendation method that integrates user behavior. Addressing the problems in existing technologies where access control and multi-tenant isolation are not integrated into the recommendation process, resulting in "visible but unusable" recommendation results and risks of unauthorized access, this invention proposes a technical solution that integrates access control attributes with vector retrieval, candidate set construction, and ranking output. This invention obtains data asset metadata and its tenant identifiers, access control rules, and maps them to generate permission tokens. It then associates asset vectors with access control attributes to construct a near-nearest neighbor index that supports filtering. It obtains target user behavior to generate user vectors and generates filtering conditions containing tenant identifiers, access constraints, and user-side permission tokens based on user-side access control information. During the ANN retrieval process, it performs skipping or pruning based on the filtering conditions to obtain a candidate asset set. In the ranking stage, it penalizes or removes inaccessible candidates based on the filtering conditions, and triggers secondary verification and incremental updates when necessary, combined with permission version numbers. This invention achieves the technical effects of satisfying tenant isolation and access control requirements throughout the entire retrieval and ranking process, improving the usability of recommendation results, and reducing the risk of unauthorized access.

[0009] This invention provides an intelligent recommendation method for data assets that combines user behavior, including:

[0010] S1. Obtain data asset metadata and access control attributes corresponding to the data assets. Based on the data asset metadata, generate asset vectors corresponding to each data asset through the asset encoder, and associate each asset vector with its corresponding access control attributes to obtain an asset vector set.

[0011] S2. Construct an approximate nearest neighbor index based on the asset vector set, enabling the approximate nearest neighbor index to perform filtering retrieval based on access control attributes during the vector retrieval process;

[0012] S3. Obtain the target user's current user behavior data. Generate a user vector through the user encoder, obtain the target user's access control information, generate filtering conditions, including tenant identifier and access constraints, and obtain the retrieval request.

[0013] S4. Based on the retrieval request, perform a filtering retrieval in the near nearest neighbor index to obtain a set of candidate assets that meet the filtering conditions;

[0014] S5. Taking the candidate asset set and current user behavior data as input, the ranking results of each candidate asset are output through the ranking model. During or after the ranking process, candidate assets that do not meet the accessibility constraints are restricted or eliminated according to the filtering conditions, and a recommended list of data assets that meet the requirements of tenant isolation and access control is output.

[0015] Optionally, S1 includes:

[0016] Obtaining the user encoder, asset encoder, and sorting model includes reading the user encoder, asset encoder, and sorting model from the storage medium for subsequent use;

[0017] Obtaining data asset metadata and access control attributes corresponding to data assets includes determining data asset metadata and access control attributes corresponding to each data asset, wherein the access control attributes include tenant identifier, access control rules, and permission tokens generated based on at least one of the tenant identifier and access control rules.

[0018] The permission token is a fixed-length identifier used to represent the scope of access, and the fixed-length identifier includes any one or more of the following: bitmap token, Bloom filter token, hash token, or role-resource pair encoded token;

[0019] The asset encoder generates asset vectors for each data asset by using the data asset metadata as input, and associates each asset vector with its corresponding access control attribute to obtain an asset vector set. This includes inputting the data asset metadata corresponding to each data asset into the asset encoder to output the asset vector of the data asset, and establishing a correspondence between the asset vector and the access control attribute of the data asset, thereby aggregating the asset vectors to obtain the asset vector set.

[0020] Optionally, S2 includes:

[0021] Constructing an approximate nearest neighbor index using the asset vector set as input includes: generating an index entry for each asset vector in the asset vector set, wherein the index entry includes a data asset identifier, the asset vector, and an access control attribute associated with the asset vector;

[0022] The access control attributes include at least the tenant identifier, access control rules, and permission tokens;

[0023] Based on all the index entries, a vector index structure for approximate nearest neighbor retrieval is constructed, wherein the vector index structure is a hierarchical small-world graph (HNSW) index structure or an inverted file (IVF) index structure.

[0024] Simultaneously, a filtering index structure is constructed based on the access control attributes. The filtering index structure is at least used to determine the range of searchable index entries based on the tenant identifier, and to filter index entries that meet the access control requirements from the range of searchable index entries based on at least one of the access control rules and permission tokens.

[0025] The filtering index structure includes a partitioned structure built by tenant identifier, and at least one of an inverted index, bitmap, or bucketed structure built by permission token;

[0026] The vector index structure is associated with and stored with the filter index structure to obtain the approximate nearest neighbor index. This ensures that during subsequent filtering retrieval, only index entries that meet the filtering conditions are used for approximate nearest neighbor retrieval and retrieval results are output during the retrieval process of the vector index structure.

[0027] The filtering retrieval includes skipping or pruning index entries that do not meet the filtering conditions during the graph traversal process of HNSW or the inverted linked list scan process of IVF, based on the filtering conditions.

[0028] Furthermore, a permission version number is set for the access control attribute, and when a change in the access control attribute of a data asset is detected in the permission management system, an incremental update is performed on the access control attribute and filtering index structure of the corresponding index entry based on the permission version number.

[0029] When performing a filtering retrieval, if the permission version number of the filtering condition is inconsistent with the permission version number of the index entry, the index entry is marked as an entry to be verified and the access control attributes are re-fetched or a second verification is performed before deciding whether to output it.

[0030] Furthermore, when the vector index structure is a hierarchical small-world graph (HNSW) index structure, filterable tags corresponding to access control attributes are maintained for the HNSW entry point and each layer node.

[0031] Before graph traversal begins, a set of entry points that satisfy the filtering conditions is selected based on the filtering conditions.

[0032] Furthermore, during the traversal, when there are multiple candidate adjacent nodes available for expansion, adjacent nodes with filterable tags that satisfy the filtering conditions are expanded first, so as to reduce the number of times nodes that do not meet the filtering conditions are visited.

[0033] Optionally, S3 includes:

[0034] Acquiring the target user's current user behavior data and the target user's current access control information includes collecting the target user's search, click, preview, download, call, or query behavior records for data assets within a preset time window as the current user behavior data, and obtaining the target user's tenant identifier and access control rules from the permission management system as the current access control information;

[0035] Generating a user vector using the user encoder by taking the current user behavior data and the current access control information as input includes organizing the current user behavior data into a user-side input in chronological order and inputting it together with the current access control information into the user encoder to output the user vector.

[0036] Generating filtering conditions based on the current access control information includes extracting the tenant identifier from the current access control information and determining access constraints based on the access control rules to form the filtering conditions, and generating a user-side permission token based on at least one of the tenant identifier and the access control rules, and adding the user-side permission token to the filtering conditions.

[0037] The user vector is combined with the filtering conditions to obtain the search request.

[0038] Optionally, S4 includes:

[0039] Using the search request as input to call the approximate nearest neighbor index to perform a filtered search includes: inputting the filtering conditions in the search request into the approximate nearest neighbor index, so that the approximate nearest neighbor index determines a set of target index entries that satisfy the tenant identifier and the accessibility constraints based on the filtering conditions in the index entries;

[0040] Determining the target set of index entries includes filtering index entries based on the matching relationship between the user-side permission tokens in the filtering conditions and the permission tokens in the index entries;

[0041] Within the scope defined by the target index entry set, an approximate nearest neighbor search is performed in the approximate nearest neighbor index using the user vector in the search request as the query vector, to obtain search hit entries that meet the preset conditions in terms of similarity with the user vector;

[0042] Data asset identifiers are extracted from the retrieved entries to form the candidate asset set.

[0043] Optionally, S5 includes:

[0044] The process of ranking the candidate asset set, the current user behavior data, and the filtering conditions using the ranking model as input includes: obtaining the asset vector corresponding to each data asset from the candidate asset set, and inputting the current user behavior data, the asset vector, and the filtering conditions into the ranking model to output the ranking score corresponding to each data asset.

[0045] Based on the filtering conditions, it is determined whether each data asset meets the accessibility constraints, and a ranking score limit is imposed on the data assets that do not meet the accessibility constraints. The ranking score limit is to set the corresponding ranking score to a penalty score that is less than a preset recommended output threshold, or to a negative infinity approximation value.

[0046] The candidate asset set is sorted based on the ranking scores of each data asset, and when the recommendation result is output, data assets with ranking scores lower than the preset recommendation output threshold are removed from the sorted candidate asset set to obtain the recommendation result.

[0047] Optionally, the user encoder, the asset encoder, and the ranking model are obtained through the following steps:

[0048] T1. Obtain historical user behavior data, historical access control information, and data asset metadata. Generate training access control attributes based on the historical access control information, and establish the accessibility relationship between users and data assets based on the training access control attributes to obtain a training dataset. The training access control attributes include tenant identifiers, access control rules, and training permission tokens generated based on the mapping of the tenant identifiers and / or access control rules.

[0049] T2. Using the training dataset as input, aggregate the historical user behavior data according to the user identifier to generate a user behavior sequence, and construct sample pairs based on the user behavior sequence to obtain a dual-tower model training sample set. Each sample pair includes user-side input, asset-side input, and sample label. The user-side input includes the user behavior sequence and the training access control attribute corresponding to the user. The asset-side input includes data asset metadata corresponding to the data asset and the training access control attribute corresponding to the data asset.

[0050] T3. Using the training sample set of the dual-tower model as input, train the dual-tower model to obtain the user encoder and the asset encoder. Encode the user-side input through the user encoder to obtain the user vector, and encode the asset-side input through the asset encoder to obtain the asset vector. Based on the sample labels, increase the similarity between the user vector and the asset vector of positive sample pairs and decrease the similarity between the user vector and the asset vector of negative sample pairs to obtain the user encoder and the asset encoder.

[0051] T4. Using the user encoder, the asset encoder, and the training dataset as input, generate a ranking model training sample set for the user behavior records in the training dataset. For each user behavior record, use the user encoder to generate a training user vector and the asset encoder to generate a training asset vector. Construct a candidate asset set for the user behavior record based on the accessibility relationship, including data assets that satisfy the accessibility relationship and data assets that do not satisfy the accessibility relationship.

[0052] T5. The ranking model is trained using the training sample set as input to obtain the ranking model; wherein the ranking model takes the training user vector, the training asset vector, and the training access control attribute as model input, outputs the ranking score of the data assets, and imposes ranking score restrictions on data assets that do not meet the accessibility relationship based on the accessibility relationship during the training process, so that the ranking model outputs a restricted ranking score for data assets that do not meet the accessibility relationship.

[0053] Furthermore, the construction of the candidate asset set includes: inputting the training user vector and the filtering conditions generated by the training access control attributes into the approximate nearest neighbor index to perform filtering retrieval, thereby generating a candidate asset set consistent with the online stage, and using the candidate asset set for the training sample set of the ranking model.

[0054] The beneficial effects of this invention are:

[0055] 1. By associating tenant identifiers, access control rules, and permission tokens with asset vectors, and enabling near nearest neighbor retrieval to skip or prune based on filtering conditions during the retrieval process, multi-tenant isolation and access control are achieved during the recommendation recall stage, reducing invalid recommendations that are "visible but unusable" and lowering the risk of unauthorized exposure.

[0056] 2. During the ranking stage, filter conditions are incorporated into the ranking model input, and penalty scores are applied to candidate assets that do not meet accessibility constraints or they are removed at the output according to a threshold. This ensures that the recommendation results maintain high availability and stable relevance under permission constraints, thereby improving the effective hit rate of the recommendation list and the user experience.

[0057] 3. By setting permission version numbers for access control attributes and performing incremental updates to index entries and filtering index structures when permissions change, and by triggering re-fetching or secondary verification for entries with inconsistent versions during filtering retrieval, the system reduces erroneous recommendations caused by the asynchrony between permission policies and indexes, thereby improving system security and consistency. Attached Figure Description

[0058] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0059] Figure 1 This is a flowchart of a data asset intelligent recommendation method that combines user behavior, as proposed in this invention.

[0060] Figure 2 A schematic flowchart illustrating step S2 of the present invention for constructing an approximate nearest neighbor index that supports filtered retrieval. Detailed Implementation

[0061] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0062] refer to Figure 1-2 A data asset-based intelligent recommendation method that combines user behavior includes:

[0063] S1. Obtain data asset metadata and access control attributes corresponding to the data assets. Based on the data asset metadata, generate asset vectors corresponding to each data asset through the asset encoder, and associate each asset vector with its corresponding access control attributes to obtain an asset vector set.

[0064] S2. Construct an approximate nearest neighbor index based on the asset vector set, enabling the approximate nearest neighbor index to perform filtering retrieval based on access control attributes during the vector retrieval process;

[0065] S3. Obtain the target user's current user behavior data. Generate a user vector through the user encoder, obtain the target user's access control information, generate filtering conditions, including tenant identifier and access constraints, and obtain the retrieval request.

[0066] S4. Based on the retrieval request, perform a filtering retrieval in the near nearest neighbor index to obtain a set of candidate assets that meet the filtering conditions;

[0067] S5. Taking the candidate asset set and current user behavior data as input, the ranking results of each candidate asset are output through the ranking model. During or after the ranking process, candidate assets that do not meet the accessibility constraints are restricted or eliminated according to the filtering conditions, and a recommended list of data assets that meet the requirements of tenant isolation and access control is output.

[0068] In this specific embodiment, S1 includes:

[0069] The user encoder, asset encoder, and ranking model are read from a storage medium (either a local file system or object storage) and loaded into callable format. The model files are stored in the form of a "network structure description file + parameter weight file." During loading, the structure description file is parsed to construct a computation graph, and the parameter weights are loaded into GPU memory or RAM in floating-point 16 format to obtain an executable inference instance. The asset encoder uses a text Transformer structure with 12 layers, 768 hidden layer dimensions, 12 attention heads, and a maximum input length of 256 sub-tokens. A linear projection layer is concatenated at the Transformer output to compress the vector dimension. To form an asset vector output interface;

[0070] For each data asset, corresponding data asset metadata and access control attributes are determined. The data asset metadata includes asset identifier, asset name, asset type, asset description, business tag, field or indicator definition, lineage summary, update frequency, and responsible person identifier. The above metadata is concatenated into metadata text to be encoded in a fixed field order, and excessive parts are truncated to the maximum input length to ensure input determinism. The access control attributes include tenant identifier, access control rules, and permission tokens generated based on the mapping of the tenant identifier and access control rules. The tenant identifier is a string-type unique tenant identifier and is written into the tenant field of the access control attributes. The access control rules are represented by an Access Control List (ACL) structure, and each rule consists of a subject identifier, a resource identifier, an action set, and a permission effect. The subject identifier is either a user identifier or a role identifier, the resource identifier is the data asset identifier, the action set is a subset of the five types of actions: "search, preview, download, call, and query", and the permission effect is either allow or deny, and the final access constraints of the data asset for each subject are obtained by merging rules with denial as the priority.

[0071] A fixed-length permission token is generated based on the tenant identifier and access control rules. This permission token is a bitmap token with a fixed length. Bit, the first bit of the bitmap Each bit corresponds to a type of "role-resource-action" permission unit, and a one-to-one correspondence is achieved through a deterministic mapping table. The mapping table is exported by the permission management system during system initialization and is stored in a fixed form with the key being a triplet of "role identifier, resource identifier, action" and the value being a bit sequence number. When generating a permission token, all triplets with the permission effect being allowed in the access control rules of the data asset are traversed and the corresponding bit sequence number is set to 1, while the remaining bits are kept to 0. This allows the permission token to determine whether the user-side permission token covers the action set required by the asset-side permission token through bit operations during subsequent matching.

[0072] For each data asset, the metadata text The asset encoder is input to generate asset vectors, and the asset vectors are then mapped to access control attributes. The asset vector generation follows a deterministic calculation process:

[0073]

[0074] in, This indicates the sequence number of the data asset in the current batch. Indicates the first Metadata text of a data asset, Indicates will After inputting the asset encoder, the semantic representation vector is obtained by taking the output of the last Transformer layer at the first token position. This represents the weight matrix of the linear projection layer, and its element values ​​are obtained through offline training and are fixed with the model file. This represents the bias vector of the linear projection layer, and its elements are obtained through offline training and are fixed with the model file. Indicates performing an operation on the input vector. Normalization is performed to make the output asset vector satisfy the unit norm constraint, representing the first... The asset vector of each data asset;

[0075] When constructing an asset vector set, a structured record is formed using the data asset identifier as the primary key and written to an in-memory table or a persistent table. The record fields include at least the data asset identifier, asset vector, tenant identifier, access control rules, and permission token, thereby obtaining an asset vector set in which "asset vectors are associated with access control attributes one-to-one".

[0076] In this specific embodiment, S2 includes:

[0077] Each data asset is assigned an index entry using a set of asset vectors as input. Primary key and containing asset vectors Tenant identification Access control rules Permission token and permission version number ,in, It is a unit norm vector and stored in floating-point 16 format. A unique identifier for a string-type tenant. A set of rules represented by an Access Control List (ACL). For fixed-length bitmap tokens and It is stored by concatenating 32 unsigned 64-bit integers in sequence. It is a non-negative integer and is incremented by the access control attribute change event of the data asset by the permission management system;

[0078] A vector index structure is constructed based on all index entries, and a hierarchical small-world graph (HNSW) is selected as the main index structure for the approximate nearest neighbor index. The distance metric used in HNSW is cosine distance, and because... Since the vectors have been normalized, equivalent similarity comparisons are achieved by sorting the vector inner product. The HNSW construction parameter is set to the maximum number of adjacencies. , width of candidate queue during construction phase efConstruction Candidate queue width during the retrieval phase efSearch and using a fixed random seed The pseudo-random number generator assigns a hierarchy to each node to ensure that the index construction result is determined under the same input set;

[0079] A filtering index structure is constructed based on access control attributes. The filtering index structure includes a partition structure built by tenant identifier and a bitmap inverted index structure built by permission token. The partition structure maintains the information from the tenant identifier. Mapping to the set of node identifiers RoaringBitmap is used for storage to support addition, deletion, intersection, union, and difference operations. The bitmap inverted index structure stores each bit of the permission token. Maintain an inverted ranking chart and Record all satisfied The node identifier is used to support differential maintenance when the permission token is updated;

[0080] The vector index structure and the filter index structure are associated and stored as the same approximate nearest neighbor index instance. The association method is to permanently save the corresponding node in the HNSW node data structure. and node identifiers and The mapping table enables subsequent filtering and retrieval to skip or prune nodes that do not meet the filtering conditions during the graph traversal of HNSW.

[0081] To reduce invalid access overhead during filtering retrieval on HNSW, filterable tags corresponding to access control attributes are maintained at the HNSW entry point and at each layer node. These filterable tags are identified by the tenant. Signing with permission tokens Composition, in which An integer array of length 8, sorted in ascending order. The first 8 bits set to 1 are sequenced; if there are fewer than 8 bits, they are padded with -1. Used to quickly determine whether the permission is obviously not satisfied in constant time during traversal and to expand priority sorting;

[0082] Regarding entry point maintenance, each tenant is identified. Maintain entry point set The first node in the tenant partition with the highest level and the largest degree. The index is composed of several nodes and stored permanently after the index is built, allowing the selection of the entry point set directly based on the tenant identifier in the filtering conditions before the retrieval begins. As an initial candidate;

[0083] Regarding traversal pruning and expansion priority, tenant consistency judgment is first performed on any visited candidate node. Based on Determine whether each non--1 bit sequence number in the record is 1 in the user-side permission token. If the condition is not met, skip directly without proceeding to the more costly full permission matching. For those that are met, perform full permission matching and use the matching result to determine whether to participate in the candidate queue maintenance. The full permission matching satisfies the following judgment formula:

[0084]

[0085] in, This indicates the data asset sequence number corresponding to the node. This indicates the bitmap of asset-side permission tokens stored on this node, and This represents a bitmap of the user-side permission token carried in the retrieval request. Indicates bitwise AND operation and... Each bit is calculated individually. If the equation holds true, it means that the user-side permission token covers the permission unit required by the asset-side permission token, thus satisfying the accessibility constraint.

[0086] When multiple candidate adjacent nodes exist for expansion, the expansion priority follows the principle of "tenant consistency and..." Nodes that pass through quickly take precedence over those that are only tenant-consistent. The rule of "nodes that fail to pass and nodes that do not meet either condition will not be expanded" is executed, thereby reducing the number of times nodes that do not meet the filtering conditions are accessed.

[0087] Regarding the consistency of permission version numbers, User-side permission version number carried by the filter conditions This is incorporated into the index-side verification process. When the traversed node matches the tenant and permission token but an error occurs... The node is then marked as an entry to be verified, and the access control attributes corresponding to that node are retrieved from the access control management system to refresh the system. After refreshing, the tenant and permission token matching is performed again, and the node is only allowed to enter the search results if the match is successful;

[0088] Regarding incremental updates, when the access control system detects a change in the access control attribute of a data asset, it updates the access control version number accordingly. Perform in-place updates on index entries and differential updates on the filtered index structure. The differential update process starts from the old permission token. With new permission token Extract the set of bit indices that have changed from the bit difference set and perform corresponding operations. The deletion or addition of row node identifiers, and the transfer of node identifiers from / to / from the tenant identifier when the tenant identifier changes. Delete and add And synchronously recalculate the node's permission token signature. With population point set The members ensure that the selection of entry points and the expansion priority rules for subsequent filtering and retrieval remain effective.

[0089] In this specific embodiment, S3 includes:

[0090] Based on target user identifier Extracting data from the log system at the current time. The endpoint is and the length is User behavior records within a 24-day time window are used to form current user behavior data. The user behavior records are limited to six types of events: searching, clicking, previewing, downloading, calling, and querying data assets. Each record contains an event type, an event timestamp, and an associated data asset identifier or search keyword. When the event type is search, the search keyword is used as the content field of the record and the associated data asset identifier is set to empty.

[0091] Sort the behavior records within the time window by event timestamp from most recent to oldest, and extract the most recent one. The user behavior sequence is constructed by means of a user-side input token sequence and satisfies a fixed maximum length. The first two tokens are the tenant token and the permission token, respectively. The subsequent... Each token corresponds to... One record of behavior;

[0092] The current access control information corresponding to the UID is read from the access control system. The current access control information includes the tenant identifier. Access control rules The Access Control List (ACL) structure is used, and each rule consists of a subject identifier, a resource identifier, a set of actions, and a permitted effect. A rule merging strategy with denial priority is used to obtain the final access constraints for each resource. It is a non-negative integer and is incremented by the permission management system when the user's permission set changes;

[0093] based on and Generate user-side permission tokens The with steps With steps Using a deterministic mapping table with the same "role identifier, resource identifier, action" triplet position number, the generation process iterates through the... In the merged result, the corresponding bit sequence number of the triplet with the permission effect is set to 1, and the corresponding bit sequence number of the triplet with the permission effect is forcibly set to 0 to achieve the rejection priority;

[0094] Filtering conditions are generated based on the current access control information. The filtering conditions Identified by tenant User-side permission token Permission version number Together with the accessibility constraints extracted from the merging results, the accessibility constraints are constraints on the scope of resources and the scope of actions. The resource scope is jointly defined by the set of allowed data asset identifiers and the set of allowed data asset types, and the action scope is defined by the set of allowed actions. Serialized into structured parameters for subsequent filtering and retrieval using the approximate nearest neighbor index;

[0095] Call the user encoder to generate user vectors The user encoder is a sequence Transformer structure with 12 layers, 768 hidden layer dimensions, 12 attention heads, and a maximum input length of [missing information]. Its input embedding is obtained by adding three parts while maintaining a dimension of 768. The three parts are the token type embedding, content embedding, and location embedding, where the content embedding of the tenant token is determined by the tenant identifier. The permission token is obtained by looking up a table in the tenant dictionary, and the dictionary size is equal to the number of registered tenants in the system. It is also fixed in the model file. The content of the permission token is embedded from the user-side permission token. First, expand it bit by bit to a length of 2048. The vector is then input into a fully connected layer with an input dimension of 2048 and an output dimension of 768. The content embedding of the token corresponding to each behavior record is obtained by adding the event type embedding, event time embedding, and event object embedding. The event type embedding is obtained by looking up a table from a fixed dictionary of six types of events, and the event time embedding is obtained by combining the event timestamp with the current time. The time difference is obtained by binning by minutes and looking up the table, with a bin count of 1440. The event object is embedded in the behavior record associated data asset identifier when it is not empty, and the data asset is used in the step. Asset vectors generated and stored in The linear mapping layer with an input dimension of 256 and an output dimension of 768 is used to obtain a 768-dimensional semantic vector, and the parameters of the linear mapping layer are fixed with the user encoder model file. When the behavior record is a search event, the search keywords are tokenized by sub-words and input into the text Transformer with the same structure as the asset encoder to obtain a 768-dimensional semantic vector, which is then embedded as an event object.

[0096] The user vector The user encoder outputs the last layer vector at the tenant token location. Linear projection and normalization yield the result that satisfies:

[0097]

[0098] in, Represents a user vector. express Normalization operators are used to make the output vector satisfy the identity norm constraint. This represents the weight matrix of the user vector projection layer, and its element values ​​are obtained through offline training and are fixed with the model file. This represents the output vector of the last layer of the user encoder at the tenant token location. This represents the bias vector of the user vector projection layer, and its element values ​​are obtained from offline training and are fixed with the model file.

[0099] User vectors With filtration conditions Combine to generate search requests The search request Included as a structure field Accessibility constraints.

[0100] In this specific embodiment, S4 includes:

[0101] Search request Input an instance of the approximate nearest neighbor index, wherein the retrieval request Includes user vectors Tenant identification User-side permission token Permission version number And access constraints, using the above fields as filtering conditions. Parameterized input;

[0102] Approximate nearest neighbor index based on filtering conditions Determine the set of target index entries that satisfy the tenant identifier and the accessibility constraints from the index entries, specifically by first reading through the partition structure. This serves as the set of entry points for HNSW traversal, limiting the entry points to tenant partitions. To achieve tenant isolation, during graph traversal, each accessed node undergoes a four-level filtering process: "tenant consistency check, quick permission token check, complete permission token match, and permission version number consistency check." The tenant consistency check requires that the tenant identifier recorded by the node equals... The permission token is quickly verified using the node's permission token signature. right A constant-time coverage check is performed, and nodes that fail the check are skipped to prevent them from entering similarity calculation and queue maintenance. For nodes that pass the fast verification, the asset-side permission token is retrieved based on a complete permission token match. Perform a bitwise overriding relationship check and consider the node to satisfy the accessibility constraint only if the overriding relationship is true. The permission version number consistency check compares the node's permission version number after the overriding relationship is established. And in The node is then marked as an entry to be verified, and the access control attributes corresponding to the data asset are retrieved from the access control management system to refresh the data. After refreshing, the tenant consistency check is performed again to fully match the permission token, and the node is only allowed to participate in the hit output if it passes the check.

[0103] Within the target index entries that meet the above filtering criteria The HNSW approximate nearest neighbor search is performed using the query vector, and the candidate queue width is fixed during the search phase. Furthermore, when traversing and expanding adjacent nodes, similarity is calculated only for nodes that pass the screening, and a "candidate queue and optimal queue" are maintained. The similarity is achieved using cosine similarity implemented with inner product, and is maintained in descending order of score. The search returned 100 results, of which This indicates the upper limit of the candidate asset set's capacity and that the similarity meets the threshold. Only nodes that meet the "similarity satisfies a preset condition" are allowed to enter the optimal queue. The similarity is calculated as follows:

[0104]

[0105] in, Indicates the first The similarity score of each retrieved item. Indicates the sequence number of the index entry. This indicates that the user vector carried in the retrieval request is a unit norm vector. This represents the asset vector stored in the index entry, and it is a unit norm vector. Indicates the transpose operation and The vector dot product under the unit norm condition is equivalent to cosine similarity;

[0106] Extract the data asset identifier from the retrieved entries. And sort by similarity score and extract the first few. A set of candidate assets is formed. The candidate asset set It only includes data asset identifiers that simultaneously meet the constraints of tenant isolation, permission token matching, permission version verification, and similarity threshold.

[0107] In this specific embodiment, S5 includes:

[0108] Set of candidate assets Current user behavior data and filtering conditions The candidate asset set is used as input to invoke the ranking model and enter the candidate-by-candidate scoring process. It consists of several data asset identifiers and has a maximum capacity of Filtering conditions Includes tenant identifier User-side permission token Permission version number Accessibility constraints;

[0109] Identify each candidate data asset Read the asset vector corresponding to the candidate from the asset vector set. And simultaneously read its tenant identifier Permission token With permission version number As an authentication input for the sorting side, it ensures that the sorting stage can perform the same accessibility judgment on candidate assets as the recall stage.

[0110] Construct the behavior feature vector of the ranking model from the current user behavior data. The behavioral features are constructed using time windows. Recently within the day Each behavior record counts six types of events and calculates the time since the most recent behavior. The minute difference is used to obtain a numerical feature vector of length 7, which is then mapped to a fully connected network. The fully connected network has an input dimension of 7, an output dimension of 32, an activation function of ReLU, and its parameters are fixed as the sorted model file is processed.

[0111] Set filter conditions Filtered feature vectors constructed into a ranking model The tenant identifier The tenant embedding table is mapped to a 32-dimensional vector, and the user-side permission token is... The input is mapped to a 32-dimensional vector through a fully connected layer with an input dimension of 2048 and an output dimension of 32. The two are then concatenated to obtain... The tenant embedding table and the fully connected layer parameters are fixed with the sorted model file and remain read-only during inference;

[0112] The ranking model outputs a raw ranking score for each candidate. The ranking model is a multilayer perceptron structure and its input vector is composed of user vectors. Asset Vector Interaction vectors Difference vector behavioral feature vector With filtering feature vectors The multilayer perceptron, constructed in a fixed order with an input dimension of 1120, comprises three fully connected layers with layer widths of 512, 128, and 1 respectively. The first two layers use GELU activation functions, and dropout is disabled during inference to ensure deterministic output. The last layer provides a linear output to obtain a real-valued sorted score. And used for subsequent weight reduction and threshold removal;

[0113] Based on filtering conditions For each candidate, an accessibility review is performed, and a penalty is imposed on the ranking score if the review fails. The review process includes, in sequence, a tenant consistency check. Permission token overriding judgment Consistency check with permission version number When it appears This triggers a refresh by retrieving the candidate's access control attributes from the access control management system. And complete the review with the refreshed results;

[0114] The following scoring restrictions apply during review and penalty enforcement:

[0115]

[0116] in, Represents the candidate after applying accessibility constraints. The final sorting score, This indicates the accessibility indicator, and is set to 1 if and only if the candidate passes the tenant consistency, permission token coverage, and permission version number consistency checks; otherwise, it is set to 1. The ranking model represents the candidate The output is the original sort score. This represents the penalty score and takes a fixed value. To achieve an approximation of negative infinity while ensuring it is less than a preset recommended output threshold. , Indicates that the candidate is in the candidate asset set The serial number in;

[0117] All candidates are scored according to their final ranking. Sort by high to low and remove those that meet the criteria during output. The candidate data assets are then truncated from the remaining candidates. Each data asset identifier forms a recommended list of data assets that meet tenant isolation and access control requirements.

[0118] In this specific embodiment, the user encoder, asset encoder, and sorting model are obtained offline and written to the storage medium according to the described process;

[0119] First, historical user behavior data is retrieved from the log system, and the user identifier (uid) and behavior timestamp are extracted for each behavior record. Behavior type and associated data asset identifier Obtain timestamps from the access control system. Corresponding historical access control information and extracting tenant identifiers Access control rules With permission version number Simultaneously, historical data asset metadata of the data asset is obtained from the metadata system and aligned with the historical access control information to the same point in time. Based on the above and Follow the steps The deterministic mapping table generates training permission tokens and obtains user-side training permission tokens respectively. Access tokens for asset-side training Based on this, the accessibility relationship between users and data assets is established. The determination adopts the rule of "tenant consistency and permission coverage" and is solidified into a binary accessibility mark with a rejection priority policy.

[0120] Then, the time window is determined according to the user identifier (uid). Historical behavior records within a day are aggregated from most recent to oldest timestamps into a length of [length missing]. The user behavior sequence is processed and excessively long portions are truncated to ensure input determinism. Simultaneously, the training access control attributes corresponding to the user at the time of the behavior are written into the user-side input. The user-side input includes tenant identifier, access control rules, training permission token, and permission version number. The asset-side input includes data asset metadata and the corresponding training access control attributes for that asset. This constructs a dual-tower model training sample pair and generates sample labels. The sample labels assign positive labels to "data assets that the user has interacted with and that were accessible at the time," and negative labels to "data assets that the user has not interacted with or that were inaccessible at the time." Negative samples use other asset samples within the same batch as implicit negative samples, ensuring the batch size is maintained. ;

[0121] The dual-tower model is trained to obtain a user encoder and an asset encoder. The user encoder adopts a sequence Transformer structure with 12 layers, 768 hidden layers, 12 attention heads, and a maximum input length of 52. At the output, a linear projection layer maps the 768-dimensional vector to 256 dimensions and executes the algorithm. The user vector is obtained through normalization. The asset encoder adopts a text Transformer structure with 12 layers, 768 hidden layers, 12 attention heads, and a maximum input length of 256. At the output, a linear projection layer is used to map the 768-dimensional vector to 256 dimensions before execution. The asset vectors are obtained through normalization, and the AdamW optimizer is used for training with a learning rate of [missing information]. Weight decay is The training run consists of 3 epochs, and an objective function based on intra-batch comparison is used.

[0122]

[0123] in, This represents the training loss of the dual-tower model. This indicates the batch size, and is set to 256. Indicates the sample sequence number within the batch and , Indicates the candidate asset sequence number within the batch and Indicates the first The user vector is the output of a user-side input, normalized by the user encoder. Indicates the relationship with the first The asset vector is the output of the asset encoder and normalized from the asset inputs that form positive sample pairs from the user-side inputs. Indicates the first in the batch Each asset-side input corresponds to an asset vector. Indicates transpose, thus The dot product of vectors, under the condition that the vectors are normalized, is equivalent to the cosine similarity. Represents the temperature coefficient and takes Represents an exponential function;

[0124] The trained user encoder and asset encoder are solidified and used to generate training samples for the ranking model. For each user behavior record in the training dataset, a training user vector and a training asset vector are calculated, and filtering conditions are generated. These filtering conditions include a tenant identifier, a user-side training permission token, a permission version number, and access constraints parsed from access control rules. Simultaneously, a candidate asset set consistent with the online dataset is constructed. Specifically, the training user vector and the filtering conditions are input into the following steps. The approximate nearest neighbor index is used to perform filtering retrieval to output... Candidate and Furthermore, assets marked as inaccessible according to their accessibility are added to the candidate asset set to form a training candidate set in which "accessible and inaccessible candidates coexist". For each candidate, its tenant identifier, asset-side training permission token and permission version number are written to ensure that the fields required for review during the sorting stage are complete.

[0125] The ranking model is trained using a multilayer perceptron structure, and its input consists of user vectors, asset vectors, and interaction vectors. Difference vector The behavioral feature vector and the filtering feature vector are concatenated in a fixed order and combined with the steps. The network layer widths are as follows: Furthermore, the first two layers use GELU as the activation function, and dropout is enabled during training with a dropout probability of 0.1. Training uses the AdamW optimizer with a learning rate of [missing value]. Weight decay is The training rounds are 2, and during training, a score constraint is imposed on inaccessible candidates based on accessibility relationships. Specifically, the training label of inaccessible candidates is fixed to 0, and their model output scores are cropped to no greater than a threshold before loss calculation. The value enables the model to learn the ranking preference of "not recommended under permission constraints", thereby obtaining a user encoder, asset encoder and ranking model that are consistent with online filtering recall and have explicit penalty capabilities for permission constraints, and writing their parameter files to the storage medium.

[0126] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

[0127] This invention treats access control and multi-tenant isolation as intrinsic constraints in the recommendation process, rather than as post-validation independent of the recommendation process: On the asset side, asset vectors encoded from data asset metadata are associated with tenant identifiers, access control rules, and permission tokens generated from their mappings, and written into index entries; on the user side, user vectors encoded from user behavior are combined with filtering conditions generated from user permission information to form a retrieval request; in the recall phase, near nearest neighbor retrieval skips or prunes entries that do not meet the constraints during graph traversal or inverted index scanning based on the filtering conditions, preventing inaccessible assets from entering the candidate set from the source; in the ranking phase, the ranking model, combined with the filtering conditions, applies penalties or threshold removal to inaccessible candidates and can trigger secondary validation to handle permission changes as a fallback, thereby achieving "visible and usable" recommendation results, reducing the exposure of unauthorized access and the proportion of invalid recommendations, and improving the usability and security of recommendations.

[0128] Compared to conventional structures that only filter after retrieval or only authenticate before display, this invention improves the algorithm structure to address the aforementioned technical problems by focusing on permission constraints: First, permission information is tokenized into fixed-length identifiers, enabling permission matching to participate in filtering on the index side with low overhead, adapting to a filterable near nearest neighbor index structure; Second, a filtering index structure associated with the vector index is introduced into the index structure, supporting pruning during the retrieval process. Simultaneously, filterable markers can be maintained in a hierarchical small-world graph scenario, prioritizing the expansion of nodes that meet the filtering conditions, reducing invalid access and lowering the unauthorized access exposure window; Third, accessibility relation constraints are introduced during the training and ranking stages, constructing a candidate set containing accessible and inaccessible assets and imposing restricted scores on inaccessible samples, allowing the model to learn ranking preferences under permission constraints, further improving the stability and compliance of online output.

Claims

1. This invention belongs to the field of data asset management and intelligent recommendation technology. It discloses a data asset intelligent recommendation method that combines user behavior. To solve the problem of visible but unusable or unauthorized exposure caused by the lack of integration of access control and multi-tenant isolation into the recommendation chain, this invention constructs an approximate nearest neighbor index that supports filtering by associating data asset vectors with tenant identifiers, access control rules, and permission tokens. Based on user behavior and access control information, user vectors and filtering conditions are generated, and filtering and constraint learning are performed in the retrieval and sorting stages. This achieves the technical effect of outputting only the recommendation list that meets the requirements of tenant isolation and access control, improving the usability of recommendations, and reducing the risk of unauthorized exposure.

2. The intelligent recommendation method for data assets combining user behavior according to claim 1, characterized in that, S1 includes: Obtaining the user encoder, asset encoder, and sorting model includes reading the user encoder, asset encoder, and sorting model from the storage medium for subsequent use; Obtaining data asset metadata and access control attributes corresponding to data assets includes determining data asset metadata and access control attributes corresponding to each data asset, wherein the access control attributes include tenant identifier, access control rules, and permission tokens generated based on at least one of the tenant identifier and access control rules. The permission token is a fixed-length identifier used to represent the scope of access, and the fixed-length identifier includes any one or more of the following: bitmap token, Bloom filter token, hash token, or role-resource pair encoded token; The asset encoder generates asset vectors for each data asset by using the data asset metadata as input, and associates each asset vector with its corresponding access control attribute to obtain an asset vector set. This includes inputting the data asset metadata corresponding to each data asset into the asset encoder to output the asset vector of the data asset, and establishing a correspondence between the asset vector and the access control attribute of the data asset, thereby aggregating the asset vectors to obtain the asset vector set.

3. The intelligent recommendation method for data assets combining user behavior according to claim 1, characterized in that, S2 includes: Constructing an approximate nearest neighbor index using the asset vector set as input includes: generating an index entry for each asset vector in the asset vector set, wherein the index entry includes a data asset identifier, the asset vector, and an access control attribute associated with the asset vector; The access control attributes include at least the tenant identifier, access control rules, and permission tokens; Based on all the index entries, a vector index structure for approximate nearest neighbor retrieval is constructed, wherein the vector index structure is a hierarchical small-world graph (HNSW) index structure or an inverted file (IVF) index structure. Simultaneously, a filtering index structure is constructed based on the access control attributes. The filtering index structure is at least used to determine the range of searchable index entries based on the tenant identifier, and to filter index entries that meet the access control requirements from the range of searchable index entries based on at least one of the access control rules and permission tokens. The filtering index structure includes a partitioned structure built by tenant identifier, and at least one of an inverted index, bitmap, or bucketed structure built by permission token; The vector index structure is associated with and stored with the filter index structure to obtain the approximate nearest neighbor index. This ensures that during subsequent filtering retrieval, only index entries that meet the filtering conditions are used for approximate nearest neighbor retrieval and retrieval results are output during the retrieval process of the vector index structure. The filtering retrieval includes skipping or pruning index entries that do not meet the filtering conditions during the graph traversal process of HNSW or the inverted linked list scan process of IVF, based on the filtering conditions.

4. The intelligent recommendation method for data assets combining user behavior according to claim 1, characterized in that, S3 includes: Acquiring the target user's current user behavior data and the target user's current access control information includes collecting the target user's search, click, preview, download, call, or query behavior records for data assets within a preset time window as the current user behavior data, and obtaining the target user's tenant identifier and access control rules from the permission management system as the current access control information; Generating a user vector using the user encoder by taking the current user behavior data and the current access control information as input includes organizing the current user behavior data into a user-side input in chronological order and inputting it together with the current access control information into the user encoder to output the user vector. Generating filtering conditions based on the current access control information includes extracting the tenant identifier from the current access control information and determining access constraints based on the access control rules to form the filtering conditions, and generating a user-side permission token based on at least one of the tenant identifier and the access control rules, and adding the user-side permission token to the filtering conditions. The user vector is combined with the filtering conditions to obtain the search request.

5. The intelligent recommendation method for data assets combining user behavior according to claim 1, characterized in that, S4 includes: Using the search request as input to call the approximate nearest neighbor index to perform a filtered search includes: inputting the filtering conditions in the search request into the approximate nearest neighbor index, so that the approximate nearest neighbor index determines a set of target index entries that satisfy the tenant identifier and the accessibility constraints based on the filtering conditions in the index entries; Determining the target set of index entries includes filtering index entries based on the matching relationship between the user-side permission tokens in the filtering conditions and the permission tokens in the index entries; Within the scope defined by the target index entry set, an approximate nearest neighbor search is performed in the approximate nearest neighbor index using the user vector in the search request as the query vector, to obtain search hit entries that meet the preset conditions in terms of similarity with the user vector; Data asset identifiers are extracted from the retrieved entries to form the candidate asset set.

6. The intelligent recommendation method for data assets combining user behavior according to claim 1, characterized in that, S5 includes: The process of ranking the candidate asset set, the current user behavior data, and the filtering conditions using the ranking model as input includes: obtaining the asset vector corresponding to each data asset from the candidate asset set, and inputting the current user behavior data, the asset vector, and the filtering conditions into the ranking model to output the ranking score corresponding to each data asset. Based on the filtering conditions, it is determined whether each data asset meets the accessibility constraints, and a ranking score limit is imposed on the data assets that do not meet the accessibility constraints. The ranking score limit is to set the corresponding ranking score to a penalty score that is less than a preset recommended output threshold, or to a negative infinity approximation value. The candidate asset set is sorted based on the ranking scores of each data asset, and when the recommendation result is output, data assets with ranking scores lower than the preset recommendation output threshold are removed from the sorted candidate asset set to obtain the recommendation result.

7. The intelligent recommendation method for data assets combining user behavior according to claim 1, characterized in that, The user encoder, the asset encoder, and the ranking model are obtained through the following steps: T1. Obtain historical user behavior data, historical access control information, and data asset metadata. Generate training access control attributes based on the historical access control information, and establish the accessibility relationship between users and data assets based on the training access control attributes to obtain a training dataset. The training access control attributes include tenant identifiers, access control rules, and training permission tokens generated based on the mapping of the tenant identifiers and / or access control rules. T2. Using the training dataset as input, aggregate the historical user behavior data according to the user identifier to generate a user behavior sequence, and construct sample pairs based on the user behavior sequence to obtain a dual-tower model training sample set. Each sample pair includes user-side input, asset-side input, and sample label. The user-side input includes the user behavior sequence and the training access control attribute corresponding to the user. The asset-side input includes data asset metadata corresponding to the data asset and the training access control attribute corresponding to the data asset. T3. Using the training sample set of the dual-tower model as input, train the dual-tower model to obtain the user encoder and the asset encoder. Encode the user-side input through the user encoder to obtain the user vector, and encode the asset-side input through the asset encoder to obtain the asset vector. Based on the sample labels, increase the similarity between the user vector and the asset vector of positive sample pairs and decrease the similarity between the user vector and the asset vector of negative sample pairs to obtain the user encoder and the asset encoder. T4. Using the user encoder, the asset encoder, and the training dataset as input, generate a ranking model training sample set for the user behavior records in the training dataset. For each user behavior record, use the user encoder to generate a training user vector and the asset encoder to generate a training asset vector. Construct a candidate asset set for the user behavior record based on the accessibility relationship, including data assets that satisfy the accessibility relationship and data assets that do not satisfy the accessibility relationship. T5. The ranking model is trained using the training sample set as input to obtain the ranking model; wherein the ranking model takes the training user vector, the training asset vector, and the training access control attribute as model input, outputs the ranking score of the data assets, and imposes ranking score restrictions on data assets that do not meet the accessibility relationship based on the accessibility relationship during the training process, so that the ranking model outputs a restricted ranking score for data assets that do not meet the accessibility relationship.

8. The intelligent recommendation method for data assets combining user behavior according to claim 3, characterized in that, A permission version number is set for the access control attribute, and when a change in the access control attribute of a data asset is detected in the permission management system, an incremental update is performed on the access control attribute and filtering index structure of the corresponding index entry based on the permission version number. When performing a filtering retrieval, if the permission version number of the filtering condition is inconsistent with the permission version number of the index entry, the index entry is marked as an entry to be verified and the access control attributes are retrieved again or a second verification is performed before deciding whether to output it.

9. The intelligent recommendation method for data assets combining user behavior according to claim 3, characterized in that, In the case where the vector index structure is a hierarchical small-world graph (HNSW) index structure, filterable tags corresponding to access control attributes are maintained for the HNSW entry point and each layer node. Before graph traversal begins, a set of entry points that satisfy the filtering conditions is selected based on the filtering conditions. Furthermore, during the traversal, when there are multiple candidate adjacent nodes available for expansion, adjacent nodes with filterable tags that satisfy the filtering conditions are expanded first, so as to reduce the number of times nodes that do not meet the filtering conditions are visited.

10. The intelligent recommendation method for data assets combining user behavior according to claim 7, characterized in that, The construction of the candidate asset set includes: inputting the training user vector and the filtering conditions generated by the training access control attributes into the approximate nearest neighbor index to perform filtering retrieval, thereby generating a candidate asset set consistent with the online stage, and using the candidate asset set as the training sample set for the ranking model.