A project asset data aggregation method and device, a terminal device, and a medium

CN122364215BActive Publication Date: 2026-09-18NAT UNIV OF DEFENSE TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610825703.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-09
Publication Date
2026-09-18
Estimated Expiration
2046-06-09

AI Technical Summary

Technical Problem

例如,未根据资产的核心属性设定优先级权重,也未考虑多属性间的非线性耦合关系,导致相似度计算结果与资产的实际业务含义产生偏差

Benefits of technology

通过根据资产属性对资产价值的贡献度为各特征分配动态权重,使后续处理能够优先考虑核心价值属性,提升了聚合结果与资产实际业务价值的契合度,克服了传统方法对所有属性等权重处理导致的偏差。通过构建采用脱敏标识作为节点的资产图,并基于加权相似度确定边及边权重,将资产间的复杂关联关系转化为可计算的图结构,为精准聚合奠定了基础。在此基础上,采用边权重阈值筛选与价值综合得分评价相结合的方式进行粗化聚合,实现了高效降维的同时确保高价值信息不丢失;进而构建结合动态权重的图趋势滤波目标函数进行交替优化求解,实现了簇内高度相似、簇间差异显著的精细化分组;最后通过跨层信息融合,输出完整精确的聚合结果,能够有效提升数据质量,为后续资产分析、评估与决策提供高质量的数据支撑。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122364215B_ABST
    Figure CN122364215B_ABST
Patent Text Reader

Abstract

The application provides a project asset data aggregation method and device, terminal equipment and medium, comprising feature extraction on enterprise project asset data to obtain asset feature data; hierarchical security desensitization is performed on the asset feature data, and dynamic weights are assigned to each asset feature; a first asset graph is constructed; high similarity node pairs are screened based on a preset edge weight threshold, nodes with high value comprehensive scores are retained as core nodes, and redundant node information is merged to obtain a second asset graph; a graph trend filtering target function combined with dynamic weights is constructed, and the second asset graph is alternately optimized and solved to obtain a fine-grained clustering cluster set representing asset grouping relationships; core features of each clustering cluster are extracted for cross-layer information fusion, and finally, a security-verified final asset aggregation result data is output. The application can improve the aggregation accuracy of multi-attribute project asset data, effectively improve the data quality, and provide high-quality data support for subsequent asset analysis, evaluation and decision-making.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of data processing technology, specifically relating to a method, apparatus, terminal equipment, and medium for aggregating project asset data. Background Technology

[0002] With the deepening of enterprise digital transformation, project asset management has moved from fragmented and decentralized storage to a digital management stage. Currently, project assets are no longer limited to a single type but broadly cover multiple fields such as finance, physical assets, and engineering, exhibiting characteristics such as massive data volume, complex attribute dimensions, strong data correlation, and high data sensitivity. To efficiently conduct analysis and control of the entire asset lifecycle within massive amounts of data, data aggregation—including eliminating redundancy, merging similarities, and integrating relationships—has become an indispensable preliminary step.

[0003] Currently, existing technologies for data aggregation mainly fall into two categories. The first category is traditional clustering algorithms, such as K-Means and hierarchical clustering. These algorithms typically group data based on geometric distance or similarity, often considering only a single attribute or a simple linear combination of a few attributes. When processing multi-attribute project asset data, they struggle to capture the inherent relationships between attributes and cannot distinguish the priority differences between core and secondary attributes. Essentially, they are static algorithms and cannot adapt to the dynamic evolution of asset data. Therefore, when applied to massive, high-dimensional asset data, these algorithms not only lack aggregation accuracy but also experience a sharp increase in computational complexity, making it difficult to meet actual business needs. The second category is conventional graph aggregation algorithms, represented by the Louvain algorithm. While these algorithms can construct graph models using the relationships between data entities, they typically use general topological metrics when constructing edge weights or calculating node similarity, failing to deeply integrate the inherent characteristics of project asset data for optimization. For example, they do not set priority weights based on the core attributes of the asset, nor do they consider the nonlinear coupling relationships between multiple attributes, leading to deviations between the similarity calculation results and the actual business meaning of the asset.

[0004] In summary, existing technologies suffer from problems such as insufficient aggregation accuracy, low data processing efficiency, and poor adaptability to asset business characteristics when processing multi-attribute project asset data. There is an urgent need for a new technical solution that can systematically solve the above-mentioned defects. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a method, apparatus, terminal equipment and medium for project asset data aggregation, so as to improve the aggregation accuracy of multi-attribute project asset data.

[0006] In a first aspect, the present invention provides a method for aggregating project asset data, the method comprising the following steps: Feature extraction is performed on the acquired enterprise project asset data to obtain asset feature data; enterprise project asset data includes structured data, semi-structured data, and unstructured data; structured data consists of one or more of the following: book value, physical dimensions, and specifications; semi-structured data consists of business texts in daily circulation; unstructured data assets consist of one or more of the following: images, photographs, CAD engineering drawings, and technical specifications. The asset feature data is hierarchically and securely anonymized, and dynamic weights are assigned to each anonymized asset feature based on the contribution of asset attributes to asset value. A first asset graph is constructed based on asset characteristic data and dynamic weights to represent the relationships between assets. The nodes of the first asset graph use unique, de-identified asset identifiers, corresponding to the actual project asset entities of the enterprise. The edges of the first asset graph represent the aggregateable business relationships or homogeneous relationships between asset entities. The edge weights of the first asset graph are used to quantify the degree of matching between two assets in terms of their core value contribution attributes. For the first asset graph, highly similar node pairs are selected based on a preset edge weight threshold. Nodes with high comprehensive value scores are retained as core nodes, and redundant node information is merged to obtain a second asset graph with a reduced node size. A graph trend filtering objective function incorporating dynamic weights is constructed, and the graph signal of the second asset graph is alternately optimized to obtain a fine-grained cluster set representing the asset grouping relationship. The core features of each cluster in the fine-grained cluster set are extracted and cross-layer information is fused. After merging clusters that meet the preset similarity conditions, the final asset aggregation result data with security verification is output.

[0007] Optionally, dynamic weights can be assigned to each feature based on its contribution to the asset value, including: The attribute value association weighting method is adopted to calculate the asset value impact factor based on the historical average value contribution and real-time value influence of asset attributes. The weight value of each asset characteristic is then dynamically calculated based on the asset value impact factor. The expression for the asset value impact factor is as follows: ;in, This represents the adjustment coefficient. , This represents the average historical value contribution of an asset attribute. This indicates the real-time value impact of asset attributes.

[0008] Optionally, the edges and edge weights of the first asset graph are determined by calculating the weighted similarity between any two nodes, including: Through calculation formula

[0009] Obtain weighted similarity ; Represents a node and nodes Weighted similarity between them and This represents the node number index in the first asset diagram. and Representing nodes respectively and nodes The corresponding standardized asset characteristic values, Represents a node The corresponding weights of standardized asset eigenvalues Indicates the minimum value; When the weighted similarity is greater than or equal to a preset similarity threshold, an edge is constructed between nodes, and the weighted similarity is used as the edge weight of the corresponding edge.

[0010] Optionally, nodes with high overall value scores are retained as core nodes, and redundant node information is merged to obtain a second asset graph with reduced node size, including: Calculate the comprehensive value score for each node in a pair of highly similar nodes; the comprehensive value score is determined based on the weighted sum of the feature values ​​and corresponding dynamic weights of each node; the expression for the comprehensive value score is: ; Nodes with high overall value scores are retained as core nodes, while nodes with low overall value scores are removed as redundant nodes. The feature information of the redundant nodes is then merged into the core nodes using a weighted average method.

[0011] Optionally, a graph trend filtering objective function incorporating asset feature weights is constructed, and the mathematical expression for minimizing the objective function is as follows:

[0012] in, This represents the index of the asset feature view. This indicates the total number of asset feature views. This represents the total order of the graph difference operator. This represents the total number of clusters. This indicates adaptive weighting, which is positively correlated with asset value influencing factors. Indicates the first Each asset view Graph difference operator, Indicates the first Cluster indicator signal for each cluster, The L1 norm represents the smoothness of the cluster indicator signal on the asset graph. Represents the locus of a matrix. Represents the clustering indicator matrix. Indicates the first Weights of each asset view Indicates the first Enhanced asset graph adjacency matrix for each asset view Indicates the order index of the graph difference operator. This represents the cluster index.

[0013] Optionally, alternating optimization strategies include deriving a closed-form solution through the Cauchy-Schwarz inequality and sequentially optimizing the local preference weights, view weights, and clustering indicator matrix until the objective function converges.

[0014] Optionally, enterprise project asset data includes structured data, semi-structured data, and unstructured data; Feature extraction is performed on enterprise project asset data to obtain asset feature data, including: For structured, semi-structured, and unstructured data in enterprise project asset data, direct extraction, natural language processing, and convolutional neural network methods are used respectively to transform them into feature vectors in a unified format.

[0015] Secondly, the present invention provides a project asset data aggregation device, comprising: The feature extraction module is used to extract features from the acquired enterprise project asset data to obtain asset feature data. Enterprise project asset data includes structured data, semi-structured data, and unstructured data. Structured data consists of one or more of the following: book value, physical dimensions, and specifications. Semi-structured data consists of business texts in daily circulation. Unstructured data assets consist of one or more of the following: images, photographs, CAD engineering drawings, and technical specifications. The preprocessing module is used to perform hierarchical security desensitization on asset feature data and assign dynamic weights to each desensitized asset feature based on the contribution of asset attributes to asset value. The graph construction module is used to construct a first asset graph to represent the relationships between assets based on asset characteristic data and dynamic weights. The nodes of the first asset graph use a unique, de-identified asset identifier, corresponding to the actual project asset entities of the enterprise. The edges of the first asset graph represent the aggregateable business relationships or homogeneous relationships between asset entities. The edge weights of the first asset graph are used to quantify the degree of matching between two assets in terms of core value contribution attributes. The coarsening and aggregation module is used to filter highly similar node pairs based on a preset edge weight threshold for the first asset graph, retain nodes with high comprehensive value scores as core nodes and merge redundant node information to obtain a second asset graph with a reduced node size. The refinement aggregation module is used to construct a graph trend filtering objective function that incorporates dynamic weights, and to alternately optimize and solve the graph signal of the second asset graph to obtain a set of fine-grained clusters that represent the grouping relationship of assets. The aggregation module is used to extract the core features of each cluster in the fine-grained cluster set, perform cross-layer information fusion, and output the final asset aggregation result data after merging clusters that meet the preset similarity conditions, which has been verified by security.

[0016] Thirdly, the present invention provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-described method.

[0017] Fourthly, the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method.

[0018] The present invention has at least the following beneficial effects: By assigning dynamic weights to each feature based on its contribution to asset value, subsequent processing prioritizes core value attributes, improving the alignment between the aggregation results and the actual business value of the assets, and overcoming the bias caused by the traditional method of treating all attributes with equal weight. By constructing an asset graph using de-identified markers as nodes and determining edges and edge weights based on weighted similarity, the complex relationships between assets are transformed into a computable graph structure, laying the foundation for accurate aggregation. Furthermore, a coarse aggregation is performed using a combination of edge weight threshold filtering and comprehensive value score evaluation, achieving efficient dimensionality reduction while ensuring no loss of high-value information. Then, a graph trend filtering objective function incorporating dynamic weights is constructed and alternately optimized to achieve refined grouping with high intra-cluster similarity and significant inter-cluster differences. Finally, cross-layer information fusion outputs complete and accurate aggregation results, effectively improving data quality and providing high-quality data support for subsequent asset analysis, evaluation, and decision-making. Attached Figure Description

[0019] The accompanying drawings are provided to further understand the technical solutions of the present invention and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the technical solutions of the present invention, and do not constitute a limitation on the technical solutions of the present invention.

[0020] Figure 1 This is a flowchart of a project asset data aggregation method in one embodiment of this application; Figure 2 This is a structural diagram of a project asset data aggregation device according to one embodiment of this application; Figure 3 This is a structural diagram of a terminal device in one embodiment of this application. Detailed Implementation

[0021] The technical solution of the present invention will now be described in detail and completely with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0022] In the description of this invention, it should be noted that the terms "upper", "lower", "inner", "outer", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this invention.

[0023] Existing technologies generally suffer from the following drawbacks when processing multi-attribute project asset data: First, traditional clustering algorithms such as K-Means struggle to capture non-linear relationships between multiple attributes and assume equal weights for all attributes, leading to aggregation results that deviate from the actual business distribution of assets. Second, conventional graph aggregation algorithms such as Louvain only use general topological structure metrics when constructing the graph, without prioritizing based on core asset attributes, resulting in similarity calculation results that differ from business meanings. Third, existing technologies do not incorporate data security into the processing flow, posing a risk of sensitive information leakage. To address these shortcomings, this invention provides a project asset data aggregation method, which will be described in detail below through several specific embodiments.

[0024] Example 1 This embodiment provides a method for aggregating project asset data, which runs on a server terminal device with computing and storage capabilities.

[0025] In this embodiment of the invention, enterprise project asset data can be obtained from the enterprise's internal asset management system, financial system, and project archive. Enterprise project asset data includes structured data, semi-structured data, and unstructured data. Structured data consists of physical parameters or business indicators explicitly recorded in the database, such as the book value, physical dimensions, specifications, duration, and ownership code of an asset; semi-structured data consists of business texts in daily circulation, such as equipment operation and maintenance logs, fragments of procurement contracts, and transaction records; unstructured data consists of the visual or auditory representation of assets or complex document formats, such as on-site photographs of assets, CAD engineering drawings, and lengthy technical documentation.

[0026] For example, in one feasible implementation, the enterprise project asset data obtained from a company's database system is as follows: Structured data: The original asset value of the equipment with asset code AST-001 is RMB 1.2 million, the service life is 10 years, and the risk level is medium.

[0027] Semi-structured data: Asset operation and maintenance log text. The equipment is running smoothly and there are no recent fault records.

[0028] Unstructured data: images of the equipment's exterior and PDF documents of the technical specifications.

[0029] like Figure 1 As shown, the project asset data aggregation method provided by the present invention specifically includes steps 11 to 16.

[0030] Step 11: Extract features from the acquired enterprise project asset data to obtain asset feature data.

[0031] Due to the massive scale of enterprise project asset data, which contains numerous errors, empty data, and redundant data, a feasible implementation method involves preprocessing the data before feature extraction to improve data quality. This preprocessing specifically includes using an improved 3D model. The criteria are used to remove outliers and deduplicate data.

[0032] The following is based on the improved 3 The process of outlier removal based on the criteria is explained.

[0033] In actual project assets, asset values ​​typically exhibit a "right-skewed distribution" (i.e., a long-tail effect, with a few core assets having extremely high value and most assets having ordinary value), while traditional 3 The criterion (mean ± 3 standard deviations) heavily relies on the premise that the data follows a normal distribution. If the traditional 3 standard deviations are rigidly applied... The existing rules are prone to mistakenly deleting truly valuable core assets. To address this, this invention provides an improved 3D method. The criterion utilizes adaptive skewness adjustment to replace the traditional 3 The static threshold in the criterion is "mean ± 3 standard deviations".

[0034] First, for the extracted structured asset values ​​(taking the "asset value" attribute as an example), we first calculate the three basic statistical parameters of this dimension of data: mean. (The overall average level of asset value), standard deviation (Characterizing the dispersion of asset value) and skewness coefficient (This characterizes the asymmetry of asset value distribution; the formula is the ratio of the third central moment to the cube of the standard deviation). If the skewness coefficient... This indicates that asset values ​​exhibit a "right-skewed long tail" (i.e., the existence of a few extremely high-value assets).

[0035] Subsequently, breaking with tradition 3 The criteria use a fixed multiple of 3 for both the upper and lower limits, and introduce a dynamic adjustment coefficient to construct an asymmetric threshold multiplier: upper threshold multiplier. and lower limit threshold multiplier .in, , This is a penalty adjustment parameter, with a value of 0.5 or 1. It is used when the data is severely skewed to the right. When it is large, It will automatically enlarge to 4 5 Even higher. However, for asset values, extreme negative deviations are rare, therefore the lower limit threshold multiplier... The default value can be kept, such as .

[0036] Then, based on the calculated dynamic threshold multiplier, the upper and lower boundaries for outlier detection are redefined; among them, the dynamic upper boundary... Dynamic lower bound .

[0037] Finally, iterate through all project asset data, and only when the value attribute value of an asset is greater than the dynamic upper bound... or less than the dynamic lower bound Only when the data is in error is it identified as a genuine outlier (such as erroneous data caused by sensor malfunction or manual input of extra zeros) and removed.

[0038] Through this improvement 3 The criteria for outlier removal allow this invention to automatically identify whether the current batch of asset data contains high-value assets. By automatically relaxing the upper threshold, it effectively solves the problem of misjudgment of outliers caused by extreme asset values. While cleaning up truly dirty data, it retains intact the core high-quality asset data that is most valuable for enterprise analysis.

[0039] In one feasible implementation, data deduplication involves comparing data based on the unique identifier of the asset (such as asset number, unified social credit code, etc.) to directly exclude completely identical redundant and duplicate data.

[0040] For the structured, semi-structured, and unstructured data in the preprocessed enterprise project asset data, this invention uses direct extraction, natural language processing, and convolutional neural network methods to transform them into feature vectors in a unified format.

[0041] Specifically, for structured data, its numerical or enumerated attribute features are directly extracted and concatenated according to dimensions to transform it into a numerical vector, such as an asset with an original value of 120 and a usage period of 10.

[0042] For semi-structured data, natural language processing (NLP) technology is used to segment and semantically analyze text logs or contract fragments, extract core business keyword features, and map them into text feature vectors, such as "stable" and "fault-free" as text features.

[0043] For unstructured data, convolutional neural networks (CNNs) are used to extract visual features (such as contours and wear) from asset images or to extract deep text features from complex documents, ultimately transforming them into feature vectors of standard dimensions.

[0044] In practice, to eliminate the influence of different feature dimensions on the scale and facilitate subsequent calculations, all extracted feature vectors will be standardized. Instead of a fixed [0,1] rigid mapping, the standardization interval will be dynamically adjusted based on the "attribute value range" of the current input asset data. This dynamic adjustment mechanism effectively avoids the excessive compression or weakening of vastly different asset attributes (such as the value difference between million-dollar assets and hundred-dollar assets) by conventional normalization operations, ensuring that the core attribute differences between different assets are preserved while maintaining a unified scale.

[0045] Step 12: Perform hierarchical security desensitization on the asset feature data, and assign dynamic weights to each desensitized asset feature based on the contribution of asset attributes to asset value.

[0046] To prevent the leakage of ownership information, core technical parameters, or sensitive financial data, this invention performs hierarchical security anonymization on asset characteristic data. The specific process is as follows: First, based on the pre-set asset compliance strategy, sensitive information in asset characteristic data is divided into three levels: high sensitivity, medium sensitivity, and low sensitivity.

[0047] Subsequently, for highly sensitive data (such as core financial data and absolute owner ID): hash encryption is used for desensitization. A one-way hash algorithm (such as SHA-256) is used to convert plaintext features into fixed-length ciphertext, severing the direct link between the data and the real entity. For moderately sensitive data (such as core technical specifications and contact information): masking is used for desensitization. The data format is preserved according to rules, and the remaining key fields are replaced with specific symbols (e.g., masking the long string parameter "X-9982-A" with "X-"). -A”). For low-sensitivity data (such as common equipment operating status parameters): perturbation desensitization is used. A small random noise (e.g., increasing or decreasing by a very small percentage) is superimposed on the true standardized feature value, causing it to deviate from the true absolute value, but since the noise mean is zero, it does not change the statistical distribution of the feature in the overall graph structure.

[0048] Finally, the mapping relationships and keys from the above de-identification process are centrally stored in a secure area. Throughout the entire graph construction and clustering process, nodes are traversed using de-identified identifiers; data can only be reversibly restored using the key when the final result is output and the recipient is an authorized node.

[0049] After anonymization, to ensure that the attributes that truly determine the core value of the asset play a decisive role in subsequent clustering, this invention breaks through the limitation of traditional algorithms that assume equal feature weights, and assigns dynamic weights to each feature. The specific steps are as follows: The attribute value association weighting method is adopted to calculate the asset value impact factor based on the historical average value contribution and real-time value influence of asset attributes. The weight value of each asset characteristic is then dynamically calculated based on the asset value impact factor. The expression for the asset value impact factor is as follows: .in, This represents the adjustment coefficient. In practical applications, Adjustments will be made dynamically based on asset type. For example, for tangible engineering assets (based on historical experience). It can be set to 0.7. For financial digital assets (highly volatile in real-time), It can be set to 0.3. This represents the average historical value contribution of an asset attribute. It is obtained by retrieving data from a database and calculating the proportion of that attribute in determining the final appraised value of the asset in past projects. This indicates the real-time value impact of an asset attribute, calculated by combining the current project environment to determine the immediate value fluctuation impact of that attribute.

[0050] After obtaining the influence factors of all features, they are normalized to calculate the first... The final weight value of each asset feature During this process, it is strictly guaranteed that the sum of the weights of all features equals 1.

[0051] In this way, asset characteristics that contribute more to asset value (such as core financial return) are assigned larger weights, while the weights of less important peripheral characteristics are compressed. This provides a business logic-based "ruler" for calculating the similarity between subsequent nodes, ensuring that the clustering process closely revolves around the core value of the asset.

[0052] Step 13: Construct a first asset graph to represent the relationships between assets based on asset characteristic data and dynamic weights.

[0053] In this embodiment of the invention, the nodes of the first asset graph use de-identified unique asset identifiers. Specifically, each standardized "asset feature data" (i.e., a high-dimensional feature vector) is mapped to an independent node in the first asset graph. To prevent the leakage of sensitive information throughout the graph construction and computation process, all plaintext sensitive fields (such as the owner's name and core financial details) are removed from the nodes themselves, and only the hashed or encrypted "de-identified unique asset ID" is used as the physical identity marker.

[0054] The edges and edge weights of the first asset graph are determined by calculating the weighted similarity between any two nodes. Traditional algorithms typically calculate the geometric distance (such as Euclidean distance) between two feature vectors, which leads to core features and edge features being treated equally. This implementation introduces the "dynamic weights" assigned in the previous stage, applying them to any two nodes in the graph. Perform weighted difference calculations along each feature dimension. Specifically, this is done using the calculation formula...

[0055] Obtain weighted similarity ; Represents a node and nodes Weighted similarity between them and This represents the node number index in the first asset diagram. and Representing nodes respectively and nodes The corresponding standardized asset characteristic values, Represents a node The corresponding weights of standardized asset eigenvalues This represents the minimum value.

[0056] Set a similarity threshold to determine whether assets are sufficiently related. The similarity threshold is dynamically adjusted based on the currently accessed asset type. For example, for highly standardized financial assets, the similarity threshold... The similarity threshold can be set higher accordingly. For engineering assets with many unstructured features, the similarity threshold can be adjusted accordingly. The similarity can be set lower accordingly. Iterate through all node pairs, and when the weighted similarity calculated for two nodes is greater than or equal to a preset similarity threshold, construct an edge between the nodes and use this weighted similarity as the edge weight. The larger the edge weight, the closer the two assets are in their core value dimension.

[0057] Finally, the set of desensitized nodes The set of edges that satisfy the threshold condition and the corresponding set of edge weights Together they constitute the first asset diagram .

[0058] Through this specific implementation, the first asset graph constructed by the present invention is no longer a blind geometric topology network, but a value-oriented weighted graph. It ensures that in subsequent clustering operations, even if two assets differ greatly in ordinary specification parameters (low weight), as long as they are highly similar in core business tags or core value parameters (high weight), their connectivity in the graph will be extremely strong, thus being preferentially assigned to the same business cluster.

[0059] Step 14: For the first asset graph, based on the preset edge weight threshold, filter high similarity node pairs, retain the nodes with high comprehensive value scores as core nodes, and merge redundant node information to obtain the second asset graph with reduced node size.

[0060] Specifically, step 14 includes steps 14.1 to 14.4.

[0061] Step 14.1: Filter highly similar node pairs based on preset edge weight thresholds.

[0062] Regarding the first asset map For the edges in the graph, set a relatively high coarsening weight threshold. This threshold represents the critical point at which two assets are "extremely similar" or "highly redundant" in terms of physical or business attributes (e.g., two servers purchased in the same batch, with identical specifications and very similar depreciation levels).

[0063] Traverse the edge set of the first asset graph Filter out all edges that satisfy the edge weight. These are the first batch of highly similar node pairs that need to be redundantly removed and merged.

[0064] Step 14.2: Calculate the comprehensive score of node value and establish core nodes.

[0065] For each pair of highly similar nodes selected Through calculation formula Obtain the comprehensive value score of each node in the node pair. , or .Compare and Nodes with higher overall value scores are identified as core nodes with stronger representativeness and greater value retention significance (e.g., among a batch of servers of the same model, the one with the highest value, optimal operating status, and lowest risk is retained as the prototype of the group. Its feature vector is prioritized for enhancement in subsequent processing, becoming the representative node of the aggregated business cluster, used as the benchmark for maintaining strategies, valuation models, and risk profiles). Nodes with lower overall value scores are identified as redundant nodes (representing asset entities that are highly similar to core nodes but have relatively minor value, such as slightly older, less utilized equipment of the same model in the same batch. Their information is obviously redundant, and retaining them separately would lead to data duplication, statistical bias, and waste of computing resources).

[0066] Step 14.3: Perform a dual merging of features and network topology.

[0067] In order to not only avoid losing information when removing redundant nodes, but also enhance the representativeness of core nodes, this invention merges the feature information of redundant nodes into core nodes through a weighted average method.

[0068] For example, redundant nodes (let's assume) The feature information is proportionally integrated into the core node (assuming it is). In this process, the feature vectors of the core nodes are updated. The update formula is: . and These represent the percentage of the combined value score of the two nodes in the total score, i.e. The higher the value of a node, the greater the proportion of its original features retained after fusion.

[0069] Subsequently, the redundant nodes originally connected in the first asset graph were... Reconnect all external neighbor nodes to the core node. The edge weights are then merged or updated accordingly, and redundant nodes are subsequently removed from the first asset graph. and its original connecting edges.

[0070] Step 14.4, iterative convergence, output the scaled-down second asset graph.

[0071] Repeat steps 14.1 to 14.3 of the above process of screening, comparing, and merging on the current first asset graph structure. As nodes are continuously merged, highly redundant nodes in the first asset graph are continuously absorbed until the edge weights of any two connected nodes in the graph no longer satisfy the condition. When the iteration stops, output the coarsened second asset map.

[0072] By executing step 14, the first asset graph (potentially with millions of nodes) is effectively compressed into the second asset graph (a set of hundreds of thousands or tens of thousands of core nodes). This not only greatly reduces the computational time complexity of subsequent complex graph filtering and clustering at the physical level, but also ensures in terms of business logic that only low-value redundant appearances are merged, while the core data of high-value assets are not only preserved without loss, but also become richer and more robust by absorbing redundant features.

[0073] Step 15: Construct a graph trend filtering objective function that incorporates dynamic weights, and perform alternating optimization on the graph signal of the second asset graph to obtain a set of fine-grained clusters representing the asset grouping relationship.

[0074] For the scaled-down second asset graph, this invention divides its multidimensional features into different asset feature views (such as financial value view, physical attribute view, etc.), and treats the clustering labels as graph signals. Combining the dynamic weights calculated in the previous stages, the following minimum graph trend filtering objective function is constructed:

[0075] in, This represents the index of the asset feature view. This indicates the total number of asset feature views. This represents the total order of the graph difference operator. This represents the total number of clusters. This indicates adaptive weighting, which is positively correlated with asset value influencing factors. Indicates the first Each asset view Graph difference operator, Indicates the first Cluster indicator signal for each cluster, The L1 norm represents the smoothness of the cluster indicator signal on the asset graph. Represents the locus of a matrix. Represents the clustering indicator matrix. Indicates the first Weights of each asset view Indicates the first Enhanced asset graph adjacency matrix for each asset view Indicates the order index of the graph difference operator. This represents the cluster index.

[0076] By minimizing the objective function, this invention forces the clustering results to simultaneously satisfy the dual stringent conditions of "extremely similar characteristics of asset nodes within the cluster (minimum numerator)" and "extremely close physical association of asset nodes within the cluster (maximum denominator)".

[0077] Since the objective function is highly complex and non-convex, this invention adopts an "alternating optimization strategy" to derive a closed-form solution through the Cauchy-Schwarz inequality and sequentially optimize the local preference weights, view weights, and clustering indicator matrix until the objective function converges.

[0078] Specifically, the highly complex and non-convex graph trend filtering objective function is decomposed into three easily solvable subproblems, thereby achieving rapid convergence of the algorithm under massive asset data. The process is as follows: Phase 1, Fixed View Weights With indicator matrix Optimize local preference weights The original objective function is then transformed into a function that depends only on local preferences. (in the objective function) Minimize the subproblem of ).

[0079] In this embodiment of the invention, view weight This is used to measure the overall influence of different "viewpoints (data source types)" in determining the final clustering of an asset. In reality, an asset has multiple dimensions of characteristics (e.g., "financial accounting view," "equipment operation log view," "physical contract view," etc.). View weight. The weight of the "financial view" determines which category should be prioritized when clustering this batch of assets. For example, if the algorithm finds that the essential differences among these assets are mainly reflected in their financial data, then the weight of the "financial view" should be higher. It will be automatically magnified, while the weight of the "Run Log View" will be reduced. It will become smaller.

[0080] Indicator Matrix It is a dimension The matrix, where This represents the total number of asset nodes that have entered the cluster. This represents the total number of target clusters. It indicates the number of each element in the matrix. It is a discrete binary state value, when At that time, it indicates the project assets. Belongs to cluster .when At that time, it indicates the project assets. Not belonging to cluster It should be noted that during the intermediate calculations of the algorithm's alternating optimization, for ease of differentiation, the matrix elements may be continuous probability values ​​between 0 and 1, representing the probability of assignment; however, when the algorithm converges and outputs the final result, it will be discretized and ultimately fixed as a hard assignment identifier of 0 or 1.

[0081] Local preference weights This is used to capture non-linear business preferences that vary depending on specific asset clusters, ensuring that high-value, unique attributes play a dominant pulling role in the local network. For example, the "precision instrument cluster" is extremely sensitive to the feature of "minor vibration parameters," in which case the local preference weights... Extremely high. However, the "heavy-duty forging press cluster" is completely indifferent to "minor vibration parameters," at which point local preference weights... Extremely small.

[0082] For each cluster Independent optimization was performed, and a closed-form solution was derived using the Cauchy-Schwarz inequality. The closed-form solution (exact analytical solution) is obtained. This closed-form solution makes... The value of and the smoothness term of the asset chart signal ( It is inversely proportional to the asset value impact factor of the feature, and positively correlated with the asset value impact factor of the feature. This allows asset features with high value contribution to automatically obtain greater local aggregation pull. That is, if a feature (such as the core profit margin of an asset) performs very stably and has a high value contribution in the current cluster, it will be automatically given a large local weight through this closed solution, thereby more strongly "pulling" similar high-value assets together in subsequent calculations.

[0083] Phase 2, Fixed Local Weights With indicator matrix Optimize view weights The original objective function degenerates into a weight about the global view. (vector Subproblems of (elements).

[0084] Similarly, based on the Cauchy-Schwarz inequality, the denominator contains... By scaling and differentiation, the optimal value can be directly calculated. Closed-form solution. The allocation principle for this closed-form solution is: if a certain asset view (such as a "structured financial view"), under the current partitioning, maximizes the physical connectivity and internal density of asset nodes within the cluster, then the weight of that view is... This will be significantly amplified. This allows the algorithm to adapt to the differences in characteristics of different project assets.

[0085] Phase 3, Fixing Local Weights With view weight Optimize the indicator matrix At this point, all the weight parameters in the original objective function are determined, and the problem is formally transformed into a problem concerning the clustering indicator matrix. The optimization problem.

[0086] Combining the comprehensive value scores of asset nodes, a spectral clustering optimization method is employed. Using the optimized weights from the first two steps, a comprehensive graph Laplacian matrix integrating multiple scales and views is reconstructed. Eigenvalue decomposition is then performed on this comprehensive graph Laplacian matrix to obtain the leading edge... The eigenvectors corresponding to the smallest eigenvalues ​​are then discretized and updated to update the clustering indicator matrix. By rigidly adjusting the grouping of assets, we can prevent high-value but not readily apparent implicit assets from being mistakenly mixed into low-value clusters.

[0087] In practical implementation, the updated clustering indicator matrix will be used. Re-enter Phase 1 and start a new round. , , The algorithm employs alternating updates. At the end of each iteration, the total value of the current objective function is calculated. Alternating optimization terminates when the difference between the objective function values ​​of two adjacent iterations is less than a preset minimum (i.e., the algorithm has reached convergence and the cluster partitioning no longer changes substantially), or when the maximum number of iterations is reached. The final output is a clustering indicator matrix containing the optimal grouping results. This process generates a refined set of asset clusters, resulting in a fine-grained cluster set. Within this set, each cluster precisely corresponds to a group of project asset nodes that are highly consistent in terms of core business value, multi-dimensional physical attributes, and graph network connectivity.

[0088] Step 16: Extract the core features of each cluster in the fine-grained cluster set, perform cross-layer information fusion, and after merging clusters that meet the preset similarity conditions, output the final asset aggregation result data that has been verified by security.

[0089] Specifically, for each cluster in the fine-grained clustering set, the standardized feature vectors of all asset nodes within the cluster are extracted. The weighted average of the feature vectors of all nodes within the cluster is calculated using the proportion of the node's overall value score as the weight. . This represents the set of nodes contained in the current cluster. Represents a node The percentage of the overall value score in the total score within the cluster. This is the core feature vector of the cluster.

[0090] In step 14, to reduce computational complexity, this invention removes and merges a large number of highly similar redundant nodes. However, these redundant nodes may contain some long-tail information that does not affect the core value but is useful for asset profiling (such as specific models of non-core spare parts, frequency of minor maintenance logs, etc.). At this time, based on the asset association mapping, this invention re-extracts and appends or merges the key non-sensitive features of the corresponding redundant nodes that were removed into the core feature vector of the current cluster. This cross-layer mechanism ensures that the asset aggregation result has both the high efficiency brought by coarsening dimensionality reduction and does not cause the permanent loss of useful long-tail data at the lower level, achieving a balance between macro-aggregation and micro-information preservation.

[0091] Because graph trend filtering and other disclosure algorithms, when optimizing locally, may incorrectly segment assets that should belong to the same major category into multiple overly fine clusters due to minor differences in local asset features, this invention calculates the inter-cluster similarity score between the core feature vectors of any two clusters. When the inter-cluster similarity score is greater than or equal to a preset inter-cluster similarity threshold, it determines that the two clusters are highly homogeneous in macro-business logic (for example, identifying "domestic general-purpose server cluster" and "imported general-purpose server cluster" as similar as a whole). These two fine-grained clusters are then merged into a unified major asset cluster, and the core features of the merged cluster are recalculated.

[0092] Before generating the final result set, this invention traverses all merged clusters and their associated feature nodes, rigorously verifying whether they contain plaintext sensitive fields (such as the real owner's ID number, the absolute value of the real core financial transaction, etc.) that have not been hashed, masked, or perturbed. After confirming that the verification passes and there is no risk of sensitive information leakage, the final asset aggregation result dataset is encrypted using AES-256, and a high-quality aggregation result is output to the downstream asset analysis and evaluation business system through a secure channel configured with SSL / TLS protocol.

[0093] This invention achieves differentiated protection of sensitive information by extracting features and performing hierarchical security desensitization on enterprise project asset data. It preserves data usability while ensuring data security, addressing the problem of neglecting end-to-end data security in existing technologies. By assigning dynamic weights to each feature based on its contribution to asset value, subsequent processing prioritizes core value attributes, improving the alignment between the aggregation results and the actual business value of the assets, and overcoming the bias caused by the equal weighting of all attributes in traditional methods. By constructing an asset graph using desensitized identifiers as nodes and determining edges and edge weights based on weighted similarity, the complex relationships between assets are transformed into a computable graph structure, laying the foundation for accurate aggregation. Furthermore, a coarse aggregation is performed using a combination of edge weight threshold filtering and comprehensive value score evaluation, achieving efficient dimensionality reduction while ensuring no loss of high-value information. Then, a graph trend filtering objective function incorporating dynamic weights is constructed and alternately optimized to achieve refined grouping with high intra-cluster similarity and significant inter-cluster differences. Finally, through cross-layer information fusion and security verification, a complete and concise aggregation result is output. The above-mentioned multi-scale hierarchical strategies work together to significantly improve the aggregation accuracy and processing efficiency of massive multi-attribute asset data. At the same time, data security is internalized as an inherent part of the aggregation process, providing high-quality data support for subsequent asset analysis, evaluation and decision-making.

[0094] Example 2 In this embodiment of the invention, to further demonstrate the feasibility of the technical solution in a real industrial scenario, the end-to-end implementation process of the method of the present invention is described below using a "Smart Factory Project Asset Lifecycle Digital Management Platform" deployed by a large automobile manufacturing enterprise as an example. This enterprise owns more than 12,000 production equipment assets (including CNC machine tools, industrial robots, AGV logistics vehicles, production line testing equipment, etc.). Data sources include structured data from ERP / MES systems (physical parameters and business indicators such as original asset value, service life, risk level, and maintenance costs), semi-structured data such as daily operation and maintenance logs and procurement contract fragments, and unstructured data such as on-site high-definition images and equipment technical manuals. Existing technologies suffer from insufficient aggregation accuracy, low computational efficiency due to redundant data, risks of sensitive information leakage, and a disconnect between the data and the actual value contribution of the assets.

[0095] Step A involves extracting features from the acquired enterprise project asset data to obtain asset feature data.

[0096] Due to the large scale of the data and the presence of errors, null values, and redundancy, the first step is based on the improved 3 Outlier removal and data deduplication based on asset unique codes are performed using criteria (adaptively adjusting upper and lower threshold multipliers using skewness coefficients). Subsequently, for structured data, numerical and enumerated features are directly extracted and concatenated into vectors (e.g., equipment with asset code AST-001, original value 1.2 million yuan, service life 10 years, risk level medium). For semi-structured maintenance logs, natural language processing techniques are used for word segmentation and semantic analysis to extract core business keywords such as "stable operation, no recent faults" and map them into text feature vectors. For unstructured equipment images and technical manuals, convolutional neural networks are used to extract visual features such as contours and wear, as well as deep text features. All feature vectors are dynamically standardized by combining the "attribute value range" of the current input asset data, eliminating the influence of units while preserving the core differences between different assets, ultimately resulting in asset feature data in a unified format.

[0097] Step B involves performing hierarchical security desensitization on asset feature data and assigning dynamic weights to each desensitized asset feature based on the contribution of asset attributes to asset value.

[0098] To prevent the leakage of ownership information, core technical parameters, and sensitive financial data, the characteristic data is divided into three levels according to the preset asset compliance strategy: high sensitivity (core financial data, absolute ownership ID), medium sensitivity (technical specifications, contact information), and low sensitivity (ordinary operating status parameters). The data is classified into three levels of security desensitization: SHA-256 hash encryption, specific symbol mask replacement, and superimposed small random perturbations with a mean of zero. The desensitization mapping relationship and the key are centrally stored in the secure area.

[0099] After anonymization, the attribute value association weighting method is used to calculate the asset value impact factor based on the historical average value contribution Va of asset attributes and the real-time value impact Vr. ( To adjust the coefficient, a value of 0.65 is set for physical equipment assets. Then, all influencing factors are normalized to ensure that the sum of the weights is 1, thereby dynamically allocating weights to each asset characteristic (for example, the weights of core value contribution characteristics such as original asset value and risk level exceed 0.6, while the weights of marginal characteristics are significantly compressed).

[0100] Step C: Construct a first asset graph to represent the relationships between assets based on asset characteristic data and dynamic weights.

[0101] Each standardized asset feature data point is mapped to an independent node in the first asset graph. Each node uses a unique, anonymized asset identifier (physically corresponding to the company's actual project asset entity, such as device AST-001). For any two nodes... , Through calculation formula

[0102] Calculate the weighted similarity (where (Dynamic weights assigned to step B). Setting a similarity threshold. (Dynamic adjustment of equipment assets), when Edges are constructed sequentially, and the weighted similarity is used as the edge weight. The final first asset graph G1 (approximately 12,000 nodes and 850,000 edges in this embodiment) is composed of the de-identified node set V, the edge set E, and the edge weight set W. This graph physically and accurately represents the business relationships between assets based on their core value attributes.

[0103] Step D: For the first asset graph, based on the preset edge weight threshold, highly similar node pairs are selected, nodes with high comprehensive value scores are retained as core nodes, and redundant node information is merged to obtain a second asset graph with a reduced node size.

[0104] Set a higher threshold for coarsening edge weights for G1. Iterate through the edge set and filter out those that satisfy the condition. Highly similar node pairs (physically corresponding redundant equipment of the same model, purchased in the same batch, and with very similar wear levels, such as AST-001 and another similar equipment AST-0092). Calculate the comprehensive value score for each pair of nodes. ( For dynamic weights, To standardize feature values, nodes with higher scores are retained as core nodes (representing typical asset entities with the greatest value contribution and most complete business profiles within the homogeneous group), while nodes with lower scores are identified as redundant nodes. The feature information of redundant nodes is then merged into the core nodes using a weighted average method (see update formula example). , and (These represent the percentages of the combined value scores of the two nodes in the total score). Simultaneously, external neighbor edges originally connected to redundant nodes are migrated to the core node, and the redundant node and its edges are deleted. This filtering, comparison, and fusion process is iteratively executed until the edge weights of any connected nodes in the graph no longer satisfy the above conditions. The output node size of the second asset graph G2 is reduced (in this embodiment, the number of nodes is efficiently compressed from about 12,000 to about 2,800 core nodes, a reduction of about 77%).

[0105] Step E involves constructing a graph trend filtering objective function that incorporates dynamic weights, and then performing alternating optimization on the graph signal of the second asset graph to obtain a set of fine-grained clusters representing the asset grouping relationships.

[0106] The multidimensional features of G2 are divided into multiple asset feature views, including financial value view, physical attribute view, and operational status view. A graph trend filter objective function incorporating dynamic weights is constructed. Due to the high complexity and non-convexity of the objective function, an alternating optimization strategy is adopted: a closed-form solution is derived using the Cauchy-Schwarz inequality, and the local preference weights B are optimized by fixing the view weights and the indicator matrix, and then the view weights are optimized by fixing the local preference weights and the indicator matrix. Optimize the clustering indicator matrix by fixing local preference weights and view weights. The process continues until the objective function converges (the difference between two adjacent iterations is less than the preset minimum or the maximum number of iterations is reached). Finally, a fine-grained set of clusters representing the asset grouping relationships is obtained (in this embodiment, it is divided into 12 main clusters, such as "high-value precision CNC machining center cluster," "low-risk logistics AGV cluster," and "medium-risk old production line equipment cluster," etc.).

[0107] Step F involves extracting the core features of each cluster in the fine-grained cluster set, performing cross-layer information fusion, and outputting the final asset aggregation result data after merging clusters that meet the preset similarity conditions, which has been verified by security.

[0108] For each cluster in the fine-grained clustering set, the standard feature vectors of all nodes within the cluster are extracted. The weighted average is calculated using the proportion of the node's comprehensive value score as the weight, thus obtaining the core feature vector of the cluster. Simultaneously, the redundant node key non-sensitive long-tail feature information removed in step D is appended and fused into the current cluster core feature vector to achieve cross-layer information fusion. The inter-cluster similarity score between the core feature vectors of any two clusters is calculated. When the score is greater than or equal to a preset inter-cluster similarity threshold, the two fine-grained clusters are merged into a unified asset cluster, and the merged core features are recalculated. Finally, all merged clusters are traversed, and strict security checks are performed to verify whether they contain plaintext sensitive fields that have not been anonymized. After confirming no risk of leakage, the final asset aggregation result dataset is encrypted using AES-256 and output through a secure channel configured with SSL / TLS protocol (in this embodiment, the final output consists of 8 main asset aggregation groups, including the core profile, total value, risk distribution, and recommended maintenance strategies for each group).

[0109] In this smart car manufacturing factory scenario, the method of this invention improves the aggregation accuracy by approximately 35% compared to traditional K-Means and by approximately 28% compared to Louvain, reduces the node processing scale by more than 77%, ensures zero data leakage throughout the entire process, and directly supports the company's asset valuation accuracy improvement by 12%, intelligent maintenance cost reduction by 18%, and investment decision-making cycle shortening by 40%. It effectively solves the problems of insufficient aggregation accuracy, low efficiency, poor adaptability to asset business characteristics, and security risks of multi-attribute project asset data pointed out in the background technology, and provides high-quality data support for subsequent asset analysis, evaluation, and decision-making.

[0110] Example 3 like Figure 2 As shown, the present invention also provides a project asset data aggregation device, the device 200 comprising: The feature extraction module 201 is used to extract features from the acquired enterprise project asset data to obtain asset feature data. The enterprise project asset data includes structured data, semi-structured data, and unstructured data. Structured data consists of one or more of the following: book value, physical dimensions, and specifications of the asset. Semi-structured data consists of business texts in daily circulation. Unstructured data assets consist of one or more of the following: photographs, CAD engineering drawings, and technical specification documents. The preprocessing module 202 is used to perform hierarchical security desensitization on asset feature data and assign dynamic weights to each desensitized asset feature based on the contribution of asset attributes to asset value. Graph construction module 203 is used to construct a first asset graph to represent the relationship between assets based on asset feature data and dynamic weights; the nodes of the first asset graph adopt a unique asset identifier after desensitization, and the edges and edge weights of the first asset graph are determined by calculating the weighted similarity between any two nodes; The coarsening aggregation module 204 is used to filter highly similar node pairs based on a preset edge weight threshold for the first asset graph, retain nodes with high comprehensive value scores as core nodes and merge redundant node information to obtain a second asset graph with reduced node size. The refinement aggregation module 205 is used to construct a graph trend filtering objective function that incorporates dynamic weights, and to alternately optimize and solve the graph signal of the second asset graph to obtain a set of fine-grained clusters that represent the grouping relationship of assets. The aggregation module 206 is used to extract the core features of each cluster in the fine-grained cluster set, perform cross-layer information fusion, and output the final asset aggregation result data after merging clusters that meet the preset similarity conditions, which has been verified by security.

[0111] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. Their specific functions and technical effects can be found in the method embodiments section, and will not be repeated here. Those skilled in the art will understand that, for ease of description and brevity, the above-mentioned division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0112] like Figure 3 As shown, embodiments of the present invention provide a terminal device, such as... Figure 3 As shown, the terminal device D10 of this embodiment includes: at least one processor D100 ( Figure 3 The diagram shows only one processor, a memory D101, and a computer program D102 stored in the memory D101 and executable on the at least one processor D100, wherein the processor D100 executes the computer program D102 to implement the steps in any of the above method embodiments.

[0113] Specifically, when the processor D100 executes the computer program D102, it extracts features from the acquired enterprise project asset data to obtain asset feature data; it performs hierarchical security desensitization on the asset feature data and assigns dynamic weights to each desensitized asset feature based on the contribution of asset attributes to asset value; it constructs a first asset graph to represent the relationship between assets based on the asset feature data and dynamic weights; for the first asset graph, it filters highly similar node pairs based on a preset edge weight threshold, retains nodes with high comprehensive value scores as core nodes, and merges redundant node information to obtain a second asset graph with reduced node size; it constructs a graph trend filtering objective function combined with dynamic weights, and alternately optimizes and solves the graph signal of the second asset graph to obtain a fine-grained cluster set representing the grouping relationship of assets; it extracts the core features of each cluster in the fine-grained cluster set for cross-layer information fusion, and after merging clusters that meet the preset similarity conditions, it outputs the final asset aggregation result data that has been security verified. This method assigns dynamic weights to each feature based on its contribution to asset value, prioritizing core value attributes in subsequent processing. This improves the alignment between the aggregation results and the actual business value of the assets, overcoming the bias caused by the traditional method's equal weighting of all attributes. By constructing an asset graph with de-identified nodes and determining edges and edge weights based on weighted similarity, the complex relationships between assets are transformed into a computable graph structure, laying the foundation for accurate aggregation. Furthermore, a coarse aggregation is performed using a combination of edge weight threshold filtering and comprehensive value score evaluation, achieving efficient dimensionality reduction while ensuring no loss of high-value information. A graph trend filtering objective function incorporating dynamic weights is then constructed and alternately optimized to achieve refined grouping with high intra-cluster similarity and significant inter-cluster differences. Finally, cross-layer information fusion outputs complete and accurate aggregation results, effectively improving data quality and providing high-quality data support for subsequent asset analysis, evaluation, and decision-making.

[0114] The processor D100 can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.

[0115] In some embodiments, the memory D101 may be an internal storage unit of the terminal device D10, such as a hard disk or memory of the terminal device D10. In other embodiments, the memory D101 may be an external storage device of the terminal device D10, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc., equipped on the terminal device D10. Furthermore, the memory D101 may include both internal and external storage units of the terminal device D10. The memory D101 is used to store the operating system, applications, bootloader, data, and other programs, such as the program code of the computer program. The memory D101 can also be used to temporarily store data that has been output or will be output.

[0116] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps described in the various method embodiments above.

[0117] This application provides a computer program product that, when run on a terminal device, enables the terminal device to implement the steps described in the various method embodiments above.

[0118] Those skilled in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of protection of this application is limited to these examples; within the framework of this application, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of one or more embodiments of this application as described above, which are not provided in detail for the sake of brevity.

[0119] One or more embodiments in this application are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of this application. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of one or more embodiments in this application should be included within the protection scope of this application.

Claims

1. A method for aggregating project asset data, characterized in that, include: Feature extraction is performed on the acquired enterprise project asset data to obtain asset feature data; The enterprise project asset data includes structured data, semi-structured data, and unstructured data; the structured data is one or more of the following: book value, physical dimensions, and specifications of the asset; the semi-structured data is business text in daily circulation; and the unstructured data assets are one or more of the following: photographs, CAD engineering drawings, and technical specification documents. The asset feature data is hierarchically desensitized, and dynamic weights are assigned to each desensitized asset feature based on the contribution of asset attributes to asset value. A first asset graph is constructed based on the asset characteristic data and the dynamic weights to represent the relationships between assets; the nodes of the first asset graph use a unique, de-identified asset identifier, corresponding to the actual project asset entity of the enterprise. The edges of the first asset graph represent aggregable business relationships or homogeneous relationships between asset entities; The edge weights of the first asset graph are used to quantify the degree of matching between the two assets in terms of their core value contribution attributes; For the first asset graph, highly similar node pairs are filtered based on a preset edge weight threshold, and the comprehensive value score of each node in the highly similar node pairs is calculated. The overall value score is determined based on the weighted sum of the feature values ​​of each node and their corresponding dynamic weights; The expression for the comprehensive value score is as follows: ; Represents a computing node The corresponding dynamic weights, Represents a computing node eigenvalues; Nodes with high overall value scores are retained as core nodes, while nodes with low overall value scores are removed as redundant nodes. The feature information of the redundant nodes is then merged into the core nodes using a weighted average method; this includes: For each pair of highly similar nodes selected Through calculation formula Obtain the comprehensive value score of each node in the node pair. , or ; Compare and Nodes with higher overall value scores are identified as core nodes with stronger representativeness and greater value retention significance, while nodes with lower overall value scores are identified as redundant nodes. Assuming the core node is Redundant nodes are The feature information of redundant nodes is proportionally integrated into the core nodes, and the feature vector of the core nodes is updated; the update formula is: , and These represent the percentage of the combined value score of the two nodes in the total score. ; Represents the core node Feature vectors before feature fusion Indicates redundant nodes eigenvectors; This represents the updated core node feature vector obtained by fusing redundant node feature information into the core node. All external neighbor nodes that were originally connected to the redundant nodes in the first asset graph are reconnected to the core nodes, and the edge weights are merged or updated accordingly. Then, the redundant nodes and their original edges are completely deleted from the first asset graph. Repeat the screening of highly similar node pairs, comparison of comprehensive value scores, and merging of features and network topology on the current first asset graph until the edge weight of any two connected nodes in the graph is less than the preset coarsening edge weight threshold, and obtain a second asset graph with reduced node size. A graph trend filtering objective function combining the dynamic weights is constructed, and the graph signal of the second asset graph is alternately optimized and solved to obtain a fine-grained cluster set representing the asset grouping relationship; The core features of each cluster in the fine-grained cluster set are extracted and cross-layer information is fused. After merging clusters that meet the preset similarity conditions, the final asset aggregation result data that has been verified by security is output.

2. The project asset data aggregation method according to claim 1, characterized in that, The contribution of asset attributes to asset value is assigned dynamic weights to each characteristic, including: The attribute value association weighting method is adopted to calculate the asset value impact factor based on the historical average value contribution and real-time value influence of asset attributes. The weight value of each asset characteristic is then dynamically calculated based on this asset value impact factor. The expression for the asset value impact factor is as follows: ;in, This represents the adjustment coefficient. , This represents the average historical value contribution of an asset attribute. This indicates the real-time value impact of asset attributes.

3. The project asset data aggregation method according to claim 2, characterized in that, The edges and edge weights of the first asset graph are determined by calculating the weighted similarity between any two nodes, including: Through calculation formula Obtain weighted similarity ; Represents a node and nodes Weighted similarity between them and This represents the node number index in the first asset diagram. and Representing nodes respectively and nodes The corresponding standardized asset characteristic values, Represents a node The corresponding weights of standardized asset eigenvalues Indicates the minimum value; When the weighted similarity is greater than or equal to a preset similarity threshold, an edge is constructed between nodes, and the weighted similarity is used as the edge weight of the corresponding edge.

4. The project asset data aggregation method according to claim 1, characterized in that, The constructed graph trend filtering objective function incorporating the dynamic weights has the following mathematical expression for minimizing the objective function: in, This represents the index of the asset feature view. This indicates the total number of asset feature views. This represents the total order of the graph difference operator. This represents the total number of clusters. This indicates adaptive weighting, which is positively correlated with asset value influencing factors. Indicates the first Each asset view Graph difference operator, Indicates the first Cluster indicator signal for each cluster, The L1 norm represents the smoothness of the cluster indicator signal on the asset graph. Represents the locus of a matrix. Represents the clustering indicator matrix. Indicates the first Weights of each asset view Indicates the first Enhanced asset graph adjacency matrix for each asset view Indicates the order index of the graph difference operator. This represents the cluster index.

5. The project asset data aggregation method according to claim 4, characterized in that, The alternating optimization solution includes: deriving a closed-form solution through the Cauchy-Schwarz inequality, and sequentially optimizing the local preference weights, view weights, and clustering indicator matrix until the objective function converges.

6. The project asset data aggregation method according to claim 5, characterized in that, The enterprise project asset data includes structured data, semi-structured data, and unstructured data; The process of extracting features from the enterprise project asset data to obtain asset feature data includes: For the structured data, semi-structured data and unstructured data in the enterprise project asset data, direct extraction, natural language processing and convolutional neural network methods are used respectively to transform them into feature vectors in a unified format.

7. A project asset data aggregation device, characterized in that, include: The feature extraction module is used to extract features from the acquired enterprise project asset data to obtain asset feature data. The enterprise project asset data includes structured data, semi-structured data, and unstructured data; the structured data is one or more of the following: book value, physical dimensions, and specifications of the asset; the semi-structured data is business text in daily circulation; and the unstructured data assets are one or more of the following: photographs, CAD engineering drawings, and technical specification documents. The preprocessing module is used to perform hierarchical security desensitization on the asset feature data, and to assign dynamic weights to each desensitized asset feature based on the contribution of asset attributes to asset value. The graph construction module is used to construct a first asset graph to represent the relationship between assets based on the asset feature data and the dynamic weights; the nodes of the first asset graph use a unique, de-identified asset identifier, corresponding to the actual project asset entity of the enterprise. The edges of the first asset graph represent the aggregateable business relationships or homogeneous relationships between asset entities; the edge weights of the first asset graph are used to quantify the degree of matching between two assets in terms of core value contribution attributes. The coarsening aggregation module is used to filter highly similar node pairs based on a preset edge weight threshold for the first asset graph, and calculate the comprehensive value score of each node in the highly similar node pair. The overall value score is determined based on the weighted sum of the feature values ​​of each node and their corresponding dynamic weights; The expression for the comprehensive value score is as follows: ; Represents a computing node The corresponding dynamic weights, Represents a computing node eigenvalues; Nodes with high overall value scores are retained as core nodes, while nodes with low overall value scores are removed as redundant nodes. The feature information of the redundant nodes is then merged into the core nodes using a weighted average method; this includes: For each pair of highly similar nodes selected Through calculation formula Obtain the comprehensive value score of each node in the node pair. , or ; Compare and Nodes with higher overall value scores are identified as core nodes with stronger representativeness and greater value retention significance, while nodes with lower overall value scores are identified as redundant nodes. Assuming the core node is Redundant nodes are The feature information of redundant nodes is proportionally integrated into the core nodes, and the feature vector of the core nodes is updated; the update formula is: , and These represent the percentage of the combined value score of the two nodes in the total score. ; Represents the core node Feature vectors before feature fusion Indicates redundant nodes eigenvectors; This represents the updated core node feature vector obtained by fusing redundant node feature information into the core node. All external neighbor nodes that were originally connected to the redundant nodes in the first asset graph are reconnected to the core nodes, and the edge weights are merged or updated accordingly. Then, the redundant nodes and their original edges are completely deleted from the first asset graph. Repeat the screening of highly similar node pairs, comparison of comprehensive value scores, and merging of features and network topology on the current first asset graph until the edge weight of any two connected nodes in the graph is less than the preset coarsening edge weight threshold, and obtain a second asset graph with reduced node size. The refinement aggregation module is used to construct a graph trend filtering objective function that incorporates the dynamic weights, and to alternately optimize and solve the graph signal of the second asset graph to obtain a set of fine-grained clusters that characterize the asset grouping relationship; The aggregation module is used to extract the core features of each cluster in the fine-grained cluster set, perform cross-layer information fusion, and output the final asset aggregation result data after merging clusters that meet the preset similarity conditions, which has been verified by security.

8. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method as described in any one of claims 1 to 6.

9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Intelligent enterprise data asset analysis method and system based on AI identification

    CN120975397A

  • Data desensitization method and device, data desensitization gateway and storage medium

    CN121841746A