Project cost analysis method, electronic equipment, readable storage medium and program product
The method automates project cost analysis by identifying and combining data with associative relationships, enhancing efficiency and accuracy in large enterprises with dispersed data systems.
Patent Information
- Application Number
- CN202510748864.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-07-15
AI Technical Summary
In the prior art, project cost analysis is inefficient and low in accuracy, mainly due to the lack of unified management caliber for data scattered in different management systems, resulting in low efficiency in manual screening and summary and incomplete or inaccurate data.
By obtaining various types of raw data of the target project in the target system, using local sensitive hashing algorithms and similarity prediction models, a combination of associated data is generated, and a combination of data with similarity greater than the threshold is determined, an associated data set is formed, and an automatic correlation processing is performed to analyze project costs.
It improves the efficiency and accuracy of project cost analysis, realizes automated data correlation and cost analysis, and reduces the error of manual intervention.
Smart Images

Figure CN120317835A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of project management, and particularly to a project cost analysis method, an electronic device, a readable storage medium, and a program product. Background Art
[0002] Cost analysis can help decision-makers evaluate and compare the economic benefits of different solutions, thus providing certain decision-making basis for decision-makers.
[0003] In recent years, with the continuous expansion of information system construction projects in terms of quantity and scale, large and complex enterprise organizations often face a series of intractable cost management problems. In these enterprises, project cost-related data, such as project data, budget data, order data, and reimbursement data, are usually scattered in different management systems. In addition, the construction of information system projects may involve multi-year, multi-batch, and multi-functional construction. At the same time, there is usually a lack of a unified management caliber in the system dimension to which the project belongs, which further increases the difficulty of integrating and analyzing cost data.
[0004] Currently, when conducting project cost analysis, it usually requires professionals to spend a lot of time and effort to screen and summarize the data related to project costs from a large amount of data before being able to analyze the project costs more accurately. This method not only has low analysis efficiency, but also has low analysis accuracy of project costs due to the inevitable incompleteness or inaccuracy of the data summarized manually. Summary of the Invention
[0005] The main purpose of this application is to provide a project cost analysis method, an electronic device, a readable storage medium, and a program product, aiming to improve the analysis efficiency and accuracy of project costs.
[0006] To achieve the above purpose, this application provides a project cost analysis method, and the method includes:
[0007] Obtain various types of original data of a target project in a target system to which it belongs; the various types of original data include project data, system data, and product data;
[0008] Generate each associated data combination based on the data that has an association relationship between every two of the various types of original data;
[0009] Determine the similarity between the two data included in each associated data combination, and use each associated data combination with the similarity greater than a preset similarity threshold as each target combination;
[0010] Perform an association process on the data included in all the target combinations to obtain an associated data set of the target project;
[0011] Analyze the cost of the target project based on the associated dataset.
[0012] In one embodiment, the step of generating each associated data combination according to the data with an association relationship between every two of various types of original data includes:
[0013] Extract and combine the data with an association relationship between every two of various types of original data to obtain each original data combination;
[0014] Based on the locality-sensitive hashing algorithm, determine the hash fingerprints corresponding to the two data in each original data combination respectively, and calculate the Hamming distance between the hash fingerprints corresponding to the two data in each original data combination respectively;
[0015] Use the original data combinations with the Hamming distance less than the preset distance threshold as each of the associated data combinations.
[0016] In one embodiment, each type of original data includes dictionary feature information. The step of determining the hash fingerprints corresponding to the two data in each original data combination respectively based on the locality-sensitive hashing algorithm includes:
[0017] For any one of the original data combinations, obtain the dictionary feature information corresponding to the two data included in the original data combination respectively, and obtain the weights corresponding to each of the dictionary feature information;
[0018] Perform binary processing on each of the dictionary feature information to obtain the binary strings corresponding to each of the dictionary feature information;
[0019] Determine the dictionary feature vectors corresponding to each of the dictionary feature information according to the weights corresponding to each of the dictionary feature information and the binary strings corresponding to each of the dictionary feature information respectively;
[0020] Based on the dictionary feature vectors corresponding to each of the dictionary feature information, determine the hash fingerprints corresponding to the two data in the original data combination respectively.
[0021] In one embodiment, the step of extracting and combining the data with an association relationship between every two of various types of original data to obtain each original data combination includes:
[0022] Perform data cleaning processing on various types of original data to obtain various types of original data after data cleaning processing;
[0023] Extract and combine the data with an association relationship between every two of the various types of original data after data cleaning processing to obtain each of the original data combinations.
[0024] In one embodiment, all types of original data include text data feature information. The step of determining the similarity between the two data included in each of the associated data combinations includes:
[0025] For any one of the associated data combinations, obtain the text data feature information of each of the two data included in the associated data combination, and perform vectorization processing on each of the text data feature information to obtain the text feature vectors corresponding to each of the text data feature information;
[0026] Input each of the text feature vectors into a pre-trained similarity prediction model to obtain the similarity between the two data included in the associated data combination.
[0027] In one embodiment, before the step of inputting each of the text feature vectors into a pre-trained similarity prediction model, the method further includes:
[0028] Obtain an original training sample set;
[0029] Set the same label for each pair of data with an associated relationship in the original training sample set to obtain a target training sample set;
[0030] Based on the bidirectional encoder representation whitening algorithm, use the label corresponding to each data included in the target training sample set and the text data feature information of each data as model inputs to train and obtain the similarity prediction model.
[0031] In one embodiment, after the step of performing association processing on the data included in all the target combinations to obtain the associated data set of the target project, the method further includes:
[0032] Obtain the verification result of the artificial person for the associated data set;
[0033] According to the verification result and the associated data set, correct the original training sample set to obtain a corrected original training sample set;
[0034] Optimize the similarity prediction model according to the corrected original training sample set to obtain a new similarity prediction model.
[0035] In addition, to achieve the above object, the present application further provides an electronic device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor. The computer program is configured to implement the steps of the project cost analysis method as described above.
[0036] In addition, to achieve the above object, the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the project cost analysis method described above are implemented.
[0037] In addition, to achieve the above object, the present application further provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, the steps of the project cost analysis method described above are implemented.
[0038] The present application provides a project cost analysis method. First, various types of original data of a target project in a target system to which it belongs are obtained. The various types of original data include project data, system data, and product data, that is, the various types of original data include all data related to the target project in the target system. Then, according to the data with an association relationship between every two of the various types of original data, various associated data combinations are generated. Since the data with an association relationship between every two of the various types of original data may all be data related to the cost of the target project in the target system, by means of these associated data combinations generated based on the pairwise associations, the data that is truly related to the cost of the target project in the target system can be analyzed. On this basis, since in the target system, the relationship between data is closely connected with the business logic, when two data have a high similarity, it means that they are closely related in terms of business scenarios and cost influencing factors, etc. Therefore, by determining the similarity between the two data included in each associated data combination and taking the associated data combinations with a similarity greater than a preset similarity threshold as each target combination, all the data that is truly related to the cost of the target project in the target system can be determined. After that, by performing an association process on the data included in all the target combinations, all the data that is truly related to the cost of the target project in the target system can be summarized to obtain an associated data set of the target project. Subsequently, using this associated data set, the cost of the target project can be accurately analyzed.
[0039] Therefore, in summary, the technical solution of the present application provides a way to automatically associate all the data related to the cost of a project in the system to which the project belongs. Compared with the conventional method that requires manual screening and summarization, it not only has high efficiency but also high accuracy. Therefore, the technical solution of the present application improves the analysis efficiency and analysis accuracy of project costs. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present application and, together with the specification, are used to explain the principles of the present application.
[0041] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0042] Figure 1 It is a schematic flowchart of the project cost analysis method provided by the first embodiment of the present application;
[0043] Figure 2 It is a schematic diagram of the implementation principle of the locality-sensitive hashing algorithm provided by the embodiments of the present application;
[0044] Figure 3 It is a schematic flowchart of the project cost analysis method provided by the embodiments of the present application;
[0045] Figure 4 It is a schematic diagram of the module structure of the project cost analysis device provided by the embodiments of the present application;
[0046] Figure 5 It is a schematic diagram of the structure of the hardware operating environment related to the embodiments of the present application.
[0047] The implementation, functional features and advantages of the present application will be further described in combination with the embodiments and with reference to the accompanying drawings. Detailed Embodiments
[0048] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.
[0049] To better understand the technical solutions of the present application, the following will be described in detail in combination with the drawings of the specification and specific embodiments.
[0050] Cost analysis can help decision-makers evaluate and compare the economic benefits of different solutions, thus providing certain decision-making basis for decision-makers.
[0051] First of all, cost analysis helps decision-makers comprehensively and deeply evaluate and carefully compare the economic benefits of different solutions, thus providing a solid and reliable basis for their decisions. With a thorough understanding of costs, decision-makers can accurately select the optimal solution that combines economy and effectiveness from numerous solutions, achieving the maximum utilization of resources.
[0052] Secondly, delving deep into various cost components can assist decision-makers in optimizing resource allocation. For example, in human resource management, reasonably planning the staffing, job settings, and salary systems to ensure that the human input matches the project requirements; in the material procurement process, reducing procurement costs by optimizing procurement strategies, selecting high-quality suppliers, and reasonably controlling inventory levels; regarding equipment usage, scientifically arranging the purchase, lease, and maintenance plans of equipment to improve the usage efficiency and lifespan of the equipment. Through such reasonable arrangements, the overall cost of the project can be effectively reduced.
[0053] Furthermore, cost analysis plays an important role in identifying and assessing potential risks of a project. Issues such as cost overruns and budget shortages can be promptly detected through cost analysis. This keen risk identification ability enables managers to quickly take corresponding measures, such as adjusting the budget, optimizing the project plan, or strengthening cost control, so as to minimize the negative impact of risks on the project and ensure the smooth progress of the project.
[0054] In addition, by comparing the budget with the actual cost, cost analysis can effectively control the actual expenditures of a project. Timely detecting and correcting situations of budget overruns or resource waste, ensuring that the project is completed within the budget and improving the cost-effectiveness of the project.
[0055] Finally, cost analysis provides key indicators for measuring project performance. By comparing the actual cost with the expected cost, the efficiency and effectiveness of project management can be objectively evaluated. For cases where the actual cost is lower than the expected cost and the project objectives are well achieved, the successful experiences can be summarized and popularized in future similar projects; for projects with actual costs higher than the expected costs, the reasons are deeply analyzed and lessons are learned for improvement in subsequent projects.
[0056] Traditional cost analysis methods can be mainly divided into the following three categories:
[0057] The first category is the distinction between direct costs and indirect costs: Direct costs are costs that can be directly attributed to a specific project or product, such as direct labor and material costs. Indirect costs are costs that are not easily directly allocated to a specific project or product, such as management and administrative expenses, and usually need to be allocated through cost allocation methods.
[0058] The second category is the classification and allocation of expenses: Cost analysis needs to classify various expenses according to their nature and attribution, such as labor costs, material costs, equipment rental costs, etc. Allocating costs is to reasonably allocate indirect costs to each project or product in order to more accurately reflect the total cost of the project.
[0059] The third category is the depreciation and amortization treatment of costs: Depreciation refers to the reduction in the value of long-term assets during their use due to factors such as the number of years and usage volume. This reduction in value is usually amortized into each stage of the project in a certain way. Amortization means spreading a certain expenditure over a period of time. For example, software development costs can be amortized and evenly distributed to the costs of each stage within the project cycle.
[0060] In recent years, with the continuous expansion of information system construction projects in terms of quantity and scale, large and complex enterprise organizations often face a series of thorny cost management problems. In these enterprises, project cost-related data, such as project data, budget data, order data, and reimbursement data, are usually scattered in different management systems. In addition, the construction of information system projects may involve multiple years, multiple batches, and multiple functions, and at the same time, there is usually a lack of a unified management caliber in the system dimension to which the project belongs, which further increases the difficulty of integrating and analyzing cost data.
[0061] Currently, when conducting project cost analysis, it usually requires professionals to spend a lot of time and effort to screen and summarize the data related to project costs from a large amount of data before a relatively accurate analysis of the project cost can be carried out. This method not only has low analysis efficiency but also results in low accuracy in the analysis of project costs due to the inevitable incompleteness or inaccuracy of the data summarized manually.
[0062] Based on this, the present application provides a project cost analysis method. First, various types of original data of the target project in the target system to which it belongs are obtained. The various types of original data include project data, system data, and product data, that is, the various types of original data include all data related to the target project in the target system. Then, according to the data with an associated relationship between every two of the various types of original data, various associated data combinations are generated. Since the data with an associated relationship between every two of the various types of original data may all be data related to the cost of the target project in the target system, by means of these associated data combinations generated based on pairwise associations, the data truly related to the cost of the target project in the target system can be analyzed. On this basis, since in the target system, the relationships between data are closely connected with the business logic, when two data have a high similarity, it means that they are closely related in terms of business scenarios and cost influencing factors, etc. Therefore, by determining the similarity between the two data included in each associated data combination and taking the associated data combinations with a similarity greater than the preset similarity threshold as each target combination, all the data truly related to the cost of the target project in the target system can be determined. After that, by performing an association process on the data included in all the target combinations, all the data truly related to the cost of the target project in the target system can be summarized to obtain the associated data set of the target project. Subsequently, using this associated data set, the cost of the target project can be accurately analyzed.
[0063] Therefore, in summary, the technical solution of the present application provides a way to automatically associate all the data related to the cost of a project in the system to which the project belongs. Compared with the conventional method that requires manual screening and summarization, it not only has high efficiency but also high accuracy. Therefore, the technical solution of the present application improves the analysis efficiency and accuracy of project costs.
[0064] The execution subject of the project cost analysis method of the present application can be an electronic device with data processing, network communication, and program running functions, or a control system, control circuit, etc. that can implement the above functions. This embodiment does not make specific limitations on this.
[0065] Taking an electronic device as the execution subject as an example, the following describes the following embodiments.
[0066] Based on this, the present application proposes a project cost analysis method for the first embodiment. Please refer to Figure 1 , the project cost analysis method includes steps S10 to S50:
[0067] Step S10, obtain various types of original data of the target project in the target system to which it belongs; the various types of original data include project data, system data, and product data;
[0068] It should be noted that the target project is the project for which costs need to be analyzed. The number of target projects can be one or multiple, and this embodiment does not make specific limitations thereon. The target system is the system to which the target project belongs. Considering that a project can be dispersed in different systems, the number of target systems can be one or multiple, and this embodiment does not make specific limitations thereon. Various types of original data include all the data related to the target project in the target system, which usually can include, but are not limited to, project data, system data, product data, etc., and this embodiment does not make specific limitations thereon.
[0069] Among them, project data refers to the data involved in the target project itself, which can include, but are not limited to, project basic information data, project contract data, project depreciation and amortization costs, project R & D cost data, project resource allocation cost data, and / or project IT maintenance fees, etc., and this embodiment does not make specific limitations thereon.
[0070] System data refers to the data of each system involved in the target project (such as xx management system, xx construction system, etc.), which can include, but are not limited to, dictionary feature information (such as fields like year, system type, affiliated department, etc.) and text data characteristic information (such as fields like system name, system construction plan, system construction goal, etc.), and this embodiment does not make specific limitations thereon. Among them, dictionary feature information refers to data fields stored in the form of a dictionary with structured attributes (i.e., data is organized in a predefined format); text data characteristic information refers to data fields stored in free text form with unstructured attributes (i.e., data has no fixed format). Product data refers to the data of each product involved in the target project (such as R & D management platform product, ERP (Enterprise Resource Planning) platform product, etc.), which can include, but are not limited to, dictionary feature information (such as fields like year, product type, affiliated department, etc.) and text data characteristic information (such as fields like product name, product description, etc.), and this embodiment does not make specific limitations thereon.
[0071] Please refer to Table 1 below to understand the types of information contained in these three types of data: project data, system data, and product data.
[0072] Table 1:
[0073]
[0074]
[0075] In addition, it should be noted that when obtaining various types of original data of a target project in a target system to which it belongs, various forms such as data interface access, file synchronization parsing, and data template import can be used to obtain various types of original data of the target project in the target system to which it belongs. This embodiment does not make specific limitations on this.
[0076] Step S20: Generate each associated data combination based on the data with an associated relationship between every two of the various types of original data.
[0077] It should be noted that when generating each associated data combination based on the data with an associated relationship between every two of the various types of original data, the data with an associated relationship between every two of the various types of original data can be extracted and combined first to obtain each original data combination; then each original data combination can be directly used as each associated data combination. In order to improve the determination efficiency of the associated data set of the target project, certain algorithms (such as the locality-sensitive hashing algorithm) can also be used to perform certain screening on each original data combination, and then the screened original data combinations are used as each associated data combination. This embodiment does not make specific limitations on the specific implementation manner of step S20.
[0078] Step S30: Determine the similarity between the two data included in each associated data combination, and use each associated data combination with a similarity greater than a preset similarity threshold as each target combination.
[0079] It should be noted that the preset similarity threshold is used as the basis for determining whether two data are highly similar. It can be a default value or can be flexibly set according to the actual situation. This embodiment does not make specific limitations on this.
[0080] Step S40: Perform an association process on the data included in all target combinations to obtain an associated data set of the target project.
[0081] It should be noted that the association process on the data included in all target combinations can be achieved by uniformly tagging the data included in all target combinations; it can also be achieved by establishing a mapping relationship for the data included in all target combinations. This embodiment does not make specific limitations on the specific implementation manner of step S40.
[0082] Step S50: Analyze the cost of the target project based on the associated data set.
[0083] It should be noted that the associated data set is essentially the intelligent matching result of project - system - product. The associated data set can be used to establish multi - dimensional project data analysis charts such as products, systems, suppliers, input - output ratios, etc., to achieve full - perspective control of the company's full - cost data. For example, the associated data set can be used to establish an accurate investment portrait of a system or product, and analyze the rationality and scientific nature of the investment in the system and product from multiple dimensions such as suppliers, input - output ratios, and progress, so as to provide a data basis for the company's investment strategy decision - making and adjustment.
[0084] In addition, the associated data set can also assist in budget review. For example, the associated data set can clearly present the historical full - cost data and change trends of the systems and products to which the project belongs, which helps budget reviewers comprehensively evaluate whether the current budget declared for the project is reasonable and provides compliance guarantee for project budget review.
[0085] This embodiment provides a project cost analysis method. First, various types of original data of the target project in the target system to which it belongs are obtained. The various types of original data include project - type data, system - type data, and product - type data, that is, the various types of original data include all data related to the target project in the target system. Then, according to the data with an associated relationship between every two of the various types of original data, various associated data combinations are generated. Since the data with an associated relationship between every two of the various types of original data may all be the data related to the cost of the target project in the target system, by means of these associated data combinations generated based on pairwise - associated data, the data truly related to the cost of the target project in the target system can be analyzed. On this basis, since in the target system, the relationship between data is closely connected with business logic, when two data have a high degree of similarity, it means that they are closely related in terms of business scenarios and cost - influencing factors, etc. Therefore, by determining the similarity between the two data included in each associated data combination, and taking the associated data combinations with a similarity greater than the preset similarity threshold as each target combination, all the data truly related to the cost of the target project in the target system can be determined. After that, by performing an association process on the data included in all target combinations, all the data truly related to the cost of the target project in the target system can be summarized to obtain the associated data set of the target project. Subsequently, using this associated data set, the cost of the target project can be accurately analyzed.
[0086] Therefore, in summary, this embodiment provides a way to automatically associate all the data related to the cost of a project in the system to which the project belongs. Compared with the conventional method that requires manual screening and summarization, it not only has high efficiency but also high accuracy. Therefore, this embodiment improves the analysis efficiency and analysis accuracy of project costs.
[0087] Based on the above first embodiment, a second embodiment of the project cost analysis method of the present application is proposed. In the second embodiment, step S20 may include steps S21 to S23:
[0088] Step S21, extract and combine the data with an association relationship between every two of various types of original data to obtain each combination of original data;
[0089] It should be noted that generally, among various types of original data, a part of the project - type data will have an association relationship with some system - type data. At the same time, a part of the system - type data will have an association relationship with some product - type data. Exemplarily, assume that among various types of original data, the project - type data X1 in various types of original data has an association relationship with the system - type data S5, the system - type data S6, and the system - type data S n all have an association relationship, the project - type data X2 has an association relationship with the system - type data S4, and the system - type data S2 has an association relationship with the product - type data P 11 and the product - type data P n all have an association relationship. Then each combination of original data may include (X1, S5), (X1, S6), (X1, S n ), (X2, S4), (S2, P 11 ) and (S2, P n ).
[0090] In a feasible implementation manner, step S21 may include steps S211 to S212:
[0091] Step S211, perform data cleaning processing on various types of original data to obtain various types of original data after data cleaning processing;
[0092] Step S212, extract and combine the data with an association relationship between every two of various types of original data after data cleaning processing to obtain each combination of original data.
[0093] In this embodiment, by performing data cleaning processing on various types of original data to remove redundant data, problematic data, etc. in the data, thereby obtaining more standardized various types of original data; then using various types of original data after data cleaning processing to determine all the data related to the cost of the real and target projects in the target system, so as to improve the quality of the determined associated data set, and further improve the analysis accuracy of the project cost.
[0094] Step S22, based on the locality - sensitive hashing algorithm, determine the hash fingerprints corresponding to each of the two data in each combination of original data, and calculate the Hamming distance between the hash fingerprints corresponding to each of the two data in each combination of original data;
[0095] It should be noted that the locality-sensitive hashing algorithm can be the SimHash (similarity hashing) algorithm in LSH (Locality Sensitive Hashing), or the Spherical LSH (spherical locality-sensitive hashing) algorithm, etc. This embodiment does not make specific limitations on this. The hash fingerprint is a fixed-length string obtained by calculating data through a hash algorithm, which has uniqueness and representativeness and is used to identify the characteristics and content of the data. The Hamming distance is used to measure the degree of difference between two equal-length strings.
[0096] In a feasible implementation manner, all types of original data include dictionary feature information, and step S22 may include steps S221 to S224:
[0097] Step S221, for any original data combination, obtain the respective dictionary feature information of the two data included in the original data combination, and obtain the weights corresponding to the respective dictionary feature information;
[0098] It should be noted that the dictionary feature information refers to a set of information that is organized and represented in a dictionary data structure, has a specific meaning, and can describe certain attributes or characteristics of the data. When obtaining the weights corresponding to the respective dictionary feature information, various dictionary feature information can be pre-trained in advance using the locality-sensitive hashing algorithm to obtain the weights corresponding to the respective dictionary feature information, and a relational table is used for recording. Thus, the weights corresponding to the respective dictionary feature information of the two data included in the original data combination can be directly obtained by looking up the table.
[0099] Among them, the process of pre-training various dictionary feature information using the locality-sensitive hashing algorithm to obtain the weights corresponding to the respective dictionary feature information may include: first, obtain the original training sample set (the original training sample set can be a set composed of various original data combinations), and then extract the respective dictionary feature information in the original training sample set (such as year, type, affiliated department, person in charge, etc.); then, by counting the frequency of the same dictionary feature information in the original training sample set, the weights corresponding to the respective dictionary feature information can be calculated. The specific calculation process can refer to the following formulas 1 to 3:
[0100]
[0101] Among them, P(T i ) is the probability that there is an association relationship for the dictionary feature information T i n i is the number of times the dictionary feature information T iThe same frequency, where m is the total number of associated pairs in the original training sample set (for example, there is an association relationship between item class data X and system class data S, and there are 100 data pairs in the original dataset expressing the association relationship between item class data X and system class data S, then m is 100).
[0102]
[0103] Among them, P(T) is the total probability of feature values, and k is the total number of dictionary feature information contained in the original training sample set.
[0104]
[0105] Among them, w(T i ) is the weight corresponding to the dictionary feature information T i .
[0106] Exemplarily, assume that there are 100 data pairs in the original dataset expressing the association relationship between item class data X and system class data S, that is, m is 100. Taking the dictionary feature information of "year" as an example, it is statistically found that the "year" of 20 associated pairs is the same, that is, n i is 20. Then, using the above formula 1, it can be known that the probability P(T i ) of the dictionary feature information of "year" having an association relationship is 0.2; assume that there are three dictionary feature information of "year", "type", and "person in charge" in the original dataset, that is, k is 3, and their respective frequencies are 20, 25, and 30. Then, using the above formula 2, it can be known that the total probability of feature values P(T) is 0.75; then, using the above formula 3, the weight w(T i ) of the dictionary feature information of "year" can be determined to be 0.27.
[0107] Step S222: Perform binary processing on each dictionary feature information to obtain the binary string corresponding to each dictionary feature information;
[0108] It should be noted that a dictionary feature information usually includes multiple feature fields, and different feature fields usually also correspond to different weights. After the dictionary feature information is subjected to binary processing, each feature field will form a binary string, so a dictionary feature information usually corresponds to multiple binary strings.
[0109] Step S223: Determine the dictionary feature vector corresponding to each dictionary feature information according to the weight corresponding to each dictionary feature information and the binary string corresponding to each dictionary feature information;
[0110] It should be noted that when determining the dictionary feature vectors corresponding to the respective dictionary feature information based on the weights corresponding to the respective dictionary feature information and the binary strings corresponding to the respective dictionary feature information, for any dictionary feature information, by using the weight corresponding to the dictionary feature information, assignment processing is performed on the multiple binary strings corresponding to the dictionary feature information (for example, assuming the weight is w, the bit with a value of 1 in the binary string can be assigned w, and the bit with a value of 0 can be assigned -w); after accumulating the assigned binary strings, the dictionary feature vector corresponding to the dictionary feature information can be obtained.
[0111] Step S224: Based on the dictionary feature vectors corresponding to the respective dictionary feature information, determine the hash fingerprints corresponding to the two data in the original data combination respectively.
[0112] It should be noted that by performing binary processing on the dictionary feature vectors corresponding to the respective dictionary feature information, the hash fingerprints corresponding to the two data in the original data combination can be obtained respectively. The operation logic of the locality-sensitive hashing algorithm is the implementation process of the above steps S221 to S224.
[0113] Exemplarily, to facilitate understanding of the implementation process of the above steps S221 to S224, please refer to Figure 2 , specifically: for any data included in the original data combination, after extracting the dictionary feature information of any data included in the original data combination, assuming that the dictionary feature information includes three feature fields, then after performing binary processing on the dictionary feature information, three binary strings can be obtained, namely 100110, 110000, and 001001 in the figure, and then by using the weights to assign values to these three binary strings, three vectors can be obtained, namely [w1, -w1, -w1, w1, w1, -w1], [w2, w2, -w2, -w2, -w2, -w2], and [-w3, -w3, w3, -w3, -w3, w3] in the figure; then by accumulating these three vectors, the dictionary feature vector corresponding to the dictionary feature information can be obtained, namely [0.24, 0.04, -0.24, -0.16, -0.16, -0.24] in the figure; after that, by performing binary processing on the dictionary feature vector, the hash fingerprint can be obtained, namely 110000 in the figure.
[0114] Step S23: Use the original data combinations with Hamming distances less than the preset distance threshold as the respective associated data combinations.
[0115] It should be noted that the preset distance threshold is used as the basis for determining whether the two data in the original data combination have a high degree of similarity. It can be a default value or can be flexibly set by the user according to the data situation, and this embodiment does not make specific limitations on this.
[0116] Exemplarily, assume that the two hash fingerprints are respectively:
[0117] Y1 = 1000000010000000100000001000000010000000100000001000000010000000;
[0118] Y2 = 1111000010000000100000001000000010000000100000001000000010000000;
[0119] Since the second, third, and fourth digits of these two hash fingerprints are different in sequence starting from the first digit, the Hamming distance between the two can be determined to be 3.
[0120] It can be understood that if all the original data combinations are directly used as the associated data combinations, then there will be more combinations for which the similarity needs to be determined subsequently (that is, the matching calculation amount calculated according to the Cartesian product will be relatively large. The matching calculation amount between project - type data X and system - type data S is X×S, and the matching calculation amount between system - type data S and product - type data P is S×P), which will affect the analysis efficiency of project costs.
[0121] On this basis, in this embodiment, after extracting and combining the data with an associated relationship between every two of various types of original data to obtain each original data combination, the locality - sensitive hashing algorithm is used to determine the Hamming distance between the hash fingerprints corresponding to the two data in each original data combination, so as to screen out the original data combinations with relatively high data similarity from each original data combination as the associated data combinations by using the Hamming distance; thus, subsequently, when it is necessary to specifically determine the similarity between every two of the data included in the combination, since the number of combinations for which the similarity needs to be determined has been reduced to a certain extent, the determination efficiency of the target combination can be improved to a certain extent, and further, the analysis efficiency of project costs can be improved.
[0122] For example, assume that by using the locality - sensitive hashing algorithm, (X1,S1), (X1,S2), (X1,S3), and (X2,S2) are removed from each original data combination, then the Cartesian product of X×S and S×P can be streamlined to {(X1,S4),(X1,S7),(X1,S n ),(X2,S1),(X2,S 12 ),(X2,S n )...}.
[0123] Based on the above first embodiment and / or second embodiment, a third embodiment of the project cost analysis method of the present application is proposed. In the third embodiment, all types of original data include text data feature information, and step S30 may include steps S31 to S32:
[0124] Step S31, for any associated data combination, obtain the text data feature information of each of the two data included in the associated data combination, and perform vectorization processing on each text data feature information to obtain the text feature vectors corresponding to each text data feature information;
[0125] It should be noted that the text data feature information refers to the information extracted from the text data that can reflect the semantic, syntactic, structural, thematic and other aspects of the text.
[0126] Step S32, input each text feature vector into a pre-trained similarity prediction model to obtain the similarity between the two data included in the associated data combination.
[0127] Among them, the construction process of the similarity prediction model may include steps S01 to S03:
[0128] Step S01, obtain the original training sample set;
[0129] It should be noted that the original training sample set may be a set composed of each original data combination.
[0130] Step S02, set the same label for each pair of data with an associated relationship in the original training sample set to obtain the target training sample set;
[0131] For example, assume that there is an associated relationship between the project class data X1 and the system class data S6 in the original training sample set. Then, the same label L1 can be assigned to the project class data X1 and the system class data S6, so as to form the data <X1, L1> and <S6, L1>.
[0132] Step S03, based on the bidirectional encoder representation method whitening algorithm, use the label corresponding to each data included in the target training sample set and the text data feature information of each data as the model input to train and obtain the similarity prediction model.
[0133] It should be noted that the Bidirectional Encoder Representations from Transformers (BERT) is a powerful pre-trained language model that can capture semantic information in text. The BERT Whitening algorithm is a technique for optimizing the word vectors output by the BERT model. By adjusting the vector distribution, it makes the vectors of different categories more distinguishable, thereby improving the performance of the model in downstream tasks.
[0134] Additionally, it should be noted that the specific process of step S03 may include: Based on the BERT Whitening algorithm, use the labels corresponding to each data and the text data feature information of each data in the target training sample set as model inputs, and through supervised learning training, output the original model; then use the original model to batch-test the data in the original training sample set to return the average loss and prediction results; then fine-tune the parameters of the BERT Whitening algorithm according to the average loss and prediction results, and repeat the step of using the labels corresponding to each data and the text data feature information of each data in the target training sample set as model inputs based on the BERT Whitening algorithm to train and obtain the similarity prediction model.
[0135] Since the text data feature information of the data comprehensively and meticulously covers various attributes of the data such as semantics, structure, and context, it can accurately depict the unique features of the data. And the similarity prediction model has undergone repeated optimization and adjustment, learning the complex similarity relationship patterns and internal laws between the data. Therefore, in this embodiment, by using the text data feature information of the data and the pre-trained similarity prediction model, the similarity between two data can be accurately analyzed.
[0136] Based on the above first embodiment, second embodiment, and / or third embodiment, a fourth embodiment of the project cost analysis method of this application is proposed. In the fourth embodiment, after step S40, steps S41 to S43 may further be included:
[0137] Step S41, obtain the verification result of the artificial check for the associated data set;
[0138] Step S42, according to the verification result and the associated data set, correct the original training sample set to obtain the corrected original training sample set;
[0139] It should be noted that when the verification result is that the associated data set is correct, the associated data set can be added to the original training sample set to obtain the corrected original training sample set; when the verification result is that the associated data set is incorrect, the associated data can be removed from the original training sample set to obtain the corrected original training sample set.
[0140] In step S43, optimize the similarity prediction model based on the corrected original training sample set to obtain a new similarity prediction model.
[0141] It can be understood that considering that project relationships are usually relatively complex, there is a possibility of misjudgment when associating data with high similarity only through NLP (Natural Language Processing). Therefore, the determined associated data set can be verified manually. After the manual verification is completed, the verification result of the manual verification for the associated data set can be obtained; then, based on the verification result and the associated data set, the original training sample set is corrected to obtain a corrected original training sample set; afterwards, the corrected original training sample set is used to optimize the similarity prediction model, so as to obtain a similarity prediction model with higher accuracy in predicting data similarity, thereby improving the accuracy of the determined associated data set subsequently and improving the analysis accuracy of project costs.
[0142] In other embodiments, the corrected original training sample set can also be used to optimize the weights corresponding to the respective dictionary feature information obtained by pre-training various dictionary feature information using the locality-sensitive hashing algorithm mentioned above. With the optimized weights, the screening of each original data combination can be better realized, the determination efficiency and accuracy of the subsequent associated data set can be improved, and thus the analysis efficiency and analysis accuracy of project costs can be improved.
[0143] Exemplarily, to facilitate understanding of the implementation process of the project cost analysis method obtained after combining the above embodiments, please refer to Figure 3 , specifically:
[0144] First, obtain various types of original data of each target project in the target system to which it belongs. The various types of original data may include project - related data, system - related data, product - related data, etc.; then, perform data cleaning on the various types of original data to obtain the various types of original data after data cleaning; next, extract and summarize the data with an association relationship between any two of the various types of original data after data cleaning to obtain each combination of original data; then, through the locality - sensitive hashing algorithm, use the weights corresponding to each pre - trained dictionary feature information to screen out each associated data combination from each combination of original data; next, use the pre - trained similarity prediction model to determine the similarity between the two data included in each associated data combination, so as to perform intelligent matching of project - system - product, and obtain the associated data set of the target project; after that, the associated data set can be manually verified, and according to the verification results fed back manually, the weights corresponding to each dictionary feature information and the pre - training of the similarity prediction model can be performed to optimize the weights corresponding to each dictionary feature information and the similarity prediction model; finally, the associated data set can be used to analyze the cost of the target project.
[0145] It should be noted that the above examples are only used to assist in understanding the present application and do not constitute a limitation on the method for analyzing the project cost of the present application. Based on this technical concept, more forms of simple transformation are within the protection scope of the present application.
[0146] The embodiment of the present application also provides a device for analyzing project cost. Please refer to Figure 4 , and the device includes:
[0147] An original data acquisition module 10, configured to obtain various types of original data of a target project in the target system to which it belongs; the various types of original data include project - related data, system - related data, and product - related data;
[0148] An associated data combination generation module 20, configured to generate each associated data combination according to the data with an association relationship between any two of the various types of original data;
[0149] A similarity determination module 30, configured to determine the similarity between the two data included in each associated data combination, and use the associated data combinations with a similarity greater than a preset similarity threshold as each target combination;
[0150] A data association module 40, configured to perform association processing on the data included in all target combinations to obtain the associated data set of the target project;
[0151] A cost analysis module 50, configured to analyze the cost of the target project based on the associated data set.
[0152] In an embodiment, the associated data combination generation module 20 is further configured to:
[0153] Extract and combine the data with an associated relationship between each pair of various types of original data to obtain various original data combinations;
[0154] Based on the locality-sensitive hashing algorithm, determine the hash fingerprints corresponding to each of the two data in each original data combination, and calculate the Hamming distance between the hash fingerprints corresponding to each of the two data in each original data combination;
[0155] Take the original data combinations with a Hamming distance less than the preset distance threshold as the associated data combinations.
[0156] In one embodiment, each type of original data includes dictionary feature information, and the associated data combination generation module 20 is further configured to:
[0157] For any original data combination, obtain the dictionary feature information corresponding to each of the two data included in the original data combination, and obtain the weights corresponding to each dictionary feature information;
[0158] Perform binary processing on each dictionary feature information to obtain the binary string corresponding to each dictionary feature information;
[0159] According to the weights corresponding to each dictionary feature information and the binary string corresponding to each dictionary feature information, determine the dictionary feature vector corresponding to each dictionary feature information;
[0160] Based on the dictionary feature vectors corresponding to each dictionary feature information, determine the hash fingerprints corresponding to each of the two data in the original data combination.
[0161] In one embodiment, the associated data combination generation module 20 is further configured to:
[0162] Perform data cleaning processing on each type of original data to obtain each type of original data after data cleaning processing;
[0163] Extract and combine the data with an associated relationship between each pair of various types of original data after data cleaning processing to obtain various original data combinations.
[0164] In one embodiment, each type of original data includes text data feature information, and the similarity determination module 30 is further configured to:
[0165] For any associated data combination, obtain the text data feature information corresponding to each of the two data included in the associated data combination, and perform vectorization processing on each text data feature information to obtain the text feature vector corresponding to each text data feature information;
[0166] Input each text feature vector into a pre-trained similarity prediction model to obtain the similarity between the two data included in the associated data combination.
[0167] In one embodiment, the project cost analysis device further includes a model training module for:
[0168] Obtain the original training sample set;
[0169] Set the same label for each pair of related data in the original training sample set to obtain the target training sample set;
[0170] Based on the bidirectional encoder representation whitening algorithm, use the labels corresponding to the data included in the target training sample set and the text data feature information of each data as the model input to train a similarity prediction model.
[0171] In one embodiment, the project cost analysis device further includes a model optimization module for:
[0172] Obtain the verification result of the manual verification of the associated data set;
[0173] According to the verification result and the associated data set, correct the original training sample set to obtain the corrected original training sample set;
[0174] Optimize the similarity prediction model based on the corrected original training sample set to obtain a new similarity prediction model.
[0175] The project cost analysis device provided by this application adopts the project cost analysis method in the above embodiment, which can improve the analysis efficiency and accuracy of project costs. Compared with the prior art, the beneficial effects of the project cost analysis device provided by this application are the same as those of the project cost analysis method provided by the above embodiment, and other technical features in the project cost analysis device are the same as the features disclosed in the method of the above embodiment, which will not be elaborated here.
[0176] This application embodiment also provides an electronic device, which includes: at least one processor; and a memory communicatively connected to at least one processor; wherein, the memory stores instructions executable by at least one processor, and the instructions are executed by at least one processor to enable at least one processor to execute the project cost analysis method in the above embodiment.
[0177] Next, refer to Figure 5 , which shows a schematic structural diagram of an electronic device suitable for implementing the embodiment of this application. Figure 5 The electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiment of this application.
[0178] As Figure 5As shown, the electronic device may include a processing device 101 (such as a central processing unit, a graphics processing unit, etc.), which may perform various appropriate actions and processes according to a program stored in the read-only memory 102 or a program loaded from the storage device 103 into the random access memory 104. In the random access memory 104, various programs and data required for the operation of the electronic device are also stored. The processing device 101, the read-only memory 102, and the random access memory 104 are connected to each other through a bus 105. The input / output interface 106 is also connected to the bus 105. Generally, the following systems may be connected to the input / output interface 106: an input device 107 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 108 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 103 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 109. The communication device 109 may allow the electronic device to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows an electronic device with various systems, it should be understood that it is not required to implement or have all the shown systems, and instead, more or fewer systems may be implemented or had.
[0179] Specifically, according to an embodiment of the present disclosure, the process described above with reference to the flowchart may be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the method shown in the flowchart. In such an embodiment, the computer program may be downloaded and installed from a network through the communication device, or installed from the storage device 103, or installed from the read-only memory 102. When the computer program is executed by the processing device 101, the above-mentioned functions defined in the method of the embodiments of the present application are executed.
[0180] The electronic device provided by the embodiments of the present application, adopting the project cost analysis method in the above-mentioned embodiments, can improve the analysis efficiency and accuracy of project costs. Compared with the prior art, the beneficial effects of the electronic device provided by the embodiments of the present application are the same as those of the project cost analysis method provided by the above-mentioned embodiments, and other technical features in the electronic device are the same as those disclosed in the method of the above-mentioned embodiments, and will not be elaborated herein.
[0181] It should be understood that each part of the embodiments of the present application may be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics may be combined in a suitable manner in any one or more embodiments or examples.
[0182] As described above, it is only the specific implementation manner of the embodiments of the present application. However, the protection scope of the embodiments of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present application, and all of them should be covered within the protection scope of the embodiments of the present application. Therefore, the protection scope of the embodiments of the present application shall be subject to the protection scope of the above-mentioned claims.
[0183] The embodiments of the present application further provide a computer-readable storage medium storing a computer program that can be run on a processor, and the computer program is used to execute the project cost analysis method in the above embodiments.
[0184] The computer-readable storage medium provided by the embodiments of the present application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0185] The above computer-readable storage medium may be included in an electronic device; or it may exist separately without being assembled into the electronic device.
[0186] The above computer-readable storage medium carries one or more programs, which, when executed by an electronic device, cause the electronic device to: obtain various types of original data of a target project in a target system to which the target project belongs; the various types of original data include project data, system data, and product data; generate respective associated data combinations based on the data having an association relationship between any two of the various types of original data; determine the similarity between the two data included in each associated data combination, and use the associated data combinations with a similarity greater than a preset similarity threshold as respective target combinations; perform an association process on the data included in all the target combinations to obtain an associated data set of the target project; and analyze the cost of the target project based on the associated data set.
[0187] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partially on the user's computer, execute as a stand-alone software package, execute partially on the user's computer and partially on a remote computer, or execute entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., connected through the Internet using an Internet service provider).
[0188] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in an order different from that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0189] The modules involved in the embodiments of the present application can be implemented in software or in hardware. In some cases, the name of the module does not constitute a limitation on the unit itself.
[0190] The computer-readable storage medium provided by the embodiments of the present application stores computer-readable program instructions for executing the above project cost analysis method, which can improve the analysis efficiency and accuracy of project costs. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by the embodiments of the present application are the same as those of the project cost analysis method provided by the above embodiments, and will not be elaborated here.
[0191] The embodiments of the present application also provide a computer program product, including a computer program, which implements the steps of the project cost analysis method as described above when executed by a processor.
[0192] The computer program product provided by the embodiments of the present application can improve the analysis efficiency and accuracy of project costs. Compared with the prior art, the beneficial effects of the computer program product provided by the embodiments of the present application are the same as those of the project cost analysis method provided by the above embodiments, and will not be elaborated here.
[0193] The above are only the preferred embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent scope of the present application.
Claims
1. A project cost analysis method, characterized in that, The method includes: Obtaining various types of original data of a target project in a target system to which the target project belongs; the various types of original data include project - type data, system - type data, and product - type data; Generating respective associated data combinations based on the data with an associated relationship between every two of the various types of original data; Determining the similarity between the two data included in each of the associated data combinations, and taking the associated data combinations with the similarity greater than a preset similarity threshold as respective target combinations; Performing an association process on the data included in all the target combinations to obtain an associated data set of the target project; Analyzing the cost of the target project based on the associated data set.
2. The project cost analysis method according to claim 1, wherein The step of generating respective associated data combinations based on the data with an associated relationship between every two of the various types of original data includes: Extracting and combining the data with an associated relationship between every two of the various types of original data to obtain respective original data combinations; Based on the locality - sensitive hashing algorithm, determining the respective hash fingerprints corresponding to the two data in each of the original data combinations, and calculating the Hamming distance between the respective hash fingerprints corresponding to the two data in each of the original data combinations; Taking the original data combinations with the Hamming distance less than a preset distance threshold as the respective associated data combinations.
3. The project cost analysis method according to claim 2, characterized in that Each of the various types of original data includes dictionary feature information. The step of determining the respective hash fingerprints corresponding to the two data in each of the original data combinations based on the locality - sensitive hashing algorithm includes: For any one of the original data combinations, obtaining the respective dictionary feature information of the two data included in the original data combination, and obtaining the weights corresponding to the respective dictionary feature information; Performing binary processing on the respective dictionary feature information to obtain the binary strings corresponding to the respective dictionary feature information; Determining the respective dictionary feature vectors corresponding to the respective dictionary feature information according to the weights corresponding to the respective dictionary feature information and the binary strings corresponding to the respective dictionary feature information; Based on the respective dictionary feature vectors corresponding to the respective dictionary feature information, determining the respective hash fingerprints corresponding to the two data in the original data combination.
4. The project cost analysis method according to claim 2, wherein The step of extracting and combining the data with an associated relationship between every two of the various types of original data to obtain respective original data combinations includes: Performing data cleaning processing on the various types of original data to obtain the various types of original data after data cleaning processing; Extracting and combining the data with an associated relationship between every two of the various types of original data after data cleaning processing to obtain the respective original data combinations.
5. The project cost analysis method according to any one of claims 1 to 4, characterized in that Each of the various types of original data includes text data feature information. The step of determining the similarity between the two data included in each of the associated data combinations includes: For any one of the associated data combinations, obtaining the respective text data feature information of the two data included in the associated data combination, and performing vectorization processing on the respective text data feature information to obtain the respective text feature vectors corresponding to the respective text data feature information; Inputting the respective text feature vectors into a pre - trained similarity prediction model to obtain the similarity between the two data included in the associated data combination.
6. The project cost analysis method according to claim 5, wherein Before the step of inputting each of the text feature vectors into a pre-trained similarity prediction model, the method further includes: Obtain an original training sample set; Set the same label for each pair of related data in the original training sample set to obtain a target training sample set; Based on the bidirectional encoder representation whitening algorithm, use the labels corresponding to each data and the text data feature information of each data included in the target training sample set as model inputs to train the similarity prediction model.
7. The project cost analysis method according to claim 6, characterized in that After the step of performing an association process on the data included in all the target combinations to obtain the association data set of the target project, the method further includes: Obtain the verification result of an artificial person for the association data set; According to the verification result and the association data set, correct the original training sample set to obtain a corrected original training sample set; Optimize the similarity prediction model according to the corrected original training sample set to obtain a new similarity prediction model.
8. An electronic device, characterized in that, The electronic device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the project cost analysis method according to any one of claims 1 to 7.
9. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, it implements the steps of the project cost analysis method according to any one of claims 1 to 7.
10. A computer program product, characterized in that, The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps of the project cost analysis method according to any one of claims 1 to 7.
Citation Information
Cited By
Individual asset cost allocation and collection analysis method, equipment and medium
CN120725285A