Method and system for classifying engineering bill of materials data based on multi-feature fingerprints
By using a multi-feature fingerprinting method, the problem of inaccurate classification results in the automatic classification of bill of quantities data was solved, and the reliability and efficiency of comparative analysis of quantities, unit prices and cost composition were improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NORTHWEST ENGINEERING CORPORATION LIMITED
- Filing Date
- 2026-06-24
- Publication Date
- 2026-07-21
AI Technical Summary
In the process of automatically classifying bill of quantities data in engineering cost data processing, existing technologies are prone to inaccurate classification results, which affect the reliability of comparative analysis of quantities, unit prices, and cost composition. This is mainly because different descriptions of the same bill of quantities content, differences in engineering parts, and differences in attribute data are not effectively identified and processed.
By using a multi-feature fingerprint-based method, engineering list data is obtained and processed for name normalization, part identification, and attribute standardization to generate multi-feature fingerprints, which are then matched with standard list items to improve the accuracy and efficiency of classification.
It enables efficient and accurate classification of bill of quantities data, reduces misclassification caused by synonyms with different names or the same name for different parts, and improves the reliability of comparative analysis of quantities, unit prices, and cost composition.
Smart Images

Figure CN122432343A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of data processing technology, and more specifically, to a method and system for classifying engineering bill of quantities data based on multi-feature fingerprints. Background Technology
[0002] In the data processing scenarios of engineering cost estimation in hydropower, construction, and other projects, price limit documents, quotation documents, and contract list documents are usually stored in spreadsheet format, with data for different parts of the project often distributed across different worksheets. When conducting cost analysis, engineers need to collect and compare the same or similar list contents from multiple similar projects to determine the reasonableness of the quantities, unit prices, and cost composition.
[0003] Existing processing methods typically rely on manual compilation, tabular function matching, or text similarity matching. Manual compilation is inefficient and prone to errors; tabular function matching mainly depends on identical names, making it difficult to handle situations where the same item in the bill of quantities has different descriptions; text similarity matching focuses primarily on the name itself, easily overlooking the engineering part to which the bill of quantities data belongs, leading to the merging of bill of quantities items with the same name from different parts. Therefore, existing technologies are prone to inaccurate classification results during the automatic classification of engineering bill of quantities data, which in turn affects the reliability of subsequent comparative analysis of quantities, unit prices, and cost composition. Summary of the Invention
[0004] The purpose of this disclosure is to provide a method, system, electronic device, and computer-readable storage medium for classifying engineering list data based on multi-feature fingerprints. This method can process engineering list names, source data of engineering parts, and engineering attribute data to generate multi-feature fingerprints to determine target standard list items, thereby improving the efficiency and accuracy of engineering list data classification.
[0005] Other features and advantages of this disclosure will become apparent from the following detailed description, or may be learned in part from practice of this disclosure.
[0006] According to a first aspect of the present disclosure, a method for classifying engineering list data based on multi-feature fingerprints is provided, comprising: acquiring engineering list data to be classified and a list classification knowledge base; wherein the engineering list data to be classified includes an engineering list name, engineering part source data, and engineering attribute data, and the list classification knowledge base includes multiple standard list items and list description mapping data associated with each of the standard list items; using the list description mapping data, performing name normalization processing on the engineering list name to obtain a standard list name corresponding to the engineering list data to be classified; performing part identification processing on the engineering part source data to obtain engineering part information corresponding to the engineering list data to be classified, and performing attribute standardization processing on the engineering attribute data to obtain list attribute information corresponding to the engineering list data to be classified; performing fingerprint fusion processing on the standard list name, the engineering part information, and the list attribute information to obtain a multi-feature fingerprint corresponding to the engineering list data to be classified; performing matching processing on the multi-feature fingerprint and the multiple standard list items, and determining a target standard list item corresponding to the engineering list data to be classified from the multiple standard list items according to the matching processing result.
[0007] According to a second aspect of the present disclosure, a system for classifying engineering list data based on multi-feature fingerprints is provided. The system includes: a data acquisition module for acquiring engineering list data to be classified and a list classification knowledge base; wherein the engineering list data to be classified includes an engineering list name, engineering location source data, and engineering attribute data, and the list classification knowledge base includes multiple standard list items and list description mapping data associated with each standard list item; a name normalization module for performing name normalization processing on the engineering list name using the list description mapping data to obtain the standard list name corresponding to the engineering list data to be classified; and a location attribute module for processing the engineering list data... The project location data is processed for location identification to obtain the project location information corresponding to the project list data to be classified. The project attribute data is then processed for attribute standardization to obtain the list attribute information corresponding to the project list data to be classified. A fingerprint fusion module is used to perform fingerprint fusion processing on the standard list name, the project location information, and the list attribute information to obtain a multi-feature fingerprint corresponding to the project list data to be classified. A classification matching module is used to match the multi-feature fingerprint with the multiple standard list items, and based on the matching results, determine the target standard list item corresponding to the project list data to be classified from among the multiple standard list items.
[0008] According to a third aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory storing computer-readable instructions that, when executed by the processor, implement the engineering bill of quantities data classification method based on multi-feature fingerprints as described in the first aspect.
[0009] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the engineering list data classification method based on multi-feature fingerprints as described in the first aspect.
[0010] The technical solutions provided in this disclosure may have the following beneficial effects: According to the engineering list data classification method based on multi-feature fingerprints in this example embodiment, on the one hand, the engineering list names are normalized through list description mapping data, so that engineering list names with different descriptions can be converted into standard list names, thereby reducing scattered classification caused by synonyms with different names; on the other hand, by performing part identification processing on the source data of engineering parts and attribute standardization processing on the engineering attribute data, the engineering list data to be classified can carry both engineering part information and list attribute information, thereby reducing misclassification caused by different parts with the same name or inconsistent attributes; furthermore, by performing fingerprint fusion processing on standard list names, engineering part information, and list attribute information, and matching processing with multiple standard list items based on multi-feature fingerprints, the determination process of target standard list items has multi-dimensional matching basis, thereby improving the reliability of matching results; thus, it has the advantage of improving the efficiency and accuracy of engineering list data classification.
[0011] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0012] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure. It is obvious that the drawings described below are merely some embodiments of this disclosure, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.
[0013] Figure 1 The illustration shows a flowchart of a multi-feature fingerprint-based engineering inventory data classification method according to some embodiments of the present disclosure.
[0014] Figure 2 The illustration shows a schematic diagram of an engineering part keyword matching library according to some embodiments of the present disclosure.
[0015] Figure 3 A schematic diagram of an engineering part numbering rule library according to some embodiments of the present disclosure is shown.
[0016] Figure 4 A schematic diagram of an engineering part content feature library according to some embodiments of the present disclosure is shown.
[0017] Figure 5 A schematic diagram illustrating a comparison of bill of quantities results according to some embodiments of the present disclosure is shown.
[0018] Figure 6 The illustration shows a schematic diagram of the comparison results of the list unit prices according to some embodiments of the present disclosure.
[0019] Figure 7 The illustration shows a schematic diagram of the comparison results of bill of quantities costs according to some embodiments of the present disclosure.
[0020] Figure 8 The diagram illustrates a block diagram of a multi-feature fingerprint-based bill of materials data classification system according to some embodiments of the present disclosure.
[0021] Figure 9 The schematic diagram illustrates the structural schematic of a computer system of an electronic device according to some embodiments of the present disclosure.
[0022] In the accompanying drawings, the same or corresponding reference numerals indicate the same or corresponding parts. Detailed Implementation
[0023] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this specification. Rather, they are merely examples of systems and methods consistent with some aspects of this specification as detailed in the appended claims.
[0024] The terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of this specification. The singular forms “a,” “the,” and “the” as used in this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.
[0025] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be more thorough and complete, and will fully convey the concept of the exemplary embodiments to those skilled in the art.
[0026] Furthermore, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. Numerous specific details are provided in the following description to give a thorough understanding of embodiments of this disclosure. However, those skilled in the art will recognize that the technical solutions of this disclosure can be practiced without one or more of the specific details, or other methods, components, systems, steps, etc., can be employed. In other instances, well-known methods, systems, implementations, or operations are not shown or described in detail to avoid obscuring various aspects of this disclosure.
[0027] Furthermore, the accompanying drawings are for illustrative purposes only and are not necessarily drawn to scale. The block diagrams shown in the drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor systems and / or microcontroller systems.
[0028] In large-scale engineering construction, price limit documents, quotation documents, and contract list documents for hydropower projects are usually stored in spreadsheet format. These documents often contain multiple worksheets, each corresponding to a different part of the project. Each worksheet records data such as the name, unit, quantity, unit price, and total price of the corresponding list items. When conducting cost analysis, engineers typically need to collect and compare data on the same parts and similar or identical list items from multiple similar projects to determine the reasonableness of the quantities, unit prices, and cost composition.
[0029] Common processing methods in related technologies include manual comparison, tabular function matching, fuzzy text matching, and machine learning classification. Manual comparison relies on line-by-line review and copying, which is inefficient and prone to errors. Tabular function matching typically relies on identical list names, making it difficult to identify the correspondence between different descriptions such as anchor bolts, mortar anchor bolts, and system anchor bolts. Fuzzy text matching primarily judges based on name similarity, easily overlooking the engineering location to which a list item belongs, leading to the incorrect merging of list items with the same name from different locations. Machine learning classification relies on a large number of labeled samples, and the classification criteria are not intuitive enough to meet the verification needs in engineering cost data processing. Therefore, these technologies struggle to simultaneously address differences in list name descriptions, engineering location, and attribute data during engineering list data processing, easily leading to inaccurate list item aggregation and affecting the reliability of subsequent comparative analysis of quantities, unit prices, and cost composition.
[0030] To address all or part of the technical problems in the aforementioned related technologies, this disclosure provides an example embodiment of a method for classifying engineering bill of quantities data based on multi-feature fingerprints. This method can be implemented by an engineering bill of quantities data classification system based on multi-feature fingerprints. Figure 1 The illustration schematically depicts a flowchart of a multi-feature fingerprint-based method for classifying bill of quantities data according to some embodiments of the present disclosure. Reference Figure 1 As shown, the engineering inventory data classification method based on multi-feature fingerprints may include the following steps: Step S110: Obtain the list of projects to be classified and the list classification knowledge base; wherein, the list of projects to be classified includes the list name, source data of project parts and project attribute data, and the list classification knowledge base includes multiple standard list items and list description mapping data associated with each standard list item; Step S120: Using the list representation mapping data, perform name normalization processing on the project list names to obtain the standard list names corresponding to the project list data to be classified. Step S130: Perform part identification processing on the source data of engineering parts to obtain the engineering part information corresponding to the engineering list data to be classified, and perform attribute standardization processing on the engineering attribute data to obtain the list attribute information corresponding to the engineering list data to be classified. Step S140: Perform fingerprint fusion processing on the standard list name, project location information and list attribute information to obtain the multi-feature fingerprint corresponding to the project list data to be classified. Step S150: Perform matching processing on the multi-feature fingerprint and multiple standard list items, and determine the target standard list item corresponding to the project list data to be classified from among the multiple standard list items based on the matching processing results.
[0031] According to the engineering list data classification method based on multi-feature fingerprints in this example embodiment, on the one hand, the engineering list names are normalized through list description mapping data, so that engineering list names with different descriptions can be converted into standard list names, thereby reducing scattered classification caused by synonyms with different names; on the other hand, by performing part identification processing on the source data of engineering parts and attribute standardization processing on the engineering attribute data, the engineering list data to be classified can carry both engineering part information and list attribute information, thereby reducing misclassification caused by different parts with the same name or inconsistent attributes; furthermore, by performing fingerprint fusion processing on standard list names, engineering part information, and list attribute information, and matching processing with multiple standard list items based on multi-feature fingerprints, the determination process of target standard list items has multi-dimensional matching basis, thereby improving the reliability of matching results; thus, it has the advantage of improving the efficiency and accuracy of engineering list data classification.
[0032] The following will further explain the engineering bill of quantities data classification method based on multi-feature fingerprints in this example embodiment.
[0033] In step S110, the project list data to be classified and the list classification knowledge base are obtained; wherein, the project list data to be classified includes the project list name, project location source data and project attribute data, and the list classification knowledge base includes multiple standard list items and list description mapping data associated with each standard list item.
[0034] The project list data to be classified represents project list records that require classification. The project list name represents the name field in the project list data describing the project content. The project location source data represents data reflecting the project location to which the project list data belongs. The project attribute data represents data related to units of measurement, quantities, unit prices, or specifications in the project list data to be classified. The list classification knowledge base represents a set of data used to provide reference for classifying project list data. The standard list item represents a standardized list item that serves as the classification target. The list description mapping data represents the correspondence between different list name descriptions and standard list items.
[0035] For example, the bill of quantities data to be classified can be a record in a quotation document for a hydropower project. The bill of quantities name of this record can be "mortar anchor rod φ25L=4.5m". The source data of the project location can include the worksheet name "water retaining project" where the bill of quantities record is located, the file path "a hydropower station / water retaining project / quotation document.xlsx" or the bill of quantities number. The project attribute data can include the unit "piece", the quantity "12500", the unit price "185.50", and the specification "φ25L=4.5m". The bill of quantities classification knowledge base can include the standard bill of quantities item "anchor rod", as well as the bill of quantities description mapping data between "mortar anchor rod", "system anchor rod", "anchor bar" and the standard bill of quantities item "anchor rod".
[0036] Specifically, the project list data to be classified can be read from the project list file, quotation file or contract list file, and a pre-built list classification knowledge base can be invoked to establish a corresponding processing basis between the project list data to be classified and the standardized classification basis.
[0037] In step S120, the project list names are normalized using the list representation mapping data to obtain the standard list names corresponding to the project list data to be classified.
[0038] Among these, "name standardization" refers to the process of converting project list names with different expressions into a unified name format. "Standard list name" refers to the standardized name corresponding to a standard list item.
[0039] For example, when the bill of quantities name is "mortar anchor φ25L=4.5m", the correspondence between "mortar anchor" and the standard bill of quantities item "anchor" can be determined based on the bill of quantities description mapping data, thereby unifying the bill of quantities name to the standard bill of quantities name "anchor". Similarly, when the bill of quantities name is "shotcrete", it can be unified to the standard bill of quantities name "shotcrete" based on the bill of quantities description mapping data.
[0040] Specifically, based on the mapping data of the bill of quantities description, the correspondence between the project bill of quantities names and standard bill of quantities items can be determined, and the project bill of quantities names can be converted into the corresponding standard bill of quantities names. This reduces the likelihood of the same project content being classified into different categories due to different name descriptions.
[0041] In step S130, the source data of the engineering parts is processed for part identification to obtain the engineering part information corresponding to the engineering list data to be classified, and the engineering attribute data is processed for attribute standardization to obtain the list attribute information corresponding to the engineering list data to be classified.
[0042] Specifically, the part identification process refers to determining the engineering part to which the engineering list data to be classified belongs based on the source data of the engineering part. Engineering part information can represent the engineering location or component corresponding to the engineering list data to be classified. Attribute standardization processing refers to converting engineering attribute data into a unified attribute expression format. List attribute information can represent the attribute data obtained after attribute standardization processing.
[0043] For example, if the source data for the engineering part includes the worksheet name "Water Retaining Project", then the engineering part information can be determined as "Water Retaining Project"; if the worksheet name is unclear, but the file path contains "Water Retaining Project", the engineering part information can also be determined based on the file path. For engineering attribute data, if the unit is "strip" and the corresponding list content belongs to the anchor bolt category, then the unit can be unified as "pole"; if the specification is "φ25L=4.5m", then this specification can be included as part of the list attribute information.
[0044] Specifically, the project location information corresponding to the project list data to be classified can be identified from the project location source data, and the attribute content in different forms of expression in the project attribute data can be uniformly processed to obtain the list attribute information. This allows the project list data to have information on name, location, and attributes simultaneously, reducing misjudgments caused by classifying solely based on name.
[0045] In step S140, fingerprint fusion processing is performed on the standard list name, project location information, and list attribute information to obtain the multi-feature fingerprint corresponding to the project list data to be classified.
[0046] Fingerprint fusion processing can be described as combining multiple dimensions of inventory features to form a unified feature representation. Multi-feature fingerprints can represent inventory data features jointly characterized by standard inventory name, project location information, and inventory attribute information.
[0047] For example, for a bill of quantities data to be classified, with the standard bill of quantities name "anchor bolt," the engineering location information "water-retaining project," and the bill of quantities attribute information including the unit "bolt," the engineering quantity level, and the specification "φ25L=4.5m," the above information can be fused to form a corresponding multi-feature fingerprint. Correspondingly, even if another bill of quantities data also has the standard bill of quantities name "anchor bolt," but its engineering location information is "water-draining project," the multi-feature fingerprint corresponding to this bill of quantities data can still be distinguished from the anchor bolt bill of quantities data in the "water-retaining project" category.
[0048] Specifically, the standard list name, project location information, and list attribute information can be used as different feature sources for the project list data to be classified. These feature sources can then be fused to obtain a multi-feature fingerprint. This allows for the formation of a unified matching basis based on the relationships between names, locations, and attributes.
[0049] In step S150, the multi-feature fingerprint is matched with multiple standard list items, and based on the matching results, the target standard list item corresponding to the engineering list data to be classified is determined from among the multiple standard list items.
[0050] The matching process can be described as classifying data based on the correspondence between multi-feature fingerprints and standard list items. The matching result represents the matching status between the multi-feature fingerprints and multiple standard list items. The target standard list item represents the final determined standard list item corresponding to the project list data to be classified.
[0051] For example, for a list of projects to be classified that contains "water-retaining works," "anchor bolts," "roots," and corresponding quantity levels in its multi-feature fingerprint, the multi-feature fingerprint can be matched with multiple standard list items in the list classification knowledge base. If a certain standard list item has a higher degree of correspondence with the multi-feature fingerprint in terms of name, location, and attributes, then that standard list item can be identified as the target standard list item.
[0052] Specifically, multi-feature fingerprints can be used as the classification basis for the bill of quantities data to be categorized, and matched with multiple standard bill of quantities items in the bill of quantities classification knowledge base. The target standard bill of quantities item is then determined based on the matching results. This allows for the classification of bill of quantities data under the combined constraints of name, location, and attributes, improving the accuracy and consistency of the classification results.
[0053] The following is a detailed explanation of the steps described above.
[0054] In related technologies, keyword matching, tabular function matching, or ordinary text similarity matching are commonly used to process engineering bill of quantities names. These methods primarily rely on the literal similarity of the bill of quantities names themselves, making it difficult to identify differences in bill of quantities expressions from different sources, different specification versions, or different engineering practices. For example, expressions such as mortar anchor bolt, system anchor bolt, anchor bar, and rock bolt (anchor bolt) may all correspond to anchor bolts in engineering meaning, but related technologies easily identify these expressions as different bill of quantities items, leading to the fragmented aggregation of similar bill of quantities data. Furthermore, relying solely on manually maintaining a simple thesaurus, while capable of handling some known synonyms, makes it difficult to distinguish the source relationships between standard names, field colloquialisms, regionally differentiated terms, and historical terms, and also makes it difficult to continuously expand and uniformly manage newly added expressions. Therefore, this disclosure proposes a method for constructing a bill of quantities classification knowledge base based on multi-level engineering domain synonyms. For example, the structure of the bill of quantities classification knowledge base can be shown in Table 1 below.
[0055] Table 1 Specifically, the process of building a list-based knowledge base may include the following technical steps: The first step is to obtain the standard data of the engineering bill of quantities, and then extract the bill of quantities number, bill of quantities name, bill of quantities unit and the location to which the bill of quantities belongs from the standard data of the engineering bill of quantities to obtain multiple standard bill of quantities items.
[0056] The standard data in the bill of quantities can represent data derived from national standards, industry specifications, or pre-defined standard bill of quantities systems, used to form the standard terminology layer in Table 1. The bill of quantities number represents the coding information of the standard bill of quantities item within the standard bill of quantities system. The bill of quantities name represents the standardized name corresponding to the standard bill of quantities item. The bill of quantities unit represents the unit of measurement corresponding to the standard bill of quantities item. The bill of quantities location represents the engineering location to which the standard bill of quantities item applies.
[0057] Specifically, standard list contents can be extracted from standard documents such as the "Regulations for the Compilation of Design Estimates for Hydropower Projects" (NB / T 11408-2023) and the "Regulations for the Calculation of Design Quantities for Hydropower Projects" (NB / T 11808-2025). These extracted standard list contents are then organized into standard list items that include the list number, list name, list unit, and the location to which the list belongs. For example, standard list contents with the list name "anchor bolt," list unit "root," and location "water-retaining structure" can be organized into a single standard list item. This forms the core classification object in the list classification knowledge base, ensuring that subsequent list names from different sources have a unified target.
[0058] The second step is to obtain the source data of the bill of quantities description and extract the list name description from the source data of the bill of quantities description to obtain the list description data to be mapped.
[0059] The project list description source data can represent data sources including non-standard list names, customary list names, regionally different list names, or historical list names. List name descriptions can represent the specific names used in engineering practice to describe list content. List description data to be mapped can represent a set of list name descriptions that have not yet been associated with standard list items.
[0060] Specifically, list item names can be extracted from historical project Excel files, on-site survey data, expert experience records, regional standards, cross-regional project documents, foreign-related project documents, old standard documents, or historical project archives. For example, "mortar anchor bolt" can be extracted from historical project Excel files, "rockbolt" from foreign-related project documents, and "anchor bar" from old standard documents, and these list item names can be used as the list item description data to be mapped. This provides a data foundation for subsequently establishing the mapping relationship between different list item names and standard list items.
[0061] The third step is to classify the data of the list to be mapped based on the source information of the data, and obtain at least one of the following: engineering custom data, regional difference data, and historical data.
[0062] The source information can indicate the data origin, usage scenario, or reason for the formation of the list description data to be mapped. Engineering custom description data can represent customary list name descriptions from construction sites, construction units, or historical projects, corresponding to the engineering colloquialism layer in Table 1. Regional difference description data can represent differentiated list name descriptions from different regions, different units, or international projects, corresponding to the dialect difference layer in Table 1. Historical description data can represent historical list name descriptions from old specifications or historical project archives, corresponding to the historical data layer in Table 1.
[0063] Specifically, based on the source information of the list description data to be mapped, list descriptions from field surveys, historical project data, or expert experience can be categorized as engineering convention descriptions; list descriptions from regional standards, cross-regional project data, or international project data can be categorized as regionally different descriptions; and list descriptions from older standard documents or historical project archives can be categorized as historical descriptions. For example, "mortar anchor bolt" can be categorized as engineering convention description data, "rock bolt" as regionally different description data, and "anchor bar" as historical description data. This allows for hierarchical management of list descriptions from different sources according to source type, avoiding mixed storage of description data from different sources.
[0064] The fourth step involves associating at least one of the engineering practice description data, regional difference description data, and historical description data with the corresponding standard list item to obtain list description mapping data associated with each standard list item.
[0065] Among these, the association processing can refer to the process of establishing a correspondence between the list names from different sources and the corresponding standard list items. The list description mapping data can represent the mapping relationship between engineering convention description data, regional difference description data, or historical description data and standard list items.
[0066] Specifically, the standard terminology layer in Table 1 can be used as the core layer, and the engineering colloquialism layer, dialect difference layer, and historical data layer can be used as mapping layers. This allows the list item descriptions in the engineering colloquialism layer, dialect difference layer, and historical data layer to be linked to the standard list item in the standard terminology layer through synonym relationships. For example, the engineering colloquialism data "mortar anchor," the regional difference data "rock bolt," and the historical data "anchor bar" can all be associated with the standard list item "anchor." Similarly, the engineering colloquialism data "normal concrete," the historical data "sprayed concrete," and "plain sprayed concrete" can all be associated with the standard list item "sprayed concrete." This creates a many-to-one mapping relationship where multiple mapping layer descriptions point to the same standard list item, enabling different list item descriptions to be aggregated into a unified standard list item.
[0067] The fifth step is to construct a list classification knowledge base based on multiple standard list items and the list description mapping data associated with each standard list item.
[0068] The list classification knowledge base can represent a data set that stores multiple standard list items and their corresponding list description mapping data according to a preset hierarchical structure. For example, the list classification knowledge base can be stored in JSON format, where JSON can represent JavaScript Object Notation, a lightweight data interchange format.
[0069] For example, the list categorization knowledge base can be stored using the following structure: { "knowledge_base": { Description: "A multi-level engineering field thesaurus" "structure": { "core_layer": "Standard Terminology Layer", "mapping_layers": ["Common Engineering Names Layer", "Dialect Differences Layer", "Historical Data Layer"], "mapping_rule": "Many-to-one mapping to standard terminology level" }, "terms": [ { "standard_term": { "term": "open excavation" "definition": "The official name stipulated by national standards and industry specifications", Source: "Regulations for the Preparation of Design Estimates for Hydropower Projects" (NB / T 11408-2023) and "Regulations for the Calculation of Design Quantities for Hydropower Projects" (NB / T 11808-2025) }, "mappings": { "common_names": [ { "term": "excavation", "source": "on-site surveys, historical project data, expert experience" } ], "regional_variants": [], "historical_variants": [] } }, { "standard_term": { "term": "sprayed concrete", "definition": "The official name stipulated by national standards and industry specifications", Source: "Regulations for the Preparation of Design Estimates for Hydropower Projects" (NB / T 11408-2023) and "Regulations for the Calculation of Design Quantities for Hydropower Projects" (NB / T 11808-2025) }, "mappings": { "common_names": [ { "term": "normal concrete", "source": "on-site surveys, historical project data, expert experience" } ], "regional_variants": [], "historical_variants": [ { "term": "shotcrete", "source": "Historical project archives, old version specification documents" }, { "term": "plain sprayed concrete", "source": "Historical project archives, old version specification documents" } ] } }, { "standard_term": { "term": "anchor bolt", "definition": "The official name stipulated by national standards and industry specifications", Source: "Regulations for the Preparation of Design Estimates for Hydropower Projects" (NB / T 11408-2023) and "Regulations for the Calculation of Design Quantities for Hydropower Projects" (NB / T 11808-2025) }, "mappings": { "common_names": [ { "term": "mortar anchor bolt", "source": "on-site surveys, historical project data, expert experience" } ], "regional_variants": [ { "term": "rock bolt", "source": "Regional regulations, cross-regional project data", Note: Commonly used in international projects } ], "historical_variants": [ { "term": "anchor bar", "source": "Historical project archives, old version specification documents" } ] } } ] } } Specifically, in the above JSON structure, `knowledge_base` is used to record the overall data of the list classification knowledge base; `description` is used to record the knowledge base description; `structure` is used to record the hierarchical structure and mapping rules of the knowledge base; `core_layer` is used to record the standard terminology layer in Table 1; `mapping_layers` is used to record the engineering colloquialism layer, dialect difference layer, and historical data layer in Table 1; `mapping_rule` is used to record the rules for many-to-one mapping to the standard terminology layer; and `terms` is used to record multiple standard list items and their corresponding list description mapping data.
[0070] Furthermore, each term's data unit can include a standard_term and mappings. The standard_term corresponds to the standard terminology layer in Table 1, serving as the core layer and classification target in the list classification knowledge base, used to record the name, definition, and data source of standard list items. Mappings corresponds to the mapping layer in Table 1, used to record list name descriptions from different sources corresponding to the standard_term. Specifically, common_names corresponds to the engineering colloquialism layer, used to record engineering customary descriptions; regional_variants corresponds to the dialectal variation layer, used to record regional variation descriptions; and historical_variants corresponds to the historical data layer, used to record historical descriptions. Each mapping item can record a term and a source to retain the specific list name description and its data source; in some embodiments, supplementary explanatory information can also be recorded via notes.
[0071] For example, in the data unit corresponding to the standard list item "anchor bolt", the standard_term records the standard name "anchor bolt", the common_names in mappings records "mortar anchor bolt", the regional_variants records "rock bolt", and the historical_variants records "anchor bar", so that "mortar anchor bolt", "rock bolt", and "anchor bar" all point to "anchor bolt" in the standard_term through mappings. In the data unit corresponding to the standard list item "shotcrete", the common_names records "normal concrete", and the historical_variants records "shotcrete" and "plain shotcrete", so that the above descriptions all point to the standard list item "shotcrete".
[0072] Furthermore, in other embodiments of this disclosure, the list classification knowledge base can be updated, and index data can be constructed based on the updated list classification knowledge base. The update process includes: obtaining manual review results; determining the list descriptions to be added based on the manual review results; calculating the description similarity between the description to be added and existing list descriptions in the list classification knowledge base; if the description similarity meets the inclusion criteria, associating the description to be added with the corresponding standard list item to obtain updated list description mapping data; updating the list classification knowledge base based on the updated list description mapping data, and updating the corresponding matching weights and index data based on the updated list description mapping data.
[0073] The manual review result refers to the result obtained by users or administrators after confirming the correspondence between the data of the project list to be classified and the standard list items. The list description to be added refers to the list name description that has been confirmed by the manual review result to be mapped to a standard list item but has not yet been recorded in the list classification knowledge base. Description similarity refers to the degree of similarity between the list description to be added and existing list descriptions. Inclusion conditions refer to the conditions used to determine whether the list description to be added can be added to the list classification knowledge base, such as the similarity not reaching the duplication threshold and confirmation by the administrator.
[0074] Specifically, during manual review, if a user confirms that an original list name should be categorized into a standard list item, that original list name can be used as a list description to be added. Further, the similarity between the list description to be added and existing list descriptions in the list classification knowledge base can be calculated to determine whether the list description to be added already exists or is highly repetitive with existing descriptions. If the list description to be added does not duplicate existing list descriptions and has been confirmed by the administrator, it can be added to the mapping layer under the corresponding standard list item, and its data source can be recorded. For example, if manual review confirms that a system anchor in a project should be categorized into a standard list item anchor, the system anchor can be added as new engineering convention description data to the list description mapping data corresponding to the anchor. This allows the list classification knowledge base to be continuously updated with the results of manual review, reducing duplicate reviews of similar list descriptions in subsequent classifications.
[0075] Furthermore, after updating the bill of quantities classification knowledge base, the matching weights can be updated based on the usage frequency of the bill of quantities description mapping data. Specifically, if a bill of quantities description is used multiple times in engineering bill of quantities classification and has been manually reviewed and confirmed, its matching priority in the name matching process can be increased; if a bill of quantities description has not been used for a long time or has been found to be mismatched after review, its matching priority can be decreased. Thus, the bill of quantities description mapping data in the bill of quantities classification knowledge base can not only be expanded but also dynamically adjusted according to actual usage.
[0076] Furthermore, to support efficient retrieval of the list classification knowledge base, multi-level index data can be constructed based on the list classification knowledge base. Multi-level index data can include at least one of the following: list description index, engineering part index, and rule matching index.
[0077] The list description index can represent a data structure that uses the list name as the retrieval entry point and points to the corresponding standard list item. For example, the list description index may include the following mapping relationship: Mortar anchor bolt: [code:B1-03,name:anchor bolt,site:water retaining project]; System anchor bolts: [code:B1-03, name: anchor bolt, site: water-retaining project]; rock bolt: [code:B1-03, name: anchor bolt, site: water-retaining project]; Anchor bar: [code:B1-03,name:anchor rod,site:water barrier project].
[0078] Among them, "rock bolt" can be used to describe an anchor bolt.
[0079] An engineering part index can represent a data structure that uses an engineering part as the retrieval entry point and points to multiple standard list items under that engineering part. For example, an engineering part index may include the following mapping relationship: Water-retaining works: [B1-01, B1-02, B1-03]; Drainage works: [B2-01, B2-02, B2-03].
[0080] A rule matching index can represent a data structure that matches list name descriptions based on preset text rules. For example, a rule matching index may include pre-compiled rules for matching list descriptions such as anchor bolts, mortar anchor bolts, or earthwork excavation. Pre-compiled rules can represent pre-built and loaded text matching rules, reducing the processing overhead of repeatedly generating rules for each match.
[0081] Furthermore, when performing name matching based on the list classification knowledge base, the input engineering list name can be preprocessed to obtain the list name to be matched and the list specification information; then, precise matching can be performed based on the list description index; if no matching result is obtained, candidate retrieval can be performed based on keywords; if candidate results still need to be supplemented, rule matching can be performed based on the rule matching index; finally, the candidate standard list items are sorted according to the similarity between the candidate results and the list name to be matched and the degree of matching of the list specification information.
[0082] For example, for the bill of quantities name "mortar anchor bolt φ25L=4.5m", the specification information φ25L=4.5m can be retained during text preprocessing, and the core name "mortar anchor bolt" can be extracted. Subsequently, "mortar anchor bolt" can be searched in the bill of quantities description index to obtain the corresponding standard bill of quantities item "anchor bolt". If the input name does not directly match the bill of quantities description index, the keyword "anchor bolt" can be extracted, and candidate standard bill of quantities items containing "anchor bolt" can be retrieved; supplementary matching can also be performed using rules in the rule matching index. Afterwards, the candidate standard bill of quantities items can be sorted based on name similarity and specification information matching degree, and the top-ranked candidate results can be used for subsequent name normalization processing.
[0083] Through the aforementioned continuous learning mechanism and multi-level indexing mechanism, on the one hand, newly added list descriptions confirmed during manual review can be promptly incorporated into the list classification knowledge base, enabling the list classification knowledge base to expand as engineering project data accumulates; on the other hand, the efficiency of name matching and candidate retrieval can be improved through list description index, engineering part index, and rule matching index; furthermore, the matching priority can be adjusted based on usage frequency to improve the stability and accuracy of subsequent name normalization processing.
[0084] In the embodiments of this disclosure, the names of the engineering list are normalized using the list representation mapping data to obtain the standard list names corresponding to the engineering list data to be classified. This can be done through the following technical steps: The first step is to preprocess the project list names to obtain the list names and specifications to be matched.
[0085] The text preprocessing can refer to the cleaning, splitting, and normalization of the bill of quantities names. The bill of quantities name to be matched can refer to the core name extracted from the bill of quantities names and used for matching with the bill of quantities description mapping data. The bill of quantities specification information can refer to the data in the bill of quantities names used to describe specifications, models, dimensions, or parameters.
[0086] Specifically, the project list name can be obtained, and then subjected to null value checks, leading and trailing spaces removal, special character processing, and format standardization. If the project list name is valid, specifications, model numbers, or dimensions are extracted from it to obtain the list specification information. The portion of the name other than the specification information is then designated as the list name to be matched. For example, when the project list name is "mortar anchor φ25L=4.5m", the specification information "φ25L=4.5m" can be extracted, and "mortar anchor" can be designated as the list name to be matched. This reduces interference from specification content in name matching and retains specification information for subsequent sorting and judgment.
[0087] The second step is to perform list description matching processing in the list description mapping data based on the list name to be matched, and obtain the candidate standard list items corresponding to the list name to be matched.
[0088] The list description matching process can be described as the process of searching for the corresponding standard list item in the list description mapping data based on the list name to be matched. The candidate standard list item can be described as a standard list item that may correspond to the list data of the project to be classified, which is initially determined based on the list name to be matched.
[0089] Specifically, the names of the lists to be matched can be matched with the list name expressions in the list expression mapping data. Matching methods can include at least one of exact matching, inclusion matching, and similarity matching. Exact matching means finding a list name expression that is exactly the same as the list name to be matched; inclusion matching means finding list name expressions that contain or are contained within the list name to be matched, when no exact match result is found; similarity matching means calculating the text similarity between the list name to be matched and each list name expression in the list expression mapping data, and determining candidate standard list items based on the text similarity. For example, when the list name to be matched is "mortar anchor," its corresponding candidate standard list item can be determined as "anchor" based on the list expression mapping data. Thus, list names with different expression forms can be initially grouped into the corresponding standard list item range.
[0090] The third step is to sort the candidate standard list items based on their similarity according to the list name and list specification information, and obtain the sorting results corresponding to the candidate standard list items.
[0091] The similarity ranking process can be described as sorting the candidate standard list items according to their degree of matching with the names in the list to be matched. The ranking result can represent the result obtained by arranging multiple candidate standard list items according to their degree of matching.
[0092] Specifically, the name similarity between the name of the list to be matched and the corresponding list name of each candidate standard list item can be calculated separately. Then, the candidate standard list items are ranked based on the matching degree between the list specification information and the specification information of the candidate standard list items. For example, when the list name to be matched is "mortar anchor bolt" and the list specification information is "φ25L=4.5m", the candidate standard list item whose name corresponds to "anchor bolt" and whose specification information is closer can be selected first. This improves the correspondence between the ranked results and the actual list content when multiple candidate results exist.
[0093] The fourth step is to determine the target candidate standard list item from the candidate standard list items based on the sorting results, and to determine the list name corresponding to the target candidate standard list item as the standard list name.
[0094] The target candidate standard list item can represent the final candidate item determined based on the ranking results among the candidate standard list items. The standard list name can represent the normalized name corresponding to the target candidate standard list item.
[0095] Specifically, the highest-ranked candidate standard list item can be selected from the ranking results as the target candidate standard list item, and the list name corresponding to the target candidate standard list item can be determined as the standard list name. For example, if the project list name is "mortar anchor φ25L=4.5m", after text preprocessing, the list name to be matched "mortar anchor" and the list specification information "φ25L=4.5m" are obtained. If the target candidate standard list item corresponding to "mortar anchor" is determined to be "anchor" in the list description mapping data, then "anchor" can be determined as the standard list name.
[0096] By using the above method, non-core content that affects name matching in the project list name can be removed first, then candidate search can be performed based on the list description mapping data, and the candidate results can be sorted by combining name similarity and specification information, thereby improving the accuracy of the standard list name determination process.
[0097] Furthermore, in this embodiment of the disclosure, the engineering part identification processing is performed on the source data of the engineering parts to obtain the engineering part information corresponding to the engineering list data to be classified. This can be done through the following technical steps: The first step is to perform part name matching processing on the worksheet names in the source data of the engineering parts to obtain the worksheet part identification results and the corresponding first identification confidence scores.
[0098] Here, "Worksheet Name" refers to the name of the spreadsheet page containing the project list data to be categorized. "Part Name Matching" refers to the process of matching the worksheet name with a pre-established keyword matching library for project parts. "Worksheet Part Identification Result" represents the project parts identified based on the worksheet name. "First Identification Confidence" indicates the reliability of the worksheet part identification result.
[0099] Specifically, a keyword matching library for engineering parts can be pre-built. (Reference) Figure 2 As shown, the engineering component keyword matching library allows configuration of component codes, precise keywords, included keywords, matching modes, and matching weights for engineering components such as water-retaining projects, water-spraying projects, water diversion projects, and power generation projects. For example, the component code for a water-retaining project could be B1. Precise keywords could include water-retaining project, water-retaining structure, dam project, etc. Included keywords could include water-retaining, dam, upper reservoir, lower reservoir, dam, etc., and matching modes could include expressions used to identify fields such as dam, reservoir, dam, and reservoir. Water-spraying projects, water diversion projects, and power generation projects can also be configured with corresponding keywords and matching modes in the same way.
[0100] When performing worksheet name matching, worksheet names can be identified in the following order: exact match, containment match, pattern match, and combined match. For example, if the worksheet name is "Water Retaining Project," it can be directly matched using exact keywords, and the corresponding part code is determined to be B1; if the worksheet name is "Upper Reservoir Dam," it can be matched using the keywords "dam" or "upper reservoir"; if the worksheet name includes the field "dam," it can be identified through pattern matching. Therefore, when the worksheet name clearly reflects the project location, the worksheet location identification result can be directly obtained, and the corresponding first identification confidence score can be generated. For example, the worksheet location identification results and corresponding first identification confidence scores obtained after performing location name matching processing on the worksheet names in the project location source data are shown in Table 2 below.
[0101] Table 2 The second step involves parsing the file path information in the source data of the engineering parts when the worksheet part identification result is empty or the first identification confidence level is lower than the part identification confidence level threshold, and then performing path parsing and part name matching to obtain the path part identification result and the corresponding second identification confidence level.
[0102] The part identification confidence threshold represents the threshold used to determine whether the part identification result meets reliability requirements. File path information represents the storage path of the file containing the project list data to be classified. Path parsing represents the process of splitting the file path information into multiple path-level fields. Path part identification result represents the project part identified based on the file path information. Second identification confidence represents the reliability of the path part identification result.
[0103] Specifically, when the worksheet name fails to identify the project location, or the initial identification confidence is low, file path information can be parsed. For example, for the file path D:\Projects\a hydropower station\water-retaining project\price limit.xlsx, it can be split into multiple path level fields according to directory separators, and a reverse scan can be performed starting from the directory closest to the file. Since path levels closer to the file generally better reflect the location of the file, different depth weights can be assigned to different path levels, and a matching score can be calculated by combining keyword length or keyword matching method.
[0104] For example, if the path-level field "water-retaining project" contains the keyword "water-retaining," then a water-retaining project can be matched; if the path-level field "dam" contains the English keyword "dam," then a water-retaining project can also be matched; if the path-level field "spillway or stilling basin" contains a feature word corresponding to "water-discharging project," then a water-discharging project can be matched. Furthermore, a path matching score can be calculated based on the path level depth, the matched keywords, and the corresponding matching weight, and a second identification confidence level can be determined based on the path matching score. Thus, when the worksheet name is not standardized or the worksheet name lacks part information, the project part can be identified by supplementing the identification through file path information. For example, when performing path parsing and part name matching processing on file path information, the correspondence between the path matching score and the second identification confidence level can be shown in Table 3 below.
[0105] Table 3 The third step involves performing numbering rule matching on the numbering information in the source data of the engineering parts when the path part identification result is empty or the second identification confidence level is lower than the part identification confidence level threshold, in order to obtain the numbered part identification result and the corresponding third identification confidence level.
[0106] The numbering information can represent coded data reflecting the attribution of a project location, such as the list number, project number, or serial number in the project list data to be classified. Numbering rule matching processing can represent the process of identifying project locations based on number prefixes, number formats, or number value ranges. The numbered location identification result can represent the project location identified based on the numbering information. The third identification confidence level can represent the reliability of the numbered location identification result.
[0107] Specifically, a rule library for numbering engineering parts can be pre-built. (Reference) Figure 3 As shown, the engineering component numbering rule base allows configuration of component codes, numbering matching patterns, and numbering ranges for different engineering components. For example, the component code for a water-retaining project can be B1, and its numbering matching pattern can include starting with 1., A1 or a1, DS-, or DAM-, with a numbering range of 10000 to 19999; the component code for a water-discharge project can be B2, and its numbering matching pattern can include starting with 2., B1 or b1, SP-, or SLW-, with a numbering range of 20000 to 29999; the component code for a water diversion project can be B3, and its numbering matching pattern can include starting with 3., C1 or c1, IN-, or INT-, with a numbering range of 30000 to 39999.
[0108] When performing numbering rule matching, the numbering column in the bill of quantities data to be classified can be identified first, and multiple valid numbers can be extracted from the numbering column as analysis samples. Then, each valid number is matched based on the numbering matching pattern and numbering range, and the number of matches or matching scores corresponding to each project part are counted. Finally, the identification result of the numbered part and the third identification confidence level are determined based on the number of matches or matching scores. For example, if most valid numbers fall within the range of 10000 to 19999, or if most valid numbers meet the numbering rule starting with 1., the identification result of the numbered part can be determined as a water-retaining project. Thus, when the worksheet name and file path cannot reliably identify the project part, the numbering pattern within the bill of quantities data can be used for supplementary identification. For example, the identification result of the numbered part and the corresponding third identification confidence level obtained after numbering rule matching of the numbering information can be shown in Table 4 below.
[0109] Table 4 The fourth step involves performing part feature statistical processing on the list content in the engineering part source data when the part identification result is empty or the third identification confidence level is lower than the part identification confidence level threshold, in order to obtain the content part identification result and the corresponding fourth identification confidence level.
[0110] The list content can refer to the list name, project description, quota information, or other text content in the list data to be classified. Location feature statistical processing refers to the statistical and scoring of the list content based on the content characteristics corresponding to different project locations. Content location identification results represent the project locations identified based on the list content. The fourth identification confidence level represents the reliability of the content location identification results.
[0111] Specifically, a feature library of engineering components can be pre-built. (Reference) Figure 4 As shown, the engineering component content feature library allows for the configuration of keywords, quota ranges, and typical list items for different engineering components. For example, water-retaining engineering can be configured with keywords such as dam body, dam, seepage prevention, core wall, rockfill, panel, toe slab, grouting, and curtain wall, along with corresponding quota ranges and typical list items; water-discharge engineering can be configured with keywords such as spillway, flood discharge, stilling basin, jet dam, gate, hoist, spillway, and energy dissipation; water diversion engineering can be configured with keywords such as tunnel, intake, surge tank, pressure steel pipe, gate well, and water diversion channel; and power generation engineering can be configured with keywords such as powerhouse, unit, turbine, generator, crane beam, spiral casing, and tailrace pipe.
[0112] When performing statistical processing of component features, text fields can be extracted from the list content, and the frequency or matching weight of keywords for each engineering component in the list content can be counted. Simultaneously, a comprehensive score can be calculated for each engineering component by combining the quota range and typical list items. For example, if terms such as spillway, flood discharge, and stilling basin appear multiple times in the list content, and the number falls within the quota range corresponding to the spillway project, the comprehensive score for the spillway project can be increased, and the content component identification result can be determined as a spillway project. Therefore, when the first three sources cannot reliably identify the engineering component, the component can be inferred from the component features of the list content itself. For example, the content component identification results and corresponding fourth identification confidence scores obtained after performing statistical processing of the list content are shown in Table 5 below.
[0113] Table 5 The fifth step is to determine the project location information corresponding to the project list data to be classified based on at least one of the worksheet location identification results, path location identification results, number location identification results, and content location identification results, as well as the corresponding identification confidence level.
[0114] The engineering part information may include at least one of the following: part name, part code, part source, full path, key parent, and extraction confidence score. The part source indicates which data source—worksheet name, file path, number, or list content—identified the engineering part information. The extraction confidence score indicates the degree of credibility of the finally determined engineering part information.
[0115] Specifically, the engineering part information can be determined according to the progressive order of worksheet name, file path information, number information, and list content. If any identification result is not empty and the corresponding identification confidence level is not lower than the part identification confidence level threshold, the identification result can be identified as the engineering part information. If multiple valid identification results exist, the identification confidence levels of each result can be compared, and the identification result with the highest confidence level can be selected as the engineering part information, or multiple identification results can be fused to determine the engineering part information. For example, if the worksheet part identification result is empty, the path part identification result is a water-retaining project, and the second identification confidence level is 0.9, then the engineering part information can be identified as a water-retaining project, and the corresponding part code B1 and the identification source as file path information can be recorded.
[0116] The above method enables a four-level progressive process for identifying engineering parts: direct identification via worksheet name, supplementary identification via file path, supplementary identification via numbering rules, and reverse identification via list content. This allows for the determination of relatively reliable engineering part information from various sources, even when worksheet names are not standardized, file paths are missing, or numbering information is incomplete. This provides a part-dimensional foundation for subsequent classification of engineering list data.
[0117] In some embodiments, the engineering attribute data is subjected to attribute standardization processing to obtain the list attribute information corresponding to the list data of the project to be classified. This can be done through the following technical steps: The first step is to perform unit mapping processing on the unit data in the project attribute data to obtain standard unit data.
[0118] Here, unit data can represent the data used to characterize the units of measurement in the project list data to be classified. Unit mapping processing can represent the process of converting units in different forms of expression into a unified unit expression form. Standard unit data can represent the unified units of measurement obtained after unit mapping processing.
[0119] Specifically, the unit field can be found in the project attribute data. The unit field can include at least one of the following: unit, unit of measurement, unit, u, etc. If no unit field is found, the unit data can be set to null. Furthermore, unit data can be mapped based on preset unit mapping rules. For example, meters can be mapped to m, kilometers to km, square meters to m², cubic meters to m³, tons to t, and kilograms to kg. For counting units, mapping can be combined with the project list name or standard list name. For example, in an anchor bolt list, a strip can be mapped to a root. This reduces attribute differences caused by different representations of the same unit.
[0120] The second step is to extract the numerical values and classify the quantities in the project attribute data to obtain the project quantity data.
[0121] Here, quantity data refers to the data used to characterize the quantity of work in the bill of quantities data to be classified. Numerical extraction refers to the process of extracting calculable values from the quantity data. Quantity classification refers to the process of determining the quantity level of work based on its numerical value. Quantity level data refers to the classification results corresponding to the quantity data.
[0122] Specifically, the quantity field can be searched for in the project attribute data. The quantity field can include at least one of the following: quantity, project quantity, project number, quantity, qty, etc. If the quantity field is not found, or the content of the quantity field cannot be converted into a numerical value, the quantity data can be set to null. Further, thousands separators, spaces, and unit appended characters in the quantity data can be cleaned up, and the numerical portion can be extracted. Then, according to a preset quantity class division rule, the quantity data can be divided into the corresponding quantity class data. For example, if the quantity data is 12500, it can be divided into the ten-thousands level according to the class division rule. Thus, the specific numerical value can be converted into a quantity-based attribute that can participate in subsequent attribute matching.
[0123] The third step is to extract the unit price data from the project attribute data and perform hierarchical division to obtain the unit price hierarchical data.
[0124] Here, unit price data refers to the data used to characterize the unit price in the bill of quantities data to be classified. Hierarchical classification processing refers to the process of determining the unit price level to which the unit price belongs based on the numerical value of the unit price data. Unit price level data represents the price classification result corresponding to the unit price data.
[0125] Specifically, the unit price field can be found in the project attribute data. The unit price field can include at least one of the following: unit price, unit_price, price, p, etc. Here, unit_price can represent the unit price, and price can represent the price. If the unit price field is not found, or the content of the unit price field cannot be converted into a numerical value, the unit price data can be set to null. Furthermore, currency symbols, thousands separators, spaces, etc., in the unit price data can be cleaned up, and the numerical portion can be extracted. Then, according to the preset unit price level division rules, the unit price data can be divided into corresponding unit price level data. Thus, the unit price data can be converted into a unified unit price level expression, facilitating comprehensive judgment in conjunction with other attributes.
[0126] The fourth step is to extract the specifications from the project attribute data to obtain standard specification data.
[0127] Specifically, specification data can represent the data used to characterize specifications, dimensions, or technical parameters in the project list data to be classified. Specification extraction processing can represent the process of extracting specification content from project attribute data or name processing results. Standard specification data can represent the specification expression result obtained after specification extraction processing.
[0128] Specifically, specification data can be synchronized first from the specification information already extracted during the name normalization process. If specification data cannot be obtained from the specification information, specification fields can be searched from the project attribute data. Specification fields can include at least one of the following: specification, specification model, model, spec, etc., where spec can represent specification. Furthermore, specification data can be formatted and standardized, such as removing redundant spaces, standardizing length unit expressions, or standardizing parameter connectors. For example, specification data φ25L=4.5m can be recorded as standard specification data. Thus, specification information that distinguishes the contents of the list can be retained.
[0129] The fifth step is to generate the list attribute information corresponding to the list data to be classified based on at least one of the following: standard unit data, engineering quantity data, unit price level data, and standard specification data.
[0130] Among them, the list attribute information can represent the attribute expression result composed of at least one of standard unit data, engineering quantity level data, unit price level data and standard specification data, which is used to characterize the attribute features of the list data of the project to be classified.
[0131] Specifically, standard unit data, quantity data, unit price level data, and standard specification data can be combined according to preset attribute fields to obtain the list attribute information. For example, for the list data to be classified as mortar anchor rod φ25L=4.5m, if the unit data is a strip, the quantity data is 12500, the unit price data is 185.50, and the specification data is φ25L=4.5m, then the unit can be mapped as the root, the quantity can be divided into ten-thousand levels, the unit price can be divided into the corresponding unit price level, and φ25L=4.5m can be used as the standard specification data to generate the corresponding list attribute information.
[0132] By using the above method, different field names, unit expressions, numerical forms, and specification expressions in the project attribute data can be uniformly converted into list attribute information, so that the project list data to be classified has a unified attribute expression basis, thereby reducing the impact of inconsistent attribute expressions on subsequent classification processing.
[0133] Furthermore, fingerprint fusion processing is performed on the standard list name, project location information, and list attribute information to obtain multi-feature fingerprints corresponding to the project list data to be classified. This process includes the following technical steps: The first step is to extract the project location identifier from the project location information and extract the standard unit data and project quantity data from the list attribute information.
[0134] The engineering component identifier can represent the identification information used to distinguish different engineering components in the engineering component information, such as component code, component number, or component name code. For example, the engineering component identifier corresponding to a water-retaining project can be B1, the engineering component identifier corresponding to a water-discharging project can be B2, and the engineering component identifier corresponding to a water-diversion project can be B3. The engineering quantity data can represent the quantity result obtained by dividing the engineering quantity data according to its numerical size, such as thousands, tens of thousands, or other preset quantities.
[0135] Specifically, the part code can be extracted from the engineering part information as the engineering part identifier, and the standard unit data and engineering quantity data can be extracted from the list attribute information. For example, if the engineering part information includes the part name "water-retaining project" and the part code B1, and the list attribute information includes the standard unit data root and the engineering quantity data in the ten-thousands level, then the engineering part identifier B1, the standard unit data root, and the engineering quantity data in the ten-thousands level can be extracted. Thus, the core fields involved in fingerprint generation can be extracted from the engineering part information and the list attribute information.
[0136] The second step involves combining and processing the engineering location identifier, standard list name, standard unit data, and engineering quantity data to obtain the fingerprint basic data.
[0137] Among them, the fingerprint basic data can represent the data obtained by combining multiple core fields according to a preset field order. The combination processing can represent the processing of splicing or arranging the engineering part identifier, standard list name, standard unit data, and engineering quantity data according to a preset connection method.
[0138] Specifically, fingerprint base data can be obtained by combining the following in the order of project location identifier, standard list name, standard unit data, and project quantity data. For example, for project location identifier B1, standard list name "anchor," standard unit data "root," and project quantity data "ten thousand," fingerprint base data B1|anchor|root|ten thousand level can be obtained. For project location identifier B2, standard list name "anchor," standard unit data "root," and project quantity data "thousand level," fingerprint base data B2|anchor|root|thousand level can be obtained. Therefore, even if different list data have the same standard list name, different fingerprint base data can be formed using the project location identifier, standard unit data, and project quantity data.
[0139] The third step is to encode the basic fingerprint data to obtain the fingerprint identifiers corresponding to the project list data to be classified.
[0140] The encoding process can refer to the process of converting fingerprint base data into a fixed-format identifier. The fingerprint identifier can represent identification information used to uniquely identify a combination of features in the project list data to be classified. For example, the encoding process can employ a hashing method, which can represent the process of converting input data into a fixed-length encoded result.
[0141] Specifically, the fingerprint base data can be hashed to obtain the fingerprint identifier corresponding to the project list data to be classified. For example, B1|anchor bolt|root|ten thousand level can be hashed to obtain the fingerprint identifier b1c3d5e7f9g1h3i5; B2|anchor bolt|root|thousand level can be hashed to obtain the fingerprint identifier c2d4e6f8g0h2i4j6. Thus, the combination of multiple dimensions of fields can be converted into fingerprint identifiers that are easy to store, compare, and retrieve.
[0142] The fourth step involves associating the fingerprint identifier, standard list name, project location information, and list attribute information to obtain the multi-feature fingerprint corresponding to the project list data to be classified.
[0143] The association processing can refer to the process of establishing a correspondence between the fingerprint identifier and the name information, location information, and attribute information on which the fingerprint identifier was generated. A multi-feature fingerprint can represent a data structure that uses the fingerprint identifier as an index and associates it with the standard list name, engineering location information, and list attribute information.
[0144] Specifically, the fingerprint identifier can be used as the primary identifier, and the standard list name, project location information, and list attribute information can be encapsulated as associated fields to obtain the multi-feature fingerprint corresponding to the project list data to be classified. For example, the fused multi-feature fingerprint can be represented using the following structure: { "fingerprint_id": "b1c3d5e7f9g1h3i5", "site_fingerprint": { "site_name": "Water Retaining Project", "site_code": "B1", "full_path": "a certain hydropower station > water-retaining project > dam > slope", "key_parents": ["water barrier project", "dam"], "confidence": 1.0 }, "text_fingerprint": { "original_name": "Mortar Anchor Rod φ25 L=4.5m", "standardized_name": "anchor bolt", "synonym_matches": ["anchor bolts"], "key_terms": ["anchor bolt"], "text_hash": "a1b2c3d4e5f6g7h8" }, "attribute_fingerprint": { "unit": "root", "unit_normalized": "root", "quantity": 12500, "quantity_level": "ten thousand level", "price": 185.50, "price_level": "Mid-to-low price", "specification": "φ25 L=4.5m", "quota_code": "10151" } } In the above structure, fingerprint_id can represent a fingerprint identifier; site_fingerprint can represent a structured field corresponding to the project location information; text_fingerprint can represent a structured field corresponding to the standard list name and its source name; and attribute_fingerprint can represent a structured field corresponding to the list attribute information. For example, when the project list data to be classified is a mortar anchor bolt φ25L=4.5m in a water-retaining project, the project location identifier B1, the standard list name anchor bolt, the standard unit data root, and the project quantity data (in ten thousand levels) can be combined and encoded to obtain a fingerprint identifier. This fingerprint identifier is then associated with the project location information, name information, and list attribute information to form a corresponding multi-feature fingerprint.
[0145] The above method can integrate the standard list name, project location information, and list attribute information into a unified multi-feature fingerprint, so that each project list data to be classified has a feature expression that can simultaneously reflect the name, location, and attributes, thereby providing a stable data foundation for subsequent matching processing.
[0146] In some embodiments, the matching process between multi-feature fingerprints and multiple standard list items includes the following technical steps: The first step is to obtain the engineering part identifier, standard list name, standard unit data, and engineering quantity data associated with the fingerprint identifier based on the fingerprint identifier in the multi-feature fingerprint.
[0147] Among them, the data associated with the fingerprint identifier can represent the feature fields that establish a corresponding relationship with the fingerprint identifier during the fingerprint fusion process. The engineering part identifier, standard list name, standard unit data, and engineering quantity data can serve as the basis for subsequent part filtering, name matching, and attribute verification.
[0148] Specifically, after generating a multi-feature fingerprint, the fingerprint identifier can be used as an index to read the engineering location identifier, standard list name, standard unit data, and engineering quantity data from the multi-feature fingerprint. For example, if the fingerprint identifier is obtained by encoding B1|anchor bolt|root|ten thousand, then the engineering location identifier B1, the standard list name anchor bolt, the standard unit data root, and the engineering quantity data ten thousand can be obtained based on this fingerprint identifier. Thus, the key fields that have already been fused in the multi-feature fingerprint can be extracted again, providing input data for subsequent hierarchical matching.
[0149] The second step is to perform part selection processing on multiple standard list items based on the engineering part identification to obtain candidate standard list items.
[0150] The part selection process can be described as the process of selecting standard list items related to the engineering part from multiple standard list items based on the engineering part identifier. Candidate standard list items can be described as the standard list items retained after part selection.
[0151] Specifically, the project location identifier can be compared with the location information corresponding to multiple standard list items, and the standard list items corresponding to the project location identifier can be retained first. For example, if the project location identifier is B1, and B1 corresponds to a water-retaining project, then the standard list items under the water-retaining project can be selected as candidate standard list items. If the location information of a standard list item is completely consistent with the project location identifier, the standard list item can be directly added to the candidate set; if the location information of a standard list item has a hierarchical inclusion relationship with the project location identifier or belongs to the same location category, the standard list item can also be used as a candidate standard list item. Thus, the matching range of standard list items can be narrowed down from the location dimension, reducing mutual interference between list items with the same name in different project locations.
[0152] The third step is to perform name matching processing on the candidate standard list items based on the standard list name to obtain the name matching result.
[0153] Name matching processing refers to determining whether candidate standard list items match in terms of name based on the standard list name. The name matching result indicates the name correspondence between candidate standard list items and standard list names.
[0154] Specifically, the standard list name can be matched with the list names corresponding to the candidate standard list items. Matching methods can include at least one of full match, synonym match, inclusion relationship match, and character similarity match. For example, if the standard list name and the list name of the candidate standard list item are exactly the same, they can be determined to be a full match in the name dimension; if the standard list name and the list name of the candidate standard list item have a synonym relationship, the name matching result can be determined based on list description mapping data or a synonym inverted index; if the standard list name only has an inclusion relationship with the list name of the candidate standard list item, name matching can be performed by combining keywords; if a direct match cannot be achieved through the above methods, the character similarity between the two can be calculated to obtain supplementary matching results.
[0155] For example, if the standard list name is "anchor bolt," and the candidate standard list items include names such as anchor bolt, mortar anchor bolt, and anchor pile, then the corresponding name matching results can be obtained through exact matching, synonym matching, or inclusion relationship matching. Therefore, within the candidate range after location filtering, potentially matching standard list items can be further determined based on the name dimension.
[0156] The fourth step is to perform attribute validation on the candidate standard list items corresponding to the name matching results based on the standard unit data and engineering quantity data, so as to obtain the standard list items to be scored.
[0157] The attribute validation process involves determining the consistency of attributes for candidate standard list items corresponding to name matching results based on standard unit data and engineering quantity data. The standard list items to be scored represent the standard list items retained after location filtering, name matching, and attribute validation.
[0158] Specifically, the standard unit data can be compared with the unit data corresponding to the candidate standard list items, and the quantity data can be compared with the quantity data corresponding to the candidate standard list items. If the units are completely identical or convertible, the unit dimension is considered to meet the attribute verification requirements; if the quantity data are completely identical or belong to adjacent quantities, the quantity dimension is considered to meet the attribute verification requirements. For example, if the standard unit data is the root, and the unit corresponding to the candidate standard list item is also the root, then the unit verification passes; if the quantity data is in the tens of thousands, and the quantity corresponding to the candidate standard list item is also in the tens of thousands or adjacent to the tens of thousands, then the quantity verification passes.
[0159] Furthermore, if a candidate standard list item matches in name dimension, but its standard unit data or engineering quantity data differs significantly from the engineering list data to be classified, then the candidate standard list item can be removed from subsequent processing objects, or its priority as a standard list item to be scored can be reduced. Thus, based on name matching, further validation can be performed using attribute dimensions to reduce false matches caused solely by identical names, and to form a set of standard list items to be scored for subsequent scoring processing.
[0160] Furthermore, based on the matching results, the target standard list item corresponding to the project list data to be classified can be determined from multiple standard list items through the following technical steps: The first step is to calculate the matching degree of the part, the matching degree of the name, and the matching degree of the attribute of the standard list item to be scored, based on the engineering part identifier, the standard list name, the standard unit data, and the engineering quantity data.
[0161] Among them, the part matching degree can represent the degree of matching between the part identifier corresponding to the project list data to be classified and the part corresponding to the standard list item to be scored. The name matching degree can represent the degree of matching between the standard list name and the list name corresponding to the standard list item to be scored. The attribute matching degree can represent the degree of matching between the standard unit data, the project quantity data and the corresponding attributes of the standard list item to be scored.
[0162] Specifically, the matching degree of a component can be determined based on the relationship between the component identifier and the corresponding component in the list of criteria to be scored. For example, if the component names are exactly the same, it can be determined as a complete match; if the source component contains the target component or the target component contains the source component, it can be determined as a hierarchical containment; if there is a keyword correspondence between the two components, it can be determined as a keyword match; if the two components belong to the same component category, it can be determined as a component family match; if the two components have no obvious relationship, it can be determined as a mismatch. For example, the calculation rules for the matching degree of a component can be shown in Table 6 below.
[0163] Table 6 Furthermore, the name matching degree can be determined based on the relationship between the standard list name and the corresponding list name of the standard list item to be scored. For example, if the two names are exactly the same, it can be determined as a complete match; if the two have a synonym relationship in the data mapped through the list description, it can be determined as a synonym match; if the two have a name inclusion relationship, it can be determined as an inclusion relationship match; if the two do not match the above relationships but have a certain degree of text similarity, it can be determined as a character similarity match. For example, the calculation rules for name matching degree can be shown in Table 7 below.
[0164] Table 7 Furthermore, the attribute matching degree can be determined based on standard unit data and quantity data. For example, if the standard unit data and the corresponding unit of the item in the standard list to be scored are exactly the same, the unit dimension can be determined to have a high degree of matching; if the two units are convertible, the unit dimension can be determined to have a medium degree of matching; if the two units are different and not convertible, the attribute matching degree can be reduced. For quantity data, if the quantities are exactly the same, the quantity dimension can be determined to have a high degree of matching; if the quantities are adjacent, the quantity dimension can be determined to have a medium degree of matching; if the quantities differ significantly, the attribute matching degree can be reduced. If the list attribute information also includes quota data, the similarity of the quota data can also be used to assist in determining the attribute matching degree. For example, the calculation rules for attribute matching degree can be shown in Table 8 below.
[0165] Table 8 The second step is to weight the matching degree of the part, the matching degree of the name, and the matching degree of the attribute to obtain the matching score corresponding to the item in the list of criteria to be scored.
[0166] Weighted processing refers to the comprehensive calculation of the matching degree corresponding to multiple matching dimensions according to preset weights. The matching score represents the comprehensive matching degree between the items in the list of criteria to be scored and the data in the list of projects to be classified.
[0167] Specifically, part matching, name matching, and attribute matching can be used as different matching dimensions, and corresponding weights can be assigned to each dimension. For example, part matching can be given a higher weight to distinguish between items with the same name in different engineering parts; name matching can be given a medium weight to identify items with the same name but different meanings; and attribute matching can be given a medium weight to help determine whether attributes such as units and quantities are consistent. For example, a comprehensive processing method can be used, with part matching accounting for 40%, name matching for 30%, and attribute matching for 30%, to obtain the matching score corresponding to the item in the list to be scored.
[0168] For example, when the list of projects to be classified consists of anchor bolts in a water-retaining project, if a certain item in the list of standards to be scored is an anchor bolt in a water-retaining project, and the unit and magnitude match, then this item in the list of standards to be scored has a high degree of matching in the dimensions of location, name, and attribute, and its matching score is high. If another item in the list of standards to be scored is an anchor bolt in a water-drainage project, although the dimensions of name and attribute may match, the dimension of location does not match or the degree of matching is low, so its matching score is lower than that of the standard list item corresponding to the anchor bolt in the water-retaining project.
[0169] The third step is to generate matching processing results based on the matching scores, and then determine the target standard list items in the list of standards to be scored based on the matching processing results.
[0170] The matching result can represent the sorting, grading, or filtering results obtained based on the matching scores corresponding to each item in the list of criteria to be scored. The target criterion list item can represent the criterion list item that is finally determined in the list of criteria to be scored and corresponds to the list of engineering items to be classified.
[0171] Specifically, multiple criteria items to be scored can be sorted according to their matching scores, and the matching result can be determined based on the sorting results. For example, the criterion item with the highest matching score can be determined as the target criterion item; alternatively, if the matching score reaches a preset high confidence threshold, the corresponding criterion item can be directly determined as the target criterion item; if the matching score is within a preset review range, the corresponding criterion item can be used as the review result; and if the matching score is below a preset matching threshold, it can be determined that there is no reliable matching result.
[0172] By using the above method, the matching results of the three dimensions of location, name, and attribute can be converted into comparable matching scores. Target standard list items can then be determined based on these matching scores, thereby reducing misclassification caused by matching only by name and improving the reliability of the engineering list data classification results.
[0173] Furthermore, the confidence level of the matching results can be graded based on the matching score to obtain the corresponding confidence level; and the processing method corresponding to the matching results can be determined based on the confidence level. For example, the confidence level can include at least one of automatic matching, recommended matching, and manual review. Automatic matching can mean that when the matching score reaches a high confidence range, the corresponding item in the list of criteria to be scored is directly identified as the target criterion list item; recommended matching can mean that when the matching score is in a medium confidence range, the corresponding item in the list of criteria to be scored is output as a recommended result for manual confirmation; manual review can mean that when the matching score is below a preset matching range, a list of data awaiting manual judgment is output.
[0174] For example, if the matching score is greater than or equal to 80, the confidence level can be set to automatic matching, and the item in the list of criteria to be scored with the highest matching score can be set as the target criterion list item; if the matching score is between 60 and 79, the confidence level can be set to recommended matching, and the item in the list of criteria to be scored corresponding to the matching score can be output as the recommended result; if the matching score is less than 60, the confidence level can be set to manual review, and the list of projects to be classified can be output to the manual processing flow.
[0175] Furthermore, after obtaining the target standard list items, the original list data has been categorized into the corresponding engineering parts and standard list items, forming a standardized data structure. However, the categorized list data may still be distributed across different engineering parts, different list items, or different engineering files, making it difficult to intuitively reflect the differences in quantity, unit price, and total cost of the same list item across different engineering parts; at the same time, if only a single list data item is viewed, it is difficult to promptly detect anomalies in unit price, quantity, or total cost. Therefore, in this embodiment of the disclosure, the categorized engineering list data can also be subjected to cross-part multi-dimensional comparative analysis, and corresponding cross-part comparison results can be generated to improve the efficiency of engineering list data review and analysis. Specifically, this can be achieved through the following technical steps: The first step is to collect and process the project list data to be classified based on the target standard list items and project location information to obtain the classified list data.
[0176] In this context, "aggregation processing" refers to the process of organizing and summarizing the project list data to be classified according to the target standard list items and project location information. Classified list data refers to list data that has already established a correspondence with the target standard list items and project location information.
[0177] Specifically, the project list data to be categorized can be uniformly classified according to the target standard list items, while retaining the corresponding project location information, project attribute data, and source project document information. For example, if multiple power plant documents contain list data with anchor bolts as the target standard list item, this list data can be aggregated according to its corresponding water-retaining works, water-discharging works, or water-diversion works, etc., to obtain categorized list data. This allows list data originally scattered across different project documents and worksheets to form a unified data foundation.
[0178] The second step is to group the classified list data corresponding to different engineering parts according to the target standard list items to obtain cross-part list data groups.
[0179] Grouping can refer to the process of grouping classified list data according to target standard list items and engineering location information. Cross-location list data group can refer to a set of list data belonging to the same target standard list item but corresponding to different engineering location information.
[0180] Specifically, the target standard list item can be used as the main index to group list data belonging to the same target standard list item into the same data group, and further divided according to engineering location information. For example, the list data for earthwork cut-out, rock cut-out, anchor bolts, and concrete corresponding to water-retaining works, water-spraying works, and water diversion works in different power stations can be grouped separately to obtain cross-location list data groups. This provides a data organization format for horizontal comparison of the same list item across different engineering locations.
[0181] The third step is to perform statistical processing on the engineering attribute data in the cross-part list data group to obtain the comparative index data corresponding to the target standard list items under different engineering part information.
[0182] Statistical processing can refer to the summarization, calculation, or comparison of numerical attributes in cross-part list data groups. Comparative indicator data can represent indicator data used to characterize differences between different parts of the project. For example, comparative indicator data may include at least one of the following: quantity, unit price, total price, percentage, difference value, mean, or ranking result.
[0183] Specifically, for engineering attribute data in cross-part bill of quantities data groups, separate statistics can be performed on quantities, unit prices, total costs, and percentages. For example, in quantity comparison, the quantities of the same target standard bill of quantities item can be calculated across different engineering parts or different power plants; in unit price comparison, the unit price of the same target standard bill of quantities item can be calculated across different engineering parts or different power plants; and in total cost comparison, the corresponding total price or total cost can be determined based on the quantities and unit prices. (Reference) Figure 5 As shown, it can generate a comparison of bill of quantities quantities, which can be used to display the quantity data of each item in the bill of quantities for different engineering parts in different power plants; for reference Figure 6 As shown, it can generate unit price comparison results for the bill of quantities, which can be used to display the unit price differences of the same bill of quantities item in different power plants; Reference Figure 7 As shown, it can generate a comparison of bill of quantities costs, which can be used to show the differences in cost of the same bill of quantities items in different power plants.
[0184] The fourth step is to generate cross-regional comparison results based on the comparison indicator data. These cross-regional comparison results can represent the results of a list-based comparative analysis generated from the comparison indicator data. The output format of the cross-regional comparison results can include at least one of the following: tables, bar charts, line charts, pie charts, or percentage statistics.
[0185] Specifically, cross-part comparison results can be generated based on different comparison dimensions. For example, quantity comparison can output the quantity and percentage, which can be displayed in tables or bar charts; unit price comparison can output the weighted average unit price and unit price difference, which can be displayed in tables or line charts; cost comparison can output the total price and percentage, which can be displayed in tables or pie charts; percentage analysis can output the cost percentage or quantity percentage corresponding to each part of the project. Thus, the categorized bill of quantities data can be converted into comparison results that are easy to review and analyze.
[0186] The fifth step is to review and process the cross-part comparison results, and update the list classification knowledge base based on the review and processing results.
[0187] Specifically, "audit processing" can refer to the process of confirming abnormal data, low-confidence classification data, or classification data pending confirmation in cross-department comparison results. "Audit processing result" can refer to the result obtained after manual or system confirmation of the cross-department comparison results. "Update processing" can refer to the process of supplementing, correcting, or adjusting the weights of the list description mapping data in the list classification knowledge base based on the audit processing result.
[0188] Specifically, abnormal unit prices, abnormal quantities, abnormal costs, or low-confidence classification results in cross-part comparisons can be marked, and the marking results can be output to the review end. If the review process confirms that a certain original list name should be classified into a certain target standard list item, then that original list name can be added to the list classification knowledge base as new list description mapping data; if the review process confirms that a certain mapping relationship is inaccurate, then the corresponding list description mapping data can be corrected or its matching weight reduced. Thus, the review results from cross-part comparison analysis can be fed back to the list classification knowledge base, enabling the list classification knowledge base to be continuously optimized as project data accumulates.
[0189] Furthermore, in the exemplary embodiments of this disclosure, an engineering bill of quantities data classification system based on multi-feature fingerprints is also provided. (Refer to...) Figure 8 As shown, the engineering inventory data classification system 800 based on multi-feature fingerprints includes: a data acquisition module 810, a name normalization module 820, a part attribute module 830, a fingerprint fusion module 840, and a classification matching module 850. Wherein: The data acquisition module 810 can be used to acquire the list of works to be classified and the list classification knowledge base; wherein, the list of works to be classified includes the list name, source data of the work parts and the work attribute data, and the list classification knowledge base includes multiple standard list items and list description mapping data associated with each standard list item; The name normalization module 820 can be used to normalize the names of the engineering list using the list representation mapping data, and obtain the standard list name corresponding to the engineering list data to be classified. The part attribute module 830 can be used to perform part identification processing on the source data of engineering parts, obtain the engineering part information corresponding to the engineering list data to be classified, and perform attribute standardization processing on the engineering attribute data to obtain the list attribute information corresponding to the engineering list data to be classified. The fingerprint fusion module 840 can be used to perform fingerprint fusion processing on the standard list name, engineering part information and list attribute information to obtain the multi-feature fingerprint corresponding to the engineering list data to be classified. The classification and matching module 850 can be used to match multiple feature fingerprints with multiple standard list items, and determine the target standard list item corresponding to the engineering list data to be classified from among the multiple standard list items based on the matching results.
[0190] The specific details of each module in the above engineering bill of quantities data classification system based on multi-feature fingerprints have been described in detail in the corresponding engineering bill of quantities data classification method based on multi-feature fingerprints, so they will not be repeated here.
[0191] It should be noted that although several modules or units of the engineering bill of quantities data classification system based on multi-feature fingerprints have been mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of this disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0192] Furthermore, in an exemplary embodiment of this disclosure, an electronic device capable of implementing the above-described method for classifying engineering bill of quantities data based on multi-feature fingerprints is also provided.
[0193] Those skilled in the art will understand that various aspects of this disclosure can be implemented as a system, method, or program product. Therefore, various aspects of this disclosure can be embodied in the following forms: a completely hardware embodiment, a completely software embodiment (including firmware, microcode, etc.), or an embodiment combining hardware and software aspects, collectively referred to herein as a "circuit," "module," or "system."
[0194] The following reference Figure 9 To describe an electronic device 900 according to such an embodiment of the present disclosure. Figure 9 The electronic device 900 shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments disclosed herein.
[0195] like Figure 9 As shown, the electronic device 900 is presented in the form of a general-purpose computing device. The components of the electronic device 900 may include, but are not limited to: at least one processing unit 910, at least one storage unit 920, a bus 930 connecting different system components (including storage unit 920 and processing unit 910), and a display unit 940.
[0196] The storage unit stores program code that can be executed by the processing unit 910, causing the processing unit 910 to perform the steps described in the "Exemplary Methods" section of this specification according to various exemplary embodiments of this disclosure. The storage unit 920 may include a readable medium in the form of a volatile storage unit, such as a random access memory unit (RAM) 921 and / or a cache memory unit 922, and may further include a read-only memory unit (ROM) 923.
[0197] Storage unit 920 may also include a program / utility 924 having a set (at least one) program module 925, such program module 925 including but not limited to: operating system, one or more application programs, other program modules and program data, each of these examples or some combination thereof may include an implementation of a network environment.
[0198] Bus 930 can represent one or more of several types of bus structures, including a memory cell bus or memory cell controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of the various bus structures.
[0199] Electronic device 900 can also communicate with one or more external devices 970 (e.g., keyboard, pointing device, Bluetooth device, etc.), and with one or more devices that enable a user to interact with electronic device 900, and / or with any device that enables electronic device 900 to communicate with one or more other computing devices (e.g., router, modem, etc.). This communication can be performed via input / output (I / O) interface 950. Furthermore, electronic device 900 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 960. As shown, network adapter 960 communicates with other modules of electronic device 900 via bus 930. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with electronic device 900, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0200] Through the description of the above embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, terminal system, or network device, etc.) to execute the methods according to the embodiments of this disclosure.
[0201] In exemplary embodiments of this disclosure, a computer-readable storage medium is also provided, on which a program product capable of implementing the methods described above is stored. In some possible embodiments, various aspects of this disclosure may also be implemented as a program product including program code that, when the program product is run on a terminal device, causes the terminal device to perform the steps described in the "Exemplary Methods" section of this specification according to various exemplary embodiments of this disclosure.
[0202] The program product for implementing the above-described multi-feature fingerprint-based engineering list data classification method according to embodiments of this disclosure may employ a portable compact disc read-only memory (CD-ROM) and include program code, and may run on a terminal device, such as a personal computer. However, the program product of this disclosure is not limited thereto. In this document, the readable storage medium may be any tangible medium containing or storing a program that may be used by or in conjunction with an instruction execution system, apparatus, or device.
[0203] The program product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or any combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: electrical connections with one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0204] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium, capable of sending, propagating, or transmitting programs for use by or in conjunction with an instruction execution system, apparatus, or device.
[0205] The program code contained on the readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.
[0206] Program code for performing the operations of this disclosure can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java and C++, and conventional procedural programming languages such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0207] Furthermore, the above figures are merely illustrative of the processes included in the method according to exemplary embodiments of this disclosure and are not intended to be limiting. It is readily understood that the processes shown in the above figures do not indicate or limit the temporal order of these processes. Additionally, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.
[0208] From the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, touch terminal, or network device, etc.) to execute the methods according to the embodiments of this disclosure.
[0209] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and embodiments are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the claims.
[0210] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
Claims
1. A method for classifying engineering bill of quantities data based on multi-feature fingerprints, characterized in that, The method includes: Obtain the project list data to be classified and the list classification knowledge base; wherein, the project list data to be classified includes project list name, project location source data and project attribute data, and the list classification knowledge base includes multiple standard list items and list description mapping data associated with each standard list item; Using the list representation mapping data, the names of the project list are normalized to obtain the standard list names corresponding to the project list data to be classified. The source data of the engineering parts is processed for part identification to obtain the engineering part information corresponding to the list data of the engineering to be classified, and the engineering attribute data is processed for attribute standardization to obtain the list attribute information corresponding to the list data of the engineering to be classified. Fingerprint fusion processing is performed on the standard list name, the project location information, and the list attribute information to obtain a multi-feature fingerprint corresponding to the project list data to be classified. The multi-feature fingerprint is matched with the multiple standard list items, and based on the matching results, the target standard list item corresponding to the project list data to be classified is determined from the multiple standard list items.
2. The method for classifying engineering bill of quantities data based on multi-feature fingerprints according to claim 1, characterized in that, Also includes: Based on the target standard list items and the project location information, the project list data to be classified is collected and processed to obtain the classification list data; According to the target standard list items, the classified list data corresponding to different engineering parts information are grouped to obtain cross-part list data groups; Statistical processing is performed on the engineering attribute data in the cross-part list data group to obtain the comparative index data corresponding to the target standard list item under different engineering part information; Generate cross-site comparison results based on the aforementioned comparison index data; The cross-part comparison results are reviewed and processed, and the list classification knowledge base is updated based on the review and processing results.
3. The method for classifying engineering bill of quantities data based on multi-feature fingerprints according to claim 1, characterized in that, The construction process of the list classification knowledge base includes: Obtain standard data for the engineering list, and extract the list number, list name, list unit, and list location from the standard data to obtain multiple standard list items; Obtain the source data of the bill of quantities description, and extract the list name description from the source data of the bill of quantities description to obtain the list description data to be mapped; Based on the source information of the data to be mapped, the data to be mapped is classified and processed to obtain at least one of engineering habitual description data, regional difference description data, and historical description data. At least one of the engineering habit description data, the regional difference description data, and the historical description data is associated with the corresponding standard list item to obtain the list description mapping data associated with each standard list item; The list classification knowledge base is constructed based on the multiple standard list items and the list representation mapping data associated with each of the standard list items.
4. The method for classifying engineering bill of quantities data based on multi-feature fingerprints according to claim 1, characterized in that, The step of using the list representation mapping data to perform name normalization processing on the project list names to obtain the standard list names corresponding to the project list data to be classified includes: The project list name is preprocessed to obtain the list name and list specification information to be matched; Based on the name of the list to be matched, a list description matching process is performed in the list description mapping data to obtain candidate standard list items corresponding to the name of the list to be matched. Based on the name of the list to be matched and the list specification information, the candidate standard list items are sorted by similarity to obtain the sorting result corresponding to the candidate standard list items; Based on the sorting results, a target candidate standard list item is determined from the candidate standard list items, and the list name corresponding to the target candidate standard list item is determined as the standard list name.
5. The method for classifying engineering bill of quantities data based on multi-feature fingerprints according to claim 1, characterized in that, The step of performing part identification processing on the source data of the engineering parts to obtain the engineering part information corresponding to the engineering list data to be classified includes: Perform part name matching processing on the worksheet names in the source data of the engineering parts to obtain the worksheet part identification results and the corresponding first identification confidence level; If the worksheet part identification result is empty or the first identification confidence level is lower than the part identification confidence level threshold, the file path information in the engineering part source data is parsed and the part name is matched to obtain the path part identification result and the corresponding second identification confidence level. If the path part identification result is empty or the second identification confidence level is lower than the part identification confidence level threshold, the numbering information in the source data of the engineering part is subjected to numbering rule matching processing to obtain the numbered part identification result and the corresponding third identification confidence level; If the identification result of the numbered part is empty or the third identification confidence level is lower than the identification confidence threshold of the part, the list content in the source data of the engineering part is subjected to part feature statistical processing to obtain the content part identification result and the corresponding fourth identification confidence level. Based on at least one of the worksheet part identification results, the path part identification results, the number part identification results, and the content part identification results, and the corresponding identification confidence level, the project part information corresponding to the project list data to be classified is determined.
6. The method for classifying engineering bill of quantities data based on multi-feature fingerprints according to claim 1, characterized in that, The step of performing attribute standardization processing on the project attribute data to obtain the list attribute information corresponding to the project list data to be classified includes: Perform unit mapping processing on the unit data in the engineering attribute data to obtain standard unit data; Numerical extraction and magnitude classification processing are performed on the engineering quantity data in the engineering attribute data to obtain engineering magnitude data; The unit price data in the project attribute data is numerically extracted and hierarchically divided to obtain unit price hierarchical data; The specification data in the engineering attribute data is extracted to obtain standard specification data. Based on at least one of the standard unit data, the project quantity data, the unit price level data, and the standard specification data, generate the list attribute information corresponding to the project list data to be classified.
7. The method for classifying engineering bill of quantities data based on multi-feature fingerprints according to claim 1, characterized in that, The fingerprint fusion processing of the standard list name, the project location information, and the list attribute information yields a multi-feature fingerprint corresponding to the project list data to be classified, including: Extract the project location identifier from the project location information, and extract standard unit data and project quantity data from the list attribute information; The fingerprint base data is obtained by combining the engineering part identifier, the standard list name, the standard unit data, and the engineering quantity data. The fingerprint data is encoded to obtain the fingerprint identifier corresponding to the project list data to be classified. The fingerprint identifier, the standard list name, the project location information, and the list attribute information are associated to obtain the multi-feature fingerprint corresponding to the project list data to be classified.
8. The method for classifying engineering bill of quantities data based on multi-feature fingerprints according to claim 7, characterized in that, The matching process between the multi-feature fingerprint and the multiple standard list items includes: Based on the fingerprint identifier in the multi-feature fingerprint, obtain the engineering part identifier, the standard list name, the standard unit data, and the engineering quantity data associated with the fingerprint identifier; Based on the engineering part identifier, the multiple standard list items are filtered by part to obtain candidate standard list items; Based on the names of the standard list, name matching processing is performed on the candidate standard list items to obtain name matching results; Based on the standard unit data and the engineering quantity data, attribute verification processing is performed on the candidate standard list items corresponding to the name matching results to obtain the standard list items to be scored.
9. The method for classifying engineering bill of quantities data based on multi-feature fingerprints according to claim 8, characterized in that, The step of determining the target standard list item corresponding to the project list data to be classified from the plurality of standard list items based on the matching processing results includes: Based on the engineering part identifier, the standard list name, the standard unit data, and the engineering quantity data, calculate the part matching degree, name matching degree, and attribute matching degree corresponding to the standard list item to be scored; The matching scores of the part matching degree, the name matching degree, and the attribute matching degree are weighted to obtain the matching scores corresponding to the list of criteria items to be scored. The matching result is generated based on the matching score, and the target standard list item is determined from the list of standards to be scored based on the matching result.
10. A system for classifying engineering bill of quantities data based on multi-feature fingerprints, used to implement the engineering bill of quantities data classification method based on multi-feature fingerprints as described in any one of claims 1 to 9, characterized in that, The system includes: The data acquisition module is used to acquire the project list data to be classified and the list classification knowledge base; wherein, the project list data to be classified includes the project list name, project location source data and project attribute data, and the list classification knowledge base includes multiple standard list items and list description mapping data associated with each standard list item; The name normalization module is used to normalize the names of the project list using the list description mapping data, so as to obtain the standard list name corresponding to the project list data to be classified. The part attribute module is used to perform part identification processing on the source data of the engineering part to obtain the engineering part information corresponding to the engineering list data to be classified, and to perform attribute standardization processing on the engineering attribute data to obtain the list attribute information corresponding to the engineering list data to be classified. The fingerprint fusion module is used to perform fingerprint fusion processing on the standard list name, the project part information and the list attribute information to obtain the multi-feature fingerprint corresponding to the project list data to be classified. The classification and matching module is used to perform matching processing on the multi-feature fingerprint and the multiple standard list items, and to determine the target standard list item corresponding to the project list data to be classified from the multiple standard list items based on the matching processing results.