A material classification and attribute extraction method for the construction industry based on large language model
By using large language models to classify and extract materials in the construction industry, the problems of inefficiency of traditional methods and information islands are solved, and efficient and accurate material information processing and the ability to quickly adapt to industry changes is achieved.
Patent Information
- Application Number
- CN202411898504.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-23
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-12-23
AI Technical Summary
The material classification and attribute extraction methods in the traditional construction industry are inefficient and difficult to expand, and the information island phenomenon is serious, making it difficult to quickly adapt to industry changes.
A method based on a large language model is adopted to build traditional databases and vector databases through standard classification system data sets and full-category data sets, and material classification and attribute extraction are carried out in combination with the semantic understanding ability of the large language model to realize dual retrieval and semantic analysis.
It improves the efficiency and accuracy of material classification and attribute extraction, enhances the robustness and adaptability of the workflow, and can effectively deal with complex and irregular inputs.
Smart Images

Figure CN119378553B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular to a material classification and attribute extraction method for the construction industry based on a large language model. Background Art
[0002] With the rapid development of artificial intelligence technology, Large Language Model (LLM) has been widely used in many industries. As an industry involving multiple fields, complex processes and huge amounts of data, the construction industry has an increasing demand for material classification and attribute extraction. In traditional methods, material classification and attribute extraction mainly rely on manual experience or rule-driven systems. These methods usually have the following problems: low efficiency, manual classification and attribute extraction require a lot of manpower and time, especially when dealing with complex and large amounts of data, the efficiency is difficult to meet the needs of rapid advancement of modern construction projects; difficult to expand, rule-driven systems usually require a large number of rules to be defined in advance to deal with different types of materials and attributes.
[0003] However, with the continuous emergence of new materials and technologies, the cost of updating rules is high and the cycle is long, making it difficult to quickly adapt to changes in the industry; information islands. In actual operations, building material information is often scattered in different data sources and platforms. The information island effect makes it impossible to effectively integrate data and it is difficult to grasp the classification and attribute information of materials from a global perspective. In response to these problems, the work of classifying and extracting materials and attributes in the construction industry using the semantic understanding ability of large models came into being.
[0004] The large language model has a strong semantic understanding ability and can understand complex text information. It can classify materials and extract attributes by directly analyzing the irregular input of users without manually defining complex rules. This not only reduces the system maintenance cost, but also improves the scalability and flexibility of the system, and significantly improves work efficiency. Therefore, a material classification and attribute extraction method for the construction industry based on a large language model is urgently needed. Summary of the invention
[0005] The purpose of the present invention is to provide a material classification and attribute extraction method for the construction industry based on a large language model to at least solve some of the above-mentioned technical problems.
[0006] To achieve the above purpose, the technical solution adopted by the present invention is as follows:
[0007] A method for material classification and attribute extraction in the construction industry based on a large language model comprises the following steps:
[0008] S1. Construct a traditional database and a vector database based on the standard classification system data set and the full category data set. The vector database includes a standard classification name vector database, a standard classification sample vector database, a full category classification name vector database, and a full category classification sample vector database;
[0009] S2. Input irregular text, and search the irregular text based on the large language model using the standard classification name vector database, the standard classification sample vector database, the full category classification name vector database, and the full category classification sample vector database to obtain search information; perform preliminary matching and secondary matching on the search information based on the large language model to obtain the material classification name;
[0010] S3. Extract the attributes of material classification names from the traditional database based on the large language model to obtain material attributes.
[0011] Furthermore, the traditional database includes a standard classification name table, a standard classification sample table, a full category classification name table and a full category classification sample table.
[0012] Furthermore, the standard classification name table is constructed based on the standard classification system data set, including the first-level classification name, second-level classification name, third-level classification name, fourth-level classification name in the classification system data set, and the classification attributes, classification attribute values, classification definitions and example samples corresponding to the classification names at each level; the standard classification sample table is constructed based on the standard classification system data set, including the standard classification samples in the standard classification system data set, the first-level classification name, second-level classification name, third-level classification name, fourth-level classification name corresponding to each standard classification sample, and the classification attributes, classification attribute values, and classification definitions corresponding to the classification names at each level; the full category classification name table is constructed based on the full category data set, including the first-level classification name, second-level classification name, third-level classification name, fourth-level classification name in the full category data set, and the classification attributes, classification attribute values, classification definitions and example samples corresponding to the classification names at each level; the full category classification sample table is constructed based on the full category data set, including the full category classification samples in the full category data set, the first-level classification name, second-level classification name, third-level classification name, fourth-level classification name corresponding to each full category classification sample, and the classification attributes, classification attribute values, and classification definitions corresponding to the classification names at each level.
[0013] Furthermore, each level of classification name is separated by a separator.
[0014] Furthermore, the standard classification name vector database contains all the first-level classification names, second-level classification names, third-level classification names, and fourth-level classification names in the standard classification system data set; the standard classification sample vector database contains all the standard classification samples in the standard classification system data set; the all-category classification name vector database contains all the first-level classification names, second-level classification names, third-level classification names, and fourth-level classification names in the all-category data set; the all-category classification sample vector database contains all the all-category classification samples in the all-category data set.
[0015] Further, the S2 includes: S21, inputting irregular text; S22, extracting the material name of the irregular text through the large language model and the preset first prompt project, searching the material name in the standard classification name vector database and the full category classification name vector database respectively, obtaining similar classification names, and searching the retrieved classification names in the standard classification name table and the full category classification name table of the traditional database respectively for corresponding classification attributes, classification attribute values, classification definitions and example samples; at the same time, searching the irregular text in the standard classification sample vector database and the full category classification sample vector database respectively to obtain similar classification samples, the similar classification samples are standard classification samples and full category classification samples similar to the irregular text, and searching the retrieved classification samples in the standard classification sample table and the full category classification sample table of the traditional database respectively for corresponding classification names, classification attributes, classification attribute values, and classification definitions; S23, The classification name retrieved each time and the corresponding classification attributes, classification attribute values, classification definitions and example samples are merged into one retrieval information, the irregular text and the retrieval information are input into the large language model, and the preset second Prompt project is used to perform language matching on each retrieval information to obtain a preliminary classification name; S24, each preliminary classification name is searched in the standard classification name vector database, and it is confirmed that the preliminary classification name is in the classification name of the standard classification system data set, and then the classification attributes, classification attribute values, classification definitions and example samples corresponding to the preliminary classification name are found in the standard classification name table of the traditional database, and each preliminary classification name and the corresponding classification attributes, classification attribute values, classification definitions and example samples are merged into one classification information, the irregular text and the classification information are input into the large language model, and the preset third Prompt project is used to perform language matching on all classification information to obtain a secondary classification name, and the secondary classification name is the material classification name.
[0016] Furthermore, the search information and classification information start with a start mark and end with an end mark.
[0017] Furthermore, the S3 includes: S31, retrieving all classification attributes of the S2 material classification name in the standard classification name table of the traditional database; S32, using the retrieved classification attributes as a template, utilizing a large language model and adopting a preset fourth prompt project to analyze the irregular text, extracting classification attribute values of all classification attributes in the template from the irregular text, and obtaining all classification attributes and corresponding classification attribute values as outputs of material attributes.
[0018] Furthermore, in S32, if a classification attribute in the template does not have a corresponding classification attribute value extracted from the irregular text, the classification attribute value of the current classification attribute is assigned as a blank.
[0019] Compared with the prior art, the present invention has the following beneficial effects:
[0020] The present invention utilizes the semantic understanding ability and information retrieval mechanism of a large language model to integrate multi-source data, thereby achieving efficient and accurate material classification and attribute extraction. First, a traditional database and a vector database are constructed based on the standard classification system data set and the full category data set to ensure that a wide range of and accurate classification information is covered; then, a vector database is combined with a traditional database for double retrieval to recall classification samples and classification information related to the input information; then, this information is sorted, and the semantic analysis ability of a large language model is used to match the sorted information with the user input to accurately determine the material classification name; finally, based on the material classification name, the user input is parsed in combination with the traditional database and the semantic analysis ability of the large language model is used again to extract material attributes, thereby ensuring the consistency and completeness of the classification name and classification attribute extraction. This method not only improves the efficiency of material classification and attribute extraction in the construction industry, but also enhances the robustness and adaptability of the workflow, and can cope with complex and irregular inputs. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 The figure is a flow chart of the method of the present invention. DETAILED DESCRIPTION
[0022] In order to make the purpose, technical solutions and advantages of the present invention clearer, the present invention is further described in detail below in conjunction with the accompanying drawings, examples and embodiments. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present invention. In the description of the present invention, it should be noted that the terms "first", "second" and "third" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance.
[0023] like Figure 1As shown, the present invention provides a method for classifying and extracting materials in the construction industry based on a large language model, comprising the following steps:
[0024] S1. Construct a traditional database and a vector database based on the standard classification system data set and the full category data set. The vector database includes a standard classification name vector database, a standard classification sample vector database, a full category classification name vector database, and a full category classification sample vector database;
[0025] S2. Input irregular text, and search the irregular text based on the large language model using the standard classification name vector database, the standard classification sample vector database, the full category classification name vector database, and the full category classification sample vector database to obtain search information; perform preliminary matching and secondary matching on the search information based on the large language model to obtain the material classification name;
[0026] S3. Extract the attributes of material classification names from the traditional database based on the large language model to obtain material attributes.
[0027] The present invention combines prompt word engineering (Prompt) and information retrieval augmented generation (RAG) technology. Prompt word engineering (Prompt) is implemented based on a large language model, and the entire classification process is embodied as the application of RAG technology. However, in the application of traditional RAG technology, most of them only use vector databases to retrieve recall information, while the present invention uses vector databases combined with traditional databases to retrieve recall information, which not only improves the efficiency of material classification and attribute extraction in the construction industry, but also enhances the robustness and adaptability of the workflow.
[0028] First, the present invention constructs a traditional database and a vector database based on the standard classification system data set and the full category data set. The standard classification system data set is a selected data with a small amount; while the full category data set is relatively extensive, but it is not 100% correct and may contain incorrect classifications. Therefore, the traditional database and the vector database are constructed through the two data sets to achieve effective integration of multi-source data.
[0029] Preferably, the traditional database includes a standard classification name table, a standard classification sample table, a full category classification name table and a full category classification sample table.
[0030] More preferably, the standard classification name table is constructed based on the standard classification system dataset, and includes the first-level classification name, second-level classification name, third-level classification name, fourth-level classification name in the classification system dataset, as well as the classification attributes, classification attribute values, classification definitions, and example samples corresponding to each level of classification name. The primary key of the standard classification name table is the first-level classification name, second-level classification name, third-level classification name, and fourth-level classification name in the classification system dataset, and the other keys are the classification attributes, classification attribute values, classification definitions, and example samples corresponding to each level of classification name. If there is no corresponding information for the other keys, it is a blank (null).
[0031] The following is a specific example:
[0032] Classification name: Insurance, insulation and electrothermal materials <;> Insurance equipment <;> Insurance equipment <;> Fuse;
[0033] Classification attribute: Rated current: 5a;
[0034] Classification definition: First-level classification definition: Includes various switches and sockets \t Second-level classification definition: Includes insurance materials such as fuses, insurance belts, insurance pieces, insurance covers, and insurance racks;
[0035] Example samples: Example 1: CAITONG / Caitong / Fuse \t Example 2: Foshan Fuse / Fuse wire Fuse \t Example 3: CHINT Fuse / Fuse wire Household knife switch insurance...
[0036] Among them, insurance, insulation and electrothermal materials are the first-level classification name, insurance equipment is the second-level classification name, insurance equipment is the third-level classification name, and fuse is the fourth-level classification name; rated current is a specific classification attribute, and 5a is the classification attribute value corresponding to the rated current; the classification definition is the definition corresponding to each level of classification name; Example 1: CAITONG / Caitong / Fuse \t is a specific sample, and Example 2, Example 3, etc. are other specific samples. More preferably, each level of classification name in the standard classification name table is separated by the delimiter <;>.
[0037] The standard classification sample table is constructed based on the standard classification system dataset, and includes the standard classification samples in the standard classification system dataset, the first-level classification name, second-level classification name, third-level classification name, fourth-level classification name corresponding to each standard classification sample, as well as the classification attributes, classification attribute values, and classification definitions corresponding to each level of classification name. The primary key of the standard classification sample is the standard classification sample, and the standard classification sample is the sample marked in the standard classification system dataset. The other keys are the first-level classification name, second-level classification name, third-level classification name, fourth-level classification name corresponding to each standard classification sample, as well as the classification attributes, classification attribute values, and classification definitions corresponding to each level of classification name. Similarly, if there is no corresponding information for the other keys, it is a blank (null).
[0038] The following is a specific example:
[0039] Standard classification sample: CAITONG / Caitong / Fuse;
[0040] Classification name: Insurance, insulation and electrothermal materials<;>Insurance equipment<;>Insurance equipment<;>Fuse;
[0041] Classification attribute: Rated current: 5a;
[0042] Classification definition: First-level classification definition: Includes various switches and sockets\tSecond-level classification definition: Includes insurance materials such as fuses, insurance belts, insurance pieces, insurance covers, and insurance racks.
[0043] Among them, CAITONG / Caitong / Fuse\tis a sample marked in the standard classification system dataset; Insurance, insulation and electrothermal materials is the first-level classification name, Insurance equipment is the second-level classification name, Insurance equipment is the third-level classification name, and Fuse is the fourth-level classification name; Rated current is a specific classification attribute, and 5a is the classification attribute value corresponding to the rated current. More preferably, each level of classification name in the standard classification sample table is separated by the delimiter <;>.
[0044] The full-category classification name table is constructed based on the full-category dataset, including the first-level classification name, second-level classification name, third-level classification name, fourth-level classification name in the full-category dataset, as well as the classification attributes, classification attribute values, classification definitions, and example samples corresponding to each level of classification name. The full-category classification name table has the same table structure as the standard classification name table, but the full-category classification name table has a more comprehensive classification compared to the standard classification name table.
[0045] The full-category classification sample table is constructed based on the full-category dataset, including the full-category classification samples in the full-category dataset, the first-level classification name, second-level classification name, third-level classification name, fourth-level classification name corresponding to each full-category classification sample, as well as the classification attributes, classification attribute values, and classification definitions corresponding to each level of classification name. The full-category classification sample table has the same table structure as the standard classification sample table, but the full-category classification samples in the full-category classification sample table are samples in the full-category dataset, and its samples are more extensive.
[0046] Preferably, the standard classification name vector database contains all the first-level classification names, second-level classification names, third-level classification names, and fourth-level classification names in the standard classification system data set; the standard classification sample vector database contains all the standard classification samples in the standard classification system data set; the all-category classification name vector database contains all the first-level classification names, second-level classification names, third-level classification names, and fourth-level classification names in the all-category data set; the all-category classification sample vector database contains all the all-category classification samples in the all-category data set.
[0047] Since irregular text is an irregular text description input by the user, these inputs may include material information of various formats and contents. Therefore, we first perform semantic analysis on the irregular text input by the user, and use the powerful semantic understanding ability of the large language model (LLM model) to extract the material name from it. This step is guided by the Prompt project of the LLM model, so that the LLM model can accurately identify and extract the core name information of the material.
[0048] The following is a specific embodiment, which is implemented by presetting the first prompt project:
[0049] <task>
[0050] Extract material name from input
[0051] < / task>
[0052] <description>
[0053] This task requires identifying and extracting material names from a given text input. The output should not contain any additional content and should not include thinking.
[0054] < / description>
[0055] <example>
[0056] Input: Floor spring door M1034 single door - floor spring + lock unit: set.
[0057] Output: Ground spring door
[0058] <instructions>
[0059] 1. First, analyze the input text and identify all possible material names.
[0060] 2. If there are multiple material names, please separate them with ',' and then output the material names.
[0061] 3. Ensure that the output only contains the material name and does not contain any additional content such as attribute specifications.
[0062] 4. It is strictly forbidden to output any explanation, analysis, reasoning, additional explanation or thought process.
[0063] < / instructions> .
[0064] After obtaining the material name, a double search is performed. On the one hand, according to the extracted material name, the material name is searched in the standard classification name vector database and the full category classification name vector database respectively to obtain similar classification names, and the retrieved classification names are searched in the standard classification name table and the full category classification name table of the traditional database to find the corresponding classification attributes, classification attribute values, classification definitions and example samples; on the other hand, the irregular text is searched in the standard classification sample vector database and the full category classification sample vector database respectively to obtain similar classification samples, and the similar classification samples are standard classification samples and full category classification samples similar to the irregular text, and the retrieved classification samples are searched in the standard classification sample table and the full category classification sample table of the traditional database to find the corresponding classification name, classification attribute, classification attribute value, and classification definition. Through the above four searches of double search, multiple classification information most relevant to the input can be widely recalled to ensure wide and accurate coverage.
[0065] Then, the category name and corresponding category attributes, category attribute values, category definitions and sample samples retrieved each time are combined into a search information. Each search information begins with a start mark " <start>"Begins with" and ends with" <end>" is the end, and "<class_start> "to"<class_end> "The content between is the grade name of the material,<attribute_start> "to"<attribute_end> "The content between is the material's classification attribute and classification attribute value, "<define_start> "to"<define_end> "The content between is the material classification definition,<samples_start> "to"<samples_end> The content between " is an example sample. Then the irregular text and search information are input into the large language model, and the preset second prompt project is used to perform language matching on each search information to obtain a preliminary classification name. The semantic matching process in this step is designed to ensure that each classification judgment is based on a comprehensive analysis, thereby providing highly accurate classification output.
[0066] After obtaining the preliminary classification name, each preliminary classification name is searched in the standard classification name vector database to confirm that the preliminary classification name is in the classification name of the standard classification system data set, and then the classification attributes, classification attribute values, classification definitions and example samples corresponding to the preliminary classification name are found in the standard classification name table of the traditional database, and each preliminary classification name and the corresponding classification attributes, classification attribute values, classification definitions and example samples are merged into one classification information, and the irregular text and classification information are input into the large language model, and the preset third prompt project is used to perform language matching on all classification information to obtain a secondary classification name, which is the material classification name.
[0067] Finally, the extraction of material attributes is based on the above material classification results. The semantic understanding ability of the LLM model is used in combination with the standard attribute information in the traditional database to ensure the accuracy of the extracted attributes. Specifically: S31. Retrieve all the classification attributes of the S2 material classification name in the standard classification name table of the traditional database; S32. Use the retrieved classification attributes as a template, use the large language model and the preset fourth prompt project to analyze the irregular text, extract the classification attribute values of all classification attributes in the template from the irregular text, and obtain all classification attributes and corresponding classification attribute values as material attributes. If a classification attribute in the template does not have a corresponding classification attribute value extracted from the irregular text, the classification attribute value of the current classification attribute is assigned to a blank (null).
[0068] The following is a specific embodiment of S32, which is implemented by presetting the fourth Prompt project:
[0069] <instruction>
[0070] <task-description>
[0071] The task is to assign values to the classification attributes in the template according to the user input, using the retrieved classification attributes as templates. If there is no attribute value corresponding to the classification attribute in the template in the user input, the attribute value of the classification attribute should be set to null.
[0072] The template is as follows:
[0073] <template> {{template}}< / template>
[0074]
[0075] <instructions>
[0076] 1. Receive user input as a parameter, analyze the user input, and identify information corresponding to the attribute classification in the template.
[0077] 2. If the template's classification attribute is "nan" and the corresponding value is "null", there is no need to extract any information and null is directly output.
[0078] 3. Assign the identified attribute value to the corresponding classification attribute in the template. If there is no information corresponding to the template classification attribute in the user input, the attribute value should be set to null.
[0079] 4. The output is given in the following json format: {'key1': 'value1', 'key2': 'value2',...}.
[0080] 5. Make sure the output is in json format, make sure the output does not contain thoughts, and make sure the output does not have any extra content.
[0081] < / instructions> .
[0082] Finally, it should be noted that the above embodiments are only preferred embodiments of the present invention to illustrate the technical solutions of the present invention, rather than limiting them, and certainly not limiting the patent scope of the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions recorded in the aforementioned embodiments, or replace some or all of the technical features therein by equivalents. These modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention. In other words, any changes or modifications made to the main design concept and spirit of the present invention that have no substantive significance, and the technical problems they solve are still consistent with the present invention, should be included in the protection scope of the present invention. In addition, the direct or indirect application of the technical solutions of the present invention in other related technical fields is also included in the patent protection scope of the present invention.< / instruction> < / end> < / start> < / example>
Claims
1. A method for material classification and attribute extraction in the construction industry based on a large language model, characterized in that: The following steps are involved: S1. Construct a traditional database and a vector database based on the standard classification system data set and the full category data set. The vector database includes a standard classification name vector database, a standard classification sample vector database, a full category classification name vector database, and a full category classification sample vector database; S2. Input irregular text, and search the irregular text based on the large language model using the standard classification name vector database, the standard classification sample vector database, the full category classification name vector database, and the full category classification sample vector database to obtain search information; perform preliminary matching and secondary matching on the search information based on the large language model to obtain the material classification name; S3. Extract the attributes of material classification names from the traditional database based on the large language model to obtain material attributes; The S2 comprises: S21, inputting irregular text; S22, extracting the material name of the irregular text through the large language model and the preset first prompt project, searching the material name in the standard classification name vector database and the full category classification name vector database respectively, obtaining similar classification names, and searching the retrieved classification names in the standard classification name table and the full category classification name table of the traditional database respectively for corresponding classification attributes, classification attribute values, classification definitions and example samples; at the same time, searching the irregular text in the standard classification sample vector database and the full category classification sample vector database respectively, obtaining similar classification samples, wherein the similar classification samples are standard classification samples and full category classification samples similar to the irregular text, and searching the retrieved classification samples in the standard classification sample table and the full category classification sample table of the traditional database respectively for corresponding classification names, classification attributes, classification attribute values and classification definitions; S23, The retrieved classification name and the corresponding classification attribute, classification attribute value, classification definition and example sample are combined into one retrieval information, the irregular text and the retrieval information are input into the large language model, and the preset second Prompt project is used to perform language matching on each retrieval information to obtain a preliminary classification name; S24, each preliminary classification name is searched in the standard classification name vector database, and it is confirmed that the preliminary classification name is in the classification name of the standard classification system data set, and then the classification attribute, classification attribute value, classification definition and example sample corresponding to the preliminary classification name are found in the standard classification name table of the traditional database, and each preliminary classification name and the corresponding classification attribute, classification attribute value, classification definition and example sample are combined into one classification information, the irregular text and the classification information are input into the large language model, and the preset third Prompt project is used to perform language matching on all classification information to obtain a secondary classification name, and the secondary classification name is the material classification name.
2. According to claim 1, a method for material classification and attribute extraction in the construction industry based on a large language model is characterized in that: The traditional database includes a standard classification name table, a standard classification sample table, a full category classification name table and a full category classification sample table.
3. According to claim 2, a method for material classification and attribute extraction in the construction industry based on a large language model is characterized in that: The standard classification name table is constructed based on the standard classification system data set, including the first-level classification name, second-level classification name, third-level classification name, fourth-level classification name in the classification system data set, and classification attributes, classification attribute values, classification definitions and example samples corresponding to the classification names at each level; The standard classification sample table is constructed based on the standard classification system data set, including the standard classification samples in the standard classification system data set, the first-level classification name, second-level classification name, third-level classification name, fourth-level classification name corresponding to each standard classification sample, and the classification attributes, classification attribute values, and classification definitions corresponding to each level of classification name; the full category classification name table is constructed based on the full category data set, including the first-level classification name, second-level classification name, third-level classification name, fourth-level classification name, and classification attributes, classification attribute values, classification definitions, and example samples corresponding to each level of classification name in the full category data set; the full category classification sample table is constructed based on the full category data set, including the full category classification samples in the full category data set, the first-level classification name, second-level classification name, third-level classification name, fourth-level classification name corresponding to each full category classification sample, and the classification attributes, classification attribute values, and classification definitions corresponding to each level of classification name.
4. According to claim 3, a method for material classification and attribute extraction in the construction industry based on a large language model is characterized in that: Each level of classification name is separated by a separator.
5. According to claim 3, a method for material classification and attribute extraction in the construction industry based on a large language model is characterized in that: The standard classification name vector database contains all the first-level classification names, second-level classification names, third-level classification names, and fourth-level classification names in the standard classification system data set; the standard classification sample vector database contains all the standard classification samples in the standard classification system data set; the all-category classification name vector database contains all the first-level classification names, second-level classification names, third-level classification names, and fourth-level classification names in the all-category data set; the all-category classification sample vector database contains all the all-category classification samples in the all-category data set.
6. The method for material classification and attribute extraction in the construction industry based on a large language model according to claim 1 is characterized in that: The retrieval information and classification information start with a start mark and end with an end mark.
7. The method for material classification and attribute extraction in the construction industry based on a large language model according to claim 5 is characterized in that: The S3 includes: S31, retrieving all classification attributes of the S2 material classification name in the standard classification name table of the traditional database; S32, using the retrieved classification attributes as a template, utilizing the large language model and adopting the preset fourth prompt project to analyze the irregular text, extracting the classification attribute values of all classification attributes in the template from the irregular text, and obtaining all classification attributes and corresponding classification attribute values as outputs of material attributes.
8. The method for material classification and attribute extraction in the construction industry based on a large language model according to claim 7 is characterized in that: In S32, if a classification attribute in the template does not have a corresponding classification attribute value extracted from the irregular text, the classification attribute value of the current classification attribute is assigned as a blank.
Citation Information
Patent Citations
Interior decoration material classification method
CN107909086A
Building industry material data identification method and device and electronic equipment
CN116738343A
Cited By
Multi-Agent collaborative building material classification and extraction method and system
CN121786201A
Building material classification and extraction method and system based on multi-agent cooperation
CN121786201B