Project file processing method and device, equipment, medium and product
Through a large language model, the structured analysis and classification of project project files is solved, and the problems of low efficiency and high error rate of traditional manual analysis methods are realized, and the precise classification and efficient management of project files are realized.
Patent Information
- Application Number
- CN202510165464.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-06-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional manual analysis methods are inefficient and have high error rates when processing complex and large-scale engineering project files, making it difficult to meet the management needs of large-scale engineering projects.
By using a large language model to perform structured analysis, semantic reasoning and contextual analysis of project files, classify them in combination with preset engineering project categories, and generate a verification result report.
It realizes accurate classification and management of project files, improves file processing efficiency, and ensures the effectiveness and accuracy of classification results.
Smart Images

Figure CN120104801A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of computer science and engineering management, and in particular to a project file processing method, device, equipment, medium and product. Background Art
[0002] As the complexity and scale of engineering projects increase, the types of bidding documents in project management are diverse and complex, and parsing and classifying these documents has become an important requirement in engineering management. Traditional manual analysis methods face problems such as low efficiency, high error rate, and inability to quickly process large amounts of data. As the scale of projects continues to expand, manual methods are difficult to meet the needs.
[0003] Therefore, how to use large language models to accurately classify project files and ensure the validity of the classification results through verification, so as to help improve the management efficiency of project files, is a problem that needs to be solved urgently. Summary of the invention
[0004] The present invention provides a project file processing method, device, equipment, medium and product, so as to accurately classify project files by using a large language model, and ensure the validity of the classification results through verification, so as to improve the management efficiency of project files.
[0005] According to one aspect of the present invention, a project file processing method is provided, comprising:
[0006] In response to a request for processing a target project file, the target project file is determined, and a structured analysis and screening process is performed on the target project file to obtain target text content and target text type that meet the screening conditions;
[0007] Perform semantic reasoning and context analysis on the target text content, and integrate the target text content according to the target text type to obtain the analyzed text;
[0008] Based on the preset large language model, the parsed text is classified to obtain the engineering project category and classification confidence to which the target project file belongs, so as to verify the target project file and generate a verification result report for the target project file.
[0009] According to another aspect of the present invention, there is provided a project file processing device, comprising:
[0010] A screening module, for responding to a processing request for a target project file, determining the target project file, and performing structured parsing and screening processing on the target project file to obtain target text content and target text type that meet the screening conditions;
[0011] The integration module is used to perform semantic reasoning and context analysis on the target text content, and integrate the target text content according to the target text type to obtain the parsed text;
[0012] The verification module is used to classify the parsed text based on the preset large language model, obtain the engineering project category and classification confidence to which the target project file belongs, so as to verify the target project file and generate a verification result report for the target project file.
[0013] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0014] at least one processor; and
[0015] a memory communicatively connected to the at least one processor; wherein,
[0016] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the project file processing method described in any embodiment of the present invention.
[0017] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the project file processing method described in any embodiment of the present invention when executed.
[0018] According to another aspect of the present invention, a computer program product is provided. The computer program product includes a computer program. When the computer program is executed by a processor, the project file processing method of any embodiment of the present invention is implemented.
[0019] The technical solution of the embodiment of the present invention determines the target project file in response to a processing request for the target project file, and performs structured parsing and screening processing on the target project file to obtain the target text content and target text type that meet the screening conditions; performs semantic reasoning and context parsing on the target text content to integrate the target text content in combination with the target text type to obtain the parsed text; based on the preset large language model, the parsed text is classified to obtain the engineering project category and classification confidence to which the target project file belongs, so as to verify the target project file and generate a verification result report for the target project file. By using the large language model to accurately classify the project files and ensuring the validity of the classification results through verification, it can help improve the management efficiency of the project files.
[0020] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present invention, nor are they intended to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0022] Figure 1 is a flow chart of a project file processing method provided in Embodiment 1 of the present invention;
[0023] Figure 2 is a flow chart of a project file processing method provided by Embodiment 2 of the present invention;
[0024] Figure 3 is a structural block diagram of a project file processing device provided by Embodiment 3 of the present invention;
[0025] Figure 4 It is a schematic diagram of the structure of an electronic device provided in Embodiment 4 of the present invention. DETAILED DESCRIPTION
[0026] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
[0027] It should be noted that the terms "first", "second", "target", "candidate", "alternative", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices. The acquisition, storage, use, processing, etc. of data in the technical solution of this application comply with the relevant provisions of national laws and regulations.
[0028] Embodiment 1
[0029] Figure 1 1 is a flowchart of a project file processing method provided in the first embodiment of the present invention; this embodiment is applicable to the case where a project file management system classifies and verifies a newly entered project file. The method can be executed by a project file processing device. The project file processing device can be implemented in the form of hardware and / or software. The project file processing device can be configured in an electronic device and executed by the project file management system, such as Figure 1 As shown, the project file processing method includes:
[0030] S101 . In response to a request for processing a target project file, determine the target project file, and perform structured parsing and screening processing on the target project file to obtain target text content and target text type that meet screening conditions.
[0031] The target project file refers to the project file that needs to be parsed, classified and verified and is sent to the project file management system. The target project file can be a project bidding document. The file format of the target project file can be PDF (Portable Document Format) and DOCX (Document Open XML Format). The screening condition refers to the condition for judging and screening the file content in the target project file. Specifically, it can be to filter out irrelevant information (such as page numbers, headers, copyright statements, etc.) by judging the length, structure, keywords and other features of the text, so as to ensure that the subsequent analysis module receives valid and accurate text data.
[0032] Optionally, according to the file type of the target project file, a corresponding text parsing library can be used to perform structured parsing on the target project file; if the file type is PDF, a PDF parsing library such as pdf plumber is used to extract text content page by page to perform structured parsing on the target project file; if the file type is DOCX, a DOCX parsing library such as python-docx library is used to extract text content page by page to perform structured parsing on the target project file.
[0033] Optionally, the parsed text content can be judged according to preset conditions to filter out text paragraphs that meet the conditions, i.e., the target text content; the preset conditions include at least one of the following: whether the number of paragraph characters is less than a specified threshold, whether the number of paragraph characters exceeds a reasonable range, whether the paragraph contains a colon or a period, whether it contains common engineering project terms, whether there are continuous separators, whether it is table content, keyword matching, whether it contains common bidding information expressions, whether it contains redundant information, and whether it contains invalid content.
[0034] Optionally, the target project file is subjected to structured parsing and screening processing to obtain target text content and target text type that meet the screening criteria, including: performing structured parsing on the target project file to obtain candidate text content and candidate text types in the target project file; and screening the candidate text content based on the candidate text content and candidate text type to obtain target text content and target text type that meet the screening criteria.
[0035] The candidate text content refers to all text content included in the target project file, and the candidate text type can be the text type of each candidate text content, and the candidate text type can be paragraph, table, title or page number, etc. The target text content refers to the text content in the candidate text content that meets the screening conditions. The target text type can be a body paragraph, table or title paragraph.
[0036] The screening conditions include at least one of the following: paragraph judgment conditions, specific mark recognition conditions, anomaly detection conditions, and redundant content filtering conditions. The paragraph judgment conditions may include length judgment conditions and structure judgment conditions, the specific mark recognition conditions may include keyword matching conditions and punctuation format analysis conditions, and the redundant content filtering conditions may include page information removal conditions and separator removal conditions.
[0037] Exemplarily, based on the length judgment condition, the number of characters of each candidate text content can be counted. If the number of characters is less than a certain threshold (such as 20 characters), it is judged as an invalid paragraph (such as a page number or header) and the candidate text content is deleted; if the number of characters is greater than a certain threshold (such as 500 characters), it is judged as a paragraph with splicing errors or abnormal format, and further inspection is required to achieve the screening of candidate text content and obtain the target text content.
[0038] Exemplarily, based on specific mark recognition conditions, it is possible to identify whether each candidate text content contains project-related information based on predefined keywords (such as "construction scope", "bidding unit", "construction period", "contract amount", etc.). If so, the candidate text content is determined as the target text content. Structured symbols such as colons, vertical lines, project numbers, etc. in the candidate text content can also be identified. If the number of symbols is greater than a preset threshold or the distribution is irregular, it can be determined that the candidate text content is invalid information and is discarded and not determined as the target text content.
[0039] Exemplarily, based on the redundant content filtering conditions, page information such as page numbers, headers, footers, etc. (such as "Page X", "Copyright", etc.) in the candidate text content can be eliminated, and a large number (such as more than 5 times) of separators (such as "...") contained in the candidate text content can also be eliminated. If so, it can be determined that the candidate text content is invalid content and discarded from the target text content.
[0040] Optionally, paragraphs may be accurately segmented based on document features such as punctuation marks, blank lines, indents, etc., to obtain candidate text content whose candidate text type is paragraph.
[0041] Optionally, if the candidate text type is a table, the row and column contents of the table can be extracted using the corresponding parsing method according to the file type of the target project file and stored in the form of a two-dimensional array, that is, structured parsing. For example, for PDF files, the table area can be identified through geometric relationships, and for DOCX files, the table can be extracted through the document structure. In addition, line breaks and nested tables in the cell contents of the table can also be processed to expand them into flat data.
[0042] Optionally, if the filter condition is a redundant content filter condition, you can use a preset regular expression to match and clean invalid characters (such as page separators, headers and footers, and blank paragraphs) to automatically remove extra spaces, line breaks, separators, etc.
[0043] Exemplarily, the target text content whose target text type is paragraph can be stored in the form of a paragraph list, with each paragraph marked with its logical position (such as page number, paragraph number), and the target text content whose target text type is table can be stored in row and column format, with each cell content separated by a mark symbol.
[0044] Optionally, the screening condition can also be an anomaly detection condition. Specifically, blank pages, garbled characters, format anomalies and other problems in the file can be detected based on the anomaly detection condition, and skip or repair operations can be performed. At the same time, error prompts are provided in the final verification result report to ensure the completeness and accuracy of the parsing results. For example, if a page full of spaces or no characters is detected in the target project file, a skip operation can be performed. If garbled characters are detected in the target project file, a repair operation can be performed by encoding conversion or regular replacement correction. If there are paragraphs or tables that cannot be parsed during the parsing process using the text parsing library, log output can be used for manual inspection.
[0045] Exemplarily, through the text parsing library, key information such as paragraphs, titles, tables, page numbers, etc. in the target project file can be extracted to obtain candidate text content and candidate text types. Furthermore, based on screening conditions, such as redundant content filtering conditions, useless characters (such as headers, footers, page numbers, etc.) in the target project file can be cleaned up. Paragraphs can also be segmented according to the logical structure of the document to ensure the integrity of the content and obtain the target text content and target text type.
[0046] Exemplarily, if the candidate text type corresponding to the candidate text content is a table, the data in the table can be extracted and converted into a structured format (such as using a separator "|" to mark the cell boundaries of the table) to obtain the target text content corresponding to the paragraph of candidate text content. Optionally, multi-layer nested table processing can also be performed to associate the candidate text content whose candidate text type is a table with the paragraph content.
[0047] Optionally, if the target text type is a paragraph, the text content, page number and paragraph number of each paragraph can be saved based on a preset format to generate the corresponding target text content. If the target text type is a table, the table can be stored in the form of a two-dimensional array, marked with its source page number and paragraph number, to generate the corresponding target text content.
[0048] S102: semantic reasoning and context analysis are performed on the target text content, and the target text content is integrated and processed according to the target text type to obtain a parsed text.
[0049] Optionally, after determining the target text content and the target text type, the target text content can be further semantically reasoned to mark the content type involved in each target text content, wherein the content type can be at least one of the following: qualification conditions, project information, scoring rules, deposit information, and delivery information. Specifically, if the target text content involves information such as bidder qualifications, qualification requirements, and legal terms, the corresponding content type can be marked as qualification conditions; if the target text content is related to basic project information (such as project name, project description, project scope, budget, etc.), the corresponding content type can be marked as project information; if the target text content involves project scoring standards, evaluation systems, bid evaluation procedures, etc., the corresponding content type can be marked as scoring rules; if the target text content involves relevant terms such as bid bonds and performance bonds, the corresponding content type can be marked as deposit information; if the target text content is related to project delivery, delivery time, acceptance conditions, etc., the corresponding content type can be marked as delivery information.
[0050] Optionally, after semantic reasoning of the target text content, each target text content can be represented in a structured manner, including the following attributes: content (i.e., the specific paragraph text of the target text content), status (which can be a valid state, an invalid state, or a state requiring manual review), target text type (which can be a body paragraph, a table paragraph, a title paragraph, project information, etc.) and content type (qualification conditions, scoring criteria, delivery information, etc.).
[0051] Optionally, the relevance between paragraphs can be detected based on the content, status, target text type, and content type of the target text content. For example, when project scope and budget information are scattered in different paragraphs, the information can be completed through context to ensure that the parsed text is semantically complete.
[0052] Optionally, context analysis of the target text content may specifically include the following processes: paragraph association analysis, information integration and completion, multi-paragraph merging processing, logical consistency verification, and abnormal context processing.
[0053] Exemplarily, the semantic connection between paragraphs can be analyzed based on the keywords or pronouns in the target text content, and explicit associations (such as "the budget is 50 million yuan" and "the scope of construction includes bridges and roads") and implicit associations (such as "this project" and "a municipal project" in the previous paragraph) can be identified, that is, paragraph association analysis can be performed; the key information of multiple paragraphs can also be summarized and organized into structured units, and the missing content can be inferred from the context, such as when the budget is not clear, it can be completed from the related paragraphs, that is, information integration and completion can be performed; the information-related paragraphs can also be merged into a new paragraph block to facilitate subsequent large model analysis, while retaining the hierarchical structure of the paragraphs to support multi-dimensional classification, that is, multi-paragraph merging processing can be performed; it can also verify whether the information is logically consistent, such as whether the budget matches the scope of construction, whether the time period is reasonable, and mark the inconsistent content and record it in the exception report. For example, when the budget is less than 10 million yuan, the feedback is that the budget is too low and there may be an error, that is, a logical consistency check is performed; it can also determine whether there are obvious omissions or conflicts between paragraphs. For example, when budget information is missing, the information that cannot be determined is marked and the user is prompted to verify, that is, abnormal context processing is performed.
[0054] S103. Based on the preset large language model, the parsed text is classified to obtain the engineering project category and classification confidence to which the target project file belongs, so as to verify the target project file and generate a verification result report for the target project file.
[0055] Among them, the preset large language model can be a large language model (LLM). The engineering project category can be a single type or a composite type, and can specifically include at least one of the following: engineering construction, engineering goods, engineering services, special engineering, commercial goods, and commercial services. Exemplarily, a target project file containing project characteristic information such as "construction", "civil engineering", and "road reconstruction" can be determined as an engineering construction type, a target project file containing project characteristic information such as "procurement", "equipment", and "materials" can be determined as an engineering goods category, and a target project file containing project characteristic information such as "supervision", "consulting", and "design" can be determined as an engineering service category. Classification confidence refers to an indicator that characterizes the degree of classification reliability output by the large language model after classification.
[0056] It should be noted that if the target project file contains project characteristic information of multiple project types at the same time, it can be determined that the engineering project category is a composite type, specifically, it can be a composite type obtained by combining any single types.
[0057] Optionally, based on a preset large language model, the parsed text is classified to obtain the engineering project category and classification confidence to which the target project file belongs, including: using the preset large language model to extract keywords from the parsed text to obtain target keywords; performing keyword matching on the target keywords and candidate keywords corresponding to the candidate project categories, and classifying the parsed text according to the matching results to obtain the engineering project category and classification confidence to which the target project file belongs.
[0058] Among them, the types of target keywords include at least one of the following: project name, project characteristic information, budget amount and construction period; project characteristic information refers to characteristic keywords that can characterize the project type to which the target project file belongs, and can specifically be "construction", "civil engineering", "road reconstruction", "procurement", "equipment", "materials", "supervision", "consulting" and "design", etc.
[0059] Exemplarily, the named entity recognition (NER) capability of the large language model can be used to identify relevant fields such as "project name" and "project name" in the parsed text to obtain the target keywords corresponding to the project name, and to identify the amount expression in the parsed text (such as "the total budget is 5 million yuan") to obtain the target keywords corresponding to the budget amount, and further identify time-related information in the parsed text (such as "the construction period is 12 months") to obtain the target keywords corresponding to the construction period.
[0060] For example, if the parsed text input into the preset large language model is "This project includes construction and material procurement. The construction part involves bridge and road reconstruction, and the procurement part includes steel bars and concrete", then the corresponding engineering project category can be a composite category of engineering construction and engineering goods.
[0061] Optionally, after using a preset large language model to extract keywords, if it is found that there is no budget or scope information in the target text content, the classification result can be output as unclassifiable or missing information.
[0062] Optionally, based on the matching results, the parsed text is classified to obtain the engineering project category and classification confidence to which the target project file belongs, including: based on the matching results, determining the candidate project category corresponding to the candidate keywords that meet the matching conditions as the engineering project category to which the target project file belongs; and determining the classification confidence based on the degree of matching between the target keywords and the candidate keywords.
[0063] The candidate keywords refer to preset keywords that characterize the characteristics of different candidate item categories, and the matching result may be the similarity between the target keyword and the candidate keyword.
[0064] Optionally, if the similarity between the target keyword and the candidate keyword is greater than a preset similarity threshold, the corresponding candidate keyword can be determined as a candidate keyword that meets the matching conditions, and the candidate project category corresponding to the candidate keyword that meets the matching conditions is determined as the engineering project category to which the target project file belongs. Furthermore, based on the corresponding relationship between the preset similarity and confidence, the classification confidence of this classification result is determined.
[0065] Optionally, the target project file is verified and a verification result report for the target project file is generated, including: classification result verification, confidence verification and anomaly verification of the target project file according to the project type and classification confidence; if the verification passes, a verification result report for the target project file is generated according to the project type, classification confidence and target text content.
[0066] Exemplarily, the classification result verification can be a logical consistency check to verify whether the classification result conforms to the context logic (such as budget rationality, matching of construction scope and project type). If the budget is lower than a reasonable range (such as a large project below 5 million yuan), it is marked as an exception.
[0067] Exemplarily, confidence verification can check the confidence of the large model output, and mark the results classified below a set threshold (such as 80%) as low confidence, provide correction suggestions or prompt manual review.
[0068] Exemplarily, the anomaly check may be to mark as an anomaly when it is determined that the classification result is empty, or the classification is conflicting (such as belonging to multiple unrelated categories at the same time), and generate correction suggestions or prompts for the abnormal classification.
[0069] Optionally, a verification result report for the target project file can be generated based on the project type, classification confidence, and target keywords in the target text content. For example, the verification result report can include the following content: Basic project information: project name, budget, scope, etc. Classification result: type and confidence. Verification status: verification passed or abnormal mark. Suggestions: correction suggestions for low confidence or abnormality.
[0070] Optionally, the classification result of the target project file is verified according to the project type, including: determining the corresponding budget range according to the project type, and determining the declared budget corresponding to the target project file according to the target keywords in the parsed text; verifying the target project file according to the correlation between the declared budget and the budget range to determine the feasibility of the project, and obtaining the verification result of the classification result verification.
[0071] Optionally, based on the correlation between the declared budget and the budget range, if the declared budget is within the budget range, it can be determined that the verification result of the classification result verification is verification passed.
[0072] Optionally, after verification is completed, the verified classification results can be organized into an easy-to-read format, including type, confidence, and key information, and the verification result report can be output in multiple output formats (such as JSON or PDF).
[0073] The technical solution of the embodiment of the present invention determines the target project file in response to a processing request for the target project file, and performs structured parsing and screening processing on the target project file to obtain the target text content and target text type that meet the screening conditions; performs semantic reasoning and context parsing on the target text content to integrate the target text content in combination with the target text type to obtain the parsed text; based on the preset large language model, the parsed text is classified to obtain the engineering project category and classification confidence to which the target project file belongs, so as to verify the target project file and generate a verification result report for the target project file. By using the large language model to accurately classify the project files and ensuring the validity of the classification results through verification, it can help improve the management efficiency of the project files.
[0074] Embodiment 2
[0075] Figure 2 : is a flowchart of a project file processing method provided by the second embodiment of the present invention; based on the above embodiment, this embodiment provides a preferred example of multiple agents working together to implement the parsing, classification and verification of project files, which specifically includes the following process:
[0076] For example, see Figure 2 , the project files can be processed and the verification result report of the target project files can be obtained through the collaborative work among the file parsing agent (represented as Agent A), the condition judgment agent (represented as Agent B), the context coordination agent (represented as Agent C), the large model analysis agent (represented as Agent D) and the result verification and output agent (represented as Agent E).
[0077] Specifically, the file parsing agent (represented as Agent A) can be used to perform document parsing, that is, structured parsing, on the target project file. Furthermore, the conditional judgment agent (represented as Agent B) can be used to screen and process the target project file, and the context coordination agent (represented as Agent C) can be used to perform semantic reasoning and contextual parsing on the target text content to obtain contextual semantic information, so that the large model analysis agent (represented as Agent D) can perform semantic analysis and classification according to the contextual semantic information and state information. The output agent (represented as Agent E) can further perform result and logic verification on the classification result to obtain a verified classification result and an output report, wherein the output agent (represented as Agent E) can feed back the classification result to the large model analysis agent (represented as Agent D) to achieve model optimization, and can also feed back the classification result to the conditional judgment agent (represented as Agent B) to optimize the screening conditions. In addition, the large model analysis agent (represented as Agent D) can also feed back the low-confidence classification result to the conditional judgment agent (represented as Agent B) to optimize the screening conditions.
[0078] Optionally, the file parsing agent can be used to receive input project bidding documents (such as PDF, DOCX format) and extract key information (such as paragraphs, tables, titles, etc.) in the document through a special parsing module. The agent converts the document content into structured data for subsequent agent processing. The conditional judgment agent can be used to filter the structured data provided by the file parsing agent according to preset rules and conditions. By judging the length, structure, keywords and other features of the text, irrelevant information (such as page numbers, headers, copyright statements, etc.) is filtered out to ensure that the subsequent analysis module receives valid and accurate text data. The context coordination agent can be used to contextually integrate the filtered paragraphs to ensure the relevance of cross-paragraph information. For example, if some information (such as budget amount, project scale, etc.) is distributed across multiple paragraphs, the agent will integrate this information into a meaningful whole through semantic reasoning and context analysis, and provide it for subsequent analysis. The large model analysis agent can be used to perform in-depth analysis of the sorted paragraphs and information by using a large language model (such as Qwen 2.5 or a similar NLP model). It can understand the semantics in the text, extract key information related to the project, and classify the project type according to the analysis results. The large model analysis agent can identify the type of project, such as "engineering construction", "engineering goods", "engineering services", etc., and provide a confidence score for each classification result. The result verification and output agent can be used to verify the classification results of the large model analysis agent and check the consistency, logic and accuracy of the classification results. The agent will also generate a final classification report based on the verification results, including the classification results, confidence, analysis basis and suggestions. For low-confidence or difficult classification results, the system will issue a prompt and require manual review.
[0079] Embodiment 3
[0080] Figure 3 : is a structural block diagram of a project file processing device provided in the third embodiment of the present invention; this embodiment can be applied to the situation where the project file management system classifies and verifies the newly entered project files. The project file processing device provided in the embodiment of the present invention can execute the project file processing method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method; the project file processing device can be implemented in the form of hardware and / or software, and configured in an electronic device with a project file processing function, and executed by the project file management system, such as Figure 3 As shown, the project file processing device specifically includes:
[0081] The screening module 301 is used to respond to a processing request for a target project file, determine the target project file, and perform structured parsing and screening processing on the target project file to obtain target text content and target text type that meet the screening conditions;
[0082] An integration module 302 is used to perform semantic reasoning and context analysis on the target text content, and integrate the target text content according to the target text type to obtain a parsed text;
[0083] The verification module 303 is used to classify the parsed text based on the preset large language model, obtain the engineering project category and classification confidence to which the target project file belongs, verify the target project file, and generate a verification result report for the target project file.
[0084] The technical solution of the embodiment of the present invention determines the target project file in response to a processing request for the target project file, and performs structured parsing and screening processing on the target project file to obtain the target text content and target text type that meet the screening conditions; performs semantic reasoning and context parsing on the target text content to integrate the target text content in combination with the target text type to obtain the parsed text; based on the preset large language model, the parsed text is classified to obtain the engineering project category and classification confidence to which the target project file belongs, so as to verify the target project file and generate a verification result report for the target project file. By using the large language model to accurately classify the project files and ensuring the validity of the classification results through verification, it can help improve the management efficiency of the project files.
[0085] Furthermore, the verification module 303 may include:
[0086] An extraction unit is used to extract keywords from the parsed text using a preset large language model to obtain target keywords; the types of the target keywords include at least one of the following: project name, project feature information, budget amount, and construction period;
[0087] The classification unit is used to perform keyword matching on the target keyword and the candidate keywords corresponding to the candidate project category, so as to classify the parsed text according to the matching result and obtain the engineering project category and classification confidence to which the target project file belongs.
[0088] Furthermore, the taxonomic units are specifically used for:
[0089] According to the matching results, the candidate project category corresponding to the candidate keywords that meet the matching conditions is determined as the engineering project category to which the target project file belongs;
[0090] The classification confidence is determined based on the matching degree between the target keyword and the candidate keyword.
[0091] Furthermore, the screening module 301 is specifically used for:
[0092] Performing structural analysis on the target project file to obtain candidate text content and candidate text type in the target project file;
[0093] The candidate text content is screened according to the candidate text content and the candidate text type to obtain the target text content and the target text type that meet the screening conditions; the screening conditions include at least one of the following: paragraph judgment conditions, specific mark recognition conditions and redundant content filtering conditions.
[0094] Furthermore, the verification module 303 may include:
[0095] A verification unit is used to perform classification result verification, confidence verification and anomaly inspection on the target project file according to the project type and classification confidence;
[0096] The generation unit is used to generate a verification result report for the target project file according to the project type, classification confidence and target text content when the inspection passes.
[0097] Furthermore, the verification unit is specifically used for:
[0098] Determine the corresponding budget range according to the project type, and determine the declared budget corresponding to the target project file based on the target keywords in the parsed text;
[0099] According to the correlation between the declared budget and the budget range, the target project files are verified to determine the feasibility of the project and obtain the verification results of the classification results verification.
[0100] Embodiment 4
[0101] Figure 4 It is a schematic diagram of the structure of an electronic device provided in Embodiment 4 of the present invention. Figure 4 A schematic diagram of the structure of an electronic device 10 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein.
[0102] like Figure 4As shown, the electronic device 10 includes at least one processor 11, and a memory connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., wherein the memory stores a computer program that can be executed by at least one processor, and the processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 to the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0103] A number of components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0104] The processor 11 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as a project file processing method.
[0105] In some embodiments, the project file processing method may be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as a storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the project file processing method described above may be performed. Alternatively, in other embodiments, the processor 11 may be configured to execute the project file processing method in any other appropriate manner (e.g., by means of firmware).
[0106] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0107] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the computer program is executed by the processor, the functions / operations specified in the flow chart and / or block diagram are implemented. The computer program may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0108] In the context of the present invention, a computer readable storage medium may be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, device, or equipment. A computer readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, semiconductor systems, devices, or equipment, or any suitable combination of the foregoing. Alternatively, a computer readable storage medium may be a machine readable signal medium. A more specific example of a machine readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0109] To provide interaction with a user, the systems and techniques described herein may be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).
[0110] The systems and techniques described herein may be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0111] A computing system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The client and server relationship is generated by computer programs running on the corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services.
[0112] In one embodiment, the embodiment of the present invention further includes a computer program product, the computer program product includes a computer program, and the computer program implements the project file processing method of any embodiment of the present invention when executed by a processor.
[0113] The computer program product may be implemented in a computer program code for performing the operation of the present invention written in one or more programming languages or a combination thereof, including object-oriented programming languages and conventional procedural programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0114] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps described in the present invention can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution of the present invention can be achieved, and this document does not limit this.
[0115] The above specific implementations do not constitute a limitation on the protection scope of the present invention. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A project file processing method, characterized in that: include: In response to a request for processing a target project file, the target project file is determined, and a structured analysis and screening process is performed on the target project file to obtain target text content and target text type that meet the screening conditions; Perform semantic reasoning and context analysis on the target text content, and integrate the target text content according to the target text type to obtain the analyzed text; Based on the preset large language model, the parsed text is classified to obtain the engineering project category and classification confidence to which the target project file belongs, so as to verify the target project file and generate a verification result report for the target project file.
2. The method according to claim 1, characterized in that Based on the preset large language model, the parsed text is classified to obtain the engineering project category and classification confidence to which the target project file belongs, including: A preset large language model is used to extract keywords from the parsed text to obtain target keywords; the types of target keywords include at least one of the following: project name, project feature information, budget amount, and construction period; Keyword matching is performed on the target keyword and the candidate keywords corresponding to the candidate project category, so as to classify the parsed text according to the matching result and obtain the engineering project category and classification confidence to which the target project file belongs.
3. The method according to claim 2, characterized in that According to the matching results, the parsed text is classified to obtain the engineering project category and classification confidence to which the target project file belongs, including: According to the matching results, the candidate project category corresponding to the candidate keywords that meet the matching conditions is determined as the engineering project category to which the target project file belongs; The classification confidence is determined based on the matching degree between the target keyword and the candidate keyword.
4. The method according to claim 1, characterized in that: Perform structured parsing and screening of the target project files to obtain target text content and target text types that meet the screening criteria, including: Performing structural analysis on the target project file to obtain candidate text content and candidate text type in the target project file; The candidate text content is screened according to the candidate text content and the candidate text type to obtain the target text content and the target text type that meet the screening conditions; the screening conditions include at least one of the following: paragraph judgment conditions, specific mark recognition conditions and redundant content filtering conditions.
5. The method according to claim 1, characterized in that Verify the target project file and generate a verification result report for the target project file, including: According to the project type and classification confidence, the target project files are checked for classification results, confidence and anomaly. If the inspection passes, a verification result report for the target project file is generated based on the project type, classification confidence and target text content.
6. The method according to claim 5, characterized in that The target project files are classified and verified according to the project type, including: Determine the corresponding budget range according to the project type, and determine the declared budget corresponding to the target project file based on the target keywords in the parsed text; According to the correlation between the declared budget and the budget range, the target project files are verified to determine the feasibility of the project and obtain the verification results of the classification results verification.
7. A project file processing device, characterized in that: include: A screening module, for responding to a processing request for a target project file, determining the target project file, and performing structured parsing and screening processing on the target project file to obtain target text content and target text type that meet the screening conditions; The integration module is used to perform semantic reasoning and context analysis on the target text content, and integrate the target text content according to the target text type to obtain the parsed text; The verification module is used to classify the parsed text based on the preset large language model, obtain the engineering project category and classification confidence to which the target project file belongs, so as to verify the target project file and generate a verification result report for the target project file.
8. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the project file processing method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the project file processing method according to any one of claims 1 to 6 when executed.
10. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the computer program implements the project file processing method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Method and device for visually analyzing financial punishment data
CN111782917A
Class prediction method, device and equipment based on large language model
CN117390497A
Large language model auxiliary classification method and device, equipment and medium
CN117851598A
Text classification method and device based on large language model, medium and electronic equipment
CN117951572A
Private data classification and grading method
CN118133221A
Cited By
Charge item determination method and device based on list splitting, medium and electronic equipment
CN121998571A
Invoice split-based cost item determination method, device, medium, and electronic device
CN121998571B