Purchase file information extraction method and device based on large model, equipment and medium
Through the automatic parsing and processing of procurement files by the large language model, the problem of inefficiency of traditional manual methods is solved, intelligent data extraction and analysis is realized, and output files that meet business needs are generated.
Patent Information
- Application Number
- CN202510483520.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-07-08
AI Technical Summary
Traditional manual methods for data extraction, sorting and analysis are inefficient and error-prone, especially when facing complex large-scale procurement documents, it is difficult to meet actual needs.
Use large language models to analyze and process procurement files, combine natural language processing and data mining technology, automatically extract and generate structured key information, and generate target output files through configuration files and target templates.
It realizes intelligent information extraction of procurement files, reduces manual intervention, and improves the efficiency and accuracy of data processing, especially the processing capabilities of complex Excel files.
Smart Images

Figure CN120278140A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of procurement, and particularly to a method, device, equipment and medium for extracting information from procurement documents based on a large model. Background Art
[0002] With the rapid development of information technology, the world has entered the big data era. Various enterprises and organizations are dealing with a large amount of spreadsheet data every day, such as files in formats like Excel and CSV (Comma-Separated Values). These data are flexible in form and widely used in many scenarios such as financial statements, business records, market analysis, customer management, etc. However, with the increase in data scale and complexity, the traditional manual methods for data extraction, sorting and analysis can no longer meet the actual needs. Especially when facing complex large-scale data, manual operations are not only inefficient but also prone to errors. Therefore, how to achieve intelligent extraction of information from procurement documents is the problem to be solved currently. Summary of the Invention
[0003] In view of this, the purpose of the present invention is to provide a method, device, equipment and medium for extracting information from procurement documents based on a large model, which can achieve intelligent extraction of information from procurement documents. The specific solutions are as follows:
[0004] In a first aspect, the present application discloses a method for extracting information from procurement documents based on a large model, which is applied to a file information extraction system and includes:
[0005] Using a target configuration function to read a target configuration file to determine a preset file reading path, model calling parameters, a preset file format and a target output file name;
[0006] Reading a procurement file to be recognized based on the preset file reading path, and performing data extraction on the procurement file to be recognized to obtain table data to be recognized; the procurement file to be recognized is a table file;
[0007] Invoking a target large model using the model calling parameters to parse the table data to be recognized to obtain key information to be matched; the target large model is a trained large language model;
[0008] Processing the key information to be matched using a target template to obtain target key information, and generating a target output file based on the target template, the preset file format, the target output file name and the target key information to complete the extraction of information from the procurement file to be recognized.
[0009] Optionally, the step of using the target configuration function to read the target configuration file to determine the preset file reading path, model call parameters, preset file format, and target output file name includes:
[0010] Using the target configuration function to read the target configuration file to determine the preset file reading path, model call parameters, preset file format, name of the procurement file to be recognized, and target output file name;
[0011] Correspondingly, the step of reading the procurement file to be recognized based on the preset file reading path includes:
[0012] Determining all the tabular files corresponding to the preset file reading path;
[0013] Determining the procurement file to be recognized from all the tabular files based on the name of the procurement file to be recognized, and reading the procurement file to be recognized.
[0014] Optionally, the step of determining the procurement file to be recognized from all the tabular files based on the name of the procurement file to be recognized includes:
[0015] If the name of the procurement file to be recognized is empty, determining the first tabular file under the preset file reading path as the procurement file to be recognized.
[0016] Optionally, the step of reading the procurement file to be recognized based on the preset file reading path and performing data extraction on the procurement file to be recognized to obtain the table data to be recognized includes:
[0017] Reading the procurement file to be recognized based on the preset file reading path;
[0018] Performing preprocessing on the procurement file to be recognized to obtain the processed procurement file;
[0019] Using the preset file reading function to perform data extraction on the processed procurement file to obtain the original table data;
[0020] Processing the original table data based on the first target format to obtain the corresponding table data to be recognized;
[0021] Among them, the step of performing preprocessing on the procurement file to be recognized to obtain the processed procurement file includes:
[0022] Replacing the first target character in the procurement file to be recognized with a preset character;
[0023] Deleting the second target character in the procurement file to be recognized;
[0024] Processing the null values in the procurement file to be recognized based on the preset processing method to obtain the processed procurement file;
[0025] Among them, the first target format includes the DataFrame format; the first target character includes the line break character, and the second target character includes the blank character; the preset processing methods include the first processing method and the second processing method. The first processing method includes converting the null values in the to-be-identified procurement document into preset values, and the second processing method includes marking the null values in the to-be-identified procurement document as missing values.
[0026] Optionally, the step of using the model call parameter to call the target large model to parse the to-be-identified tabular data to obtain the to-be-matched key information includes:
[0027] Performing format conversion on the to-be-identified tabular data based on the second target format to obtain the to-be-parsed data;
[0028] Determining all target title data in the to-be-parsed data and the target data blocks corresponding to each of the target title data based on the structure information in the to-be-parsed data;
[0029] Using the model call parameter to call the target large model to parse each of the target data blocks to obtain corresponding parsing results;
[0030] Performing format conversion on each of the parsing results based on the third target format to obtain the to-be-matched key information;
[0031] Among them, the second target format includes the Markdown format, and the third target format includes the JSON format.
[0032] Optionally, the step of using the target template to process the to-be-matched key information to obtain the target key information includes:
[0033] Traversing and parsing the to-be-matched key information based on the preset field name and preset structure to obtain the to-be-adapted information;
[0034] Performing format adaptation on the to-be-adapted information based on the target template to obtain the to-be-verified information;
[0035] Verifying and cleaning the data type of the to-be-verified information based on the preset data cleaning rules to obtain the target key information.
[0036] Optionally, the step of generating the target output file based on the target template, the preset file format, the target output file name, and the target key information includes:
[0037] Determining the mapping relationship between each marked cell in the target template and each of the preset field names based on the key-value mapping method;
[0038] Populate the target key information corresponding to each of the preset field names into the target template based on the mapping relationship to obtain a populated file;
[0039] Generate a target output file based on the preset file format, the target output file name, and the populated file.
[0040] In a second aspect, the present application discloses a procurement document information extraction device based on a large model, which is applied to a document information extraction system and includes:
[0041] A configuration file reading module, configured to read a target configuration file using a target configuration function to determine a preset file reading path, model call parameters, a preset file format, and a target output file name;
[0042] A data extraction module, configured to read a procurement document to be recognized based on the preset file reading path, and perform data extraction on the procurement document to be recognized to obtain table data to be recognized; the procurement document to be recognized is a table file;
[0043] A data parsing module, configured to call a target large model using the model call parameters to parse the table data to be recognized to obtain key information to be matched; the target large model is a trained large language model;
[0044] A file generation module, configured to process the key information to be matched using a target template to obtain target key information, and generate a target output file based on the target template, the preset file format, the target output file name, and the target key information, so as to complete the information extraction of the procurement document to be recognized.
[0045] In a third aspect, the present application discloses an electronic device, including:
[0046] A memory, configured to store a computer program;
[0047] A processor, configured to execute the computer program to implement the foregoing procurement document information extraction method based on a large model.
[0048] In a fourth aspect, the present application discloses a computer-readable storage medium, configured to store a computer program, wherein the computer program, when executed by a processor, implements the foregoing procurement document information extraction method based on a large model.
[0049] In this application, when extracting information from procurement documents, the document information extraction system uses a target configuration function to read a target configuration file to determine a preset file reading path, model call parameters, a preset file format, and a target output file name; reads a procurement document to be recognized based on the preset file reading path, and performs data extraction on the procurement document to be recognized to obtain table data to be recognized; the procurement document to be recognized is a table file; calls a target large model using the model call parameters to parse the table data to be recognized to obtain key information to be matched; the target large model is a trained large language model; processes the key information to be matched using a target template to obtain target key information, and generates a target output file based on the target template, the preset file format, the target output file name, and the target key information, so as to complete the information extraction of the procurement document to be recognized. It can be seen that this application uses a trained large language model with powerful semantic understanding capabilities, combines natural language processing and data mining technologies to process the procurement document to be recognized read based on the preset file reading path, enabling users to achieve intelligent data extraction and analysis without having to master complex programming skills or manually operate Excel. Especially for Excel files with complex data and relationships, the large model can automatically process complex logic across tables and multiple levels through its powerful semantic understanding capabilities, thereby generating a target output file that matches the target template and providing highly accurate information extraction results for the procurement document to be recognized. Description of the Drawings
[0050] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on the provided drawings without creative efforts.
[0051] Figure 1 It is a flowchart of a method for extracting procurement document information based on a large model disclosed in this application;
[0052] Figure 2 It is a schematic diagram of a specific process for extracting procurement document information based on a large model disclosed in this application;
[0053] Figure 3 It is a schematic diagram of a specific process for extracting procurement document information based on a large model disclosed in this application;
[0054] Figure 4 It is a schematic diagram of the structure of a device for extracting procurement document information based on a large model disclosed in this application;
[0055] Figure 5 A structural diagram of an electronic device disclosed in this application. Detailed implementation manners
[0056] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0057] With the rapid development of information technology, the world has entered the big data era. Various enterprises and organizations are dealing with a large amount of spreadsheet data every day, such as files in Excel, CSV and other formats. These data forms are flexible and widely used in many scenarios such as financial statements, business records, market analysis, customer management, etc. However, with the increase in data scale and complexity, the traditional manual methods for data extraction, sorting and analysis can no longer meet the actual needs. Especially when facing complex large-scale data, manual operations are not only inefficient but also prone to errors. To solve the above technical problems, this application discloses a method for extracting procurement document information based on a large model, which can realize intelligent extraction of information from procurement documents.
[0058] See Figure 1 As shown in the figure, an embodiment of the present invention discloses a method for extracting procurement document information based on a large model, which is applied to a document information extraction system and includes:
[0059] Step S11: Use a target configuration function to read a target configuration file to determine a preset file reading path, model call parameters, a preset file format, and a target output file name.
[0060] In this embodiment, the process of the document information extraction system using the target configuration function to read the target configuration file is the preparatory work before data reading of the procurement document. By reading the target configuration file, relevant configuration information can be determined, including the preset file reading path, model call parameters, preset file format, target output file name, etc. Therefore, the administrator can flexibly adjust the parameter configuration in different operating environments by modifying the configuration file, so as to achieve seamless switching between different application scenarios. For example, when the document information extraction system extracts procurement document information on different servers, variables such as API (Application Programming Interface) tokens and preset file reading paths can be adjusted by modifying the configuration file.
[0061] Step S12: Read the procurement document to be recognized based on the preset file reading path, and perform data extraction on the procurement document to be recognized to obtain the table data to be recognized; the procurement document to be recognized is a table file.
[0062] In this embodiment, in addition to including the preset file reading path, model call parameters, preset file format, and target output file name, the configuration file may also include the name of the procurement document to be recognized. Therefore, the process of the file information extraction system reading the procurement document to be recognized based on the preset file reading path may specifically include: determining all table files corresponding to the preset file reading path, and then determining the procurement document to be recognized from all the table files under the current preset file reading path based on the name of the procurement document to be recognized, and reading the procurement document to be recognized. That is to say, the file information extraction system will read the path and name of the Excel file to be recognized from the configuration file. For complex Excel files commonly found in business scenarios, which may contain multiple worksheets, the user can specify the name of the worksheet to be processed through the configuration file. It can be understood that if the name of the procurement document to be recognized is not in the configuration file, that is, the name of the procurement document to be recognized in the configuration file is empty, the first table file under the preset file reading path will be defaultly determined as the procurement document to be recognized. That is to say, if the worksheet name is not specified, the file information extraction system will defaultly read the data in the first worksheet. It should be noted that in addition to obtaining the path and name of the procurement document to be recognized from the configuration file, the file information extraction system can also read these data from the user's input.
[0063] In this embodiment, the file information extraction system reads the procurement document to be recognized based on the preset file reading path, and performs data extraction on the procurement document to be recognized to obtain the table data to be recognized, which may specifically include: reading the procurement document to be recognized based on the preset file reading path; preprocessing the procurement document to be recognized to obtain the processed procurement document; using the preset file reading function to perform data extraction on the processed procurement document to obtain the original table data; processing the original table data based on the first target format to obtain the corresponding table data to be recognized.
[0064] In a specific embodiment, the document information extraction system uses the pd.read_excel function in the Pandas library to read the specified procurement document to be recognized, and automatically processes line breaks, special characters, and white space characters in the Excel table when reading data, converting the original table data into a DataFrame format to obtain the table data to be recognized, ensuring that all data in the table can be correctly parsed for subsequent processing. If there are cases where multiple cells in the Excel file are merged, the pd.read_excel function will only read the value in the upper left corner of the merged cell, and the other areas will be NaN. Considering that the original cell may be multiple lines, after parsing, it is necessary to prevent the content of these multiple lines from being misaligned and misrecognized as the next line. Therefore, the document information extraction system will mark and replace the line breaks in the cell in advance through appropriate parameter settings and preprocess them into one line to ensure that the data in the merged cell can be correctly read. The final Markdown text can be typeset in the correct format; at the same time, due to the large number of NaNs affecting model recognition, it is also necessary to process the NaNs.
[0065] In this embodiment, the process of preprocessing the procurement document to be recognized to obtain the processed procurement document may specifically include: replacing the first target character in the procurement document to be recognized with a preset character; deleting the second target character in the procurement document to be recognized; processing the null values in the procurement document to be recognized based on a preset processing method to obtain the processed procurement document; where the first target format includes the DataFrame format; the first target character includes line breaks, and the second target character includes white space characters; the preset processing method includes a first processing method and a second processing method, the first processing method includes converting the null values in the procurement document to be recognized into preset values, and the second processing method includes marking the null values in the procurement document to be recognized as missing values. That is to say, the document information extraction system will replace special characters (such as " " and the line break "\n") in the Excel table. For example, the line break is replaced with "00huanhang00" to avoid parsing errors caused by these characters in the subsequent large model parsing process; remove the white space characters in the cell to ensure the neatness of the data; process the null values, convert the empty cells into the default values preset by the document information extraction system or mark them as missing values, so as to ensure the smooth progress of data processing in the subsequent steps.
[0066] It is understandable that in each link of data reading, the file information extraction system will track the whole process through log records. The log file information extraction system is configured to output to the console and the log file, which is convenient for troubleshooting problems later. The log records include the timestamp of each operation, the file path, the name of the form read, the status of success or failure, and the error information, etc. During the operation of the file information extraction system, if an error or a reading failure occurs, the file information extraction system will record the specific error reason in the log, such as "file does not exist" or "invalid form name". These log information can help the maintainers of the file information extraction system quickly troubleshoot problems and improve the stability, reliability and traceability of the file information extraction system.
[0067] Step S13: Use the model call parameter to call the target large model to parse the to-be-recognized table data to obtain the to-be-matched key information; the target large model is a trained large language model.
[0068] In this embodiment, in order to ensure that the target large model can accurately parse the to-be-recognized table data, the file information extraction system first needs to preprocess the read DataFrame data and construct a data format suitable for the input of the large model. Therefore, the process of using the model call parameter to call the target large model to parse the to-be-recognized table data to obtain the to-be-matched key information can specifically include: converting the format of the to-be-recognized table data based on the second target format to obtain the to-be-parsed data; determining all target header data and the target data blocks corresponding to each target header data in the to-be-parsed data based on the structure information in the to-be-parsed data; using the model call parameter to call the target large model to parse each target data block to obtain the corresponding parsing results; converting the format of each parsing result based on the third target format to obtain the to-be-matched key information; wherein, the second target format includes the Markdown format, and the third target format includes the JSON format.
[0069] In a specific implementation, the file information extraction system converts the table data to be recognized in DataFrame format into data to be parsed in Markdown format. The Markdown format can effectively preserve the structural information of the table, such as column names and corresponding column values. Therefore, all target header data in the data to be parsed and the target data blocks corresponding to each target header data can be determined based on the structural information in the data to be parsed, and the target large model is called using the model call parameters to parse each target data block to obtain the corresponding parsing results. By using delimiters such as commas and line breaks, the file information extraction system ensures that each column of data can be correctly mapped to its corresponding column name, thus avoiding the risk of data confusion. For example, assume that an Excel table contains columns such as "Project Name", "Procurement Entity", "Department Name", etc. The file information extraction system will convert these column names into headers in Markdown format and separate the data in each row with commas to form a complete target data block. This target data block will be used as the input to the target large model, and the target large model extracts the key information in the table by analyzing this structured data, that is, the key information to be matched. After the table data to be recognized is converted into data to be parsed, that is, in a format suitable as the input to the target large model, the file information extraction system will interact with the target large model through the API and send the constructed data to be parsed to the large model for parsing. The task of the target large model is to perform intelligent analysis and processing on each target data block in the data to be parsed according to the preset instructions. For example, the target large model can extract key information such as project name, procurement entity, purchase order number, project leader, delivery date, material name, etc. from each target data block and convert this information into structured JSON format key information to be matched. Among them, the API called by the target large model includes a JSON object for setting the request header and sending data. The authentication information in the request header is provided by the authentication token (token) in the configuration file of the file information extraction system, while the JSON object contains the data and questions input by the user. The result of the target large model's parsing will be returned in JSON format, and the file information extraction system will parse the returned key information to be matched to extract the required key-value pair information.
[0070] It is understandable that, in order to improve the reliability of large model parsing, the document information extraction system has introduced an error handling and retry mechanism. In actual operation, the output results of large models may have problems where the format does not meet expectations, such as incomplete JSON data or missing certain key information. For this reason, during the parsing process (i.e., the process of parsing Markdown-formatted information to be parsed into key information to be matched), the document information extraction system will verify whether the results returned by the model meet the expected format, such as checking whether the returned JSON object contains complete key-value pair information. If it is found that the results output by the target large model are incomplete or incorrect, the document information extraction system will automatically retry until the number of retries reaches the preset maximum number of retries or the output results contain complete key-value pair information. Each time a retry is performed, the document information extraction system will fine-tune the input data, such as adjusting the structure or format of the data, so as to improve the success rate of parsing. Based on this retry mechanism, the document information extraction system can ensure to the greatest extent that the parsing results of the large model are accurate, making the data processing process smoother and more automated, and further reducing the need for manual intervention. Once the target large model successfully parses the data, the document information extraction system will convert the parsing results into a standard JSON format to obtain the key information to be matched, and save the obtained key information to be matched to the specified JSON file storage directory. The generated JSON file not only has a clear structure, but also can well adapt to subsequent processing and data transmission. During the process of saving the JSON file, the document information extraction system will also perform further format processing on the data, such as removing redundant white space characters and replacing "00huanhang00" with a line break again, ensuring that the output data format is standardized and easy to read.
[0071] Step S14: Process the key information to be matched using the target template to obtain target key information, and generate a target output file based on the target template, the preset file format, the target output file name, and the target key information, so as to complete the information extraction of the procurement document to be identified.
[0072] In this embodiment, the document information extraction system will load a pre-prepared target template, and then process the key information to be matched using the target template to obtain target key information, which may specifically include: traversing and parsing the key information to be matched based on the preset field name and preset structure to obtain information to be adapted; performing format adaptation on the information to be adapted based on the target template to obtain information to be verified; verifying and cleaning the data type of the information to be verified based on the preset data cleaning rules to obtain target key information.
[0073] In a specific implementation, the file information extraction system extracts specific fields related to the business from the JSON file generated by the aforementioned target large model. These fields usually include project name, procurement entity, purchase order number, project leader, delivery date, material name, etc. During the process of extracting fields, the file information extraction system traverses and parses the data saved in the JSON file according to the preset field names and preset structures. For some business scenarios, such as material procurement, the file information extraction system may need to extract data corresponding to multiple preset field names, that is, the information to be adapted. These data may exist in the form of a list in the JSON file, and the file information extraction system will extract these list data and prepare for further processing.
[0074] After obtaining the information to be adapted, the file information extraction system performs format adaptation and conversion on these information to be adapted based on the target template to ensure that these data can be correctly filled into the target template. For example, for list-type fields (such as the material name list), the file information extraction system will convert it into a comma-separated string format for correct display in Excel cells. Similarly, for numerical data (such as budget amount), the file information extraction system will add currency symbols or adjust the decimal point precision as needed to ensure that the data display meets business requirements. In addition, the file information extraction system will also perform formatting processing on date-type data and convert it into a standard date format (such as "YYYY-MM-DD", year-month-day) for correct display and calculation in Excel. If the data format in some fields does not meet expectations, such as a numerical field containing non-numerical characters, the file information extraction system will record the corresponding error log and skip the processing of that field to avoid errors in subsequent steps.
[0075] It should be noted that after completing the data adaptation and conversion, the file information extraction system will also perform field verification and data cleaning on the converted information to be verified to ensure that the data format and content of all fields meet business requirements. During the field verification process, the file information extraction system will check whether the data type of each field is correct and whether it conforms to the expected format. If it is found that there are problems with the data in some fields, the file information extraction system will correct the data or mark it as abnormal according to the preset rules. The specific cleaning methods (that is, the preset data cleaning rules) include:
[0076] 1. Perform single-valued processing when the data types do not match. For example:
[0077] # Original model output
[0078] "Remarks": ["Emergency procurement"],
[0079] # Result after processing
[0080] "Remarks": "Emergency procurement", # Convert to single value
[0081] 2. If the data of certain fields is empty (such as NaN or empty string), the file information extraction system will fill it with an empty string or other appropriate values to avoid errors in subsequent processing caused by null values. For example:
[0082] "Warranty period": [], # Empty list
[0083] "Brand": ["nan"], # String representation of NaN after pandas conversion
[0084] # Result after processing
[0085] "Warranty period": "", # Empty string
[0086] "Brand": "" # Empty string
[0087] 3. For duplicate data, retain the first element among multiple elements in the list. The first element of the list can be selected by default. For example:
[0088] # Output of the original model
[0089] "Brand": ["A", "A"],
[0090] # Result after processing
[0091] "Brand": "A" # Retain the first element
[0092] 4. Remove the "*" in the key name. For example, correct "Project * Name" to "Project Name".
[0093] In this embodiment, after verifying and cleaning the information to be verified, the target key information can be obtained, and then the target output file can be generated. Generating the target output file based on the target template, preset file format, target output file name, and target key information can specifically include: determining the mapping relationship between each marked cell in the target template and each preset field name based on the key-value mapping method; filling the target key information corresponding to each preset field name into the target template to obtain the filled file; generating the target output file based on the preset file format, target output file name, and the filled file.
[0094] In a specific implementation, specific marked cells have been set in the target template, and these marked cells are used to indicate the filling positions of various types of data. The file information extraction system can process template files in multiple formats, including complex Excel files containing multiple worksheets. When loading the target template, the file information extraction system determines the filling position of each field according to the layout structure of the target template to match the target key information in JSON format obtained in the previous process with these filling positions. When performing the matching, the file information extraction system corresponds the fields in the JSON data to the cells in the Excel template by means of key-value mapping. The key-value mapping table defines the specific positions of each field in the Excel template. For example, the "project name" field in the JSON data may correspond to cell A1 in the Excel template, while the "procurement entity" field corresponds to cell A2. Through this mapping relationship, the file information extraction system can accurately fill the data of each extracted field into the corresponding cell, thus obtaining the filled file.
[0095] It can be understood that in an actual Excel template file, there may be merged cells or complex table structures. The file information extraction system can automatically identify the merged cells and perform correct data filling according to the cell merging rules. If the data of some fields need to be filled into multiple cells, the file information extraction system can also automatically split the data and fill it into the corresponding cells row by row or column by column. For example, when there are multiple records for the material name, the file information extraction system will fill these records into different rows one by one to ensure that each record has an independent display position. In addition, in some cases, there may be duplicate data in the target key information, such as multiple material names being the same but the quantities being different. The file information extraction system can identify these duplicate data and process them according to business rules. For example, the file information extraction system can choose to merge the duplicate material names and accumulate the quantities to ensure that there is no redundant data in the final output file. For cases where there are null values, the file information extraction system will fill in default values or skip these cells according to user configuration to ensure that the data in the filled file is complete and meets the requirements.
[0096] In this embodiment, after obtaining the filled file, the file information extraction system can generate a target output file based on a preset file format, a target output file name, and the filled file. The target output file name can be specified by the administrator through a configuration file, that is, the naming rule of the target output file can be specified by modifying the configuration file. For example, the naming rule can be set to name the file flexibly according to the project name, generation time, etc., to ensure that each generated file has a unique identifier. The path and naming format of the generated file can be flexibly configured to facilitate user management and archiving of the output file. At the same time, the preset file format can also be specified by the configuration file to convert the filled Excel file into other formats according to user needs, such as PDF format, CSV format, XML format, and JSON format, etc. Among them, the PDF format is usually used to generate static and uneditable reports, contracts, or protocol files. The file information extraction system automatically converts the filled Excel file into the PDF format, which can ensure that the data will not be tampered with during the export process and can maintain the fixed layout of the document when printed; the CSV format is a lightweight text format mainly used for data exchange and import, and is often used in data analysis, database import, or file information extraction system integration. The file information extraction system exports the tabular data in Excel as a CSV file, which can facilitate the interaction with other data file information extraction systems. For example, an enterprise may need to export the data of purchase orders in CSV format for the ERP file information extraction system to import and use; while formats such as XML and JSON are usually used for data exchange or archiving with other software file information extraction systems, and are particularly suitable for scenarios that require further data processing and analysis.
[0097] In a specific implementation, to ensure the uniqueness and traceability of each generated file, the file information extraction system supports flexible file naming rules. Users can customize the file naming rules according to business requirements. The common naming format can include timestamp, project name or number, and version number, etc. Specifically, the uniqueness of the file name is ensured by adding the timestamp of the generated file. For example, "Purchase Order_2024-09-06_10:35.xlsx". This naming method not only ensures the uniqueness of the file but also helps users quickly locate the generation time of the file. For files related to certain specific projects, users can choose to include the project name or number as part of the file name. For example: "Project_No. 12345_Contract File.xlsx". In addition, the file information extraction system supports automatically incrementing the version number each time a file is generated, so that users can perform version control and historical tracking. For example, when the same file is generated multiple times, the file information extraction system can name it as: "Contract_Version 1.xlsx", "Contract_Version 2.xlsx", etc., which is convenient for users to manage different versions of the file. Through flexible file naming rules and version control, the file information extraction system can help users more effectively manage and track the output files, avoiding problems such as file overwriting and data loss.
[0098] It should be noted that the filled file and the target output file generated in the foregoing process can be automatically saved to a specified directory according to user configuration. The file information extraction system supports saving files classified by project, date, type, etc. For example, purchase orders can be automatically saved to the "Purchase Order" directory, while financial statements can be saved to the "Financial Statement" directory. The file information extraction system also supports automatically archiving files to a specified cloud storage service, such as AWS S3, Google Drive, or local network storage (NAS, Network Attached Storage), for users to perform long-term storage and backup management. At the same time, the file information extraction system also supports automatically archiving the generated files, especially in the scenario of version control. Users can configure the archiving rules. When the number of files in the same project reaches a certain quantity or time threshold, the file information extraction system will automatically archive the old version files to the historical record directory to save storage space and keep the working directory clean.
[0099] In this embodiment, to ensure the security of the output file, the file information extraction system also provides a file permission management function. The administrator can set different file access permissions according to business requirements. For example, some sensitive contract files may only be accessible to specific users or departments. By integrating an access control list (ACL) or role-based permissions management into the file information extraction system, different access permissions can be set for different users. In addition, the file information extraction system supports encrypting the output file. Especially for the generated PDF files, password protection can be set to prevent unauthorized users from viewing or modifying the file content. For files stored in the cloud, the file information extraction system also supports integration with the encryption transfer protocols of cloud storage services (such as HTTPS (HyperText Transfer Protocol Secure) and SFTP (Secure File Transfer Protocol)) to ensure the security of the file during transmission and storage.
[0100] In this embodiment, during the file generation and saving process, the file information extraction system will also automatically generate detailed log records, covering multiple aspects such as generation time, generated file path, file size, generation status, etc. The generation time records the specific time of each file generation, the generated file path saves the path information of the file so that users can quickly locate the generated file, the file size is used to record the size of the generated file to help users with file management and performance monitoring, and the generation status is the success or failure status of the generated file. If the file generation fails, the log will contain detailed error information for developers or users to troubleshoot. The log file is usually saved in text format and supports automatic rotation and compression archiving according to time to ensure efficient log file management and conservation of storage space. It can be understood that this log file is the operation log of the file information extraction system for locating program exception problems, and a time threshold can be set and the old logs can be automatically archived according to the set time threshold to prevent a single log file from becoming too large. It can be understood that through the log records executed in the foregoing process, all problems encountered during the process of the file information extraction system from obtaining the original procurement file to be recognized to the saving of the target output file are recorded in the final log file, establishing an error detection mechanism. For example, if a data format mismatch occurs when writing data to an Excel template, the file information extraction system will record the specific error information, including the data content, the cell position where it occurred, etc., and stop subsequent processing. This error detection mechanism can help users quickly locate problems and make corresponding adjustments during the next processing.
[0101] In this embodiment, to cope with unexpected interruptions during the file generation process, the file information extraction system designs a fault recovery mechanism. If the file information extraction system fails to process due to some unexpected situations (such as server restart, network connection interruption) in the middle of generating a file, the file information extraction system will automatically save the current progress and continue processing from the previous progress after the fault recovery. For example, when processing an Excel file containing a large amount of data, if the file information extraction system is unexpectedly interrupted after generating some forms, when the file information extraction system resumes, it will continue to generate the data of the remaining forms without having to start the entire file processing again. After each file generation is completed, the file information extraction system will automatically perform file verification to ensure the integrity and consistency of the file. For example, the file information extraction system will perform data consistency checks on the generated Excel file to ensure that all expected fields have been correctly filled. For the target output file, the file information extraction system will verify whether the file has been successfully converted to ensure that there is no format damage or content loss. After the verification is completed, the file information extraction system will record the verification result in the log and mark the file as successfully generated. Based on the above error handling mechanism and fault recovery mechanism, the file information extraction system can improve stability in the actual operating environment, avoid file generation failures caused by network failures or data problems, and improve the robustness of the file information extraction system and the user experience.
[0102] It can be seen that this application uses a large language model trained with powerful semantic understanding capabilities, combines natural language processing and data mining technologies to process the to-be-identified procurement file read based on the preset file reading path, enabling users to realize intelligent extraction and analysis of data without mastering complex programming skills or manually operating Excel. Especially for Excel files with complex data and relationships, the large model can automatically process complex logic across tables and multiple levels through its powerful semantic understanding capabilities, thereby generating a target output file that matches the target template and providing highly accurate information extraction results for the to-be-identified procurement file for users.
[0103] Based on the previous embodiment, it can be known that this application discloses a method for extracting procurement file information based on a large model, which can realize intelligent extraction of procurement file information. Next, the specific process of extracting procurement file information based on a large model will be described.
[0104] See Figure 2 As shown, when the file information extraction system in this embodiment extracts the information contained in the to-be-identified procurement file, it mainly includes six parts: raw data reading, parsing data using the target large model, data format conversion, template filling, target output file generation, and error handling and fault recovery. As Figure 3As shown in the figure, the original data reading extracts data from the original Excel file and performs preliminary processing, providing the basis of the original data for subsequent large model parsing and format conversion. Utilizing the target large model to parse data is the core part of the entire file information extraction process, responsible for analyzing and parsing the data read from the Excel file. Through interaction with the large model, this module can automatically extract the key information in the table and convert it into a structured JSON format for subsequent processing and application. After generating the JSON format output, data format conversion further processes and adapts the JSON data generated by the large model to ensure that these data can match the format requirements in the Excel template. Template filling maps the fields in the parsed and processed JSON data to the cells in the Excel template one by one through key-value mapping and fills them into the preset Excel template, thus realizing automated data filling to generate an output file that meets business requirements. Generating the target output file saves the filled Excel file as a new file and can also provide functions such as format conversion, version control, and logging, ensuring that the generated file can meet the diverse needs of users and the file management is clear and easy to operate. This step is the final output link of the system, ensuring that the data is complete, the format is accurate, and it has scalability. Error handling and fault recovery are mainly implemented based on the system's log file and also include the verification of the output file. Among them, for error handling and fault recovery, three checkpoints can be specifically set, namely converting Excel data to Markdown data, model access (i.e., utilizing the target large model to parse data), and the generation of the final target output file. For each checkpoint, the current code execution status and file status will be saved and recorded.
[0105] It can be seen that this application uses a large language model trained with powerful semantic understanding capabilities, combines natural language processing and data mining technologies to process the to-be-identified procurement file read based on the preset file reading path, enabling users to achieve intelligent data extraction and analysis without mastering complex programming skills or manually operating Excel. Especially for Excel files with complex data and relationships, the large model can automatically process complex logic across tables and multiple levels through its powerful semantic understanding capabilities, thereby generating a target output file that matches the target template and providing highly accurate information extraction results for the to-be-identified procurement file for users.
[0106] See Figure 4 As shown in the figure, this application discloses a procurement file information extraction device based on a large model, which is applied to a file information extraction system and includes:
[0107] A configuration file reading module 11, which is used to read a target configuration file by using a target configuration function to determine a preset file reading path, model calling parameters, a preset file format, and a target output file name;
[0108] A data extraction module 12, which is used to read a procurement file to be recognized based on the preset file reading path, and perform data extraction on the procurement file to be recognized to obtain table data to be recognized; the procurement file to be recognized is a table file;
[0109] A data parsing module 13, which is used to call a target large model by using the model calling parameters to parse the table data to be recognized to obtain key information to be matched; the target large model is a trained large language model;
[0110] A file generation module 14, which is used to process the key information to be matched by using a target template to obtain target key information, and generate a target output file based on the target template, the preset file format, the target output file name, and the target key information, so as to complete the information extraction of the procurement file to be recognized.
[0111] It can be seen that this application uses a trained large language model with powerful semantic understanding ability, combines natural language processing and data mining technologies to process the procurement file to be recognized read based on the preset file reading path, so that users do not need to master complex programming skills or manually operate Excel, and can realize intelligent data extraction and analysis. Especially for Excel files with complex data and relationships, the large model can automatically process complex logics across tables and multiple levels through its powerful semantic understanding ability, so as to generate a target output file that matches the target template, and provide highly accurate information extraction results of the procurement file to be recognized for users.
[0112] In a specific embodiment, the configuration file reading module 11 may specifically include:
[0113] A configuration file reading sub-module, which is used to read a target configuration file by using a target configuration function to determine a preset file reading path, model calling parameters, a preset file format, a procurement file name to be recognized, and a target output file name;
[0114] Correspondingly, the data extraction module 12 may specifically include:
[0115] A table file determination sub-module, which is used to determine all table files corresponding to the preset file reading path;
[0116] A first procurement file determination sub-module, which is used to determine a procurement file to be recognized from all the table files based on the procurement file name to be recognized, and read the procurement file to be recognized.
[0117] In a specific embodiment, the first purchase document determination sub-module may specifically include:
[0118] The purchase document determination unit is configured to determine the first table file under the preset file reading path as the to-be-identified purchase document if the name of the to-be-identified purchase document is empty.
[0119] In a specific embodiment, the data extraction module 12 may specifically include:
[0120] The second purchase document determination sub-module is configured to read the to-be-identified purchase document based on the preset file reading path;
[0121] The purchase document preprocessing sub-module is configured to preprocess the to-be-identified purchase document to obtain a processed purchase document;
[0122] The table data extraction sub-module is configured to extract data from the processed purchase document by using a preset file reading function to obtain original table data;
[0123] The table data processing sub-module is configured to process the original table data based on a first target format to obtain corresponding to-be-identified table data;
[0124] Among them, the purchase document preprocessing sub-module may specifically include:
[0125] The character replacement unit is configured to replace a first target character in the to-be-identified purchase document with a preset character;
[0126] The character deletion unit is configured to delete a second target character in the to-be-identified purchase document;
[0127] The null value processing unit is configured to process the null values in the to-be-identified purchase document based on a preset processing method to obtain a processed purchase document;
[0128] Among them, the first target format includes the DataFrame format; the first target character includes a line break character, the second target character includes a blank character; the preset processing method includes a first processing method and a second processing method, the first processing method includes converting the null values in the to-be-identified purchase document into preset values, and the second processing method includes marking the null values in the to-be-identified purchase document as missing values.
[0129] In a specific embodiment, the data parsing module 13 may specifically include:
[0130] The first format conversion sub-module is configured to perform format conversion on the to-be-identified table data based on a second target format to obtain to-be-parsed data;
[0131] A target data block determination sub-module, configured to determine all target header data in the data to be parsed and target data blocks corresponding to each of the target header data based on the structure information in the data to be parsed;
[0132] A parsing result acquisition sub-module, configured to call a target large model by using the model call parameter to parse each of the target data blocks to obtain corresponding parsing results;
[0133] A second format conversion sub-module, configured to perform format conversion on each of the parsing results based on a third target format to obtain key information to be matched;
[0134] Wherein, the second target format includes Markdown format, and the third target format includes JSON format.
[0135] In a specific embodiment, the file generation module 14 may specifically include:
[0136] A to-be-adapted information determination sub-module, configured to traverse and parse the key information to be matched based on a preset field name and a preset structure to obtain to-be-adapted information;
[0137] A format adaptation sub-module, configured to perform format adaptation on the to-be-adapted information based on a target template to obtain information to be verified;
[0138] A data cleaning sub-module, configured to verify and clean the data type of the information to be verified based on a preset data cleaning rule to obtain target key information.
[0139] In a specific embodiment, the file generation module 14 may specifically include:
[0140] A mapping relationship determination sub-module, configured to determine a mapping relationship between each marked cell in the target template and each of the preset field names based on a key-value mapping method;
[0141] A template filling sub-module, configured to fill the target key information corresponding to each of the preset field names into the target template based on the mapping relationship to obtain a filled file;
[0142] A file generation sub-module, configured to generate a target output file based on the preset file format, the target output file name, and the filled file.
[0143] Furthermore, an embodiment of the present application also discloses an electronic device, Figure 5 It is a structural diagram of an electronic device 20 shown according to an exemplary embodiment, and the content in the figure cannot be considered as any limitation on the usage scope of the present application.
[0144] Figure 5 This is a schematic structural diagram of an electronic device 20 provided by an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the procurement document information extraction method based on a large model disclosed in any of the foregoing embodiments. Additionally, the electronic device 20 in this embodiment may specifically be an electronic computer.
[0145] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of the present application, and no specific limitation is imposed on it here; the input / output interface 25 is used to obtain external input data or output data to the outside, and its specific interface type can be selected according to specific application requirements, and no specific limitation is made here.
[0146] In addition, as a carrier for resource storage, the memory 22 can be a read-only memory, a random access memory, a magnetic disk, or an optical disc, etc., and the resources stored thereon can include an operation file information extraction system 221, a computer program 222, etc., and the storage method can be temporary storage or permanent storage.
[0147] Among them, the operation file information extraction system 221 is used to manage and control each hardware device and the computer program 222 on the electronic device 20, and it can be Windows Server, Netware, Unix, Linux, etc. The computer program 222 may further include a computer program capable of completing other specific tasks in addition to the computer program capable of implementing the procurement document information extraction method based on a large model executed by the electronic device 20 disclosed in any of the foregoing embodiments.
[0148] Furthermore, the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the procurement document information extraction method based on a large model disclosed above. For the specific steps of this method, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details are not repeated here.
[0149] In this specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For related parts, reference can be made to the description in the method section.
[0150] Those skilled in the art can further realize that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0151] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be directly implemented by hardware, software modules executed by a processor, or a combination of both. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.
[0152] Finally, it should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.
[0153] The technical solutions provided in this application have been introduced in detail above. Specific examples have been used herein to elaborate on the principles and implementation manners of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, based on the idea of this application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to this application.
Claims
1. A method for extracting procurement document information based on a large model, characterized in that, Applied to a file information extraction system, including: Using a target configuration function to read a target configuration file to determine a preset file reading path, model call parameters, a preset file format, and a target output file name; Reading a to-be-identified procurement file based on the preset file reading path, and performing data extraction on the to-be-identified procurement file to obtain to-be-identified table data; the to-be-identified procurement file is a table file; Using the model call parameters to call a target large model to parse the to-be-identified table data to obtain to-be-matched key information; the target large model is a trained large language model; Using a target template to process the to-be-matched key information to obtain target key information, and generating a target output file based on the target template, the preset file format, the target output file name, and the target key information to complete the information extraction of the to-be-identified procurement file.
2. The procurement document information extraction method based on a large model according to claim 1, wherein The using a target configuration function to read a target configuration file to determine a preset file reading path, model call parameters, a preset file format, and a target output file name includes: Using a target configuration function to read a target configuration file to determine a preset file reading path, model call parameters, a preset file format, a to-be-identified procurement file name, and a target output file name; Correspondingly, the reading a to-be-identified procurement file based on the preset file reading path includes: Determining all table files corresponding to the preset file reading path; Determining the to-be-identified procurement file from all the table files based on the to-be-identified procurement file name, and reading the to-be-identified procurement file.
3. The procurement document information extraction method based on a large model according to claim 2, wherein, The determining the to-be-identified procurement file from all the table files based on the to-be-identified procurement file name includes: If the to-be-identified procurement file name is empty, determining the first table file under the preset file reading path as the to-be-identified procurement file.
4. The procurement document information extraction method based on a large model according to claim 1, wherein The reading a to-be-identified procurement file based on the preset file reading path, and performing data extraction on the to-be-identified procurement file to obtain to-be-identified table data includes: Reading the to-be-identified procurement file based on the preset file reading path; Performing preprocessing on the to-be-identified procurement file to obtain a processed procurement file; Using a preset file reading function to perform data extraction on the processed procurement file to obtain original table data; Processing the original table data based on a first target format to obtain corresponding to-be-identified table data; Wherein, the performing preprocessing on the to-be-identified procurement file to obtain a processed procurement file includes: Replacing a first target character in the to-be-identified procurement file with a preset character; Deleting a second target character in the to-be-identified procurement file; Processing null values in the to-be-identified procurement file based on a preset processing method to obtain a processed procurement file; Among them, the first target format includes the DataFrame format; the first target character includes the line break character, and the second target character includes the whitespace character; the preset processing methods include the first processing method and the second processing method. The first processing method includes converting the null values in the to-be-recognized procurement document into preset values, and the second processing method includes marking the null values in the to-be-recognized procurement document as missing values.
5. The procurement document information extraction method based on a large model according to claim 1, wherein The step of using the model call parameters to call the target large model to parse the to-be-recognized tabular data to obtain the to-be-matched key information includes: Performing format conversion on the to-be-recognized tabular data based on the second target format to obtain the to-be-parsed data; Determining all target header data in the to-be-parsed data and the target data blocks corresponding to each of the target header data based on the structure information in the to-be-parsed data; Using the model call parameters to call the target large model to parse each of the target data blocks to obtain the corresponding parsing results; Performing format conversion on each of the parsing results based on the third target format to obtain the to-be-matched key information; Among them, the second target format includes the Markdown format, and the third target format includes the JSON format.
6. The method for extracting procurement document information based on a large model according to any one of claims 1 to 5, characterized in that The step of using the target template to process the to-be-matched key information to obtain the target key information includes: Traversing and parsing the to-be-matched key information based on the preset field names and preset structure to obtain the to-be-adapted information; Performing format adaptation on the to-be-adapted information based on the target template to obtain the to-be-verified information; Verifying and cleaning the data types of the to-be-verified information based on the preset data cleaning rules to obtain the target key information.
7. The method for extracting procurement document information based on a large model according to claim 6, wherein The step of generating the target output file based on the target template, the preset file format, the target output file name, and the target key information includes: Determining the mapping relationship between each marked cell in the target template and each of the preset field names based on the key-value mapping method; Filling the target key information corresponding to each of the preset field names into the target template based on the mapping relationship to obtain the filled file; Generating the target output file based on the preset file format, the target output file name, and the filled file.
8. An information extraction device for procurement documents based on a large model, characterized in that, Applied to a file information extraction system, it includes: A configuration file reading module, which is used to read the target configuration file by using the target configuration function to determine the preset file reading path, model call parameters, preset file format, and target output file name; A data extraction module, which is used to read the to-be-recognized procurement document based on the preset file reading path and perform data extraction on the to-be-recognized procurement document to obtain the to-be-recognized tabular data; the to-be-recognized procurement document is a tabular file; A data parsing module, which is used to use the model call parameters to call the target large model to parse the to-be-recognized tabular data to obtain the to-be-matched key information; the target large model is a trained large language model; A file generation module, configured to process the to-be-matched key information by using a target template to obtain target key information, and generate a target output file based on the target template, the preset file format, the target output file name, and the target key information, so as to complete the information extraction of the to-be-identified procurement document.
9. An electronic device, characterized in that, Comprising: A memory, configured to store a computer program; A processor, configured to execute the computer program to implement the procurement document information extraction method based on a large model according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, For storing a computer program, wherein when the computer program is executed by a processor, the procurement document information extraction method based on a large model according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Format conversion method and device for heterogeneous data and medium
CN121092506A