Job fund collection information extraction method and system based on large language model
Through the information extraction method based on the large language model and the HBaseX-Pack database, the efficiency and security issues in the data exchange between the tax system and the union system are solved, and efficient and secure union funding data processing and analysis are achieved.
Patent Information
- Application Number
- CN202510780213.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-06-12
AI Technical Summary
In the prior art, the data exchange efficiency between the tax system and the trade union system is low, and data consistency and security are difficult to ensure, which affects the quality and efficiency of trade union fund collection work.
Using a large language model-based information extraction method, multi-level prompt words and noise node judgment are designed, combined with HBaseX-Pack distributed database and digital envelope technology, efficient data classification and secure transmission are achieved.
It improves data processing efficiency, ensures data consistency and security, reduces data loss and error, can quickly identify underpaid units and push reminders, improving the quality and efficiency of union fund collection work.
Smart Images

Figure CN120296799A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of information extraction, and relates to a method and system for extracting information on the collection of trade union funds based on a large language model. Background Art
[0002] Tax registration units declare and pay trade union funds through the tax collection system. After reconciliation among the trade union, tax authorities, and banks at the end of each month, the transfer of collected funds is carried out. Currently, the data on the collection of trade union funds by tax authorities mainly relies on the communication between tax and trade union departments in various cities through the regular transmission of excel sheets. In this way, the transmission of collected data information is not timely, and the systematic and comprehensive analysis and utilization of the collected data cannot be carried out, which affects the quality and efficiency of the work on the collection of trade union funds; problems in application: the volume of data exchanged between the trade union system and the tax system is large, and it is impossible to efficiently and quickly process big data; the data consistency between the trade union system and the tax system, and too many intermediate links are likely to cause data loss or errors; the data transmission security between the trade union system and the tax system, data security is crucial, and if the data is obtained by illegal elements, serious consequences are likely to occur. Summary of the Invention
[0003] To solve the technical problems existing in the above background art, on the one hand, the present invention provides a method for extracting information on the collection of trade union funds based on a large language model, including: Obtain the fund data file and store it using database tools, and extract the fund text from the fund data file using extraction tools; for the fund data file in the form of transmitting an excel sheet, before extracting the fund text information, first convert the document in the excel sheet into a LA-DOM tree structure, then judge the noise nodes according to the attributes of each node of the LA-DOM tree and the statistical information of its subtree, and finally realize the extraction of the main body information; Design the first-level prompt words, input the first-level prompt words and the main body information into a preset large language model, and the large language model outputs the key information of the fund text under the guidance of the first-level prompt words; For each fund in the key information of the fund text, design the second-level prompt words, where the second-level prompt words include the fund name and the corresponding fund category, input the second-level prompt words into the large language model, and the large language model outputs the fund category coding result of each fund under the guidance of the second-level prompt words.
[0004] Further, for each fund in the key information of the fund text, design the second-level prompt words, where the second-level prompt words include the fund name and the corresponding fund category, and input the second-level prompt words into the large language model.
[0005] Further, the large language model outputs the funding category coding results for each fund under the guidance of the second-level prompt words, including: Construct a mapping relationship correspondence table between the fund name and the fund category, between the fund category and the code, and between the fund name and the code; design second-level prompt words for each fund in the key information of the fund text based on the mapping relationship correspondence table; input the second-level prompt words into the large language model, and the large language model outputs the funding category coding results for each fund under the guidance of the second-level prompt words.
[0006] Further, for the funds for which the large language model cannot output the funding category coding results, repeatedly design broader third-level prompt words and input them into the large language model until the large language model can output the funding category coding results for the remaining funds under the guidance of the third-level prompt words; post-process the funding category coding results for each fund output by the large language; integrate the key information of the fund text and the funding category coding results for each fund and output.
[0007] In a second aspect, the present invention provides a trade union fund collection information extraction system based on a large language model, including: A fund data acquisition module, configured to: acquire a fund data file and store it using a database tool, and extract fund text from the fund data file using an extraction tool; A key information extraction module, configured to: design first-level prompt words, input the first-level prompt words and the main body information into a preset large language model, and the large language model outputs the key information of the fund text under the guidance of the first-level prompt words; A data classification module, configured to: for each fund in the key information of the fund text, design second-level prompt words, where the second-level prompt words include the fund name and the corresponding fund category, input the second-level prompt words into the large language model, and the large language model outputs the funding category coding results for each fund under the guidance of the second-level prompt words.
[0008] The beneficial effects of the present invention are: 1. The present invention provides a method for extracting information on the collection of trade union funds based on a large language model. By obtaining the fund data file and storing it using database tools, the extraction tool is used to extract the fund text from the fund data file. For the storage of fund data, excellent database tools are selected, and the HBaseX-Pack distributed columnar database is chosen. Partition operations are performed on the massive data, and extensive indexes are established to solve the problem of the system freezing due to the excessive amount of data pushed at one time. The pushed data is put into the message middleware for peak shaving processing. The data is stored in the log file without any processing to prevent data loss caused by data processing. The system can perform the warehousing operation according to the concurrency configured based on the server performance.
[0009] 2. The present invention uses a database server to communicate and couple to a database storing multiple data records. The front end retrieves metadata through the central metadata retrieval interface. The central metadata retrieval interface finds the appropriate metadata plugin through the metadata plugin management to obtain the corresponding technical metadata, ensuring the consistency of metadata and the timeliness of obtaining metadata changes. This solution reduces the information error caused by inconsistent metadata and the information lag caused by untimely changes and inability to obtain. It unifies the caliber problem to ensure the consistency of the data caliber among the tax system, the actual business of the trade union, and this system. Otherwise, it may lead to data loss due to inconsistent calibers.
[0010] 3. The present invention designs the first-level prompt words, inputs the first-level prompt words and the main body information into a preset large language model. The large language model outputs the key information of the fund text under the guidance of the first-level prompt words and classifies it to obtain the fund information, which can effectively analyze the underpaid units and the non-paid units. It is necessary to determine the payment cycle according to the nature of the unit, and then calculate the payable amount based on the payment amount, payment date, and the total salary of the unit. After comparison and analysis, the underpaid units, underpaid amounts, and non-paid units are obtained, and the list is pushed to the tax system for reminder.
[0011] The advantages of the additional aspects of the present invention will be partially given in the following description, partially will become obvious from the following description, or will be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The specification drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention.
[0013] Figure 1 It is a flowchart of a method for extracting information on the collection of trade union funds based on a large language model of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0014] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0015] It should be noted that the following detailed description is illustrative and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used in this embodiment have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention pertains.
[0016] It should be noted that the terms used herein are merely for describing specific embodiments and are not intended to limit the exemplary embodiments of the present invention. As used herein, unless the context clearly dictates otherwise, the singular form is also intended to include the plural form. Additionally, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0017] In the present invention, terms such as "upper", "lower", "left", "right", "front", "rear", "vertical", "horizontal", "side", "bottom", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings. They are only relationship terms determined for the convenience of describing the structural relationship of each component or element of the present invention and do not specifically refer to any component or element of the present invention and should not be construed as limiting the present invention.
[0018] In the present invention, terms such as "fixed connection", "connected", "connected to" should be understood in a broad sense, which may mean a fixed connection, an integral connection or a detachable connection; it may be directly connected or indirectly connected through an intermediate medium. For those skilled in the relevant scientific research or technology in this field, the specific meaning of the above terms in the present invention can be determined according to specific circumstances and should not be construed as limiting the present invention.
[0019] Embodiment 1, as Figure 1 shown, this embodiment provides a method for extracting trade union funds collection information based on a large language model, including: Obtaining a funds data file and storing it using a database tool, and extracting funds text from the funds data file using an extraction tool; Designing a first-level prompt, and inputting the first-level prompt and main information into a preset large language model, and the large language model outputs key information of the funds text under the guidance of the first-level prompt; For each fund in the key information of the fund text, design a second-level prompt word, where the second-level prompt word includes the fund name and the corresponding fund category, and input the second-level prompt word into the large language model, and the large language model outputs the fund category coding result of each fund under the guidance of the second-level prompt word; for each fund in the key information of the fund text, design a second-level prompt word, where the second-level prompt word includes the fund name and the corresponding fund category, and input the second-level prompt word into the large language model, and the large language model outputs the fund category coding result of each fund under the guidance of the second-level prompt word, including: Construct a mapping relationship correspondence table for the fund name to the fund category, the fund category to the code, and the fund name to the code; design a second-level prompt word for each fund in the key information of the fund text based on the mapping relationship correspondence table; input the second-level prompt word into the large language model, and the large language model outputs the fund category coding result of each fund under the guidance of the second-level prompt word; for the funds for which the large language model cannot output the fund category coding result, repeat to design a more general third-level prompt word and input it into the large language model until the large language model can output the fund category coding result of the remaining funds under the guidance of the third-level prompt word; post-process the fund category coding result of each fund output by the large language model; Integrate the key information of the fund text and the fund category coding result of each fund and output.
[0020] Among them, obtain the fund data file and store it using database tools, and use the extraction tool to extract the fund text from the fund data file specifically as follows: For the fund data file, the form of passing an excel table is usually used for communication. The key to extracting information from the excel table lies in how to efficiently extract a large amount of data while ensuring the timeliness and effectiveness of data transmission. A noise judgment and main body information judgment are established, and noise information and main body information are identified based on statistical data, and information extraction is performed through the feature marking method of noise and main body information.
[0021] The fund data file usually contains parts such as title, time, region, regional code, fund expenditure amount, ID, etc., and the rest are irrelevant information, such as project introduction, precautions, etc., which are called "noise information". Noise information is non-essential information. From the perspective of information extraction, it is information without extraction value. Before extracting the fund text information, first convert the document in the excel table into a LA-DOM tree structure, then judge the noise nodes according to the attributes of each node in the LA-DOM tree and the statistical information of its subtree, and finally realize the extraction of the main body information. To realize the conversion of the document in the excel table to LA-DOM, mainly two aspects of processing are required: node generation and node attribute addition.
[0022] The node generation process classifies according to the tag type, generates LA-DOM nodes from node tags, and converts attribute tags into attribute values. The content of the document in the Excel sheet to be processed is passed into the conversion program as a string. The conversion program scans the string from start to end. When encountering the start tag of a node tag, a node is generated and added to the child nodes of the current node. At the same time, the current node pointer moves down, and the newly generated node becomes the current node. When encountering the end tag of a node tag, the current node pointer traces back to point to the parent node of the current node; when encountering an attribute tag, it is used as a node attribute and added to the current node together with other attributes of the node tag.
[0023] The attribute stack is generated during the generation process of LA-DOM tree nodes and is destroyed when the LA-DOM tree is established. Its function is to record the attributes of nodes at each layer, and finally assign the attributes to text nodes as the basis for subsequent analysis work. The generation rule of the attribute stack is as follows: when encountering the start tag of a node tag, while generating a hierarchical node, an attribute element is generated. This attribute element first copies the attributes at the top of the attribute stack, then adds the current new attribute to the attribute element, and finally the attribute element is pushed onto the stack; when encountering the end tag of a node tag, when the current node pointer traces back, the top element of the attribute stack also pops out at the same time; when encountering a text node, the top element of the attribute stack is used as an attribute and assigned to the text node, and it does not pop out. According to this rule, when the LA-DOM tree is established, all text nodes are added with attribute values according to the hierarchical relationship, and at this time the attribute stack is empty.
[0024] The algorithm flow for converting the document in the Excel sheet into an LA-DOM tree: if (start tag) { if (it is a node tag) { generate a new node; extract the tag attributes and push them onto the attribute stack; add this node as a subtree of the current node in the DOM tree; the current node points to the new node}; else if (it is an attribute tag) { record the attribute and push it onto the attribute stack}; else { irrelevant tag, skip directly}; else if (it is an end tag) { if (it is a node tag) { pop the top element of the attribute stack; trace back to find the matching start tag; if (the matching start tag is found) { the tag is closed; set it as the current node} else { extra tag, skip directly}}; else { it is not a node tag, skip directly}}; else { / / text information; generate a text node; add the attributes at the top of the attribute stack to the text node; add the leaf node as a child node of the current node in the DOM tree}.
[0025] After completing the conversion of the document in the Excel sheet into an LA-DOM tree, it is necessary to mark the noise nodes, which is also the core part of the extraction of funding information and directly affects the accuracy of the information extraction result.
[0026] Statistics such as the number of direct non-linked leaf nodes, the number of characters in direct non-linked leaf nodes, the number of direct linked leaf nodes, the number of characters in direct linked leaf nodes, the total number of linked child nodes, the total number of characters in linked child nodes, the total number of non-linked child nodes, and the total number of characters in non-linked child nodes are added to the nodes of the LA-DOM tree. Using these statistics as the judgment criteria, noise nodes are marked.
[0027] The statistical information in the nodes is named as follows: DLN: The number of direct linked child nodes. DLT: The number of characters in direct linked child nodes. DUN: The number of direct non-linked child nodes. DUT: Direct non-linked child nodes. TLN: The total number of linked child nodes. TLT: The total number of characters in linked child nodes. TUN: The total number of non-linked child nodes. TUT: The total number of characters in non-linked child nodes. Text containing links is usually noise information in news web pages, while a large amount of non-linked text is usually funding text information. Among them, DLN, DLT, TLN, and TLT are collectively called noise eigenvalue, and DUN, DUT, TUN, and TUT are collectively called information eigenvalue. The larger the information eigenvalue of a node, the more likely it is to be the main information of the news; on the contrary, if the noise eigenvalue is larger, the more likely it is to be noise information.
[0028] The algorithm for establishing node statistical information is as follows: postTraversal(Node currentNode){ if (has child nodes) { Node child = currentNode.getFirstChild() postTraversal(child); while (child.hasNextSibling) { child = child.getNextSibling(); postTraversal(child);} The noise node judgment rule is designed as follows: (1) DUN = 0 DUT = 0 When the number of direct non-linked child nodes or the number of characters in non-linked child nodes is 0, that is, there are no non-linked text nodes under the current node, it is judged as a noise node.
[0029] (2) DLT >= DUT When the number of characters in direct linked child nodes is greater than the number of characters in direct non-linked child nodes, it is judged as a noise node because in the main part of the funding text, most of the text is non-linked text.
[0030] (3) DLN * a > DUN When the number of direct linked child nodes is greater than the number of direct non-linked child nodes, it is judged as a noise node. a is an adjustment factor, and according to statistics, when the value is 2.6, the judgment is more accurate.
[0031] (4) When the ratio of the number of characters of the directly linked child nodes (DLT / DLN) to the number of directly linked child nodes is greater than the ratio of the number of characters of the directly non-linked child nodes (DUT / DUN) to the number of directly non-linked child nodes, it is determined as a noise node.
[0032] (5) When DUT / DUN < threshold, and the value of DUT / DUN is less than a certain threshold, it is determined as noise. According to previous statistics, a better effect is achieved when the threshold takes the value of 9.7.
[0033] For the pictures of funds and the pictures converted from the scanned copies in PDF files, it is a schematic flowchart of the process for identifying the fund text in the pictures of funds and the pictures converted from the scanned copies in PDF files in the embodiments of the present application. This method at least includes the following steps: Preprocess the picture. First, use non-local means denoising to reduce the noise in the picture and retain the clarity of the text. Subsequently, use the following formula to enhance the contrast and brightness of the picture:
[0034] Where, is the pixel value of the picture after contrast and brightness enhancement processing, is the pixel value of the original picture, and are the minimum and maximum pixel values of the picture respectively, and α and β are adjustment parameters used to adjust the contrast and brightness.
[0035] Local Binary Pattern (LBP) is a method for describing the local texture features of an image. For each pixel, by comparing its gray value with the gray values of surrounding pixels, a binary number is generated, and then the LBP value is obtained.
[0036] For each text region, a method based on pixel intensity and shape features can be used to extract possible characters, which is achieved through connected component analysis or pixel-based operations. The specific methods include: connecting the pixels in the text region into character shapes, and determining the boundaries of the characters according to the relative positions and pixel intensities of the pixels. Extract the characters by detecting specific patterns or contours of the character shapes. Perform simple pattern matching or rule matching on the extracted characters. For example, some basic rules can be defined to identify numbers, letters, or specific symbols. These rules can be determined based on the shape, size, and pixel distribution of the characters to identify the characters. Reconstruct the recognized and extracted characters into complete text lines or paragraphs according to their layout order in the picture. This can be completed by connecting the characters and sorting them according to their positions in the picture. Finally, combine the reconstructed text lines or paragraphs into the fund text of the entire picture. The fund text can be organized in the format of text lines or paragraphs and output in the form of a text string.
[0037] For the storage of funds data, an excellent database tool is selected, and the HBaseX-Pack distributed columnar database is chosen; the massive data is partitioned, a wide range of indexes are established, and a caching mechanism is built; the data is sampled, data mining is carried out, and the massive data is associated and stored; the HBaseX-Pack provides high-performance random read and write operations externally; every day, the data of the previous day is regularly aggregated and synchronized and archived to other low-performance but low-cost databases. The selected HBaseX-Pack is a low-cost one-stop data processing platform built based on HBase and the HBase ecosystem, and the HBaseX-Pack supports HBase APIs (including RestServer and ThriftServer), relational PhoenixSQL, time series OpenTSDB, full-text Solr, spatio-temporal GeoMesa, graph HGraph, and analysis Spark on HBase. The HBaseX-Pack can achieve a full-process closed loop from data processing, storage to analysis. To solve the problem that the system freezes due to the excessive amount of data pushed at one time, the pushed data is put into the message middleware for peak shaving. The data is stored in the log file without any processing to prevent data loss caused by data processing. The system can configure the concurrency according to the server performance for the inbound operation.
[0038] The database server is communicatively coupled to a database that stores multiple data records, and these data records can be accessed by multiple data access systems communicatively coupled to the database server. The front end retrieves metadata through the central metadata retrieval interface. The central metadata retrieval interface obtains the corresponding technical metadata by finding the appropriate metadata plug-in through the metadata plug-in management. The central metadata retrieval interface obtains management metadata and business metadata by querying the metadata management system DB. The central metadata retrieval interface summarizes the metadata and returns it to the front end. Based on the central metadata retrieval interface, metadata retrieval is carried out, data assembly is performed in the retrieval interface, and the assembled metadata is returned to the front end for rendering, ensuring the consistency of metadata and the timeliness of obtaining metadata changes. This solution reduces the information error caused by metadata inconsistency and the information lag caused by untimely changes and inability to obtain. It unifies the data caliber issue to ensure the consistency of the data calibers among the tax system, the actual business of the labor union, and this system. Otherwise, it may lead to data loss due to inconsistent calibers.
[0039] Regarding the problem of missed pushes, by counting the total number of pushed data, the total number of inbound data, and the total of tax-pushed data, the situation of missed pushes can be avoided when the data is consistent. If a missed push occurs, it is necessary to quickly locate the specific data of the missed push. The data can be hashed by the payment amount + payment unit + payment year and then compared with the pushed data to complete the quick location.
[0040] Analyze the units with underpayment and non-payment. It is necessary to determine the payment cycle according to the nature of the unit, and then calculate the payable amount based on the payment amount, payment date and the total salary of the unit. After comparison and analysis, obtain the units with underpayment, the amount of underpayment and the non-paying units, and push the list to the tax system for reminder.
[0041] Many of the information transmitted in the business system is sensitive. Therefore, it is necessary to prevent it from being illegally obtained and tampered with during online transmission, so transmission encryption must be carried out. The secure encryption of information transmission can be achieved by digital envelope technology. Digital envelope adopts two cryptographic systems: single-key cryptosystem and public-key cryptosystem. Digital envelope technology means that the information sender first encrypts the information using symmetric cryptography, then encrypts the symmetric key using the public key of the recipient, and then sends it to the recipient. When the information recipient wants to decrypt the information, they must first decrypt the symmetric key with their own private key before they can use the symmetric key to decrypt the obtained information.
[0042] To ensure the security of the transmitted information, the symmetric key used for each information transmission is different. Digital envelope technology was developed to solve the problem of changing keys each time. Digital envelope technology combines the advantages of symmetric encryption technology and public key technology, overcomes the difficulties of symmetric key distribution and the long encryption time of public key encryption, and uses two levels of encryption to obtain the flexibility of public key technology and the efficiency of symmetric key technology. After adopting digital envelope technology, even if the encrypted file is illegally intercepted by others, it is impossible to decrypt the file.
[0043] Embodiment 2 This embodiment provides a trade union fund collection information extraction system based on a large language model, including: A fund data acquisition module, configured to: acquire a fund data file and store it using database tools, and extract fund text from the fund data file using extraction tools; A key information extraction module, configured to: design a first-level prompt word, input the first-level prompt word and the main information into a preset large language model, and the large language model outputs the key information of the fund text under the guidance of the first-level prompt word; A data classification module, configured to: for each fund in the key information of the fund text, design a second-level prompt word, the second-level prompt word includes the fund name and the corresponding fund category, input the second-level prompt word into the large language model, and the large language model outputs the fund category coding result of each fund under the guidance of the second-level prompt word.
[0044] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for extracting information on the collection of trade union funds based on a large language model, characterized in that, Including: Obtain the funding data file and store it using database tools, and extract the funding text from the funding data file using extraction tools; For the funding data file, adopt the form of passing an excel table. Before extracting the funding text information, first convert the document in the excel table into a LA-DOM tree structure, then judge the noise nodes according to the attributes of each node in the LA-DOM tree and the statistical information of its subtree, and finally realize the extraction of the main information; Design the first-level prompt words, input the first-level prompt words and the main information into a preset large language model, and the large language model outputs the key information of the funding text under the guidance of the first-level prompt words; For each funding in the key information of the funding text, design the second-level prompt words. The second-level prompt words include the funding name and the corresponding funding category, and input the second-level prompt words into the large language model. The large language model outputs the funding category coding result of each funding under the guidance of the second-level prompt words.
2. The method for extracting trade union fund collection information based on a large language model according to claim 1, wherein For each funding in the key information of the funding text, design the second-level prompt words. The second-level prompt words include the funding name and the corresponding funding category, and input the second-level prompt words into the large language model.
3. The method for extracting trade union fund collection information based on a large language model according to claim 1, wherein, The large language model outputs the funding category coding result of each funding under the guidance of the second-level prompt words, including: Construct a mapping relationship correspondence table for the funding name to the funding category, the funding category to the coding, and the funding name to the coding; design the second-level prompt words for each funding in the key information of the funding text based on the mapping relationship correspondence table; input the second-level prompt words into the large language model, and the large language model outputs the funding category coding result of each funding under the guidance of the second-level prompt words.
4. The method for extracting trade union fund collection information based on a large language model according to claim 1, wherein, For the funds for which the large language model cannot output the funding category coding result, repeatedly design more general third-level prompt words and input them into the large language model until the large language model can output the funding category coding result of the remaining funds under the guidance of the third-level prompt words; post-process the funding category coding result of each funding output by the large language; Integrate the key information of the funding text and the funding category coding result of each funding and output.
5. The method for extracting information on the collection of trade union funds based on a large language model according to claim 1, wherein, In the node generation process, according to the label classification, generate LA-DOM nodes from the node labels, and convert the attribute labels into attribute values; the content of the document in the excel table to be processed is passed into the conversion program as a string. The conversion program scans the string from beginning to end. When encountering the start tag of the node label, generate a node and add it to the child nodes of the current node. At the same time, move the current node pointer down, and use the newly generated node as the current node; when encountering the end tag of the node label, the current node pointer backtracks and points to the parent node of the current node; when encountering the attribute label, it is used as a node attribute and added to the current node together with other attributes of the node label.
6. The method for extracting trade union fund collection information based on a large language model according to claim 5, wherein, The attribute stack is generated during the generation process of LA-DOM tree nodes and is destroyed when the LA-DOM tree is established. Its function is to record the attributes of nodes at each layer and finally assign the attributes to text nodes as the basis for subsequent analysis. The generation rule of the attribute stack is as follows: when encountering the start tag of the node tag, an attribute element is generated while generating the hierarchical node. The attribute element first copies the top attribute of the attribute stack, then adds the current new attribute to the attribute element, and finally pushes the attribute element into the stack. When encountering the end tag of the node tag, the top element of the attribute stack is also popped out when the current node pointer is backtracked. When encountering a text node, the top element of the attribute stack is assigned to the text node as an attribute and is not popped out. According to this rule, when the LA-DOM tree is established, all text nodes are added with attribute values according to the hierarchical relationship, and the attribute stack is empty at this time.
7. A trade union fund collection information extraction system based on a large language model, characterized in that, include: The funding data acquisition module is configured to: acquire funding data files and store them using a database tool, and extract funding text from the funding data files using an extraction tool; The key information extraction module is configured to: design a first-level prompt word, input the first-level prompt word and main information into a preset large language model, and the large language model outputs the key information of the expense text under the guidance of the first-level prompt word; The data classification module is configured as follows: for each expense in the key information of the expense text, a second-level prompt word is designed, wherein the second-level prompt word includes the expense name and the corresponding expense category, and the second-level prompt word is input into the large language model, and the large language model outputs the expense category coding result of each expense under the guidance of the second-level prompt word.
Citation Information
Patent Citations
Noise data cleaning method based on semantic ontology
CN101986296A
Method and system for generating Target Link data dictionary hierarchical tree
CN104281604A
List downloading template creating method and device, terminal and readable storage medium
CN110019970A
Webpage denoising method
CN112347353A
Web content information extraction method based on DOM (Document Object Model) tree and row-column segmentation
CN113158626A