Method for extracting bidding and tendering data information by using generative large model

The extraction of bidding data information through a generative large model solves the problems of data structure information loss and manual labeling dependence in the existing methods, and achieves high accuracy and low cost information extraction effects.

CN119917756AInactive Publication Date: 2025-05-02SHANGHAI ICEKREDIT INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510413245.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-05-02
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing bidding data processing methods will lose table and list structure information when converting HTML format data into plain text, and require a large amount of manual annotation data to train the model, resulting in high costs and fluctuations in the annotation quality.

Method used

Generative big model is used to extract bidding data information, and data extraction is achieved by obtaining HTML format data, cleaning data, building structured prompt words, generating question and answer data sets, fine-tuning the generative big model and applying low-rank adaptive technology.

Benefits of technology

It improves the accuracy of information extraction, avoids the loss of data structure information, reduces dependence on manual labeling data, and reduces labor costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119917756A_ABST
    Figure CN119917756A_ABST
Patent Text Reader

Abstract

The invention designs a method for extracting bidding and tendering data information by using a generative large model. The method comprises the following steps: S1, acquiring bidding and tendering data in an html format from the Internet; s2, clearing the bidding and tendering data; s3, acquiring a Few-shot sample based on the information extraction task, and constructing a plurality of structured cue words; s4, combining Few-shot samples and cues, using a generative large model to generate bidding and tendering information extraction question and answer data sets in batches, performing manual random spot check and verifying data validity, and performing manual correction on problematic data; s5, inputting the question and answer data obtained in the step S4 into a large generative model for fine tuning, and then setting a low-rank adaptive technology to generate the large generative model for bidding and tendering information extraction; and S6, embedding to-be-extracted information into cue words, and inputting the cue words into the fine-tuned large production model to obtain an information extraction result. Through the design, the accuracy of information extraction is improved, and the problem that table and list structure information is lost in a traditional method is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of data processing, and in particular relates to a method for extracting bidding data information by using a generative large model. Background Art

[0002] The existing bidding data processing method mainly converts the bidding data in HTML format into plain text, and then uses rule-based or manually annotated data training models such as BERT to extract information.

[0003] When HTML data is converted into plain text, structural information such as tables and lists will be destroyed, resulting in incomplete or inaccurate final extraction results. Existing methods usually require a large amount of manually annotated data to train the model, which is labor-intensive and time-consuming, and fluctuations in annotation quality will also affect the performance of the final model. Summary of the invention

[0004] In order to solve the above problems, this application designs a method for extracting bidding data information using a generative big model, in order to fine-tune the generative big model, improve the accuracy of information extraction, and extract key information without losing the data structure, avoiding the problem of loss of table and list structure information in traditional methods.

[0005] A method for extracting bidding data information using a generative large model comprises the following steps: Step S1, obtaining bidding data in HTML format from the Internet; Step S2, cleaning up the bidding data; Step S3, obtaining a few-shot sample based on the information extraction task and constructing multiple structured prompt words; Step S4: Combine the few-shot samples and prompt words, use the generative large model to batch generate the bidding information extraction question and answer dataset, manually randomly check and verify the data validity, and manually correct the problematic data; Step S5, input the question-answer data obtained in step S4 into the generative large model for fine-tuning, and then set the low-rank adaptive technology to generate a generative large model for bidding information extraction; Step S6: embed the information to be extracted into the prompt word and input it into the fine-tuned production model to obtain the information extraction result.

[0006] In the step S1, the bidding data includes announcement documents related to the bidding.

[0007] In step S2, the method for cleaning up the bidding data includes the following steps: Step S21, removing the cascading style sheet, scripting language and comments of the bidding data in HTML format; Step S22, eliminating redundant html tag attributes; Step S23: Perform lossless structural compression on the cleaned bidding data.

[0008] The lossless structure compression method comprises: Step S231, merging the multi-layer nested tags in the bidding data processed in step S22; Step S232: remove tags with empty content.

[0009] In step S3, the method of obtaining a few-shot sample based on the information extraction task and constructing a plurality of structured prompt words includes: Step S31, constructing a Few-shot project according to a predefined information extraction task; Step S32: design prompt words based on the few-shot examples.

[0010] In step S5, the question-answer data obtained in step S4 is input into the generative large model for fine-tuning, and then a low-rank adaptive technology is set to generate a generative large model for extracting bidding information, including: Step S51, low-rank decomposition of weights: Perform low-rank matrix decomposition on the weights of the generative large model, and decompose the weight matrix W into low-rank matrices A and B, that is, ∆W=AB, where AϵR d×r and BϵR r×d ; Step S52, parameter update: During the parameter update process, LoRA only updates the low-rank matrices A and B, and does not change the original weight W; Step S53, model fine-tuning: Use the question-answering dataset obtained in step S4 and the Lora technology to fine-tune the generative large model.

[0011] The advantages and effects of this application are as follows: The present application designs a method for extracting bidding data information using a generative big model, including: step S1, obtaining bidding data in HTML format from the Internet; step S2, cleaning the bidding data; step S3, obtaining a few-shot sample based on the information extraction task and constructing multiple structured prompt words; step S4, combining the few-shot samples and the prompt words, using the generative big model to batch generate a bidding information extraction question and answer data set, manually randomly sampling and verifying the validity of the data, and manually correcting the problematic data; step S5, inputting the question and answer data obtained in step S4 into the generative big model for fine-tuning, and then setting the low-rank adaptive technology to generate a generative big model for bidding information extraction; step S6, embedding the information to be extracted into the prompt words, and inputting it into the fine-tuned production big model to obtain the information extraction result. Through the above design, the present application can realize fine-tuning training of the generative large model, improve the accuracy of information extraction, and extract key information without losing the data structure, avoiding the problem of loss of table and list structure information in traditional methods. In addition, since the present application only requires a small number of labeled samples, the generative large model can automatically learn and understand the structure and content of the bidding data, reducing the reliance on manually labeled data during the data labeling process and reducing labor costs.

[0012] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application so that it can be implemented in accordance with the contents of the specification, and to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the following is a detailed description of the preferred embodiments of the present application in conjunction with the accompanying drawings as follows.

[0013] Based on the detailed description of the specific embodiments of the present application in combination with the accompanying drawings below, those skilled in the art will become more aware of the above and other objects, advantages and features of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings required for use in the embodiments or the prior art descriptions are briefly introduced below. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can also be obtained based on these drawings without creative work. In all drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, each element or part is not necessarily drawn according to the actual scale.

[0015] Figure 1 A flowchart of a method for extracting bidding data information using a generative large model designed for this application; Figure 2An example diagram of prompt words designed for this application is combined with a Few-shot example. DETAILED DESCRIPTION

[0016] To make the purpose, technical scheme and advantages of the embodiment of the present application clearer, the technical scheme in the embodiment of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiment of the present application. Obviously, the described embodiment is a part of the embodiment of the present application, rather than all of the embodiments. In the following description, specific details such as specific configuration and components are provided only to help fully understand the embodiments of the present application. Therefore, it should be clear to those skilled in the art that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present application. In addition, for clarity and brevity, the description of known functions and structures is omitted in the embodiment.

[0017] It should be understood that the references to "one embodiment" or "this embodiment" throughout the specification mean that the specific features, structures, or characteristics associated with the embodiment are included in at least one embodiment of the present application. Therefore, the references to "one embodiment" or "this embodiment" appearing throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.

[0018] In addition, the present application may repeat reference numerals and / or letters in different examples. This repetition is for the purpose of simplicity and clarity, and does not in itself indicate the relationship between the various embodiments and / or settings discussed.

[0019] The term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, B exists alone, and A and B exist at the same time. The term " / and" in this article describes another type of association object relationship, indicating that there can be two relationships. For example, A / and B can mean: A exists alone, and A and B exist alone. In addition, the character " / " in this article generally indicates that the previous and next associated objects are in an "or" relationship.

[0020] The term "at least one" in this article is merely a description of the association relationship of associated objects, indicating that there may be three relationships. For example, at least one of A and B can mean: A exists alone, A and B exist at the same time, and B exists alone.

[0021] It should also be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusions. Example

[0022] Please refer to Figure 1 This embodiment mainly introduces a specific design of a method for extracting bidding data information using a generative large model, including the following steps: Step S1, obtaining bidding data in HTML format from the Internet; Furthermore, in step S1, the bidding data includes announcement documents related to the bidding.

[0023] Step S2: Since the original HTML document is too long, directly inputting the model will exceed the context length of the large model; therefore, a rule-based HTML cleaning method is designed to clean the bidding data; Furthermore, in step S2, the method for cleaning the bidding data includes the following steps: Step S21, removing the cascading style sheet, scripting language and comments of the bidding data in HTML format; Step S22, eliminating redundant html tag attributes; Step S23: Considering that there may be structural redundancy in the original HTML document, compression operation can be performed without losing semantics, and lossless structural compression can be performed on the cleaned bidding data.

[0024] Furthermore, in step S23, the lossless structure compression method includes: Step S231: merge the multi-layer nested tags in the bidding data processed in step S22, such as " some text " is simplified to " some text ”; Step S232: Remove tags with empty content, such as " ”.

[0025] Step S3, obtaining a few-shot sample based on the information extraction task and constructing multiple structured prompt words; Furthermore, in step S3, the method of obtaining a few-shot sample based on the information extraction task and constructing a plurality of structured prompt words includes: Step S31, constructing a Few-shot project: according to the pre-defined information extraction task, manually annotating n (n<=10) bidding data as examples, the data source is the public bidding information; Step S32: Design prompt words based on the Few-shot example, such as Figure 2 .

[0026] Step S4: Combine the few-shot samples and prompt words, use a generative large model such as the Qwen model to batch generate a bidding information extraction question and answer dataset, manually randomly check and verify the data validity, and manually correct the problematic data; Step S5: Input the question-answer data obtained in step S4 into a generative large model such as the Qwen model for fine-tuning, and then set the low-rank adaptive (LoRA) technology, that is, perform low-rank decomposition through the weight matrix of the pre-trained model to approximate the incremental parameters in the process of full parameter fine-tuning, and generate a generative large model for bidding information extraction.

[0027] Furthermore, in step S5, the question-answer data obtained in step S4 is input into the generative large model for fine-tuning, and then a low-rank adaptive technology is set to generate a generative large model for extracting bidding information, including: Step S51, low-rank decomposition of weights: Perform low-rank matrix decomposition on the weights of the generative large model Qwen, and decompose the weight matrix W into low-rank matrices A and B, that is, ∆W=AB, where AϵR d×r and BϵR r×d ; Step S52, parameter update: During the parameter update process, LoRA only updates the low-rank matrices A and B, and does not change the original weight W; Step S53, model fine-tuning: Use the question-answering dataset obtained in step S4 and the Lora technology to fine-tune the generative large model.

[0028] Step S6: embed the information to be extracted into the prompt word and input it into the fine-tuned production model to obtain the information extraction result.

[0029] The present application designs a method for extracting bidding data information using a generative big model, which can realize fine-tuning training of the generative big model, improve the accuracy of information extraction, and extract key information without losing the data structure, thus avoiding the problem of loss of table and list structure information in traditional methods. Furthermore, since the present application only requires a small number of labeled samples, the generative big model can automatically learn and understand the structure and content of the bidding data, reducing the reliance on manually labeled data during the data labeling process and reducing labor costs.

[0030] The above description is only the preferred embodiment of the present invention, which does not limit the protection scope of the present invention. For those skilled in the art, the present invention can be modified and varied in various ways. Within the spirit and principle of the present invention, any change, modification, replacement, integration and parameter change of these embodiments by conventional substitution or capable of achieving the same function without departing from the principle and spirit of the present invention shall fall within the protection scope of the present invention.

Claims

1. A method for extracting bidding data information using a generative large model, characterized in that: The following steps are involved: Step S1, obtaining bidding data in HTML format from the Internet; Step S2, cleaning up the bidding data; Step S3, obtaining a few-shot sample based on the information extraction task and constructing multiple structured prompt words; Step S4: Combine the few-shot samples and prompt words, use the generative large model to batch generate the bidding information extraction question and answer dataset, manually randomly check and verify the data validity, and manually correct the problematic data; Step S5, input the question-answer data obtained in step S4 into the generative large model for fine-tuning, and then set the low-rank adaptive technology to generate a generative large model for bidding information extraction; Step S6: embed the information to be extracted into the prompt word and input it into the fine-tuned production model to obtain the information extraction result.

2. According to claim 1, a method for extracting bidding data information using a generative large model is characterized in that: In the step S1, the bidding data includes announcement documents related to the bidding.

3. According to claim 1, a method for extracting bidding data information using a generative large model is characterized in that: In step S2, the method for cleaning up the bidding data includes the following steps: Step S21, removing the cascading style sheet, scripting language and comments of the bidding data in HTML format; Step S22, eliminating redundant html tag attributes; Step S23: Perform lossless structural compression on the cleaned bidding data.

4. According to claim 3, a method for extracting bidding data information using a generative large model is characterized in that: The lossless structure compression method comprises: Step S231, merging the multi-layer nested tags in the bidding data processed in step S22; Step S232: remove tags with empty content.

5. According to claim 1, a method for extracting bidding data information using a generative large model is characterized in that: In step S3, the method of obtaining a few-shot sample based on the information extraction task and constructing a plurality of structured prompt words includes: Step S31, constructing a Few-shot project according to a predefined information extraction task; Step S32: design prompt words based on the few-shot examples.

6. The method for extracting bidding data information using a generative large model according to claim 1 is characterized in that: In step S5, the question-answer data obtained in step S4 is input into the generative large model for fine-tuning, and then a low-rank adaptive technology is set to generate a generative large model for extracting bidding information, including: Step S51, low-rank decomposition of weights: Perform low-rank matrix decomposition on the weights of the generative large model, and decompose the weight matrix W into low-rank matrices A and B, that is, ∆W=AB, where AϵR d×r and BϵR r×d ; Step S52, parameter update: During the parameter update process, LoRA only updates the low-rank matrices A and B, and does not change the original weight W; Step S53, model fine-tuning: Use the question-answering dataset obtained in step S4 and the Lora technology to fine-tune the generative large model.

Citation Information

Patent Citations

  • Bid invitation data filling treatment method and device, equipment and medium

    CN118862889A

  • Method for generating and utilizing a training dataset for deep learning based generative ai system using super-large ai

    KR102570178B1

  • Artificial Intelligence Generation of Advertisements

    US20200273062A1

  • Systems, methods, kits, and apparatuses for generative artificial intelligence, graphical neural networks, transformer models, and converging technology stacks in value chain networks

    WO2024226801A2