Bidding document intelligent abstract extraction method and system based on image recognition

Through image recognition and natural language processing technology, an abstract extraction model for bid documents was established, which solved the problem that existing technology was difficult to deal with bid documents in different industries, realized the intelligent and automated abstract extraction of bid documents, and improved the accuracy and adaptability of extraction.

CN119938908AActive Publication Date: 2025-05-06NANTONG INST OF TECH +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510421859.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-05-06
Estimated Expiration
2045-04-07

AI Technical Summary

Technical Problem

The prior art is difficult to cover all scenarios when processing bid documents from different industries, and it is difficult to obtain high-quality training data, resulting in poor performance of models when facing new types or format documents.

Method used

Through image recognition and natural language processing technology, bid file information is entered and abstract extraction model is established, bid file keywords are extracted, keyword sorting and tuning are performed to ensure the smoothness and semantic consistency of the abstract.

Benefits of technology

The intelligent and automated abstract extraction of bid documents is realized, the accuracy and adaptability of extraction are improved, and bid documents in different fields and formats can be better handled.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938908A_ABST
    Figure CN119938908A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent document extraction, and discloses an intelligent bidding document abstract extraction method and system based on image recognition, and the method comprises the steps: selecting an application field to which a bidding document belongs according to the bidding document, inputting bidding document information through image scanning, and building an abstract extraction model to extract a bidding document keyword; the method comprises the following steps: establishing a bidding file abstract through keyword sorting, checking whether the abstract is smooth or not and whether the abstract conforms to the semantics of an initial bidding file or not by a tuning and correcting unit, if the abstract is smooth, continuing checking whether the abstract conforms to the semantics of the initial bidding file or not, and if the abstract is not smooth, performing synonym replacement and root association on the keywords, and resorting the keywords; if the initial bidding file semantics are met, intelligent abstract extraction of the bidding file is completed, if the initial bidding file semantics are not met, whether the keyword is correct or not is checked again, intelligent abstract extraction is achieved, complete algorithm logic is provided, and the accuracy of abstract extraction of the bidding file is guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the technical field of intelligent file extraction, and discloses a method and a system for intelligent abstract extraction of bidding documents based on image recognition. Background Art

[0002] Image recognition technology has developed rapidly in recent years, especially driven by deep learning algorithms, and the accuracy of image recognition has been significantly improved. Intelligent summary extraction of bidding documents is an application field of image recognition technology, which mainly involves automatically extracting key information from documents in the form of scanned copies or pictures to generate document summaries or for subsequent data processing. However, there are still many shortcomings. For example, bidding documents may have multiple formats and layouts, including complex elements such as tables and graphics, which poses a challenge to image recognition. High-quality training data is essential for building accurate models. However, in actual applications, it is not easy to obtain a large amount of well-annotated training data. Bidding documents in different industries differ in terms of terminology, format, etc. A single model is difficult to cover all scenarios and needs to be adjusted or retrained for different fields.

[0003] For example, the existing Chinese patent application with publication number CN118504559A discloses a method and system for intelligent extraction of legal and regulatory annotation files, including collecting text with legal and regulatory annotation basis as input original legal and regulatory annotation basis text, performing data preprocessing on the original legal and regulatory annotation basis text; realizing feature construction based on feature engineering, at least constituting the title feature, the table text feature, the non-table text feature and the symbolic feature; using the constructed features to extract key information of the original legal and regulatory annotation basis text, and automatically identifying key entity information in the legal and regulatory text through text scanning, splitting, feature comparison, regular matching, etc., based on the extracted key information. Compared with the prior art, the invention can improve the accuracy of 2D gaze point estimation.

[0004] The data set of the training model of the above patent may not be diverse enough, resulting in poor performance of the model when faced with new types or formats of documents. For example, different documents may have different terminology usage habits, format layout, etc. Although natural language processing technology has made great progress, it still has certain limitations in understanding and extracting complex semantics. Legal and regulatory texts often contain a large number of professional terms and complex logical relationships, which places high demands on the machine's understanding ability and relies on fixed rules to extract information. Such methods may fail when faced with flexible and changeable text formats. For example, regular expressions are very useful in some scenarios, but they may become inflexible when faced with complex text structures. Summary of the invention

[0005] The purpose of this section is to summarize some aspects of embodiments of the present invention and briefly introduce some preferred embodiments. Some simplifications or omissions may be made in this section and the specification abstract and the invention title of this application to avoid blurring the purpose of this section, the specification abstract and the invention title, and such simplifications or omissions cannot be used to limit the scope of the present invention.

[0006] In order to solve the above technical problems, the main purpose of the present invention is to provide a bidding document intelligent summary extraction method based on image recognition, comprising:

[0007] S1. Select the application field to which the bidding document belongs according to the bidding document (describe the application field in detail and give examples);

[0008] S2. Input the bidding document information through image scanning (specifically, the technical means of image scanning can be realized by conventional technical means), and establish a summary extraction model to extract keywords of the bidding document;

[0009] S3. Establish a summary of the bidding document by sorting keywords, and use the optimization and correction unit to check whether the summary is fluent and conforms to the semantics of the initial bidding document;

[0010] S4. If the abstract is fluent, continue to check whether it complies with the semantics of the initial bidding document. If the abstract is not fluent, perform synonym replacement and root association on the keywords and re-order the keywords;

[0011] S5. If the bidding document semantics are consistent, the intelligent extraction summary of the bidding document is completed. If the bidding document semantics are not consistent, the keywords are rechecked to see if they are correct.

[0012] As a preferred solution of the method for intelligent abstract extraction of bidding documents based on image recognition of the present invention, wherein:

[0013] The image scanning and input bidding document information converts the document into an editable electronic text format through image recognition and natural language processing;

[0014] The image recognition is used to extract text content from an editable electronic text format;

[0015] The natural language processing is used to initially understand the extracted text.

[0016] As a preferred solution of the method for intelligent abstract extraction of bidding documents based on image recognition of the present invention, wherein:

[0017] The abstract extraction model includes a word density analysis unit, a word priority screening unit, a word priority sorting unit and an abstract generation verification unit;

[0018] The word density analysis unit is used to extract and check the density of word occurrences, and extract word features through a feature extraction function;

[0019] The word priority screening unit calculates the priority weights of the words by setting the word priority sorting logic;

[0020] The word priority sorting unit is used to receive word priority weights and sort the word priorities;

[0021] The summary generation and verification unit is used to generate a summary of the bidding document according to the word priority ranking, and perform integrity verification on the summary of the bidding document.

[0022] As a preferred solution of the method for intelligent abstract extraction of bidding documents based on image recognition of the present invention, wherein:

[0023] The word density analysis unit extracts text content from the editable electronic text format and performs regional segmentation, establishes a two-dimensional coordinate system of size a×a, numbers the grids from left to right and from top to bottom, and calculates the density of word occurrences in each part of the grid. The word density analysis unit calculates the expression as follows:

[0024]

[0025] in, is the word density of the grid with horizontal coordinate i and vertical coordinate j in the electronic text format, is the k-th word weight, is the length of the kth word, is the distance factor of the kth word in the grid (i, j), is the total number of words;

[0026] The distance factor calculation expression is:

[0027]

[0028] in, is the distance factor, e is the exponential constant, is the horizontal coordinate position of word k, is the horizontal coordinate of the grid center position, is the ordinate position of word k, is the vertical coordinate position of the center of the grid, α is the distance coefficient, which is related to the word spacing;

[0029] The distance factor is a value between 0 and 1, indicating that the importance of a word changes with its position.

[0030] As a preferred solution of the method for intelligent abstract extraction of bidding documents based on image recognition of the present invention, wherein:

[0031] The word priority screening unit collects the density and position of the grid words in the area segmentation, extracts features from the collected density and position of the grid words, calculates weights based on the features extracted from the features, and finally sorts and outputs the grid words according to the priority weights of all the grid words;

[0032] The feature extraction is used to extract the word frequency and inverse text frequency index of the grid word, the grid word length and the grid word position weight.

[0033] As a preferred solution of the method for intelligent abstract extraction of bidding documents based on image recognition of the present invention, wherein:

[0034] The calculation expression of the weight calculation is as follows:

[0035]

[0036] in, is the grid word priority weight, is the frequency and inverse text frequency index of the grid words, is the length of word k, is the position weight of word k;

[0037] The position weight calculation expression is as follows:

[0038]

[0039] in, is the position weight coefficient of the grid word as the title, is the position weight coefficient of the grid word as the first sentence, is the position weight coefficient of the grid word in the sentence, is the position of the kth word in the bidding document, For other cases;

[0040] Normalizing the grid word priority weight, the grid word frequency and inverse text frequency index, the word length and the word position weight by maximum normalization;

[0041] The maximum normalization calculation expression is as follows:

[0042]

[0043] in, is the data after maximum normalization processing, is the input data that needs to be normalized, min[] is the minimum value function, and max[] is the maximum value function;

[0044] The word priority sorting unit is used to receive the word priority weights after maximum normalization and sort the word priorities.

[0045] As a preferred solution of the method for intelligent abstract extraction of bidding documents based on image recognition of the present invention, wherein:

[0046] The keyword sorting is based on the extracted grid words and their weights, and the grid words are sorted in descending order of weight;

[0047] The step of establishing the bid document summary includes selecting the first N keywords from the sorted keyword list and constructing a logically coherent sentence or paragraph as the summary;

[0048] Setting a screening threshold by ranking the word priorities to screen out high-priority words as the keywords;

[0049] The correction function and the pruning function are constructed by the tuning correction unit to perform grammatical check and semantic consistency check on the summary. The correction function is used to correct the symmetry, compactness, vanishing distance and orthogonality problems when extracting image features. The pruning function compensates the word density analysis unit through the database and text image information.

[0050] As a preferred solution of the method for intelligent abstract extraction of bidding documents based on image recognition of the present invention, wherein:

[0051] If the abstract is fluent, the semantic consistency of the abstract is checked again through the semantic similarity algorithm. If the abstract is not fluent, synonym replacement, root association and re-ordering are performed to make the abstract fluent.

[0052] If it does not conform to the semantics of the initial bidding document, the keywords of the bidding document will be screened again through the word density analysis unit and the word priority sorting unit to recheck whether there are any missing or wrong keywords;

[0053] The semantic similarity algorithm measures the similarity of two vectors by cosine similarity. The calculation expression is as follows:

[0054]

[0055] in, is the similarity between the abstract and the text, is the cosine similarity, is the vector representation of the summary, is the vector representation of the bidding document;

[0056] The synonym replacement finds a set of synonyms for the keyword k, and selects one of them to replace the original keyword;

[0057] The root association searches for a specific word, root(t), and replaces it with synonyms and derivatives of the root.

[0058] The bidding document intelligent summary extraction system based on image recognition includes:

[0059] The document identification module includes a scanning unit for scanning the bidding document information and a classification unit for defining the technical field to which the bidding document belongs;

[0060] An input module, including an input unit for image scanning and input, an image recognition unit for extracting text content from an editable electronic text format, and a natural language processing unit for preliminary understanding of the extracted text;

[0061] A summary extraction model, including a word density analysis unit, a word priority screening unit, a word priority sorting unit and a summary generation verification unit;

[0062] The feedback module includes a tuning correction unit, a semantic similarity algorithm, synonym replacement, and root association.

[0063] As a preferred solution of the bidding document intelligent abstract extraction system based on image recognition of the present invention, wherein:

[0064] The word density analysis unit is used to extract and check the density of word occurrences, and extract word features through a feature extraction function;

[0065] The word priority screening unit calculates the priority weights of the words by setting the word priority sorting logic;

[0066] The word priority sorting unit is used to receive word priority weights and sort the word priorities;

[0067] The summary generation and verification unit is used to generate a summary of the bidding document according to the word priority ranking, and to perform integrity verification on the summary of the bidding document.

[0068] Beneficial effects of the present invention:

[0069] The summary extraction model solves the problem that different documents may have different terminology usage habits. The semantic intelligent segmentation and intelligent feature extraction of image recognition are used to realize the extraction of key information of bidding documents, enhance the ability to understand document images, realize the intelligent and automated extraction, and improve the accuracy of extraction. BRIEF DESCRIPTION OF THE DRAWINGS

[0070] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative labor. Among them:

[0071] Figure 1 It is a flow chart of the method for intelligent abstract extraction of bidding documents based on image recognition of the present invention;

[0072] Figure 2 It is a work flow chart of a word density analysis unit of the bidding document intelligent abstract extraction method based on image recognition of the present invention;

[0073] Figure 3 This is a system composition diagram of the bidding document intelligent summary extraction system based on image recognition of the present invention;

[0074] Figure 4 It is the overall structure diagram of the bidding document intelligent summary extraction system based on image recognition of the present invention;

[0075] Figure 5 It is an internal structure diagram of the bidding document intelligent summary extraction system based on image recognition of the present invention;

[0076] Figure 6 This is a physical diagram of the structure of the bidding document intelligent summary extraction system based on image recognition of the present invention;

[0077] Figure 7 It is a module diagram of the bidding document intelligent summary extraction system based on image recognition of the present invention.

[0078] Figure numerals: 1. Cabinet; 2. Main floor door; 3. Lower floor door; 4. Roller floor; 5. Wireless communication module; 6. Keyboard; 7. Recognition camera; 8. Laser printer; 9. Printer output tray; 10. Host; 11. File recognition module; 12. Input module; 13. Lightning protection and power distribution unit; 14. Data processing unit; 15. Storage drawer. DETAILED DESCRIPTION

[0079] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the accompanying drawings.

[0080] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0081] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The term "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive with other embodiments.

[0082] Embodiment 1

[0083] like Figure 1 As shown, the bidding document intelligent summary extraction method based on image recognition includes:

[0084] S1. Select the application field to which the bidding document belongs according to the bidding document;

[0085] In this embodiment, the bidding documents may belong to multiple fields such as construction engineering, information technology, medical equipment, energy development, environmental protection projects, etc. By analyzing key information such as project descriptions, technical requirements, contract terms, etc. in the bidding documents, it is possible to preliminarily determine the field to which they belong; for example, a bidding document contains a large amount of content about "building structure design", "construction material selection" and "construction schedule", which can be judged to belong to the field of construction engineering; by determining the application field to which the bidding document belongs, it is helpful for subsequent information processing and summary extraction to target specific fields more accurately.

[0086] S2, input bidding document information through image scanning, and establish a summary extraction model to extract bidding document keywords;

[0087] In this embodiment, the bidding document information is input by image scanning. The recognition accuracy of modern OCR (optical character recognition) technology has reached a relatively high level. The image scanning technology of the present invention uses OCR technology to convert paper bidding documents or electronic bidding documents in formats such as PDF into editable text formats. OCR technology converts characters in images into computer-readable text data by recognizing their shapes and arrangements. The bidding document information is input by image scanning, which greatly reduces the time and cost of manual input. A summary extraction model is established to extract bidding document keywords, and keywords that can summarize the main information of the document are extracted by analyzing the text content of the bidding document. By condensing a large amount of text information into a small number of keywords, it is convenient to quickly understand the main content of the bidding document. The extraction of keywords helps non-professionals to quickly understand the key points of the bidding document. For example, keywords such as "project name", "bid amount", "technical requirements" and "construction period" are extracted from a bidding document.

[0088] S3. Establish a summary of the bidding document by sorting keywords, and use the optimization and correction unit to check whether the summary is fluent and conforms to the semantics of the initial bidding document;

[0089] In this embodiment, the keywords are sorted according to their importance and relevance, and then combined into a summary according to a certain logical structure (such as chronological order, logical hierarchy, etc.). The basis for sorting can be the frequency of occurrence, position, correlation with other keywords in the document, etc. Through reasonable sorting and combination, the summary can be easier to understand; for example, the extracted keywords are combined into a summary in the order of project name, bid amount, technical requirements, and construction period.

[0090] Specifically, the tuning and correction unit is used to check the syntax, semantics and logic of the generated summary. By comparing the summary with the text content of the initial bid document, it is determined whether the summary is fluent and whether it accurately conveys the semantics of the initial bid document. By correcting the summary, it is ensured that the summary accurately conveys the semantics of the initial bid document to avoid misunderstandings or ambiguities caused by inaccurate summary. For example, if the "technical requirements" part in the summary deviates from the description in the initial bid document, the tuning and correction unit will make corrections to ensure the accuracy of the summary.

[0091] S4. If the abstract is fluent, continue to check whether it conforms to the semantics of the initial bidding document. If the abstract is not fluent, perform synonym replacement and root association on the keywords and re-order the keywords;

[0092] In this embodiment, when the abstract is not fluent or does not conform to the semantics of the initial bidding document, synonym replacement and root association techniques can be used to optimize keywords. Synonymous replacement is to replace inappropriate keywords with their synonyms or near synonyms; root association is to associate other related words based on the root or affix of the keyword to enrich the content of the abstract; through synonym replacement and root association, the content of the abstract can be flexibly adjusted to make the abstract more fluent and easy to understand; for example, if the word "construction period" in the abstract makes the abstract incoherent, it can be replaced with synonyms such as "construction period" or "construction period".

[0093] S5. If the bidding document semantics are consistent, the intelligent extraction summary of the bidding document is completed. If the bidding document semantics are not consistent, the keywords are rechecked to see if they are correct.

[0094] After synonym replacement and root association, it is necessary to recheck whether the keywords are correct and whether they still conform to the semantics of the initial bidding document; by rechecking, ensure that the keywords in the abstract are correct, ensure that the quality of the abstract meets the requirements, and can accurately convey the semantics of the initial bidding document, making the abstract more credible and reliable, which helps to improve the overall quality of the bidding document.

[0095] Furthermore, the bidding document information input by image scanning is converted into an editable electronic text format through image recognition and natural language processing; image recognition is used to extract text content from the editable electronic text format; and natural language processing is used to initially understand the extracted text.

[0096] In this embodiment, paper bidding documents are converted into digital images by using a scanner or camera. These images are usually high-resolution to ensure that the text is clearly legible. The images are pre-processed, including denoising, binarization (converting the image to black and white so that the text can be more clearly identified) and correction (such as rotation correction to ensure that the text is neatly arranged). High-quality scanning can accurately capture every detail in the document and convert the document into an editable electronic text format for easy storage and processing; the OCR technology in the image recognition technology is used to extract text content from the pre-processed image, and the text in the image is converted into editable text for subsequent processing and analysis; the extracted text content is preliminarily understood through natural language processing. The natural language processing technology can identify keywords, phrases and sentence structures in the document, so as to preliminarily understand the theme and main points of the document.

[0097] Furthermore, the abstract extraction model includes a word density analysis unit, a word priority screening unit, a word priority sorting unit and an abstract generation verification unit;

[0098] The word density analysis unit is used to extract and check the density of word occurrences, and extract word features through a feature extraction function;

[0099] The word priority screening unit calculates the priority weights of the words by setting the word priority sorting logic;

[0100] The word priority sorting unit is used to receive word priority weights and sort the word priorities;

[0101] The summary generation and verification unit is used to generate a summary of the bidding document according to the word priority ranking, and to perform integrity verification on the summary of the bidding document.

[0102] Further, such as Figure 2 The word density analysis unit extracts text content from the editable electronic text format to perform regional segmentation, establishes a two-dimensional coordinate system of size a×a, numbers the grids from left to right and from top to bottom, and calculates the density of word occurrences in each part of the grid. The word density analysis unit calculates the expression as follows:

[0103]

[0104] in, is the word density of the grid with horizontal coordinate i and vertical coordinate j in the electronic text format, is the kth word weight, is the length of the kth word, is the distance factor of the kth word in the grid (i, j), is the total number of words;

[0105] Furthermore, Used to indicate that all words in the grid are included;

[0106] The distance factor calculation expression is:

[0107]

[0108] in, is the distance factor, e is the exponential constant, is the horizontal coordinate position of word k, is the horizontal coordinate of the grid center position, is the ordinate position of word k, is the vertical coordinate position of the center of the grid, α is the distance coefficient, which is related to the word spacing;

[0109] The distance factor is a value between 0 and 1, indicating that the importance of a word changes with its position.

[0110] Furthermore, the word priority screening unit collects the density and position of the grid words in the area segmentation, extracts features from the collected density and position of the grid words, calculates weights based on the features extracted from the features, and finally sorts and outputs the grid words according to the priority weights of all the grid words;

[0111] The feature extraction is used to extract the word frequency and inverse text frequency index of the grid word, the grid word length and the grid word position weight.

[0112] In this embodiment, the word priority screening unit first performs regional segmentation on the text or image content, divides the overall content into multiple grids, identifies and collects all grid words in each grid, and through regional segmentation, the text or image content can be analyzed more carefully and more details can be captured; multiple features are extracted from the collected grid words, including word frequency, inverse text frequency index, grid word length and grid word position weight. By extracting multiple features, the importance and priority of the grid words can be more comprehensively evaluated, and screening combined with multiple features can more accurately identify key information; a priority weight is calculated for each grid word based on the extracted features, and the importance of the grid word can be quantified into a specific numerical value through weight calculation, which is convenient for comparison and sorting; all grid words are sorted according to the calculated priority weights, and output according to the sorting results. The sorting is usually in descending order, that is, the grid words with the highest weight are ranked first. The sorted results can intuitively show which grid words are more important, so that users can quickly obtain key information.

[0113] Furthermore, the calculation expression of the weight calculation is as follows:

[0114]

[0115] in, is the grid word priority weight, is the frequency and inverse text frequency index of the grid words, is the length of word k, is the position weight of word k;

[0116] The position weight calculation expression is as follows:

[0117]

[0118] in, is the position weight coefficient of the grid word as the title, is the position weight coefficient of the grid word as the first sentence, is the position weight coefficient of the grid word in the sentence, is the position of the kth word in the bidding document, For other cases;

[0119] Normalizing the grid word priority weight, the grid word frequency and inverse text frequency index, the word length and the word position weight by maximum normalization;

[0120] The maximum normalization calculation expression is as follows:

[0121]

[0122] in, is the data after maximum normalization processing, is the input data that needs to be normalized, min[] is the minimum value function, and max[] is the maximum value function;

[0123] The word priority sorting unit is used to receive the word priority weights after maximum normalization and sort the word priorities.

[0124] Further, keyword sorting sorts the grid words in descending order of weight based on the extracted grid words and their weights;

[0125] The step of establishing the bid document summary includes selecting the first N keywords from the sorted keyword list and constructing a logically coherent sentence or paragraph as the summary;

[0126] Setting a screening threshold by ranking the word priorities to screen out high-priority words as the keywords;

[0127] The correction function and the pruning function are constructed by the tuning correction unit to perform grammatical check and semantic consistency check on the summary. The correction function is used to correct the symmetry, compactness, vanishing distance and orthogonality problems when extracting image features. The pruning function compensates the word density analysis unit through the database and text image information.

[0128] In this embodiment, the keywords are sorted from high to low according to the weight value. Through the sorting, the keywords with higher weight can be processed preferentially, thereby improving the efficiency of information processing and text analysis; the first N (N is a preset value) keywords are selected from the sorted keyword list as the core content of the summary, and a logically coherent and information-complete sentence or paragraph is constructed as the summary based on these keywords. When constructing the summary, the relevance and contextual logic between the keywords need to be considered to ensure the accuracy and readability of the summary; by setting a screening threshold in the word priority sorting, high-priority keywords are screened out. Through the screening threshold, keywords with lower weight can be removed, thereby optimizing the quality of the keyword list; the tuning and correction unit performs grammatical check and semantic consistency check on the summary by constructing a correction function and a pruning function. The correction function is used to solve the symmetry, compactness, vanishing distance and orthogonality problems during image feature extraction to ensure that the image information in the summary is accurate. The pruning function compensates the word density analysis unit through the database and text image information to improve the accuracy and completeness of the summary.

[0129] Furthermore, if the abstract is fluent, the semantic consistency of the abstract is checked again through the semantic similarity algorithm. If the abstract is not fluent, synonym replacement, root association and re-ordering are performed to make the abstract fluent.

[0130] If it does not conform to the semantics of the initial bidding document, the keywords of the bidding document will be screened again through the word density analysis unit and the word priority sorting unit to recheck whether there are any missing or wrong keywords;

[0131] The semantic similarity algorithm measures the similarity of two vectors by cosine similarity. The calculation expression is as follows:

[0132]

[0133] in, is the similarity between the abstract and the text, is the cosine similarity, is the vector representation of the summary, is the vector representation of the bidding document;

[0134] The synonym replacement finds a set of synonyms for the keyword k, and selects one of them to replace the original keyword;

[0135] The root association searches for a specific word, root(t), and replaces it with synonyms and derivatives of the root.

[0136] Embodiment 2

[0137] like Figure 3 As shown, the bidding document intelligent summary extraction system based on image recognition includes:

[0138] The document identification module includes a scanning unit for scanning the bidding document information and a classification unit for defining the technical field to which the bidding document belongs;

[0139] An input module, including an input unit for image scanning and input, an image recognition unit for extracting text content from an editable electronic text format, and a natural language processing unit for preliminary understanding of the extracted text;

[0140] A summary extraction model, including a word density analysis unit, a word priority screening unit, a word priority sorting unit and a summary generation verification unit;

[0141] The feedback module includes a tuning correction unit, a semantic similarity algorithm, synonym replacement, and root association;

[0142] The correction function and the pruning function are constructed by the tuning correction unit to perform grammatical check and semantic consistency check on the summary. The correction function is used to correct the symmetry, compactness, vanishing distance and orthogonality problems when extracting image features. The pruning function compensates the word density analysis unit through the database and text image information.

[0143] The word density analysis unit is used to extract and check the density of word occurrences, and extract word features through a feature extraction function;

[0144] The word priority screening unit calculates the priority weights of the words by setting the word priority sorting logic;

[0145] The word priority sorting unit is used to receive word priority weights and sort the word priorities;

[0146] The summary generation and verification unit is used to generate a summary of the bidding document according to the word priority ranking, and perform integrity verification on the summary of the bidding document.

[0147] Embodiment 3

[0148] like Figure 4 As shown, the structure diagram of the bidding document intelligent summary extraction system based on image recognition includes: a cabinet 1, a main layer hatch 2, a lower layer hatch 3 and a roller floor 4. The cabinet 1 is used to support the display, and the lower layer hatch 3 is used to store Figure 5 The lightning protection and power distribution unit 13, the data processing unit 14 and the storage drawer 15 shown;

[0149] Furthermore, the main floor door 2 is used to store the host 10, the file identification module 11 and the input module 12;

[0150] like Figure 5 As shown, it also includes a wireless communication module 5, a keyboard 6, an identification camera 7, a laser printer 8, and a printer paper output tray 9.

[0151] Among them, the keyboard 6 is used to control the input of the bidding documents, the recognition camera 7 is used to recognize the content of the paper bidding documents, and the laser printer 8 is used to print the summary generation result.

[0152] like Figure 6 Physical diagram of the structure of the bidding document intelligent summary extraction system based on image recognition.

[0153] like Figure 7 As shown, the bidding document intelligent summary extraction system inputs the office supplies bidding document, enters it into the system, parses the file size and records the entry time, and previews the original text. The office supplies bidding document is initially identified and processed by marking the file scanning and keyword extraction. The bidding document summary generation process is monitored through the processing progress. Finally, the final summary generation result of the office supplies bidding document is generated from the summary result.

[0154] Importantly, it should be noted that the construction and arrangement of the present application shown in a plurality of different exemplary embodiments are only exemplary. Although only two embodiments are described in detail in this disclosure, it should be readily understood by those who refer to this disclosure that many modifications are possible, for example, the size, scale, structure, shape and ratio of various elements, and parameter values ​​(e.g., temperature, pressure, etc.), installation arrangement, use of materials, color, directional changes, etc., without substantially departing from the novel teachings and advantages of the subject matter described in the application. For example, the element shown as integrally formed can be composed of multiple parts or elements, the position of the element can be inverted or otherwise changed, and the nature or number or position of the discrete element can be changed or changed. Therefore, all such modifications are intended to be included in the scope of the present invention. The order or sequence of any process or method steps can be changed or reordered according to alternative embodiments. Any "device plus function" clause is intended to cover the structure of the execution function described in this article, and is not only structurally equivalent but also equivalent structure. Without departing from the scope of the present invention, other replacements, modifications, changes and omissions can be made in the design, operating conditions and arrangement of the exemplary embodiment. Therefore, the invention is not limited to a specific embodiment, but extends to numerous modifications still falling within the scope of the appended claims.

[0155] Additionally, in order to provide a concise description of exemplary embodiments, all features of an actual embodiment may not be described (ie, those features that are not relevant to the best mode presently contemplated for carrying out the invention or those features that are not relevant to implementing the invention).

[0156] It should be understood that in the development of any actual implementation, as in any engineering or design project, numerous implementation-specific decisions may be made. Such a development effort may be complex and time-consuming, but for those of ordinary skill having the benefit of this disclosure, the development effort will be a routine task of design, fabrication, and production without undue experimentation.

[0157] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. The intelligent summary extraction method of bidding documents based on image recognition is characterized by: include: S1. Select the application field to which the bidding document belongs according to the bidding document; S2, input bidding document information through image scanning, and establish a summary extraction model to extract bidding document keywords; S3. Establish a summary of the bidding document by sorting keywords, and use the optimization and correction unit to check whether the summary is fluent and conforms to the semantics of the initial bidding document; S4. If the abstract is fluent, continue to check whether it conforms to the semantics of the initial bidding document. If the abstract is not fluent, perform synonym replacement and root association on the keywords and re-order the keywords; S5. If the bidding document semantics are consistent, the intelligent extraction summary of the bidding document is completed. If the bidding document semantics are not consistent, the keywords are rechecked to see if they are correct.

2. The method for extracting intelligent abstracts from bidding documents based on image recognition according to claim 1 is characterized in that: The image scanning and input bidding document information converts the document into an editable electronic text format through image recognition and natural language processing; The image recognition is used to extract text content from an editable electronic text format; The natural language processing is used to initially understand the extracted text.

3. The method for extracting intelligent abstracts from bidding documents based on image recognition according to claim 2 is characterized in that: The abstract extraction model includes a word density analysis unit, a word priority screening unit, a word priority sorting unit and an abstract generation verification unit; The word density analysis unit is used to extract and check the density of word occurrences, and extract word features through a feature extraction function; The word priority screening unit calculates the priority weights of the words by setting the word priority sorting logic; The word priority sorting unit is used to receive word priority weights and sort the word priorities; The summary generation and verification unit is used to generate a summary of the bidding document according to the word priority ranking, and to perform integrity verification on the summary of the bidding document.

4. The method for extracting intelligent abstracts from bidding documents based on image recognition according to claim 3 is characterized by: The word density analysis unit extracts text content from the editable electronic text format and performs regional segmentation, establishes a two-dimensional coordinate system of size a×a, numbers the grids from left to right and from top to bottom, and calculates the density of word occurrences in each part of the grid. The word density analysis unit calculates the expression as follows: ; in, is the word density of the grid with horizontal coordinate i and vertical coordinate j in the electronic text format, is the k-th word weight, is the length of the kth word, is the distance factor of the kth word in the grid (i, j), is the total number of words; The distance factor calculation expression is: ; in, is the distance factor, e is the exponential constant, is the horizontal coordinate position of word k, is the horizontal coordinate of the grid center position, is the ordinate position of word k, is the vertical coordinate position of the grid center, α is the distance coefficient; The distance factor is a value between 0 and 1, indicating that the importance of a word changes with its position.

5. The method for extracting intelligent abstracts from bidding documents based on image recognition according to claim 4 is characterized in that: The word priority screening unit collects the density and position of the grid words in the area segmentation, extracts features from the collected density and position of the grid words, calculates weights based on the features extracted from the features, and finally sorts and outputs the grid words according to the priority weights of all the grid words; The feature extraction is used to extract the word frequency and inverse text frequency index of the grid word, the grid word length and the grid word position weight.

6. The method for extracting intelligent abstracts from bidding documents based on image recognition according to claim 5 is characterized in that: The calculation expression of the weight calculation is as follows: ; in, is the grid word priority weight, is the frequency and inverse text frequency index of the grid words, is the length of word k, is the position weight of word k; The position weight calculation expression is as follows: ; in, is the position weight coefficient of the grid word as the title, is the position weight coefficient of the grid word as the first sentence, is the position weight coefficient of the grid word in the sentence, is the position of the kth word in the bidding document, For other cases; Normalizing the grid word priority weight, the grid word frequency and inverse text frequency index, the word length and the word position weight by maximum normalization; The maximum normalization calculation expression is as follows: ; in, is the data after maximum normalization processing, is the input data that needs to be normalized, min[] is the minimum value function, and max[] is the maximum value function; The word priority sorting unit is used to receive the word priority weights after maximum normalization and sort the word priorities.

7. The method for extracting intelligent abstracts from bidding documents based on image recognition according to claim 6 is characterized in that: The keyword sorting is based on the extracted grid words and their weights, and the grid words are sorted in descending order of weight; The step of establishing the bid document summary includes selecting the first N keywords from the sorted keyword list and constructing a logically coherent sentence or paragraph as the summary; Setting a screening threshold by ranking the word priorities to screen out high-priority words as the keywords; The correction function and the pruning function are constructed by the tuning correction unit to perform grammatical check and semantic consistency check on the summary. The correction function is used to correct the symmetry, compactness, vanishing distance and orthogonality problems when extracting image features. The pruning function compensates the word density analysis unit through the database and text image information.

8. The method for extracting intelligent abstracts from bidding documents based on image recognition according to claim 7 is characterized in that: If the abstract is fluent, the semantic consistency of the abstract is checked again through the semantic similarity algorithm. If the abstract is not fluent, synonym replacement, root association and re-ordering are performed to make the abstract fluent. If it does not conform to the semantics of the initial bidding document, the keywords of the bidding document will be screened again through the word density analysis unit and the word priority sorting unit to recheck whether there are any missing or wrong keywords; The semantic similarity algorithm measures the similarity of two vectors by cosine similarity. The calculation expression is as follows: ; in, is the similarity between the abstract and the text, is the cosine similarity, is the vector representation of the summary, is the vector representation of the bidding document; The synonym replacement finds a set of synonyms for the keyword k, and selects one of them to replace the original keyword; The root association searches for a specific word, root(t), and replaces it with synonyms and derivatives of the root.

9. A system for intelligent abstract extraction of bidding documents based on image recognition, used to implement the method for intelligent abstract extraction of bidding documents based on image recognition as claimed in any one of claims 1 to 8, characterized in that: include: The document identification module includes a scanning unit for scanning the bidding document information and a classification unit for defining the technical field to which the bidding document belongs; An input module, including an input unit for image scanning and input, an image recognition unit for extracting text content from an editable electronic text format, and a natural language processing unit for preliminary understanding of the extracted text; A summary extraction model, including a word density analysis unit, a word priority screening unit, a word priority sorting unit and a summary generation verification unit; The feedback module includes a tuning correction unit, a semantic similarity algorithm, synonym replacement, and root association.

10. The bidding document intelligent summary extraction system based on image recognition according to claim 9 is characterized by: The word density analysis unit is used to extract and check the density of word occurrences, and extract word features through a feature extraction function; The word priority screening unit calculates the priority weights of the words by setting the word priority sorting logic; The word priority sorting unit is used to receive word priority weights and sort the word priorities; The summary generation and verification unit is used to generate a summary of the bidding document according to the word priority ranking, and to perform integrity verification on the summary of the bidding document.

Citation Information

Patent Citations

  • Intelligent extraction method and system for law and regulation annotation files

    CN118504559A

  • Course target achievement condition evaluation rationality evaluation method based on semantic analysis

    CN116362927A

  • File uploading method with function of abstracting index-information in real-time and web-storage system using the same

    KR100905434B1